Home LiteratureArticle Details
PMID: 9521927 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S. Research Support, U.S. Gov't, P.H.S. Review

GAIA: framework annotation of genomic sequence.

Genome research ·Vol. 8 ·No. 3 ·1998-03-00 ·Pages 234-50

Bailey LC, Fischer S, Schug J, Crabtree J, Gibson M, Overton GC

Abstract

As increasing amounts of genomic sequence from many organisms become available, and as DNA sequences become a primary reagent in biologic investigations, the role of annotation as a prospective guide for laboratory experiments will expand rapidly. Here we describe a process of high-throughput, reliable annotation, called framework annotation, which is designed to provide a foundation for initial biologic characterization of previously unexamined sequence. To examine this concept in practice, we have constructed Genome Annotation and Information Analysis (GAIA), a prototype software architecture that implements several elements important for framework annotation. The center of GAIA consists of an annotation database and the associated data management subsystem that forms the software bus along which other components communicate. The schema for this database defines three principal concepts: (1) Entries, consisting of sequence and associated historical data; (2) Features, comprising information of biologic interest; and (3) Experiments, describing the evidence that supports Features. The database permits tracking of annotation results over time, as well as assessment of the reliability of particular results. New framework annotation is produced by CARTA, a set of autonomous sensors that perform automatic analyses and assert results into the annotation database. These results are available via a Web-based query interface that uses graphical Java applets as well as text-based HTML pages to display data at different levels of resolution and permit interactive exploration of annotation. We present results for initial application of framework annotation to a set of test sequences, demonstrating its effectiveness in providing a starting point for biologic investigation, and discuss ways in which the current prototype can be improved. The prototype is available for public use and comment at http://www.cbil.upenn.edu/gaia.

MeSH Terms
Amino Acid Sequence Animals Base Sequence Computational Biology/methods Databases, Factual Human Genome Project Humans Molecular Sequence Data Online Systems Sequence Analysis, DNA/methods Software
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Bailey L C
Computational Biology and Informatics Laboratory, Department of Genetics, University of Pennsylvania School of Medicine, Philadelphia, Pennsylvania 19104-6021, USA. bailey@www.cbil.upenn.edu
Fischer S
Schug J
Crabtree J
Gibson M
Overton G C
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
1998-03-00
Pages
234-50
Language
English
Region
United States
NLM ID
9518021
Subset
IM
Grants
NHGRI NIH HHS · 1R011HG0153901 · United States
NHGRI NIH HHS · R01HG0145001 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com