Home LiteratureArticle Details
PMID: 11997345 Published · ppublish English Letter

Large-scale protein annotation through gene ontology.

Genome research ·Vol. 12 ·No. 5 ·2002-05-00 ·Pages 785-94

Xie H, Wasserman A, Levine Z, Novik A, Grebinskiy V, Shoshan A, Mintz L

Abstract

Recent progress in genomic sequencing, computational biology, and ontology development has presented an opportunity to investigate biological systems from a unique perspective, that is, examining genomes and transcriptomes through the multiple and hierarchical structure of Gene Ontology (GO). We report here our development of GO Engine, a computational platform for GO annotation, and analysis of the resultant GO annotations of human proteins. Protein annotation was centered on sequence homology with GO-annotated proteins and protein domain analysis. Text information analysis and a multiparameter cellular localization predictive tool were also used to increase the annotation accuracy, and to predict novel annotations. The majority of proteins corresponding to full-length mRNA in GenBank, and the majority of proteins in the NR database (nonredundant database of proteins) were annotated with one or more GO nodes in each of the three GO categories. The annotations of GenBank and SWISS-PROT proteins are available to the public at the GO Consortium web site.

MeSH Terms
Animals Computational Biology/methods Databases, Genetic Databases, Protein Genome, Human Humans Multigene Family Proteins/classification,genetics,physiology Sequence Analysis, Protein Sequence Homology, Amino Acid
Chemicals
Proteins
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Xie Hanqing
Compugen Inc., Jamesburg, New Jersey 08831, USA. han@cgen.com
Wasserman Alon
Levine Zurit
Novik Amit
Grebinskiy Vladimir
Shoshan Avi
Mintz Liat
References (19)
19 references, click to expand
  1. ProtoMap: automatic classification of protein sequences, a hierarchy of protein families, and local maps of the protein space.
    Proteins. 1999 Nov 15;37(3):360-78 PMID: 10591097
  2. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  3. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  4. The InterPro database, an integrated documentation resource for protein families, domains and functional sites.
    Nucleic Acids Res. 2001 Jan 1;29(1):37-40 PMID: 11125043
  5. Ontology acquisition from on-line knowledge sources.
    Proc AMIA Symp. 2000;:497-501 PMID: 11079933
  6. All human genes of the uteroglobin family are localized on chromosome 11q12.2 and form a dense cluster.
    Ann N Y Acad Sci. 2000;923:25-42 PMID: 11193762
  7. The sequence of the human genome.
    Science. 2001 Feb 16;291(5507):1304-51 PMID: 11181995
  8. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  9. Ancient origin of the Hox gene cluster.
    Nat Rev Genet. 2001 Jan;2(1):33-8 PMID: 11253066
  10. Comparative DNA sequence analysis of mouse and human protocadherin gene clusters.
    Genome Res. 2001 Mar;11(3):389-404 PMID: 11230163
  11. A literature network of human genes for high-throughput analysis of gene expression.
    Nat Genet. 2001 May;28(1):21-8 PMID: 11326270
  12. Creating the gene ontology resource: design and implementation.
    Genome Res. 2001 Aug;11(8):1425-33 PMID: 11483584
  13. A draft annotation and overview of the human genome.
    Genome Biol. 2001;2(7):RESEARCH0025 PMID: 11516338
  14. Clustering protein sequences--structure prediction by transitive homology.
    Bioinformatics. 2001 Oct;17(10):935-41 PMID: 11673238
  15. The Ensembl genome database project.
    Nucleic Acids Res. 2002 Jan 1;30(1):38-41 PMID: 11752248
  16. Saccharomyces Genome Database (SGD) provides secondary gene annotation using the Gene Ontology (GO).
    Nucleic Acids Res. 2002 Jan 1;30(1):69-72 PMID: 11752257
  17. The FlyBase database of the Drosophila genome projects and community literature.
    Nucleic Acids Res. 2002 Jan 1;30(1):106-8 PMID: 11752267
  18. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  19. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
2002-05-00
Pages
785-94
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC186564
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com