Home LiteratureArticle Details
PMID: 20011109 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

Annotation error in public databases: misannotation of molecular function in enzyme superfamilies.

PLoS computational biology ·Vol. 5 ·No. 12 ·2009-12-00 ·Pages e1000605

Schnoes AM, Brown SD, Dodevski I, Babbitt PC

Abstract

Due to the rapid release of new data from genome sequencing projects, the majority of protein sequences in public databases have not been experimentally characterized; rather, sequences are annotated using computational analysis. The level of misannotation and the types of misannotation in large public databases are currently unknown and have not been analyzed in depth. We have investigated the misannotation levels for molecular function in four public protein sequence databases (UniProtKB/Swiss-Prot, GenBank NR, UniProtKB/TrEMBL, and KEGG) for a model set of 37 enzyme families for which extensive experimental information is available. The manually curated database Swiss-Prot shows the lowest annotation error levels (close to 0% for most families); the two other protein sequence databases (GenBank NR and TrEMBL) and the protein sequences in the KEGG pathways database exhibit similar and surprisingly high levels of misannotation that average 5%-63% across the six superfamilies studied. For 10 of the 37 families examined, the level of misannotation in one or more of these databases is >80%. Examination of the NR database over time shows that misannotation has increased from 1993 to 2005. The types of misannotation that were found fall into several categories, most associated with "overprediction" of molecular function. These results suggest that misannotation in enzyme superfamilies containing multiple families that catalyze different reactions is a larger problem than has been recognized. Strategies are suggested for addressing some of the systematic problems contributing to these high levels of misannotation.

MeSH Terms
Biocatalysis Database Management Systems Databases, Protein
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Schnoes Alexandra M
Graduate Group in Biophysics, University of California San Francisco, San Francisco, California, United States of America.
Brown Shoshana D
Dodevski Igor
Babbitt Patricia C
References (68)
68 references, click to expand
  1. The Gene Ontology project in 2008.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D440-4 PMID: 17984083
  2. The Mouse Genome Database (MGD): mouse biology and model systems.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D724-8 PMID: 18158299
  3. Errors in genome annotation.
    Trends Genet. 1999 Apr;15(4):132-3 PMID: 10203816
  4. Sources of systematic error in functional annotation of genomes: domain rearrangement, non-orthologous gene displacement and operon disruption.
    In Silico Biol. 1998;1(1):55-67 PMID: 11471243
  5. Gene Ontology annotations at SGD: new data sources and annotation methods.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D577-81 PMID: 17982175
  6. Evolution of structure and function in the o-succinylbenzoate synthase/N-acylamino acid racemase family of the enolase superfamily.
    J Mol Biol. 2006 Jun 30;360(1):228-50 PMID: 16740275
  7. The Pfam protein families database.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D281-8 PMID: 18039703
  8. Percolation of annotation errors through hierarchically structured protein sequence databases.
    Math Biosci. 2005 Feb;193(2):223-34 PMID: 15748731
  9. Errors in genome reviews.
    Science. 1998 Sep 4;281(5382):1457 PMID: 9750114
  10. Annotating proteins with generalized functional linkages.
    Proc Natl Acad Sci U S A. 2008 Nov 18;105(46):17700-5 PMID: 19004787
  11. Righting the wrongs.
    EMBO Rep. 2003 Sep;4(9):829-31 PMID: 12949580
  12. Estimating the annotation error rate of curated GO database sequence annotations.
    BMC Bioinformatics. 2007 May 22;8:170 PMID: 17519041
  13. Toward an online repository of Standard Operating Procedures (SOPs) for (meta)genomic annotation.
    OMICS. 2008 Jun;12(2):137-41 PMID: 18416670
  14. Multidimensional annotation of the Escherichia coli K-12 genome.
    Nucleic Acids Res. 2007;35(22):7577-90 PMID: 17940092
  15. Genome re-annotation: a wiki solution?
    Genome Biol. 2007;8(1):102 PMID: 17274839
  16. Novel hopanoid cyclases from the environment.
    Environ Microbiol. 2007 Sep;9(9):2175-88 PMID: 17686016
  17. Enzyme function less conserved than anticipated.
    J Mol Biol. 2002 Apr 26;318(2):595-608 PMID: 12051862
  18. The COG database: an updated version includes eukaryotes.
    BMC Bioinformatics. 2003 Sep 11;4:41 PMID: 12969510
  19. Go hunting in sequence databases but watch out for the traps.
    Trends Genet. 1996 Oct;12(10):425-7 PMID: 8909140
  20. Exploring inconsistencies in genome-wide protein function annotations: a machine learning approach.
    BMC Bioinformatics. 2007 Aug 03;8:284 PMID: 17683567
  21. Retrieving sequences of enzymes experimentally characterized but erroneously annotated : the case of the putrescine carbamoyltransferase.
    BMC Genomics. 2004 Aug 02;5(1):52 PMID: 15287962
  22. Intrinsic errors in genome annotation.
    Trends Genet. 2001 Aug;17(8):429-31 PMID: 11485799
  23. 'Going wrong with confidence': misleading sequence analyses of CiaB and clpX.
    Mol Microbiol. 1999 Oct;34(1):195 PMID: 10540297
  24. InterPro: the integrative protein signature database.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D211-5 PMID: 18940856
  25. Protein function prediction--the power of multiplicity.
    Trends Biotechnol. 2009 Apr;27(4):210-9 PMID: 19251332
  26. PRINTS and its automatic supplement, prePRINTS.
    Nucleic Acids Res. 2003 Jan 1;31(1):400-2 PMID: 12520033
  27. The Catalytic Site Atlas: a resource of catalytic sites and residues identified in enzymes using structural data.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D129-33 PMID: 14681376
  28. How well is enzyme function conserved as a function of pairwise sequence identity?
    J Mol Biol. 2003 Oct 31;333(4):863-82 PMID: 14568541
  29. Leveraging enzyme structure-function relationships for functional inference and experimental design: the structure-function linkage database.
    Biochemistry. 2006 Feb 28;45(8):2545-55 PMID: 16489747
  30. Evolution of function in protein superfamilies, from a structural perspective.
    J Mol Biol. 2001 Apr 6;307(4):1113-43 PMID: 11286560
  31. Ig-like domains on bacteriophages: a tale of promiscuity and deceit.
    J Mol Biol. 2006 Jun 2;359(2):496-507 PMID: 16631788
  32. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  33. Preserving accuracy in GenBank.
    Science. 2008 Mar 21;319(5870):1616 PMID: 18356505
  34. Structure-based functional motif identifies a potential disulfide oxidoreductase active site in the serine/threonine protein phosphatase-1 subfamily.
    FASEB J. 1999 Oct;13(13):1866-74 PMID: 10506591
  35. Cloning and characterization of glyoxalase I from soybean.
    Arch Biochem Biophys. 2000 Feb 15;374(2):261-8 PMID: 10666306
  36. Detecting protein function and protein-protein interactions from genome sequences.
    Science. 1999 Jul 30;285(5428):751-3 PMID: 10427000
  37. What we do not know about sequence analysis and sequence databases.
    Bioinformatics. 1998;14(9):753-4 PMID: 10366280
  38. Cytoscape: a software environment for integrated models of biomolecular interaction networks.
    Genome Res. 2003 Nov;13(11):2498-504 PMID: 14597658
  39. The past, present and future of genome-wide re-annotation.
    Genome Biol. 2002;3(2):COMMENT2001 PMID: 11864365
  40. GenBank.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D26-31 PMID: 18940867
  41. Divergent evolution of enzymatic function: mechanistically diverse superfamilies and functionally distinct suprafamilies.
    Annu Rev Biochem. 2001;70:209-46 PMID: 11395407
  42. Functional classification using phylogenomic inference.
    PLoS Comput Biol. 2006 Jun 30;2(6):e77 PMID: 16846248
  43. Protein function space: viewing the limits or limited by our view?
    Curr Opin Struct Biol. 2007 Jun;17(3):362-9 PMID: 17574832
  44. Comprehensive site-directed mutagenesis of L-2-halo acid dehalogenase to probe catalytic amino acid residues.
    J Biochem. 1995 Jun;117(6):1317-22 PMID: 7490277
  45. DNA data. Proposal to 'Wikify' GenBank meets stiff resistance.
    Science. 2008 Mar 21;319(5870):1598-9 PMID: 18356493
  46. The Universal Protein Resource (UniProt) 2009.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D169-74 PMID: 18836194
  47. Using sequence similarity networks for visualization of relationships across diverse protein superfamilies.
    PLoS One. 2009;4(2):e4345 PMID: 19190775
  48. SCOPEC: a database of protein catalytic domains.
    Bioinformatics. 2004 Aug 4;20 Suppl 1:i130-6 PMID: 15262791
  49. Whole-genome sequence annotation: 'Going wrong with confidence'.
    Mol Microbiol. 1999 May;32(4):886-7 PMID: 10361291
  50. Quantitative assessment of protein function prediction from metagenomics shotgun sequences.
    Proc Natl Acad Sci U S A. 2007 Aug 28;104(35):13913-8 PMID: 17717083
  51. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  52. A gold standard set of mechanistically diverse enzyme superfamilies.
    Genome Biol. 2006;7(1):R8 PMID: 16507141
  53. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  54. The use of gene clusters to infer functional coupling.
    Proc Natl Acad Sci U S A. 1999 Mar 16;96(6):2896-901 PMID: 10077608
  55. Whole proteome analysis of post-translational modifications: applications of mass-spectrometry for proteogenomic annotation.
    Genome Res. 2007 Sep;17(9):1362-77 PMID: 17690205
  56. Phylogenomics: improving functional predictions for uncharacterized genes by evolutionary analysis.
    Genome Res. 1998 Mar;8(3):163-7 PMID: 9521918
  57. KEGG for linking genomes to life and the environment.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D480-4 PMID: 18077471
  58. Modeling the percolation of annotation errors in a database of protein sequences.
    Bioinformatics. 2002 Dec;18(12):1641-9 PMID: 12490449
  59. Manual curation is not sufficient for annotation of genomic databases.
    Bioinformatics. 2007 Jul 1;23(13):i41-8 PMID: 17646325
  60. OrthoMCL-DB: querying a comprehensive multi-species collection of ortholog groups.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D363-8 PMID: 16381887
  61. Predicting protein function from sequence and structure.
    Nat Rev Mol Cell Biol. 2007 Dec;8(12):995-1005 PMID: 18037900
  62. Assigning protein functions by comparative genome analysis: protein phylogenetic profiles.
    Proc Natl Acad Sci U S A. 1999 Apr 13;96(8):4285-8 PMID: 10200254
  63. Conservation of gene order: a fingerprint of proteins that physically interact.
    Trends Biochem Sci. 1998 Sep;23(9):324-8 PMID: 9787636
  64. The PROSITE database.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D227-30 PMID: 16381852
  65. Protein annotation at genomic scale: the current status.
    Chem Rev. 2007 Aug;107(8):3448-66 PMID: 17658902
  66. Representing structure-function relationships in mechanistically diverse enzyme superfamilies.
    Pac Symp Biocomput. 2005;:358-69 PMID: 15759641
  67. MUSCLE: multiple sequence alignment with high accuracy and high throughput.
    Nucleic Acids Res. 2004 Mar 19;32(5):1792-7 PMID: 15034147
  68. Analysis of genomic context: prediction of functional associations from conserved bidirectionally transcribed gene pairs.
    Nat Biotechnol. 2004 Jul;22(7):911-7 PMID: 15229555
Article Info
Journal
PLoS computational biology
Abbr.
PLoS Comput Biol
ISSN
1553-7358
Published
2009-12-00
Epub
2009-00-11
Pages
e1000605
Language
English
Region
United States
NLM ID
101238922
PMCID
PMC2781113
Subset
IM
Grants
NIGMS NIH HHS · GM071790 · United States
Howard Hughes Medical Institute · United States
NIGMS NIH HHS · P01 GM071790 · United States
NIGMS NIH HHS · R01 GM060595 · United States
NIGMS NIH HHS · GM60595 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com