Home LiteratureArticle Details
PMID: 11879526 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Re-annotation of genome microbial coding-sequences: finding new genes and inaccurately annotated genes.

BMC bioinformatics ·Vol. 3 ·2002-00-00 ·Pages 5

Bocs S, Danchin A, Médigue C

Abstract

Analysis of any newly sequenced bacterial genome starts with the identification of protein-coding genes. Despite the accumulation of multiple complete genome sequences, which provide useful comparisons with close relatives among other organisms during the annotation process, accurate gene prediction remains quite difficult. A major reason for this situation is that genes are tightly packed in prokaryotes, resulting in frequent overlap. Thus, detection of translation initiation sites and/or selection of the correct coding regions remain difficult unless appropriate biological knowledge (about the structure of a gene) is imbedded in the approach. We have developed a new program that automatically identifies biologically significant candidate genes in a bacterial genome. Twenty-six complete prokaryotic genomes were analyzed using this tool, and the accuracy of gene finding was assessed by comparison with existing annotations. This analysis revealed that, despite the enormous effort of genome program annotators, a small but not negligible number of genes annotated within the framework of sequencing projects are likely to be partially inaccurate or plainly wrong. Moreover, the analysis of several putative new genes shows that, as expected, many short genes have escaped annotation. In most cases, these new genes revealed frameshifts that could be either artifacts or genuine frameshifts. Some entirely unexpected new genes have also been identified. This allowed us to get a more complete picture of prokaryotic genomes. The results of this procedure are progressively integrated into the SWISS-PROT reference databank. The results described in the present study show that our procedure is very satisfactory in terms of gene finding accuracy. Except in few cases, discrepancies between our results and annotations provided by individual authors can be accounted for by the nature of each annotation process or by specific characteristics of some genomes. This stresses that close cooperation between scientists, regular update and curation of the findings in databases are clearly required to reduce the level of errors in genome annotation (and hence in reducing the unfortunate spreading of errors through centralized data libraries).

MeSH Terms
Computational Biology/methods Databases, Protein Gene Transfer, Horizontal/genetics Genes, Archaeal/physiology Genes, Bacterial/physiology Genetic Variation Genome, Archaeal Genome, Bacterial Gram-Negative Bacteria/genetics Gram-Positive Bacteria/genetics Open Reading Frames/genetics Predictive Value of Tests Reading Frames/genetics Software Spirochaetaceae/genetics
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Bocs Stéphanie
Laboratoire Génome et Informatique, Université de Versailles, 91034 Evry Cedex, France. scbocs@infobiogen.fr
Danchin Antoine
Médigue Claudine
References (42)
42 references, click to expand
  1. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  2. Molecular archaeology of the Escherichia coli genome.
    Proc Natl Acad Sci U S A. 1998 Aug 4;95(16):9413-7 PMID: 9689094
  3. Lateral gene transfer and the nature of bacterial innovation.
    Nature. 2000 May 18;405(6784):299-304 PMID: 10830951
  4. Repeat-associated phase variable genes in the complete genome sequence of Neisseria meningitidis strain MC58.
    Mol Microbiol. 2000 Jul;37(1):207-15 PMID: 10931317
  5. Re-annotating the Mycoplasma pneumoniae genome sequence: adding value, function and reading frames.
    Nucleic Acids Res. 2000 Sep 1;28(17):3278-88 PMID: 10954595
  6. Artemis: sequence visualization and annotation.
    Bioinformatics. 2000 Oct;16(10):944-5 PMID: 11120685
  7. Database resources of the National Center for Biotechnology Information.
    Nucleic Acids Res. 2001 Jan 1;29(1):11-6 PMID: 11125038
  8. The COG database: new developments in phylogenetic classification of proteins from complete genomes.
    Nucleic Acids Res. 2001 Jan 1;29(1):22-8 PMID: 11125040
  9. Using the COG database to improve gene recognition in complete genomes.
    Genetica. 2000;108(1):9-17 PMID: 11145426
  10. Towards understanding the first genome sequence of a crenarchaeon by genome annotation using clusters of orthologous groups of proteins (COGs).
    Genome Biol. 2000;1(5):RESEARCH0009 PMID: 11178258
  11. On the total number of genes and their length distribution in complete microbial genomes.
    Trends Genet. 2001 Aug;17(8):425-8 PMID: 11485798
  12. Intrinsic errors in genome annotation.
    Trends Genet. 2001 Aug;17(8):429-31 PMID: 11485799
  13. The codon preference plot: graphic analysis of protein coding sequences and prediction of gene expression.
    Nucleic Acids Res. 1984 Jan 11;12(1 Pt 2):539-49 PMID: 6694906
  14. Evidence for horizontal gene transfer in Escherichia coli speciation.
    J Mol Biol. 1991 Dec 20;222(4):851-6 PMID: 1762151
  15. Linkage map of Escherichia coli K-12, edition 10: the physical map.
    Microbiol Mol Biol Rev. 1998 Sep;62(3):985-1019 PMID: 9729612
  16. Codon usages in different gene classes of the Escherichia coli genome.
    Mol Microbiol. 1998 Sep;29(6):1341-55 PMID: 9781873
  17. Pfam 3.1: 1313 multiple alignments and profile HMMs match the majority of proteins.
    Nucleic Acids Res. 1999 Jan 1;27(1):260-2 PMID: 9847196
  18. Imagene: an integrated computer environment for sequence annotation and analysis.
    Bioinformatics. 1999 Jan;15(1):2-15 PMID: 10068688
  19. Automated genome sequence analysis and annotation.
    Bioinformatics. 1999 May;15(5):391-412 PMID: 10366660
  20. Complete genome sequence of an aerobic hyper-thermophilic crenarchaeon, Aeropyrum pernix K1.
    DNA Res. 1999 Apr 30;6(2):83-101, 145-52 PMID: 10382966
  21. A novel tRNA-associated locus (trl) from Helicobacter pylori is co-transcribed with tRNA(Gly) and reveals genetic diversity.
    Microbiology. 1999 Jun;145 ( Pt 6):1289-98 PMID: 10411255
  22. The minimal gene complement of Mycoplasma genitalium.
    Science. 1995 Oct 20;270(5235):397-403 PMID: 7569993
  23. Detection of new genes in a bacterial genome using Markov models for three gene classes.
    Nucleic Acids Res. 1995 Sep 11;23(17):3554-62 PMID: 7567469
  24. Finding genes by computer: the state of the art.
    Trends Genet. 1996 Aug;12(8):316-20 PMID: 8783942
  25. Selfish operons: horizontal transfer may drive the evolution of gene clusters.
    Genetics. 1996 Aug;143(4):1843-60 PMID: 8844169
  26. Fully automated genome analysis that reflects user needs and preferences. A detailed introduction to the MAGPIE system architecture.
    Biochimie. 1996;78(5):302-10 PMID: 8905148
  27. Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. II. Sequence determination of the entire genome and assignment of potential protein-coding regions.
    DNA Res. 1996 Jun 30;3(3):109-36 PMID: 8905231
  28. Complete sequence analysis of the genome of the bacterium Mycoplasma pneumoniae.
    Nucleic Acids Res. 1996 Nov 15;24(22):4420-49 PMID: 8948633
  29. Protein evolution viewed through Escherichia coli protein sequences: introducing the notion of a structural segment of homology, the module.
    J Mol Biol. 1997 May 23;268(5):857-68 PMID: 9180377
  30. Genotator: a workbench for sequence annotation.
    Genome Res. 1997 Jul;7(7):754-62 PMID: 9253604
  31. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  32. The complete genome sequence of Escherichia coli K-12.
    Science. 1997 Sep 5;277(5331):1453-62 PMID: 9278503
  33. Microbial gene identification using interpolated Markov models.
    Nucleic Acids Res. 1998 Jan 15;26(2):544-8 PMID: 9421513
  34. GAIA: framework annotation of genomic sequence.
    Genome Res. 1998 Mar;8(3):234-50 PMID: 9521927
  35. Detecting and analyzing DNA sequencing errors: toward a higher quality of the Bacillus subtilis genome sequence.
    Genome Res. 1999 Nov;9(11):1116-27 PMID: 10568751
  36. The ancient regulatory-protein family of WD-repeat proteins.
    Nature. 1994 Sep 22;371(6495):297-300 PMID: 8090199
  37. Combining diverse evidence for gene recognition in completely sequenced bacterial genomes.
    Nucleic Acids Res. 1998 Jun 15;26(12):2941-7 PMID: 9611239
  38. The complete genome of the hyperthermophilic bacterium Aquifex aeolicus.
    Nature. 1998 Mar 26;392(6674):353-8 PMID: 9537320
  39. Large scale bacterial gene discovery by similarity search.
    Nat Genet. 1994 Jun;7(2):205-14 PMID: 7920643
  40. Deciphering the biology of Mycobacterium tuberculosis from the complete genome sequence.
    Nature. 1998 Jun 11;393(6685):537-44 PMID: 9634230
  41. Complete sequence and gene organization of the genome of a hyper-thermophilic archaebacterium, Pyrococcus horikoshii OT3.
    DNA Res. 1998 Apr 30;5(2):55-76 PMID: 9679194
  42. Complete DNA sequence of a serogroup A strain of Neisseria meningitidis Z2491.
    Nature. 2000 Mar 30;404(6777):502-6 PMID: 10761919
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2002-00-00
Epub
2002-00-05
Pages
5
Language
English
Region
England
NLM ID
100965194
PMCID
PMC77393
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com