Home LiteratureArticle Details
PMID: 9611239 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Combining diverse evidence for gene recognition in completely sequenced bacterial genomes.

Nucleic acids research ·Vol. 26 ·No. 12 ·1998-06-15 ·Pages 2941-7

Frishman D, Mironov A, Mewes HW, Gelfand M

Abstract

Analysis of a newly sequenced bacterial genome starts with identification of protein-coding genes. Functional assignment of proteins requires the exact knowledge of protein N-termini. We present a new program ORPHEUS that identifies candidate genes and accurately predicts gene starts. The analysis starts with a database similarity search and identification of reliable gene fragments. The latter are used to derive statistical characteristics of protein-coding regions and ribosome-binding sites and to predict the complete set of genes in the analyzed genome. In a test on Bacillus subtilis and Escherichia coli genomes, the program correctly identified 93.3% (resp. 96.3%) of experimentally annotated genes longer than 100 codons described in the PIR-International database, and for these genes 96.3% (83.9%) of starts were predicted exactly. Furthermore, 98.9% (99.1%) of genes longer than 100 codons annotated in GenBank were found, and 92.9% (75.7%) of predicted starts coincided with the feature table description. Finally, for the complete gene complements of B.subtilis and E.coli , including genes shorter than 100 codons, gene prediction accuracy was 88.9 and 87.1%, respectively, with 94.2 and 76.7% starts coinciding with the existing annotation.

MeSH Terms
Algorithms Bacillus subtilis/genetics Bacterial Proteins/genetics Codon, Initiator Databases, Factual Escherichia coli/genetics Genome, Bacterial Open Reading Frames Sequence Alignment/methods Sequence Analysis, DNA Software
Chemicals
Bacterial Proteins Codon, Initiator
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Frishman D
Munich Information Center for Protein Sequences (MIPS) of the German National Center for Health and Environment (GSF), Am Klopferspitz 18a, 82152 Martinsried, Germany. frishman@mips.biochem.mpg.de
Mironov A
Mewes H W
Gelfand M
References (33)
33 references, click to expand
  1. Comparison of DNA sequences with protein sequences.
    Genomics. 1997 Nov 15;46(1):24-36 PMID: 9403055
  2. The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1998.
    Nucleic Acids Res. 1998 Jan 1;26(1):38-42 PMID: 9399796
  3. GeneMark.hmm: new solutions for gene finding.
    Nucleic Acids Res. 1998 Feb 15;26(4):1107-15 PMID: 9461475
  4. Fast comparison of a DNA sequence with a protein sequence database.
    Microb Comp Genomics. 1996;1(4):281-91 PMID: 9689213
  5. Deriving ribosomal binding site (RBS) statistical models from unannotated DNA sequences and the use of the RBS model for N-terminal prediction.
    Pac Symp Biocomput. 1998;:279-90 PMID: 9697189
  6. Microbial gene identification using interpolated Markov models.
    Nucleic Acids Res. 1998 Jan 15;26(2):544-8 PMID: 9421513
  7. The 3'-terminal sequence of Escherichia coli 16S ribosomal RNA: complementarity to nonsense triplets and ribosome binding sites.
    Proc Natl Acad Sci U S A. 1974 Apr;71(4):1342-6 PMID: 4598299
  8. Markedly unbiased codon usage in Bacillus subtilis.
    Gene. 1985;40(1):145-50 PMID: 3937765
  9. Information content of binding sites on nucleotide sequences.
    J Mol Biol. 1986 Apr 5;188(3):415-31 PMID: 3525846
  10. Synonymous codon usage in Bacillus subtilis reflects both translational selection and mutational biases.
    Nucleic Acids Res. 1987 Oct 12;15(19):8023-40 PMID: 3118331
  11. What constitutes the signal for the initiation of protein synthesis on Escherichia coli mRNAs?
    J Mol Biol. 1988 Nov 5;204(1):79-94 PMID: 2464068
  12. Sequence and transcription mapping of Bacillus subtilis competence genes comB and comA, one of which is related to a family of bacterial regulatory determinants.
    J Bacteriol. 1989 Oct;171(10):5362-75 PMID: 2507523
  13. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  14. What's in a genome?
    Nature. 1992 Jul 23;358(6384):287 PMID: 1641000
  15. Assessment of protein coding measures.
    Nucleic Acids Res. 1992 Dec 25;20(24):6441-50 PMID: 1480466
  16. Identification of protein coding regions by database similarity search.
    Nat Genet. 1993 Mar;3(3):266-72 PMID: 8485583
  17. Quantitative analysis of ribosome binding sites in E.coli.
    Nucleic Acids Res. 1994 Apr 11;22(7):1287-95 PMID: 8165145
  18. RNA sequence analysis using covariance models.
    Nucleic Acids Res. 1994 Jun 11;22(11):2079-88 PMID: 8029015
  19. Large scale bacterial gene discovery by similarity search.
    Nat Genet. 1994 Jun;7(2):205-14 PMID: 7920643
  20. Intrinsic and extrinsic approaches for detecting genes in a bacterial genome.
    Nucleic Acids Res. 1994 Nov 11;22(22):4756-67 PMID: 7984428
  21. A hidden Markov model that finds genes in E. coli DNA.
    Nucleic Acids Res. 1994 Nov 11;22(22):4768-78 PMID: 7984429
  22. Prediction of function in DNA sequence analysis.
    J Comput Biol. 1995 Spring;2(1):87-115 PMID: 7497122
  23. Sequencing and analysis of bacterial genomes.
    Curr Biol. 1996 Apr 1;6(4):404-16 PMID: 8723345
  24. Complete genome sequence of the methanogenic archaeon, Methanococcus jannaschii.
    Science. 1996 Aug 23;273(5278):1058-73 PMID: 8688087
  25. SRS: information retrieval system for molecular biology data banks.
    Methods Enzymol. 1996;266:114-28 PMID: 8743681
  26. Finding genes by computer: the state of the art.
    Trends Genet. 1996 Aug;12(8):316-20 PMID: 8783942
  27. Gene recognition via spliced sequence alignment.
    Proc Natl Acad Sci U S A. 1996 Aug 20;93(17):9061-6 PMID: 8799154
  28. The N-end rule: functions, mysteries, uses.
    Proc Natl Acad Sci U S A. 1996 Oct 29;93(22):12142-9 PMID: 8901547
  29. Identification of prokaryotic and eukaryotic signal peptides and prediction of their cleavage sites.
    Protein Eng. 1997 Jan;10(1):1-6 PMID: 9051728
  30. Prediction of complete gene structures in human genomic DNA.
    J Mol Biol. 1997 Apr 25;268(1):78-94 PMID: 9149143
  31. The complete genome sequence of Escherichia coli K-12.
    Science. 1997 Sep 5;277(5331):1453-62 PMID: 9278503
  32. The complete genome sequence of the gram-positive bacterium Bacillus subtilis.
    Nature. 1997 Nov 20;390(6657):249-56 PMID: 9384377
  33. The PIR-International Protein Sequence Database.
    Nucleic Acids Res. 1998 Jan 1;26(1):27-32 PMID: 9399794
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
0305-1048
Published
1998-06-15
Pages
2941-7
Language
English
Region
England
NLM ID
0411011
PMCID
PMC147632
Subset
IM
Corrections
ErratumIn
-
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com