Abstract
Analysis of a newly sequenced bacterial genome starts with identification of protein-coding genes. Functional assignment of proteins requires the exact knowledge of protein N-termini. We present a new program ORPHEUS that identifies candidate genes and accurately predicts gene starts. The analysis starts with a database similarity search and identification of reliable gene fragments. The latter are used to derive statistical characteristics of protein-coding regions and ribosome-binding sites and to predict the complete set of genes in the analyzed genome. In a test on Bacillus subtilis and Escherichia coli genomes, the program correctly identified 93.3% (resp. 96.3%) of experimentally annotated genes longer than 100 codons described in the PIR-International database, and for these genes 96.3% (83.9%) of starts were predicted exactly. Furthermore, 98.9% (99.1%) of genes longer than 100 codons annotated in GenBank were found, and 92.9% (75.7%) of predicted starts coincided with the feature table description. Finally, for the complete gene complements of B.subtilis and E.coli , including genes shorter than 100 codons, gene prediction accuracy was 88.9 and 87.1%, respectively, with 94.2 and 76.7% starts coinciding with the existing annotation.
MeSH Terms
Algorithms
Bacillus subtilis/genetics
Bacterial Proteins/genetics
Codon, Initiator
Databases, Factual
Escherichia coli/genetics
Genome, Bacterial
Open Reading Frames
Sequence Alignment/methods
Sequence Analysis, DNA
Software
Chemicals
Bacterial Proteins
Codon, Initiator
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Frishman D
Munich Information Center for Protein Sequences (MIPS) of the German National Center for Health and Environment (GSF), Am Klopferspitz 18a, 82152 Martinsried, Germany. frishman@mips.biochem.mpg.de
Mironov A
Mewes H W
Gelfand M
References (33)
33 references, click to expand
-
Comparison of DNA sequences with protein sequences.
Genomics. 1997 Nov 15;46(1):24-36
PMID: 9403055
-
The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1998.
Nucleic Acids Res. 1998 Jan 1;26(1):38-42
PMID: 9399796
-
GeneMark.hmm: new solutions for gene finding.
Nucleic Acids Res. 1998 Feb 15;26(4):1107-15
PMID: 9461475
-
Fast comparison of a DNA sequence with a protein sequence database.
Microb Comp Genomics. 1996;1(4):281-91
PMID: 9689213
-
Deriving ribosomal binding site (RBS) statistical models from unannotated DNA sequences and the use of the RBS model for N-terminal prediction.
Pac Symp Biocomput. 1998;:279-90
PMID: 9697189
-
Microbial gene identification using interpolated Markov models.
Nucleic Acids Res. 1998 Jan 15;26(2):544-8
PMID: 9421513
-
The 3'-terminal sequence of Escherichia coli 16S ribosomal RNA: complementarity to nonsense triplets and ribosome binding sites.
Proc Natl Acad Sci U S A. 1974 Apr;71(4):1342-6
PMID: 4598299
-
Markedly unbiased codon usage in Bacillus subtilis.
Gene. 1985;40(1):145-50
PMID: 3937765
-
Information content of binding sites on nucleotide sequences.
J Mol Biol. 1986 Apr 5;188(3):415-31
PMID: 3525846
-
Synonymous codon usage in Bacillus subtilis reflects both translational selection and mutational biases.
Nucleic Acids Res. 1987 Oct 12;15(19):8023-40
PMID: 3118331
-
What constitutes the signal for the initiation of protein synthesis on Escherichia coli mRNAs?
J Mol Biol. 1988 Nov 5;204(1):79-94
PMID: 2464068
-
Sequence and transcription mapping of Bacillus subtilis competence genes comB and comA, one of which is related to a family of bacterial regulatory determinants.
J Bacteriol. 1989 Oct;171(10):5362-75
PMID: 2507523
-
Basic local alignment search tool.
J Mol Biol. 1990 Oct 5;215(3):403-10
PMID: 2231712
-
What's in a genome?
Nature. 1992 Jul 23;358(6384):287
PMID: 1641000
-
Assessment of protein coding measures.
Nucleic Acids Res. 1992 Dec 25;20(24):6441-50
PMID: 1480466
-
Identification of protein coding regions by database similarity search.
Nat Genet. 1993 Mar;3(3):266-72
PMID: 8485583
-
Quantitative analysis of ribosome binding sites in E.coli.
Nucleic Acids Res. 1994 Apr 11;22(7):1287-95
PMID: 8165145
-
RNA sequence analysis using covariance models.
Nucleic Acids Res. 1994 Jun 11;22(11):2079-88
PMID: 8029015
-
Large scale bacterial gene discovery by similarity search.
Nat Genet. 1994 Jun;7(2):205-14
PMID: 7920643
-
Intrinsic and extrinsic approaches for detecting genes in a bacterial genome.
Nucleic Acids Res. 1994 Nov 11;22(22):4756-67
PMID: 7984428
-
A hidden Markov model that finds genes in E. coli DNA.
Nucleic Acids Res. 1994 Nov 11;22(22):4768-78
PMID: 7984429
-
Prediction of function in DNA sequence analysis.
J Comput Biol. 1995 Spring;2(1):87-115
PMID: 7497122
-
Sequencing and analysis of bacterial genomes.
Curr Biol. 1996 Apr 1;6(4):404-16
PMID: 8723345
-
Complete genome sequence of the methanogenic archaeon, Methanococcus jannaschii.
Science. 1996 Aug 23;273(5278):1058-73
PMID: 8688087
-
SRS: information retrieval system for molecular biology data banks.
Methods Enzymol. 1996;266:114-28
PMID: 8743681
-
Finding genes by computer: the state of the art.
Trends Genet. 1996 Aug;12(8):316-20
PMID: 8783942
-
Gene recognition via spliced sequence alignment.
Proc Natl Acad Sci U S A. 1996 Aug 20;93(17):9061-6
PMID: 8799154
-
The N-end rule: functions, mysteries, uses.
Proc Natl Acad Sci U S A. 1996 Oct 29;93(22):12142-9
PMID: 8901547
-
Identification of prokaryotic and eukaryotic signal peptides and prediction of their cleavage sites.
Protein Eng. 1997 Jan;10(1):1-6
PMID: 9051728
-
Prediction of complete gene structures in human genomic DNA.
J Mol Biol. 1997 Apr 25;268(1):78-94
PMID: 9149143
-
The complete genome sequence of Escherichia coli K-12.
Science. 1997 Sep 5;277(5331):1453-62
PMID: 9278503
-
The complete genome sequence of the gram-positive bacterium Bacillus subtilis.
Nature. 1997 Nov 20;390(6657):249-56
PMID: 9384377
-
The PIR-International Protein Sequence Database.
Nucleic Acids Res. 1998 Jan 1;26(1):27-32
PMID: 9399794