Abstract
We further investigated the statistical features of the three classes of Escherichia coli genes that have been previously delineated by factorial correspondence analysis and dynamic clustering methods. A phased Markov model for a nucleotide sequence of each gene class was developed and employed for gene prediction using the GeneMark program. The protein-coding region prediction accuracy was determined for class-specific Markov models of different orders when the programs implementing these models were applied to gene sequences from the same or other classes. It is shown that at least two training sets and two program versions derived for different classes of E. coli genes are necessary in order to achieve a high accuracy of coding region prediction for uncharacterized sequences. Some annotated E. coli genes from Class I and Class III are shown to be spurious, whereas many open reading frames (ORFs) that have not been annotated in GenBank as genes are predicted to encode proteins. The amino acid sequences of the putative products of these ORFs initially did not show similarity to already known proteins. However, conserved regions have been identified in several of them by screening the latest entries in protein sequence databases and applying methods for motif search, while some other of these new genes have been identified in independent experiments.
MeSH Terms
Algorithms
Amino Acid Sequence
Bacterial Proteins/genetics
Base Composition
Base Sequence
DNA, Bacterial/chemistry
Escherichia coli/genetics
Genes, Bacterial
Molecular Sequence Data
Open Reading Frames
Sequence Alignment
Sequence Homology, Amino Acid
Statistics as Topic
Chemicals
Bacterial Proteins
DNA, Bacterial
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Borodovsky M
School of Biology, Georgia Institute of Technology, Atlanta 30332, USA.
McIninch J D
Koonin E V
Rudd K E
Médigue C
Danchin A
References (22)
22 references, click to expand
-
Codon catalog usage is a genome strategy modulated for gene expressivity.
Nucleic Acids Res. 1981 Jan 10;9(1):r43-74
PMID: 7208352
-
A simple tool to search for sequence motifs that are conserved in BLAST outputs.
Comput Appl Biosci. 1994 Jul;10(4):457-9
PMID: 7804881
-
Measurements of the effects that coding for a protein has on a DNA sequence and their use for finding genes.
Nucleic Acids Res. 1984 Jan 11;12(1 Pt 2):551-67
PMID: 6364041
-
The codon Adaptation Index--a measure of directional synonymous codon usage bias, and its potential applications.
Nucleic Acids Res. 1987 Feb 11;15(3):1281-95
PMID: 3547335
-
Codon usage and tRNA content in unicellular and multicellular organisms.
Mol Biol Evol. 1985 Jan;2(1):13-34
PMID: 3916708
-
Merging of distance matrices and classification by dynamic clustering.
Comput Appl Biosci. 1988 Nov;4(4):453-8
PMID: 3208179
-
Codon preference and primary sequence structure in protein-coding regions.
Bull Math Biol. 1989;51(1):95-115
PMID: 2706404
-
Evidence for horizontal gene transfer in Escherichia coli speciation.
J Mol Biol. 1991 Dec 20;222(4):851-6
PMID: 1762151
-
Analysis of the Escherichia coli genome: DNA sequence of the region from 84.5 to 86.5 minutes.
Science. 1992 Aug 7;257(5071):771-8
PMID: 1379743
-
Mutational analysis of the Escherichia coli serB promoter region reveals transcriptional linkage to a downstream gene.
Gene. 1992 Oct 12;120(1):1-9
PMID: 1327967
-
First and second moment of counts of words in random texts generated by Markov chains.
Comput Appl Biosci. 1992 Oct;8(5):433-41
PMID: 1422876
-
Assessment of protein coding measures.
Nucleic Acids Res. 1992 Dec 25;20(24):6441-50
PMID: 1480466
-
Contamination of cDNA sequences in databases.
Science. 1993 Mar 19;259(5102):1677-8
PMID: 8456288
-
DNA sequence and analysis of 136 kilobases of the Escherichia coli genome: organizational symmetry around the origin of replication.
Genomics. 1993 Jun;16(3):551-61
PMID: 7686882
-
Analysis of the Escherichia coli genome. III. DNA sequence of the region from 87.2 to 89.2 minutes.
Nucleic Acids Res. 1993 Jul 25;21(15):3391-8
PMID: 8346018
-
Analysis of the Escherichia coli genome. IV. DNA sequence of the region from 89.2 to 92.8 minutes.
Nucleic Acids Res. 1993 Nov 25;21(23):5408-17
PMID: 8265357
-
Issues in searching molecular sequence databases.
Nat Genet. 1994 Feb;6(2):119-29
PMID: 8162065
-
Large scale bacterial gene discovery by similarity search.
Nat Genet. 1994 Jun;7(2):205-14
PMID: 7920643
-
New genes in old sequence: a strategy for finding genes in the bacterial genome.
Trends Biochem Sci. 1994 Aug;19(8):309-13
PMID: 7940673
-
Intrinsic and extrinsic approaches for detecting genes in a bacterial genome.
Nucleic Acids Res. 1994 Nov 11;22(22):4756-67
PMID: 7984428
-
A hidden Markov model that finds genes in E. coli DNA.
Nucleic Acids Res. 1994 Nov 11;22(22):4768-78
PMID: 7984429
-
Codon usage in bacteria: correlation with gene expressivity.
Nucleic Acids Res. 1982 Nov 25;10(22):7055-74
PMID: 6760125