Home LiteratureArticle Details
PMID: 7567469 Published · ppublish English Comparative Study Journal Article Research Support, U.S. Gov't, P.H.S.

Detection of new genes in a bacterial genome using Markov models for three gene classes.

Nucleic acids research ·Vol. 23 ·No. 17 ·1995-09-11 ·Pages 3554-62

Borodovsky M, McIninch JD, Koonin EV, Rudd KE, Médigue C, Danchin A

Abstract

We further investigated the statistical features of the three classes of Escherichia coli genes that have been previously delineated by factorial correspondence analysis and dynamic clustering methods. A phased Markov model for a nucleotide sequence of each gene class was developed and employed for gene prediction using the GeneMark program. The protein-coding region prediction accuracy was determined for class-specific Markov models of different orders when the programs implementing these models were applied to gene sequences from the same or other classes. It is shown that at least two training sets and two program versions derived for different classes of E. coli genes are necessary in order to achieve a high accuracy of coding region prediction for uncharacterized sequences. Some annotated E. coli genes from Class I and Class III are shown to be spurious, whereas many open reading frames (ORFs) that have not been annotated in GenBank as genes are predicted to encode proteins. The amino acid sequences of the putative products of these ORFs initially did not show similarity to already known proteins. However, conserved regions have been identified in several of them by screening the latest entries in protein sequence databases and applying methods for motif search, while some other of these new genes have been identified in independent experiments.

MeSH Terms
Algorithms Amino Acid Sequence Bacterial Proteins/genetics Base Composition Base Sequence DNA, Bacterial/chemistry Escherichia coli/genetics Genes, Bacterial Molecular Sequence Data Open Reading Frames Sequence Alignment Sequence Homology, Amino Acid Statistics as Topic
Chemicals
Bacterial Proteins DNA, Bacterial
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Borodovsky M
School of Biology, Georgia Institute of Technology, Atlanta 30332, USA.
McIninch J D
Koonin E V
Rudd K E
Médigue C
Danchin A
References (22)
22 references, click to expand
  1. Codon catalog usage is a genome strategy modulated for gene expressivity.
    Nucleic Acids Res. 1981 Jan 10;9(1):r43-74 PMID: 7208352
  2. A simple tool to search for sequence motifs that are conserved in BLAST outputs.
    Comput Appl Biosci. 1994 Jul;10(4):457-9 PMID: 7804881
  3. Measurements of the effects that coding for a protein has on a DNA sequence and their use for finding genes.
    Nucleic Acids Res. 1984 Jan 11;12(1 Pt 2):551-67 PMID: 6364041
  4. The codon Adaptation Index--a measure of directional synonymous codon usage bias, and its potential applications.
    Nucleic Acids Res. 1987 Feb 11;15(3):1281-95 PMID: 3547335
  5. Codon usage and tRNA content in unicellular and multicellular organisms.
    Mol Biol Evol. 1985 Jan;2(1):13-34 PMID: 3916708
  6. Merging of distance matrices and classification by dynamic clustering.
    Comput Appl Biosci. 1988 Nov;4(4):453-8 PMID: 3208179
  7. Codon preference and primary sequence structure in protein-coding regions.
    Bull Math Biol. 1989;51(1):95-115 PMID: 2706404
  8. Evidence for horizontal gene transfer in Escherichia coli speciation.
    J Mol Biol. 1991 Dec 20;222(4):851-6 PMID: 1762151
  9. Analysis of the Escherichia coli genome: DNA sequence of the region from 84.5 to 86.5 minutes.
    Science. 1992 Aug 7;257(5071):771-8 PMID: 1379743
  10. Mutational analysis of the Escherichia coli serB promoter region reveals transcriptional linkage to a downstream gene.
    Gene. 1992 Oct 12;120(1):1-9 PMID: 1327967
  11. First and second moment of counts of words in random texts generated by Markov chains.
    Comput Appl Biosci. 1992 Oct;8(5):433-41 PMID: 1422876
  12. Assessment of protein coding measures.
    Nucleic Acids Res. 1992 Dec 25;20(24):6441-50 PMID: 1480466
  13. Contamination of cDNA sequences in databases.
    Science. 1993 Mar 19;259(5102):1677-8 PMID: 8456288
  14. DNA sequence and analysis of 136 kilobases of the Escherichia coli genome: organizational symmetry around the origin of replication.
    Genomics. 1993 Jun;16(3):551-61 PMID: 7686882
  15. Analysis of the Escherichia coli genome. III. DNA sequence of the region from 87.2 to 89.2 minutes.
    Nucleic Acids Res. 1993 Jul 25;21(15):3391-8 PMID: 8346018
  16. Analysis of the Escherichia coli genome. IV. DNA sequence of the region from 89.2 to 92.8 minutes.
    Nucleic Acids Res. 1993 Nov 25;21(23):5408-17 PMID: 8265357
  17. Issues in searching molecular sequence databases.
    Nat Genet. 1994 Feb;6(2):119-29 PMID: 8162065
  18. Large scale bacterial gene discovery by similarity search.
    Nat Genet. 1994 Jun;7(2):205-14 PMID: 7920643
  19. New genes in old sequence: a strategy for finding genes in the bacterial genome.
    Trends Biochem Sci. 1994 Aug;19(8):309-13 PMID: 7940673
  20. Intrinsic and extrinsic approaches for detecting genes in a bacterial genome.
    Nucleic Acids Res. 1994 Nov 11;22(22):4756-67 PMID: 7984428
  21. A hidden Markov model that finds genes in E. coli DNA.
    Nucleic Acids Res. 1994 Nov 11;22(22):4768-78 PMID: 7984429
  22. Codon usage in bacteria: correlation with gene expressivity.
    Nucleic Acids Res. 1982 Nov 25;10(22):7055-74 PMID: 6760125
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
0305-1048
Published
1995-09-11
Pages
3554-62
Language
English
Region
England
NLM ID
0411011
PMCID
PMC307237
Subset
IM
Grants
NIGMS NIH HHS · GM47853 · United States
NHGRI NIH HHS · HG00783 · United States
Databases
PIR
S28499
SWISSPROT
P11036, P24554, P27398, P30864, P37572, P37598
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com