Home LiteratureArticle Details
PMID: 2251125 Published · ppublish English Journal Article

Statistical analysis of nucleotide sequences.

Nucleic acids research ·Vol. 18 ·No. 22 ·1990-11-25 ·Pages 6641-7

Stückle EE, Emmrich C, Grob U, Nielsen PJ

Abstract

In order to scan nucleic acid databases for potentially relevant but as yet unknown signals, we have developed an improved statistical model for pattern analysis of nucleic acid sequences by modifying previous methods based on Markov chains. We demonstrate the importance of selecting the appropriate parameters in order for the method to function at all. The model allows the simultaneous analysis of several short sequences with unequal base frequencies and Markov order k not equal to 0 as is usually the case in databases. As a test of these modifications, we show that in E. coli sequences there is a bias against palindromic hexamers which correspond to known restriction enzyme recognition sites.

MeSH Terms
Base Sequence Databases, Factual Models, Statistical Molecular Sequence Data Pattern Recognition, Automated Repetitive Sequences, Nucleic Acid
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Stückle E E
Max-Planck-Institut für Immunbiologie, Freiburg, FRG.
Emmrich C
Grob U
Nielsen P J
References (27)
27 references, click to expand
  1. GenBank: current status and future directions.
    Methods Enzymol. 1990;183:3-22 PMID: 2314279
  2. K-tuple frequency analysis: from intron/exon discrimination to T-cell epitope mapping.
    Methods Enzymol. 1990;183:237-52 PMID: 1690334
  3. Prediction of the frequencies of restriction endonuclease recognition sequences using di- and mononucleotide frequencies.
    Biotechniques. 1988 Jan;6(1):34-40 PMID: 2908508
  4. Efficient algorithms for folding and comparing nucleic acid sequences.
    Nucleic Acids Res. 1982 Jan 11;10(1):197-206 PMID: 6174935
  5. Characterization of translational initiation sites in E. coli.
    Nucleic Acids Res. 1982 May 11;10(9):2971-96 PMID: 7048258
  6. Statistical characterization of nucleic acid sequence functional domains.
    Nucleic Acids Res. 1983 Apr 11;11(7):2205-20 PMID: 6835847
  7. A Markov analysis of DNA sequences.
    J Theor Biol. 1983 Oct 21;104(4):633-45 PMID: 6316035
  8. A comprehensive set of sequence analysis programs for the VAX.
    Nucleic Acids Res. 1984 Jan 11;12(1 Pt 1):387-95 PMID: 6546423
  9. Doublet frequencies in evolutionary distinct groups.
    Nucleic Acids Res. 1984 Feb 10;12(3):1749-63 PMID: 6583663
  10. Restriction and modification enzymes and their recognition sequences.
    Nucleic Acids Res. 1985;13 Suppl:r165-200 PMID: 2987885
  11. Markov chain analysis finds a significant influence of neighboring bases on the occurrence of a base in eucaryotic nuclear DNA sequences both protein-coding and noncoding.
    J Mol Evol. 1984-1985;21(3):278-88 PMID: 6443131
  12. Heuristic informational analysis of sequences.
    Nucleic Acids Res. 1986 Jan 10;14(1):179-96 PMID: 3753763
  13. Dinucleotide frequencies in different reading frame positions of coding mammalian DNA sequences.
    Biomed Biochim Acta. 1986;45(6):737-48 PMID: 3463303
  14. Highly recurring sequence elements identified in eukaryotic DNAs by computer analysis are often homologous to regulatory sequences or protein binding sites.
    Nucleic Acids Res. 1987 Feb 25;15(4):1835-51 PMID: 3822840
  15. Nucleotide quartets in the vicinity of eukaryotic transcriptional initiation sites: some DNA and chromatin structural implications.
    DNA. 1987 Feb;6(1):13-22 PMID: 3829887
  16. Mono- through hexanucleotide composition of the Escherichia coli genome: a Markov chain analysis.
    Nucleic Acids Res. 1987 Mar 25;15(6):2611-26 PMID: 3550699
  17. The effect of codon usage on the oligonucleotide composition of the E. coli genome and identification of over- and underrepresented sequences by Markov chain analysis.
    Nucleic Acids Res. 1987 Mar 25;15(6):2627-38 PMID: 3550700
  18. Theoretical molecular biology: prospectives and perspectives.
    J Theor Biol. 1987 Mar 21;125(2):219-35 PMID: 3657210
  19. The GenBank genetic sequence data bank.
    Nucleic Acids Res. 1988 Mar 11;16(5):1861-3 PMID: 3353225
  20. The EMBL data library.
    Nucleic Acids Res. 1988 Mar 11;16(5):1865-7 PMID: 3353226
  21. The protein identification resource (PIR).
    Nucleic Acids Res. 1988 Mar 11;16(5):1869-71 PMID: 3353227
  22. Mono- through hexanucleotide composition of the sense strand of yeast DNA: a Markov chain analysis.
    Nucleic Acids Res. 1988 Jul 25;16(14B):7145-58 PMID: 3043378
  23. Co-localization of rare oligonucleotides and regulatory elements in mammalian upstream gene regions.
    J Mol Biol. 1988 Sep 20;203(2):385-90 PMID: 3199439
  24. Linguistics of nucleotide sequences. I: The significance of deviations from mean statistical characteristics and prediction of the frequencies of occurrence of words.
    J Biomol Struct Dyn. 1989 Apr;6(5):1013-26 PMID: 2531596
  25. Linguistics of nucleotide sequences: morphology and comparison of vocabularies.
    J Biomol Struct Dyn. 1986 Aug;4(1):11-21 PMID: 3078230
  26. EMBL Data Library.
    Methods Enzymol. 1990;183:23-31 PMID: 2314277
  27. Protein sequence database.
    Methods Enzymol. 1990;183:31-49 PMID: 2314280
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
0305-1048
Published
1990-11-25
Pages
6641-7
Language
English
Region
England
NLM ID
0411011
PMCID
PMC332623
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com