Home LiteratureArticle Details
PMID: 18505553 Published · epublish English Journal Article

Dinucleotide controlled null models for comparative RNA gene prediction.

BMC bioinformatics ·Vol. 9 ·2008-05-27 ·Pages 248

Gesell T, Washietl S

Abstract

Comparative prediction of RNA structures can be used to identify functional noncoding RNAs in genomic screens. It was shown recently by Babak et al. [BMC Bioinformatics. 8:33] that RNA gene prediction programs can be biased by the genomic dinucleotide content, in particular those programs using a thermodynamic folding model including stacking energies. As a consequence, there is need for dinucleotide-preserving control strategies to assess the significance of such predictions. While there have been randomization algorithms for single sequences for many years, the problem has remained challenging for multiple alignments and there is currently no algorithm available. We present a program called SISSIz that simulates multiple alignments of a given average dinucleotide content. Meeting additional requirements of an accurate null model, the randomized alignments are on average of the same sequence diversity and preserve local conservation and gap patterns. We make use of a phylogenetic substitution model that includes overlapping dependencies and site-specific rates. Using fast heuristics and a distance based approach, a tree is estimated under this model which is used to guide the simulations. The new algorithm is tested on vertebrate genomic alignments and the effect on RNA structure predictions is studied. In addition, we directly combined the new null model with the RNAalifold consensus folding algorithm giving a new variant of a thermodynamic structure based RNA gene finding program that is not biased by the dinucleotide content. SISSIz implements an efficient algorithm to randomize multiple alignments preserving dinucleotide content. It can be used to get more accurate estimates of false positive rates of existing programs, to produce negative controls for the training of machine learning based programs, or as standalone RNA gene finding program. Other applications in comparative genomics that require randomization of multiple alignments can be considered. SISSIz is available as open source C code that can be compiled for every major platform and downloaded here: http://sourceforge.net/projects/sissiz.

MeSH Terms
Algorithms Animals Base Composition Computational Biology/methods Humans Markov Chains RNA, Untranslated/chemistry Sequence Alignment Sequence Analysis, RNA/methods Sequence Homology, Nucleic Acid Software
Chemicals
RNA, Untranslated
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Gesell Tanja
Center for Integrative Bioinformatics Vienna, Max F. Perutz Laboratories, Dr. Bohr-Gasse 9, A-1030 Vienna, Austria. tanja.gesell@univie.ac.at
Washietl Stefan
References (51)
51 references, click to expand
  1. Simulating efficiently the evolution of DNA sequences.
    Comput Appl Biosci. 1995 Feb;11(1):111-5 PMID: 7796269
  2. Noncoding RNA gene detection using comparative sequence analysis.
    BMC Bioinformatics. 2001;2:8 PMID: 11801179
  3. A "Long Indel" model for evolutionary sequence alignment.
    Mol Biol Evol. 2004 Mar;21(3):529-40 PMID: 14694074
  4. Dating of the human-ape splitting by a molecular clock of mitochondrial DNA.
    J Mol Evol. 1985;22(2):160-74 PMID: 3934395
  5. An evolutionary model for maximum likelihood alignment of DNA sequences.
    J Mol Evol. 1991 Aug;33(2):114-24 PMID: 1920447
  6. Protein evolution with dependence among codons due to tertiary structure.
    Mol Biol Evol. 2003 Oct;20(10):1692-704 PMID: 12885968
  7. Pseudo-likelihood for non-reversible nucleotide substitution models with neighbour dependent rates.
    Stat Appl Genet Mol Biol. 2006;5:Article18 PMID: 17049029
  8. Statistical alignment based on fragment insertion and deletion models.
    Bioinformatics. 2003 Mar 1;19(4):490-9 PMID: 12611804
  9. Structural RNA has lower folding energy than random RNA of the same dinucleotide frequency.
    RNA. 2005 May;11(5):578-91 PMID: 15840812
  10. Identification and classification of conserved RNA secondary structures in the human genome.
    PLoS Comput Biol. 2006 Apr;2(4):e33 PMID: 16628248
  11. In silico sequence evolution with site-specific interactions along phylogenetic trees.
    Bioinformatics. 2006 Mar 15;22(6):716-22 PMID: 16332711
  12. A nucleotide substitution model with nearest-neighbour interactions.
    Bioinformatics. 2004 Aug 4;20 Suppl 1:i216-23 PMID: 15262802
  13. Genome-wide discovery and verification of novel structured RNAs in Plasmodium falciparum.
    Genome Res. 2008 Feb;18(2):281-92 PMID: 18096748
  14. Fast and reliable prediction of noncoding RNAs.
    Proc Natl Acad Sci U S A. 2005 Feb 15;102(7):2454-9 PMID: 15665081
  15. Seq-Gen: an application for the Monte Carlo simulation of DNA sequence evolution along phylogenetic trees.
    Comput Appl Biosci. 1997 Jun;13(3):235-8 PMID: 9183526
  16. Simultaneous statistical multiple alignment and phylogeny reconstruction.
    Syst Biol. 2005 Aug;54(4):548-61 PMID: 16085574
  17. Structured RNAs in the ENCODE selected regions of the human genome.
    Genome Res. 2007 Jun;17(6):852-64 PMID: 17568003
  18. Non-coding RNAs in Ciona intestinalis.
    Bioinformatics. 2005 Sep 1;21 Suppl 2:ii77-8 PMID: 16204130
  19. Thousands of corresponding human and mouse genomic regions unalignable in primary sequence contain common RNA structure.
    Genome Res. 2006 Jul;16(7):885-9 PMID: 16751343
  20. Mapping of conserved RNA secondary structures predicts thousands of functional noncoding RNAs in the human genome.
    Nat Biotechnol. 2005 Nov;23(11):1383-90 PMID: 16273071
  21. A dependent-rates model and an MCMC-based methodology for the maximum-likelihood analysis of sequences with overlapping reading frames.
    Mol Biol Evol. 2001 May;18(5):763-76 PMID: 11319261
  22. Aligning multiple genomic sequences with the threaded blockset aligner.
    Genome Res. 2004 Apr;14(4):708-15 PMID: 15060014
  23. Identification of differentially expressed small non-coding RNAs in the legume endosymbiont Sinorhizobium meliloti by comparative genomics.
    Mol Microbiol. 2007 Dec;66(5):1080-91 PMID: 17971083
  24. BIONJ: an improved version of the NJ algorithm based on a simple model of sequence data.
    Mol Biol Evol. 1997 Jul;14(7):685-95 PMID: 9254330
  25. The UCSC Genome Browser Database: 2008 update.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D773-9 PMID: 18086701
  26. Phylogenetic estimation of context-dependent substitution rates by maximum likelihood.
    Mol Biol Evol. 2004 Mar;21(3):468-88 PMID: 14660683
  27. No evidence that mRNAs have lower folding free energies than random sequences with the same dinucleotide distribution.
    Nucleic Acids Res. 1999 Dec 15;27(24):4816-22 PMID: 10572183
  28. MSARI: multiple sequence alignments for statistical detection of RNA secondary structure.
    Proc Natl Acad Sci U S A. 2004 Aug 17;101(33):12102-7 PMID: 15304649
  29. A stochastic model for the evolution of autocorrelated DNA sequences.
    Mol Phylogenet Evol. 1994 Sep;3(3):240-7 PMID: 7529616
  30. CMfinder--a covariance model based RNA motif finding algorithm.
    Bioinformatics. 2006 Feb 15;22(4):445-52 PMID: 16357030
  31. An updated and comprehensive rRNA phylogeny of (crown) eukaryotes based on rate-calibrated evolutionary distances.
    J Mol Evol. 2000 Dec;51(6):565-76 PMID: 11116330
  32. Identification of novel Drosophila melanogaster microRNAs.
    PLoS One. 2007 Nov 28;2(11):e1265 PMID: 18043761
  33. Identification of cyanobacterial non-coding RNAs by comparative genome analysis.
    Genome Biol. 2005;6(9):R73 PMID: 16168080
  34. Significance of nucleotide sequence alignments: a method for random sequence permutation that preserves dinucleotide and codon usage.
    Mol Biol Evol. 1985 Nov;2(6):526-38 PMID: 3870875
  35. Inching toward reality: an improved likelihood model of sequence evolution.
    J Mol Evol. 1992 Jan;34(1):3-16 PMID: 1556741
  36. Prediction of structured non-coding RNAs in the genomes of the nematodes Caenorhabditis elegans and Caenorhabditis briggsae.
    J Exp Zool B Mol Dev Evol. 2006 Jul 15;306(4):379-92 PMID: 16425273
  37. The covariation between TpA deficiency, CpG deficiency, and G+C content of human isochores is due to a mathematical artifact.
    Mol Biol Evol. 2000 Nov;17(11):1620-5 PMID: 11070050
  38. A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood.
    Syst Biol. 2003 Oct;52(5):696-704 PMID: 14530136
  39. Detection of non-coding RNAs on the basis of predicted secondary structure formation free energy change.
    BMC Bioinformatics. 2006 Mar 27;7:173 PMID: 16566836
  40. A new method for calculating evolutionary substitution rates.
    J Mol Evol. 1984;20(1):86-93 PMID: 6429346
  41. Calculation of folding energies of single-stranded nucleic acid sequences: conceptual issues.
    J Theor Biol. 2007 Oct 21;248(4):745-53 PMID: 17698086
  42. DNA sequence evolution with neighbor-dependent mutation.
    J Comput Biol. 2003;10(3-4):313-22 PMID: 12935330
  43. RNAs everywhere: genome-wide annotation of structured RNAs.
    J Exp Zool B Mol Dev Evol. 2007 Jan 15;308(1):1-25 PMID: 17171697
  44. Computational RNomics of drosophilids.
    BMC Genomics. 2007 Nov 08;8:406 PMID: 17996037
  45. Use of tiling array data and RNA secondary structure predictions to identify noncoding RNA genes.
    BMC Genomics. 2007 Jul 23;8:244 PMID: 17645787
  46. Considerations in the identification of functional RNA structural elements in genomic alignments.
    BMC Bioinformatics. 2007 Jan 30;8:33 PMID: 17263882
  47. Prediction of structural noncoding RNAs with RNAz.
    Methods Mol Biol. 2007;395:503-26 PMID: 17993695
  48. Secondary structure prediction for aligned RNA sequences.
    J Mol Biol. 2002 Jun 21;319(5):1059-66 PMID: 12079347
  49. Rfam: annotating non-coding RNAs in complete genomes.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D121-4 PMID: 15608160
  50. Consensus folding of aligned sequences as a new measure for the detection of functional RNAs by comparative genomics.
    J Mol Biol. 2004 Sep 3;342(1):19-30 PMID: 15313604
  51. Annotating noncoding RNA genes.
    Annu Rev Genomics Hum Genet. 2007;8:279-98 PMID: 17506659
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2008-05-27
Epub
2008-00-27
Pages
248
Language
English
Region
England
NLM ID
100965194
PMCID
PMC2453142
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com