Abstract
Accurate identification of novel, functional noncoding (nc) RNA features in genome sequence has proven more difficult than for exons. Current algorithms identify and score potential RNA secondary structures on the basis of thermodynamic stability, conservation, and/or covariance in sequence alignments. Neither the algorithms nor the information gained from the individual inputs have been independently assessed. Furthermore, due to issues in modelling background signal, it has been difficult to gauge the precision of these algorithms on a genomic scale, in which even a seemingly small false-positive rate can result in a vast excess of false discoveries. We developed a shuffling algorithm, shuffle-pair.pl, that simultaneously preserves dinucleotide frequency, gaps, and local conservation in pairwise sequence alignments. We used shuffle-pair.pl to assess precision and recall of six ncRNA search tools (MSARI, QRNA, ddbRNA, RNAz, Evofold, and several variants of simple thermodynamic stability on a test set of 3046 alignments of known ncRNAs. Relative to mononucleotide shuffling, preservation of dinucleotide content in shuffling the alignments resulted in a drastic increase in estimated false-positive detection rates for ncRNA elements, precluding evaluation of higher order alignments, which cannot not be adequately shuffled maintaining both dinucleotides and alignment structure. On pairwise alignments, none of the covariance-based tools performed markedly better than thermodynamic scoring alone. Although the high false-positive rates call into question the veracity of any individual predicted secondary structural element in our analysis, we nevertheless identified intriguing global trends in human genome alignments. The distribution of ncRNA prediction scores in 75-base windows overlapping UTRs, introns, and intergenic regions analyzed using both thermodynamic stability and EvoFold (which has no thermodynamic component) was significantly higher for real than shuffled sequence, while the distribution for coding sequences was lower than that of corresponding shuffles. Accurate prediction of novel RNA structural elements in genome sequence remains a difficult problem, and development of an appropriate negative-control strategy for multiple alignments is an important practical challenge. Nonetheless, the general trends we observed for the distributions of predicted ncRNAs across genomic features are biologically meaningful, supporting the presence of secondary structural elements in many 3' UTRs, and providing evidence for evolutionary selection against secondary structures in coding regions.
MeSH Terms
Algorithms
Base Sequence
Chromosome Mapping/methods
Conserved Sequence/genetics
Molecular Sequence Data
RNA/chemistry,genetics
Sequence Alignment/methods
Sequence Analysis, RNA/methods
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Babak Tomas
Banting and Best Department of Medical Research, Donnelly Centre for Cellular and Biomolecular Research, 160 College St, Toronto, ON M5S 3E1 Canada. tomas.babak@utoronto.ca <tomas.babak@utoronto.ca>
Blencowe Benjamin J
Hughes Timothy R
References (33)
33 references, click to expand
-
No evidence that mRNAs have lower folding free energies than random sequences with the same dinucleotide distribution.
Nucleic Acids Res. 1999 Dec 15;27(24):4816-22
PMID: 10572183
-
Mapping of conserved RNA secondary structures predicts thousands of functional noncoding RNAs in the human genome.
Nat Biotechnol. 2005 Nov;23(11):1383-90
PMID: 16273071
-
Identification of novel small RNAs using comparative genomics and microarrays.
Genes Dev. 2001 Jul 1;15(13):1637-51
PMID: 11445539
-
Novel small RNA-encoding genes in the intergenic regions of Escherichia coli.
Curr Biol. 2001 Jun 26;11(12):941-50
PMID: 11448770
-
Computational identification of noncoding RNAs in E. coli by comparative genomics.
Curr Biol. 2001 Sep 4;11(17):1369-73
PMID: 11553332
-
A computational approach to identify genes for functional RNAs in genomic sequences.
Nucleic Acids Res. 2001 Oct 1;29(19):3928-38
PMID: 11574674
-
BLAT--the BLAST-like alignment tool.
Genome Res. 2002 Apr;12(4):656-64
PMID: 11932250
-
Computational genomics of noncoding RNA genes.
Cell. 2002 Apr 19;109(2):137-40
PMID: 12007398
-
Initial sequencing and comparative analysis of the mouse genome.
Nature. 2002 Dec 5;420(6915):520-62
PMID: 12466850
-
Vienna RNA secondary structure server.
Nucleic Acids Res. 2003 Jul 1;31(13):3429-31
PMID: 12824340
-
Widespread selection for local RNA secondary structure in coding regions of bacterial genes.
Genome Res. 2003 Sep;13(9):2042-51
PMID: 12952875
-
ddbRNA: detection of conserved secondary structures in multiple alignments.
Bioinformatics. 2003 Sep 1;19(13):1606-11
PMID: 12967955
-
Noncoding RNA gene detection using comparative sequence analysis.
BMC Bioinformatics. 2001;2:8
PMID: 11801179
-
Benchmarking tools for the alignment of functional noncoding DNA.
BMC Bioinformatics. 2004 Jan 21;5:6
PMID: 14736341
-
Consensus folding of aligned sequences as a new measure for the detection of functional RNAs by comparative genomics.
J Mol Biol. 2004 Sep 3;342(1):19-30
PMID: 15313604
-
MSARI: multiple sequence alignments for statistical detection of RNA secondary structure.
Proc Natl Acad Sci U S A. 2004 Aug 17;101(33):12102-7
PMID: 15304649
-
Insertion mutagenesis to increase secondary structure within the 5' noncoding region of a eukaryotic mRNA reduces translational efficiency.
Cell. 1985 Mar;40(3):515-26
PMID: 2982496
-
The involvement of mRNA secondary structure in protein synthesis.
Biochem Cell Biol. 1987 Jun;65(6):576-81
PMID: 3322328
-
Significance of nucleotide sequence alignments: a method for random sequence permutation that preserves dinucleotide and codon usage.
Mol Biol Evol. 1985 Nov;2(6):526-38
PMID: 3870875
-
The UCSC Genome Browser Database: update 2006.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D590-8
PMID: 16381938
-
Identification and classification of conserved RNA secondary structures in the human genome.
PLoS Comput Biol. 2006 Apr;2(4):e33
PMID: 16628248
-
Thousands of corresponding human and mouse genomic regions unalignable in primary sequence contain common RNA structure.
Genome Res. 2006 Jul;16(7):885-9
PMID: 16751343
-
Detection of non-coding RNAs on the basis of predicted secondary structure formation free energy change.
BMC Bioinformatics. 2006;7:173
PMID: 16566836
-
Compilation of tRNA sequences and sequences of tRNA genes.
Nucleic Acids Res. 1998 Jan 1;26(1):148-53
PMID: 9399820
-
NONCODE: an integrated knowledge database of non-coding RNAs.
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D112-5
PMID: 15608158
-
Rfam: annotating non-coding RNAs in complete genomes.
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D121-4
PMID: 15608160
-
NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D501-4
PMID: 15608248
-
Fast and reliable prediction of noncoding RNAs.
Proc Natl Acad Sci U S A. 2005 Feb 15;102(7):2454-9
PMID: 15665081
-
Structural RNA has lower folding energy than random RNA of the same dinucleotide frequency.
RNA. 2005 May;11(5):578-91
PMID: 15840812
-
Pairwise local structural alignment of RNA sequences with sequence similarity less than 40%.
Bioinformatics. 2005 May 1;21(9):1815-24
PMID: 15657094
-
DINAMelt web server for nucleic acid melting prediction.
Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W577-81
PMID: 15980540
-
Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes.
Genome Res. 2005 Aug;15(8):1034-50
PMID: 16024819
-
Initial sequencing and analysis of the human genome.
Nature. 2001 Feb 15;409(6822):860-921
PMID: 11237011