Home LiteratureArticle Details
PMID: 17053093 Published · ppublish English Journal Article Research Support, N.I.H., Extramural

ESPERR: learning strong and weak signals in genomic sequence alignments to identify functional elements.

Genome research ·Vol. 16 ·No. 12 ·2006-12-00 ·Pages 1596-604

Taylor J, Tyekucheva S, King DC, Hardison RC, Miller W, Chiaromonte F

Abstract

Genomic sequence signals - such as base composition, presence of particular motifs, or evolutionary constraint - have been used effectively to identify functional elements. However, approaches based only on specific signals known to correlate with function can be quite limiting. When training data are available, application of computational learning algorithms to multispecies alignments has the potential to capture broader and more informative sequence and evolutionary patterns that better characterize a class of elements. However, effective exploitation of patterns in multispecies alignments is impeded by the vast number of possible alignment columns and by a limited understanding of which particular strings of columns may characterize a given class. We have developed a computational method, called ESPERR (evolutionary and sequence pattern extraction through reduced representations), which uses training examples to learn encodings of multispecies alignments into reduced forms tailored for the prediction of chosen classes of functional elements. ESPERR produces a greatly improved Regulatory Potential score, which can discriminate regulatory regions from neutral sites with excellent accuracy ( approximately 94%). This score captures strong signals (GC content and conservation), as well as subtler signals (with small contributions from many different alignment patterns) that characterize the regulatory elements in our training set. ESPERR is also effective for predicting other classes of functional elements, as we show for DNaseI hypersensitive sites and highly conserved regions with developmental enhancer activity. Our software, training data, and genome-wide predictions are available from our Web site (http://www.bx.psu.edu/projects/esperr).

MeSH Terms
Algorithms Base Composition Base Pairing Computational Biology Conserved Sequence Deoxyribonuclease I/immunology Enhancer Elements, Genetic Evolution, Molecular Genomics Molecular Sequence Data ROC Curve Regulatory Sequences, Nucleic Acid Sequence Alignment/methods
Chemicals
Deoxyribonuclease I
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Taylor James
Center for Comparative Genomics and Bioinformatics, The Pennsylvania State University, University Park, Pennsylvania 16802, USA. james@bx.psu.edu
Tyekucheva Svitlana
King David C
Hardison Ross C
Miller Webb
Chiaromonte Francesca
References (26)
26 references, click to expand
  1. Using multiple alignments to improve gene prediction.
    J Comput Biol. 2006 Mar;13(2):379-93 PMID: 16597247
  2. Comprehensive analysis of transcriptional promoter structure and function in 1% of the human genome.
    Genome Res. 2006 Jan;16(1):1-10 PMID: 16344566
  3. Experimental validation of predicted mammalian erythroid cis-regulatory modules.
    Genome Res. 2006 Dec;16(12):1480-92 PMID: 17038566
  4. Combining phylogenetic and hidden Markov models in biosequence analysis.
    J Comput Biol. 2004;11(2-3):413-28 PMID: 15285899
  5. Comparison of site-specific rate-inference methods for protein sequences: empirical Bayesian methods are superior.
    Mol Biol Evol. 2004 Sep;21(9):1781-91 PMID: 15201400
  6. Dating of the human-ape splitting by a molecular clock of mitochondrial DNA.
    J Mol Evol. 1985;22(2):160-74 PMID: 3934395
  7. Functional analysis of eve stripe 2 enhancer evolution in Drosophila: rules governing conservation and change.
    Development. 1998 Mar;125(5):949-58 PMID: 9449677
  8. Differences in the chromatin structure and cis-element organization of the human and mouse GATA1 loci: implications for cis-element identification.
    Blood. 2004 Nov 15;104(10):3106-16 PMID: 15265794
  9. Assessing computational tools for the discovery of transcription factor binding sites.
    Nat Biotechnol. 2005 Jan;23(1):137-44 PMID: 15637633
  10. Subtree power analysis and species selection for comparative genomics.
    Proc Natl Acad Sci U S A. 2005 May 31;102(22):7900-5 PMID: 15911755
  11. Predicting the in vivo signature of human gene regulatory sequences.
    Bioinformatics. 2005 Jun;21 Suppl 1:i338-43 PMID: 15961476
  12. Models of sequence evolution for DNA sequences containing gaps.
    Mol Biol Evol. 2001 Apr;18(4):481-90 PMID: 11264399
  13. Integrating genomic homology into gene structure prediction.
    Bioinformatics. 2001;17 Suppl 1:S140-8 PMID: 11473003
  14. Evolution of transcription factor binding sites in Mammalian gene regulatory regions: conservation and turnover.
    Mol Biol Evol. 2002 Jul;19(7):1114-21 PMID: 12082130
  15. Initial sequencing and comparative analysis of the mouse genome.
    Nature. 2002 Dec 5;420(6915):520-62 PMID: 12466850
  16. The UCSC Genome Browser Database.
    Nucleic Acids Res. 2003 Jan 1;31(1):51-4 PMID: 12519945
  17. Distinguishing regulatory DNA from neutral sites.
    Genome Res. 2003 Jan;13(1):64-72 PMID: 12529307
  18. Turnover of binding sites for transcription factors involved in early Drosophila development.
    Gene. 2003 May 22;310:215-20 PMID: 12801649
  19. Identification and characterization of multi-species conserved sequences.
    Genome Res. 2003 Dec;13(12):2507-18 PMID: 14656959
  20. Unbiased mapping of transcription factor binding sites along human chromosomes 21 and 22 points to widespread regulation of noncoding RNAs.
    Cell. 2004 Feb 20;116(4):499-509 PMID: 14980218
  21. Regulatory potential scores from genome-wide three-way alignments of human, mouse, and rat.
    Genome Res. 2004 Apr;14(4):700-7 PMID: 15060013
  22. Aligning multiple genomic sequences with the threaded blockset aligner.
    Genome Res. 2004 Apr;14(4):708-15 PMID: 15060014
  23. Into the heart of darkness: large-scale clustering of human non-coding DNA.
    Bioinformatics. 2004 Aug 4;20 Suppl 1:i40-8 PMID: 15262779
  24. Evaluation of regulatory potential and conservation scores for detecting cis-regulatory modules in aligned mammalian genome sequences.
    Genome Res. 2005 Aug;15(8):1051-60 PMID: 16024817
  25. Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes.
    Genome Res. 2005 Aug;15(8):1034-50 PMID: 16024819
  26. Unbiased location analysis of E2F1-binding sites suggests a widespread role for E2F1 in the human genome.
    Genome Res. 2006 May;16(5):595-605 PMID: 16606705
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
2006-12-00
Epub
2006-00-19
Pages
1596-604
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC1665643
Subset
IM
Grants
NHGRI NIH HHS · HG02238 · United States
NIDDK NIH HHS · R56 DK065806 · United States
NIDDK NIH HHS · R01 DK065806 · United States
NIDDK NIH HHS · DK65806 · United States
NHGRI NIH HHS · R01 HG002238 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com