Home LiteratureArticle Details
PMID: 21152010 Published · epublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

Identifying a high fraction of the human genome to be under selective constraint using GERP++.

PLoS computational biology ·Vol. 6 ·No. 12 ·2010-12-02 ·Pages e1001025

Davydov EV, Goode DL, Sirota M, Cooper GM, Sidow A, Batzoglou S

Abstract

Computational efforts to identify functional elements within genomes leverage comparative sequence information by looking for regions that exhibit evidence of selective constraint. One way of detecting constrained elements is to follow a bottom-up approach by computing constraint scores for individual positions of a multiple alignment and then defining constrained elements as segments of contiguous, highly scoring nucleotide positions. Here we present GERP++, a new tool that uses maximum likelihood evolutionary rate estimation for position-specific scoring and, in contrast to previous bottom-up methods, a novel dynamic programming approach to subsequently define constrained elements. GERP++ evaluates a richer set of candidate element breakpoints and ranks them based on statistical significance, eliminating the need for biased heuristic extension techniques. Using GERP++ we identify over 1.3 million constrained elements spanning over 7% of the human genome. We predict a higher fraction than earlier estimates largely due to the annotation of longer constrained elements, which improves one to one correspondence between predicted elements with known functional sequences. GERP++ is an efficient and effective tool to provide both nucleotide- and element-level constraint scores within deep multiple sequence alignments.

MeSH Terms
Algorithms Animals Genome, Human/genetics Genomics/methods Humans Mammals/genetics Models, Genetic Phylogeny Sequence Alignment/methods Sequence Analysis, DNA Software User-Computer Interface
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Davydov Eugene V
Department of Computer Science, Stanford University, Stanford, California, United States of America.
Goode David L
Sirota Marina
Cooper Gregory M
Sidow Arend
Batzoglou Serafim
References (18)
18 references, click to expand
  1. A simple method for estimating evolutionary rates of base substitutions through comparative studies of nucleotide sequences.
    J Mol Evol. 1980 Dec;16(2):111-20 PMID: 7463489
  2. Dating of the human-ape splitting by a molecular clock of mitochondrial DNA.
    J Mol Evol. 1985;22(2):160-74 PMID: 3934395
  3. CONTRAST: a discriminative, phylogeny-free approach to multiple informant de novo gene prediction.
    Genome Biol. 2007;8(12):R269 PMID: 18096039
  4. Analyses of deep mammalian sequence alignments and constraint predictions for 1% of the human genome.
    Genome Res. 2007 Jun;17(6):760-74 PMID: 17567995
  5. The UCSC Genome Browser database: update 2010.
    Nucleic Acids Res. 2010 Jan;38(Database issue):D613-9 PMID: 19906737
  6. Identifying novel constrained elements by exploiting biased substitution patterns.
    Bioinformatics. 2009 Jun 15;25(12):i54-62 PMID: 19478016
  7. Identification and characterization of multi-species conserved sequences.
    Genome Res. 2003 Dec;13(12):2507-18 PMID: 14656959
  8. The ENCODE (ENCyclopedia Of DNA Elements) Project.
    Science. 2004 Oct 22;306(5696):636-40 PMID: 15499007
  9. Evidence for a selectively favourable reduction in the mutation rate of the X chromosome.
    Nature. 1997 Mar 27;386(6623):388-92 PMID: 9121553
  10. Evolutionary trees from DNA sequences: a maximum likelihood approach.
    J Mol Evol. 1981;17(6):368-76 PMID: 7288891
  11. Identification and analysis of functional elements in 1% of the human genome by the ENCODE pilot project.
    Nature. 2007 Jun 14;447(7146):799-816 PMID: 17571346
  12. The UCSC Known Genes.
    Bioinformatics. 2006 May 1;22(9):1036-46 PMID: 16500937
  13. Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes.
    Genome Res. 2005 Aug;15(8):1034-50 PMID: 16024819
  14. Aligning multiple genomic sequences with the threaded blockset aligner.
    Genome Res. 2004 Apr;14(4):708-15 PMID: 15060014
  15. Distribution and intensity of constraint in mammalian genomic sequence.
    Genome Res. 2005 Jul;15(7):901-13 PMID: 15965027
  16. The human genome browser at UCSC.
    Genome Res. 2002 Jun;12(6):996-1006 PMID: 12045153
  17. Detection of nonneutral substitution rates on mammalian phylogenies.
    Genome Res. 2010 Jan;20(1):110-21 PMID: 19858363
  18. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
Article Info
Journal
PLoS computational biology
Abbr.
PLoS Comput Biol
ISSN
1553-7358
Published
2010-12-02
Epub
2010-00-02
Pages
e1001025
Language
English
Region
United States
NLM ID
101238922
PMCID
PMC2996323
Subset
IM
Grants
NLM NIH HHS · K22 LM008261 · United States
NLM NIH HHS · T15 LM007033 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com