Home LiteratureArticle Details
PMID: 30371827 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

CADD: predicting the deleteriousness of variants throughout the human genome.

Nucleic acids research ·Vol. 47 ·No. D1 ·2019-00-08 ·Pages D886-D894

Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M

Abstract

Combined Annotation-Dependent Depletion (CADD) is a widely used measure of variant deleteriousness that can effectively prioritize causal variants in genetic analyses, particularly highly penetrant contributors to severe Mendelian disorders. CADD is an integrative annotation built from more than 60 genomic features, and can score human single nucleotide variants and short insertion and deletions anywhere in the reference assembly. CADD uses a machine learning model trained on a binary distinction between simulated de novo variants and variants that have arisen and become fixed in human populations since the split between humans and chimpanzees; the former are free of selective pressure and may thus include both neutral and deleterious alleles, while the latter are overwhelmingly neutral (or, at most, weakly deleterious) by virtue of having survived millions of years of purifying selection. Here we review the latest updates to CADD, including the most recent version, 1.4, which supports the human genome build GRCh38. We also present updates to our website that include simplified variant lookup, extended documentation, an Application Program Interface and improved mechanisms for integrating CADD scores into other tools or applications. CADD scores, software and documentation are available at https://cadd.gs.washington.edu.

MeSH Terms
Databases, Nucleic Acid Genetic Variation Genome, Human Humans Machine Learning Molecular Sequence Annotation
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Rentzsch Philipp
Berlin Institute of Health (BIH), 10178 Berlin, Germany. | Charité - Universitätsmedizin Berlin, 10117 Berlin, Germany.
Witten Daniela
Department of Statistics and Biostatistics, University of Washington, Seattle, WA 98195, USA.
Cooper Gregory M
HudsonAlpha Institute for Biotechnology, Huntsville, AL 35806, USA.
Shendure Jay
Department of Genome Sciences, University of Washington, Seattle, WA 98195, USA. | Brotman Baty Institute for Precision Medicine, Seattle, WA 98195, USA.
Kircher Martin
Berlin Institute of Health (BIH), 10178 Berlin, Germany. | Charité - Universitätsmedizin Berlin, 10117 Berlin, Germany. | Department of Genome Sciences, University of Washington, Seattle, WA 98195, USA.
References (57)
57 references, click to expand
  1. Human Gene Mutation Database (HGMD): 2003 update.
    Hum Mutat. 2003 Jun;21(6):577-81 PMID: 12754702
  2. SIFT: Predicting amino acid changes that affect protein function.
    Nucleic Acids Res. 2003 Jul 1;31(13):3812-4 PMID: 12824425
  3. Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes.
    Genome Res. 2005 Aug;15(8):1034-50 PMID: 16024819
  4. Targeted capture and massively parallel sequencing of 12 human exomes.
    Nature. 2009 Sep 10;461(7261):272-6 PMID: 19684571
  5. Detection of nonneutral substitution rates on mammalian phylogenies.
    Genome Res. 2010 Jan;20(1):110-21 PMID: 19858363
  6. High-resolution analysis of DNA regulatory elements by synthetic saturation mutagenesis.
    Nat Biotechnol. 2009 Dec;27(12):1173-5 PMID: 19915551
  7. A method and server for predicting damaging missense mutations.
    Nat Methods. 2010 Apr;7(4):248-9 PMID: 20354512
  8. Single-nucleotide evolutionary constraint scores highlight disease-causing mutations.
    Nat Methods. 2010 Apr;7(4):250-1 PMID: 20354513
  9. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data.
    Nucleic Acids Res. 2010 Sep;38(16):e164 PMID: 20601685
  10. Identifying a high fraction of the human genome to be under selective constraint using GERP++.
    PLoS Comput Biol. 2010 Dec 02;6(12):e1001025 PMID: 21152010
  11. Tabix: fast retrieval of sequence features from generic TAB-delimited files.
    Bioinformatics. 2011 Mar 1;27(5):718-9 PMID: 21208982
  12. Needles in stacks of needles: finding disease-causal variants in a wealth of genomic data.
    Nat Rev Genet. 2011 Aug 18;12(9):628-40 PMID: 21850043
  13. Massively parallel functional dissection of mammalian enhancers in vivo.
    Nat Biotechnol. 2012 Feb 26;30(3):265-70 PMID: 22371081
  14. ClinVar: public archive of relationships among sequence variation and human phenotype.
    Nucleic Acids Res. 2014 Jan;42(Database issue):D980-5 PMID: 24234437
  15. A general framework for estimating the relative pathogenicity of human genetic variants.
    Nat Genet. 2014 Mar;46(3):310-5 PMID: 24487276
  16. DANN: a deep learning approach for annotating the pathogenicity of genetic variants.
    Bioinformatics. 2015 Mar 1;31(5):761-3 PMID: 25338716
  17. Integrating functional data to prioritize causal variants in statistical fine-mapping studies.
    PLoS Genet. 2014 Oct 30;10(10):e1004722 PMID: 25357204
  18. Approximation to the distribution of fitness effects across functional categories in human segregating polymorphisms.
    PLoS Genet. 2014 Nov 06;10(11):e1004697 PMID: 25375159
  19. RNA splicing. The human splicing code reveals new insights into the genetic determinants of disease.
    Science. 2015 Jan 9;347(6218):1254806 PMID: 25525159
  20. Evaluation of CADD Scores in Curated Mismatch Repair Gene Variants Yields a Model for Clinical Validation and Prioritization.
    Hum Mutat. 2015 Jul;36(7):712-9 PMID: 25871441
  21. A method to predict the impact of regulatory variants from DNA sequence.
    Nat Genet. 2015 Aug;47(8):955-61 PMID: 26075791
  22. Predicting effects of noncoding variants with deep learning-based sequence model.
    Nat Methods. 2015 Oct;12(10):931-4 PMID: 26301843
  23. dbNSFP v3.0: A One-Stop Database of Functional Predictions and Annotations for Human Nonsynonymous and Splice-Site SNVs.
    Hum Mutat. 2016 Mar;37(3):235-41 PMID: 26555599
  24. A spectral approach integrating functional genomic annotations for coding and noncoding variants.
    Nat Genet. 2016 Feb;48(2):214-20 PMID: 26727659
  25. The mutation significance cutoff: gene-level thresholds for variant predictions.
    Nat Methods. 2016 Feb;13(2):109-10 PMID: 26820543
  26. Ensembl comparative genomics resources.
    Database (Oxford). 2016 May 02;2016:null PMID: 27141089
  27. The Ensembl Variant Effect Predictor.
    Genome Biol. 2016 Jun 06;17(1):122 PMID: 27268795
  28. TP53 Variations in Human Cancers: New Lessons from the IARC TP53 Database and Genomics Data.
    Hum Mutat. 2016 Sep;37(9):865-76 PMID: 27328919
  29. Analysis of protein-coding genetic variation in 60,706 humans.
    Nature. 2016 Aug 17;536(7616):285-91 PMID: 27535533
  30. A Whole-Genome Analysis Framework for Effective Identification of Pathogenic Regulatory Variants in Mendelian Disease.
    Am J Hum Genet. 2016 Sep 1;99(3):595-606 PMID: 27569544
  31. REVEL: An Ensemble Method for Predicting the Pathogenicity of Rare Missense Variants.
    Am J Hum Genet. 2016 Oct 6;99(4):877-885 PMID: 27666373
  32. M-CAP eliminates a majority of variants of uncertain significance in clinical exomes at high sensitivity.
    Nat Genet. 2016 Dec;48(12):1581-1586 PMID: 27776117
  33. SVScore: an impact prediction tool for structural variation.
    Bioinformatics. 2017 Apr 1;33(7):1083-1085 PMID: 28031184
  34. GAVIN: Gene-Aware Variant INterpretation for medical sequencing.
    Genome Biol. 2017 Jan 16;18(1):6 PMID: 28093075
  35. IMHOTEP-a composite score integrating popular tools for predicting the functional consequences of non-synonymous sequence variants.
    Nucleic Acids Res. 2017 Feb 17;45(3):e13 PMID: 28180317
  36. Impacts of Neanderthal-Introgressed Sequences on the Landscape of Human Gene Expression.
    Cell. 2017 Feb 23;168(5):916-927.e12 PMID: 28235201
  37. Fast, scalable prediction of deleterious noncoding variants from functional and population genomic data.
    Nat Genet. 2017 Apr;49(4):618-624 PMID: 28288115
  38. Ensembl core software resources: storage and programmatic access for DNA sequence and genome annotation.
    Database (Oxford). 2017 Jan 1;2017(1): PMID: 28365736
  39. Characterization of pathogenic SORL1 genetic variants for association with Alzheimer's disease: a clinical interpretation strategy.
    Eur J Hum Genet. 2017 Aug;25(8):973-981 PMID: 28537274
  40. Genomic diagnosis for children with intellectual disability and/or developmental delay.
    Genome Med. 2017 May 30;9(1):43 PMID: 28554332
  41. Using the Neandertal genome to study the evolution of small insertions and deletions in modern humans.
    BMC Evol Biol. 2017 Aug 4;17(1):179 PMID: 28778150
  42. Variant Interpretation: Functional Assays to the Rescue.
    Am J Hum Genet. 2017 Sep 7;101(3):315-325 PMID: 28886340
  43. DNA sequencing at 40: past, present and future.
    Nature. 2017 Oct 19;550(7676):345-353 PMID: 29019985
  44. Deep learning of the regulatory grammar of yeast 5' untranslated regions from 500,000 random sequences.
    Genome Res. 2017 Dec;27(12):2015-2024 PMID: 29097404
  45. The UCSC Genome Browser database: 2018 update.
    Nucleic Acids Res. 2018 Jan 4;46(D1):D762-D769 PMID: 29106570
  46. Evaluation of in silico algorithms for use with ACMG/AMP clinical variant interpretation guidelines.
    Genome Biol. 2017 Nov 28;18(1):225 PMID: 29179779
  47. Functional mapping and annotation of genetic associations with FUMA.
    Nat Commun. 2017 Nov 28;8(1):1826 PMID: 29184056
  48. Quantitative Missense Variant Effect Prediction Using Large-Scale Mutagenesis Data.
    Cell Syst. 2018 Jan 24;6(1):116-124.e3 PMID: 29226803
  49. The human noncoding genome defined by genetic diversity.
    Nat Genet. 2018 Mar;50(3):333-337 PMID: 29483654
  50. Demographic History and Genetic Adaptation in the Himalayan Region Inferred from Genome-Wide SNP Genotypes of 49 Populations.
    Mol Biol Evol. 2018 Aug 1;35(8):1916-1933 PMID: 29796643
  51. Deep learning sequence-based ab initio prediction of variant effects on expression and disease risk.
    Nat Genet. 2018 Aug;50(8):1171-1179 PMID: 30013180
  52. Predicting the clinical impact of human mutation with deep neural networks.
    Nat Genet. 2018 Aug;50(8):1161-1170 PMID: 30038395
  53. Accurate classification of BRCA1 variants with saturation genome editing.
    Nature. 2018 Oct;562(7726):217-222 PMID: 30209399
  54. Predicting variant deleteriousness in non-human species: applying the CADD approach in mouse.
    BMC Bioinformatics. 2018 Oct 12;19(1):373 PMID: 30314430
  55. PopViz: a webserver for visualizing minor allele frequencies and damage prediction scores of human genetic variations.
    Bioinformatics. 2018 Dec 15;34(24):4307-4309 PMID: 30535305
  56. Amino acid difference formula to help explain protein evolution.
    Science. 1974 Sep 6;185(4154):862-4 PMID: 4843792
  57. Base-calling of automated sequencer traces using phred. II. Error probabilities.
    Genome Res. 1998 Mar;8(3):186-94 PMID: 9521922
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2019-00-08
Pages
D886-D894
Language
English
Region
England
NLM ID
0411011
PMCID
PMC6323892
Subset
IM
Grants
NCI NIH HHS · R01 CA197139 · United States
NHGRI NIH HHS · UM1 HG006493 · United States
NHGRI NIH HHS · U54 HG006493 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com