Home LiteratureArticle Details
PMID: 25338716 Published · ppublish English Journal Article Research Support, N.I.H., Extramural

DANN: a deep learning approach for annotating the pathogenicity of genetic variants.

Bioinformatics (Oxford, England) ·Vol. 31 ·No. 5 ·2015-03-01 ·Pages 761-3

Quang D, Chen Y, Xie X

Abstract

Annotating genetic variants, especially non-coding variants, for the purpose of identifying pathogenic variants remains a challenge. Combined annotation-dependent depletion (CADD) is an algorithm designed to annotate both coding and non-coding variants, and has been shown to outperform other annotation algorithms. CADD trains a linear kernel support vector machine (SVM) to differentiate evolutionarily derived, likely benign, alleles from simulated, likely deleterious, variants. However, SVMs cannot capture non-linear relationships among the features, which can limit performance. To address this issue, we have developed DANN. DANN uses the same feature set and training data as CADD to train a deep neural network (DNN). DNNs can capture non-linear relationships among features and are better suited than SVMs for problems with a large number of samples and features. We exploit Compute Unified Device Architecture-compatible graphics processing units and deep learning techniques such as dropout and momentum training to accelerate the DNN training. DANN achieves about a 19% relative reduction in the error rate and about a 14% relative increase in the area under the curve (AUC) metric over CADD's SVM methodology. All data and source code are available at https://cbcl.ics.uci.edu/public_data/DANN/.

MeSH Terms
Algorithms Area Under Curve Computer Graphics Genetic Variation/genetics Genome, Human Humans Molecular Sequence Annotation Neural Networks, Computer Selection, Genetic Support Vector Machine
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Quang Daniel
Department of Computer Science and Center for Complex Biological Systems, University of California, Irvine, CA 92697, USA Department of Computer Science and Center for Complex Biological Systems, University of California, Irvine, CA 92697, USA.
Chen Yifei
Department of Computer Science and Center for Complex Biological Systems, University of California, Irvine, CA 92697, USA.
Xie Xiaohui
Department of Computer Science and Center for Complex Biological Systems, University of California, Irvine, CA 92697, USA Department of Computer Science and Center for Complex Biological Systems, University of California, Irvine, CA 92697, USA.
References (3)
3 references, click to expand
  1. One-stop shop for disease genes.
    Nature. 2012 Nov 8;491(7423):171 PMID: 23135443
  2. Analysis of 6,515 exomes reveals the recent origin of most human protein-coding variants.
    Nature. 2013 Jan 10;493(7431):216-20 PMID: 23201682
  3. A general framework for estimating the relative pathogenicity of human genetic variants.
    Nat Genet. 2014 Mar;46(3):310-5 PMID: 24487276
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2015-03-01
Epub
2014-00-22
Pages
761-3
Language
English
Region
England
NLM ID
9808944
PMCID
PMC4341060
Subset
IM
Grants
NIBIB NIH HHS · T32 EB009418 · United States
NHGRI NIH HHS · R01 HG006870 · United States
NIGMS NIH HHS · P50 GM076516 · United States
NIBIB NIH HHS · EB009418 · United States
NHGRI NIH HHS · R01HG006870 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com