Home LiteratureArticle Details
PMID: 21478889 Published · ppublish English Comparative Study Journal Article Research Support, N.I.H., Extramural

A framework for variation discovery and genotyping using next-generation DNA sequencing data.

Nature genetics ·Vol. 43 ·No. 5 ·2011-05-00 ·Pages 491-8

DePristo MA, Banks E, Poplin R, Garimella KV, Maguire JR, Hartl C, Philippakis AA, del Angel G, Rivas MA, Hanna M, McKenna A, Fennell TJ, Kernytsky AM, Sivachenko AY, Cibulskis K, Gabriel SB, Altshuler D, Daly MJ

Abstract

Recent advances in sequencing technology make it possible to comprehensively catalog genetic variation in population samples, creating a foundation for understanding human disease, ancestry and evolution. The amounts of raw data produced are prodigious, and many computational steps are required to translate this output into high-quality variant calls. We present a unified analytic framework to discover and genotype variation among multiple samples simultaneously that achieves sensitive and specific results across five sequencing technologies and three distinct, canonical experimental designs. Our process includes (i) initial read mapping; (ii) local realignment around indels; (iii) base quality score recalibration; (iv) SNP discovery and genotyping to find all potential variants; and (v) machine learning to separate true segregating variation from machine artifacts common to next-generation sequencing technologies. We here discuss the application of these tools, instantiated in the Genome Analysis Toolkit, to deep whole-genome, whole-exome capture and multi-sample low-pass (∼4×) 1000 Genomes Project datasets.

MeSH Terms
Data Interpretation, Statistical Databases, Nucleic Acid Exons Genetic Variation Genetics, Population/methods,statistics & numerical data Genome, Human Genotype Humans Polymorphism, Single Nucleotide Sequence Alignment/methods,statistics & numerical data Sequence Analysis, DNA/methods,statistics & numerical data Software
Authors & Affiliations
18 authors, click to expand affiliations / ORCID
DePristo Mark A
Program in Medical and Population Genetics, Broad Institute of Harvard and MIT, Cambridge, Massachusetts, USA. depristo@broadinstitute.org
Banks Eric
Poplin Ryan
Garimella Kiran V
Maguire Jared R
Hartl Christopher
Philippakis Anthony A
del Angel Guillermo
Rivas Manuel A
Hanna Matt
McKenna Aaron
Fennell Tim J
Kernytsky Andrew M
Sivachenko Andrey Y
Cibulskis Kristian
Gabriel Stacey B
Altshuler David
Daly Mark J
References (38)
38 references, click to expand
  1. SSAHA: a fast search method for large DNA databases.
    Genome Res. 2001 Oct;11(10):1725-9 PMID: 11591649
  2. Single nucleotide variation analysis in 65 candidate genes for CNS disorders in a representative sample of the European population.
    Genome Res. 2003 Oct;13(10):2271-6 PMID: 14525928
  3. Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
    Nucleic Acids Res. 2008 Sep;36(16):e105 PMID: 18660515
  4. SOAP2: an improved ultrafast tool for short read alignment.
    Bioinformatics. 2009 Aug 1;25(15):1966-7 PMID: 19497933
  5. A probabilistic approach for SNP discovery in high-throughput human resequencing data.
    Genome Res. 2009 Sep;19(9):1542-52 PMID: 19605794
  6. Base-calling of automated sequencer traces using phred. II. Error probabilities.
    Genome Res. 1998 Mar;8(3):186-94 PMID: 9521922
  7. Accurate SNP and mutation detection by targeted custom microarray-based genomic enrichment of short-fragment sequencing libraries.
    Nucleic Acids Res. 2010 Jun;38(10):e116 PMID: 20164091
  8. Mapping short DNA sequencing reads and calling variants using mapping quality scores.
    Genome Res. 2008 Nov;18(11):1851-8 PMID: 18714091
  9. A SNP discovery method to assess variant allele probability from next-generation resequencing data.
    Genome Res. 2010 Feb;20(2):273-80 PMID: 20019143
  10. Mapping human genetic diversity in Asia.
    Science. 2009 Dec 11;326(5959):1541-5 PMID: 20007900
  11. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  12. Searching for SNPs with cloud computing.
    Genome Biol. 2009;10(11):R134 PMID: 19930550
  13. Sequence and structural variation in a human genome uncovered by short-read, massively parallel ligation sequencing using two-base encoding.
    Genome Res. 2009 Sep;19(9):1527-41 PMID: 19546169
  14. Sequencing of 50 human exomes reveals adaptation to high altitude.
    Science. 2010 Jul 2;329(5987):75-8 PMID: 20595611
  15. The complete genome of an individual by massively parallel DNA sequencing.
    Nature. 2008 Apr 17;452(7189):872-6 PMID: 18421352
  16. The mutation spectrum revealed by paired genome sequences from a lung cancer patient.
    Nature. 2010 May 27;465(7297):473-7 PMID: 20505728
  17. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  18. A map of human genome variation from population-scale sequencing.
    Nature. 2010 Oct 28;467(7319):1061-73 PMID: 20981092
  19. Variation in genome-wide mutation rates within and between human families.
    Nat Genet. 2011 Jun 12;43(7):712-4 PMID: 21666693
  20. Targeted capture and massively parallel sequencing of 12 human exomes.
    Nature. 2009 Sep 10;461(7261):272-6 PMID: 19684571
  21. A comprehensive catalogue of somatic mutations from a human cancer genome.
    Nature. 2010 Jan 14;463(7278):191-6 PMID: 20016485
  22. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data.
    Genome Res. 2010 Sep;20(9):1297-303 PMID: 20644199
  23. Discovery and genotyping of genome structural polymorphism by sequencing on a population scale.
    Nat Genet. 2011 Mar;43(3):269-76 PMID: 21317889
  24. The landscape of somatic copy-number alteration across human cancers.
    Nature. 2010 Feb 18;463(7283):899-905 PMID: 20164920
  25. Simultaneous genotype calling and haplotype phasing improves genotype accuracy and reduces false-positive associations for genome-wide association studies.
    Am J Hum Genet. 2009 Dec;85(6):847-61 PMID: 19931040
  26. Quality scores and SNP detection in sequencing-by-synthesis systems.
    Genome Res. 2008 May;18(5):763-70 PMID: 18212088
  27. Solution hybrid selection with ultra-long oligonucleotides for massively parallel targeted sequencing.
    Nat Biotechnol. 2009 Feb;27(2):182-9 PMID: 19182786
  28. Genomewide comparison of DNA sequences between humans and chimpanzees.
    Am J Hum Genet. 2002 Jun;70(6):1490-7 PMID: 11992255
  29. Accurate whole human genome sequencing using reversible terminator chemistry.
    Nature. 2008 Nov 6;456(7218):53-9 PMID: 18987734
  30. High quality SNP calling using Illumina data at shallow coverage.
    Bioinformatics. 2010 Apr 15;26(8):1029-35 PMID: 20190250
  31. SNP detection for massively parallel whole-genome resequencing.
    Genome Res. 2009 Jun;19(6):1124-32 PMID: 19420381
  32. A draft sequence of the Neandertal genome.
    Science. 2010 May 7;328(5979):710-722 PMID: 20448178
  33. Human genome sequencing using unchained base reads on self-assembling DNA nanoarrays.
    Science. 2010 Jan 1;327(5961):78-81 PMID: 19892942
  34. Analysis of genetic inheritance in a family quartet by whole-genome sequencing.
    Science. 2010 Apr 30;328(5978):636-9 PMID: 20220176
  35. Adjust quality scores from alignment and improve sequencing accuracy.
    Nucleic Acids Res. 2004 Sep 30;32(17):5183-91 PMID: 15459287
  36. VarScan: variant detection in massively parallel sequencing of individual and pooled samples.
    Bioinformatics. 2009 Sep 1;25(17):2283-5 PMID: 19542151
  37. The Sequence Alignment/Map format and SAMtools.
    Bioinformatics. 2009 Aug 15;25(16):2078-9 PMID: 19505943
  38. Exome sequencing identifies the cause of a mendelian disorder.
    Nat Genet. 2010 Jan;42(1):30-5 PMID: 19915526
Article Info
Journal
Nature genetics
Abbr.
Nat Genet
ISSN
1546-1718
Published
2011-05-00
Epub
2011-00-10
Pages
491-8
Language
English
Region
United States
NLM ID
9216904
PMCID
PMC3083463
Subset
IM
Grants
NHGRI NIH HHS · U54 HG003067-01 · United States
NHGRI NIH HHS · U01 HG005208-01 · United States
NIDDK NIH HHS · P30 DK043351 · United States
NHGRI NIH HHS · 54 HG003067 · United States
NHGRI NIH HHS · U54 HG003067 · United States
NHGRI NIH HHS · U01 HG005208 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com