Home LiteratureArticle Details
PMID: 21801405 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

Addressing challenges in the production and analysis of illumina sequencing data.

BMC genomics ·Vol. 12 ·2011-07-29 ·Pages 382

Kircher M, Heyn P, Kelso J

Abstract

Advances in DNA sequencing technologies have made it possible to generate large amounts of sequence data very rapidly and at substantially lower cost than capillary sequencing. These new technologies have specific characteristics and limitations that require either consideration during project design, or which must be addressed during data analysis. Specialist skills, both at the laboratory and the computational stages of project design and analysis, are crucial to the generation of high quality data from these new platforms. The Illumina sequencers (including the Genome Analyzers I/II/IIe/IIx and the new HiScan and HiSeq) represent a widely used platform providing parallel readout of several hundred million immobilized sequences using fluorescent-dye reversible-terminator chemistry. Sequencing library quality, sample handling, instrument settings and sequencing chemistry have a strong impact on sequencing run quality. The presence of adapter chimeras and adapter sequences at the end of short-insert molecules, as well as increased error rates and short read lengths complicate many computational analyses. We discuss here some of the factors that influence the frequency and severity of these problems and provide solutions for circumventing these. Further, we present a set of general principles for good analysis practice that enable problems with sequencing runs to be identified and dealt with.

MeSH Terms
Artifacts DNA Primers/genetics Gene Library Humans Image Processing, Computer-Assisted Quality Control Sequence Analysis, DNA/methods Statistics as Topic/methods
Chemicals
DNA Primers
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Kircher Martin
Max Planck Institute for Evolutionary Anthropology, Department of Evolutionary Genetics Deutscher Platz 6 04103 Leipzig, Germany.
Heyn Patricia
Kelso Janet
References (43)
43 references, click to expand
  1. Swift: primary data analysis for the Illumina Solexa sequencing platform.
    Bioinformatics. 2009 Sep 1;25(17):2194-9 PMID: 19549630
  2. Solid-phase reversible immobilization for the isolation of PCR products.
    Nucleic Acids Res. 1995 Nov 25;23(22):4742-3 PMID: 8524672
  3. A large genome center's improvements to the Illumina sequencing system.
    Nat Methods. 2008 Dec;5(12):1005-10 PMID: 19034268
  4. SOAP2: an improved ultrafast tool for short read alignment.
    Bioinformatics. 2009 Aug 1;25(15):1966-7 PMID: 19497933
  5. Massively parallel exon capture and library-free resequencing across 16 genomes.
    Nat Methods. 2009 May;6(5):315-6 PMID: 19349981
  6. Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
    Genome Res. 2008 May;18(5):821-9 PMID: 18349386
  7. TopHat: discovering splice junctions with RNA-Seq.
    Bioinformatics. 2009 May 1;25(9):1105-11 PMID: 19289445
  8. A complete mtDNA genome of an early modern human from Kostenki, Russia.
    Curr Biol. 2010 Feb 9;20(3):231-6 PMID: 20045327
  9. TileQC: a system for tile-based quality control of Solexa data.
    BMC Bioinformatics. 2008 May 28;9:250 PMID: 18507856
  10. Next-generation sequencing transforms today's biology.
    Nat Methods. 2008 Jan;5(1):16-8 PMID: 18165802
  11. Expression profiling of microRNAs by deep sequencing.
    Brief Bioinform. 2009 Sep;10(5):490-7 PMID: 19332473
  12. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  13. Removal of deaminated cytosines and detection of in vivo methylation in ancient DNA.
    Nucleic Acids Res. 2010 Apr;38(6):e87 PMID: 20028723
  14. Multiplex amplification of large sets of human exons.
    Nat Methods. 2007 Nov;4(11):931-6 PMID: 17934468
  15. A map of human genome variation from population-scale sequencing.
    Nature. 2010 Oct 28;467(7319):1061-73 PMID: 20981092
  16. The Neandertal genome and ancient DNA authenticity.
    EMBO J. 2009 Sep 2;28(17):2494-502 PMID: 19661919
  17. Probabilistic base calling of Solexa sequencing data.
    BMC Bioinformatics. 2008 Oct 13;9:431 PMID: 18851737
  18. Improved base calling for the Illumina Genome Analyzer using machine learning strategies.
    Genome Biol. 2009;10(8):R83 PMID: 19682367
  19. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  20. Sequencing technologies - the next generation.
    Nat Rev Genet. 2010 Jan;11(1):31-46 PMID: 19997069
  21. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data.
    Genome Res. 2010 Sep;20(9):1297-303 PMID: 20644199
  22. Illumina sequencing library preparation for highly multiplexed target capture and sequencing.
    Cold Spring Harb Protoc. 2010 Jun;2010(6):pdb.prot5448 PMID: 20516186
  23. High-throughput DNA sequencing--concepts and limitations.
    Bioessays. 2010 Jun;32(6):524-36 PMID: 20486139
  24. Next-generation DNA sequencing methods.
    Annu Rev Genomics Hum Genet. 2008;9:387-402 PMID: 18576944
  25. Next-generation DNA sequencing techniques.
    N Biotechnol. 2009 Apr;25(4):195-203 PMID: 19429539
  26. DNA recombination during PCR.
    Nucleic Acids Res. 1990 Apr 11;18(7):1687-91 PMID: 2186361
  27. Template-switching during DNA synthesis by Thermus aquaticus DNA polymerase I.
    Nucleic Acids Res. 1995 Jun 11;23(11):2049-57 PMID: 7596836
  28. Reducing the impact of PCR-mediated recombination in molecular evolution and environmental studies using a new-generation high-fidelity DNA polymerase.
    Biotechniques. 2009 Oct;47(4):857-66 PMID: 19852769
  29. BTA, a novel reagent for DNA attachment on glass and efficient generation of solid-phase amplified DNA colonies.
    Nucleic Acids Res. 2006 Feb 09;34(3):e22 PMID: 16473845
  30. Genetic history of an archaic hominin group from Denisova Cave in Siberia.
    Nature. 2010 Dec 23;468(7327):1053-60 PMID: 21179161
  31. Alta-Cyclic: a self-optimizing base caller for next-generation sequencing.
    Nat Methods. 2008 Aug;5(8):679-82 PMID: 18604217
  32. Accurate whole human genome sequencing using reversible terminator chemistry.
    Nature. 2008 Nov 6;456(7218):53-9 PMID: 18987734
  33. Targeted investigation of the Neandertal genome by array-based sequence capture.
    Science. 2010 May 7;328(5979):723-5 PMID: 20448179
  34. Biases in Illumina transcriptome sequencing caused by random hexamer priming.
    Nucleic Acids Res. 2010 Jul;38(12):e131 PMID: 20395217
  35. SNP detection for massively parallel whole-genome resequencing.
    Genome Res. 2009 Jun;19(6):1124-32 PMID: 19420381
  36. A draft sequence of the Neandertal genome.
    Science. 2010 May 7;328(5979):710-722 PMID: 20448178
  37. BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing.
    Genome Res. 2009 Oct;19(10):1884-95 PMID: 19661376
  38. DNA damage promotes jumping between templates during enzymatic amplification.
    J Biol Chem. 1990 Mar 15;265(8):4718-21 PMID: 2307682
  39. TagDust--a program to eliminate artifacts from next generation sequencing data.
    Bioinformatics. 2009 Nov 1;25(21):2839-40 PMID: 19737799
  40. The Sequence Alignment/Map format and SAMtools.
    Bioinformatics. 2009 Aug 15;25(16):2078-9 PMID: 19505943
  41. Target-enrichment strategies for next-generation sequencing.
    Nat Methods. 2010 Feb;7(2):111-8 PMID: 20111037
  42. De novo fragment assembly with short mate-paired reads: Does the read length matter?
    Genome Res. 2009 Feb;19(2):336-46 PMID: 19056694
  43. Fast mapping of short sequences with mismatches, insertions and deletions using index structures.
    PLoS Comput Biol. 2009 Sep;5(9):e1000502 PMID: 19750212
Article Info
Journal
BMC genomics
Abbr.
BMC Genomics
ISSN
1471-2164
Published
2011-07-29
Epub
2011-00-29
Pages
382
Language
English
Region
England
NLM ID
100965258
PMCID
PMC3163567
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com