Abstract
Amplification artifacts introduced during library preparation for the Illumina Genome Analyzer increase the likelihood that an appreciable proportion of these sequences will be duplicates and cause an uneven distribution of read coverage across the targeted sequencing regions. As a consequence, these unfavorable features result in difficulties in genome assembly and variation analysis from the short reads, particularly when the sequences are from genomes with base compositions at the extremes of high or low G+C content. Here we present an amplification-free method of library preparation, in which the cluster amplification step, rather than the PCR, enriches for fully ligated template strands, reducing the incidence of duplicate sequences, improving read mapping and single nucleotide polymorphism calling and aiding de novo assembly. We illustrate this by generating and analyzing DNA sequences from extremely (G+C)-poor (Plasmodium falciparum), (G+C)-neutral (Escherichia coli) and (G+C)-rich (Bordetella pertussis) genomes.
MeSH Terms
Base Composition
Base Sequence
Chromosome Mapping/methods
Gene Library
Molecular Sequence Data
Nucleic Acid Amplification Techniques
Polymorphism, Single Nucleotide/genetics
Sequence Analysis, DNA/methods
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Kozarewa Iwanka
The Wellcome Trust Sanger Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge, UK.
Ning Zemin
Quail Michael A
Sanders Mandy J
Berriman Matthew
Turner Daniel J
References (17)
17 references, click to expand
-
Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
Nucleic Acids Res. 2008 Sep;36(16):e105
PMID: 18660515
-
A large genome center's improvements to the Illumina sequencing system.
Nat Methods. 2008 Dec;5(12):1005-10
PMID: 19034268
-
Primer-directed enzymatic amplification of DNA with a thermostable DNA polymerase.
Science. 1988 Jan 29;239(4839):487-91
PMID: 2448875
-
The establishment of genomic DNA libraries for the human malaria parasite Plasmodium falciparum and identification of individual clones by hybridisation.
Mol Biochem Parasitol. 1982 Jun;5(6):391-400
PMID: 6213858
-
Identification of non-amplifying CYP21 genes when using PCR-based diagnosis of 21-hydroxylase deficiency in congenital adrenal hyperplasia (CAH) affected pedigrees.
Hum Mol Genet. 1996 Dec;5(12):2039-48
PMID: 8968761
-
Allele drop-out can occur in alleles differing by a single nucleotide and is not alleviated by preamplification or minor template increments.
Genet Test. 1998;2(4):351-5
PMID: 10464616
-
Large fragments of Plasmodium falciparum DNA can be stable when cloned in yeast artificial chromosomes.
Mol Biochem Parasitol. 1991 Feb;44(2):207-11
PMID: 2052022
-
Accurate whole human genome sequencing using reversible terminator chemistry.
Nature. 2008 Nov 6;456(7218):53-9
PMID: 18987734
-
Mapping short DNA sequencing reads and calling variants using mapping quality scores.
Genome Res. 2008 Nov;18(11):1851-8
PMID: 18714091
-
The genome of Plasmodium falciparum. I: DNA base composition.
Nucleic Acids Res. 1982 Jan 22;10(2):539-46
PMID: 6278419
-
PCR bias toward the wild-type k-ras and p53 sequences: implications for PCR detection of mutations and cancer diagnosis.
Biotechniques. 1998 Oct;25(4):684-91
PMID: 9793653
-
SSAHA: a fast search method for large DNA databases.
Genome Res. 2001 Oct;11(10):1725-9
PMID: 11591649
-
Quantification of PCR bias caused by a single nucleotide polymorphism in SMN gene dosage analysis.
J Mol Diagn. 2002 Nov;4(4):185-90
PMID: 12411585
-
Characterization of yeast artificial chromosomes from Plasmodium falciparum: construction of a stable, representative library and cloning of telomeric DNA fragments.
Genomics. 1992 Oct;14(2):332-9
PMID: 1427849
-
Construction and characterization of a Plasmodium vivax genomic library in yeast artificial chromosomes.
Genomics. 1997 Jun 15;42(3):467-73
PMID: 9205119
-
Genome sequence of the human malaria parasite Plasmodium falciparum.
Nature. 2002 Oct 3;419(6906):498-511
PMID: 12368864
-
Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
Genome Res. 2008 May;18(5):821-9
PMID: 18349386