Abstract
As next-generation sequence (NGS) production continues to increase, analysis is becoming a significant bottleneck. However, in situations where information is required only for specific sequence variants, it is not necessary to assemble or align whole genome data sets in their entirety. Rather, NGS data sets can be mined for the presence of sequence variants of interest by localized assembly, which is a faster, easier, and more accurate approach. We present TASR, a streamlined assembler that interrogates very large NGS data sets for the presence of specific variants by only considering reads within the sequence space of input target sequences provided by the user. The NGS data set is searched for reads with an exact match to all possible short words within the target sequence, and these reads are then assembled stringently to generate a consensus of the target and flanking sequence. Typically, variants of a particular locus are provided as different target sequences, and the presence of the variant in the data set being interrogated is revealed by a successful assembly outcome. However, TASR can also be used to find unknown sequences that flank a given target. We demonstrate that TASR has utility in finding or confirming genomic mutations, polymorphisms, fusions and integration events. Targeted assembly is a powerful method for interrogating large data sets for the presence of sequence variants of interest. TASR is a fast, flexible and easy to use tool for targeted assembly.
MeSH Terms
Breast Neoplasms/genetics
Female
Genome, Human
HeLa Cells
Humans
Polymorphism, Single Nucleotide
Sequence Analysis
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Warren René L
Genome Sciences Centre, British Columbia Cancer Agency, Vancouver, British Columbia, Canada. rwarren@bcgsc.ca
Holt Robert A
References (18)
18 references, click to expand
-
Fast and accurate short read alignment with Burrows-Wheeler transform.
Bioinformatics. 2009 Jul 15;25(14):1754-60
PMID: 19451168
-
A map of human genome variation from population-scale sequencing.
Nature. 2010 Oct 28;467(7319):1061-73
PMID: 20981092
-
Ancient human genome sequence of an extinct Palaeo-Eskimo.
Nature. 2010 Feb 11;463(7282):757-62
PMID: 20148029
-
Profiling the T-cell receptor beta-chain repertoire by massively parallel sequencing.
Genome Res. 2009 Oct;19(10):1817-24
PMID: 19541912
-
The case for cloud computing in genome informatics.
Genome Biol. 2010;11(5):207
PMID: 20441614
-
Building the sequence map of the human pan-genome.
Nat Biotechnol. 2010 Jan;28(1):57-63
PMID: 19997067
-
De novo assembly of human genomes with massively parallel short read sequencing.
Genome Res. 2010 Feb;20(2):265-72
PMID: 20019144
-
Expression of the TMPRSS2:ERG fusion gene predicts cancer recurrence after surgery for localised prostate cancer.
Br J Cancer. 2007 Dec 17;97(12):1690-5
PMID: 17971772
-
Mapping short DNA sequencing reads and calling variants using mapping quality scores.
Genome Res. 2008 Nov;18(11):1851-8
PMID: 18714091
-
The Sequence Alignment/Map format and SAMtools.
Bioinformatics. 2009 Aug 15;25(16):2078-9
PMID: 19505943
-
ABySS: a parallel assembler for short read sequence data.
Genome Res. 2009 Jun;19(6):1117-23
PMID: 19251739
-
Mutational evolution in a lobular breast tumour profiled at single nucleotide resolution.
Nature. 2009 Oct 8;461(7265):809-13
PMID: 19812674
-
Assembling millions of short DNA sequences using SSAKE.
Bioinformatics. 2007 Feb 15;23(4):500-1
PMID: 17158514
-
Extending assembly of short DNA sequences to handle error.
Bioinformatics. 2007 Nov 1;23(21):2942-4
PMID: 17893086
-
Profiling model T-cell metagenomes with short reads.
Bioinformatics. 2009 Feb 15;25(4):458-64
PMID: 19136549
-
Profiling the HeLa S3 transcriptome using randomly primed cDNA and massively parallel short-read sequencing.
Biotechniques. 2008 Jul;45(1):81-94
PMID: 18611170
-
Deep RNA sequencing analysis of readthrough gene fusions in human prostate adenocarcinoma and reference samples.
BMC Med Genomics. 2011 Jan 24;4:11
PMID: 21261984
-
SNVMix: predicting single nucleotide variants from next-generation sequencing of tumors.
Bioinformatics. 2010 Mar 15;26(6):730-6
PMID: 20130035