Home LiteratureArticle Details
PMID: 21589938 Published · epublish English Journal Article

Targeted assembly of short sequence reads.

PloS one ·Vol. 6 ·No. 5 ·2011-05-11 ·Pages e19816

Warren RL, Holt RA

Abstract

As next-generation sequence (NGS) production continues to increase, analysis is becoming a significant bottleneck. However, in situations where information is required only for specific sequence variants, it is not necessary to assemble or align whole genome data sets in their entirety. Rather, NGS data sets can be mined for the presence of sequence variants of interest by localized assembly, which is a faster, easier, and more accurate approach. We present TASR, a streamlined assembler that interrogates very large NGS data sets for the presence of specific variants by only considering reads within the sequence space of input target sequences provided by the user. The NGS data set is searched for reads with an exact match to all possible short words within the target sequence, and these reads are then assembled stringently to generate a consensus of the target and flanking sequence. Typically, variants of a particular locus are provided as different target sequences, and the presence of the variant in the data set being interrogated is revealed by a successful assembly outcome. However, TASR can also be used to find unknown sequences that flank a given target. We demonstrate that TASR has utility in finding or confirming genomic mutations, polymorphisms, fusions and integration events. Targeted assembly is a powerful method for interrogating large data sets for the presence of sequence variants of interest. TASR is a fast, flexible and easy to use tool for targeted assembly.

MeSH Terms
Breast Neoplasms/genetics Female Genome, Human HeLa Cells Humans Polymorphism, Single Nucleotide Sequence Analysis
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Warren René L
Genome Sciences Centre, British Columbia Cancer Agency, Vancouver, British Columbia, Canada. rwarren@bcgsc.ca
Holt Robert A
References (18)
18 references, click to expand
  1. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  2. A map of human genome variation from population-scale sequencing.
    Nature. 2010 Oct 28;467(7319):1061-73 PMID: 20981092
  3. Ancient human genome sequence of an extinct Palaeo-Eskimo.
    Nature. 2010 Feb 11;463(7282):757-62 PMID: 20148029
  4. Profiling the T-cell receptor beta-chain repertoire by massively parallel sequencing.
    Genome Res. 2009 Oct;19(10):1817-24 PMID: 19541912
  5. The case for cloud computing in genome informatics.
    Genome Biol. 2010;11(5):207 PMID: 20441614
  6. Building the sequence map of the human pan-genome.
    Nat Biotechnol. 2010 Jan;28(1):57-63 PMID: 19997067
  7. De novo assembly of human genomes with massively parallel short read sequencing.
    Genome Res. 2010 Feb;20(2):265-72 PMID: 20019144
  8. Expression of the TMPRSS2:ERG fusion gene predicts cancer recurrence after surgery for localised prostate cancer.
    Br J Cancer. 2007 Dec 17;97(12):1690-5 PMID: 17971772
  9. Mapping short DNA sequencing reads and calling variants using mapping quality scores.
    Genome Res. 2008 Nov;18(11):1851-8 PMID: 18714091
  10. The Sequence Alignment/Map format and SAMtools.
    Bioinformatics. 2009 Aug 15;25(16):2078-9 PMID: 19505943
  11. ABySS: a parallel assembler for short read sequence data.
    Genome Res. 2009 Jun;19(6):1117-23 PMID: 19251739
  12. Mutational evolution in a lobular breast tumour profiled at single nucleotide resolution.
    Nature. 2009 Oct 8;461(7265):809-13 PMID: 19812674
  13. Assembling millions of short DNA sequences using SSAKE.
    Bioinformatics. 2007 Feb 15;23(4):500-1 PMID: 17158514
  14. Extending assembly of short DNA sequences to handle error.
    Bioinformatics. 2007 Nov 1;23(21):2942-4 PMID: 17893086
  15. Profiling model T-cell metagenomes with short reads.
    Bioinformatics. 2009 Feb 15;25(4):458-64 PMID: 19136549
  16. Profiling the HeLa S3 transcriptome using randomly primed cDNA and massively parallel short-read sequencing.
    Biotechniques. 2008 Jul;45(1):81-94 PMID: 18611170
  17. Deep RNA sequencing analysis of readthrough gene fusions in human prostate adenocarcinoma and reference samples.
    BMC Med Genomics. 2011 Jan 24;4:11 PMID: 21261984
  18. SNVMix: predicting single nucleotide variants from next-generation sequencing of tumors.
    Bioinformatics. 2010 Mar 15;26(6):730-6 PMID: 20130035
Article Info
Journal
PloS one
Abbr.
PLoS One
ISSN
1932-6203
Published
2011-05-11
Epub
2011-00-11
Pages
e19816
Language
English
Region
United States
NLM ID
101285081
PMCID
PMC3092772
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com