Home LiteratureArticle Details
PMID: 27043002 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

Near-optimal probabilistic RNA-seq quantification.

Nature biotechnology ·Vol. 34 ·No. 5 ·2016-00-00 ·Pages 525-7

Bray NL, Pimentel H, Melsted P, Pachter L

Abstract

We present kallisto, an RNA-seq quantification program that is two orders of magnitude faster than previous approaches and achieves similar accuracy. Kallisto pseudoaligns reads to a reference, producing a list of transcripts that are compatible with each read while avoiding alignment of individual bases. We use kallisto to analyze 30 million unaligned paired-end RNA-seq reads in <10 min on a standard laptop computer. This removes a major computational bottleneck in RNA-seq analysis.

MeSH Terms
Algorithms Computer Simulation Data Interpretation, Statistical High-Throughput Nucleotide Sequencing/methods Models, Statistical Pattern Recognition, Automated/methods RNA/genetics Reproducibility of Results Sensitivity and Specificity Sequence Alignment/methods Sequence Analysis, RNA/methods Software
Chemicals
RNA
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Bray Nicolas L
Innovative Genomics Initiative, University of California, Berkeley, California, USA.
Pimentel Harold
Department of Computer Science, University of California, Berkeley, California, USA.
Melsted Páll ORCID
Faculty of Industrial Engineering, Mechanical Engineering and Computer Science, University of Iceland, Reykjavik, Iceland.
Pachter Lior
Department of Computer Science, University of California, Berkeley, California, USA. | Department of Mathematics, University of California, Berkeley, California, USA. | Department of Molecular &Cell Biology, University of California, Berkeley, California, USA.
References (16)
16 references, click to expand
  1. HTSeq--a Python framework to work with high-throughput sequencing data.
    Bioinformatics. 2015 Jan 15;31(2):166-9 PMID: 25260700
  2. Streaming fragment assignment for real-time analysis of sequencing experiments.
    Nat Methods. 2013 Jan;10(1):71-3 PMID: 23160280
  3. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome.
    BMC Bioinformatics. 2011 Aug 04;12:323 PMID: 21816040
  4. How to apply de Bruijn graphs to genome assembly.
    Nat Biotechnol. 2011 Nov 08;29(11):987-91 PMID: 22068540
  5. RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
    Genome Res. 2008 Sep;18(9):1509-17 PMID: 18550803
  6. TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions.
    Genome Biol. 2013 Apr 25;14(4):R36 PMID: 23618408
  7. Transcriptome and genome sequencing uncovers functional variation in humans.
    Nature. 2013 Sep 26;501(7468):506-11 PMID: 24037378
  8. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  9. A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control Consortium.
    Nat Biotechnol. 2014 Sep;32(9):903-14 PMID: 25150838
  10. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
    Nat Biotechnol. 2010 May;28(5):511-5 PMID: 20436464
  11. De novo assembly and genotyping of variants using colored de Bruijn graphs.
    Nat Genet. 2012 Jan 08;44(2):226-32 PMID: 22231483
  12. Sequence census methods for functional genomics.
    Nat Methods. 2008 Jan;5(1):19-21 PMID: 18165803
  13. EMSAR: estimation of transcript abundance from RNA-seq data by mappability-based segmentation and reclustering.
    BMC Bioinformatics. 2015 Sep 03;16:278 PMID: 26335049
  14. Improving RNA-Seq expression estimates by correcting for fragment bias.
    Genome Biol. 2011;12(3):R22 PMID: 21410973
  15. Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms.
    Nat Biotechnol. 2014 May;32(5):462-4 PMID: 24752080
  16. Snakemake--a scalable bioinformatics workflow engine.
    Bioinformatics. 2012 Oct 1;28(19):2520-2 PMID: 22908215
Article Info
Journal
Nature biotechnology
Abbr.
Nat Biotechnol
ISSN
1546-1696
Published
2016-00-00
Epub
2016-00-04
Pages
525-7
Language
English
Region
United States
NLM ID
9604648
Subset
IM
Grants
NHGRI NIH HHS · R01 HG006129 · United States
Corrections
ErratumIn
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com