Home LiteratureArticle Details
PMID: 23104886 Published · ppublish English Evaluation Study Journal Article Research Support, N.I.H., Extramural

STAR: ultrafast universal RNA-seq aligner.

Bioinformatics (Oxford, England) ·Vol. 29 ·No. 1 ·2013-01-01 ·Pages 15-21

Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, Batut P, Chaisson M, Gingeras TR

Abstract

Accurate alignment of high-throughput RNA-seq data is a challenging and yet unsolved problem because of the non-contiguous transcript structure, relatively short read lengths and constantly increasing throughput of the sequencing technologies. Currently available RNA-seq aligners suffer from high mapping error rates, low mapping speed, read length limitation and mapping biases. To align our large (>80 billon reads) ENCODE Transcriptome RNA-seq dataset, we developed the Spliced Transcripts Alignment to a Reference (STAR) software based on a previously undescribed RNA-seq alignment algorithm that uses sequential maximum mappable seed search in uncompressed suffix arrays followed by seed clustering and stitching procedure. STAR outperforms other aligners by a factor of >50 in mapping speed, aligning to the human genome 550 million 2 × 76 bp paired-end reads per hour on a modest 12-core server, while at the same time improving alignment sensitivity and precision. In addition to unbiased de novo detection of canonical junctions, STAR can discover non-canonical splices and chimeric (fusion) transcripts, and is also capable of mapping full-length RNA sequences. Using Roche 454 sequencing of reverse transcription polymerase chain reaction amplicons, we experimentally validated 1960 novel intergenic splice junctions with an 80-90% success rate, corroborating the high precision of the STAR mapping strategy. STAR is implemented as a standalone C++ code. STAR is free open source software distributed under GPLv3 license and can be downloaded from http://code.google.com/p/rna-star/.

MeSH Terms
Algorithms Cluster Analysis Gene Expression Profiling Genome, Human Humans RNA Splicing Sequence Alignment/methods Sequence Analysis, RNA/methods Software
Authors & Affiliations
9 authors, click to expand affiliations / ORCID
Dobin Alexander
Cold Spring Harbor Laboratory, Cold Spring Harbor, NY, USA. dobin@cshl.edu
Davis Carrie A
Schlesinger Felix
Drenkow Jorg
Zaleski Chris
Jha Sonali
Batut Philippe
Chaisson Mark
Gingeras Thomas R
References (21)
21 references, click to expand
  1. Pre-mRNA splicing: where and when in the nucleus.
    Trends Cell Biol. 2011 Jun;21(6):336-43 PMID: 21514162
  2. Pre-mRNA splicing in the new millennium.
    Curr Opin Cell Biol. 2001 Jun;13(3):302-9 PMID: 11343900
  3. Comparative analysis of RNA-Seq alignment algorithms and the RNA-Seq unified mapper (RUM).
    Bioinformatics. 2011 Sep 15;27(18):2518-28 PMID: 21775302
  4. Landscape of transcription in human cells.
    Nature. 2012 Sep 6;489(7414):101-8 PMID: 22955620
  5. Detection of splice junctions from paired-end RNA-seq data by SpliceMap.
    Nucleic Acids Res. 2010 Aug;38(14):4570-8 PMID: 20371516
  6. TopHat: discovering splice junctions with RNA-Seq.
    Bioinformatics. 2009 May 1;25(9):1105-11 PMID: 19289445
  7. GENCODE: the reference human genome annotation for The ENCODE Project.
    Genome Res. 2012 Sep;22(9):1760-74 PMID: 22955987
  8. Alignment of whole genomes.
    Nucleic Acids Res. 1999 Jun 1;27(11):2369-76 PMID: 10325427
  9. BLAT--the BLAST-like alignment tool.
    Genome Res. 2002 Apr;12(4):656-64 PMID: 11932250
  10. Mauve: multiple alignment of conserved genomic sequence with rearrangements.
    Genome Res. 2004 Jul;14(7):1394-403 PMID: 15231754
  11. MapSplice: accurate mapping of RNA-seq reads for splice junction discovery.
    Nucleic Acids Res. 2010 Oct;38(18):e178 PMID: 20802226
  12. PASSion: a pattern growth algorithm-based pipeline for splice junction detection in paired-end RNA-Seq data.
    Bioinformatics. 2012 Feb 15;28(4):479-86 PMID: 22219203
  13. Versatile and open software for comparing large genomes.
    Genome Biol. 2004;5(2):R12 PMID: 14759262
  14. progressiveMauve: multiple genome alignment with gene gain, loss and rearrangement.
    PLoS One. 2010 Jun 25;5(6):e11147 PMID: 20593022
  15. ChIP-seq guidelines and practices of the ENCODE and modENCODE consortia.
    Genome Res. 2012 Sep;22(9):1813-31 PMID: 22955991
  16. Fast and SNP-tolerant detection of complex variants and splicing in short reads.
    Bioinformatics. 2010 Apr 1;26(7):873-81 PMID: 20147302
  17. Fast algorithms for large-scale genome alignment and comparison.
    Nucleic Acids Res. 2002 Jun 1;30(11):2478-83 PMID: 12034836
  18. Optimal spliced alignments of short sequence reads.
    Bioinformatics. 2008 Aug 15;24(16):i174-80 PMID: 18689821
  19. Direct detection of DNA methylation during single-molecule, real-time sequencing.
    Nat Methods. 2010 Jun;7(6):461-5 PMID: 20453866
  20. Transcriptome analysis by strand-specific sequencing of complementary DNA.
    Nucleic Acids Res. 2009 Oct;37(18):e123 PMID: 19620212
  21. An integrated semiconductor device enabling non-optical genome sequencing.
    Nature. 2011 Jul 20;475(7356):348-52 PMID: 21776081
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2013-01-01
Epub
2012-00-25
Pages
15-21
Language
English
Region
England
NLM ID
9808944
PMCID
PMC3530905
Subset
IM
Grants
NHGRI NIH HHS · U54HG004557 · United States
Databases
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com