Home LiteratureArticle Details
PMID: 26335049 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

EMSAR: estimation of transcript abundance from RNA-seq data by mappability-based segmentation and reclustering.

BMC bioinformatics ·Vol. 16 ·2015-09-03 ·Pages 278

Lee S, Seo CH, Alver BH, Lee S, Park PJ

Abstract

RNA-seq has been widely used for genome-wide expression profiling. RNA-seq data typically consists of tens of millions of short sequenced reads from different transcripts. However, due to sequence similarity among genes and among isoforms, the source of a given read is often ambiguous. Existing approaches for estimating expression levels from RNA-seq reads tend to compromise between accuracy and computational cost. We introduce a new approach for quantifying transcript abundance from RNA-seq data. EMSAR (Estimation by Mappability-based Segmentation And Reclustering) groups reads according to the set of transcripts to which they are mapped and finds maximum likelihood estimates using a joint Poisson model for each optimal set of segments of transcripts. The method uses nearly all mapped reads, including those mapped to multiple genes. With an efficient transcriptome indexing based on modified suffix arrays, EMSAR minimizes the use of CPU time and memory while achieving accuracy comparable to the best existing methods. EMSAR is a method for quantifying transcripts from RNA-seq data with high accuracy and low computational cost. EMSAR is available at https://github.com/parklab/emsar.

MeSH Terms
Base Sequence Gene Expression Profiling/methods Genome/genetics Protein Isoforms/genetics RNA/genetics Sequence Analysis, RNA/methods Transcriptome
Chemicals
Protein Isoforms RNA
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Lee Soohyun
Department of Biomedical Informatics, Harvard Medical School, Boston, MA, USA.
Seo Chae Hwa
Emerging Technology Center, DNA link, Seoul, South Korea.
Alver Burak Han
Department of Biomedical Informatics, Harvard Medical School, Boston, MA, USA.
Lee Sanghyuk
Emerging Technology Center, DNA link, Seoul, South Korea. | Ewha Womans University, Seoul, Korea.
Park Peter J
Department of Biomedical Informatics, Harvard Medical School, Boston, MA, USA. peter_park@hms.harvard.edu. | Informatics Program, Boston Children's Hospital and Division of Genetics, Brigham and Women's Hospital, Boston, MA, USA. peter_park@hms.harvard.edu.
References (26)
26 references, click to expand
  1. A global view of gene activity and alternative splicing by deep sequencing of the human transcriptome.
    Science. 2008 Aug 15;321(5891):956-60 PMID: 18599741
  2. Streaming fragment assignment for real-time analysis of sequencing experiments.
    Nat Methods. 2013 Jan;10(1):71-3 PMID: 23160280
  3. Transcriptome analysis by strand-specific sequencing of complementary DNA.
    Nucleic Acids Res. 2009 Oct;37(18):e123 PMID: 19620212
  4. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.
    Bioinformatics. 2010 Jan 1;26(1):139-40 PMID: 19910308
  5. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  6. A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control Consortium.
    Nat Biotechnol. 2014 Sep;32(9):903-14 PMID: 25150838
  7. Differential expression analysis for sequence count data.
    Genome Biol. 2010;11(10):R106 PMID: 20979621
  8. The MicroArray Quality Control (MAQC) project shows inter- and intraplatform reproducibility of gene expression measurements.
    Nat Biotechnol. 2006 Sep;24(9):1151-61 PMID: 16964229
  9. Large scale real-time PCR validation on gene expression measurements from two commercial long-oligonucleotide microarrays.
    BMC Genomics. 2006;7:59 PMID: 16551369
  10. Statistical inferences for isoform expression in RNA-Seq.
    Bioinformatics. 2009 Apr 15;25(8):1026-32 PMID: 19244387
  11. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
    Nat Biotechnol. 2010 May;28(5):511-5 PMID: 20436464
  12. Immunohistological localization of the adhesion molecules L1, N-CAM, and MAG in the developing and adult optic nerve of mice.
    J Comp Neurol. 1989 Jun 15;284(3):451-62 PMID: 2474006
  13. A comprehensive evaluation of normalization methods for Illumina high-throughput RNA sequencing data analysis.
    Brief Bioinform. 2013 Nov;14(6):671-83 PMID: 22988256
  14. voom: Precision weights unlock linear model analysis tools for RNA-seq read counts.
    Genome Biol. 2014;15(2):R29 PMID: 24485249
  15. Estimation of alternative splicing isoform frequencies from RNA-Seq data.
    Algorithms Mol Biol. 2011 Apr 19;6(1):9 PMID: 21504602
  16. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome.
    BMC Bioinformatics. 2011;12:323 PMID: 21816040
  17. Accurate quantification of transcriptome from RNA-Seq data by effective length normalization.
    Nucleic Acids Res. 2011 Jan;39(2):e9 PMID: 21059678
  18. Genome sequence of the palaeopolyploid soybean.
    Nature. 2010 Jan 14;463(7278):178-83 PMID: 20075913
  19. Comprehensive comparative analysis of strand-specific RNA sequencing methods.
    Nat Methods. 2010 Sep;7(9):709-15 PMID: 20711195
  20. Alternative isoform regulation in human tissue transcriptomes.
    Nature. 2008 Nov 27;456(7221):470-6 PMID: 18978772
  21. Modelling and simulating generic RNA-Seq experiments with the flux simulator.
    Nucleic Acids Res. 2012 Nov 1;40(20):10073-83 PMID: 22962361
  22. A strand-specific library preparation protocol for RNA sequencing.
    Methods Enzymol. 2011;500:79-98 PMID: 21943893
  23. Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms.
    Nat Biotechnol. 2014 May;32(5):462-4 PMID: 24752080
  24. Accurate estimation of expression levels of homologous genes in RNA-seq experiments.
    J Comput Biol. 2011 Mar;18(3):459-68 PMID: 21385047
  25. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  26. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2.
    Genome Biol. 2014;15(12):550 PMID: 25516281
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2015-09-03
Epub
2015-00-03
Pages
278
Language
English
Region
England
NLM ID
100965194
PMCID
PMC4559005
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com