Home LiteratureArticle Details
PMID: 23975260 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Count-based differential expression analysis of RNA sequencing data using R and Bioconductor.

Nature protocols ·Vol. 8 ·No. 9 ·2013-09-00 ·Pages 1765-86

Anders S, McCarthy DJ, Chen Y, Okoniewski M, Smyth GK, Huber W, Robinson MD

Abstract

RNA sequencing (RNA-seq) has been rapidly adopted for the profiling of transcriptomes in many areas of biology, including studies into gene regulation, development and disease. Of particular interest is the discovery of differentially expressed genes across different conditions (e.g., tissues, perturbations) while optionally adjusting for other systematic factors that affect the data-collection process. There are a number of subtle yet crucial aspects of these analyses, such as read counting, appropriate treatment of biological variability, quality control checks and appropriate setup of statistical modeling. Several variations have been presented in the literature, and there is a need for guidance on current best practices. This protocol presents a state-of-the-art computational and statistical RNA-seq differential expression analysis workflow largely based on the free open-source R language and Bioconductor software and, in particular, on two widely used tools, DESeq and edgeR. Hands-on time for typical small experiments (e.g., 4-10 samples) can be <1 h, with computation time <1 d using a standard desktop PC.

MeSH Terms
Base Sequence Computational Biology/methods Gene Expression Profiling/methods Sequence Analysis, RNA/methods Software Workflow
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Anders Simon
Genome Biology Unit, European Molecular Biology Laboratory, Heidelberg, Germany.
McCarthy Davis J
Chen Yunshun
Okoniewski Michal
Smyth Gordon K
Huber Wolfgang
Robinson Mark D
References (58)
58 references, click to expand
  1. Differential expression in RNA-seq: a matter of depth.
    Genome Res. 2011 Dec;21(12):2213-23 PMID: 21903743
  2. GC-content normalization for RNA-Seq data.
    BMC Bioinformatics. 2011 Dec 17;12:480 PMID: 22177264
  3. Small-sample estimation of negative binomial dispersion, with applications to SAGE data.
    Biostatistics. 2008 Apr;9(2):321-32 PMID: 17728317
  4. The Arabidopsis nucleosome remodeler DDM1 allows DNA methyltransferases to access H1-containing heterochromatin.
    Cell. 2013 Mar 28;153(1):193-205 PMID: 23540698
  5. Savant Genome Browser 2: visualization and analysis for population-scale genomics.
    Nucleic Acids Res. 2012 Jul;40(Web Server issue):W615-21 PMID: 22638571
  6. Differential expression analysis for sequence count data.
    Genome Biol. 2010;11(10):R106 PMID: 20979621
  7. Epigenetic expansion of VHL-HIF signal output drives multiorgan metastasis in renal cancer.
    Nat Med. 2013 Jan;19(1):50-6 PMID: 23223005
  8. Capturing heterogeneity in gene expression studies by surrogate variable analysis.
    PLoS Genet. 2007 Sep;3(9):1724-35 PMID: 17907809
  9. The Sequence Alignment/Map format and SAMtools.
    Bioinformatics. 2009 Aug 15;25(16):2078-9 PMID: 19505943
  10. Integrative Genomics Viewer (IGV): high-performance genomics data visualization and exploration.
    Brief Bioinform. 2013 Mar;14(2):178-92 PMID: 22517427
  11. Tools for mapping high-throughput sequencing data.
    Bioinformatics. 2012 Dec 15;28(24):3169-77 PMID: 23060614
  12. Conservation of an RNA regulatory map between Drosophila and mammals.
    Genome Res. 2011 Feb;21(2):193-202 PMID: 20921232
  13. A powerful and flexible approach to the analysis of RNA sequence count data.
    Bioinformatics. 2011 Oct 1;27(19):2672-8 PMID: 21810900
  14. Copy-number-aware differential analysis of quantitative DNA sequencing data.
    Genome Res. 2012 Dec;22(12):2489-96 PMID: 22879430
  15. A comprehensive comparison of RNA-Seq-based transcriptome analysis from reads to differential gene expression and cross-comparison with microarrays: a case study in Saccharomyces cerevisiae.
    Nucleic Acids Res. 2012 Nov 1;40(20):10084-97 PMID: 22965124
  16. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
    Nat Biotechnol. 2010 May;28(5):511-5 PMID: 20436464
  17. Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments.
    BMC Bioinformatics. 2010 Feb 18;11:94 PMID: 20167110
  18. Gene Expression Omnibus: NCBI gene expression and hybridization array data repository.
    Nucleic Acids Res. 2002 Jan 1;30(1):207-10 PMID: 11752295
  19. easyRNASeq: a bioconductor package for processing RNA-Seq data.
    Bioinformatics. 2012 Oct 1;28(19):2532-3 PMID: 22847932
  20. Sequencing technology does not eliminate biological variability.
    Nat Biotechnol. 2011 Jul 11;29(7):572-3 PMID: 21747377
  21. Differential oestrogen receptor binding is associated with clinical outcome in breast cancer.
    Nature. 2012 Jan 04;481(7381):389-93 PMID: 22217937
  22. Statistical design and analysis of RNA sequencing data.
    Genetics. 2010 Jun;185(2):405-16 PMID: 20439781
  23. Differential expression analysis of multifactor RNA-Seq experiments with respect to biological variation.
    Nucleic Acids Res. 2012 May;40(10):4288-97 PMID: 22287627
  24. Full-length transcriptome assembly from RNA-Seq data without a reference genome.
    Nat Biotechnol. 2011 May 15;29(7):644-52 PMID: 21572440
  25. Independent filtering increases detection power for high-throughput experiments.
    Proc Natl Acad Sci U S A. 2010 May 25;107(21):9546-51 PMID: 20460310
  26. Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks.
    Nat Protoc. 2012 Mar 01;7(3):562-78 PMID: 22383036
  27. Identifying differentially expressed transcripts from RNA-seq data with biological variation.
    Bioinformatics. 2012 Jul 1;28(13):1721-8 PMID: 22563066
  28. Bayesian analysis of RNA sequencing data by estimating multiple shrinkage priors.
    Biostatistics. 2013 Jan;14(1):113-28 PMID: 22988280
  29. Unproductive splicing of SR genes associated with highly conserved and ultraconserved DNA elements.
    Nature. 2007 Apr 19;446(7138):926-9 PMID: 17361132
  30. The Subread aligner: fast, accurate and scalable read mapping by seed-and-vote.
    Nucleic Acids Res. 2013 May 1;41(10):e108 PMID: 23558742
  31. Preferred analysis methods for single genomic regions in RNA sequencing revealed by processing the shape of coverage.
    Nucleic Acids Res. 2012 May;40(9):e63 PMID: 22210855
  32. Foxp3 exploits a pre-existent enhancer landscape for regulatory T cell lineage specification.
    Cell. 2012 Sep 28;151(1):153-66 PMID: 23021222
  33. A scaling normalization method for differential expression analysis of RNA-seq data.
    Genome Biol. 2010;11(3):R25 PMID: 20196867
  34. Moderated statistical tests for assessing differences in tag abundance.
    Bioinformatics. 2007 Nov 1;23(21):2881-7 PMID: 17881408
  35. RNA-Seq: a revolutionary tool for transcriptomics.
    Nat Rev Genet. 2009 Jan;10(1):57-63 PMID: 19015660
  36. Proteomic analysis reveals new cardiac-specific dystrophin-associated proteins.
    PLoS One. 2012;7(8):e43515 PMID: 22937058
  37. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.
    Bioinformatics. 2010 Jan 1;26(1):139-40 PMID: 19910308
  38. TopHat: discovering splice junctions with RNA-Seq.
    Bioinformatics. 2009 May 1;25(9):1105-11 PMID: 19289445
  39. ShortRead: a bioconductor package for input, quality assessment and exploration of high-throughput sequence data.
    Bioinformatics. 2009 Oct 1;25(19):2607-8 PMID: 19654119
  40. Bioconductor: open software development for computational biology and bioinformatics.
    Genome Biol. 2004;5(10):R80 PMID: 15461798
  41. Fast and SNP-tolerant detection of complex variants and splicing in short reads.
    Bioinformatics. 2010 Apr 1;26(7):873-81 PMID: 20147302
  42. A new shrinkage estimator for dispersion improves differential expression detection in RNA-seq data.
    Biostatistics. 2013 Apr;14(2):232-43 PMID: 23001152
  43. Sex-specific and lineage-specific alternative splicing in primates.
    Genome Res. 2010 Feb;20(2):180-9 PMID: 20009012
  44. Rev-Erbs repress macrophage gene expression by inhibiting enhancer-directed transcription.
    Nature. 2013 Jun 27;498(7455):511-5 PMID: 23728303
  45. A comparison of methods for differential expression analysis of RNA-seq data.
    BMC Bioinformatics. 2013 Mar 09;14:91 PMID: 23497356
  46. Differential gene expression in the siphonophore Nanomia bijuga (Cnidaria) assessed with multiple next-generation sequencing workflows.
    PLoS One. 2011;6(7):e22953 PMID: 21829563
  47. STAR: ultrafast universal RNA-seq aligner.
    Bioinformatics. 2013 Jan 1;29(1):15-21 PMID: 23104886
  48. How to map billions of short reads onto genomes.
    Nat Biotechnol. 2009 May;27(5):455-7 PMID: 19430453
  49. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  50. baySeq: empirical Bayesian methods for identifying differential expression in sequence count data.
    BMC Bioinformatics. 2010 Aug 10;11:422 PMID: 20698981
  51. Reproducible research: a bioinformatics case study.
    Stat Appl Genet Mol Biol. 2005;4:Article2 PMID: 16646837
  52. Detecting differential expression in RNA-sequence data using quasi-likelihood with shrunken dispersion estimates.
    Stat Appl Genet Mol Biol. 2012 Oct 22;11(5): PMID: 23104842
  53. Detecting differential usage of exons from RNA-seq data.
    Genome Res. 2012 Oct;22(10):2008-17 PMID: 22722343
  54. MapSplice: accurate mapping of RNA-seq reads for splice junction discovery.
    Nucleic Acids Res. 2010 Oct;38(18):e178 PMID: 20802226
  55. Using control genes to correct for unwanted variation in microarray data.
    Biostatistics. 2012 Jul;13(3):539-52 PMID: 22101192
  56. Removing technical variability in RNA-seq data using conditional quantile normalization.
    Biostatistics. 2012 Apr;13(2):204-16 PMID: 22285995
  57. Savant: genome browser for high-throughput sequencing data.
    Bioinformatics. 2010 Aug 15;26(16):1938-44 PMID: 20562449
  58. Tackling the widespread and critical impact of batch effects in high-throughput data.
    Nat Rev Genet. 2010 Oct;11(10):733-9 PMID: 20838408
Article Info
Journal
Nature protocols
Abbr.
Nat Protoc
ISSN
1750-2799
Published
2013-09-00
Epub
2013-00-22
Pages
1765-86
Language
English
Region
England
NLM ID
101284307
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com