Home LiteratureArticle Details
PMID: 22177264 Published · epublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

GC-content normalization for RNA-Seq data.

BMC bioinformatics ·Vol. 12 ·2011-12-17 ·Pages 480

Risso D, Schwartz K, Sherlock G, Dudoit S

Abstract

Transcriptome sequencing (RNA-Seq) has become the assay of choice for high-throughput studies of gene expression. However, as is the case with microarrays, major technology-related artifacts and biases affect the resulting expression measures. Normalization is therefore essential to ensure accurate inference of expression levels and subsequent analyses thereof. We focus on biases related to GC-content and demonstrate the existence of strong sample-specific GC-content effects on RNA-Seq read counts, which can substantially bias differential expression analysis. We propose three simple within-lane gene-level GC-content normalization approaches and assess their performance on two different RNA-Seq datasets, involving different species and experimental designs. Our methods are compared to state-of-the-art normalization procedures in terms of bias and mean squared error for expression fold-change estimation and in terms of Type I error and p-value distributions for tests of differential expression. The exploratory data analysis and normalization methods proposed in this article are implemented in the open-source Bioconductor R package EDASeq. Our within-lane normalization procedures, followed by between-lane normalization, reduce GC-content bias and lead to more accurate estimates of expression fold-changes and tests of differential expression. Such results are crucial for the biological interpretation of RNA-Seq experiments, where downstream analyses can be sensitive to the supplied lists of genes.

MeSH Terms
Base Composition Gene Expression Profiling Saccharomyces cerevisiae/genetics Sequence Analysis, RNA/methods Transcriptome
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Risso Davide
Division of Biostatistics and Department of Statistics, University of California, Berkeley, USA.
Schwartz Katja
Sherlock Gavin
Dudoit Sandrine
References (31)
31 references, click to expand
  1. Understanding mechanisms underlying human gene expression variation with RNA sequencing.
    Nature. 2010 Apr 1;464(7289):768-72 PMID: 20220758
  2. EGO-1, a C. elegans RdRP, modulates gene expression via production of mRNA-templated short antisense RNAs.
    Curr Biol. 2011 Mar 22;21(6):449-59 PMID: 21396820
  3. Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
    Nucleic Acids Res. 2008 Sep;36(16):e105 PMID: 18660515
  4. Polygenic and directional regulatory evolution across pathways in Saccharomyces.
    Proc Natl Acad Sci U S A. 2010 Mar 16;107(11):5058-63 PMID: 20194736
  5. Exploration, normalization, and summaries of high density oligonucleotide array probe level data.
    Biostatistics. 2003 Apr;4(2):249-64 PMID: 12925520
  6. Normalization for cDNA microarray data: a robust composite method addressing single and multiple slide systematic variation.
    Nucleic Acids Res. 2002 Feb 15;30(4):e15 PMID: 11842121
  7. Control-free calling of copy number alterations in deep-sequencing data using GC-content normalization.
    Bioinformatics. 2011 Jan 15;27(2):268-9 PMID: 21081509
  8. Rnnotator: an automated de novo transcriptome assembly pipeline from stranded RNA-Seq reads.
    BMC Genomics. 2010 Nov 24;11:663 PMID: 21106091
  9. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  10. Synthetic spike-in standards for RNA-seq experiments.
    Genome Res. 2011 Sep;21(9):1543-51 PMID: 21816910
  11. Improving RNA-Seq expression estimates by correcting for fragment bias.
    Genome Biol. 2011;12(3):R22 PMID: 21410973
  12. Linear models and empirical bayes methods for assessing differential expression in microarray experiments.
    Stat Appl Genet Mol Biol. 2004;3:Article3 PMID: 16646809
  13. Transcriptome analysis by strand-specific sequencing of complementary DNA.
    Nucleic Acids Res. 2009 Oct;37(18):e123 PMID: 19620212
  14. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  15. Transcript length bias in RNA-seq data confounds systems biology.
    Biol Direct. 2009 Apr 16;4:14 PMID: 19371405
  16. Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments.
    BMC Bioinformatics. 2010 Feb 18;11:94 PMID: 20167110
  17. Modeling non-uniformity in short-read rates in RNA-Seq data.
    Genome Biol. 2010;11(5):R50 PMID: 20459815
  18. Sensitive and accurate detection of copy number variants using read depth of coverage.
    Genome Res. 2009 Sep;19(9):1586-92 PMID: 19657104
  19. Gene ontology analysis for RNA-seq: accounting for selection bias.
    Genome Biol. 2010;11(2):R14 PMID: 20132535
  20. A scaling normalization method for differential expression analysis of RNA-seq data.
    Genome Biol. 2010;11(3):R25 PMID: 20196867
  21. A rapid and simple method for preparation of RNA from Saccharomyces cerevisiae.
    Nucleic Acids Res. 1990 May 25;18(10):3091-2 PMID: 2190191
  22. Impact of chromatin structures on DNA processing for genomic analyses.
    PLoS One. 2009 Aug 20;4(8):e6700 PMID: 19693276
  23. RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
    Genome Res. 2008 Sep;18(9):1509-17 PMID: 18550803
  24. Differential expression analysis for sequence count data.
    Genome Biol. 2010;11(10):R106 PMID: 20979621
  25. The transcriptional landscape of the yeast genome defined by RNA sequencing.
    Science. 2008 Jun 6;320(5881):1344-9 PMID: 18451266
  26. RNA-Seq: a revolutionary tool for transcriptomics.
    Nat Rev Genet. 2009 Jan;10(1):57-63 PMID: 19015660
  27. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.
    Bioinformatics. 2010 Jan 1;26(1):139-40 PMID: 19910308
  28. Accurate whole human genome sequencing using reversible terminator chemistry.
    Nature. 2008 Nov 6;456(7218):53-9 PMID: 18987734
  29. Biases in Illumina transcriptome sequencing caused by random hexamer priming.
    Nucleic Acids Res. 2010 Jul;38(12):e131 PMID: 20395217
  30. Bias detection and correction in RNA-Sequencing data.
    BMC Bioinformatics. 2011 Jul 19;12:290 PMID: 21771300
  31. The MicroArray Quality Control (MAQC) project shows inter- and intraplatform reproducibility of gene expression measurements.
    Nat Biotechnol. 2006 Sep;24(9):1151-61 PMID: 16964229
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2011-12-17
Epub
2011-00-17
Pages
480
Language
English
Region
England
NLM ID
100965194
PMCID
PMC3315510
Subset
IM
Grants
NIAID NIH HHS · R01 AI077737 · United States
NHGRI NIH HHS · R01 HG03468 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com