Home LiteratureArticle Details
PMID: 21771300 Published · epublish English Journal Article Research Support, N.I.H., Extramural

Bias detection and correction in RNA-Sequencing data.

BMC bioinformatics ·Vol. 12 ·2011-07-19 ·Pages 290

Zheng W, Chung LM, Zhao H

Abstract

High throughput sequencing technology provides us unprecedented opportunities to study transcriptome dynamics. Compared to microarray-based gene expression profiling, RNA-Seq has many advantages, such as high resolution, low background, and ability to identify novel transcripts. Moreover, for genes with multiple isoforms, expression of each isoform may be estimated from RNA-Seq data. Despite these advantages, recent work revealed that base level read counts from RNA-Seq data may not be randomly distributed and can be affected by local nucleotide composition. It was not clear though how the base level read count bias may affect gene level expression estimates. In this paper, by using five published RNA-Seq data sets from different biological sources and with different data preprocessing schemes, we showed that commonly used estimates of gene expression levels from RNA-Seq data, such as reads per kilobase of gene length per million reads (RPKM), are biased in terms of gene length, GC content and dinucleotide frequencies. We directly examined the biases at the gene-level, and proposed a simple generalized-additive-model based approach to correct different sources of biases simultaneously. Compared to previously proposed base level correction methods, our method reduces bias in gene-level expression estimates more effectively. Our method identifies and corrects different sources of biases in gene-level expression measures from RNA-Seq data, and provides more accurate estimates of gene expression levels from RNA-Seq. This method should prove useful in meta-analysis of gene expression levels using different platforms or experimental protocols.

MeSH Terms
Gene Expression Profiling Humans Meta-Analysis as Topic RNA/genetics,metabolism Sequence Analysis, RNA/methods Yeasts/genetics
Chemicals
RNA
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Zheng Wei
Biostatistics Resource, Keck Laboratory, Yale University, 300 George Street, New Haven, Connecticut, 06510, USA.
Chung Lisa M
Zhao Hongyu
References (33)
33 references, click to expand
  1. Understanding mechanisms underlying human gene expression variation with RNA sequencing.
    Nature. 2010 Apr 1;464(7289):768-72 PMID: 20220758
  2. A high-resolution recombination map of the human genome.
    Nat Genet. 2002 Jul;31(3):241-7 PMID: 12053178
  3. Massively parallel signature sequencing (MPSS) as a tool for in-depth quantitative gene expression profiling in all organisms.
    Brief Funct Genomic Proteomic. 2002 Feb;1(1):95-104 PMID: 15251069
  4. CpG doublets, CpG islands and Alu repeats in long human DNA sequences from different isochore families.
    Gene. 1998 Dec 11;224(1-2):123-7 PMID: 9931467
  5. Modeling non-uniformity in short-read rates in RNA-Seq data.
    Genome Biol. 2010;11(5):R50 PMID: 20459815
  6. Novel low abundance and transient RNAs in yeast revealed by tiling microarrays and ultra high-throughput sequencing are not conserved across closely related yeast species.
    PLoS Genet. 2008 Dec;4(12):e1000299 PMID: 19096707
  7. A scaling normalization method for differential expression analysis of RNA-seq data.
    Genome Biol. 2010;11(3):R25 PMID: 20196867
  8. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  9. Complementary DNA sequencing: expressed sequence tags and human genome project.
    Science. 1991 Jun 21;252(5013):1651-6 PMID: 2047873
  10. Improving RNA-Seq expression estimates by correcting for fragment bias.
    Genome Biol. 2011;12(3):R22 PMID: 21410973
  11. Statistical issues in the analysis of Illumina data.
    BMC Bioinformatics. 2008 Feb 06;9:85 PMID: 18254947
  12. The beginning of the end for microarrays?
    Nat Methods. 2008 Jul;5(7):585-7 PMID: 18587314
  13. Transcript length bias in RNA-seq data confounds systems biology.
    Biol Direct. 2009 Apr 16;4:14 PMID: 19371405
  14. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
    Nat Biotechnol. 2010 May;28(5):511-5 PMID: 20436464
  15. Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments.
    BMC Bioinformatics. 2010 Feb 18;11:94 PMID: 20167110
  16. Statistical inferences for isoform expression in RNA-Seq.
    Bioinformatics. 2009 Apr 15;25(8):1026-32 PMID: 19244387
  17. RNA-Seq gene expression estimation with read mapping uncertainty.
    Bioinformatics. 2010 Feb 15;26(4):493-500 PMID: 20022975
  18. Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
    Nucleic Acids Res. 2008 Sep;36(16):e105 PMID: 18660515
  19. A code for transcription initiation in mammalian genomes.
    Genome Res. 2008 Jan;18(1):1-12 PMID: 18032727
  20. Using the transcriptome to annotate the genome.
    Nat Biotechnol. 2002 May;20(5):508-12 PMID: 11981567
  21. Length bias correction for RNA-seq data in gene set analyses.
    Bioinformatics. 2011 Mar 1;27(5):662-9 PMID: 21252076
  22. Relationship between gene expression and GC-content in mammals: statistical significance and biological relevance.
    Hum Mol Genet. 2005 Feb 1;14(3):421-7 PMID: 15590696
  23. RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
    Genome Res. 2008 Sep;18(9):1509-17 PMID: 18550803
  24. Serial analysis of gene expression.
    Science. 1995 Oct 20;270(5235):484-7 PMID: 7570003
  25. Deep sequencing-based expression analysis shows major advances in robustness, resolution and inter-lab portability over five microarray platforms.
    Nucleic Acids Res. 2008 Dec;36(21):e141 PMID: 18927111
  26. The transcriptional landscape of the yeast genome defined by RNA sequencing.
    Science. 2008 Jun 6;320(5881):1344-9 PMID: 18451266
  27. RNA-Seq: a revolutionary tool for transcriptomics.
    Nat Rev Genet. 2009 Jan;10(1):57-63 PMID: 19015660
  28. Evaluation of DNA microarray results with quantitative gene expression platforms.
    Nat Biotechnol. 2006 Sep;24(9):1115-22 PMID: 16964225
  29. TopHat: discovering splice junctions with RNA-Seq.
    Bioinformatics. 2009 May 1;25(9):1105-11 PMID: 19289445
  30. FRT-seq: amplification-free, strand-specific transcriptome sequencing.
    Nat Methods. 2010 Feb;7(2):130-2 PMID: 20081834
  31. Biases in Illumina transcriptome sequencing caused by random hexamer priming.
    Nucleic Acids Res. 2010 Jul;38(12):e131 PMID: 20395217
  32. SAMStat: monitoring biases in next generation sequencing data.
    Bioinformatics. 2011 Jan 1;27(1):130-1 PMID: 21088025
  33. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2011-07-19
Epub
2011-00-19
Pages
290
Language
English
Region
England
NLM ID
100965194
PMCID
PMC3149584
Subset
IM
Grants
NCRR NIH HHS · RR19895 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com