Abstract
After mapping, RNA-Seq data can be summarized by a sequence of read counts commonly modeled as Poisson variables with constant rates along each transcript, which actually fit data poorly. We suggest using variable rates for different positions, and propose two models to predict these rates based on local sequences. These models explain more than 50% of the variations and can lead to improved estimates of gene and isoform expressions for both Illumina and Applied Biosystems data.
MeSH Terms
Animals
Apolipoproteins E/genetics
Base Sequence
Databases, Nucleic Acid
Embryo, Mammalian/metabolism
Exons/genetics
Gene Expression Profiling
Gene Expression Regulation
Humans
Linear Models
Mice
Models, Genetic
Poisson Distribution
Protein Isoforms/genetics,metabolism
RNA/genetics
Sequence Analysis, RNA/methods
Statistics, Nonparametric
Chemicals
Apolipoproteins E
Protein Isoforms
RNA
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Li Jun
Department of Statistics, Stanford University, Sequoia Hall, 390 Serra Mall, Stanford, CA 94305, USA. junli07@stanford.edu
Jiang Hui
Wong Wing Hung
References (28)
28 references, click to expand
-
Efficient mapping of Applied Biosystems SOLiD sequence data to a reference genome for functional genomic applications.
Bioinformatics. 2008 Dec 1;24(23):2776-7
PMID: 18842598
-
Profiling the HeLa S3 transcriptome using randomly primed cDNA and massively parallel short-read sequencing.
Biotechniques. 2008 Jul;45(1):81-94
PMID: 18611170
-
Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
Nucleic Acids Res. 2008 Sep;36(16):e105
PMID: 18660515
-
Toward a universal microarray: prediction of gene expression through nearest-neighbor probe sequence identification.
Nucleic Acids Res. 2007;35(15):e99
PMID: 17686789
-
Summaries of Affymetrix GeneChip probe level data.
Nucleic Acids Res. 2003 Feb 15;31(4):e15
PMID: 12582260
-
Highly integrated single-base resolution maps of the epigenome in Arabidopsis.
Cell. 2008 May 2;133(3):523-36
PMID: 18423832
-
Solving the riddle of the bright mismatches: labeling and effective binding in oligonucleotide arrays.
Phys Rev E Stat Nonlin Soft Matter Phys. 2003 Jul;68(1 Pt 1):011906
PMID: 12935175
-
Dynamic repertoire of a eukaryotic transcriptome surveyed at single-nucleotide resolution.
Nature. 2008 Jun 26;453(7199):1239-43
PMID: 18488015
-
Revealing global regulatory features of mammalian alternative splicing using a quantitative microarray platform.
Mol Cell. 2004 Dec 22;16(6):929-41
PMID: 15610736
-
Alternative isoform regulation in human tissue transcriptomes.
Nature. 2008 Nov 27;456(7221):470-6
PMID: 18978772
-
RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
Genome Res. 2008 Sep;18(9):1509-17
PMID: 18550803
-
Probe signal correction for differential methylation hybridization experiments.
BMC Bioinformatics. 2008 Oct 23;9:453
PMID: 18947421
-
Biases in Illumina transcriptome sequencing caused by random hexamer priming.
Nucleic Acids Res. 2010 Jul;38(12):e131
PMID: 20395217
-
Mapping and quantifying mammalian transcriptomes by RNA-Seq.
Nat Methods. 2008 Jul;5(7):621-8
PMID: 18516045
-
Initial sequencing and analysis of the human genome.
Nature. 2001 Feb 15;409(6822):860-921
PMID: 11237011
-
Model-based analysis of tiling-arrays for ChIP-chip.
Proc Natl Acad Sci U S A. 2006 Aug 15;103(33):12457-62
PMID: 16895995
-
Cross-hybridization modeling on Affymetrix exon arrays.
Bioinformatics. 2008 Dec 15;24(24):2887-93
PMID: 18984598
-
Initial sequencing and comparative analysis of the mouse genome.
Nature. 2002 Dec 5;420(6915):520-62
PMID: 12466850
-
Hybridization interactions between probesets in short oligo microarrays lead to spurious correlations.
BMC Bioinformatics. 2006 Jun 02;7:276
PMID: 16749918
-
Deep surveying of alternative splicing complexity in the human transcriptome by high-throughput sequencing.
Nat Genet. 2008 Dec;40(12):1413-5
PMID: 18978789
-
An integrated software system for analyzing ChIP-chip and ChIP-seq data.
Nat Biotechnol. 2008 Nov;26(11):1293-300
PMID: 18978777
-
Statistical inferences for isoform expression in RNA-Seq.
Bioinformatics. 2009 Apr 15;25(8):1026-32
PMID: 19244387
-
Stem cell transcriptome profiling via massive-scale mRNA sequencing.
Nat Methods. 2008 Jul;5(7):613-9
PMID: 18516046
-
SeqMap: mapping massive amount of oligonucleotides to the genome.
Bioinformatics. 2008 Oct 15;24(20):2395-6
PMID: 18697769
-
The transcriptional landscape of the yeast genome defined by RNA sequencing.
Science. 2008 Jun 6;320(5881):1344-9
PMID: 18451266
-
RNA-Seq: a revolutionary tool for transcriptomics.
Nat Rev Genet. 2009 Jan;10(1):57-63
PMID: 19015660
-
The new paradigm of flow cell sequencing.
Genome Res. 2008 Jun;18(6):839-46
PMID: 18519653
-
Model-based analysis of two-color arrays (MA2C).
Genome Biol. 2007;8(8):R178
PMID: 17727723