Abstract
Gene expression measurements have successfully been used for building prognostic signatures, i.e for identifying a short list of important genes that can predict patient outcome. Mostly microarray measurements have been considered, and there is little advice available for building multivariable risk prediction models from RNA-Seq data. We specifically consider penalized regression techniques, such as the lasso and componentwise boosting, which can simultaneously consider all measurements and provide both, multivariable regression models for prediction and automated variable selection. However, they might be affected by the typical skewness, mean-variance-dependency or extreme values of RNA-Seq covariates and therefore could benefit from transformations of the latter. In an analytical part, we highlight preferential selection of covariates with large variances, which is problematic due to the mean-variance dependency of RNA-Seq data. In a simulation study, we compare different transformations of RNA-Seq data for potentially improving detection of important genes. Specifically, we consider standardization, the log transformation, a variance-stabilizing transformation, the Box-Cox transformation, and rank-based transformations. In addition, the prediction performance for real data from patients with kidney cancer and acute myeloid leukemia is considered. We show that signature size, identification performance, and prediction performance critically depend on the choice of a suitable transformation. Rank-based transformations perform well in all scenarios and can even outperform complex variance-stabilizing approaches. Generally, the results illustrate that the distribution and potential transformations of RNA-Seq data need to be considered as a critical step when building risk prediction models by penalized regression techniques.
MeSH Terms
Algorithms
Carcinoma, Renal Cell/genetics,metabolism,mortality,pathology
Female
Gene Expression
Humans
Kidney Neoplasms/genetics,metabolism,mortality,pathology
Leukemia, Myeloid, Acute/genetics,metabolism,mortality,pathology
Male
Multivariate Analysis
Neoplasm Proteins/genetics,metabolism
Probability
Prognosis
RNA/genetics,metabolism
Risk
Sequence Analysis, RNA
Survival Analysis
Transcriptome
Chemicals
Neoplasm Proteins
RNA
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Zwiener Isabella
Center for Thrombosis and Hemostasis (CTH), University Medical Center Mainz, Mainz, Germany ; Institute of Medical Biostatistics, Epidemiology and Informatics (IMBEI), University Medical Center Mainz, Mainz, Germany.
Frisch Barbara
Institute of Medical Biostatistics, Epidemiology and Informatics (IMBEI), University Medical Center Mainz, Mainz, Germany.
Binder Harald
Institute of Medical Biostatistics, Epidemiology and Informatics (IMBEI), University Medical Center Mainz, Mainz, Germany.
References (30)
30 references, click to expand
-
Differential expression in RNA-seq: a matter of depth.
Genome Res. 2011 Dec;21(12):2213-23
PMID: 21903743
-
Efron-type measures of prediction error for survival analysis.
Biometrics. 2007 Dec;63(4):1283-7
PMID: 17651459
-
Generalized additive modeling with implicit variable selection by likelihood-based boosting.
Biometrics. 2006 Dec;62(4):961-71
PMID: 17156269
-
Transcript length bias in RNA-seq data confounds systems biology.
Biol Direct. 2009 Apr 16;4:14
PMID: 19371405
-
Rank-based inverse normal transformations are increasingly used, but are they merited?
Behav Genet. 2009 Sep;39(5):580-95
PMID: 19526352
-
Predicting survival from microarray data--a comparative study.
Bioinformatics. 2007 Aug 15;23(16):2080-7
PMID: 17553857
-
A scaling normalization method for differential expression analysis of RNA-seq data.
Genome Biol. 2010;11(3):R25
PMID: 20196867
-
A new shrinkage estimator for dispersion improves differential expression detection in RNA-seq data.
Biostatistics. 2013 Apr;14(2):232-43
PMID: 23001152
-
A comparison of methods for differential expression analysis of RNA-seq data.
BMC Bioinformatics. 2013 Mar 09;14:91
PMID: 23497356
-
An overview of techniques for linking high-dimensional molecular data to time-to-event endpoints by risk prediction models.
Biom J. 2011 Mar;53(2):170-89
PMID: 21328602
-
Development of strategies for SNP detection in RNA-seq data: application to lymphoblastoid cell lines and evaluation using 1000 Genomes data.
PLoS One. 2013;8(3):e58815
PMID: 23555596
-
The lasso method for variable selection in the Cox model.
Stat Med. 1997 Feb 28;16(4):385-95
PMID: 9044528
-
Comparative RNA-Seq and microarray analysis of gene expression changes in B-cell lymphomas of Canis familiaris.
PLoS One. 2013 Apr 04;8(4):e61088
PMID: 23593398
-
Allowing for mandatory covariates in boosting estimation of sparse high-dimensional survival models.
BMC Bioinformatics. 2008 Jan 10;9:14
PMID: 18186927
-
An FLT3 gene-expression signature predicts clinical outcome in normal karyotype AML.
Blood. 2008 May 1;111(9):4490-5
PMID: 18309032
-
Bayesian analysis of RNA sequencing data by estimating multiple shrinkage priors.
Biostatistics. 2013 Jan;14(1):113-28
PMID: 22988280
-
Cross-validation in survival analysis.
Stat Med. 1993 Dec 30;12(24):2305-14
PMID: 8134734
-
RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
Genome Res. 2008 Sep;18(9):1509-17
PMID: 18550803
-
Differential expression analysis for sequence count data.
Genome Biol. 2010;11(10):R106
PMID: 20979621
-
The transcriptional landscape of the yeast genome defined by RNA sequencing.
Science. 2008 Jun 6;320(5881):1344-9
PMID: 18451266
-
RNA-Seq: a revolutionary tool for transcriptomics.
Nat Rev Genet. 2009 Jan;10(1):57-63
PMID: 19015660
-
L1 penalized estimation in the Cox proportional hazards model.
Biom J. 2010 Feb;52(1):70-84
PMID: 19937997
-
Use of pretransformation to cope with extreme values in important candidate features.
Biom J. 2011 Jul;53(4):673-88
PMID: 21626533
-
Cross-validated Cox regression on microarray gene expression data.
Stat Med. 2006 Sep 30;25(18):3201-16
PMID: 16143967
-
Mapping and quantifying mammalian transcriptomes by RNA-Seq.
Nat Methods. 2008 Jul;5(7):621-8
PMID: 18516045
-
baySeq: empirical Bayesian methods for identifying differential expression in sequence count data.
BMC Bioinformatics. 2010 Aug 10;11:422
PMID: 20698981
-
Tailoring sparse multivariable regression techniques for prognostic single-nucleotide polymorphism signatures.
Stat Med. 2013 May 10;32(10):1778-91
PMID: 22786659
-
Normalization, testing, and false discovery rate estimation for RNA-sequencing data.
Biostatistics. 2012 Jul;13(3):523-38
PMID: 22003245
-
S-MART, a software toolbox to aid RNA-Seq data analysis.
PLoS One. 2011;6(10):e25988
PMID: 21998740
-
Finding consistent patterns: a nonparametric approach for identifying differential expression in RNA-Seq data.
Stat Methods Med Res. 2013 Oct;22(5):519-36
PMID: 22127579