Home LiteratureArticle Details
PMID: 24416353 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

Transforming RNA-Seq data to improve the performance of prognostic gene signatures.

PloS one ·Vol. 9 ·No. 1 ·2014-00-00 ·Pages e85150

Zwiener I, Frisch B, Binder H

Abstract

Gene expression measurements have successfully been used for building prognostic signatures, i.e for identifying a short list of important genes that can predict patient outcome. Mostly microarray measurements have been considered, and there is little advice available for building multivariable risk prediction models from RNA-Seq data. We specifically consider penalized regression techniques, such as the lasso and componentwise boosting, which can simultaneously consider all measurements and provide both, multivariable regression models for prediction and automated variable selection. However, they might be affected by the typical skewness, mean-variance-dependency or extreme values of RNA-Seq covariates and therefore could benefit from transformations of the latter. In an analytical part, we highlight preferential selection of covariates with large variances, which is problematic due to the mean-variance dependency of RNA-Seq data. In a simulation study, we compare different transformations of RNA-Seq data for potentially improving detection of important genes. Specifically, we consider standardization, the log transformation, a variance-stabilizing transformation, the Box-Cox transformation, and rank-based transformations. In addition, the prediction performance for real data from patients with kidney cancer and acute myeloid leukemia is considered. We show that signature size, identification performance, and prediction performance critically depend on the choice of a suitable transformation. Rank-based transformations perform well in all scenarios and can even outperform complex variance-stabilizing approaches. Generally, the results illustrate that the distribution and potential transformations of RNA-Seq data need to be considered as a critical step when building risk prediction models by penalized regression techniques.

MeSH Terms
Algorithms Carcinoma, Renal Cell/genetics,metabolism,mortality,pathology Female Gene Expression Humans Kidney Neoplasms/genetics,metabolism,mortality,pathology Leukemia, Myeloid, Acute/genetics,metabolism,mortality,pathology Male Multivariate Analysis Neoplasm Proteins/genetics,metabolism Probability Prognosis RNA/genetics,metabolism Risk Sequence Analysis, RNA Survival Analysis Transcriptome
Chemicals
Neoplasm Proteins RNA
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Zwiener Isabella
Center for Thrombosis and Hemostasis (CTH), University Medical Center Mainz, Mainz, Germany ; Institute of Medical Biostatistics, Epidemiology and Informatics (IMBEI), University Medical Center Mainz, Mainz, Germany.
Frisch Barbara
Institute of Medical Biostatistics, Epidemiology and Informatics (IMBEI), University Medical Center Mainz, Mainz, Germany.
Binder Harald
Institute of Medical Biostatistics, Epidemiology and Informatics (IMBEI), University Medical Center Mainz, Mainz, Germany.
References (30)
30 references, click to expand
  1. Differential expression in RNA-seq: a matter of depth.
    Genome Res. 2011 Dec;21(12):2213-23 PMID: 21903743
  2. Efron-type measures of prediction error for survival analysis.
    Biometrics. 2007 Dec;63(4):1283-7 PMID: 17651459
  3. Generalized additive modeling with implicit variable selection by likelihood-based boosting.
    Biometrics. 2006 Dec;62(4):961-71 PMID: 17156269
  4. Transcript length bias in RNA-seq data confounds systems biology.
    Biol Direct. 2009 Apr 16;4:14 PMID: 19371405
  5. Rank-based inverse normal transformations are increasingly used, but are they merited?
    Behav Genet. 2009 Sep;39(5):580-95 PMID: 19526352
  6. Predicting survival from microarray data--a comparative study.
    Bioinformatics. 2007 Aug 15;23(16):2080-7 PMID: 17553857
  7. A scaling normalization method for differential expression analysis of RNA-seq data.
    Genome Biol. 2010;11(3):R25 PMID: 20196867
  8. A new shrinkage estimator for dispersion improves differential expression detection in RNA-seq data.
    Biostatistics. 2013 Apr;14(2):232-43 PMID: 23001152
  9. A comparison of methods for differential expression analysis of RNA-seq data.
    BMC Bioinformatics. 2013 Mar 09;14:91 PMID: 23497356
  10. An overview of techniques for linking high-dimensional molecular data to time-to-event endpoints by risk prediction models.
    Biom J. 2011 Mar;53(2):170-89 PMID: 21328602
  11. Development of strategies for SNP detection in RNA-seq data: application to lymphoblastoid cell lines and evaluation using 1000 Genomes data.
    PLoS One. 2013;8(3):e58815 PMID: 23555596
  12. The lasso method for variable selection in the Cox model.
    Stat Med. 1997 Feb 28;16(4):385-95 PMID: 9044528
  13. Comparative RNA-Seq and microarray analysis of gene expression changes in B-cell lymphomas of Canis familiaris.
    PLoS One. 2013 Apr 04;8(4):e61088 PMID: 23593398
  14. Allowing for mandatory covariates in boosting estimation of sparse high-dimensional survival models.
    BMC Bioinformatics. 2008 Jan 10;9:14 PMID: 18186927
  15. An FLT3 gene-expression signature predicts clinical outcome in normal karyotype AML.
    Blood. 2008 May 1;111(9):4490-5 PMID: 18309032
  16. Bayesian analysis of RNA sequencing data by estimating multiple shrinkage priors.
    Biostatistics. 2013 Jan;14(1):113-28 PMID: 22988280
  17. Cross-validation in survival analysis.
    Stat Med. 1993 Dec 30;12(24):2305-14 PMID: 8134734
  18. RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
    Genome Res. 2008 Sep;18(9):1509-17 PMID: 18550803
  19. Differential expression analysis for sequence count data.
    Genome Biol. 2010;11(10):R106 PMID: 20979621
  20. The transcriptional landscape of the yeast genome defined by RNA sequencing.
    Science. 2008 Jun 6;320(5881):1344-9 PMID: 18451266
  21. RNA-Seq: a revolutionary tool for transcriptomics.
    Nat Rev Genet. 2009 Jan;10(1):57-63 PMID: 19015660
  22. L1 penalized estimation in the Cox proportional hazards model.
    Biom J. 2010 Feb;52(1):70-84 PMID: 19937997
  23. Use of pretransformation to cope with extreme values in important candidate features.
    Biom J. 2011 Jul;53(4):673-88 PMID: 21626533
  24. Cross-validated Cox regression on microarray gene expression data.
    Stat Med. 2006 Sep 30;25(18):3201-16 PMID: 16143967
  25. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  26. baySeq: empirical Bayesian methods for identifying differential expression in sequence count data.
    BMC Bioinformatics. 2010 Aug 10;11:422 PMID: 20698981
  27. Tailoring sparse multivariable regression techniques for prognostic single-nucleotide polymorphism signatures.
    Stat Med. 2013 May 10;32(10):1778-91 PMID: 22786659
  28. Normalization, testing, and false discovery rate estimation for RNA-sequencing data.
    Biostatistics. 2012 Jul;13(3):523-38 PMID: 22003245
  29. S-MART, a software toolbox to aid RNA-Seq data analysis.
    PLoS One. 2011;6(10):e25988 PMID: 21998740
  30. Finding consistent patterns: a nonparametric approach for identifying differential expression in RNA-Seq data.
    Stat Methods Med Res. 2013 Oct;22(5):519-36 PMID: 22127579
Article Info
Journal
PloS one
Abbr.
PLoS One
ISSN
1932-6203
Published
2014-00-00
Epub
2014-00-08
Pages
e85150
Language
English
Region
United States
NLM ID
101285081
PMCID
PMC3885686
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com