Home LiteratureArticle Details
PMID: 25336500 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

Statistical significance of variables driving systematic variation in high-dimensional data.

Bioinformatics (Oxford, England) ·Vol. 31 ·No. 4 ·2015-02-15 ·Pages 545-54

Chung NC, Storey JD

Abstract

There are a number of well-established methods such as principal component analysis (PCA) for automatically capturing systematic variation due to latent variables in large-scale genomic data. PCA and related methods may directly provide a quantitative characterization of a complex biological variable that is otherwise difficult to precisely define or model. An unsolved problem in this context is how to systematically identify the genomic variables that are drivers of systematic variation captured by PCA. Principal components (PCs) (and other estimates of systematic variation) are directly constructed from the genomic variables themselves, making measures of statistical significance artificially inflated when using conventional methods due to over-fitting. We introduce a new approach called the jackstraw that allows one to accurately identify genomic variables that are statistically significantly associated with any subset or linear combination of PCs. The proposed method can greatly simplify complex significance testing problems encountered in genomics and can be used to identify the genomic variables significantly associated with latent variables. Using simulation, we demonstrate that our method attains accurate measures of statistical significance over a range of relevant scenarios. We consider yeast cell-cycle gene expression data, and show that the proposed method can be used to straightforwardly identify genes that are cell-cycle regulated with an accurate measure of statistical significance. We also analyze gene expression data from post-trauma patients, allowing the gene expression data to provide a molecularly driven phenotype. Using our method, we find a greater enrichment for inflammatory-related gene sets compared to the original analysis that uses a clinically defined, although likely imprecise, phenotype. The proposed method provides a useful bridge between large-scale quantifications of systematic variation and gene-level significance analyses. An R software package, called jackstraw, is available in CRAN. jstorey@princeton.edu.

MeSH Terms
Algorithms Computer Simulation Data Interpretation, Statistical Gene Expression Profiling Genes, cdc Genetic Variation Genomics/methods Humans Inflammation/genetics Microarray Analysis Models, Statistical Phenotype Principal Component Analysis Saccharomyces cerevisiae/genetics Saccharomyces cerevisiae Proteins/genetics Software Stress Disorders, Post-Traumatic/genetics
Chemicals
Saccharomyces cerevisiae Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Chung Neo Christopher
Lewis-Sigler Institute for Integrative Genomics and Department of Molecular Biology, Princeton University, Princeton, NJ 08544, USA.
Storey John D
Lewis-Sigler Institute for Integrative Genomics and Department of Molecular Biology, Princeton University, Princeton, NJ 08544, USA Lewis-Sigler Institute for Integrative Genomics and Department of Molecular Biology, Princeton University, Princeton, NJ 08544, USA.
References (25)
25 references, click to expand
  1. Dissecting inflammatory complications in critically injured patients by within-patient gene expression changes: a longitudinal clinical genomics study.
    PLoS Med. 2011 Sep;8(9):e1001093 PMID: 21931541
  2. Distinct types of diffuse large B-cell lymphoma identified by gene expression profiling.
    Nature. 2000 Feb 3;403(6769):503-11 PMID: 10676951
  3. Remarks on Parallel Analysis.
    Multivariate Behav Res. 1992 Oct 1;27(4):509-40 PMID: 26811132
  4. High-resolution timing of cell cycle-regulated gene expression.
    Proc Natl Acad Sci U S A. 2007 Oct 23;104(43):16892-7 PMID: 17827275
  5. Systematic identification of yeast cell cycle transcription factors using multiple data sources.
    BMC Bioinformatics. 2008 Dec 05;9:522 PMID: 19061501
  6. Application of genome-wide expression analysis to human health and disease.
    Proc Natl Acad Sci U S A. 2005 Mar 29;102(13):4801-6 PMID: 15781863
  7. Corrected confidence bands for functional data using principal components.
    Biometrics. 2013 Mar;69(1):41-51 PMID: 23003003
  8. Assembly of inflammation-related genes for pathway-focused genetic analysis.
    PLoS One. 2007 Oct 17;2(10):e1035 PMID: 17940599
  9. Capturing heterogeneity in gene expression studies by surrogate variable analysis.
    PLoS Genet. 2007 Sep;3(9):1724-35 PMID: 17907809
  10. Asymptotic conditional singular value decomposition for high-dimensional genomic data.
    Biometrics. 2011 Jun;67(2):344-52 PMID: 20560929
  11. Multiple organ dysfunction score: a reliable descriptor of a complex clinical outcome.
    Crit Care Med. 1995 Oct;23(10):1638-52 PMID: 7587228
  12. Logic of the yeast metabolic cycle: temporal compartmentalization of cellular processes.
    Science. 2005 Nov 18;310(5751):1152-8 PMID: 16254148
  13. Singular value decomposition for genome-wide expression data processing and modeling.
    Proc Natl Acad Sci U S A. 2000 Aug 29;97(18):10101-6 PMID: 10963673
  14. The Forkhead transcription factor Hcm1 regulates chromosome segregation genes and fills the S-phase gap in the transcriptional circuitry of the cell cycle.
    Genes Dev. 2006 Aug 15;20(16):2266-78 PMID: 16912276
  15. Use of a cDNA microarray to analyse gene expression patterns in human cancer.
    Nat Genet. 1996 Dec;14(4):457-60 PMID: 8944026
  16. Analysis of population structure: a unifying framework and novel methods based on sparse factor analysis.
    PLoS Genet. 2010 Sep 16;6(9):e1001117 PMID: 20862358
  17. Fundamental patterns underlying gene expression profiles: simplicity from complexity.
    Proc Natl Acad Sci U S A. 2000 Jul 18;97(15):8409-14 PMID: 10890920
  18. Comprehensive identification of cell cycle-regulated genes of the yeast Saccharomyces cerevisiae by microarray hybridization.
    Mol Biol Cell. 1998 Dec;9(12):3273-97 PMID: 9843569
  19. A genome-wide transcriptional analysis of the mitotic cell cycle.
    Mol Cell. 1998 Jul;2(1):65-73 PMID: 9702192
  20. Principal components analysis corrects for stratification in genome-wide association studies.
    Nat Genet. 2006 Aug;38(8):904-9 PMID: 16862161
  21. A penalized matrix decomposition, with applications to sparse principal components and canonical correlation analysis.
    Biostatistics. 2009 Jul;10(3):515-34 PMID: 19377034
  22. Estimating confidence intervals for principal component loadings: a comparison between the bootstrap and asymptotic results.
    Br J Math Stat Psychol. 2007 Nov;60(Pt 2):295-314 PMID: 17971271
  23. Principal components analysis to summarize microarray experiments: application to sporulation time series.
    Pac Symp Biocomput. 2000;:455-66 PMID: 10902193
  24. Association mapping, using a mixture model for complex traits.
    Genet Epidemiol. 2002 Aug;23(2):181-96 PMID: 12214310
  25. A general framework for multiple testing dependence.
    Proc Natl Acad Sci U S A. 2008 Dec 2;105(48):18718-23 PMID: 19033188
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2015-02-15
Epub
2014-00-21
Pages
545-54
Language
English
Region
England
NLM ID
9808944
PMCID
PMC4325543
Subset
IM
Grants
NHGRI NIH HHS · R01 HG002913 · United States
NHGRI NIH HHS · R01 HG006448 · United States
NHGRI NIH HHS · HG002913 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com