Home LiteratureArticle Details
PMID: 19377034 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

A penalized matrix decomposition, with applications to sparse principal components and canonical correlation analysis.

Biostatistics (Oxford, England) ·Vol. 10 ·No. 3 ·2009-07-00 ·Pages 515-34

Witten DM, Tibshirani R, Hastie T

Abstract

We present a penalized matrix decomposition (PMD), a new framework for computing a rank-K approximation for a matrix. We approximate the matrix X as circumflexX = sigma(k=1)(K) d(k)u(k)v(k)(T), where d(k), u(k), and v(k) minimize the squared Frobenius norm of X - circumflexX, subject to penalties on u(k) and v(k). This results in a regularized version of the singular value decomposition. Of particular interest is the use of L(1)-penalties on u(k) and v(k), which yields a decomposition of X using sparse vectors. We show that when the PMD is applied using an L(1)-penalty on v(k) but not on u(k), a method for sparse principal components results. In fact, this yields an efficient algorithm for the "SCoTLASS" proposal (Jolliffe and others 2003) for obtaining sparse principal components. This method is demonstrated on a publicly available gene expression data set. We also establish connections between the SCoTLASS method for sparse principal component analysis and the method of Zou and others (2006). In addition, we show that when the PMD is applied to a cross-products matrix, it results in a method for penalized canonical correlation analysis (CCA). We apply this penalized CCA method to simulated data and to a genomic data set consisting of gene expression and DNA copy number measurements on the same set of samples.

MeSH Terms
Algorithms Biometry/methods Breast Neoplasms/genetics Chromosomes, Human, Pair 1/genetics DNA, Neoplasm/genetics Data Interpretation, Statistical Female Gene Dosage Genomics/statistics & numerical data Humans Models, Statistical Principal Component Analysis/methods
Chemicals
DNA, Neoplasm
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Witten Daniela M
Department of Statistics, Stanford University, Stanford, CA 94305, USA. dwitten@stanford.edu
Tibshirani Robert
Hastie Trevor
References (12)
12 references, click to expand
  1. Missing value estimation methods for DNA microarrays.
    Bioinformatics. 2001 Jun;17(6):520-5 PMID: 11395428
  2. Impact of DNA amplification on gene expression patterns in breast cancer.
    Cancer Res. 2002 Nov 1;62(21):6240-5 PMID: 12414653
  3. Learning the parts of objects by non-negative matrix factorization.
    Nature. 1999 Oct 21;401(6755):788-91 PMID: 10548103
  4. Genetic analysis of genome-wide variation in human gene expression.
    Nature. 2004 Aug 12;430(7001):743-7 PMID: 15269782
  5. Sparse canonical correlation analysis with application to genomic data integration.
    Stat Appl Genet Mol Biol. 2009;8:Article 1 PMID: 19222376
  6. Relative impact of nucleotide and copy number variation on gene expression phenotypes.
    Science. 2007 Feb 9;315(5813):848-53 PMID: 17289997
  7. Microarray analysis reveals a major direct role of DNA copy number alteration in the transcriptional program of human breast tumors.
    Proc Natl Acad Sci U S A. 2002 Oct 1;99(20):12963-8 PMID: 12297621
  8. Genomic and transcriptional aberrations linked to breast cancer pathophysiologies.
    Cancer Cell. 2006 Dec;10(6):529-41 PMID: 17157792
  9. Genome-wide sparse canonical correlation of gene expression with genotypes.
    BMC Proc. 2007;1 Suppl 1:S119 PMID: 18466460
  10. Genome-wide associations of gene expression variation in humans.
    PLoS Genet. 2005 Dec;1(6):e78 PMID: 16362079
  11. Quantifying the association between gene expressions and DNA-markers by penalized canonical correlation analysis.
    Stat Appl Genet Mol Biol. 2008;7(1):Article3 PMID: 18241193
  12. Spatial smoothing and hot spot detection for CGH data using the fused lasso.
    Biostatistics. 2008 Jan;9(1):18-29 PMID: 17513312
Article Info
Journal
Biostatistics (Oxford, England)
Abbr.
Biostatistics
ISSN
1468-4357
Published
2009-07-00
Epub
2009-00-17
Pages
515-34
Language
English
Region
England
NLM ID
100897327
PMCID
PMC2697346
Subset
IM
Grants
NCI NIH HHS · 2R01 CA 72028-07 · United States
NHLBI NIH HHS · N01-HV-28183 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com