Abstract
We consider the statistical analysis of population structure using genetic data. We show how the two most widely used approaches to modeling population structure, admixture-based models and principal components analysis (PCA), can be viewed within a single unifying framework of matrix factorization. Specifically, they can both be interpreted as approximating an observed genotype matrix by a product of two lower-rank matrices, but with different constraints or prior distributions on these lower-rank matrices. This opens the door to a large range of possible approaches to analyzing population structure, by considering other constraints or priors. In this paper, we introduce one such novel approach, based on sparse factor analysis (SFA). We investigate the effects of the different types of constraint in several real and simulated data sets. We find that SFA produces similar results to admixture-based models when the samples are descended from a few well-differentiated ancestral populations and can recapitulate the results of PCA when the population structure is more "continuous," as in isolation-by-distance models.
MeSH Terms
Africa
Computer Simulation
Databases, Genetic
Europe
Factor Analysis, Statistical
Genetics, Population/methods
Genotype
Humans
India
Models, Genetic
Population Dynamics
Principal Component Analysis
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Engelhardt Barbara E
Computer Science Department, University of Chicago, Chicago, Illinois, USA. engelhardt@uchicago.edu
Stephens Matthew
Conflict of Interest
The authors have declared that no competing interests exist.
References (27)
27 references, click to expand
-
Use of unlinked genetic markers to detect population stratification in association studies.
Am J Hum Genet. 1999 Jul;65(1):220-8
PMID: 10364535
-
Using DNA to track the origin of the largest ivory seizure since the 1989 trade ban.
Proc Natl Acad Sci U S A. 2007 Mar 6;104(10):4228-33
PMID: 17360505
-
Factor analysis for gene regulatory networks and transcription factor activity profiles.
BMC Bioinformatics. 2007 Feb 23;8:61
PMID: 17319944
-
Reconstructing genetic ancestry blocks in admixed individuals.
Am J Hum Genet. 2006 Jul;79(1):1-12
PMID: 16773560
-
High-Dimensional Sparse Factor Modeling: Applications in Gene Expression Genomics.
J Am Stat Assoc. 2008 Dec 1;103(484):1438-1456
PMID: 21218139
-
A flexible and accurate genotype imputation method for the next generation of genome-wide association studies.
PLoS Genet. 2009 Jun;5(6):e1000529
PMID: 19543373
-
Case-control studies of association in structured or admixed populations.
Theor Popul Biol. 2001 Nov;60(3):227-37
PMID: 11855957
-
Genes mirror geography within Europe.
Nature. 2008 Nov 6;456(7218):98-101
PMID: 18758442
-
Interpreting principal component analyses of spatial population genetic variation.
Nat Genet. 2008 May;40(5):646-9
PMID: 18425127
-
Correlation between genetic and geographic structure in Europe.
Curr Biol. 2008 Aug 26;18(16):1241-8
PMID: 18691889
-
A worldwide survey of haplotype variation and linkage disequilibrium in the human genome.
Nat Genet. 2006 Nov;38(11):1251-60
PMID: 17057719
-
Principal components analysis corrects for stratification in genome-wide association studies.
Nat Genet. 2006 Aug;38(8):904-9
PMID: 16862161
-
Estimation of individual admixture: analytical and study design considerations.
Genet Epidemiol. 2005 May;28(4):289-301
PMID: 15712363
-
Association mapping, using a mixture model for complex traits.
Genet Epidemiol. 2002 Aug;23(2):181-96
PMID: 12214310
-
Reconstructing Indian population history.
Nature. 2009 Sep 24;461(7263):489-94
PMID: 19779445
-
Learning the parts of objects by non-negative matrix factorization.
Nature. 1999 Oct 21;401(6755):788-91
PMID: 10548103
-
The Population Reference Sample, POPRES: a resource for population, disease, and pharmacological genetics research.
Am J Hum Genet. 2008 Sep;83(3):347-58
PMID: 18760391
-
Genetic structure of the purebred domestic dog.
Science. 2004 May 21;304(5674):1160-4
PMID: 15155949
-
Population structure and eigenanalysis.
PLoS Genet. 2006 Dec;2(12):e190
PMID: 17194218
-
Generating samples under a Wright-Fisher neutral model of genetic variation.
Bioinformatics. 2002 Feb;18(2):337-8
PMID: 11847089
-
Evidence for gradients of human genetic diversity within and among continents.
Genome Res. 2004 Sep;14(9):1679-85
PMID: 15342553
-
Genetic structure of human populations.
Science. 2002 Dec 20;298(5602):2381-5
PMID: 12493913
-
Inference of population structure using multilocus genotype data.
Genetics. 2000 Jun;155(2):945-59
PMID: 10835412
-
Fast model-based estimation of ancestry in unrelated individuals.
Genome Res. 2009 Sep;19(9):1655-64
PMID: 19648217
-
A genealogical interpretation of principal components analysis.
PLoS Genet. 2009 Oct;5(10):e1000686
PMID: 19834557
-
A penalized matrix decomposition, with applications to sparse principal components and canonical correlation analysis.
Biostatistics. 2009 Jul;10(3):515-34
PMID: 19377034
-
Inference of population structure using multilocus genotype data: linked loci and correlated allele frequencies.
Genetics. 2003 Aug;164(4):1567-87
PMID: 12930761