Abstract
For large-scale genotyping studies, it is common for most subjects to have some missing genetic markers, even if the missing rate per marker is low. This compromises association analyses, with varying numbers of subjects contributing to analyses when performing single-marker or multi-marker analyses. In this paper, we consider eight methods to infer missing genotypes, including two haplotype reconstruction methods (local expectation maximization-EM, and fastPHASE), two k-nearest neighbor methods (original k-nearest neighbor, KNN, and a weighted k-nearest neighbor, wtKNN), three linear regression methods (backward variable selection, LM.back, least angle regression, LM.lars, and singular value decomposition, LM.svd), and a regression tree, Rtree. We evaluate the accuracy of them using single nucleotide polymorphism (SNP) data from the HapMap project, under a variety of conditions and parameters. We find that fastPHASE has the lowest error rates across different analysis panels and marker densities. LM.lars gives slightly less accurate estimate of missing genotypes than fastPHASE, but has better performance than the other methods.
MeSH Terms
Genetic Markers
Genetics, Population/statistics & numerical data
Genotype
Haplotypes
Humans
Linear Models
Linkage Disequilibrium
Models, Genetic
Models, Statistical
Polymorphism, Single Nucleotide
Statistics, Nonparametric
Chemicals
Genetic Markers
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Yu Zhaoxia
Department of Statistics, University of California, Irvine, CA 92697, USA. yu.zhaoxia@ics.uci.edu
Schaid Daniel J
References (29)
29 references, click to expand
-
Impact of missing genotype data on Monte-Carlo simulation based haplotype analysis.
Hum Hered. 2005;59(4):185-9
PMID: 16015028
-
Meiotic recombination hotspots.
Annu Rev Genet. 1995;29:423-44
PMID: 8825482
-
Principal components analysis corrects for stratification in genome-wide association studies.
Nat Genet. 2006 Aug;38(8):904-9
PMID: 16862161
-
Imputation methods to improve inference in SNP association studies.
Genet Epidemiol. 2006 Dec;30(8):690-702
PMID: 16986162
-
A comparison of bayesian methods for haplotype reconstruction from population genotype data.
Am J Hum Genet. 2003 Nov;73(5):1162-9
PMID: 14574645
-
A haplotype map of the human genome.
Nature. 2005 Oct 27;437(7063):1299-320
PMID: 16255080
-
Partition-ligation-expectation-maximization algorithm for haplotype inference with single-nucleotide polymorphisms.
Am J Hum Genet. 2002 Nov;71(5):1242-7
PMID: 12452179
-
Accuracy of haplotype frequency estimation for biallelic loci, via the expectation-maximization algorithm for unphased diploid genotype data.
Am J Hum Genet. 2000 Oct;67(4):947-59
PMID: 10954684
-
A comparison of phasing algorithms for trios and unrelated individuals.
Am J Hum Genet. 2006 Mar;78(3):437-50
PMID: 16465620
-
A new multipoint method for genome-wide association studies by imputation of genotypes.
Nat Genet. 2007 Jul;39(7):906-13
PMID: 17572673
-
Imputation-based analysis of association studies: candidate regions and quantitative traits.
PLoS Genet. 2007 Jul;3(7):e114
PMID: 17676998
-
The Interaction of Selection and Linkage. I. General Considerations; Heterotic Models.
Genetics. 1964 Jan;49(1):49-67
PMID: 17248194
-
A fast and flexible statistical model for large-scale population genotype data: applications to inferring missing genotypes and haplotypic phase.
Am J Hum Genet. 2006 Apr;78(4):629-44
PMID: 16532393
-
An E-M algorithm and testing strategy for multiple-locus haplotypes.
Am J Hum Genet. 1995 Mar;56(3):799-810
PMID: 7887436
-
Estimation and tests of haplotype-environment interaction when linkage phase is ambiguous.
Hum Hered. 2003;55(1):56-65
PMID: 12890927
-
Bayesian haplotype inference for multiple linked single-nucleotide polymorphisms.
Am J Hum Genet. 2002 Jan;70(1):157-69
PMID: 11741196
-
Fine genetic mapping using haplotype analysis and the missing data problem.
Ann Hum Genet. 1998 Jan;62(Pt 1):55-60
PMID: 9659978
-
Maximum-likelihood estimation of molecular haplotype frequencies in a diploid population.
Mol Biol Evol. 1995 Sep;12(5):921-7
PMID: 7476138
-
Testing untyped alleles (TUNA)-applications to genome-wide association studies.
Genet Epidemiol. 2006 Dec;30(8):718-27
PMID: 16986160
-
HAPLO: a program using the EM algorithm to estimate the frequencies of multi-site haplotypes.
J Hered. 1995 Sep-Oct;86(5):409-11
PMID: 7560877
-
Haplotype analysis in the presence of informatively missing genotype data.
Genet Epidemiol. 2006 May;30(4):290-300
PMID: 16528706
-
Singular value decomposition for genome-wide expression data processing and modeling.
Proc Natl Acad Sci U S A. 2000 Aug 29;97(18):10101-6
PMID: 10963673
-
Score tests for association between traits and haplotypes when linkage phase is ambiguous.
Am J Hum Genet. 2002 Feb;70(2):425-34
PMID: 11791212
-
Accounting for decay of linkage disequilibrium in haplotype inference and missing-data imputation.
Am J Hum Genet. 2005 Mar;76(3):449-62
PMID: 15700229
-
Bayesian mapping of genotype x expression interactions in quantitative and qualitative traits.
Heredity (Edinb). 2006 Jul;97(1):4-18
PMID: 16670709
-
A new statistical method for haplotype reconstruction from population data.
Am J Hum Genet. 2001 Apr;68(4):978-89
PMID: 11254454
-
Missing value estimation methods for DNA microarrays.
Bioinformatics. 2001 Jun;17(6):520-5
PMID: 11395428
-
Haplotype and missing data inference in nuclear families.
Genome Res. 2004 Aug;14(8):1624-32
PMID: 15256514
-
Multiple imputation of missing genotype data for unrelated individuals.
Ann Hum Genet. 2006 May;70(Pt 3):372-81
PMID: 16674559