Abstract
In a previous paper we have shown that, when DNA samples for cases and controls are prepared in different laboratories prior to high-throughput genotyping, scoring inaccuracies can lead to differential misclassification and, consequently, to increased false-positive rates. Different DNA sourcing is often unavoidable in large-scale disease association studies of multiple case and control sets. Here, we describe methodological improvements to minimise such biases. These fall into two categories: improvements to the basic clustering methods for identifying genotypes from fluorescence intensities, and use of "fuzzy" calls in association tests in order to make appropriate allowance for call uncertainty. We find that the main improvement is a modification of the calling algorithm that links the clustering of cases and controls while allowing for different DNA sourcing. We also find that, in the presence of different DNA sourcing, biases associated with missing data can increase the false-positive rate. Therefore, we propose the use of "fuzzy" calls to deal with uncertain genotypes that would otherwise be labeled as missing.
MeSH Terms
Algorithms
Bias
Case-Control Studies
Computer Simulation
Databases, Genetic
Epidemiologic Methods
Female
Gene Frequency
Genetic Predisposition to Disease/genetics
Genotype
Humans
Male
Polymorphism, Single Nucleotide/genetics
United Kingdom
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Plagnol Vincent
Juvenile Diabetes Research Foundation/Wellcome Trust Diabetes and Inflammation Laboratory, Department of Medical Genetics, Cambridge Institute for Medical Research, University of Cambridge, Cambridge, United Kingdom. vincent.plagnol@cimr.cam.ac.uk
Cooper Jason D
Todd John A
Clayton David G
Conflict of Interest
Competing interests. The authors have declared that no competing interests exist.
References (9)
9 references, click to expand
-
Optimal genotype determination in highly multiplexed SNP data.
Eur J Hum Genet. 2006 Feb;14(2):207-15
PMID: 16306880
-
Cohort profile: 1958 British birth cohort (National Child Development Study).
Int J Epidemiol. 2006 Feb;35(1):34-41
PMID: 16155052
-
Multiplexed genotyping with sequence-tagged molecular inversion probes.
Nat Biotechnol. 2003 Jun;21(6):673-8
PMID: 12730666
-
Detecting disease associations due to linkage disequilibrium using haplotype tags: a class of tests and the determinants of statistical power.
Hum Hered. 2003;56(1-3):18-31
PMID: 14614235
-
Incorporating genotyping uncertainty in haplotype inference for single-nucleotide polymorphisms.
Am J Hum Genet. 2004 Mar;74(3):495-510
PMID: 14966673
-
Highly multiplexed molecular inversion probe genotyping: over 10,000 targeted SNPs genotyped in a single tube assay.
Genome Res. 2005 Feb;15(2):269-75
PMID: 15687290
-
Genome-wide association studies: theoretical and practical concerns.
Nat Rev Genet. 2005 Feb;6(2):109-18
PMID: 15716907
-
A haplotype map of the human genome.
Nature. 2005 Oct 27;437(7063):1299-320
PMID: 16255080
-
Population structure, differential bias and genomic control in a large-scale, case-control association study.
Nat Genet. 2005 Nov;37(11):1243-6
PMID: 16228001