Home LiteratureArticle Details
PMID: 19837654 Published · ppublish English Journal Article

PICNIC: an algorithm to predict absolute allelic copy number variation with microarray cancer data.

Biostatistics (Oxford, England) ·Vol. 11 ·No. 1 ·2010-01-00 ·Pages 164-75

Greenman CD, Bignell G, Butler A, Edkins S, Hinton J, Beare D, Swamy S, Santarius T, Chen L, Widaa S, Futreal PA, Stratton MR

Abstract

High-throughput oligonucleotide microarrays are commonly employed to investigate genetic disease, including cancer. The algorithms employed to extract genotypes and copy number variation function optimally for diploid genomes usually associated with inherited disease. However, cancer genomes are aneuploid in nature leading to systematic errors when using these techniques. We introduce a preprocessing transformation and hidden Markov model algorithm bespoke to cancer. This produces genotype classification, specification of regions of loss of heterozygosity, and absolute allelic copy number segmentation. Accurate prediction is demonstrated with a combination of independent experimental techniques. These methods are exemplified with affymetrix genome-wide SNP6.0 data from 755 cancer cell lines, enabling inference upon a number of features of biological interest. These data and the coded algorithm are freely available for download.

MeSH Terms
Algorithms Alleles Aneuploidy Bayes Theorem Bias Cell Line, Tumor DNA Copy Number Variations/genetics Genes, Tumor Suppressor Genetic Testing Genotype Humans Internet Loss of Heterozygosity/genetics Markov Chains Models, Statistical Neoplasms/diagnosis,genetics Oligonucleotide Array Sequence Analysis/methods Polymorphism, Single Nucleotide/genetics Polyploidy Reproducibility of Results Sensitivity and Specificity Software
Authors & Affiliations
12 authors, click to expand affiliations / ORCID
Greenman Chris D
Cancer Genome Project, Wellcome Trust Sanger Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SA, UK. cdg@sanger.ac.uk
Bignell Graham
Butler Adam
Edkins Sarah
Hinton Jon
Beare Dave
Swamy Sajani
Santarius Thomas
Chen Lina
Widaa Sara
Futreal P Andy
Stratton Michael R
References (33)
33 references, click to expand
  1. Estimating genome-wide copy number using allele-specific mixture models.
    J Comput Biol. 2008 Sep;15(7):857-66 PMID: 18707534
  2. CARAT: a novel method for allelic detection of DNA copy number changes using high density oligonucleotide arrays.
    BMC Bioinformatics. 2006 Feb 21;7:83 PMID: 16504045
  3. QuantiSNP: an Objective Bayes Hidden-Markov Model to detect and accurately map copy number variation using SNP genotyping data.
    Nucleic Acids Res. 2007;35(6):2013-25 PMID: 17341461
  4. High-resolution genomic profiling of chromosomal aberrations using Infinium whole-genome genotyping.
    Genome Res. 2006 Sep;16(9):1136-48 PMID: 16899659
  5. High-resolution analysis of DNA copy number using oligonucleotide microarrays.
    Genome Res. 2004 Feb;14(2):287-95 PMID: 14762065
  6. PLASQ: a generalized linear model-based procedure to determine allelic dosage in cancer cells from SNP array data.
    Biostatistics. 2007 Apr;8(2):323-36 PMID: 16787995
  7. Hidden Markov models for the assessment of chromosomal alterations using high-throughput SNP arrays.
    Ann Appl Stat. 2008 Jun 1;2(2):687-713 PMID: 19609370
  8. Genotyping and annotation of Affymetrix SNP arrays.
    Nucleic Acids Res. 2006;34(14):e100 PMID: 16899450
  9. Breaking the waves: improved detection of copy number variation from microarray-based comparative genomic hybridization.
    Genome Biol. 2007;8(10):R228 PMID: 17961237
  10. A hierarchical clustering method for estimating copy number variation.
    Biostatistics. 2007 Jul;8(3):632-53 PMID: 17060368
  11. GenoSNP: a variational Bayes within-sample SNP genotyping algorithm that does not require a reference population.
    Bioinformatics. 2008 Oct 1;24(19):2209-14 PMID: 18653518
  12. Circular binary segmentation for the analysis of array-based DNA copy number data.
    Biostatistics. 2004 Oct;5(4):557-72 PMID: 15475419
  13. Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls.
    Nature. 2007 Jun 7;447(7145):661-78 PMID: 17554300
  14. Analysis of array CGH data for cancer studies using fused quantile regression.
    Bioinformatics. 2007 Sep 15;23(18):2470-6 PMID: 17644559
  15. Inferring loss-of-heterozygosity from unpaired tumors using high-density oligonucleotide SNP arrays.
    PLoS Comput Biol. 2006 May;2(5):e41 PMID: 16699594
  16. Integrating copy number polymorphisms into array CGH analysis using a robust HMM.
    Bioinformatics. 2006 Jul 15;22(14):e431-9 PMID: 16873504
  17. Allele-specific amplification in cancer revealed by SNP array analysis.
    PLoS Comput Biol. 2005 Nov;1(6):e65 PMID: 16322765
  18. Exploration, normalization, and genotype calls of high-density oligonucleotide SNP array data.
    Biostatistics. 2007 Apr;8(2):485-99 PMID: 17189563
  19. A multi-array multi-SNP genotyping algorithm for Affymetrix SNP microarrays.
    Bioinformatics. 2007 Jun 15;23(12):1459-67 PMID: 17459966
  20. Assessment of algorithms for high throughput detection of genomic copy number variation in oligonucleotide microarray data.
    BMC Bioinformatics. 2007 Oct 02;8:368 PMID: 17910767
  21. Array painting reveals a high frequency of balanced translocations in breast cancer cell lines that break in cancer-relevant genes.
    Oncogene. 2008 May 22;27(23):3345-59 PMID: 18084325
  22. Common deletion polymorphisms in the human genome.
    Nat Genet. 2006 Jan;38(1):86-92 PMID: 16468122
  23. Analysis of molecular inversion probe performance for allele copy number determination.
    Genome Biol. 2007;8(11):R246 PMID: 18028543
  24. Flexible and accurate detection of genomic copy-number changes from aCGH.
    PLoS Comput Biol. 2007 Jun;3(6):e122 PMID: 17590078
  25. Continuous-index hidden Markov modelling of array CGH copy number data.
    Bioinformatics. 2007 Apr 15;23(8):1006-14 PMID: 17309894
  26. PennCNV: an integrated hidden Markov model designed for high-resolution copy number variation detection in whole-genome SNP genotyping data.
    Genome Res. 2007 Nov;17(11):1665-74 PMID: 17921354
  27. SNiPer-HD: improved genotype calling accuracy by an expectation-maximization algorithm for high-density SNP arrays.
    Bioinformatics. 2007 Jan 1;23(1):57-63 PMID: 17062589
  28. Major copy proportion analysis of tumor samples using SNP arrays.
    BMC Bioinformatics. 2008 Apr 21;9:204 PMID: 18426588
  29. A genotype calling algorithm for affymetrix SNP arrays.
    Bioinformatics. 2006 Jan 1;22(1):7-12 PMID: 16267090
  30. Characterizing the cancer genome in lung adenocarcinoma.
    Nature. 2007 Dec 6;450(7171):893-8 PMID: 17982442
  31. Aneuploidy and cancer.
    Nature. 2004 Nov 18;432(7015):338-41 PMID: 15549096
  32. A Hidden Markov Model to estimate population mixture and allelic copy-numbers in cancers using Affymetrix SNP arrays.
    BMC Bioinformatics. 2007 Nov 09;8:434 PMID: 17996079
  33. Integrated genotype calling and association analysis of SNPs, common copy number polymorphisms and rare CNVs.
    Nat Genet. 2008 Oct;40(10):1253-60 PMID: 18776909
Article Info
Journal
Biostatistics (Oxford, England)
Abbr.
Biostatistics
ISSN
1468-4357
Published
2010-01-00
Epub
2009-00-15
Pages
164-75
Language
English
Region
England
NLM ID
100897327
PMCID
PMC2800165
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com