Home LiteratureArticle Details
PMID: 11983868 Published · ppublish English Journal Article

Selection bias in gene extraction on the basis of microarray gene-expression data.

Ambroise C, McLachlan GJ

Abstract

In the context of cancer diagnosis and treatment, we consider the problem of constructing an accurate prediction rule on the basis of a relatively small number of tumor tissue samples of known type containing the expression data on very many (possibly thousands) genes. Recently, results have been presented in the literature suggesting that it is possible to construct a prediction rule from only a few genes such that it has a negligible prediction error rate. However, in these results the test error or the leave-one-out cross-validated error is calculated without allowance for the selection bias. There is no allowance because the rule is either tested on tissue samples that were used in the first instance to select the genes being used in the rule or because the cross-validation of the rule is not external to the selection process; that is, gene selection is not performed in training the rule at each stage of the cross-validation process. We describe how in practice the selection bias can be assessed and corrected for by either performing a cross-validation or applying the bootstrap external to the selection process. We recommend using 10-fold rather than leave-one-out cross-validation, and concerning the bootstrap, we suggest using the so-called .632+ bootstrap error estimate designed to handle overfitted prediction rules. Using two published data sets, we demonstrate that when correction is made for the selection bias, the cross-validated error is no longer zero for a subset of only a few genes.

MeSH Terms
Discriminant Analysis Gene Expression Linear Models Oligonucleotide Array Sequence Analysis/methods Selection Bias
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Ambroise Christophe
Laboratoire Heudiasyc, Unité Mixte de Recherche/Centre National de la Recherche Scientifique 6599, 60200 Compiègne, France.
McLachlan Geoffrey J
References (16)
16 references, click to expand
  1. Knowledge-based analysis of microarray gene expression data by using support vector machines.
    Proc Natl Acad Sci U S A. 2000 Jan 4;97(1):262-7 PMID: 10618406
  2. Molecular classification of cancer: class discovery and class prediction by gene expression monitoring.
    Science. 1999 Oct 15;286(5439):531-7 PMID: 10521349
  3. Tissue classification with gene expression profiles.
    J Comput Biol. 2000;7(3-4):559-83 PMID: 11108479
  4. Support vector machine classification and validation of cancer tissue samples using microarray expression data.
    Bioinformatics. 2000 Oct;16(10):906-14 PMID: 11120680
  5. Analysis of molecular profile data using generative and discriminative methods.
    Physiol Genomics. 2000 Dec 18;4(2):109-126 PMID: 11120872
  6. Identifying marker genes in transcription profiling data using a mixture of feature relevance experts.
    Physiol Genomics. 2001 Mar 8;5(2):99-111 PMID: 11242594
  7. Recursive partitioning for tumor classification with gene expression microarray data.
    Proc Natl Acad Sci U S A. 2001 Jun 5;98(12):6730-5 PMID: 11381113
  8. Feature (gene) selection in gene expression-based tumor classification.
    Mol Genet Metab. 2001 Jul;73(3):239-47 PMID: 11461191
  9. Gene expression patterns of breast carcinomas distinguish tumor subclasses with clinical implications.
    Proc Natl Acad Sci U S A. 2001 Sep 11;98(19):10869-74 PMID: 11553815
  10. Predicting the clinical status of human breast cancer by using gene expression profiles.
    Proc Natl Acad Sci U S A. 2001 Sep 25;98(20):11462-7 PMID: 11562467
  11. Bringing out the best features of expression data.
    Genome Res. 2001 Nov;11(11):1801-2 PMID: 11691842
  12. Biomarker identification by feature wrappers.
    Genome Res. 2001 Nov;11(11):1878-87 PMID: 11691853
  13. A mixture model-based approach to the clustering of microarray expression data.
    Bioinformatics. 2002 Mar;18(3):413-22 PMID: 11934740
  14. Cluster analysis and display of genome-wide expression patterns.
    Proc Natl Acad Sci U S A. 1998 Dec 8;95(25):14863-8 PMID: 9843981
  15. Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays.
    Proc Natl Acad Sci U S A. 1999 Jun 8;96(12):6745-50 PMID: 10359783
  16. Coupled two-way clustering analysis of gene microarray data.
    Proc Natl Acad Sci U S A. 2000 Oct 24;97(22):12079-84 PMID: 11035779
Article Info
Journal
Proceedings of the National Academy of Sciences of the United States of America
Abbr.
Proc Natl Acad Sci U S A
ISSN
0027-8424
Published
2002-05-14
Epub
2002-00-30
Pages
6562-6
Language
English
Region
United States
NLM ID
7505876
PMCID
PMC124442
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com