Home LiteratureArticle Details
PMID: 12611800 Published · ppublish English Comparative Study Evaluation Study Journal Article Research Support, U.S. Gov't, Non-P.H.S. Validation Study

Comparisons and validation of statistical clustering techniques for microarray gene expression data.

Bioinformatics (Oxford, England) ·Vol. 19 ·No. 4 ·2003-03-01 ·Pages 459-66

Datta S, Datta S

Abstract

With the advent of microarray chip technology, large data sets are emerging containing the simultaneous expression levels of thousands of genes at various time points during a biological process. Biologists are attempting to group genes based on the temporal pattern of their expression levels. While the use of hierarchical clustering (UPGMA) with correlation 'distance' has been the most common in the microarray studies, there are many more choices of clustering algorithms in pattern recognition and statistics literature. At the moment there do not seem to be any clear-cut guidelines regarding the choice of a clustering algorithm to be used for grouping genes based on their expression profiles. In this paper, we consider six clustering algorithms (of various flavors!) and evaluate their performances on a well-known publicly available microarray data set on sporulation of budding yeast and on two simulated data sets. Among other things, we formulate three reasonable validation strategies that can be used with any clustering algorithm when temporal observations or replications are present. We evaluate each of these six clustering methods with these validation measures. While the 'best' method is dependent on the exact validation strategy and the number of clusters to be used, overall Diana appears to be a solid performer. Interestingly, the performance of correlation-based hierarchical clustering and model-based clustering (another method that has been advocated by a number of researchers) appear to be on opposite extremes, depending on what validation measure one employs. Next it is shown that the group means produced by Diana are the closest and those produced by UPGMA are the farthest from a model profile based on a set of hand-picked genes. S+ codes for the partial least squares based clustering are available from the authors upon request. All other clustering methods considered have S+ implementation in the library MASS. S+ codes for calculating the validation measures are available from the authors upon request. The sporulation data set is publicly available at http://cmgm.stanford.edu/pbrown/sporulation

MeSH Terms
Algorithms Cluster Analysis Computer Simulation Gene Expression Profiling/methods Gene Expression Regulation/genetics,physiology Models, Genetic Models, Statistical Oligonucleotide Array Sequence Analysis/methods Pattern Recognition, Automated Reproducibility of Results Sensitivity and Specificity Sequence Alignment/methods Sequence Analysis, DNA/methods Software Spores, Fungal/physiology Yeasts/genetics,physiology
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Datta Susmita
Department of Mathematics and Statistics and Department of Biology, Georgia State University, Atlanta, GA 30303, USA. sdatta@mathstat.gsu.edu
Datta Somnath
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4803
Published
2003-03-01
Pages
459-66
Language
English
Region
England
NLM ID
9808944
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com