Home LiteratureArticle Details
PMID: 14978222 Published · epublish English Comparative Study Journal Article

LSimpute: accurate estimation of missing values in microarray data with least squares methods.

Nucleic acids research ·Vol. 32 ·No. 3 ·2004-02-20 ·Pages e34

Bø TH, Dysvik B, Jonassen I

Abstract

Microarray experiments generate data sets with information on the expression levels of thousands of genes in a set of biological samples. Unfortunately, such experiments often produce multiple missing expression values, normally due to various experimental problems. As many algorithms for gene expression analysis require a complete data matrix as input, the missing values have to be estimated in order to analyze the available data. Alternatively, genes and arrays can be removed until no missing values remain. However, for genes or arrays with only a small number of missing values, it is desirable to impute those values. For the subsequent analysis to be as informative as possible, it is essential that the estimates for the missing gene expression values are accurate. A small amount of badly estimated missing values in the data might be enough for clustering methods, such as hierachical clustering or K-means clustering, to produce misleading results. Thus, accurate methods for missing value estimation are needed. We present novel methods for estimation of missing values in microarray data sets that are based on the least squares principle, and that utilize correlations between both genes and arrays. For this set of methods, we use the common reference name LSimpute. We compare the estimation accuracy of our methods with the widely used KNNimpute on three complete data matrices from public data sets by randomly knocking out data (labeling as missing). From these tests, we conclude that our LSimpute methods produce estimates that consistently are more accurate than those obtained using KNNimpute. Additionally, we examine a more classic approach to missing value estimation based on expectation maximization (EM). We refer to our EM implementations as EMimpute, and the estimate errors using the EMimpute methods are compared with those our novel methods produce. The results indicate that on average, the estimates from our best performing LSimpute method are at least as accurate as those from the best EMimpute algorithm.

MeSH Terms
Algorithms Computational Biology/methods Gene Expression Profiling Internet Linear Models Oligonucleotide Array Sequence Analysis/standards Reproducibility of Results Software
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Bø Trond Hellem
Department of Informatics, BCCS, University of Bergen, HIB, N5020 Bergen, Norway. trondb@ii.uib.no
Dysvik Bjarte
Jonassen Inge
References (15)
15 references, click to expand
  1. Interpreting patterns of gene expression with self-organizing maps: methods and application to hematopoietic differentiation.
    Proc Natl Acad Sci U S A. 1999 Mar 16;96(6):2907-12 PMID: 10077610
  2. Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays.
    Proc Natl Acad Sci U S A. 1999 Jun 8;96(12):6745-50 PMID: 10359783
  3. Molecular classification of cancer: class discovery and class prediction by gene expression monitoring.
    Science. 1999 Oct 15;286(5439):531-7 PMID: 10521349
  4. Distinct types of diffuse large B-cell lymphoma identified by gene expression profiling.
    Nature. 2000 Feb 3;403(6769):503-11 PMID: 10676951
  5. Systematic variation in gene expression patterns in human cancer cell lines.
    Nat Genet. 2000 Mar;24(3):227-35 PMID: 10700174
  6. Principal components analysis to summarize microarray experiments: application to sporulation time series.
    Pac Symp Biocomput. 2000;:455-66 PMID: 10902193
  7. Molecular portraits of human breast tumours.
    Nature. 2000 Aug 17;406(6797):747-52 PMID: 10963602
  8. Singular value decomposition for genome-wide expression data processing and modeling.
    Proc Natl Acad Sci U S A. 2000 Aug 29;97(18):10101-6 PMID: 10963673
  9. Gene expression data analysis.
    FEBS Lett. 2000 Aug 25;480(1):17-24 PMID: 10967323
  10. Genomic expression programs in the response of yeast cells to environmental changes.
    Mol Biol Cell. 2000 Dec;11(12):4241-57 PMID: 11102521
  11. Missing value estimation methods for DNA microarrays.
    Bioinformatics. 2001 Jun;17(6):520-5 PMID: 11395428
  12. Judging the quality of gene expression-based clustering methods using gene annotation.
    Genome Res. 2002 Oct;12(10):1574-81 PMID: 12368250
  13. A gene-expression program reflecting the innate immune response of cultured intestinal epithelial cells to infection by Listeria monocytogenes.
    Genome Biol. 2003;4(1):R2 PMID: 12537547
  14. The transcriptional program of sporulation in budding yeast.
    Science. 1998 Oct 23;282(5389):699-705 PMID: 9784122
  15. Cluster analysis and display of genome-wide expression patterns.
    Proc Natl Acad Sci U S A. 1998 Dec 8;95(25):14863-8 PMID: 9843981
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2004-02-20
Epub
2004-00-20
Pages
e34
Language
English
Region
England
NLM ID
0411011
PMCID
PMC374359
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com