Home LiteratureArticle Details
PMID: 12925520 Published · ppublish English Journal Article Research Support, U.S. Gov't, P.H.S.

Exploration, normalization, and summaries of high density oligonucleotide array probe level data.

Biostatistics (Oxford, England) ·Vol. 4 ·No. 2 ·2003-04-00 ·Pages 249-64

Irizarry RA, Hobbs B, Collin F, Beazer-Barclay YD, Antonellis KJ, Scherf U, Speed TP

Abstract

In this paper we report exploratory analyses of high-density oligonucleotide array data from the Affymetrix GeneChip system with the objective of improving upon currently used measures of gene expression. Our analyses make use of three data sets: a small experimental study consisting of five MGU74A mouse GeneChip arrays, part of the data from an extensive spike-in study conducted by Gene Logic and Wyeth's Genetics Institute involving 95 HG-U95A human GeneChip arrays; and part of a dilution study conducted by Gene Logic involving 75 HG-U95A GeneChip arrays. We display some familiar features of the perfect match and mismatch probe (PM and MM) values of these data, and examine the variance-mean relationship with probe-level data from probes believed to be defective, and so delivering noise only. We explain why we need to normalize the arrays to one another using probe level intensities. We then examine the behavior of the PM and MM using spike-in data and assess three commonly used summary measures: Affymetrix's (i) average difference (AvDiff) and (ii) MAS 5.0 signal, and (iii) the Li and Wong multiplicative model-based expression index (MBEI). The exploratory data analyses of the probe level data motivate a new summary measure that is a robust multi-array average (RMA) of background-adjusted, normalized, and log-transformed PM values. We evaluate the four expression summary measures using the dilution study data, assessing their behavior in terms of bias, variance and (for MBEI and RMA) model fit. Finally, we evaluate the algorithms in terms of their ability to detect known levels of differential expression using the spike-in data. We conclude that there is no obvious downside to using RMA and attaching a standard error (SE) to this quantity using a linear model which removes probe-specific affinities.

MeSH Terms
Algorithms Animals DNA Probes/genetics Data Interpretation, Statistical Gene Expression Profiling/statistics & numerical data Humans Linear Models Mice Normal Distribution Oligonucleotide Array Sequence Analysis/methods Reproducibility of Results Statistics, Nonparametric
Chemicals
DNA Probes
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Irizarry Rafael A
Department of Biostatistics, Johns Hopkins University, Baltimore, MD 21205, USA. rafa@jhu.edu
Hobbs Bridget
Collin Francois
Beazer-Barclay Yasmin D
Antonellis Kristen J
Scherf Uwe
Speed Terence P
Article Info
Journal
Biostatistics (Oxford, England)
Abbr.
Biostatistics
ISSN
1465-4644
Published
2003-04-00
Pages
249-64
Language
English
Region
England
NLM ID
100897327
Subset
IM
Grants
NHLBI NIH HHS · U01 HL66583 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com