Abstract
A decade ago, the Gene Expression Omnibus (GEO) database was established at the National Center for Biotechnology Information (NCBI). The original objective of GEO was to serve as a public repository for high-throughput gene expression data generated mostly by microarray technology. However, the research community quickly applied microarrays to non-gene-expression studies, including examination of genome copy number variation and genome-wide profiling of DNA-binding proteins. Because the GEO database was designed with a flexible structure, it was possible to quickly adapt the repository to store these data types. More recently, as the microarray community switches to next-generation sequencing technologies, GEO has again adapted to host these data sets. Today, GEO stores over 20,000 microarray- and sequence-based functional genomics studies, and continues to handle the majority of direct high-throughput data submissions from the research community. Multiple mechanisms are provided to help users effectively search, browse, download and visualize the data at the level of individual genes or entire studies. This paper describes recent database enhancements, including new search and data representation tools, as well as a brief review of how the community uses GEO data. GEO is freely accessible at http://www.ncbi.nlm.nih.gov/geo/.
MeSH Terms
Databases, Genetic
Gene Expression Profiling
Genomics
Oligonucleotide Array Sequence Analysis
User-Computer Interface
Authors & Affiliations
15 authors, click to expand affiliations / ORCID
Barrett Tanya
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, 45 Center Drive, Bethesda, MD 20892, USA. barrett@ncbi.nlm.nih.gov
Troup Dennis B
Wilhite Stephen E
Ledoux Pierre
Evangelista Carlos
Kim Irene F
Tomashevsky Maxim
Marshall Kimberly A
Phillippy Katherine H
Sherman Patti M
Muertter Rolf N
Holko Michelle
Ayanbule Oluwabukunmi
Yefanov Andrey
Soboleva Alexandra
References (22)
22 references, click to expand
-
TranscriptomeBrowser: a powerful and flexible toolbox to explore productively the transcriptional landscape of the Gene Expression Omnibus database.
PLoS One. 2008;3(12):e4001
PMID: 19104654
-
Global reconstruction of the human metabolic network based on genomic and bibliomic data.
Proc Natl Acad Sci U S A. 2007 Feb 6;104(6):1777-82
PMID: 17267599
-
Network-based elucidation of human disease similarities reveals common functional modules enriched for pluripotent drug targets.
PLoS Comput Biol. 2010 Feb 05;6(2):e1000662
PMID: 20140234
-
An 8-gene qRT-PCR-based gene expression score that has prognostic value in early breast cancer.
BMC Cancer. 2010 Jun 28;10:336
PMID: 20584321
-
NCBI Epigenomics: a new public resource for exploring epigenomic data sets.
Nucleic Acids Res. 2011 Jan;39(Database issue):D908-12
PMID: 21075792
-
Archiving next generation sequencing data.
Nucleic Acids Res. 2010 Jan;38(Database issue):D870-1
PMID: 19965774
-
Bayesian approach to transforming public gene expression repositories into disease diagnosis databases.
Proc Natl Acad Sci U S A. 2010 Apr 13;107(15):6823-8
PMID: 20360561
-
Minimum information about a microarray experiment (MIAME)-toward standards for microarray data.
Nat Genet. 2001 Dec;29(4):365-71
PMID: 11726920
-
Detection of gene orthology from gene co-expression and protein interaction networks.
BMC Bioinformatics. 2010 Apr 29;11 Suppl 3:S7
PMID: 20438654
-
Gene expression prediction by soft integration and the elastic net-best performance of the DREAM3 gene expression challenge.
PLoS One. 2010 Feb 16;5(2):e9134
PMID: 20169069
-
GEOGLE: context mining tool for the correlation between gene expression and the phenotypic distinction.
BMC Bioinformatics. 2009 Aug 25;10:264
PMID: 19703314
-
Gene Expression Omnibus: NCBI gene expression and hybridization array data repository.
Nucleic Acids Res. 2002 Jan 1;30(1):207-10
PMID: 11752295
-
Simultaneous analysis of distinct Omics data sets with integration of biological knowledge: Multiple Factor Analysis approach.
BMC Genomics. 2009 Jan 20;10:32
PMID: 19154582
-
Database resources of the National Center for Biotechnology Information.
Nucleic Acids Res. 2009 Jan;37(Database issue):D5-15
PMID: 18940862
-
Predicting environmental chemical factors associated with disease-related gene expression data.
BMC Med Genomics. 2010 May 06;3:17
PMID: 20459635
-
Microarray standards at last.
Nature. 2002 Sep 26;419(6905):323
PMID: 12352992
-
MARQ: an online tool to mine GEO for experiments with similar or opposite gene expression signatures.
Nucleic Acids Res. 2010 Jul;38(Web Server issue):W228-32
PMID: 20513648
-
Gene expression-based classification of non-small cell lung carcinomas and survival prediction.
PLoS One. 2010 Apr 22;5(4):e10312
PMID: 20421987
-
ArrayExpress update--from an archive of functional genomics experiments to the atlas of gene expression.
Nucleic Acids Res. 2009 Jan;37(Database issue):D868-72
PMID: 19015125
-
GEOquery: a bridge between the Gene Expression Omnibus (GEO) and BioConductor.
Bioinformatics. 2007 Jul 15;23(14):1846-7
PMID: 17496320
-
NCBI GEO: archive for high-throughput functional genomic data.
Nucleic Acids Res. 2009 Jan;37(Database issue):D885-90
PMID: 18940857
-
Reducing the algorithmic variability in transcriptome-based inference.
Bioinformatics. 2010 May 1;26(9):1185-91
PMID: 20212019