Abstract
Systems biologists work with many kinds of data, from many different sources, using a variety of software tools. Each of these tools typically excels at one type of analysis, such as of microarrays, of metabolic networks and of predicted protein structure. A crucial challenge is to combine the capabilities of these (and other forthcoming) data resources and tools to create a data exploration and analysis environment that does justice to the variety and complexity of systems biology data sets. A solution to this problem should recognize that data types, formats and software in this high throughput age of biology are constantly changing. In this paper we describe the Gaggle -a simple, open-source Java software environment that helps to solve the problem of software and database integration. Guided by the classic software engineering strategy of separation of concerns and a policy of semantic flexibility, it integrates existing popular programs and web resources into a user-friendly, easily-extended environment. We demonstrate that four simple data types (names, matrices, networks, and associative arrays) are sufficient to bring together diverse databases and software. We highlight some capabilities of the Gaggle with an exploration of Helicobacter pylori pathogenesis genes, in which we identify a putative ricin-like protein -a discovery made possible by simultaneous data exploration using a wide range of publicly available data and a variety of popular bioinformatics software tools. We have integrated diverse databases (for example, KEGG, BioCyc, String) and software (Cytoscape, DataMatrixViewer, R statistical environment, and TIGR Microarray Expression Viewer). Through this loose coupling of diverse software and databases the Gaggle enables simultaneous exploration of experimental data (mRNA and protein abundance, protein-protein and protein-DNA interactions), functional associations (operon, chromosomal proximity, phylogenetic pattern), metabolic pathways (KEGG) and Pubmed abstracts (STRING web resource), creating an exploratory environment useful to 'web browser and spreadsheet biologists', to statistically savvy computational biologists, and those in between. The Gaggle uses Java RMI and Java Web Start technologies and can be found at http://gaggle.systemsbiology.net.
MeSH Terms
Computational Biology/methods
Database Management Systems
Databases, Factual
Information Storage and Retrieval/methods
Programming Languages
Software
Systems Integration
User-Computer Interface
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Shannon Paul T
Institute for Systems Biology, Seattle, WA 98103, USA. pshannon@systemsbiology.org
Reiss David J
Bonneau Richard
Baliga Nitin S
References (23)
23 references, click to expand
-
The Stanford Microarray Database accommodates additional microarray platforms and data formats.
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D580-2
PMID: 15608265
-
STRING: known and predicted protein-protein associations, integrated and transferred across organisms.
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D433-7
PMID: 15608232
-
Altered states: involvement of phosphorylated CagA in the induction of host cellular growth changes by Helicobacter pylori.
Proc Natl Acad Sci U S A. 1999 Dec 7;96(25):14559-64
PMID: 10588744
-
The protein-protein interaction map of Helicobacter pylori.
Nature. 2001 Jan 11;409(6817):211-5
PMID: 11196647
-
Helicobacter pylori infection.
N Engl J Med. 2002 Oct 10;347(15):1175-86
PMID: 12374879
-
BioMOBY: an open source biological web services proposal.
Brief Bioinform. 2002 Dec;3(4):331-41
PMID: 12511062
-
The KEGG database.
Novartis Found Symp. 2002;247:91-101; discussion 101-3, 119-28, 244-52
PMID: 12539951
-
The systems biology markup language (SBML): a medium for representation and exchange of biochemical network models.
Bioinformatics. 2003 Mar 1;19(4):524-31
PMID: 12611808
-
TM4: a free, open-source system for microarray data management and analysis.
Biotechniques. 2003 Feb;34(2):374-8
PMID: 12613259
-
A life scientist's gateway to distributed data management and computing: the PathPort/ToolBus framework.
OMICS. 2003 Spring;7(1):79-88
PMID: 12831562
-
Automated prediction of CASP-5 structures using the Robetta server.
Proteins. 2003;53 Suppl 6:524-33
PMID: 14579342
-
Cytoscape: a software environment for integrated models of biomolecular interaction networks.
Genome Res. 2003 Nov;13(11):2498-504
PMID: 14597658
-
caCORE: a common infrastructure for cancer informatics.
Bioinformatics. 2003 Dec 12;19(18):2404-12
PMID: 14668224
-
Prolinks: a database of protein functional linkages derived from coevolution.
Genome Biol. 2004;5(5):R35
PMID: 15128449
-
Bioconductor: open software development for computational biology and bioinformatics.
Genome Biol. 2004;5(10):R80
PMID: 15461798
-
Identification, characterization, and spatial localization of two flagellin species in Helicobacter pylori flagella.
J Bacteriol. 1991 Feb;173(3):937-46
PMID: 1704004
-
Comparative ultrastructural and functional studies of Helicobacter pylori and Helicobacter mustelae flagellin mutants: both flagellin subunits, FlaA and FlaB, are necessary for full motility in Helicobacter species.
J Bacteriol. 1995 Jun;177(11):3010-20
PMID: 7768796
-
Integrated access to metabolic and genomic data.
J Comput Biol. 1996 Spring;3(1):191-212
PMID: 8697237
-
Colonization of gnotobiotic piglets by Helicobacter pylori deficient in two flagellin genes.
Infect Immun. 1996 Jul;64(7):2445-8
PMID: 8698465
-
The role of lipopolysaccharide in Helicobacter pylori pathogenesis.
Aliment Pharmacol Ther. 1996 Apr;10 Suppl 1:39-50
PMID: 8730258
-
The morphological transition of Helicobacter pylori cells from spiral to coccoid is preceded by a substantial modification of the cell wall.
J Bacteriol. 1999 Jun;181(12):3710-5
PMID: 10368145
-
Taverna: a tool for the composition and enactment of bioinformatics workflows.
Bioinformatics. 2004 Nov 22;20(17):3045-54
PMID: 15201187
-
Querying and computing with BioCyc databases.
Bioinformatics. 2005 Aug 15;21(16):3454-5
PMID: 15961440