Abstract
By combining the pairwise interactions between proteins, as predicted by the conserved co-occurrence of their genes in operons, we obtain protein interaction networks. Here we study the properties of such networks to identify functional modules: sets of proteins that together are involved in a biological process. The complete network contains 3,033 orthologous groups of proteins in 38 genomes. It consists of one giant component, containing 1,611 orthologous groups, and of 516 small disjointed clusters that, on average, contain only 2.7 orthologous groups. These small clusters have a homogeneous functional composition and thus represent functional modules in themselves. Analysis of the giant component reveals that it is a scale-free, small-world network with a high degree of local clustering (C = 0.6). It consists of locally highly connected subclusters that are connected to each other by linker proteins. The linker proteins tend to have multiple functions, or are involved in multiple processes and have an above average probability of being essential. By splitting up the giant component at these linker proteins, we identify 265 subclusters that tend to have a homogeneous functional composition. The rare functional inhomogeneities in our subclusters reflect the mixing of different types of (molecular) functions in a single cellular process, exemplified by subclusters containing both metabolic enzymes as well as the transcription factors that regulate them. Comparative genome analysis, thus, allows identification of a level of functional interaction between that of pairwise interactions, and of the complete genome.
MeSH Terms
Animals
Databases as Topic
Genome
Models, Genetic
Multigene Family
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Snel Berend
European Molecular Biology Laboratory, Meyerhofstrasse 1, 69117 Heidelberg, Germany. snel@embl-heidelberg.de
Bork Peer
Huynen Martijn A
References (31)
31 references, click to expand
-
Distinguishing homologous from analogous proteins.
Syst Zool. 1970 Jun;19(2):99-113
PMID: 5449325
-
Genome alignment, evolution of prokaryotic genome organization, and prediction of gene function using genomic context.
Genome Res. 2001 Mar;11(3):356-72
PMID: 11230160
-
Conservation of gene order: a fingerprint of proteins that physically interact.
Trends Biochem Sci. 1998 Sep;23(9):324-8
PMID: 9787636
-
The archaeal flagellum: a different kind of prokaryotic motility structure.
FEMS Microbiol Rev. 2001 Apr;25(2):147-74
PMID: 11250034
-
KEGG: kyoto encyclopedia of genes and genomes.
Nucleic Acids Res. 2000 Jan 1;28(1):27-30
PMID: 10592173
-
Emergence of scaling in random networks
Science. 1999 Oct 15;286(5439):509-12
PMID: 10521342
-
Predicting regulons and their cis-regulatory motifs by comparative genomics.
Nucleic Acids Res. 2000 Nov 15;28(22):4523-30
PMID: 11071941
-
A comprehensive two-hybrid analysis to explore the yeast protein interactome.
Proc Natl Acad Sci U S A. 2001 Apr 10;98(8):4569-74
PMID: 11283351
-
The use of gene clusters to infer functional coupling.
Proc Natl Acad Sci U S A. 1999 Mar 16;96(6):2896-901
PMID: 10077608
-
Assigning protein functions by comparative genome analysis: protein phylogenetic profiles.
Proc Natl Acad Sci U S A. 1999 Apr 13;96(8):4285-8
PMID: 10200254
-
Generating protein interaction maps from incomplete data: application to fold assignment.
Bioinformatics. 2001;17 Suppl 1:S149-56
PMID: 11473004
-
Predicting protein function by genomic context: quantitative evaluation and qualitative inferences.
Genome Res. 2000 Aug;10(8):1204-10
PMID: 10958638
-
Detecting protein function and protein-protein interactions from genome sequences.
Science. 1999 Jul 30;285(5428):751-3
PMID: 10427000
-
SMART, a simple modular architecture research tool: identification of signaling domains.
Proc Natl Acad Sci U S A. 1998 May 26;95(11):5857-64
PMID: 9600884
-
Proteome Analysis Database: online application of InterPro and CluSTr for the functional classification of proteins in whole genomes.
Nucleic Acids Res. 2001 Jan 1;29(1):44-8
PMID: 11125045
-
The COG database: new developments in phylogenetic classification of proteins from complete genomes.
Nucleic Acids Res. 2001 Jan 1;29(1):22-8
PMID: 11125040
-
Genes linked by fusion events are generally of the same functional category: a systematic analysis of 30 microbial genomes.
Proc Natl Acad Sci U S A. 2001 Jul 3;98(14):7940-5
PMID: 11438739
-
STRING: a web-server to retrieve and display the repeatedly occurring neighbourhood of a gene.
Nucleic Acids Res. 2000 Sep 15;28(18):3442-4
PMID: 10982861
-
Protein interaction maps for complete genomes based on gene fusion events.
Nature. 1999 Nov 4;402(6757):86-90
PMID: 10573422
-
Gene context conservation of a higher order than operons.
Trends Biochem Sci. 2000 Oct;25(10):474-9
PMID: 11050428
-
Prediction of co-regulated genes in Bacillus subtilis on the basis of upstream elements conserved across three closely related species.
Genome Biol. 2001;2(11):RESEARCH0048
PMID: 11737947
-
Measuring genome evolution.
Proc Natl Acad Sci U S A. 1998 May 26;95(11):5849-56
PMID: 9600883
-
Requirement of nickel metabolism proteins HypA and HypB for full activity of both hydrogenase and urease in Helicobacter pylori.
Mol Microbiol. 2001 Jan;39(1):176-82
PMID: 11123699
-
Functional characterization of the S. cerevisiae genome by gene deletion and parallel analysis.
Science. 1999 Aug 6;285(5429):901-6
PMID: 10436161
-
Predicting function: from genes to genomes and back.
J Mol Biol. 1998 Nov 6;283(4):707-25
PMID: 9790834
-
Evolution, language and analogy in functional genomics.
Trends Genet. 2001 Jul;17(7):414-8
PMID: 11418223
-
The large-scale organization of metabolic networks.
Nature. 2000 Oct 5;407(6804):651-4
PMID: 11034217
-
Scale-free behavior in protein domain networks.
Mol Biol Evol. 2001 Sep;18(9):1694-702
PMID: 11504849
-
Collective dynamics of 'small-world' networks.
Nature. 1998 Jun 4;393(6684):440-2
PMID: 9623998
-
A network of protein-protein interactions in yeast.
Nat Biotechnol. 2000 Dec;18(12):1257-61
PMID: 11101803
-
The yeast protein interaction network evolves rapidly and contains few redundant duplicate genes.
Mol Biol Evol. 2001 Jul;18(7):1283-92
PMID: 11420367