Abstract
The modeling of complex systems, as disparate as the World Wide Web and the cellular metabolism, as networks has recently uncovered a set of generic organizing principles: Most of these systems are scale-free while at the same time modular, resulting in a hierarchical architecture. The structure of the protein domain network, where individual domains correspond to nodes and their co-occurrences in a protein are interpreted as links, also falls into this category, suggesting that domains involved in the maintenance of increasingly developed, multicellular organisms accumulate links. Here, we take the next step by studying link based properties of the protein domain co-occurrence networks of the eukaryotes S. cerevisiae, C. elegans, D. melanogaster, M. musculus and H. sapiens. We construct the protein domain co-occurrence networks from the PFAM database and analyze them by applying a k-core decomposition method that isolates the globally central (highly connected domains in the central cores) from the locally central (highly connected domains in the peripheral cores) protein domains through an iterative peeling process. Furthermore, we compare the subnetworks thus obtained to the physical domain interaction network of S. cerevisiae. We find that the innermost cores of the domain co-occurrence networks gradually grow with increasing degree of evolutionary development in going from single cellular to multicellular eukaryotes. The comparison of the cores across all the organisms under consideration uncovers patterns of domain combinations that are predominately involved in protein functions such as cell-cell contacts and signal transduction. Analyzing a weighted interaction network of PFAM domains of yeast, we find that domains having only a few partners frequently interact with these, while the converse is true for domains with a multitude of partners. Combining domain co-occurrence and interaction information, we observe that the co-occurrence of domains in the innermost cores (globally central domains) strongly coincides with physical interaction. The comparison of the multicellular eukaryotic domain co-occurrence networks with the single celled of S. cerevisiae (the overlap network) uncovers small, connected network patterns. We hypothesize that these patterns, consisting of the domains and links preserved through evolution, may constitute nucleation kernels for the evolutionary increase in proteome complexity. Combining co-occurrence and physical interaction data we argue that the driving force behind domain fusions is a collective effect caused by the number of interactions and not the individual interaction frequency.
MeSH Terms
Animals
Cluster Analysis
Computational Biology
Databases, Protein
Evolution, Molecular
Genomics
Humans
Models, Genetic
Models, Statistical
Protein Binding
Protein Interaction Mapping
Protein Structure, Tertiary
Proteins/chemistry
Proteomics/methods
Saccharomyces cerevisiae/metabolism
Sequence Analysis, Protein
Signal Transduction
Species Specificity
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Wuchty Stefan
Northwestern Institute of Complexity, Northwestern University, Evanston, IL, USA. s-wuchty@northwestern.edu
Almaas Eivind
References (28)
28 references, click to expand
-
Apoptotic molecular machinery: vastly increased complexity in vertebrates revealed by genome comparisons.
Science. 2001 Feb 16;291(5507):1279-84
PMID: 11181990
-
Small worlds in RNA structures.
Nucleic Acids Res. 2003 Feb 1;31(3):1108-17
PMID: 12560509
-
Emergence of scaling in random networks
Science. 1999 Oct 15;286(5439):509-12
PMID: 10521342
-
The small world of metabolism.
Nat Biotechnol. 2000 Nov;18(11):1121-2
PMID: 11062388
-
Modular organization of cellular networks.
Proc Natl Acad Sci U S A. 2003 Feb 4;100(3):1128-33
PMID: 12538875
-
Network biology: understanding the cell's functional organization.
Nat Rev Genet. 2004 Feb;5(2):101-13
PMID: 14735121
-
Hierarchical organization of modularity in metabolic networks.
Science. 2002 Aug 30;297(5586):1551-5
PMID: 12202830
-
Evidence for dynamically organized modularity in the yeast protein-protein interaction network.
Nature. 2004 Jul 1;430(6995):88-93
PMID: 15190252
-
The architecture of complex weighted networks.
Proc Natl Acad Sci U S A. 2004 Mar 16;101(11):3747-52
PMID: 15007165
-
Domains in proteins: definitions, location, and structural principles.
Methods Enzymol. 1985;115:420-30
PMID: 4079796
-
Detecting protein function and protein-protein interactions from genome sequences.
Science. 1999 Jul 30;285(5428):751-3
PMID: 10427000
-
Mapping protein family interactions: intramolecular and intermolecular protein family interaction repertoires in the PDB and yeast.
J Mol Biol. 2001 Mar 30;307(3):929-38
PMID: 11273711
-
Integrative approach for computationally inferring protein domain interactions.
Bioinformatics. 2003 May 22;19(8):923-9
PMID: 12761053
-
Lethality and centrality in protein networks.
Nature. 2001 May 3;411(6833):41-2
PMID: 11333967
-
Domain combinations in archaeal, eubacterial and eukaryotic proteomes.
J Mol Biol. 2001 Jul 6;310(2):311-25
PMID: 11428892
-
The Pfam protein families database.
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D138-41
PMID: 14681378
-
Evolution and topology in the yeast protein interaction network.
Genome Res. 2004 Jul;14(7):1310-4
PMID: 15231746
-
Subnetwork hierarchies of biochemical pathways.
Bioinformatics. 2003 Mar 1;19(4):532-8
PMID: 12611809
-
The sequence of the human genome.
Science. 2001 Feb 16;291(5507):1304-51
PMID: 11181995
-
InterDom: a database of putative interacting protein domains for validating predicted protein interactions and complexes.
Nucleic Acids Res. 2003 Jan 1;31(1):251-4
PMID: 12519994
-
The small world inside large metabolic networks.
Proc Biol Sci. 2001 Sep 7;268(1478):1803-10
PMID: 11522199
-
The large-scale organization of metabolic networks.
Nature. 2000 Oct 5;407(6804):651-4
PMID: 11034217
-
Scale-free behavior in protein domain networks.
Mol Biol Evol. 2001 Sep;18(9):1694-702
PMID: 11504849
-
Collective dynamics of 'small-world' networks.
Nature. 1998 Jun 4;393(6684):440-2
PMID: 9623998
-
Interaction and domain networks of yeast.
Proteomics. 2002 Dec;2(12):1715-23
PMID: 12469341
-
The multiplicity of domains in proteins.
Annu Rev Biochem. 1995;64:287-314
PMID: 7574483
-
The Proteome Analysis database: a tool for the in silico analysis of whole proteomes.
Nucleic Acids Res. 2003 Jan 1;31(1):414-7
PMID: 12520037
-
Initial sequencing and analysis of the human genome.
Nature. 2001 Feb 15;409(6822):860-921
PMID: 11237011