Abstract
Power distributions appear in numerous biological, physical and other contexts, which appear to be fundamentally different. In biology, power laws have been claimed to describe the distributions of the connections of enzymes and metabolites in metabolic networks, the number of interactions partners of a given protein, the number of members in paralogous families, and other quantities. In network analysis, power laws imply evolution of the network with preferential attachment, i.e. a greater likelihood of nodes being added to pre-existing hubs. Exploration of different types of evolutionary models in an attempt to determine which of them lead to power law distributions has the potential of revealing non-trivial aspects of genome evolution. A simple model of evolution of the domain composition of proteomes was developed, with the following elementary processes: i) domain birth (duplication with divergence), ii) death (inactivation and/or deletion), and iii) innovation (emergence from non-coding or non-globular sequences or acquisition via horizontal gene transfer). This formalism can be described as a birth, death and innovation model (BDIM). The formulas for equilibrium frequencies of domain families of different size and the total number of families at equilibrium are derived for a general BDIM. All asymptotics of equilibrium frequencies of domain families possible for the given type of models are found and their appearance depending on model parameters is investigated. It is proved that the power law asymptotics appears if, and only if, the model is balanced, i.e. domain duplication and deletion rates are asymptotically equal up to the second order. It is further proved that any power asymptotic with the degree not equal to -1 can appear only if the hypothesis of independence of the duplication/deletion rates on the size of a domain family is rejected. Specific cases of BDIMs, namely simple, linear, polynomial and rational models, are considered in details and the distributions of the equilibrium frequencies of domain families of different size are determined for each case. We apply the BDIM formalism to the analysis of the domain family size distributions in prokaryotic and eukaryotic proteomes and show an excellent fit between these empirical data and a particular form of the model, the second-order balanced linear BDIM. Calculation of the parameters of these models suggests surprisingly high innovation rates, comparable to the total domain birth (duplication) and elimination rates, particularly for prokaryotic genomes. We show that a straightforward model of genome evolution, which does not explicitly include selection, is sufficient to explain the observed distributions of domain family sizes, in which power laws appear as asymptotic. However, for the model to be compatible with the data, there has to be a precise balance between domain birth, death and innovation rates, and this is likely to be maintained by selection. The developed approach is oriented at a mathematical description of evolution of domain composition of proteomes, but a simple reformulation could be applied to models of other evolving networks with preferential attachment.
MeSH Terms
Animals
Arabidopsis/genetics
Bacillus subtilis/genetics
Biological Evolution
Caenorhabditis elegans/genetics
Computer Simulation
Death
Drosophila melanogaster/genetics
Escherichia coli/genetics
Genetic Variation/genetics
Humans
Mathematical Computing
Methanobacteriaceae/genetics
Models, Biological
Parturition
Protein Structure, Tertiary/genetics
Saccharomyces cerevisiae/genetics
Sulfolobus solfataricus/genetics
Thermotoga maritima/genetics
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Karev Georgy P
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA. karev@ncbi.nlm.nih.gov
Wolf Yuri I
Rzhetsky Andrey Y
Berezovskaya Faina S
Koonin Eugene V
References (33)
33 references, click to expand
-
Organization of growing random networks.
Phys Rev E Stat Nonlin Soft Matter Phys. 2001 Jun;63(6 Pt 2):066123
PMID: 11415189
-
Scale invariance in biology: coincidence or footprint of a universal mechanism?
Biol Rev Camb Philos Soc. 2001 May;76(2):161-209
PMID: 11396846
-
The frequency distribution of gene family sizes in complete genomes.
Mol Biol Evol. 1998 May;15(5):583-9
PMID: 9580988
-
The impact of comparative genomics on our understanding of evolution.
Cell. 2000 Jun 9;101(6):573-6
PMID: 10892642
-
Comparison of the complete protein sets of worm and yeast: orthology and divergence.
Science. 1998 Dec 11;282(5396):2022-8
PMID: 9851918
-
Emergence of scaling in random networks
Science. 1999 Oct 15;286(5439):509-12
PMID: 10521342
-
Lineage-specific gene expansions in bacterial and archaeal genomes.
Genome Res. 2001 Apr;11(4):555-65
PMID: 11282971
-
Widespread protein sequence similarities: origins of Escherichia coli genes.
J Bacteriol. 1995 Mar;177(6):1585-8
PMID: 7883716
-
Metabolism and evolution of Haemophilus influenzae deduced from a whole-genome comparison with Escherichia coli.
Curr Biol. 1996 Mar 1;6(3):279-91
PMID: 8805245
-
SMART: identification and annotation of domains from signalling and extracellular protein sequences.
Nucleic Acids Res. 1999 Jan 1;27(1):229-32
PMID: 9847187
-
Lethality and centrality in protein networks.
Nature. 2001 May 3;411(6833):41-2
PMID: 11333967
-
Protein family and fold occurrence in genomes: power-law behaviour and evolutionary model.
J Mol Biol. 2001 Nov 2;313(4):673-81
PMID: 11697896
-
CDD: a database of conserved domain alignments with links to domain three-dimensional structure.
Nucleic Acids Res. 2002 Jan 1;30(1):281-3
PMID: 11752315
-
The dominance of the population by a selected few: power-law behaviour applies to a wide variety of genomic properties.
Genome Biol. 2002 Jul 25;3(8):RESEARCH0040
PMID: 12186647
-
Reconstructing/deconstructing the earliest eukaryotes: how comparative genomics can help.
Cell. 2001 Nov 16;107(4):419-25
PMID: 11719183
-
Sequence similarity analysis of Escherichia coli proteins: functional and evolutionary implications.
Proc Natl Acad Sci U S A. 1995 Dec 5;92(25):11921-5
PMID: 8524875
-
Birth of scale-free molecular networks and the number of distinct DNA and protein domains per genome.
Bioinformatics. 2001 Oct;17(10):988-96
PMID: 11673244
-
Comparative genomics of the eukaryotes.
Science. 2000 Mar 24;287(5461):2204-15
PMID: 10731134
-
Estimating the number of protein folds and families from complete genome data.
J Mol Biol. 2000 Jun 16;299(4):897-905
PMID: 10843846
-
Automatic clustering of orthologs and in-paralogs from pairwise species comparisons.
J Mol Biol. 2001 Dec 14;314(5):1041-52
PMID: 11743721
-
Classes of small-world networks.
Proc Natl Acad Sci U S A. 2000 Oct 10;97(21):11149-52
PMID: 11005838
-
The large-scale organization of metabolic networks.
Nature. 2000 Oct 5;407(6804):651-4
PMID: 11034217
-
Scale-free behavior in protein domain networks.
Mol Biol Evol. 2001 Sep;18(9):1694-702
PMID: 11504849
-
Error and attack tolerance of complex networks
Nature. 2000 Jul 27;406(6794):378-82
PMID: 10935628
-
Amelioration of bacterial genomes: rates of change and exchange.
J Mol Evol. 1997 Apr;44(4):383-97
PMID: 9089078
-
The Pfam protein families database.
Nucleic Acids Res. 2002 Jan 1;30(1):276-80
PMID: 11752314
-
Scaling properties of scale-free evolving networks: continuous approach.
Phys Rev E Stat Nonlin Soft Matter Phys. 2001 May;63(5 Pt 2):056125
PMID: 11414979
-
Horizontal gene transfer in prokaryotes: quantification and classification.
Annu Rev Microbiol. 2001;55:709-42
PMID: 11544372
-
The role of lineage-specific gene family expansion in the evolution of eukaryotes.
Genome Res. 2002 Jul;12(7):1048-59
PMID: 12097341
-
The COG database: a tool for genome-scale analysis of protein functions and evolution.
Nucleic Acids Res. 2000 Jan 1;28(1):33-6
PMID: 10592175
-
Lineage-specific loss and divergence of functionally linked genes in eukaryotes.
Proc Natl Acad Sci U S A. 2000 Oct 10;97(21):11319-24
PMID: 11016957
-
Gene duplications in H. influenzae.
Nature. 1995 Nov 9;378(6553):140
PMID: 7477316
-
Initial sequencing and analysis of the human genome.
Nature. 2001 Feb 15;409(6822):860-921
PMID: 11237011