Abstract
The sequencing of genomes provides us with an inventory of the 'molecular parts' in nature, such as protein families and folds, and their functions in living organisms. Through the analysis of such inventories, it has been shown that different genomes have very different usage of parts; for example, the common folds in the worm are very different from those in Escherichia coli. Despite these differences, we find that the genomic occurrence of generalized parts follows a well-known mathematical framework called the power law, with a few parts occurring many times and most occurring only a few times. This observation is true in a wide variety of genomic contexts. Earlier studies found power laws in a few specific cases, such as the occurrence of protein families. Here, we find many further cases of power-law behavior, for example in the occurrence of pseudogenes and in levels of gene expression. We show comprehensively that this behavior applies across many different genomes, for many different types of parts (DNA words, InterPro families, protein superfamilies and folds, pseudogene families and pseudomotifs), and for the many disparate attributes associated with these parts (their functions, interactions and expression levels). Power-law behavior provides a concise mathematical description of an important biological feature: the sheer dominance of a few members over the overall population. We present this behavior in a unified framework and propose that all these observations are connected to an underlying DNA duplication process as genomes evolved to their current state.
MeSH Terms
Animals
Caenorhabditis elegans/genetics
Computational Biology/methods,statistics & numerical data
DNA, Helminth/genetics
Gene Frequency/genetics
Genes, Dominant/genetics
Genes, Duplicate/genetics
Genes, Helminth/genetics
Genetics, Behavioral
Genome
Multigene Family/genetics
Oligonucleotides/genetics
Protein Binding/genetics
Protein Folding
Protein Interaction Mapping/statistics & numerical data
Saccharomyces cerevisiae/genetics
Saccharomyces cerevisiae Proteins/biosynthesis,physiology
Chemicals
DNA, Helminth
Oligonucleotides
Saccharomyces cerevisiae Proteins
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Luscombe Nicholas M
Department of Molecular Biophysics and Biochemistry, Yale University, New Haven, CT 06520-8114, USA.
Qian Jiang
Zhang Zhaolei
Johnson Ted
Gerstein Mark
References (30)
30 references, click to expand
-
Analysis of the yeast transcriptome with structural and functional categories: characterizing highly expressed proteins.
Nucleic Acids Res. 2000 Mar 15;28(6):1481-8
PMID: 10684945
-
The frequency distribution of gene family sizes in complete genomes.
Mol Biol Evol. 1998 May;15(5):583-9
PMID: 9580988
-
Emergence of scaling in random networks
Science. 1999 Oct 15;286(5439):509-12
PMID: 10521342
-
SCOP: a structural classification of proteins database.
Nucleic Acids Res. 2000 Jan 1;28(1):257-9
PMID: 10592240
-
Protein fold recognition using sequence profiles and its application in structural genomics.
Adv Protein Chem. 2000;54:245-75
PMID: 10829230
-
Integrative database analysis in structural genomics.
Nat Struct Biol. 2000 Nov;7 Suppl:960-3
PMID: 11104000
-
Molecular fossils in the human genome: identification and analysis of the pseudogenes in chromosomes 21 and 22.
Genome Res. 2002 Feb;12(2):272-80
PMID: 11827946
-
Proteome Analysis Database: online application of InterPro and CluSTr for the functional classification of proteins in whole genomes.
Nucleic Acids Res. 2001 Jan 1;29(1):44-8
PMID: 11125045
-
Mapping protein family interactions: intramolecular and intermolecular protein family interaction repertoires in the PDB and yeast.
J Mol Biol. 2001 Mar 30;307(3):929-38
PMID: 11273711
-
Predictions of gene family distributions in microbial genomes: evolution by gene duplication and modification.
Phys Rev Lett. 2000 Sep 18;85(12):2641-4
PMID: 10978127
-
Protein family and fold occurrence in genomes: power-law behaviour and evolutionary model.
J Mol Biol. 2001 Nov 2;313(4):673-81
PMID: 11697896
-
Digging for dead genes: an analysis of the characteristics of the pseudogene population in the Caenorhabditis elegans genome.
Nucleic Acids Res. 2001 Feb 1;29(3):818-30
PMID: 11160906
-
The InterPro database, an integrated documentation resource for protein families, domains and functional sites.
Nucleic Acids Res. 2001 Jan 1;29(1):37-40
PMID: 11125043
-
Assignment of homology to genome sequences using a library of hidden Markov models that represent all proteins of known structure.
J Mol Biol. 2001 Nov 2;313(4):903-19
PMID: 11697912
-
Zipf's law, the central limit theorem, and the random division of the unit interval.
Phys Rev E Stat Phys Plasmas Fluids Relat Interdiscip Topics. 1996 Jul;54(1):220-223
PMID: 9965063
-
Birth of scale-free molecular networks and the number of distinct DNA and protein domains per genome.
Bioinformatics. 2001 Oct;17(10):988-96
PMID: 11673244
-
Explaining "linguistic features" of noncoding DNA.
Science. 1996 Jan 5;271(5245):14-5
PMID: 8539587
-
Evolution of function in protein superfamilies, from a structural perspective.
J Mol Biol. 2001 Apr 6;307(4):1113-43
PMID: 11286560
-
PartsList: a web-based system for dynamically ranking protein folds based on disparate attributes, including whole-genome expression and interaction information.
Nucleic Acids Res. 2001 Apr 15;29(8):1750-64
PMID: 11292848
-
The large-scale organization of metabolic networks.
Nature. 2000 Oct 5;407(6804):651-4
PMID: 11034217
-
The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.
J Mol Biol. 1999 Apr 23;288(1):147-64
PMID: 10329133
-
Error and attack tolerance of complex networks
Nature. 2000 Jul 27;406(6794):378-82
PMID: 10935628
-
Noncoding DNA, Zipf's law, and language.
Science. 1995 May 12;268(5212):789
PMID: 7754361
-
Oligonucleotide frequencies in DNA follow a Yule distribution.
Comput Chem. 1996 Mar;20(1):35-8
PMID: 16718864
-
No signs of hidden language in noncoding DNA.
Phys Rev Lett. 1996 Mar 11;76(11):1977
PMID: 10060572
-
Comment on "Linguistic features of noncoding DNA sequences"
Phys Rev Lett. 1996 Mar 11;76(11):1978
PMID: 10060573
-
Can Zipf distinguish language from noise in noncoding DNA?
Phys Rev Lett. 1996 Mar 11;76(11):1976
PMID: 10060571
-
Linguistic features of noncoding DNA sequences.
Phys Rev Lett. 1994 Dec 5;73(23):3169-72
PMID: 10057305
-
The yeast protein interaction network evolves rapidly and contains few redundant duplicate genes.
Mol Biol Evol. 2001 Jul;18(7):1283-92
PMID: 11420367
-
A structural census of genomes: comparing bacterial, eukaryotic, and archaeal genomes in terms of protein structure.
J Mol Biol. 1997 Dec 12;274(4):562-76
PMID: 9417935