Abstract
Accurate detection of protein families allows assignment of protein function and the analysis of functional diversity in complete genomes. Recently, we presented a novel algorithm called TribeMCL for the detection of protein families that is both accurate and efficient. This method allows family analysis to be carried out on a very large scale. Using TribeMCL, we have generated a resource called TRIBES that contains protein family information, comprising annotations, protein sequence alignments and phylogenetic distributions describing 311 257 proteins from 83 completely sequenced genomes. The analysis of at least 60 934 detected protein families reveals that, with the essential families excluded, paralogy levels are similar between prokaryotes, irrespective of genome size. The number of essential families is estimated to be between 366 and 426. We also show that the currently known space of protein families is scale free and discuss the implications of this distribution. In addition, we show that smaller families are often formed by shorter proteins and discuss the reasons for this intriguing pattern. Finally, we analyse the functional diversity of protein families in entire genome sequences. The TRIBES protein family resource is accessible at http://www.ebi.ac.uk/research/cgg/tribes/.
MeSH Terms
Algorithms
Amino Acid Sequence
Cluster Analysis
Databases, Protein
Genome
Phylogeny
Proteins/chemistry,classification,genetics
Sequence Alignment
Sequence Analysis, Protein/methods
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Enright Anton J
Computational Genomics Group, The European Bioinformatics Institute, EMBL Cambridge Outstation, Cambridge CB10 1SD, UK.
Kunin Victor
Ouzounis Christos A
References (27)
27 references, click to expand
-
Functional classes in the three domains of life.
J Mol Evol. 1999 Nov;49(5):551-7
PMID: 10552036
-
Universal protein families and the functional content of the last universal common ancestor.
J Mol Evol. 1999 Oct;49(4):413-23
PMID: 10485999
-
The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
Nucleic Acids Res. 2000 Jan 1;28(1):45-8
PMID: 10592178
-
ProtoMap: automatic classification of protein sequences and hierarchy of protein families.
Nucleic Acids Res. 2000 Jan 1;28(1):49-55
PMID: 10592179
-
The GeneQuiz web server: protein functional analysis through the Web.
Trends Biochem Sci. 2000 Jan;25(1):33-5
PMID: 10637611
-
IMPALA: matching a protein sequence against a collection of PSI-BLAST-constructed position-specific score matrices.
Bioinformatics. 1999 Dec;15(12):1000-11
PMID: 10745990
-
Protein function in the post-genomic era.
Nature. 2000 Jun 15;405(6788):823-6
PMID: 10866208
-
GeneRAGE: a robust algorithm for sequence clustering and domain detection.
Bioinformatics. 2000 May;16(5):451-7
PMID: 10871267
-
Practical limits of function prediction.
Proteins. 2000 Oct 1;41(1):98-107
PMID: 10944397
-
CAST: an iterative algorithm for the complexity analysis of sequence tracts. Complexity analysis of sequence tracts.
Bioinformatics. 2000 Oct;16(10):915-22
PMID: 11120681
-
The COG database: new developments in phylogenetic classification of proteins from complete genomes.
Nucleic Acids Res. 2001 Jan 1;29(1):22-8
PMID: 11125040
-
CluSTr: a database of clusters of SWISS-PROT+TrEMBL proteins.
Nucleic Acids Res. 2001 Jan 1;29(1):33-6
PMID: 11125042
-
Mining the draft human genome.
Nature. 2001 Feb 15;409(6822):827-8
PMID: 11236999
-
The Pfam protein families database.
Nucleic Acids Res. 2002 Jan 1;30(1):276-80
PMID: 11752314
-
SYSTERS, GeneNest, SpliceNest: exploring sequence space from genome to protein.
Nucleic Acids Res. 2002 Jan 1;30(1):299-300
PMID: 11752319
-
An efficient algorithm for large-scale detection of protein families.
Nucleic Acids Res. 2002 Apr 1;30(7):1575-84
PMID: 11917018
-
Studying genomes through the aeons: protein families, pseudogenes and proteome evolution.
J Mol Biol. 2002 May 17;318(5):1155-74
PMID: 12083509
-
Domains, motifs and clusters in the protein universe.
Curr Opin Chem Biol. 2003 Feb;7(1):5-11
PMID: 12547420
-
Myriads of protein families, and still counting.
Genome Biol. 2003;4(2):401
PMID: 12620116
-
GeneTRACE-reconstruction of gene content of ancestral species.
Bioinformatics. 2003 Jul 22;19(11):1412-6
PMID: 12874054
-
COmplete GENome Tracking (COGENT): a flexible data environment for computational genomics.
Bioinformatics. 2003 Jul 22;19(11):1451-2
PMID: 12874064
-
Similar amino acid sequences: chance or common ancestry?
Science. 1981 Oct 9;214(4517):149-59
PMID: 7280687
-
The relation between the divergence of sequence and structure in proteins.
EMBO J. 1986 Apr;5(4):823-6
PMID: 3709526
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694
-
A genomic perspective on protein families.
Science. 1997 Oct 24;278(5338):631-7
PMID: 9381173
-
Automated genome sequence analysis and annotation.
Bioinformatics. 1999 May;15(5):391-412
PMID: 10366660
-
Global transposon mutagenesis and a minimal Mycoplasma genome.
Science. 1999 Dec 10;286(5447):2165-9
PMID: 10591650