Home LiteratureArticle Details
PMID: 12888524 Published · ppublish English Journal Article

Protein families and TRIBES in genome sequence space.

Nucleic acids research ·Vol. 31 ·No. 15 ·2003-08-01 ·Pages 4632-8

Enright AJ, Kunin V, Ouzounis CA

Abstract

Accurate detection of protein families allows assignment of protein function and the analysis of functional diversity in complete genomes. Recently, we presented a novel algorithm called TribeMCL for the detection of protein families that is both accurate and efficient. This method allows family analysis to be carried out on a very large scale. Using TribeMCL, we have generated a resource called TRIBES that contains protein family information, comprising annotations, protein sequence alignments and phylogenetic distributions describing 311 257 proteins from 83 completely sequenced genomes. The analysis of at least 60 934 detected protein families reveals that, with the essential families excluded, paralogy levels are similar between prokaryotes, irrespective of genome size. The number of essential families is estimated to be between 366 and 426. We also show that the currently known space of protein families is scale free and discuss the implications of this distribution. In addition, we show that smaller families are often formed by shorter proteins and discuss the reasons for this intriguing pattern. Finally, we analyse the functional diversity of protein families in entire genome sequences. The TRIBES protein family resource is accessible at http://www.ebi.ac.uk/research/cgg/tribes/.

MeSH Terms
Algorithms Amino Acid Sequence Cluster Analysis Databases, Protein Genome Phylogeny Proteins/chemistry,classification,genetics Sequence Alignment Sequence Analysis, Protein/methods
Chemicals
Proteins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Enright Anton J
Computational Genomics Group, The European Bioinformatics Institute, EMBL Cambridge Outstation, Cambridge CB10 1SD, UK.
Kunin Victor
Ouzounis Christos A
References (27)
27 references, click to expand
  1. Functional classes in the three domains of life.
    J Mol Evol. 1999 Nov;49(5):551-7 PMID: 10552036
  2. Universal protein families and the functional content of the last universal common ancestor.
    J Mol Evol. 1999 Oct;49(4):413-23 PMID: 10485999
  3. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  4. ProtoMap: automatic classification of protein sequences and hierarchy of protein families.
    Nucleic Acids Res. 2000 Jan 1;28(1):49-55 PMID: 10592179
  5. The GeneQuiz web server: protein functional analysis through the Web.
    Trends Biochem Sci. 2000 Jan;25(1):33-5 PMID: 10637611
  6. IMPALA: matching a protein sequence against a collection of PSI-BLAST-constructed position-specific score matrices.
    Bioinformatics. 1999 Dec;15(12):1000-11 PMID: 10745990
  7. Protein function in the post-genomic era.
    Nature. 2000 Jun 15;405(6788):823-6 PMID: 10866208
  8. GeneRAGE: a robust algorithm for sequence clustering and domain detection.
    Bioinformatics. 2000 May;16(5):451-7 PMID: 10871267
  9. Practical limits of function prediction.
    Proteins. 2000 Oct 1;41(1):98-107 PMID: 10944397
  10. CAST: an iterative algorithm for the complexity analysis of sequence tracts. Complexity analysis of sequence tracts.
    Bioinformatics. 2000 Oct;16(10):915-22 PMID: 11120681
  11. The COG database: new developments in phylogenetic classification of proteins from complete genomes.
    Nucleic Acids Res. 2001 Jan 1;29(1):22-8 PMID: 11125040
  12. CluSTr: a database of clusters of SWISS-PROT+TrEMBL proteins.
    Nucleic Acids Res. 2001 Jan 1;29(1):33-6 PMID: 11125042
  13. Mining the draft human genome.
    Nature. 2001 Feb 15;409(6822):827-8 PMID: 11236999
  14. The Pfam protein families database.
    Nucleic Acids Res. 2002 Jan 1;30(1):276-80 PMID: 11752314
  15. SYSTERS, GeneNest, SpliceNest: exploring sequence space from genome to protein.
    Nucleic Acids Res. 2002 Jan 1;30(1):299-300 PMID: 11752319
  16. An efficient algorithm for large-scale detection of protein families.
    Nucleic Acids Res. 2002 Apr 1;30(7):1575-84 PMID: 11917018
  17. Studying genomes through the aeons: protein families, pseudogenes and proteome evolution.
    J Mol Biol. 2002 May 17;318(5):1155-74 PMID: 12083509
  18. Domains, motifs and clusters in the protein universe.
    Curr Opin Chem Biol. 2003 Feb;7(1):5-11 PMID: 12547420
  19. Myriads of protein families, and still counting.
    Genome Biol. 2003;4(2):401 PMID: 12620116
  20. GeneTRACE-reconstruction of gene content of ancestral species.
    Bioinformatics. 2003 Jul 22;19(11):1412-6 PMID: 12874054
  21. COmplete GENome Tracking (COGENT): a flexible data environment for computational genomics.
    Bioinformatics. 2003 Jul 22;19(11):1451-2 PMID: 12874064
  22. Similar amino acid sequences: chance or common ancestry?
    Science. 1981 Oct 9;214(4517):149-59 PMID: 7280687
  23. The relation between the divergence of sequence and structure in proteins.
    EMBO J. 1986 Apr;5(4):823-6 PMID: 3709526
  24. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  25. A genomic perspective on protein families.
    Science. 1997 Oct 24;278(5338):631-7 PMID: 9381173
  26. Automated genome sequence analysis and annotation.
    Bioinformatics. 1999 May;15(5):391-412 PMID: 10366660
  27. Global transposon mutagenesis and a minimal Mycoplasma genome.
    Science. 1999 Dec 10;286(5447):2165-9 PMID: 10591650
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2003-08-01
Pages
4632-8
Language
English
Region
England
NLM ID
0411011
PMCID
PMC169885
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com