Home LiteratureArticle Details
PMID: 11917018 Published · ppublish English Journal Article

An efficient algorithm for large-scale detection of protein families.

Nucleic acids research ·Vol. 30 ·No. 7 ·2002-04-01 ·Pages 1575-84

Enright AJ, Van Dongen S, Ouzounis CA

Abstract

Detection of protein families in large databases is one of the principal research objectives in structural and functional genomics. Protein family classification can significantly contribute to the delineation of functional diversity of homologous proteins, the prediction of function based on domain architecture or the presence of sequence motifs as well as comparative genomics, providing valuable evolutionary insights. We present a novel approach called TRIBE-MCL for rapid and accurate clustering of protein sequences into families. The method relies on the Markov cluster (MCL) algorithm for the assignment of proteins into families based on precomputed sequence similarity information. This novel approach does not suffer from the problems that normally hinder other protein sequence clustering algorithms, such as the presence of multi-domain proteins, promiscuous domains and fragmented proteins. The method has been rigorously tested and validated on a number of very large databases, including SwissProt, InterPro, SCOP and the draft human genome. Our results indicate that the method is ideally suited to the rapid and accurate detection of protein families on a large scale. The method has been used to detect and categorise protein families within the draft human genome and the resulting families have been used to annotate a large proportion of human proteins.

MeSH Terms
Algorithms Amino Acid Sequence Databases, Protein Genome, Human Humans Internet Molecular Sequence Data Proteins/genetics Sequence Alignment Sequence Homology, Amino Acid Transcription Factor TFIIB Transcription Factors/genetics
Chemicals
Proteins Transcription Factor TFIIB Transcription Factors
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Enright A J
Computational Genomics Group, The European Bioinformatics Institute, EMBL Cambridge Outstation, Cambridge CB10 1SD, UK. anton@ebi.ac.uk
Van Dongen S
Ouzounis C A
References (47)
47 references, click to expand
  1. An insight into domain combinations.
    Bioinformatics. 2001;17 Suppl 1:S83-9 PMID: 11472996
  2. Domain combinations in archaeal, eubacterial and eukaryotic proteomes.
    J Mol Biol. 2001 Jul 6;310(2):311-25 PMID: 11428892
  3. Strain-specific genes of Helicobacter pylori: distribution, function and dynamics.
    Nucleic Acids Res. 2001 Nov 1;29(21):4395-404 PMID: 11691927
  4. Aspects of molecular evolution.
    Annu Rev Genet. 1973;7:343-80 PMID: 4593308
  5. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  6. Evolutionarily mobile modules in proteins.
    Sci Am. 1993 Oct;269(4):50-6 PMID: 8235550
  7. Eukaryotes have "two-component" signal transducers.
    Res Microbiol. 1994 Jun-Aug;145(5-6):481-6 PMID: 7855435
  8. The multiplicity of domains in proteins.
    Annu Rev Biochem. 1995;64:287-314 PMID: 7574483
  9. The emergence of major cellular processes in evolution.
    FEBS Lett. 1996 Jul 22;390(2):119-23 PMID: 8706840
  10. Hidden Markov models.
    Curr Opin Struct Biol. 1996 Jun;6(3):361-5 PMID: 8804822
  11. Computational comparisons of model genomes.
    Trends Biotechnol. 1996 Aug;14(8):280-5 PMID: 8987458
  12. On the classification and evolution of protein modules.
    J Protein Chem. 1997 Jul;16(5):545-51 PMID: 9246642
  13. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  14. A cyanobacterial phytochrome two-component light sensory system.
    Science. 1997 Sep 5;277(5331):1505-8 PMID: 9278513
  15. Domain identification by clustering sequence alignments.
    Proc Int Conf Intell Syst Mol Biol. 1997;5:124-30 PMID: 9322026
  16. Gene families: the taxonomy of protein paralogs and chimeras.
    Science. 1997 Oct 24;278(5338):609-14 PMID: 9381171
  17. The challenges of genome sequence annotation or "the devil is in the details".
    Nat Biotechnol. 1997 Nov;15(12):1222-3 PMID: 9359093
  18. The ProDom database of protein domain families.
    Nucleic Acids Res. 1998 Jan 1;26(1):323-6 PMID: 9399865
  19. Eukaryotic transcription factors.
    Curr Opin Struct Biol. 1998 Feb;8(1):41-8 PMID: 9519295
  20. Predicting function: from genes to genomes and back.
    J Mol Biol. 1998 Nov 6;283(4):707-25 PMID: 9790834
  21. The PROSITE database, its status in 1999.
    Nucleic Acids Res. 1999 Jan 1;27(1):215-9 PMID: 9847184
  22. PRINTS prepares for the new millennium.
    Nucleic Acids Res. 1999 Jan 1;27(1):220-5 PMID: 9847185
  23. Profile hidden Markov models.
    Bioinformatics. 1998;14(9):755-63 PMID: 9918945
  24. The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.
    J Mol Biol. 1999 Apr 23;288(1):147-64 PMID: 10329133
  25. Detecting protein function and protein-protein interactions from genome sequences.
    Science. 1999 Jul 30;285(5428):751-3 PMID: 10427000
  26. The origin and evolution of protein superfamilies.
    Fed Proc. 1976 Aug;35(10):2132-8 PMID: 181273
  27. Protein interaction maps for complete genomes based on gene fusion events.
    Nature. 1999 Nov 4;402(6757):86-90 PMID: 10573422
  28. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  29. SCOP: a structural classification of proteins database.
    Nucleic Acids Res. 2000 Jan 1;28(1):257-9 PMID: 10592240
  30. The Pfam protein families database.
    Nucleic Acids Res. 2000 Jan 1;28(1):263-6 PMID: 10592242
  31. Comparative genomics of the eukaryotes.
    Science. 2000 Mar 24;287(5461):2204-15 PMID: 10731134
  32. Global properties of the metabolic map of Escherichia coli.
    Genome Res. 2000 Apr;10(4):568-76 PMID: 10779499
  33. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  34. Protein function in the post-genomic era.
    Nature. 2000 Jun 15;405(6788):823-6 PMID: 10866208
  35. GeneRAGE: a robust algorithm for sequence clustering and domain detection.
    Bioinformatics. 2000 May;16(5):451-7 PMID: 10871267
  36. Recent developments and future directions in computational genomics.
    FEBS Lett. 2000 Aug 25;480(1):42-8 PMID: 10967327
  37. Two-component signal transduction.
    Annu Rev Biochem. 2000;69:183-215 PMID: 10966457
  38. Towards a covering set of protein family profiles.
    Prog Biophys Mol Biol. 2000;73(5):321-37 PMID: 11063778
  39. CAST: an iterative algorithm for the complexity analysis of sequence tracts. Complexity analysis of sequence tracts.
    Bioinformatics. 2000 Oct;16(10):915-22 PMID: 11120681
  40. The COG database: new developments in phylogenetic classification of proteins from complete genomes.
    Nucleic Acids Res. 2001 Jan 1;29(1):22-8 PMID: 11125040
  41. The InterPro database, an integrated documentation resource for protein families, domains and functional sites.
    Nucleic Acids Res. 2001 Jan 1;29(1):37-40 PMID: 11125043
  42. Genomes OnLine Database (GOLD): a monitor of genome projects world-wide.
    Nucleic Acids Res. 2001 Jan 1;29(1):126-7 PMID: 11125068
  43. Genome sequences and great expectations.
    Genome Biol. 2001;2(1):INTERACTIONS0001 PMID: 11178275
  44. Transcription-associated protein families are primarily taxon-specific.
    Bioinformatics. 2001 Jan;17(1):95-7 PMID: 11222266
  45. Mining the draft human genome.
    Nature. 2001 Feb 15;409(6822):827-8 PMID: 11236999
  46. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  47. BioLayout--an automatic graph layout algorithm for similarity visualization.
    Bioinformatics. 2001 Sep;17(9):853-4 PMID: 11590107
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2002-04-01
Pages
1575-84
Language
English
Region
England
NLM ID
0411011
PMCID
PMC101833
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com