Home LiteratureArticle Details
PMID: 12952885 Published · ppublish English Comparative Study Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, P.H.S.

OrthoMCL: identification of ortholog groups for eukaryotic genomes.

Genome research ·Vol. 13 ·No. 9 ·2003-09-00 ·Pages 2178-89

Li L, Stoeckert CJ, Roos DS

Abstract

The identification of orthologous groups is useful for genome annotation, studies on gene/protein evolution, comparative genomics, and the identification of taxonomically restricted sequences. Methods successfully exploited for prokaryotic genome analysis have proved difficult to apply to eukaryotes, however, as larger genomes may contain multiple paralogous genes, and sequence information is often incomplete. OrthoMCL provides a scalable method for constructing orthologous groups across multiple eukaryotic taxa, using a Markov Cluster algorithm to group (putative) orthologs and paralogs. This method performs similarly to the INPARANOID algorithm when applied to two genomes, but can be extended to cluster orthologs from multiple species. OrthoMCL clusters are coherent with groups identified by EGO, but improved recognition of "recent" paralogs permits overlapping EGO groups representing the same gene to be merged. Comparison with previously assigned EC annotations suggests a high degree of reliability, implying utility for automated eukaryotic genome annotation. OrthoMCL has been applied to the proteome data set from seven publicly available genomes (human, fly, worm, yeast, Arabidopsis, the malaria parasite Plasmodium falciparum, and Escherichia coli). A Web interface allows queries based on individual genes or user-defined phylogenetic patterns (http://www.cbil.upenn.edu/gene-family). Analysis of clusters incorporating P. falciparum genes identifies numerous enzymes that were incompletely annotated in first-pass annotation of the parasite genome.

MeSH Terms
Animals Arabidopsis/genetics Caenorhabditis elegans/genetics Computational Biology/methods Drosophila melanogaster/genetics Eukaryotic Cells/chemistry,metabolism Genome Genome, Fungal Genome, Plant Genome, Protozoan Humans Internet Plasmodium falciparum/genetics Saccharomyces cerevisiae/genetics Sequence Homology, Nucleic Acid Software
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Li Li
Department of Biology and Genetics, Center for Bioinformatics, and Genomics Institute, University of Pennsylvania, Philadelphia, Pennsylvania 19104, USA.
Stoeckert Christian J
Roos David S
References (33)
33 references, click to expand
  1. Human and nematode orthologs--lessons from the analysis of 1800 human genes and the proteome of Caenorhabditis elegans.
    Gene. 1999 Sep 30;238(1):163-70 PMID: 10570994
  2. Comparison of the complete protein sets of worm and yeast: orthology and divergence.
    Science. 1998 Dec 11;282(5396):2022-8 PMID: 9851918
  3. The COG database: a tool for genome-scale analysis of protein functions and evolution.
    Nucleic Acids Res. 2000 Jan 1;28(1):33-6 PMID: 10592175
  4. The TIGR gene indices: reconstruction and representation of expressed gene sequences.
    Nucleic Acids Res. 2000 Jan 1;28(1):141-5 PMID: 10592205
  5. Searching for drug targets in microbial genomes.
    Curr Opin Biotechnol. 1999 Dec;10(6):571-8 PMID: 10600691
  6. Comparative genomics of the eukaryotes.
    Science. 2000 Mar 24;287(5461):2204-15 PMID: 10731134
  7. Homology a personal view on some of the problems.
    Trends Genet. 2000 May;16(5):227-31 PMID: 10782117
  8. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  9. The COG database: new developments in phylogenetic classification of proteins from complete genomes.
    Nucleic Acids Res. 2001 Jan 1;29(1):22-8 PMID: 11125040
  10. The TIGR Gene Indices: analysis of gene transcript sequences in highly sampled eukaryotic species.
    Nucleic Acids Res. 2001 Jan 1;29(1):159-64 PMID: 11125077
  11. Using the COG database to improve gene recognition in complete genomes.
    Genetica. 2000;108(1):9-17 PMID: 11145426
  12. A plastid segregation defect in the protozoan parasite Toxoplasma gondii.
    EMBO J. 2001 Feb 1;20(3):330-9 PMID: 11157740
  13. Towards understanding the first genome sequence of a crenarchaeon by genome annotation using clusters of orthologous groups of proteins (COGs).
    Genome Biol. 2000;1(5):RESEARCH0009 PMID: 11178258
  14. Creating the gene ontology resource: design and implementation.
    Genome Res. 2001 Aug;11(8):1425-33 PMID: 11483584
  15. Automatic clustering of orthologs and in-paralogs from pairwise species comparisons.
    J Mol Biol. 2001 Dec 14;314(5):1041-52 PMID: 11743721
  16. Cross-referencing eukaryotic genomes: TIGR Orthologous Gene Alignments (TOGA).
    Genome Res. 2002 Mar;12(3):493-502 PMID: 11875039
  17. An efficient algorithm for large-scale detection of protein families.
    Nucleic Acids Res. 2002 Apr 1;30(7):1575-84 PMID: 11917018
  18. Predicting gene ontology functions from ProDom and CDD protein domains.
    Genome Res. 2002 Apr;12(4):648-55 PMID: 11932249
  19. A hot story from comparative genomics: reverse gyrase is the only hyperthermophile-specific protein.
    Trends Genet. 2002 May;18(5):236-7 PMID: 12047940
  20. Clustering of proximal sequence space for the identification of protein families.
    Bioinformatics. 2002 Jul;18(7):908-21 PMID: 12117788
  21. The Plasmodium genome database.
    Nature. 2002 Oct 3;419(6906):490-2 PMID: 12368860
  22. Genome sequence of the human malaria parasite Plasmodium falciparum.
    Nature. 2002 Oct 3;419(6906):498-511 PMID: 12368864
  23. Genome sequence and comparative analysis of the model rodent malaria parasite Plasmodium yoelii yoelii.
    Nature. 2002 Oct 3;419(6906):512-9 PMID: 12368865
  24. PlasmoDB: the Plasmodium genome resource. A database integrating experimental and computational data.
    Nucleic Acids Res. 2003 Jan 1;31(1):212-5 PMID: 12519984
  25. Distinguishing homologous from analogous proteins.
    Syst Zool. 1970 Jun;19(2):99-113 PMID: 5449325
  26. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  27. The multiplicity of domains in proteins.
    Annu Rev Biochem. 1995;64:287-314 PMID: 7574483
  28. A plastid of probable green algal origin in Apicomplexan parasites.
    Science. 1997 Mar 7;275(5305):1485-9 PMID: 9045615
  29. Gene families: the taxonomy of protein paralogs and chimeras.
    Science. 1997 Oct 24;278(5338):609-14 PMID: 9381171
  30. A genomic perspective on protein families.
    Science. 1997 Oct 24;278(5338):631-7 PMID: 9381173
  31. A plastid organelle as a drug target in apicomplexan parasites.
    Nature. 1997 Nov 27;390(6658):407-9 PMID: 9389481
  32. Large-scale taxonomic profiling of eukaryotic model organisms: a comparison of orthologous proteins encoded by the human, fly, nematode, and yeast genomes.
    Genome Res. 1998 Jun;8(6):590-8 PMID: 9647634
  33. The apicoplast as a potential therapeutic target in Toxoplasma and other apicomplexan parasites: some additional thoughts.
    Parasitol Today. 1999 Jan;15(1):41 PMID: 10234180
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
2003-09-00
Pages
2178-89
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC403725
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com