Home LiteratureArticle Details
PMID: 17708678 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

Automated protein subfamily identification and classification.

PLoS computational biology ·Vol. 3 ·No. 8 ·2007-08-00 ·Pages e160

Brown DP, Krishnamurthy N, Sjölander K

Abstract

Function prediction by homology is widely used to provide preliminary functional annotations for genes for which experimental evidence of function is unavailable or limited. This approach has been shown to be prone to systematic error, including percolation of annotation errors through sequence databases. Phylogenomic analysis avoids these errors in function prediction but has been difficult to automate for high-throughput application. To address this limitation, we present a computationally efficient pipeline for phylogenomic classification of proteins. This pipeline uses the SCI-PHY (Subfamily Classification in Phylogenomics) algorithm for automatic subfamily identification, followed by subfamily hidden Markov model (HMM) construction. A simple and computationally efficient scoring scheme using family and subfamily HMMs enables classification of novel sequences to protein families and subfamilies. Sequences representing entirely novel subfamilies are differentiated from those that can be classified to subfamilies in the input training set using logistic regression. Subfamily HMM parameters are estimated using an information-sharing protocol, enabling subfamilies containing even a single sequence to benefit from conservation patterns defining the family as a whole or in related subfamilies. SCI-PHY subfamilies correspond closely to functional subtypes defined by experts and to conserved clades found by phylogenetic analysis. Extensive comparisons of subfamily and family HMM performances show that subfamily HMMs dramatically improve the separation between homologous and non-homologous proteins in sequence database searches. Subfamily HMMs also provide extremely high specificity of classification and can be used to predict entirely novel subtypes. The SCI-PHY Web server at http://phylogenomics.berkeley.edu/SCI-PHY/ allows users to upload a multiple sequence alignment for subfamily identification and subfamily HMM construction. Biologists wishing to provide their own subfamily definitions can do so. Source code is available on the Web page. The Berkeley Phylogenomics Group PhyloFacts resource contains pre-calculated subfamily predictions and subfamily HMMs for more than 40,000 protein families and domains at http://phylogenomics.berkeley.edu/phylofacts/.

MeSH Terms
Algorithms Amino Acid Sequence Artificial Intelligence Markov Chains Molecular Sequence Data Pattern Recognition, Automated/methods Proteins/chemistry,classification Reproducibility of Results Sensitivity and Specificity Sequence Alignment/methods Sequence Analysis, Protein/methods
Chemicals
Proteins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Brown Duncan P
Department of Bioengineering, University of California, Berkeley, California, United States of America.
Krishnamurthy Nandini
Sjölander Kimmen
References (64)
64 references, click to expand
  1. Detection of conserved segments in proteins: iterative scanning of sequence databases with alignment blocks.
    Proc Natl Acad Sci U S A. 1994 Dec 6;91(25):12091-5 PMID: 7991589
  2. FIGENIX: intelligent automation of genomic annotation: expertise integration in a new software platform.
    BMC Bioinformatics. 2005 Aug 05;6:198 PMID: 16083500
  3. Combining local-structure, fold-recognition, and new fold methods for protein structure prediction.
    Proteins. 2003;53 Suppl 6:491-6 PMID: 14579338
  4. Hidden Markov models for detecting remote protein homologies.
    Bioinformatics. 1998;14(10):846-56 PMID: 9927713
  5. Practical limits of function prediction.
    Proteins. 2000 Oct 1;41(1):98-107 PMID: 10944397
  6. Automated genome sequence analysis and annotation.
    Bioinformatics. 1999 May;15(5):391-412 PMID: 10366660
  7. Local alignment statistics.
    Methods Enzymol. 1996;266:460-80 PMID: 8743700
  8. Percolation of annotation errors through hierarchically structured protein sequence databases.
    Math Biosci. 2005 Feb;193(2):223-34 PMID: 15748731
  9. Sequence comparisons using multiple sequences detect three times as many remote homologues as pairwise methods.
    J Mol Biol. 1998 Dec 11;284(4):1201-10 PMID: 9837738
  10. RIO: analyzing proteomes by automated phylogenomics using resampled inference of orthologs.
    BMC Bioinformatics. 2002 May 16;3:14 PMID: 12028595
  11. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  12. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.
    Bioinformatics. 2006 Jul 1;22(13):1658-9 PMID: 16731699
  13. CDD: a Conserved Domain Database for protein classification.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D192-6 PMID: 15608175
  14. Position-based sequence weights.
    J Mol Biol. 1994 Nov 4;243(4):574-8 PMID: 7966282
  15. PipeAlign: A new toolkit for protein family analysis.
    Nucleic Acids Res. 2003 Jul 1;31(13):3829-32 PMID: 12824430
  16. Automatic clustering of orthologs and in-paralogs from pairwise species comparisons.
    J Mol Biol. 2001 Dec 14;314(5):1041-52 PMID: 11743721
  17. Environmental genome shotgun sequencing of the Sargasso Sea.
    Science. 2004 Apr 2;304(5667):66-74 PMID: 15001713
  18. The ASTRAL compendium for protein structure and sequence analysis.
    Nucleic Acids Res. 2000 Jan 1;28(1):254-6 PMID: 10592239
  19. The Pfam protein families database.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D138-41 PMID: 14681378
  20. The closest BLAST hit is often not the nearest neighbor.
    J Mol Evol. 2001 Jun;52(6):540-2 PMID: 11443357
  21. Evolution of the SNF2 family of proteins: subfamilies with distinct sequences and functions.
    Nucleic Acids Res. 1995 Jul 25;23(14):2715-23 PMID: 7651832
  22. Collecting and harvesting biological data: the GPCRDB and NucleaRDB information systems.
    Nucleic Acids Res. 2001 Jan 1;29(1):346-9 PMID: 11125133
  23. Semi-supervised protein classification using cluster kernels.
    Bioinformatics. 2005 Aug 1;21(15):3241-7 PMID: 15905279
  24. Subfamily hmms in functional genomics.
    Pac Symp Biocomput. 2005;:322-33 PMID: 15759638
  25. Protein molecular function prediction by Bayesian phylogenomics.
    PLoS Comput Biol. 2005 Oct;1(5):e45 PMID: 16217548
  26. GoFigure: automated Gene Ontology annotation.
    Bioinformatics. 2003 Dec 12;19(18):2484-5 PMID: 14668239
  27. Analysis and prediction of functional sub-types from protein sequence alignments.
    J Mol Biol. 2000 Oct 13;303(1):61-76 PMID: 11021970
  28. Intrinsic errors in genome annotation.
    Trends Genet. 2001 Aug;17(8):429-31 PMID: 11485799
  29. OntoBlast function: From sequence similarities directly to potential functional annotations by ontology terms.
    Nucleic Acids Res. 2003 Jul 1;31(13):3799-803 PMID: 12824422
  30. Modulation of pulmonary innate immunity during bacterial infection: animal studies.
    Arch Immunol Ther Exp (Warsz). 2002;50(3):159-67 PMID: 12098931
  31. The prediction of protein function at CASP6.
    Proteins. 2005;61 Suppl 7:201-13 PMID: 16187363
  32. Phylogenomic inference of protein molecular function: advances and challenges.
    Bioinformatics. 2004 Jan 22;20(2):170-9 PMID: 14734307
  33. GOblet: a platform for Gene Ontology annotation of anonymous sequence data.
    Nucleic Acids Res. 2004 Jul 1;32(Web Server issue):W313-7 PMID: 15215401
  34. Clustering of proximal sequence space for the identification of protein families.
    Bioinformatics. 2002 Jul;18(7):908-21 PMID: 12117788
  35. Leveraging enzyme structure-function relationships for functional inference and experimental design: the structure-function linkage database.
    Biochemistry. 2006 Feb 28;45(8):2545-55 PMID: 16489747
  36. Metagenomics: genomic analysis of microbial communities.
    Annu Rev Genet. 2004;38:525-52 PMID: 15568985
  37. Errors in genome annotation.
    Trends Genet. 1999 Apr;15(4):132-3 PMID: 10203816
  38. PhyloGenie: automated phylome generation and analysis.
    Nucleic Acids Res. 2004 Sep 30;32(17):5231-8 PMID: 15459293
  39. Tolerating some redundancy significantly speeds up clustering of large protein databases.
    Bioinformatics. 2002 Jan;18(1):77-82 PMID: 11836214
  40. The ASTRAL Compendium in 2004.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D189-92 PMID: 14681391
  41. Berkeley Phylogenomics Group web servers: resources for structural phylogenomic analysis.
    Nucleic Acids Res. 2007 Jul;35(Web Server issue):W27-32 PMID: 17488835
  42. Sources of systematic error in functional annotation of genomes: domain rearrangement, non-orthologous gene displacement and operon disruption.
    In Silico Biol. 1998;1(1):55-67 PMID: 11471243
  43. Profile analysis: detection of distantly related proteins.
    Proc Natl Acad Sci U S A. 1987 Jul;84(13):4355-8 PMID: 3474607
  44. GPCRDB information system for G protein-coupled receptors.
    Nucleic Acids Res. 2003 Jan 1;31(1):294-7 PMID: 12520006
  45. Automatic annotation of protein function based on family identification.
    Proteins. 2003 Nov 15;53(3):683-92 PMID: 14579359
  46. PhyloFacts: an online structural phylogenomic encyclopedia for protein functional and structural classification.
    Genome Biol. 2006;7(9):R83 PMID: 16973001
  47. Hidden Markov models in computational biology. Applications to protein modeling.
    J Mol Biol. 1994 Feb 4;235(5):1501-31 PMID: 8107089
  48. Isolation and characterization of acetoacetyl-CoA thiolase gene essential for n-decane assimilation in yeast Yarrowia lipolytica.
    Biochem Biophys Res Commun. 2001 Apr 6;282(3):832-8 PMID: 11401539
  49. UniProt: the Universal Protein knowledgebase.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D115-9 PMID: 14681372
  50. Functional classification using phylogenomic inference.
    PLoS Comput Biol. 2006 Jun 30;2(6):e77 PMID: 16846248
  51. OrthoMCL: identification of ortholog groups for eukaryotic genomes.
    Genome Res. 2003 Sep;13(9):2178-89 PMID: 12952885
  52. CASP and CAFASP experiments and their findings.
    Methods Biochem Anal. 2003;44:501-7 PMID: 12647401
  53. Dirichlet mixtures: a method for improved detection of weak but significant protein sequence homology.
    Comput Appl Biosci. 1996 Aug;12(4):327-45 PMID: 8902360
  54. Evolution of protein function, from a structural perspective.
    Curr Opin Chem Biol. 1999 Oct;3(5):548-56 PMID: 10508675
  55. The sequence of the human genome.
    Science. 2001 Feb 16;291(5507):1304-51 PMID: 11181995
  56. Profile hidden Markov models.
    Bioinformatics. 1998;14(9):755-63 PMID: 9918945
  57. GOtcha: a new method for prediction of protein function assessed by the annotation of seven genomes.
    BMC Bioinformatics. 2004 Nov 18;5:178 PMID: 15550167
  58. Classifying G-protein coupled receptors with support vector machines.
    Bioinformatics. 2002 Jan;18(1):147-59 PMID: 11836223
  59. Phylogenomics: improving functional predictions for uncharacterized genes by evolutionary analysis.
    Genome Res. 1998 Mar;8(3):163-7 PMID: 9521918
  60. Automated ortholog inference from phylogenetic trees and calculation of orthology reliability.
    Bioinformatics. 2002 Jan;18(1):92-9 PMID: 11836216
  61. A phylogenomic study of the MutS family of proteins.
    Nucleic Acids Res. 1998 Sep 15;26(18):4291-300 PMID: 9722651
  62. Secator: a program for inferring protein subfamilies from phylogenetic trees.
    Mol Biol Evol. 2001 Aug;18(8):1435-41 PMID: 11470834
  63. MUSCLE: multiple sequence alignment with high accuracy and high throughput.
    Nucleic Acids Res. 2004 Mar 19;32(5):1792-7 PMID: 15034147
  64. Automated protein function prediction--the genomic challenge.
    Brief Bioinform. 2006 Sep;7(3):225-42 PMID: 16772267
Article Info
Journal
PLoS computational biology
Abbr.
PLoS Comput Biol
ISSN
1553-7358
Published
2007-08-00
Pages
e160
Language
English
Region
United States
NLM ID
101238922
PMCID
PMC1950344
Subset
IM
Grants
NHGRI NIH HHS · R01 HG002769 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com