Home LiteratureArticle Details
PMID: 16450363 Published · ppublish English Comparative Study Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S. Research Support, U.S. Gov't, P.H.S. Validation Study

Evaluation of different biological data and computational classification methods for use in protein interaction prediction.

Proteins ·Vol. 63 ·No. 3 ·2006-05-15 ·Pages 490-500

Qi Y, Bar-Joseph Z, Klein-Seetharaman J

Abstract

Protein-protein interactions play a key role in many biological systems. High-throughput methods can directly detect the set of interacting proteins in yeast, but the results are often incomplete and exhibit high false-positive and false-negative rates. Recently, many different research groups independently suggested using supervised learning methods to integrate direct and indirect biological data sources for the protein interaction prediction task. However, the data sources, approaches, and implementations varied. Furthermore, the protein interaction prediction task itself can be subdivided into prediction of (1) physical interaction, (2) co-complex relationship, and (3) pathway co-membership. To investigate systematically the utility of different data sources and the way the data is encoded as features for predicting each of these types of protein interactions, we assembled a large set of biological features and varied their encoding for use in each of the three prediction tasks. Six different classifiers were used to assess the accuracy in predicting interactions, Random Forest (RF), RF similarity-based k-Nearest-Neighbor, Naïve Bayes, Decision Tree, Logistic Regression, and Support Vector Machine. For all classifiers, the three prediction tasks had different success rates, and co-complex prediction appears to be an easier task than the other two. Independently of prediction task, however, the RF classifier consistently ranked as one of the top two classifiers for all combinations of feature sets. Therefore, we used this classifier to study the importance of different biological datasets. First, we used the splitting function of the RF tree structure, the Gini index, to estimate feature importance. Second, we determined classification accuracy when only the top-ranking features were used as an input in the classifier. We find that the importance of different features depends on the specific prediction task and the way they are encoded. Strikingly, gene expression is consistently the most important feature for all three prediction tasks, while the protein interactions identified using the yeast-2-hybrid system were not among the top-ranking features under any condition.

MeSH Terms
Computational Biology/classification,methods Databases, Protein/classification Forecasting Protein Interaction Mapping/classification,methods
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Qi Yanjun
School of Computer Science, Carnegie Mellon University, Pittsburgh, Pennsylvania 15213, USA.
Bar-Joseph Ziv
Klein-Seetharaman Judith
References (26)
26 references, click to expand
  1. Computational methods of analysis of protein-protein interactions.
    Curr Opin Struct Biol. 2003 Jun;13(3):377-82 PMID: 12831890
  2. How reliable are experimental protein-protein interaction data?
    J Mol Biol. 2003 Apr 11;327(5):919-23 PMID: 12662919
  3. A Bayesian networks approach for predicting protein-protein interactions from genomic data.
    Science. 2003 Oct 17;302(5644):449-53 PMID: 14564010
  4. Computational discovery of gene modules and regulatory networks.
    Nat Biotechnol. 2003 Nov;21(11):1337-42 PMID: 14555958
  5. MIPS: analysis and annotation of proteins from whole genomes.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D41-4 PMID: 14681354
  6. Saccharomyces Genome Database (SGD) provides tools to identify and analyze sequences from Saccharomyces cerevisiae and related sequences from other organisms.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D311-4 PMID: 14681421
  7. Gaining confidence in high-throughput protein interaction networks.
    Nat Biotechnol. 2004 Jan;22(1):78-85 PMID: 14704708
  8. Global mapping of the yeast genetic interaction network.
    Science. 2004 Feb 6;303(5659):808-13 PMID: 14764870
  9. A statistical framework for combining and interpreting proteomic datasets.
    Bioinformatics. 2004 Mar 22;20(5):689-700 PMID: 15033876
  10. Predicting co-complexed protein pairs using genomic and proteomic data integration.
    BMC Bioinformatics. 2004 Apr 16;5:38 PMID: 15090078
  11. Protein network inference from multiple genomic data: a supervised approach.
    Bioinformatics. 2004 Aug 4;20 Suppl 1:i363-70 PMID: 15262821
  12. Transcriptional regulatory code of a eukaryotic genome.
    Nature. 2004 Sep 2;431(7004):99-104 PMID: 15343339
  13. Information assessment on predicting protein-protein interactions.
    BMC Bioinformatics. 2004 Oct 18;5:154 PMID: 15491499
  14. A probabilistic functional network of yeast genes.
    Science. 2004 Nov 26;306(5701):1555-8 PMID: 15567862
  15. Random forest similarity for protein-protein interaction prediction from multiple sources.
    Pac Symp Biocomput. 2005;:531-42 PMID: 15759657
  16. KEGG: kyoto encyclopedia of genes and genomes.
    Nucleic Acids Res. 2000 Jan 1;28(1):27-30 PMID: 10592173
  17. A comprehensive analysis of protein-protein interactions in Saccharomyces cerevisiae.
    Nature. 2000 Feb 10;403(6770):623-7 PMID: 10688190
  18. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  19. A comprehensive two-hybrid analysis to explore the yeast protein interactome.
    Proc Natl Acad Sci U S A. 2001 Apr 10;98(8):4569-74 PMID: 11283351
  20. DIP, the Database of Interacting Proteins: a research tool for studying cellular networks of protein interactions.
    Nucleic Acids Res. 2002 Jan 1;30(1):303-5 PMID: 11752321
  21. Functional organization of the yeast proteome by systematic analysis of protein complexes.
    Nature. 2002 Jan 10;415(6868):141-7 PMID: 11805826
  22. Systematic identification of protein complexes in Saccharomyces cerevisiae by mass spectrometry.
    Nature. 2002 Jan 10;415(6868):180-3 PMID: 11805837
  23. Comparative assessment of large-scale data sets of protein-protein interactions.
    Nature. 2002 May 23;417(6887):399-403 PMID: 12000970
  24. Analyzing yeast protein-protein interaction data obtained from different sources.
    Nat Biotechnol. 2002 Oct;20(10):991-7 PMID: 12355115
  25. Inferring domain-domain interactions from protein-protein interactions.
    Genome Res. 2002 Oct;12(10):1540-8 PMID: 12368246
  26. Global analysis of protein expression in yeast.
    Nature. 2003 Oct 16;425(6959):737-41 PMID: 14562106
Article Info
Journal
Proteins
Abbr.
Proteins
ISSN
1097-0134
Published
2006-05-15
Pages
490-500
Language
English
Region
United States
NLM ID
8700181
PMCID
PMC3250929
Subset
IM
Grants
NLM NIH HHS · R01 LM007994 · United States
NLM NIH HHS · R01 LM007994-01A1 · United States
PHS HHS · NLM108730 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com