Home LiteratureArticle Details
PMID: 18421371 Published · epublish English Journal Article Research Support, N.I.H., Extramural

Predicting co-complexed protein pairs from heterogeneous data.

PLoS computational biology ·Vol. 4 ·No. 4 ·2008-04-18 ·Pages e1000054

Qiu J, Noble WS

Abstract

Proteins do not carry out their functions alone. Instead, they often act by participating in macromolecular complexes and play different functional roles depending on the other members of the complex. It is therefore interesting to identify co-complex relationships. Although protein complexes can be identified in a high-throughput manner by experimental technologies such as affinity purification coupled with mass spectrometry (APMS), these large-scale datasets often suffer from high false positive and false negative rates. Here, we present a computational method that predicts co-complexed protein pair (CCPP) relationships using kernel methods from heterogeneous data sources. We show that a diffusion kernel based on random walks on the full network topology yields good performance in predicting CCPPs from protein interaction networks. In the setting of direct ranking, a diffusion kernel performs much better than the mutual clustering coefficient. In the setting of SVM classifiers, a diffusion kernel performs much better than a linear kernel. We also show that combination of complementary information improves the performance of our CCPP recognizer. A summation of three diffusion kernels based on two-hybrid, APMS, and genetic interaction networks and three sequence kernels achieves better performance than the sequence kernels or diffusion kernels alone. Inclusion of additional features achieves a still better ROC(50) of 0.937. Assuming a negative-to-positive ratio of 600ratio1, the final classifier achieves 89.3% coverage at an estimated false discovery rate of 10%. Finally, we applied our prediction method to two recently described APMS datasets. We find that our predicted positives are highly enriched with CCPPs that are identified by both datasets, suggesting that our method successfully identifies true CCPPs. An SVM classifier trained from heterogeneous data sources provides accurate predictions of CCPPs in yeast. This computational method thereby provides an inexpensive method for identifying protein complexes that extends and complements high-throughput experimental data.

MeSH Terms
Amino Acid Sequence Binding Sites Computer Simulation Databases, Protein Information Storage and Retrieval/methods Models, Chemical Models, Molecular Molecular Sequence Data Multiprotein Complexes/chemistry,ultrastructure Protein Binding Protein Conformation Protein Interaction Mapping/methods Proteins/chemistry,ultrastructure Sequence Analysis, Protein
Chemicals
Multiprotein Complexes Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Qiu Jian
Department of Genome Sciences, University of Washington, Seattle, Washington, United States of America.
Noble William Stafford
References (49)
49 references, click to expand
  1. Global analysis of protein localization in budding yeast.
    Nature. 2003 Oct 16;425(6959):686-91 PMID: 14562095
  2. The Yeast Resource Center Public Data Repository.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D378-82 PMID: 15608220
  3. A high resolution protein interaction map of the yeast Mediator complex.
    Nucleic Acids Res. 2004 Oct 11;32(18):5379-91 PMID: 15477388
  4. Purification of the yeast U4/U6.U5 small nuclear ribonucleoprotein particle and identification of its proteins.
    Proc Natl Acad Sci U S A. 1999 Jun 22;96(13):7226-31 PMID: 10377396
  5. Functional Cus1p is found with Hsh155p in a multiprotein splicing factor associated with U2 snRNA.
    Mol Cell Biol. 2000 Mar;20(6):2176-85 PMID: 10688664
  6. The transcriptional program of sporulation in budding yeast.
    Science. 1998 Oct 23;282(5389):699-705 PMID: 9784122
  7. Toward a comprehensive atlas of the physical interactome of Saccharomyces cerevisiae.
    Mol Cell Proteomics. 2007 Mar;6(3):439-50 PMID: 17200106
  8. Comprehensive identification of cell cycle-regulated genes of the yeast Saccharomyces cerevisiae by microarray hybridization.
    Mol Biol Cell. 1998 Dec;9(12):3273-97 PMID: 9843569
  9. Kernel methods for predicting protein-protein interactions.
    Bioinformatics. 2005 Jun;21 Suppl 1:i38-46 PMID: 15961482
  10. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  11. Functional discovery via a compendium of expression profiles.
    Cell. 2000 Jul 7;102(1):109-26 PMID: 10929718
  12. Transcriptional regulatory networks in Saccharomyces cerevisiae.
    Science. 2002 Oct 25;298(5594):799-804 PMID: 12399584
  13. Remote homology detection: a motif based approach.
    Bioinformatics. 2003;19 Suppl 1:i26-33 PMID: 12855434
  14. Predicting co-complexed protein pairs using genomic and proteomic data integration.
    BMC Bioinformatics. 2004 Apr 16;5:38 PMID: 15090078
  15. Identification and comparative analysis of the large subunit mitochondrial ribosomal proteins of Neurospora crassa.
    FEMS Microbiol Lett. 2006 Jan;254(1):157-64 PMID: 16451194
  16. The spectrum kernel: a string kernel for SVM protein classification.
    Pac Symp Biocomput. 2002;:564-75 PMID: 11928508
  17. Assessing experimentally derived interactions in a small world.
    Proc Natl Acad Sci U S A. 2003 Apr 15;100(8):4372-6 PMID: 12676999
  18. Inferring domain-domain interactions from protein-protein interactions.
    Genome Res. 2002 Oct;12(10):1540-8 PMID: 12368246
  19. Annotation transfer between genomes: protein-protein interologs and protein-DNA regulogs.
    Genome Res. 2004 Jun;14(6):1107-18 PMID: 15173116
  20. Analyzing protein function on a genomic scale: the importance of gold-standard positives and negatives for network prediction.
    Curr Opin Microbiol. 2004 Oct;7(5):535-45 PMID: 15451510
  21. BioGRID: a general repository for interaction datasets.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D535-9 PMID: 16381927
  22. Pfam: a comprehensive database of protein domain families based on seed alignments.
    Proteins. 1997 Jul;28(3):405-20 PMID: 9223186
  23. GO::TermFinder--open source software for accessing Gene Ontology information and finding significantly enriched Gene Ontology terms associated with a list of genes.
    Bioinformatics. 2004 Dec 12;20(18):3710-5 PMID: 15297299
  24. Genetic and physical interactions between factors involved in both cell cycle progression and pre-mRNA splicing in Saccharomyces cerevisiae.
    Genetics. 2000 Dec;156(4):1503-17 PMID: 11102353
  25. Predicting protein-protein interactions using signature products.
    Bioinformatics. 2005 Jan 15;21(2):218-26 PMID: 15319262
  26. A statistical framework for genomic data fusion.
    Bioinformatics. 2004 Nov 1;20(16):2626-35 PMID: 15130933
  27. A comprehensive analysis of protein-protein interactions in Saccharomyces cerevisiae.
    Nature. 2000 Feb 10;403(6770):623-7 PMID: 10688190
  28. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  29. Global landscape of protein complexes in the yeast Saccharomyces cerevisiae.
    Nature. 2006 Mar 30;440(7084):637-43 PMID: 16554755
  30. Learning to predict protein-protein interactions from protein sequences.
    Bioinformatics. 2003 Oct 12;19(15):1875-81 PMID: 14555619
  31. Detecting protein function and protein-protein interactions from genome sequences.
    Science. 1999 Jul 30;285(5428):751-3 PMID: 10427000
  32. Functional organization of the yeast proteome by systematic analysis of protein complexes.
    Nature. 2002 Jan 10;415(6868):141-7 PMID: 11805826
  33. Proteome survey reveals modularity of the yeast cell machinery.
    Nature. 2006 Mar 30;440(7084):631-6 PMID: 16429126
  34. MIPS: a database for genomes and protein sequences.
    Nucleic Acids Res. 2000 Jan 1;28(1):37-40 PMID: 10592176
  35. High-definition macromolecular composition of yeast RNA-processing complexes.
    Mol Cell. 2004 Jan 30;13(2):225-39 PMID: 14759368
  36. What is a support vector machine?
    Nat Biotechnol. 2006 Dec;24(12):1565-7 PMID: 17160063
  37. In silico two-hybrid system for the selection of physically interacting protein pairs.
    Proteins. 2002 May 1;47(2):219-27 PMID: 11933068
  38. Global mapping of the yeast genetic interaction network.
    Science. 2004 Feb 6;303(5659):808-13 PMID: 14764870
  39. Genomic expression programs in the response of yeast cells to environmental changes.
    Mol Biol Cell. 2000 Dec;11(12):4241-57 PMID: 11102521
  40. Information assessment on predicting protein-protein interactions.
    BMC Bioinformatics. 2004 Oct 18;5:154 PMID: 15491499
  41. Toward a protein-protein interaction map of the budding yeast: A comprehensive system to examine two-hybrid interactions in all possible combinations between the yeast proteins.
    Proc Natl Acad Sci U S A. 2000 Feb 1;97(3):1143-7 PMID: 10655498
  42. Evaluation of different biological data and computational classification methods for use in protein interaction prediction.
    Proteins. 2006 May 15;63(3):490-500 PMID: 16450363
  43. Exo84p is an exocyst protein essential for secretion.
    J Biol Chem. 1999 Aug 13;274(33):23558-64 PMID: 10438536
  44. Exploiting the co-evolution of interacting proteins to discover interaction specificity.
    J Mol Biol. 2003 Mar 14;327(1):273-84 PMID: 12614624
  45. A Bayesian networks approach for predicting protein-protein interactions from genomic data.
    Science. 2003 Oct 17;302(5644):449-53 PMID: 14564010
  46. Exploring the metabolic and genetic control of gene expression on a genomic scale.
    Science. 1997 Oct 24;278(5338):680-6 PMID: 9381177
  47. Choosing negative examples for the prediction of protein-protein interactions.
    BMC Bioinformatics. 2006 Mar 20;7 Suppl 1:S2 PMID: 16723005
  48. Correlated sequence-signatures as markers of protein-protein interaction.
    J Mol Biol. 2001 Aug 24;311(4):681-92 PMID: 11518523
  49. BIND--The Biomolecular Interaction Network Database.
    Nucleic Acids Res. 2001 Jan 1;29(1):242-5 PMID: 11125103
Article Info
Journal
PLoS computational biology
Abbr.
PLoS Comput Biol
ISSN
1553-7358
Published
2008-04-18
Epub
2008-00-18
Pages
e1000054
Language
English
Region
United States
NLM ID
101238922
PMCID
PMC2275314
Subset
IM
Grants
NCRR NIH HHS · P41 RR011823 · United States
NHGRI NIH HHS · R33 HG003070 · United States
NCRR NIH HHS · P41 RR11823 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com