Home LiteratureArticle Details
PMID: 15090078 Published · epublish English Journal Article Research Support, Non-U.S. Gov't Validation Study

Predicting co-complexed protein pairs using genomic and proteomic data integration.

BMC bioinformatics ·Vol. 5 ·2004-04-16 ·Pages 38

Zhang LV, Wong SL, King OD, Roth FP

Abstract

Identifying all protein-protein interactions in an organism is a major objective of proteomics. A related goal is to know which protein pairs are present in the same protein complex. High-throughput methods such as yeast two-hybrid (Y2H) and affinity purification coupled with mass spectrometry (APMS) have been used to detect interacting proteins on a genomic scale. However, both Y2H and APMS methods have substantial false-positive rates. Aside from high-throughput interaction screens, other gene- or protein-pair characteristics may also be informative of physical interaction. Therefore it is desirable to integrate multiple datasets and utilize their different predictive value for more accurate prediction of co-complexed relationship. Using a supervised machine learning approach--probabilistic decision tree, we integrated high-throughput protein interaction datasets and other gene- and protein-pair characteristics to predict co-complexed pairs (CCP) of proteins. Our predictions proved more sensitive and specific than predictions based on Y2H or APMS methods alone or in combination. Among the top predictions not annotated as CCPs in our reference set (obtained from the MIPS complex catalogue), a significant fraction was found to physically interact according to a separate database (YPD, Yeast Proteome Database), and the remaining predictions may potentially represent unknown CCPs. We demonstrated that the probabilistic decision tree approach can be successfully used to predict co-complexed protein (CCP) pairs from other characteristics. Our top-scoring CCP predictions provide testable hypotheses for experimental validation.

MeSH Terms
Computational Biology/methods Databases, Protein Decision Trees Gene Expression Regulation, Fungal/genetics Genes, Fungal/genetics Genomics/statistics & numerical data Mass Spectrometry Predictive Value of Tests Protein Interaction Mapping/methods,statistics & numerical data Proteomics/statistics & numerical data RNA, Fungal/biosynthesis,genetics Saccharomyces cerevisiae Proteins/metabolism Two-Hybrid System Techniques
Chemicals
RNA, Fungal Saccharomyces cerevisiae Proteins
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Zhang Lan V
Department of Biological Chemistry and Molecular Pharmacology, Harvard Medical School, Boston, MA 02115, USA. lan_zhang@student.hms.harvard.edu
Wong Sharyl L
King Oliver D
Roth Frederick P
References (41)
41 references, click to expand
  1. Predicting gene function from patterns of annotation.
    Genome Res. 2003 May;13(5):896-904 PMID: 12695322
  2. Transcriptional regulatory networks in Saccharomyces cerevisiae.
    Science. 2002 Oct 25;298(5594):799-804 PMID: 12399584
  3. A Bayesian framework for combining heterogeneous data sources for gene function prediction (in Saccharomyces cerevisiae).
    Proc Natl Acad Sci U S A. 2003 Jul 8;100(14):8348-53 PMID: 12826619
  4. A combined algorithm for genome-wide prediction of protein function.
    Nature. 1999 Nov 4;402(6757):83-6 PMID: 10573421
  5. Toward a protein-protein interaction map of the budding yeast: A comprehensive system to examine two-hybrid interactions in all possible combinations between the yeast proteins.
    Proc Natl Acad Sci U S A. 2000 Feb 1;97(3):1143-7 PMID: 10655498
  6. A comprehensive analysis of protein-protein interactions in Saccharomyces cerevisiae.
    Nature. 2000 Feb 10;403(6770):623-7 PMID: 10688190
  7. Functional discovery via a compendium of expression profiles.
    Cell. 2000 Jul 7;102(1):109-26 PMID: 10929718
  8. A network of protein-protein interactions in yeast.
    Nat Biotechnol. 2000 Dec;18(12):1257-61 PMID: 11101803
  9. Gene number. What if there are only 30,000 human genes?
    Science. 2001 Feb 16;291(5507):1255-7 PMID: 11233450
  10. Networking proteins in yeast.
    Proc Natl Acad Sci U S A. 2001 Apr 10;98(8):4277-8 PMID: 11296274
  11. A comprehensive two-hybrid analysis to explore the yeast protein interactome.
    Proc Natl Acad Sci U S A. 2001 Apr 10;98(8):4569-74 PMID: 11283351
  12. A relationship between gene expression and protein interactions on the proteome scale: analysis of the bacteriophage T7 and the yeast Saccharomyces cerevisiae.
    Nucleic Acids Res. 2001 Sep 1;29(17):3513-9 PMID: 11522820
  13. Protein interaction databases.
    Curr Opin Biotechnol. 2001 Aug;12(4):334-9 PMID: 11551460
  14. Correlation between transcriptome and interactome mapping data from Saccharomyces cerevisiae.
    Nat Genet. 2001 Dec;29(4):482-6 PMID: 11694880
  15. MIPS: a database for genomes and protein sequences.
    Nucleic Acids Res. 2002 Jan 1;30(1):31-4 PMID: 11752246
  16. Relating whole-genome expression data with protein-protein interactions.
    Genome Res. 2002 Jan;12(1):37-46 PMID: 11779829
  17. Proteomics. Integrating interactomes.
    Science. 2002 Jan 11;295(5553):284-7 PMID: 11786630
  18. Functional organization of the yeast proteome by systematic analysis of protein complexes.
    Nature. 2002 Jan 10;415(6868):141-7 PMID: 11805826
  19. Systematic identification of protein complexes in Saccharomyces cerevisiae by mass spectrometry.
    Nature. 2002 Jan 10;415(6868):180-3 PMID: 11805837
  20. Subcellular localization of the yeast proteome.
    Genes Dev. 2002 Mar 15;16(6):707-19 PMID: 11914276
  21. Learning gene functional classifications from multiple data types.
    J Comput Biol. 2002;9(2):401-11 PMID: 12015889
  22. Comparative assessment of large-scale data sets of protein-protein interactions.
    Nature. 2002 May 23;417(6887):399-403 PMID: 12000970
  23. Protein interaction verification and functional annotation by integrated analysis of genome-scale data.
    Mol Cell. 2002 May;9(5):1133-43 PMID: 12049748
  24. Three yeast proteome databases: YPD, PombePD, and CalPD (MycoPathPD).
    Methods Enzymol. 2002;350:347-73 PMID: 12073323
  25. A large nucleolar U3 ribonucleoprotein required for 18S ribosomal RNA biogenesis.
    Nature. 2002 Jun 27;417(6892):967-70 PMID: 12068309
  26. Protein interactions: two methods for assessment of the reliability of high throughput observations.
    Mol Cell Proteomics. 2002 May;1(5):349-56 PMID: 12118076
  27. Analyzing yeast protein-protein interaction data obtained from different sources.
    Nat Biotechnol. 2002 Oct;20(10):991-7 PMID: 12355115
  28. Predicting phenotype from patterns of annotation.
    Bioinformatics. 2003;19 Suppl 1:i183-9 PMID: 12855456
  29. Integrating 'omic' information: a bridge between genomics and systems biology.
    Trends Genet. 2003 Oct;19(10):551-60 PMID: 14550629
  30. Greedily building protein networks with confidence.
    Bioinformatics. 2003 Oct 12;19(15):1869-74 PMID: 14555618
  31. A Bayesian networks approach for predicting protein-protein interactions from genomic data.
    Science. 2003 Oct 17;302(5644):449-53 PMID: 14564010
  32. An automated method for finding molecular complexes in large protein interaction networks.
    BMC Bioinformatics. 2003 Jan 13;4:2 PMID: 12525261
  33. Predicting protein complex membership using probabilistic network reliability.
    Genome Res. 2004 Jun;14(6):1170-5 PMID: 15140827
  34. Suppression of yeast RNA polymerase III mutations by FHL1, a gene coding for a fork head protein involved in rRNA processing.
    Mol Cell Biol. 1994 May;14(5):2905-13 PMID: 8164651
  35. SRD1, a S. cerevisiae gene affecting pre-rRNA processing contains a C2/C2 zinc finger motif.
    Nucleic Acids Res. 1994 Apr 11;22(7):1265-71 PMID: 8165142
  36. Small nucleolar RNAs direct site-specific synthesis of pseudouridine in ribosomal RNA.
    Cell. 1997 May 16;89(4):565-73 PMID: 9160748
  37. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  38. Crystallographic analysis of the recognition of a nuclear localization signal by the nuclear import factor karyopherin alpha.
    Cell. 1998 Jul 24;94(2):193-204 PMID: 9695948
  39. A genome-wide transcriptional analysis of the mitotic cell cycle.
    Mol Cell. 1998 Jul;2(1):65-73 PMID: 9702192
  40. Detecting protein function and protein-protein interactions from genome sequences.
    Science. 1999 Jul 30;285(5428):751-3 PMID: 10427000
  41. Integration of genomic datasets to predict protein complexes in yeast.
    J Struct Funct Genomics. 2002;2(2):71-81 PMID: 12836664
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2004-04-16
Epub
2004-00-16
Pages
38
Language
English
Region
England
NLM ID
100965194
PMCID
PMC419405
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com