Abstract
The goal of many shotgun proteomics experiments is to determine the protein complement of a complex biological mixture. For many mixtures, most methodological approaches fall significantly short of this goal. Existing solutions to this problem typically subdivide the task into two stages: first identifying a collection of peptides with a low false discovery rate and then inferring from the peptides a corresponding set of proteins. In contrast, we formulate the protein identification problem as a single optimization problem, which we solve using machine learning methods. This approach is motivated by the observation that the peptide and protein level tasks are cooperative, and the solution to each can be improved by using information about the solution to the other. The resulting algorithm directly controls the relevant error rate, can incorporate a wide variety of evidence and, for complex samples, provides 18-34% more protein identifications than the current state of the art approaches.
MeSH Terms
Algorithms
Amniotic Fluid/chemistry,metabolism
Artificial Intelligence
Caenorhabditis elegans Proteins/metabolism
Complex Mixtures/analysis
Databases, Protein
Humans
Laryngopharyngeal Reflux
Models, Statistical
Peptide Fragments/analysis
Proteins/analysis
Proteomics
Saccharomyces cerevisiae Proteins/metabolism
Software
Tandem Mass Spectrometry/methods
Chemicals
Caenorhabditis elegans Proteins
Complex Mixtures
Peptide Fragments
Proteins
Saccharomyces cerevisiae Proteins
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Spivak Marina
Department of Genome Sciences, University of Washington, Seattle, Washington 98195, USA.
Weston Jason
Tomazela Daniela
MacCoss Michael J
Noble William Stafford
References (28)
28 references, click to expand
-
Statistical validation of peptide identifications in large-scale proteomics using the target-decoy database search strategy and flexible mixture modeling.
J Proteome Res. 2008 Jan;7(1):286-92
PMID: 18078310
-
Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry.
Nat Methods. 2007 Mar;4(3):207-14
PMID: 17327847
-
A method for assessing the statistical significance of mass spectrometry-based protein identifications using general scoring schemes.
Anal Chem. 2003 Feb 15;75(4):768-74
PMID: 12622365
-
False discovery rates and related statistical concepts in mass spectrometry-based proteomics.
J Proteome Res. 2008 Jan;7(1):47-50
PMID: 18067251
-
Probability-based protein identification by searching sequence databases using mass spectrometry data.
Electrophoresis. 1999 Dec;20(18):3551-67
PMID: 10612281
-
Semi-supervised learning for peptide identification from shotgun proteomics datasets.
Nat Methods. 2007 Nov;4(11):923-5
PMID: 17952086
-
Prediction of error associated with false-positive rate determination for peptide identification in large-scale proteomics experiments using a combined reverse and forward peptide sequence database strategy.
J Proteome Res. 2007 Jan;6(1):392-8
PMID: 17203984
-
Improvements to the percolator algorithm for Peptide identification from shotgun proteomics data sets.
J Proteome Res. 2009 Jul;8(7):3737-45
PMID: 19385687
-
Modes of inference for evaluating the confidence of peptide identifications.
J Proteome Res. 2008 Jan;7(1):35-9
PMID: 18067248
-
Dissecting the regulatory circuitry of a eukaryotic genome.
Cell. 1998 Nov 25;95(5):717-28
PMID: 9845373
-
Protein identification false discovery rates for very large proteomics data sets generated by tandem mass spectrometry.
Mol Cell Proteomics. 2009 Nov;8(11):2405-17
PMID: 19608599
-
Rapid and accurate peptide identification from tandem mass spectra.
J Proteome Res. 2008 Jul;7(7):3022-7
PMID: 18505281
-
A statistical model for identifying proteins by tandem mass spectrometry.
Anal Chem. 2003 Sep 1;75(17):4646-58
PMID: 14632076
-
Semisupervised model-based validation of peptide identifications in mass spectrometry-based proteomics.
J Proteome Res. 2008 Jan;7(1):254-65
PMID: 18159924
-
Open mass spectrometry search algorithm.
J Proteome Res. 2004 Sep-Oct;3(5):958-64
PMID: 15473683
-
Empirical statistical model to estimate the accuracy of peptide identifications made by MS/MS and database search.
Anal Chem. 2002 Oct 15;74(20):5383-92
PMID: 12403597
-
IDPicker 2.0: Improved protein assembly with high discrimination peptide identification filtering.
J Proteome Res. 2009 Aug;8(8):3872-81
PMID: 19522537
-
Global analysis of protein expression in yeast.
Nature. 2003 Oct 16;425(6959):737-41
PMID: 14562106
-
Qscore: an algorithm for evaluating SEQUEST database search results.
J Am Soc Mass Spectrom. 2002 Apr;13(4):378-86
PMID: 11951976
-
Data management and preliminary data analysis in the pilot phase of the HUPO Plasma Proteome Project.
Proteomics. 2005 Aug;5(13):3246-61
PMID: 16104057
-
Assigning significance to peptides identified by tandem mass spectrometry using decoy databases.
J Proteome Res. 2008 Jan;7(1):29-34
PMID: 18067246
-
Machines that learn to segment images: a crucial technology for connectomics.
Curr Opin Neurobiol. 2010 Oct;20(5):653-66
PMID: 20801638
-
Evaluation of multidimensional chromatography coupled with tandem mass spectrometry (LC/LC-MS/MS) for large-scale protein analysis: the yeast proteome.
J Proteome Res. 2003 Jan-Feb;2(1):43-50
PMID: 12643542
-
Intensity-based protein identification by machine learning from a library of tandem mass spectra.
Nat Biotechnol. 2004 Feb;22(2):214-9
PMID: 14730315
-
A new algorithm for the evaluation of shotgun peptide sequencing in proteomics: support vector machine classification of peptide MS/MS spectra and SEQUEST scores.
J Proteome Res. 2003 Mar-Apr;2(2):137-46
PMID: 12716127
-
Proteomic parsimony through bipartite graph analysis improves accuracy and transparency.
J Proteome Res. 2007 Sep;6(9):3549-57
PMID: 17676885
-
Advancement in protein inference from shotgun proteomics using peptide detectability.
Pac Symp Biocomput. 2007;:409-20
PMID: 17990506
-
False discovery rates of protein identifications: a strike against the two-peptide rule.
J Proteome Res. 2009 Sep;8(9):4173-81
PMID: 19627159