Abstract
The development of liquid chromatography coupled with tandem mass spectrometry (LC-MS/MS) has made it possible to characterize phosphopeptides in an increasingly large-scale and high-throughput fashion. However, extracting confident phosphopeptide identifications from the resulting large data sets in a similar high-throughput fashion remains difficult, as does rigorously estimating the false discovery rate (FDR) of a set of phosphopeptide identifications. This article describes a data analysis pipeline designed to address these issues. The first step is to reanalyze phosphopeptide identifications that contain ambiguous assignments for the incorporated phosphate(s) to determine the most likely arrangement of the phosphate(s). The next step is to employ an expectation maximization algorithm to estimate the joint distribution of the peptide scores. A linear discriminant analysis is then performed to determine how to optimally combine peptide scores (in this case, from SEQUEST) into a discriminant score that possesses the maximum discriminating power. Based on this discriminant score, the p- and q-values for each phosphopeptide identification are calculated, and the phosphopeptide identification FDR is then estimated. This data analysis approach was applied to data from a study of irradiated human skin fibroblasts to provide a robust estimate of FDR for phosphopeptides. The Phosphopeptide FDR Estimator software is freely available for download at http://ncrr.pnl.gov/software/.
MeSH Terms
Algorithms
Bayes Theorem
Data Interpretation, Statistical
Discriminant Analysis
Fibroblasts/chemistry,cytology,radiation effects
Humans
Internet
Mass Spectrometry/statistics & numerical data
Normal Distribution
Phosphopeptides/analysis
Proteomics/methods
ROC Curve
Reproducibility of Results
Skin/cytology
Software
Chemicals
Phosphopeptides
Authors & Affiliations
10 authors, click to expand affiliations / ORCID
Du Xiuxia
Fundamental and Computational Sciences Directorate, Pacific Northwest National Laboratory, Richland, Washington 99352, USA.
Yang Feng
Manes Nathan P
Stenoien David L
Monroe Matthew E
Adkins Joshua N
States David J
Purvine Samuel O
Camp David G
Smith Richard D
References (20)
20 references, click to expand
-
Empirical statistical model to estimate the accuracy of peptide identifications made by MS/MS and database search.
Anal Chem. 2002 Oct 15;74(20):5383-92
PMID: 12403597
-
An approach to correlate tandem mass spectral data of peptides with amino acid sequences in a protein database.
J Am Soc Mass Spectrom. 1994 Nov;5(11):976-89
PMID: 24226387
-
Phosphoproteome profiling of human skin fibroblast cells in response to low- and high-dose irradiation.
J Proteome Res. 2006 May;5(5):1252-60
PMID: 16674116
-
Ubiquitin-mediated proteolysis: biological regulation via destruction.
Bioessays. 2000 May;22(5):442-51
PMID: 10797484
-
Assigning significance to peptides identified by tandem mass spectrometry using decoy databases.
J Proteome Res. 2008 Jan;7(1):29-34
PMID: 18067246
-
Integrated approach for manual evaluation of peptides identified by searching protein sequence databases with tandem mass spectra.
J Proteome Res. 2005 May-Jun;4(3):998-1005
PMID: 15952748
-
Posterior error probabilities and false discovery rates: two sides of the same coin.
J Proteome Res. 2008 Jan;7(1):40-4
PMID: 18052118
-
A probability-based approach for high-throughput protein phosphorylation analysis and site localization.
Nat Biotechnol. 2006 Oct;24(10):1285-92
PMID: 16964243
-
TANDEM: matching proteins with tandem mass spectra.
Bioinformatics. 2004 Jun 12;20(9):1466-7
PMID: 14976030
-
Statistical model for large-scale peptide identification in databases from tandem mass spectra using SEQUEST.
Anal Chem. 2004 Dec 1;76(23):6853-60
PMID: 15571333
-
Statistical significance for genomewide studies.
Proc Natl Acad Sci U S A. 2003 Aug 5;100(16):9440-5
PMID: 12883005
-
Advances in proteomics data analysis and display using an accurate mass and time tag approach.
Mass Spectrom Rev. 2006 May-Jun;25(3):450-82
PMID: 16429408
-
Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry.
Nat Methods. 2007 Mar;4(3):207-14
PMID: 17327847
-
G1/S regulatory mechanisms from yeast to man.
Prog Cell Cycle Res. 1996;2:15-27
PMID: 9552379
-
Profiling signaling polarity in chemotactic cells.
Proc Natl Acad Sci U S A. 2007 May 15;104(20):8328-33
PMID: 17494752
-
False discovery rates and related statistical concepts in mass spectrometry-based proteomics.
J Proteome Res. 2008 Jan;7(1):47-50
PMID: 18067251
-
Signaling--2000 and beyond.
Cell. 2000 Jan 7;100(1):113-27
PMID: 10647936
-
Identification of a novel mitotic phosphorylation motif associated with protein localization to the mitotic apparatus.
J Cell Sci. 2007 Nov 15;120(Pt 22):4060-70
PMID: 17971412
-
The International Protein Index: an integrated database for proteomics experiments.
Proteomics. 2004 Jul;4(7):1985-8
PMID: 15221759
-
Evaluation of multidimensional chromatography coupled with tandem mass spectrometry (LC/LC-MS/MS) for large-scale protein analysis: the yeast proteome.
J Proteome Res. 2003 Jan-Feb;2(1):43-50
PMID: 12643542