Home LiteratureArticle Details
PMID: 18187442 Published · ppublish English Journal Article Research Support, N.I.H., Extramural

Incorporating sequence information into the scoring function: a hidden Markov model for improved peptide identification.

Bioinformatics (Oxford, England) ·Vol. 24 ·No. 5 ·2008-03-01 ·Pages 674-81

Khatun J, Hamlett E, Giddings MC

Abstract

The identification of peptides by tandem mass spectrometry (MS/MS) is a central method of proteomics research, but due to the complexity of MS/MS data and the large databases searched, the accuracy of peptide identification algorithms remains limited. To improve the accuracy of identification we applied a machine-learning approach using a hidden Markov model (HMM) to capture the complex and often subtle links between a peptide sequence and its MS/MS spectrum. Our model, HMM_Score, represents ion types as HMM states and calculates the maximum joint probability for a peptide/spectrum pair using emission probabilities from three factors: the amino acids adjacent to each fragmentation site, the mass dependence of ion types and the intensity dependence of ion types. The Viterbi algorithm is used to calculate the most probable assignment between ion types in a spectrum and a peptide sequence, then a correction factor is added to account for the propensity of the model to favor longer peptides. An expectation value is calculated based on the model score to assess the significance of each peptide/spectrum match. We trained and tested HMM_Score on three data sets generated by two different mass spectrometer types. For a reference data set recently reported in the literature and validated using seven identification algorithms, HMM_Score produced 43% more positive identification results at a 1% false positive rate than the best of two other commonly used algorithms, Mascot and X!Tandem. HMM_Score is a highly accurate platform for peptide identification that works well for a variety of mass spectrometer and biological sample types. The program is freely available on ProteomeCommons via an OpenSource license. See http://bioinfo.unc.edu/downloads/ for the download link.

MeSH Terms
Algorithms Markov Chains Models, Theoretical Peptides/chemistry Spectrometry, Mass, Matrix-Assisted Laser Desorption-Ionization Tandem Mass Spectrometry
Chemicals
Peptides
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Khatun Jainab
Department of Microbiology and Immunology, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, USA.
Hamlett Eric
Giddings Morgan C
References (32)
32 references, click to expand
  1. Open mass spectrometry search algorithm.
    J Proteome Res. 2004 Sep-Oct;3(5):958-64 PMID: 15473683
  2. PepNovo: de novo peptide sequencing via probabilistic network modeling.
    Anal Chem. 2005 Feb 15;77(4):964-73 PMID: 15858974
  3. Implementation and uses of automated de novo peptide sequencing by tandem mass spectrometry.
    Anal Chem. 2001 Jun 1;73(11):2594-604 PMID: 11403305
  4. A hypergeometric probability model for protein identification and validation using tandem mass spectral data and protein sequence databases.
    Anal Chem. 2003 Aug 1;75(15):3792-8 PMID: 14572045
  5. The complete genome sequence of Escherichia coli K-12.
    Science. 1997 Sep 5;277(5331):1453-62 PMID: 9278503
  6. SCOPE: a probabilistic model for scoring tandem mass spectra against a peptide database.
    Bioinformatics. 2001;17 Suppl 1:S13-21 PMID: 11472988
  7. Error-tolerant identification of peptides in sequence databases by peptide sequence tags.
    Anal Chem. 1994 Dec 15;66(24):4390-9 PMID: 7847635
  8. De novo peptide sequencing via tandem mass spectrometry.
    J Comput Biol. 1999 Fall-Winter;6(3-4):327-42 PMID: 10582570
  9. Shotgun protein sequencing by tandem mass spectra assembly.
    Anal Chem. 2004 Dec 15;76(24):7221-33 PMID: 15595863
  10. Fragmentation characteristics of collision-induced dissociation in MALDI TOF/TOF mass spectrometry.
    Anal Chem. 2007 Apr 15;79(8):3032-40 PMID: 17367113
  11. NovoHMM: a hidden Markov model for de novo peptide sequencing.
    Anal Chem. 2005 Nov 15;77(22):7265-73 PMID: 16285674
  12. MASPIC: intensity-based tandem mass spectrometry scoring scheme that improves peptide identification at high confidence.
    Anal Chem. 2005 Dec 1;77(23):7581-93 PMID: 16316165
  13. Open source system for analyzing, validating, and storing protein identification data.
    J Proteome Res. 2004 Nov-Dec;3(6):1234-42 PMID: 15595733
  14. Statistical characterization of ion trap tandem mass spectra from doubly charged tryptic peptides.
    Anal Chem. 2003 Mar 1;75(5):1155-63 PMID: 12641236
  15. Influence of basic residue content on fragment ion peak intensities in low-energy collision-induced dissociation spectra of peptides.
    Anal Chem. 2004 Mar 1;76(5):1243-8 PMID: 14987077
  16. Genome-based peptide fingerprint scanning.
    Proc Natl Acad Sci U S A. 2003 Jan 7;100(1):20-5 PMID: 12518051
  17. The meaning and use of the area under a receiver operating characteristic (ROC) curve.
    Radiology. 1982 Apr;143(1):29-36 PMID: 7063747
  18. PEAKS: powerful software for peptide de novo sequencing by tandem mass spectrometry.
    Rapid Commun Mass Spectrom. 2003;17(20):2337-42 PMID: 14558135
  19. Mass spectrometry allows direct identification of proteins in large genomes.
    Proteomics. 2001 May;1(5):641-50 PMID: 11678034
  20. ProSight PTM: an integrated environment for protein identification and characterization by top-down mass spectrometry.
    Nucleic Acids Res. 2004 Jul 1;32(Web Server issue):W340-5 PMID: 15215407
  21. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  22. Mass spectrometry of peptides and proteins.
    Methods. 2005 Mar;35(3):211-22 PMID: 15722218
  23. Large-scale analysis of the yeast proteome by multidimensional protein identification technology.
    Nat Biotechnol. 2001 Mar;19(3):242-7 PMID: 11231557
  24. A method for assessing the statistical significance of mass spectrometry-based protein identifications using general scoring schemes.
    Anal Chem. 2003 Feb 15;75(4):768-74 PMID: 12622365
  25. Fast tandem mass spectra-based protein identification regardless of the number of spectra or potential modifications examined.
    Bioinformatics. 2005 May 15;21(10):2177-84 PMID: 15746284
  26. ProbID: a probabilistic algorithm to identify peptides through sequence database searching using tandem mass spectral data.
    Proteomics. 2002 Oct;2(10):1406-12 PMID: 12422357
  27. Automated de novo sequencing of proteins by tandem high-resolution mass spectrometry.
    Proc Natl Acad Sci U S A. 2000 Sep 12;97(19):10313-7 PMID: 10984529
  28. Probability-based protein identification by searching sequence databases using mass spectrometry data.
    Electrophoresis. 1999 Dec;20(18):3551-67 PMID: 10612281
  29. An evaluation, comparison, and accurate benchmarking of several publicly available MS/MS search algorithms: sensitivity and specificity analysis.
    Proteomics. 2005 Aug;5(13):3475-90 PMID: 16047398
  30. PepHMM: a hidden Markov model based scoring function for mass spectrometry database search.
    Anal Chem. 2006 Jan 15;78(2):432-7 PMID: 16408924
  31. An approach to correlate tandem mass spectral data of peptides with amino acid sequences in a protein database.
    J Am Soc Mass Spectrom. 1994 Nov;5(11):976-89 PMID: 24226387
  32. Peptide rearrangement during quadrupole ion trap fragmentation: added complexity to MS/MS spectra.
    Anal Chem. 2003 Mar 15;75(6):1524-35 PMID: 12659218
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2008-03-01
Epub
2008-00-10
Pages
674-81
Language
English
Region
England
NLM ID
9808944
PMCID
PMC2699941
Subset
IM
Grants
NCRR NIH HHS · R01 RR020823-01 · United States
NCRR NIH HHS · R01 RR020823-03 · United States
NHGRI NIH HHS · R01 HG003700-03 · United States
NHGRI NIH HHS · R01 HG003700-02 · United States
NHGRI NIH HHS · R01HG003700 · United States
NCRR NIH HHS · R01 RR020823-04 · United States
NCRR NIH HHS · R01 RR020823 · United States
NHGRI NIH HHS · R01 HG003700-01 · United States
NCRR NIH HHS · R01 RR020823-02 · United States
NHGRI NIH HHS · R01 HG003700 · United States
NCRR NIH HHS · R01RR020823 · United States
Corrections
CommentIn
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com