Home LiteratureArticle Details
PMID: 17608956 Published · epublish English Comparative Study Journal Article Research Support, N.I.H., Extramural

Prediction of MHC class II binding affinity using SMM-align, a novel stabilization matrix alignment method.

BMC bioinformatics ·Vol. 8 ·2007-07-04 ·Pages 238

Nielsen M, Lundegaard C, Lund O

Abstract

Antigen presenting cells (APCs) sample the extra cellular space and present peptides from here to T helper cells, which can be activated if the peptides are of foreign origin. The peptides are presented on the surface of the cells in complex with major histocompatibility class II (MHC II) molecules. Identification of peptides that bind MHC II molecules is thus a key step in rational vaccine design and developing methods for accurate prediction of the peptide:MHC interactions play a central role in epitope discovery. The MHC class II binding groove is open at both ends making the correct alignment of a peptide in the binding groove a crucial part of identifying the core of an MHC class II binding motif. Here, we present a novel stabilization matrix alignment method, SMM-align, that allows for direct prediction of peptide:MHC binding affinities. The predictive performance of the method is validated on a large MHC class II benchmark data set covering 14 HLA-DR (human MHC) and three mouse H2-IA alleles. The predictive performance of the SMM-align method was demonstrated to be superior to that of the Gibbs sampler, TEPITOPE, SVRMHC, and MHCpred methods. Cross validation between peptide data set obtained from different sources demonstrated that direct incorporation of peptide length potentially results in over-fitting of the binding prediction method. Focusing on amino terminal peptide flanking residues (PFR), we demonstrate a consistent gain in predictive performance by favoring binding registers with a minimum PFR length of two amino acids. Visualizing the binding motif as obtained by the SMM-align and TEPITOPE methods highlights a series of fundamental discrepancies between the two predicted motifs. For the DRB1*1302 allele for instance, the TEPITOPE method favors basic amino acids at most anchor positions, whereas the SMM-align method identifies a preference for hydrophobic or neutral amino acids at the anchors. The SMM-align method was shown to outperform other state of the art MHC class II prediction methods. The method predicts quantitative peptide:MHC binding affinity values, making it ideally suited for rational epitope discovery. The method has been trained and evaluated on the, to our knowledge, largest benchmark data set publicly available and covers the nine HLA-DR supertypes suggested as well as three mouse H2-IA allele. Both the peptide benchmark data set, and SMM-align prediction method (NetMHCII) are made publicly available.

MeSH Terms
Algorithms Alleles Amino Acid Motifs Amino Acid Sequence Animals Databases, Genetic Epitopes HLA-DR Antigens/chemistry,immunology Histocompatibility Antigens Class II/chemistry,immunology Humans Inhibitory Concentration 50 Mice Monte Carlo Method Peptides/chemistry,immunology Predictive Value of Tests Protein Binding Reproducibility of Results Sequence Alignment Sequence Analysis, Protein/methods
Chemicals
Epitopes HLA-DR Antigens Histocompatibility Antigens Class II Peptides
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Nielsen Morten
Center for Biological Sequence Analysis, BioCentrum-DTU, Technical University of Denmark, Lyngby, Denmark. mniel@cbs.dtu.dk
Lundegaard Claus
Lund Ole
References (24)
24 references, click to expand
  1. Amino-terminal trimming of peptides for presentation on major histocompatibility complex class II molecules.
    Proc Natl Acad Sci U S A. 1997 Jan 21;94(2):628-33 PMID: 9012835
  2. Hidden Markov model-based prediction of antigenic peptides that interact with MHC class II molecules.
    J Biosci Bioeng. 2002;94(3):264-70 PMID: 16233301
  3. Towards the in silico identification of class II restricted T-cell epitopes: a partial least squares iterative self-consistent algorithm for affinity prediction.
    Bioinformatics. 2003 Nov 22;19(17):2263-70 PMID: 14630655
  4. Measuring the accuracy of diagnostic systems.
    Science. 1988 Jun 3;240(4857):1285-93 PMID: 3287615
  5. Prediction of MHC class I binding peptides, using SVMHC.
    BMC Bioinformatics. 2002 Sep 11;3:25 PMID: 12225620
  6. SVRMHC prediction server for MHC-binding peptides.
    BMC Bioinformatics. 2006;7:463 PMID: 17059589
  7. SYFPEITHI: database for MHC ligands and peptide motifs.
    Immunogenetics. 1999 Nov;50(3-4):213-9 PMID: 10602881
  8. Selection of representative protein data sets.
    Protein Sci. 1992 Mar;1(3):409-17 PMID: 1304348
  9. Peptide length-based prediction of peptide-MHC class II binding.
    Bioinformatics. 2006 Nov 15;22(22):2761-7 PMID: 17000752
  10. ProPred: prediction of HLA-DR binding sites.
    Bioinformatics. 2001 Dec;17(12):1236-7 PMID: 11751237
  11. Amino acid substitution matrices from protein blocks.
    Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9 PMID: 1438297
  12. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  13. Improved prediction of MHC class I and class II epitopes using a novel Gibbs sampling approach.
    Bioinformatics. 2004 Jun 12;20(9):1388-97 PMID: 14962912
  14. Identifying MHC class I epitopes by predicting the TAP transport efficiency of epitope precursors.
    J Immunol. 2003 Aug 15;171(4):1741-9 PMID: 12902473
  15. Automated generation and evaluation of specific MHC binding predictive tools: ARB matrix applications.
    Immunogenetics. 2005 Jun;57(5):304-14 PMID: 15868141
  16. Generation of tissue-specific and promiscuous HLA ligand databases using DNA microarrays and virtual HLA class II matrices.
    Nat Biotechnol. 1999 Jun;17(6):555-61 PMID: 10385319
  17. A community resource benchmarking predictions of peptide binding to MHC-I molecules.
    PLoS Comput Biol. 2006 Jun 9;2(6):e65 PMID: 16789818
  18. Prediction of MHC class II-binding peptides using an evolutionary algorithm and artificial neural network.
    Bioinformatics. 1998;14(2):121-30 PMID: 9545443
  19. Prediction of MHC class II binders using the ant colony search strategy.
    Artif Intell Med. 2005 Sep-Oct;35(1-2):147-56 PMID: 16061368
  20. A roadmap for the immunomics of category A-C pathogens.
    Immunity. 2005 Feb;22(2):155-61 PMID: 15773067
  21. Sequence logos: a new way to display consensus sequences.
    Nucleic Acids Res. 1990 Oct 25;18(20):6097-100 PMID: 2172928
  22. AntiJen: a quantitative immunology database integrating functional, thermodynamic, kinetic, biophysical, and cellular data.
    Immunome Res. 2005 Oct 06;1(1):4 PMID: 16305757
  23. Reliable prediction of T-cell epitopes using neural networks with novel sequence representations.
    Protein Sci. 2003 May;12(5):1007-17 PMID: 12717023
  24. PREDBALB/c: a system for the prediction of peptide binding to H2d molecules, a haplotype of the BALB/c mouse.
    Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W180-3 PMID: 15980450
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2007-07-04
Epub
2007-00-04
Pages
238
Language
English
Region
England
NLM ID
100965194
PMCID
PMC1939856
Subset
IM
Grants
PHS HHS · HHSN266200400083C · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com