Home LiteratureArticle Details
PMID: 18247410 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

MUSTER: Improving protein sequence profile-profile alignments by using multiple sources of structure information.

Proteins ·Vol. 72 ·No. 2 ·2008-08-00 ·Pages 547-56

Wu S, Zhang Y

Abstract

We develop a new threading algorithm MUSTER by extending the previous sequence profile-profile alignment method, PPA. It combines various sequence and structure information into single-body terms which can be conveniently used in dynamic programming search: (1) sequence profiles; (2) secondary structures; (3) structure fragment profiles; (4) solvent accessibility; (5) dihedral torsion angles; (6) hydrophobic scoring matrix. The balance of the weighting parameters is optimized by a grading search based on the average TM-score of 111 training proteins which shows a better performance than using the conventional optimization methods based on the PROSUP database. The algorithm is tested on 500 nonhomologous proteins independent of the training sets. After removing the homologous templates with a sequence identity to the target >30%, in 224 cases, the first template alignment has the correct topology with a TM-score >0.5. Even with a more stringent cutoff by removing the templates with a sequence identity >20% or detectable by PSI-BLAST with an E-value <0.05, MUSTER is able to identify correct folds in 137 cases with the first model of TM-score >0.5. Dependent on the homology cutoffs, the average TM-score of the first threading alignments by MUSTER is 5.1-6.3% higher than that by PPA. This improvement is statistically significant by the Wilcoxon signed rank test with a P-value < 1.0 x 10(-13), which demonstrates the effect of additional structural information on the protein fold recognition. The MUSTER server is freely available to the academic community at http://zhang.bioinformatics.ku.edu/MUSTER.

MeSH Terms
Algorithms Databases, Protein Models, Molecular Protein Conformation Proteins/chemistry Sequence Alignment
Chemicals
Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Wu Sitao
Center for Bioinformatics and Department of Molecular Bioscience, University of Kansas, 2030 Becker Dr, Lawrence, Kansas 66047, USA.
Zhang Yang
References (62)
62 references, click to expand
  1. Critical assessment of methods of protein structure prediction (CASP)--round 6.
    Proteins. 2005;61 Suppl 7:3-7 PMID: 16187341
  2. Assessing the reliability of sequence similarities detected through hydrophobic cluster analysis.
    Proteins. 2008 Mar;70(4):1588-94 PMID: 17918727
  3. Hidden Markov models for detecting remote protein homologies.
    Bioinformatics. 1998;14(10):846-56 PMID: 9927713
  4. Residue depth: a novel parameter for the analysis of protein structure and stability.
    Structure. 1999 Jul 15;7(7):723-32 PMID: 10425675
  5. Tertiary structure predictions on a comprehensive benchmark of medium to large size proteins.
    Biophys J. 2004 Oct;87(4):2647-55 PMID: 15454459
  6. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  7. Automated structure prediction of weakly homologous proteins on a genomic scale.
    Proc Natl Acad Sci U S A. 2004 May 18;101(20):7594-9 PMID: 15126668
  8. Position-based sequence weights.
    J Mol Biol. 1994 Nov 4;243(4):574-8 PMID: 7966282
  9. Protein homology detection by HMM-HMM comparison.
    Bioinformatics. 2005 Apr 1;21(7):951-60 PMID: 15531603
  10. Real value prediction of solvent accessibility from amino acid sequence.
    Proteins. 2003 Mar 1;50(4):629-35 PMID: 12577269
  11. Template-based modeling and free modeling by I-TASSER in CASP7.
    Proteins. 2007;69 Suppl 8:108-17 PMID: 17894355
  12. Structure-based evaluation of sequence comparison and fold recognition alignment accuracy.
    J Mol Biol. 2000 Apr 7;297(4):1003-13 PMID: 10736233
  13. Alignment of protein sequences by their profiles.
    Protein Sci. 2004 Apr;13(4):1071-87 PMID: 15044736
  14. GenTHREADER: an efficient and reliable protein fold recognition method for genomic sequences.
    J Mol Biol. 1999 Apr 9;287(4):797-815 PMID: 10191147
  15. Enriching the sequence substitution matrix by structural information.
    Proteins. 2004 Jan 1;54(1):41-8 PMID: 14705022
  16. A method to identify protein sequences that fold into a known three-dimensional structure.
    Science. 1991 Jul 12;253(5016):164-70 PMID: 1853201
  17. Assessment of CASP7 predictions for template-based modeling targets.
    Proteins. 2007;69 Suppl 8:38-56 PMID: 17894352
  18. Enhanced genome annotation using structural profiles in the program 3D-PSSM.
    J Mol Biol. 2000 Jun 2;299(2):499-520 PMID: 10860755
  19. Enlarged representative set of protein structures.
    Protein Sci. 1994 Mar;3(3):522-4 PMID: 8019422
  20. Protein secondary structure prediction based on position-specific scoring matrices.
    J Mol Biol. 1999 Sep 17;292(2):195-202 PMID: 10493868
  21. Automated server predictions in CASP7.
    Proteins. 2007;69 Suppl 8:68-82 PMID: 17894354
  22. FUGUE: sequence-structure homology recognition using environment-specific substitution tables and structure-dependent gap penalties.
    J Mol Biol. 2001 Jun 29;310(1):243-57 PMID: 11419950
  23. 3D-Jury: a simple approach to improve protein structure predictions.
    Bioinformatics. 2003 May 22;19(8):1015-8 PMID: 12761065
  24. Prediction of solvent accessibility and sites of deleterious mutations from protein sequence.
    Nucleic Acids Res. 2005 Jun 03;33(10):3193-9 PMID: 15937195
  25. Assessment of predictions submitted for the CASP6 comparative modeling category.
    Proteins. 2005;61 Suppl 7:27-45 PMID: 16187345
  26. Single-body residue-level knowledge-based energy score combined with sequence-profile and secondary structure information for fold recognition.
    Proteins. 2004 Jun 1;55(4):1005-13 PMID: 15146497
  27. Scoring function for automated assessment of protein structure template quality.
    Proteins. 2004 Dec 1;57(4):702-10 PMID: 15476259
  28. A machine learning information retrieval approach to protein fold recognition.
    Bioinformatics. 2006 Jun 15;22(12):1456-63 PMID: 16547073
  29. Prediction of protein antigenic determinants from amino acid sequences.
    Proc Natl Acad Sci U S A. 1981 Jun;78(6):3824-8 PMID: 6167991
  30. A study of combined structure/sequence profiles.
    Fold Des. 1996;1(6):451-61 PMID: 9080191
  31. A simple method for displaying the hydropathic character of a protein.
    J Mol Biol. 1982 May 5;157(1):105-32 PMID: 7108955
  32. Within the twilight zone: a sensitive profile-profile comparison tool based on information theory.
    J Mol Biol. 2002 Feb 1;315(5):1257-75 PMID: 11827492
  33. Development and large scale benchmark testing of the PROSPECTOR_3 threading algorithm.
    Proteins. 2004 Aug 15;56(3):502-18 PMID: 15229883
  34. Combining local-structure, fold-recognition, and new fold methods for protein structure prediction.
    Proteins. 2003;53 Suppl 6:491-6 PMID: 14579338
  35. Fold recognition by combining sequence profiles derived from evolution and from depth-dependent structural alignment of fragments.
    Proteins. 2005 Feb 1;58(2):321-8 PMID: 15523666
  36. ORFeus: Detection of distant homology using sequence profiles and predicted secondary structure.
    Nucleic Acids Res. 2003 Jul 1;31(13):3804-7 PMID: 12824423
  37. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  38. TOUCHSTONE II: a new approach to ab initio protein structure prediction.
    Biophys J. 2003 Aug;85(2):1145-64 PMID: 12885659
  39. CAFASP3: the third critical assessment of fully automated structure prediction methods.
    Proteins. 2003;53 Suppl 6:503-16 PMID: 14579340
  40. Topology fingerprint approach to the inverse protein folding problem.
    J Mol Biol. 1992 Sep 5;227(1):227-38 PMID: 1522587
  41. RAPTOR: optimal protein threading by linear programming.
    J Bioinform Comput Biol. 2003 Apr;1(1):95-117 PMID: 15290783
  42. Profile hidden Markov models.
    Bioinformatics. 1998;14(9):755-63 PMID: 9918945
  43. Assessment of fold recognition predictions in CASP6.
    Proteins. 2005;61 Suppl 7:46-66 PMID: 16187346
  44. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features.
    Biopolymers. 1983 Dec;22(12):2577-637 PMID: 6667333
  45. Ab initio protein structure prediction using chunk-TASSER.
    Biophys J. 2007 Sep 1;93(5):1510-8 PMID: 17496016
  46. Hydrophobic cluster analysis: an efficient new way to compare and analyse amino acid sequences.
    FEBS Lett. 1987 Nov 16;224(1):149-55 PMID: 3678489
  47. PCMA: fast and accurate multiple sequence alignment based on profile consistency.
    Bioinformatics. 2003 Feb 12;19(3):427-8 PMID: 12584134
  48. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  49. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  50. Protein threading using PROSPECT: design and evaluation.
    Proteins. 2000 Aug 15;40(3):343-54 PMID: 10861926
  51. LOMETS: a local meta-threading-server for protein structure prediction.
    Nucleic Acids Res. 2007;35(10):3375-82 PMID: 17478507
  52. Protein threading by PROSPECT: a prediction experiment in CASP3.
    Protein Eng. 1999 Nov;12(11):899-907 PMID: 10585495
  53. PISCES: a protein sequence culling server.
    Bioinformatics. 2003 Aug 12;19(12):1589-91 PMID: 12912846
  54. A new approach to protein fold recognition.
    Nature. 1992 Jul 2;358(6381):86-9 PMID: 1614539
  55. TM-align: a protein structure alignment algorithm based on the TM-score.
    Nucleic Acids Res. 2005 Apr 22;33(7):2302-9 PMID: 15849316
  56. Knowledge-based protein secondary structure assignment.
    Proteins. 1995 Dec;23(4):566-79 PMID: 8749853
  57. A study of quality measures for protein threading models.
    BMC Bioinformatics. 2001;2:5 PMID: 11545673
  58. Defrosting the frozen approximation: PROSPECTOR--a new approach to threading.
    Proteins. 2001 Feb 15;42(3):319-31 PMID: 11151004
  59. Comparison of sequence profiles. Strategies for structural predictions using sequence information.
    Protein Sci. 2000 Feb;9(2):232-41 PMID: 10716175
  60. FFAS03: a server for profile--profile sequence alignments.
    Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W284-8 PMID: 15980471
  61. Ab initio modeling of small proteins by iterative TASSER simulations.
    BMC Biol. 2007 May 08;5:17 PMID: 17488521
  62. LiveBench-8: the large-scale, continuous assessment of automated protein structure prediction.
    Protein Sci. 2005 Jan;14(1):240-5 PMID: 15608124
Article Info
Journal
Proteins
Abbr.
Proteins
ISSN
1097-0134
Published
2008-08-00
Pages
547-56
Language
English
Region
United States
NLM ID
8700181
PMCID
PMC2666101
Subset
IM
Grants
NIGMS NIH HHS · R01 GM083107 · United States
NIGMS NIH HHS · R01 GM083107-01A1 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com