Home LiteratureArticle Details
PMID: 17488521 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

Ab initio modeling of small proteins by iterative TASSER simulations.

BMC biology ·Vol. 5 ·2007-05-08 ·Pages 17

Wu S, Skolnick J, Zhang Y

Abstract

Predicting 3-dimensional protein structures from amino-acid sequences is an important unsolved problem in computational structural biology. The problem becomes relatively easier if close homologous proteins have been solved, as high-resolution models can be built by aligning target sequences to the solved homologous structures. However, for sequences without similar folds in the Protein Data Bank (PDB) library, the models have to be predicted from scratch. Progress in the ab initio structure modeling is slow. The aim of this study was to extend the TASSER (threading/assembly/refinement) method for the ab initio modeling and examine systemically its ability to fold small single-domain proteins. We developed I-TASSER by iteratively implementing the TASSER method, which is used in the folding test of three benchmarks of small proteins. First, data on 16 small proteins (< 90 residues) were used to generate I-TASSER models, which had an average Calpha-root mean square deviation (RMSD) of 3.8A, with 6 of them having a Calpha-RMSD < 2.5A. The overall result was comparable with the all-atomic ROSETTA simulation, but the central processing unit (CPU) time by I-TASSER was much shorter (150 CPU days vs. 5 CPU hours). Second, data on 20 small proteins (< 120 residues) were used. I-TASSER folded four of them with a Calpha-RMSD < 2.5A. The average Calpha-RMSD of the I-TASSER models was 3.9A, whereas it was 5.9A using TOUCHSTONE-II software. Finally, 20 non-homologous small proteins (< 120 residues) were taken from the PDB library. An average Calpha-RMSD of 3.9A was obtained for the third benchmark, with seven cases having a Calpha-RMSD < 2.5A. Our simulation results show that I-TASSER can consistently predict the correct folds and sometimes high-resolution models for small single-domain proteins. Compared with other ab initio modeling methods such as ROSETTA and TOUCHSTONE II, the average performance of I-TASSER is either much better or is similar within a lower computational time. These data, together with the significant performance of automated I-TASSER server (the Zhang-Server) in the 'free modeling' section of the recent Critical Assessment of Structure Prediction (CASP)7 experiment, demonstrate new progresses in automated ab initio model generation. The I-TASSER server is freely available for academic users http://zhang.bioinformatics.ku.edu/I-TASSER.

MeSH Terms
Computational Biology/methods Databases, Protein Models, Theoretical Peptide Library Protein Structure, Tertiary Proteomics
Chemicals
Peptide Library
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Wu Sitao
Center for Bioinformatics and Department of Molecular Bioscience, University of Kansas, Lawrence, KS 66047, USA. stwu@ku.edu
Skolnick Jeffrey
Zhang Yang
References (41)
41 references, click to expand
  1. Protein structure prediction by global optimization of a potential energy function.
    Proc Natl Acad Sci U S A. 1999 May 11;96(10):5482-5 PMID: 10318909
  2. Improved recognition of native-like protein structures using a combination of sequence-dependent and sequence-independent features of proteins.
    Proteins. 1999 Jan 1;34(1):82-95 PMID: 10336385
  3. Protein secondary structure prediction based on position-specific scoring matrices.
    J Mol Biol. 1999 Sep 17;292(2):195-202 PMID: 10493868
  4. The Protein Data Bank.
    Nucleic Acids Res. 2000 Jan 1;28(1):235-42 PMID: 10592235
  5. Structural genomics and its importance for gene function analysis.
    Nat Biotechnol. 2000 Mar;18(3):283-7 PMID: 10700142
  6. Structure-based evaluation of sequence comparison and fold recognition alignment accuracy.
    J Mol Biol. 2000 Apr 7;297(4):1003-13 PMID: 10736233
  7. Protein threading using PROSPECT: design and evaluation.
    Proteins. 2000 Aug 15;40(3):343-54 PMID: 10861926
  8. Accurate reconstruction of all-atom protein representations from side-chain-based low-resolution models.
    Proteins. 2000 Oct 1;41(1):86-97 PMID: 10944396
  9. Modeling of loops in protein structures.
    Protein Sci. 2000 Sep;9(9):1753-73 PMID: 11045621
  10. Prospects for ab initio protein structural genomics.
    J Mol Biol. 2001 Mar 9;306(5):1191-9 PMID: 11237627
  11. Protein structure prediction and structural genomics.
    Science. 2001 Oct 5;294(5540):93-6 PMID: 11588250
  12. Local energy landscape flattening: parallel hyperbolic Monte Carlo sampling of protein folding.
    Proteins. 2002 Aug 1;48(2):192-201 PMID: 12112688
  13. Real value prediction of solvent accessibility from amino acid sequence.
    Proteins. 2003 Mar 1;50(4):629-35 PMID: 12577269
  14. TOUCHSTONE II: a new approach to ab initio protein structure prediction.
    Biophys J. 2003 Aug;85(2):1145-64 PMID: 12885659
  15. A graph-theory algorithm for rapid protein side-chain prediction.
    Protein Sci. 2003 Sep;12(9):2001-14 PMID: 12930999
  16. ASTRO-FOLD: a combinatorial and global optimization framework for Ab initio prediction of three-dimensional structures of proteins from the amino acid sequence.
    Biophys J. 2003 Oct;85(4):2119-46 PMID: 14507680
  17. Using multiple structure alignments, fast model building, and energetic analysis in fold recognition and homology modeling.
    Proteins. 2003;53 Suppl 6:430-5 PMID: 14579332
  18. SPICKER: a clustering approach to identify near-native protein folds.
    J Comput Chem. 2004 Apr 30;25(6):865-71 PMID: 15011258
  19. Automated structure prediction of weakly homologous proteins on a genomic scale.
    Proc Natl Acad Sci U S A. 2004 May 18;101(20):7594-9 PMID: 15126668
  20. Development and large scale benchmark testing of the PROSPECTOR_3 threading algorithm.
    Proteins. 2004 Aug 15;56(3):502-18 PMID: 15229883
  21. Tertiary structure predictions on a comprehensive benchmark of medium to large size proteins.
    Biophys J. 2004 Oct;87(4):2647-55 PMID: 15454459
  22. Scoring function for automated assessment of protein structure template quality.
    Proteins. 2004 Dec 1;57(4):702-10 PMID: 15476259
  23. Fold recognition by combining sequence profiles derived from evolution and from depth-dependent structural alignment of fragments.
    Proteins. 2005 Feb 1;58(2):321-8 PMID: 15523666
  24. Ab initio prediction of the three-dimensional structure of a de novo designed protein: a double-blind case study.
    Proteins. 2005 Feb 15;58(3):560-70 PMID: 15609306
  25. The protein structure prediction problem could be solved using the current PDB library.
    Proc Natl Acad Sci U S A. 2005 Jan 25;102(4):1029-34 PMID: 15653774
  26. TM-align: a protein structure alignment algorithm based on the TM-score.
    Nucleic Acids Res. 2005 Apr 22;33(7):2302-9 PMID: 15849316
  27. Prediction of solvent accessibility and sites of deleterious mutations from protein sequence.
    Nucleic Acids Res. 2005 Jun 03;33(10):3193-9 PMID: 15937195
  28. A new approach to protein fold recognition.
    Nature. 1992 Jul 2;358(6381):86-9 PMID: 1614539
  29. Toward high-resolution de novo structure prediction for small proteins.
    Science. 2005 Sep 16;309(5742):1868-71 PMID: 16166519
  30. Assessment of predictions submitted for the CASP6 comparative modeling category.
    Proteins. 2005;61 Suppl 7:27-45 PMID: 16187345
  31. TASSER: an automated method for the prediction of protein tertiary structures in CASP6.
    Proteins. 2005;61 Suppl 7:91-8 PMID: 16187349
  32. On the origin and highly likely completeness of single-domain protein structures.
    Proc Natl Acad Sci U S A. 2006 Feb 21;103(8):2605-10 PMID: 16478803
  33. A method to identify protein sequences that fold into a known three-dimensional structure.
    Science. 1991 Jul 12;253(5016):164-70 PMID: 1853201
  34. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  35. Prediction of protein antigenic determinants from amino acid sequences.
    Proc Natl Acad Sci U S A. 1981 Jun;78(6):3824-8 PMID: 6167991
  36. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features.
    Biopolymers. 1983 Dec;22(12):2577-637 PMID: 6667333
  37. A simple method for displaying the hydropathic character of a protein.
    J Mol Biol. 1982 May 5;157(1):105-32 PMID: 7108955
  38. Comparative protein modelling by satisfaction of spatial restraints.
    J Mol Biol. 1993 Dec 5;234(3):779-815 PMID: 8254673
  39. Knowledge-based protein secondary structure assignment.
    Proteins. 1995 Dec;23(4):566-79 PMID: 8749853
  40. Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions.
    J Mol Biol. 1997 Apr 25;268(1):209-25 PMID: 9149153
  41. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
Article Info
Journal
BMC biology
Abbr.
BMC Biol
ISSN
1741-7007
Published
2007-05-08
Epub
2007-00-08
Pages
17
Language
English
Region
England
NLM ID
101190720
PMCID
PMC1878469
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com