Home LiteratureArticle Details
PMID: 20360767 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

I-TASSER: a unified platform for automated protein structure and function prediction.

Nature protocols ·Vol. 5 ·No. 4 ·2010-04-00 ·Pages 725-38

Roy A, Kucukural A, Zhang Y

Abstract

The iterative threading assembly refinement (I-TASSER) server is an integrated platform for automated protein structure and function prediction based on the sequence-to-structure-to-function paradigm. Starting from an amino acid sequence, I-TASSER first generates three-dimensional (3D) atomic models from multiple threading alignments and iterative structural assembly simulations. The function of the protein is then inferred by structurally matching the 3D models with other known proteins. The output from a typical server run contains full-length secondary and tertiary structure predictions, and functional annotations on ligand-binding sites, Enzyme Commission numbers and Gene Ontology terms. An estimate of accuracy of the predictions is provided based on the confidence score of the modeling. This protocol provides new insights and guidelines for designing of online server systems for the state-of-the-art protein structure and function predictions. The server is available at http://zhanglab.ccmb.med.umich.edu/I-TASSER.

MeSH Terms
Amino Acid Sequence Computer Simulation Databases, Protein/statistics & numerical data Internet Models, Molecular Online Systems Proteins/chemistry,genetics,physiology Sequence Alignment/statistics & numerical data Structural Homology, Protein
Chemicals
Proteins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Roy Ambrish
Center for Computational Medicine and Bioinformatics, University of Michigan, Ann Arbor, Michigan, USA.
Kucukural Alper
Zhang Yang
References (69)
69 references, click to expand
  1. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  2. Superfamily assignments for the yeast proteome through integration of structure prediction with the gene ontology.
    PLoS Biol. 2007 Apr;5(4):e76 PMID: 17373854
  3. Hidden Markov models for detecting remote protein homologies.
    Bioinformatics. 1998;14(10):846-56 PMID: 9927713
  4. Local energy landscape flattening: parallel hyperbolic Monte Carlo sampling of protein folding.
    Proteins. 2002 Aug 1;48(2):192-201 PMID: 12112688
  5. Identification and analysis of deleterious human SNPs.
    J Mol Biol. 2006 Mar 10;356(5):1263-74 PMID: 16412461
  6. Evaluation of template-based models in CASP8 with standard measures.
    Proteins. 2009;77 Suppl 9:18-28 PMID: 19731382
  7. LOMETS: a local meta-threading-server for protein structure prediction.
    Nucleic Acids Res. 2007;35(10):3375-82 PMID: 17478507
  8. Molecular and structural basis of drift in the functions of closely-related homologous enzyme domains: implications for function annotation based on homology searches and structural genomics.
    In Silico Biol. 2009;9(1-2):S41-55 PMID: 19537164
  9. The other 90% of the protein: assessment beyond the Calphas for CASP8 template-based and high-accuracy models.
    Proteins. 2009;77 Suppl 9:29-49 PMID: 19731372
  10. REMO: A new protocol to refine full atomic protein models from C-alpha traces by optimizing hydrogen-bonding networks.
    Proteins. 2009 Aug 15;76(3):665-76 PMID: 19274737
  11. Assessment of predictions submitted for the CASP7 function prediction category.
    Proteins. 2007;69 Suppl 8:165-74 PMID: 17654548
  12. Automated structure prediction of weakly homologous proteins on a genomic scale.
    Proc Natl Acad Sci U S A. 2004 May 18;101(20):7594-9 PMID: 15126668
  13. Protein homology detection by HMM-HMM comparison.
    Bioinformatics. 2005 Apr 1;21(7):951-60 PMID: 15531603
  14. Template-based modeling and free modeling by I-TASSER in CASP7.
    Proteins. 2007;69 Suppl 8:108-17 PMID: 17894355
  15. A random mutagenesis approach to isolate dominant-negative yeast sec1 mutants reveals a functional role for domain 3a in yeast and mammalian Sec1/Munc18 proteins.
    Genetics. 2008 Sep;180(1):165-78 PMID: 18757920
  16. Protein structure prediction and analysis using the Robetta server.
    Nucleic Acids Res. 2004 Jul 1;32(Web Server issue):W526-31 PMID: 15215442
  17. Comparative modeling in structural genomics.
    Structure. 2008 Jan;16(1):14-6 PMID: 18184577
  18. A method to identify protein sequences that fold into a known three-dimensional structure.
    Science. 1991 Jul 12;253(5016):164-70 PMID: 1853201
  19. Protein structure prediction on the Web: a case study using the Phyre server.
    Nat Protoc. 2009;4(3):363-71 PMID: 19247286
  20. Assessment of CASP7 predictions for template-based modeling targets.
    Proteins. 2007;69 Suppl 8:38-56 PMID: 17894352
  21. Q-Dock: Low-resolution flexible ligand docking with pocket-specific threading restraints.
    J Comput Chem. 2008 Jul 30;29(10):1574-88 PMID: 18293308
  22. Structure prediction for CASP7 targets using extensive all-atom refinement with Rosetta@home.
    Proteins. 2007;69 Suppl 8:118-28 PMID: 17894356
  23. Critical assessment of methods of protein structure prediction-Round VII.
    Proteins. 2007;69 Suppl 8:3-9 PMID: 17918729
  24. Improvement of the GenTHREADER method for genomic fold recognition.
    Bioinformatics. 2003 May 1;19(7):874-81 PMID: 12724298
  25. Comparative protein structure modeling of genes and genomes.
    Annu Rev Biophys Biomol Struct. 2000;29:291-325 PMID: 10940251
  26. Protein secondary structure prediction based on position-specific scoring matrices.
    J Mol Biol. 1999 Sep 17;292(2):195-202 PMID: 10493868
  27. Automated server predictions in CASP7.
    Proteins. 2007;69 Suppl 8:68-82 PMID: 17894354
  28. Comparative protein modelling by satisfaction of spatial restraints.
    J Mol Biol. 1993 Dec 5;234(3):779-815 PMID: 8254673
  29. FUGUE: sequence-structure homology recognition using environment-specific substitution tables and structure-dependent gap penalties.
    J Mol Biol. 2001 Jun 29;310(1):243-57 PMID: 11419950
  30. Protein structure prediction: when is it useful?
    Curr Opin Struct Biol. 2009 Apr;19(2):145-55 PMID: 19327982
  31. Assessment of CASP7 structure predictions for template free targets.
    Proteins. 2007;69 Suppl 8:57-67 PMID: 17894330
  32. 3D-Jury: a simple approach to improve protein structure predictions.
    Bioinformatics. 2003 May 22;19(8):1015-8 PMID: 12761065
  33. Modeling and analyzing three-dimensional structures of human disease proteins.
    Pac Symp Biocomput. 2006;:439-50 PMID: 17094259
  34. Prediction of solvent accessibility and sites of deleterious mutations from protein sequence.
    Nucleic Acids Res. 2005 Jun 03;33(10):3193-9 PMID: 15937195
  35. Structure modeling of all identified G protein-coupled receptors in the human genome.
    PLoS Comput Biol. 2006 Feb;2(2):e13 PMID: 16485037
  36. Analysis of TASSER-based CASP7 protein structure prediction results.
    Proteins. 2007;69 Suppl 8:90-7 PMID: 17705276
  37. Progress and challenges in protein structure prediction.
    Curr Opin Struct Biol. 2008 Jun;18(3):342-8 PMID: 18436442
  38. Assessment of predictions submitted for the CASP6 comparative modeling category.
    Proteins. 2005;61 Suppl 7:27-45 PMID: 16187345
  39. Single-body residue-level knowledge-based energy score combined with sequence-profile and secondary structure information for fold recognition.
    Proteins. 2004 Jun 1;55(4):1005-13 PMID: 15146497
  40. Scoring function for automated assessment of protein structure template quality.
    Proteins. 2004 Dec 1;57(4):702-10 PMID: 15476259
  41. The Protein Data Bank.
    Nucleic Acids Res. 2000 Jan 1;28(1):235-42 PMID: 10592235
  42. I-TASSER server for protein 3D structure prediction.
    BMC Bioinformatics. 2008 Jan 23;9:40 PMID: 18215316
  43. Assessment of predictions submitted for the CASP7 domain prediction category.
    Proteins. 2007;69 Suppl 8:137-51 PMID: 17680686
  44. Convergent evolution of similar enzymatic function on different protein folds: the hexokinase, ribokinase, and galactokinase families of sugar kinases.
    Protein Sci. 1993 Jan;2(1):31-40 PMID: 8382990
  45. Fold recognition by combining sequence profiles derived from evolution and from depth-dependent structural alignment of fragments.
    Proteins. 2005 Feb 1;58(2):321-8 PMID: 15523666
  46. Protein structure prediction by global optimization of a potential energy function.
    Proc Natl Acad Sci U S A. 1999 May 11;96(10):5482-5 PMID: 10318909
  47. On the origin and highly likely completeness of single-domain protein structures.
    Proc Natl Acad Sci U S A. 2006 Feb 21;103(8):2605-10 PMID: 16478803
  48. A comprehensive assessment of sequence-based and template-based methods for protein contact prediction.
    Bioinformatics. 2008 Apr 1;24(7):924-31 PMID: 18296462
  49. SPICKER: a clustering approach to identify near-native protein folds.
    J Comput Chem. 2004 Apr 30;25(6):865-71 PMID: 15011258
  50. TOUCHSTONE II: a new approach to ab initio protein structure prediction.
    Biophys J. 2003 Aug;85(2):1145-64 PMID: 12885659
  51. Universal similarity measure for comparing protein structures.
    Biopolymers. 2001 Oct 15;59(5):305-9 PMID: 11514933
  52. Application of sparse NMR restraints to large-scale protein structure prediction.
    Biophys J. 2004 Aug;87(2):1241-8 PMID: 15298926
  53. I-TASSER: fully automated protein structure prediction in CASP8.
    Proteins. 2009;77 Suppl 9:100-13 PMID: 19768687
  54. Tertiary structure predictions on a comprehensive benchmark of medium to large size proteins.
    Biophys J. 2004 Oct;87(4):2647-55 PMID: 15454459
  55. A threading-based method (FINDSITE) for ligand-binding site prediction and functional annotation.
    Proc Natl Acad Sci U S A. 2008 Jan 8;105(1):129-34 PMID: 18165317
  56. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  57. An integrated in silico 3D model-driven discovery of a novel, potent, and selective amidosulfonamide 5-HT1A agonist (PRX-00023) for the treatment of anxiety and depression.
    J Med Chem. 2006 Jun 1;49(11):3116-35 PMID: 16722631
  58. Protein threading using PROSPECT: design and evaluation.
    Proteins. 2000 Aug 15;40(3):343-54 PMID: 10861926
  59. In silico pharmacology for drug discovery: applications to targets and beyond.
    Br J Pharmacol. 2007 Sep;152(1):21-37 PMID: 17549046
  60. TM-align: a protein structure alignment algorithm based on the TM-score.
    Nucleic Acids Res. 2005 Apr 22;33(7):2302-9 PMID: 15849316
  61. A new approach to protein fold recognition.
    Nature. 1992 Jul 2;358(6381):86-9 PMID: 1614539
  62. Assembly of protein tertiary structures from fragments with similar local sequences using simulated annealing and Bayesian scoring functions.
    J Mol Biol. 1997 Apr 25;268(1):209-25 PMID: 9149153
  63. Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (NC-IUBMB). Enzyme Nomenclature. Recommendations 1992. Supplement 4: corrections and additions (1997).
    Eur J Biochem. 1997 Nov 15;250(1):1-6 PMID: 9431984
  64. Pcons5: combining consensus, structural evaluation and fold recognition scores.
    Bioinformatics. 2005 Dec 1;21(23):4248-54 PMID: 16204344
  65. Large-scale assessment of the utility of low-resolution protein structures for biochemical function assignment.
    Bioinformatics. 2004 May 1;20(7):1087-96 PMID: 14764543
  66. 3D-SHOTGUN: a novel, cooperative, fold-recognition meta-predictor.
    Proteins. 2003 May 15;51(3):434-41 PMID: 12696054
  67. The PredictProtein server.
    Nucleic Acids Res. 2004 Jul 1;32(Web Server issue):W321-6 PMID: 15215403
  68. MUSTER: Improving protein sequence profile-profile alignments by using multiple sources of structure information.
    Proteins. 2008 Aug;72(2):547-56 PMID: 18247410
  69. Ab initio modeling of small proteins by iterative TASSER simulations.
    BMC Biol. 2007 May 08;5:17 PMID: 17488521
Article Info
Journal
Nature protocols
Abbr.
Nat Protoc
ISSN
1750-2799
Published
2010-04-00
Epub
2010-00-25
Pages
725-38
Language
English
Region
England
NLM ID
101284307
PMCID
PMC2849174
Subset
IM
Grants
NIGMS NIH HHS · R01 GM084222 · United States
NIGMS NIH HHS · R01 GM083107-02 · United States
NIGMS NIH HHS · GM083107 · United States
NIGMS NIH HHS · R01 GM083107 · United States
NIGMS NIH HHS · R01 GM084222-01A1 · United States
NIGMS NIH HHS · GM084222 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com