Home LiteratureArticle Details
PMID: 25701568 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

AIDA: ab initio domain assembly for automated multi-domain protein structure prediction and domain-domain interaction prediction.

Bioinformatics (Oxford, England) ·Vol. 31 ·No. 13 ·2015-07-01 ·Pages 2098-105

Xu D, Jaroszewski L, Li Z, Godzik A

Abstract

Most proteins consist of multiple domains, independent structural and evolutionary units that are often reshuffled in genomic rearrangements to form new protein architectures. Template-based modeling methods can often detect homologous templates for individual domains, but templates that could be used to model the entire query protein are often not available. We have developed a fast docking algorithm ab initio domain assembly (AIDA) for assembling multi-domain protein structures, guided by the ab initio folding potential. This approach can be extended to discontinuous domains (i.e. domains with 'inserted' domains). When tested on experimentally solved structures of multi-domain proteins, the relative domain positions were accurately found among top 5000 models in 86% of cases. AIDA server can use domain assignments provided by the user or predict them from the provided sequence. The latter approach is particularly useful for automated protein structure prediction servers. The blind test consisting of 95 CASP10 targets shows that domain boundaries could be successfully determined for 97% of targets. The AIDA package as well as the benchmark sets used here are available for download at http://ffas.burnham.org/AIDA/. adam@sanfordburnham.org Supplementary data are available at Bioinformatics online.

MeSH Terms
Algorithms Computational Biology/methods Humans Internet Models, Theoretical Protein Conformation Protein Interaction Domains and Motifs Protein Structure, Tertiary Proteins/chemistry,metabolism Sequence Analysis, Protein Software
Chemicals
Proteins
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Xu Dong
Bioinformatics and Systems Biology Program, Sanford-Burnham Medical Research Institute, 10901 North Torrey Pines Road, La Jolla, CA 92037, USA, Center for Research in Biological Systems, University of California, San Diego, 9500 Gilman Dr. La Jolla, CA 92093-0446, USA and Center of Excellence in Genomic Medicine Research (CEGMR), King Fahad Medical Research Center, King Abdulaziz University, Jeddah, Kingdom of Saudi Arabia.
Jaroszewski Lukasz
Bioinformatics and Systems Biology Program, Sanford-Burnham Medical Research Institute, 10901 North Torrey Pines Road, La Jolla, CA 92037, USA, Center for Research in Biological Systems, University of California, San Diego, 9500 Gilman Dr. La Jolla, CA 92093-0446, USA and Center of Excellence in Genomic Medicine Research (CEGMR), King Fahad Medical Research Center, King Abdulaziz University, Jeddah, Kingdom of Saudi Arabia Bioinformatics and Systems Biology Program, Sanford-Burnham Medical Research Institute, 10901 North Torrey Pines Road, La Jolla, CA 92037, USA, Center for Research in Biological Systems, University of California, San Diego, 9500 Gilman Dr. La Jolla, CA 92093-0446, USA and Center of Excellence in Genomic Medicine Research (CEGMR), King Fahad Medical Research Center, King Abdulaziz University, Jeddah, Kingdom of Saudi Arabia.
Li Zhanwen
Bioinformatics and Systems Biology Program, Sanford-Burnham Medical Research Institute, 10901 North Torrey Pines Road, La Jolla, CA 92037, USA, Center for Research in Biological Systems, University of California, San Diego, 9500 Gilman Dr. La Jolla, CA 92093-0446, USA and Center of Excellence in Genomic Medicine Research (CEGMR), King Fahad Medical Research Center, King Abdulaziz University, Jeddah, Kingdom of Saudi Arabia.
Godzik Adam
Bioinformatics and Systems Biology Program, Sanford-Burnham Medical Research Institute, 10901 North Torrey Pines Road, La Jolla, CA 92037, USA, Center for Research in Biological Systems, University of California, San Diego, 9500 Gilman Dr. La Jolla, CA 92093-0446, USA and Center of Excellence in Genomic Medicine Research (CEGMR), King Fahad Medical Research Center, King Abdulaziz University, Jeddah, Kingdom of Saudi Arabia Bioinformatics and Systems Biology Program, Sanford-Burnham Medical Research Institute, 10901 North Torrey Pines Road, La Jolla, CA 92037, USA, Center for Research in Biological Systems, University of California, San Diego, 9500 Gilman Dr. La Jolla, CA 92093-0446, USA and Center of Excellence in Genomic Medicine Research (CEGMR), King Fahad Medical Research Center, King Abdulaziz University, Jeddah, Kingdom of Saudi Arabia Bioinformatics and Systems Biology Program, Sanford-Burnham Medical Research Institute, 10901 North Torrey Pines Road, La Jolla, CA 92037, USA, Center for Research in Biological Systems, University of California, San Diego, 9500 Gilman Dr. La Jolla, CA 92093-0446, USA and Center of Excellence in Genomic Medicine Research (CEGMR), King Fahad Medical Research Center, King Abdulaziz University, Jeddah, Kingdom of Saudi Arabia.
References (29)
29 references, click to expand
  1. Combinatorial docking approach for structure prediction of large proteins and multi-molecular assemblies.
    Phys Biol. 2005 Nov;2(4):S156-65 PMID: 16280621
  2. Docking to single-domain and multiple-domain proteins: old and new challenges.
    Proteins. 2005 Aug 1;60(2):195-201 PMID: 15981268
  3. Expansion of protein domain repeats.
    PLoS Comput Biol. 2006 Aug 25;2(8):e114 PMID: 16933986
  4. Prediction of structures of multidomain proteins from structures of the individual domains.
    Protein Sci. 2007 Feb;16(2):165-75 PMID: 17189483
  5. The folding and evolution of multidomain proteins.
    Nat Rev Mol Cell Biol. 2007 Apr;8(4):319-30 PMID: 17356578
  6. Structural assembly of two-domain proteins by rigid-body docking.
    BMC Bioinformatics. 2008;9:441 PMID: 18925951
  7. Protein structure prediction on the Web: a case study using the Phyre server.
    Nat Protoc. 2009;4(3):363-71 PMID: 19247286
  8. Ab initio protein structure assembly using continuous structure fragments and optimized knowledge-based force field.
    Proteins. 2012 Jul;80(7):1715-35 PMID: 22411565
  9. How significant is a protein structure similarity with TM-score = 0.5?
    Bioinformatics. 2010 Apr 1;26(7):889-95 PMID: 20164152
  10. Inferring protein-protein interactions from multiple protein domain combinations.
    Methods Mol Biol. 2009;541:43-59 PMID: 19381530
  11. This Déjà vu feeling--analysis of multidomain protein evolution in eukaryotic genomes.
    PLoS Comput Biol. 2012;8(11):e1002701 PMID: 23166479
  12. FFAS-3D: improving fold recognition by including optimized structural features and template re-ranking.
    Bioinformatics. 2014 Mar 1;30(5):660-7 PMID: 24130308
  13. AIDA: ab initio domain assembly server.
    Nucleic Acids Res. 2014 Jul;42(Web Server issue):W308-13 PMID: 24831546
  14. Improved prediction of protein side-chain conformations with SCWRL4.
    Proteins. 2009 Dec;77(4):778-95 PMID: 19603484
  15. The Protein Data Bank.
    Nucleic Acids Res. 2000 Jan 1;28(1):235-42 PMID: 10592235
  16. Protein domain decomposition using a graph-theoretic approach.
    Bioinformatics. 2000 Dec;16(12):1091-104 PMID: 11159328
  17. Distance-scaled, finite ideal-gas reference state improves structure-derived potentials of mean force for structure selection and stability prediction.
    Protein Sci. 2002 Nov;11(11):2714-26 PMID: 12381853
  18. Evolution of the protein repertoire.
    Science. 2003 Jun 13;300(5626):1701-3 PMID: 12805536
  19. PISCES: a protein sequence culling server.
    Bioinformatics. 2003 Aug 12;19(12):1589-91 PMID: 12912846
  20. Multi-domain protein families and domain pairs: comparison with known structures and a random model of domain recombination.
    J Struct Funct Genomics. 2003;4(2-3):67-78 PMID: 14649290
  21. Conformation of polypeptides and proteins.
    Adv Protein Chem. 1968;23:283-438 PMID: 4882249
  22. Comparative protein modelling by satisfaction of spatial restraints.
    J Mol Biol. 1993 Dec 5;234(3):779-815 PMID: 8254673
  23. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  24. Protein secondary structure prediction based on position-specific scoring matrices.
    J Mol Biol. 1999 Sep 17;292(2):195-202 PMID: 10493868
  25. Scoring function for automated assessment of protein structure template quality.
    Proteins. 2004 Dec 1;57(4):702-10 PMID: 15476259
  26. Multi-domain proteins in the three kingdoms of life: orphan domains and other unassigned regions.
    J Mol Biol. 2005 Apr 22;348(1):231-43 PMID: 15808866
  27. Protein length in eukaryotic and prokaryotic proteomes.
    Nucleic Acids Res. 2005;33(10):3390-400 PMID: 15951512
  28. FFAS03: a server for profile--profile sequence alignments.
    Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W284-8 PMID: 15980471
  29. Docking protein domains in contact space.
    BMC Bioinformatics. 2006;7:310 PMID: 16790041
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2015-07-01
Epub
2015-00-19
Pages
2098-105
Language
English
Region
England
NLM ID
9808944
PMCID
PMC4481839
Subset
IM
Grants
NIGMS NIH HHS · R01 GM101457 · United States
NIGMS NIH HHS · GM101457 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com