Home LiteratureArticle Details
PMID: 19587024 Published · ppublish English Journal Article Research Support, N.I.H., Extramural

Alignment of multiple protein structures based on sequence and structure features.

Protein engineering, design & selection : PEDS ·Vol. 22 ·No. 9 ·2009-09-00 ·Pages 569-74

Madhusudhan MS, Webb BM, Marti-Renom MA, Eswar N, Sali A

Abstract

Comparing the structures of proteins is crucial to gaining insight into protein evolution and function. Here, we align the sequences of multiple protein structures by a dynamic programming optimization of a scoring function that is a sum of an affine gap penalty and terms dependent on various sequence and structure features (SALIGN). The features include amino acid residue type, residue position, residue accessible surface area, residue secondary structure state and the conformation of a short segment centered on the residue. The multiple alignment is built by following the 'guide' tree constructed from the matrix of all pairwise protein alignment scores. Importantly, the method does not depend on the exact values of various parameters, such as feature weights and gap penalties, because the optimal alignment across a range of parameter values is found. Using multiple structure alignments in the HOMSTRAD database, SALIGN was benchmarked against MUSTANG for multiple alignments as well as against TM-align and CE for pairwise alignments. On the average, SALIGN produces a 15% improvement in structural overlap over HOMSTRAD and 14% over MUSTANG, and yields more equivalent structural positions than TM-align and CE in 90% and 95% of cases, respectively. The utility of accurate multiple structure alignment is illustrated by its application to comparative protein structure modeling.

MeSH Terms
Algorithms Amino Acid Sequence Databases, Protein Protein Conformation Proteins/chemistry,genetics Sequence Alignment/methods Sequence Analysis, Protein
Chemicals
Proteins
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Madhusudhan M S
Department of Bioengineering and Therapeutic Sciences, University of California at San Francisco, San Francisco, CA 94158, USA.
Webb Benjamin M
Marti-Renom Marc A
Eswar Narayanan
Sali Andrej
References (39)
39 references, click to expand
  1. PALI: a database of alignments and phylogeny of homologous protein structures.
    Bioinformatics. 2001 Apr;17(4):375-6 PMID: 11301310
  2. MASS: multiple structural alignment by secondary structures.
    Bioinformatics. 2003;19 Suppl 1:i95-104 PMID: 12855444
  3. MAMMOTH (matching molecular models obtained from theory): an automated method for model comparison.
    Protein Sci. 2002 Nov;11(11):2606-21 PMID: 12381844
  4. HOMSTRAD: a database of protein structure alignments for homologous families.
    Protein Sci. 1998 Nov;7(11):2469-71 PMID: 9828015
  5. Protein structure alignment by incremental combinatorial extension (CE) of the optimal path.
    Protein Eng. 1998 Sep;11(9):739-47 PMID: 9796821
  6. Dali: a network tool for protein structure comparison.
    Trends Biochem Sci. 1995 Nov;20(11):478-80 PMID: 8578593
  7. OXBench: a benchmark for evaluation of protein multiple sequence alignment accuracy.
    BMC Bioinformatics. 2003 Oct 10;4:47 PMID: 14552658
  8. Comparative protein structure modeling by combining multiple templates and optimizing sequence-to-structure alignments.
    Bioinformatics. 2007 Oct 1;23(19):2558-65 PMID: 17823132
  9. A method for simultaneous alignment of multiple protein structures.
    Proteins. 2004 Jul 1;56(1):143-56 PMID: 15162494
  10. Multiple flexible structure alignment using partial order graphs.
    Bioinformatics. 2005 May 15;21(10):2362-9 PMID: 15746292
  11. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  12. Amino acid substitution matrices from protein blocks.
    Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9 PMID: 1438297
  13. Progressive combinatorial algorithm for multiple structural alignments: application to distantly related proteins.
    Proteins. 2004 May 1;55(2):436-54 PMID: 15048834
  14. Alignment of protein sequences by their profiles.
    Protein Sci. 2004 Apr;13(4):1071-87 PMID: 15044736
  15. The ASTRAL compendium for protein structure and sequence analysis.
    Nucleic Acids Res. 2000 Jan 1;28(1):254-6 PMID: 10592239
  16. CE-MC: a multiple protein structure alignment server.
    Nucleic Acids Res. 2004 Jul 1;32(Web Server issue):W100-3 PMID: 15215359
  17. Serine protease mechanism and specificity.
    Chem Rev. 2002 Dec;102(12):4501-24 PMID: 12475199
  18. Variable gap penalty for protein sequence-structure alignment.
    Protein Eng Des Sel. 2006 Mar;19(3):129-33 PMID: 16423846
  19. Using multiple templates to improve quality of homology models in automated homology modeling.
    Protein Sci. 2008 Jun;17(6):990-1002 PMID: 18441233
  20. Comparative protein modelling by satisfaction of spatial restraints.
    J Mol Biol. 1993 Dec 5;234(3):779-815 PMID: 8254673
  21. Protein structure comparison using iterated double dynamic programming.
    Protein Sci. 1999 Mar;8(3):654-65 PMID: 10091668
  22. PASS2: an automated database of protein alignments organised as structural superfamilies.
    BMC Bioinformatics. 2004 Apr 02;5:35 PMID: 15059245
  23. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  24. Definition of general topological equivalence in protein structures. A procedure involving comparison of properties and relationships through simulated annealing and dynamic programming.
    J Mol Biol. 1990 Mar 20;212(2):403-28 PMID: 2181150
  25. The PDB is a covering set of small protein structures.
    J Mol Biol. 2003 Dec 5;334(4):793-802 PMID: 14636603
  26. DBAli: a database of protein structure alignments.
    Bioinformatics. 2001 Aug;17(8):746-7 PMID: 11524379
  27. Progressive sequence alignment as a prerequisite to correct phylogenetic trees.
    J Mol Evol. 1987;25(4):351-60 PMID: 3118049
  28. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  29. Construction of phylogenetic trees.
    Science. 1967 Jan 20;155(3760):279-84 PMID: 5334057
  30. MUSTANG: a multiple structural alignment algorithm.
    Proteins. 2006 Aug 15;64(3):559-74 PMID: 16736488
  31. DBAli tools: mining the protein structure space.
    Nucleic Acids Res. 2007 Jul;35(Web Server issue):W393-7 PMID: 17478513
  32. Processing and evaluation of predictions in CASP4.
    Proteins. 2001;Suppl 5:13-21 PMID: 11835478
  33. CATH--a hierarchic classification of protein domain structures.
    Structure. 1997 Aug 15;5(8):1093-108 PMID: 9309224
  34. SUPERFAMILY--sophisticated comparative genomics, data mining, visualization and phylogeny.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D380-6 PMID: 19036790
  35. A new progressive-iterative algorithm for multiple structure alignment.
    Bioinformatics. 2005 Aug 1;21(15):3255-63 PMID: 15941743
  36. Knowledge based modelling of homologous proteins, Part I: Three-dimensional frameworks derived from the simultaneous superposition of multiple structures.
    Protein Eng. 1987 Oct-Nov;1(5):377-84 PMID: 3508286
  37. TM-align: a protein structure alignment algorithm based on the TM-score.
    Nucleic Acids Res. 2005 Apr 22;33(7):2302-9 PMID: 15849316
  38. Matt: local flexibility aids protein multiple structure alignment.
    PLoS Comput Biol. 2008 Jan;4(1):e10 PMID: 18193941
  39. Systematic analysis of the effect of multiple templates on the accuracy of comparative models of protein structure.
    BMC Struct Biol. 2008 Jul 16;8:31 PMID: 18631402
Article Info
Journal
Protein engineering, design & selection : PEDS
Abbr.
Protein Eng Des Sel
ISSN
1741-0134
Published
2009-09-00
Epub
2009-00-08
Pages
569-74
Language
English
Region
England
NLM ID
101186484
PMCID
PMC2909824
Subset
IM
Grants
NIGMS NIH HHS · R01 GM54762-11 · United States
NIGMS NIH HHS · U54 GM62529 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com