Home LiteratureArticle Details
PMID: 11829509 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

Structural clusters of evolutionary trace residues are statistically significant and common in proteins.

Journal of molecular biology ·Vol. 316 ·No. 1 ·2002-02-08 ·Pages 139-54

Madabushi S, Yao H, Marsh M, Kristensen DM, Philippi A, Sowa ME, Lichtarge O

Abstract

Given the massive increase in the number of new sequences and structures, a critical problem is how to integrate these raw data into meaningful biological information. One approach, the Evolutionary Trace, or ET, uses phylogenetic information to rank the residues in a protein sequence by evolutionary importance and then maps those ranked at the top onto a representative structure. If these residues form structural clusters, they can identify functional surfaces such as those involved in molecular recognition. Now that a number of examples have shown that ET can identify binding sites and focus mutational studies on their relevant functional determinants, we ask whether the method can be improved so as to be applicable on a large scale. To address this question, we introduce a new treatment of gaps resulting from insertions and deletions, which streamlines the selection of sequences used as input. We also introduce objective statistics to assess the significance of the total number of clusters and of the size of the largest one. As a result of the novel treatment of gaps, ET performance improves measurably. We find evolutionarily privileged clusters that are significant at the 5% level in 45 out of 46 (98%) proteins drawn from a variety of structural classes and biological functions. In 37 of the 38 proteins for which a protein-ligand complex is available, the dominant cluster contacts the ligand. We conclude that spatial clustering of evolutionarily important residues is a general phenomenon, consistent with the cooperative nature of residues that determine structure and function. In practice, these results suggest that ET can be applied on a large scale to identify functional sites in a significant fraction of the structures in the protein databank (PDB). This approach to combining raw sequences and structure to obtain detailed insights into the molecular basis of function should prove valuable in the context of the Structural Genomics Initiative.

MeSH Terms
Binding Sites Cluster Analysis Computational Biology/methods Databases, Protein Evolution, Molecular Ligands Models, Molecular Molecular Weight Phylogeny Protein Binding Protein Conformation Proteins/chemistry,metabolism Statistical Distributions Structure-Activity Relationship
Chemicals
Ligands Proteins
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Madabushi Srinivasan
Structural and Computational Biology and Molecular Biophysics Program, Baylor College of Medicine, Houston, TX 77030, USA.
Yao Hui
Marsh Mike
Kristensen David M
Philippi Anne
Sowa Mathew E
Lichtarge Olivier
Article Info
Journal
Journal of molecular biology
Abbr.
J Mol Biol
ISSN
0022-2836
Published
2002-02-08
Pages
139-54
Language
English
Region
England
NLM ID
2985088R
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com