Abstract
With a growing number of structures available in the Brookhaven Protein Data Bank, automatic methods for domain identification are required for the construction of databases. Domains are considered to be clusters of secondary structure elements. Thus, helices and strands are first clustered using intersecondary structural distances between C alpha positions, and dendrograms based on this distance measure are used to identify domains. Individual domains are recognized by a disjoint factor, which enables the automatic identification and classification into disjoint, interacting, and conjoint domains. Application to a database of 83 protein families and 18 unique structures shows that the approach provides an effective delineation of boundaries and identifies those proteins that can be considered as a single domain. A quantitative estimate of the interaction between domains has been proposed. The database of protein domains is a useful tool for understanding protein folding, for recognizing protein folds, and for understanding structure-activity relationships.
MeSH Terms
Algorithms
Aspartic Acid Endopeptidases/chemistry
Calmodulin/chemistry
Cluster Analysis
Databases, Factual
Hydroxymethylbilane Synthase/chemistry
Models, Chemical
Models, Molecular
Papain/chemistry
Porins/chemistry
Protein Structure, Secondary
Protein Structure, Tertiary
Sequence Alignment
Chemicals
Calmodulin
Porins
Hydroxymethylbilane Synthase
Papain
Aspartic Acid Endopeptidases
Endothia aspartic proteinase
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Sowdhamini R
Department of Crystallography, Birkbeck College, London, United Kingdom.
Blundell T L
References (32)
32 references, click to expand
-
Construction of phylogenetic trees.
Science. 1967 Jan 20;155(3760):279-84
PMID: 5334057
-
Parser for protein folding units.
Proteins. 1994 Jul;19(3):256-68
PMID: 7937738
-
Nucleation, rapid folding, and globular intrachain regions in proteins.
Proc Natl Acad Sci U S A. 1973 Mar;70(3):697-701
PMID: 4351801
-
Comparison of super-secondary structures in proteins.
J Mol Biol. 1973 May 15;76(2):241-56
PMID: 4737475
-
Structural patterns in globular proteins.
Nature. 1976 Jun 17;261(5561):552-8
PMID: 934293
-
The Protein Data Bank: a computer-based archival file for macromolecular structures.
J Mol Biol. 1977 May 25;112(3):535-42
PMID: 875032
-
On the conformation of proteins: towards the prediction of strand arrangements in beta-pleated sheets.
J Mol Biol. 1977 Jun 25;113(2):401-18
PMID: 886615
-
The tree structural organization of proteins.
J Mol Biol. 1978 Dec 15;126(3):315-32
PMID: 745231
-
Hierarchic organization of domains in globular proteins.
J Mol Biol. 1979 Nov 5;134(3):447-70
PMID: 537072
-
Correlation of DNA exonic regions with protein structural units in haemoglobin.
Nature. 1981 May 7;291(5810):90-2
PMID: 7231530
-
Location of structural domains in protein.
Biochemistry. 1981 Nov 10;20(23):6544-52
PMID: 7306523
-
Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features.
Biopolymers. 1983 Dec;22(12):2577-637
PMID: 6667333
-
Structure of papain refined at 1.65 A resolution.
J Mol Biol. 1984 Oct 25;179(2):233-56
PMID: 6502713
-
Domains in proteins: definitions, location, and structural principles.
Methods Enzymol. 1985;115:420-30
PMID: 4079796
-
Compact units in proteins.
Biochemistry. 1986 Sep 23;25(19):5759-65
PMID: 3778881
-
Refined structure of glutathione reductase at 1.54 A resolution.
J Mol Biol. 1987 Jun 5;195(3):701-29
PMID: 3656429
-
Elbow motion in the immunoglobulins involves a molecular ball-and-socket joint.
Nature. 1988 Sep 8;335(6186):188-90
PMID: 3412476
-
Prediction of the location of structural domains in globular proteins.
J Protein Chem. 1988 Aug;7(4):427-71
PMID: 3255372
-
Structure of a bacterial enzyme regulated by phosphorylation, isocitrate dehydrogenase.
Proc Natl Acad Sci U S A. 1989 Nov;86(22):8635-9
PMID: 2682654
-
X-ray analyses of aspartic proteinases. The three-dimensional structure at 2.1 A resolution of endothiapepsin.
J Mol Biol. 1990 Feb 20;211(4):919-41
PMID: 2179568
-
An investigation of oligopeptides linking domains in protein tertiary structures and possible candidates for general gene fusion.
J Mol Biol. 1990 Feb 20;211(4):943-58
PMID: 2313701
-
Domain flexibility in aspartic proteinases.
Proteins. 1992 Feb;12(2):158-70
PMID: 1603805
-
Structure of porphobilinogen deaminase reveals a flexible multidomain polymerase with a single catalytic site.
Nature. 1992 Sep 3;359(6390):33-9
PMID: 1522882
-
Structure of porin refined at 1.8 A resolution.
J Mol Biol. 1992 Sep 20;227(2):493-509
PMID: 1328651
-
Structure of a hinge-bending bacteriophage T4 lysozyme mutant, Ile3-->Pro.
J Mol Biol. 1992 Oct 5;227(3):917-33
PMID: 1404394
-
SETOR: hardware-lighted three-dimensional solid model representations of macromolecules.
J Mol Graph. 1993 Jun;11(2):134-8, 127-8
PMID: 8347566
-
Molecular recognition in protein families: a database of aligned three-dimensional structures of related proteins.
Biochem Soc Trans. 1993 Aug;21 ( Pt 3)(3):597-604
PMID: 8224474
-
Binary discontinuous compact protein domains.
Protein Eng. 1994 Mar;7(3):335-40
PMID: 8177882
-
Structural mechanisms for domain movements in proteins.
Biochemistry. 1994 Jun 7;33(22):6739-49
PMID: 8204609
-
Structure-based identification and clustering of protein families and superfamilies.
J Comput Aided Mol Des. 1994 Feb;8(1):5-27
PMID: 8035212
-
The closed conformation of a highly flexible protein: the structure of E. coli adenylate kinase with bound AMP and AMPPNP.
Proteins. 1994 Jul;19(3):183-98
PMID: 7937733
-
The three-dimensional structure of an enzyme molecule.
Sci Am. 1966 Nov;215(5):78-90
PMID: 5978599