Abstract
Annotation transfer is a principal process in genome annotation. It involves "transferring" structural and functional annotation to uncharacterized open reading frames (ORFs) in a newly completed genome from experimentally characterized proteins similar in sequence. To prevent errors in genome annotation, it is important that this process be robust and statistically well-characterized, especially with regard to how it depends on the degree of sequence similarity. Previously, we and others have analyzed annotation transfer in single-domain proteins. Multi-domain proteins, which make up the bulk of the ORFs in eukaryotic genomes, present more complex issues in functional conservation. Here we present a large-scale survey of annotation transfer in these proteins, using scop superfamilies to define domain folds and a thesaurus based on SWISS-PROT keywords to define functional categories. Our survey reveals that multi-domain proteins have significantly less functional conservation than single-domain ones, except when they share the exact same combination of domain folds. In particular, we find that for multi-domain proteins, approximate function can be accurately transferred with only 35% certainty for pairs of proteins sharing one structural superfamily. In contrast, this value is 67% for pairs of single-domain proteins sharing the same structural superfamily. On the other hand, if two multi-domain proteins contain the same combination of two structural superfamilies the probability of their sharing the same function increases to 80% in the case of complete coverage along the full length of both proteins, this value increases further to > 90%. Moreover, we found that only 70 of the current total of 455 structural superfamilies are found in both single and multi-domain proteins and only 14 of these were associated with the same function in both categories of proteins. We also investigated the degree to which function could be transferred between pairs of multi-domain proteins with respect to the degree of sequence similarity between them, finding that functional divergence at a given amount of sequence similarity is always about two-fold greater for pairs of multi-domain proteins (sharing similarity over a single domain) in comparison to pairs of single-domain ones, though the overall shape of the relationship is quite similar. Further information is available at http://partslist.org/func or http://bioinfo.mbb.yale.edu/partslist/func.
MeSH Terms
Computational Biology/methods
Conserved Sequence
Databases, Factual
Genomics/methods
Protein Folding
Protein Structure, Secondary/physiology
Protein Structure, Tertiary/physiology
Quantitative Structure-Activity Relationship
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Hegyi H
Department of Molecular Biophysics and Biochemistry, Yale University, New Haven, Connecticut 06520, USA.
Gerstein M
References (28)
28 references, click to expand
-
Biological function made crystal clear - annotation of hypothetical proteins via structural genomics.
Curr Opin Biotechnol. 2000 Feb;11(1):25-30
PMID: 10679350
-
The ENZYME database in 2000.
Nucleic Acids Res. 2000 Jan 1;28(1):304-5
PMID: 10592255
-
Assessing annotation transfer for genomics: quantifying the relations between protein sequence, structure and function through traditional and probabilistic scores.
J Mol Biol. 2000 Mar 17;297(1):233-49
PMID: 10704319
-
Predicting protein function from structure: unique structural features of proteases.
Proc Natl Acad Sci U S A. 2000 Apr 11;97(8):3954-8
PMID: 10759560
-
Whole-genome trees based on the occurrence of folds and orthologs: implications for comparing genomes on different levels.
Genome Res. 2000 Jun;10(6):808-18
PMID: 10854412
-
Sensitive sequence comparison as protein function predictor.
Pac Symp Biocomput. 2000;:42-53
PMID: 10902155
-
Practical limits of function prediction.
Proteins. 2000 Oct 1;41(1):98-107
PMID: 10944397
-
A Bayesian system integrating expression data with sequence patterns for localizing proteins: comprehensive application to the yeast genome.
J Mol Biol. 2000 Aug 25;301(4):1059-75
PMID: 10966805
-
From structure to function: approaches and limitations.
Nat Struct Biol. 2000 Nov;7 Suppl:991-4
PMID: 11104008
-
Digging for dead genes: an analysis of the characteristics of the pseudogene population in the Caenorhabditis elegans genome.
Nucleic Acids Res. 2001 Feb 1;29(3):818-30
PMID: 11160906
-
Evolution of function in protein superfamilies, from a structural perspective.
J Mol Biol. 2001 Apr 6;307(4):1113-43
PMID: 11286560
-
PartsList: a web-based system for dynamically ranking protein folds based on disparate attributes, including whole-genome expression and interaction information.
Nucleic Acids Res. 2001 Apr 15;29(8):1750-64
PMID: 11292848
-
The relation between the divergence of sequence and structure in proteins.
EMBO J. 1986 Apr;5(4):823-6
PMID: 3709526
-
Using the FASTA program to search protein and DNA sequence databases.
Methods Mol Biol. 1994;25:365-89
PMID: 8004177
-
SCOP: a structural classification of proteins database for the investigation of sequences and structures.
J Mol Biol. 1995 Apr 7;247(4):536-40
PMID: 7723011
-
FlyBase: a Drosophila database. The FlyBase consortium.
Nucleic Acids Res. 1997 Jan 1;25(1):63-6
PMID: 9045212
-
Recognition of analogous and homologous protein folds: analysis of sequence and structure conservation.
J Mol Biol. 1997 Jun 13;269(3):423-39
PMID: 9199410
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694
-
Predicting enzyme function from sequence: a systematic appraisal.
Proc Int Conf Intell Syst Mol Biol. 1997;5:276-83
PMID: 9322050
-
A structural census of genomes: comparing bacterial, eukaryotic, and archaeal genomes in terms of protein structure.
J Mol Biol. 1997 Dec 12;274(4):562-76
PMID: 9417935
-
Protein folds and functions.
Structure. 1998 Jul 15;6(7):875-84
PMID: 9687369
-
The CATH Database provides insights into protein structure/function relationships.
Nucleic Acids Res. 1999 Jan 1;27(1):275-9
PMID: 9847200
-
How representative are the known structures of the proteins in a complete genome? A comprehensive structural census.
Fold Des. 1998;3(6):497-512
PMID: 9889159
-
The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.
J Mol Biol. 1999 Apr 23;288(1):147-64
PMID: 10329133
-
From fold predictions to function predictions: automation of functional site conservation analysis for functional genome predictions.
Protein Sci. 1999 May;8(5):1104-15
PMID: 10338021
-
Protein folds, functions and evolution.
J Mol Biol. 1999 Oct 22;293(2):333-42
PMID: 10529349
-
The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
Nucleic Acids Res. 2000 Jan 1;28(1):45-8
PMID: 10592178
-
Finding function through structural genomics.
Curr Opin Biotechnol. 2000 Feb;11(1):31-5
PMID: 10679341