Home LiteratureArticle Details
PMID: 15276848 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

Empirical analysis of protein insertions and deletions determining parameters for the correct placement of gaps in protein sequence alignments.

Journal of molecular biology ·Vol. 341 ·No. 2 ·2004-08-06 ·Pages 617-31

Chang MS, Benner SA

Abstract

To understand how protein segments are inserted and deleted during divergent evolution, a set of pairwise alignments contained exactly one gap, and therefore arising from the first insertion-deletion (indel) event in the time separating the homologs, was examined. The alignments showed that "structure breaking" amino acids (PGDNS) were preferred within and flanking gapped regions, as are two residues with hydrophilic side-chains (QE) that frequently occur at the surface of protein folds. Conversely, hydrophobic residues (FMILYVW) occur infrequently within and flanking the gapped region. These preferences are modestly different in protein pairs separated by an episode of adaptive evolution, than in pairs diverging under strong functional constraints. Surprisingly, regions near an indel have not evolved more rapidly than the sequence pair overall, showing no evidence that an indel event must be compensated by local amino acid replacement. The gap-lengths are best approximated by a Zipfian distribution, with the probability of a gap of length L decreasing as a function of L(-1.8). These features are largely independent of the length of the gap and the extent of divergence (measured by both silent and non-silent sequence changes) separating the two proteins. Surprisingly, amino acid repeats were discovered in more than a third of the polypeptide segments in and around the gap. These correspond to repeats in the DNA sequence. This suggests that a signature of the mechanism by which indels occur in the DNA sequence remains in the encoded protein sequences. These data suggest specific tools to score gap placement in an alignment. They also suggest tools that distinguish true indels from gaps created by mistaken gene finding, including under-predicted and over-predicted introns. By providing mechanisms to identify errors, the tools will enhance the value of genome sequence databases in support of integrated paleogenomics strategies used to extract functional information in a post-genomic environment.

MeSH Terms
Amino Acid Sequence Amino Acids/chemistry Animals Databases, Protein Genetic Variation Humans Models, Genetic Molecular Sequence Data Proteins/chemistry,genetics Sequence Alignment/methods Sequence Deletion
Chemicals
Amino Acids Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Chang Mike S S
Foundation for Applied Molecular Evolution, Gainesville, FL 32601, USA.
Benner Steven A
Article Info
Journal
Journal of molecular biology
Abbr.
J Mol Biol
ISSN
0022-2836
Published
2004-08-06
Pages
617-31
Language
English
Region
England
NLM ID
2985088R
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com