Home LiteratureArticle Details
PMID: 12410938 Published · ppublish English Journal Article

The relationship of protein conservation and sequence length.

BMC evolutionary biology ·Vol. 2 ·2002-11-01 ·Pages 20

Lipman DJ, Souvorov A, Koonin EV, Panchenko AR, Tatusova TA

Abstract

In general, the length of a protein sequence is determined by its function and the wide variance in the lengths of an organism's proteins reflects the diversity of specific functional roles for these proteins. However, additional evolutionary forces that affect the length of a protein may be revealed by studying the length distributions of proteins evolving under weaker functional constraints. We performed sequence comparisons to distinguish highly conserved and poorly conserved proteins from the bacterium Escherichia coli, the archaeon Archaeoglobus fulgidus, and the eukaryotes Saccharomyces cerevisiae, Drosophila melanogaster, and Homo sapiens. For all organisms studied, the conserved and nonconserved proteins have strikingly different length distributions. The conserved proteins are, on average, longer than the poorly conserved ones, and the length distributions for the poorly conserved proteins have a relatively narrow peak, in contrast to the conserved proteins whose lengths spread over a wider range of values. For the two prokaryotes studied, the poorly conserved proteins approximate the minimal length distribution expected for a diverse range of structural folds. There is a relationship between protein conservation and sequence length. For all the organisms studied, there seems to be a significant evolutionary trend favoring shorter proteins in the absence of other, more specific functional constraints.

MeSH Terms
Animals Archaeal Proteins/chemistry Archaeoglobus fulgidus/chemistry Conserved Sequence Drosophila Proteins/chemistry Escherichia coli Proteins/chemistry Evolution, Molecular Humans Molecular Sequence Data Molecular Weight Proteins/chemistry Saccharomyces cerevisiae Proteins/chemistry Structure-Activity Relationship
Chemicals
Archaeal Proteins Drosophila Proteins Escherichia coli Proteins Proteins Saccharomyces cerevisiae Proteins
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Lipman David J
National Center for Biotechnology Information, National Institutes of Health, Bethesda, MD 20894, USA. lipman@ncbi.nlm.nih.gov
Souvorov Alexander
Koonin Eugene V
Panchenko Anna R
Tatusova Tatiana A
References (17)
17 references, click to expand
  1. Nucleotide bias causes a genomewide bias in the amino acid composition of proteins.
    Mol Biol Evol. 2000 Nov;17(11):1581-8 PMID: 11070046
  2. Evolution of genome size: new approaches to an old problem.
    Trends Genet. 2001 Jan;17(1):23-8 PMID: 11163918
  3. A simple model based on mutation and selection explains trends in codon and amino-acid usage and GC composition within and across genomes.
    Genome Biol. 2001;2(4):RESEARCH0010 PMID: 11305938
  4. Protein dispensability and rate of evolution.
    Nature. 2001 Jun 28;411(6841):1046-9 PMID: 11429604
  5. On the total number of genes and their length distribution in complete microbial genomes.
    Trends Genet. 2001 Aug;17(8):425-8 PMID: 11485798
  6. Genome sequence and gene compaction of the eukaryote parasite Encephalitozoon cuniculi.
    Nature. 2001 Nov 22;414(6862):450-3 PMID: 11719806
  7. Molecular chaperones in the cytosol: from nascent chain to folded protein.
    Science. 2002 Mar 8;295(5561):1852-8 PMID: 11884745
  8. Metabolic efficiency and amino acid composition in the proteomes of Escherichia coli and Bacillus subtilis.
    Proc Natl Acad Sci U S A. 2002 Mar 19;99(6):3695-700 PMID: 11904428
  9. Essential genes are more evolutionarily conserved than are nonessential genes in bacteria.
    Genome Res. 2002 Jun;12(6):962-8 PMID: 12045149
  10. Selection for short introns in highly expressed genes.
    Nat Genet. 2002 Aug;31(4):415-8 PMID: 12134150
  11. Mechanisms of spontaneous mutation in DNA repair-proficient Escherichia coli.
    Mutat Res. 1991 Sep-Oct;250(1-2):55-71 PMID: 1944363
  12. Threading a database of protein cores.
    Proteins. 1995 Nov;23(3):356-69 PMID: 8710828
  13. Biology's new Rosetta stone.
    Nature. 1997 Jan 2;385(6611):29-30 PMID: 8985242
  14. Amelioration of bacterial genomes: rates of change and exchange.
    J Mol Evol. 1997 Apr;44(4):383-97 PMID: 9089078
  15. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  16. Identification of homologous core structures.
    Proteins. 1999 Apr 1;35(1):70-9 PMID: 10090287
  17. Domain size distributions can predict domain boundaries.
    Bioinformatics. 2000 Jul;16(7):613-8 PMID: 11038331
Article Info
Journal
BMC evolutionary biology
Abbr.
BMC Evol Biol
ISSN
1471-2148
Published
2002-11-01
Epub
2002-00-01
Pages
20
Language
English
Region
England
NLM ID
100966975
PMCID
PMC137605
Subset
IM
Databases
RefSeq
NC_000917, NC_001133, NC_001134, NC_001135, NC_001136, NC_001137, NC_001138, NC_001139, NC_001140, NC_001141, NC_001142, NC_001143, NC_001144, NC_001145, NC_001146, NC_001147, NC_001148, NC_001224, NC_001398, NC_002142
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com