Home LiteratureArticle Details
PMID: 9159160 Published · ppublish English Comparative Study Journal Article

Positionally cloned human disease genes: patterns of evolutionary conservation and functional motifs.

Mushegian AR, Bassett DE, Boguski MS, Bork P, Koonin EV

Abstract

Positional cloning has already produced the sequences of more than 70 human genes associated with specific diseases. In addition to their medical importance, these genes are of interest as a set of human genes isolated solely on the basis of the phenotypic effect of the respective mutations. We analyzed the protein sequences encoded by the positionally cloned disease genes using an iterative strategy combining several sensitive computer methods. Comparisons to complete sequence databases and to separate databases of nematode, yeast, and bacterial proteins showed that for most of the disease gene products, statistically significant sequence similarities are detectable in each of the model organisms. Only the nematode genome encodes apparent orthologs with conserved domain architecture for the majority of the disease genes. In yeast and bacterial homologs, domain organization is typically not conserved, and sequence similarity is limited to individual domains. Generally, human genes complement mutations only in orthologous yeast genes. Most of the positionally cloned genes encode large proteins with several globular and nonglobular domains, the functions of some or all of which are not known. We detected conserved domains and motifs not described previously in a number of proteins encoded by disease genes and predicted functions for some of them. These predictions include an ATP-binding domain in the product of hereditary nonpolyposis colon cancer gene (a MutL homolog), which is conserved in the HS90 family of chaperone proteins, type II DNA topoisomerases, and histidine kinases, and a nuclease domain homologous to bacterial RNase D and the 3'-5' exonuclease domain of DNA polymerase I in the Werner syndrome gene product.

MeSH Terms
Amino Acid Sequence Animals Bacteria/genetics Biological Evolution Cloning, Molecular Colorectal Neoplasms, Hereditary Nonpolyposis/genetics Conserved Sequence DNA Polymerase I/chemistry,genetics DNA Topoisomerases, Type II/chemistry,genetics Genetic Diseases, Inborn/genetics HSP90 Heat-Shock Proteins/chemistry,genetics Histidine Kinase Humans Information Systems Molecular Sequence Data Mutation Nematoda/genetics Protein Biosynthesis Protein Kinases/chemistry,genetics Proteins/chemistry,genetics Saccharomyces cerevisiae/genetics Sequence Homology, Amino Acid
Chemicals
HSP90 Heat-Shock Proteins Proteins Protein Kinases Histidine Kinase DNA Polymerase I DNA Topoisomerases, Type II
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Mushegian A R
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA.
Bassett D E
Boguski M S
Bork P
Koonin E V
References (45)
45 references, click to expand
  1. Distinguishing homologous from analogous proteins.
    Syst Zool. 1970 Jun;19(2):99-113 PMID: 5449325
  2. Local alignment statistics.
    Methods Enzymol. 1996;266:460-80 PMID: 8743700
  3. Structure of large fragment of Escherichia coli DNA polymerase I complexed with dTMP.
    Nature. 1985 Feb 28-Mar 6;313(6005):762-6 PMID: 3883192
  4. Protein phosphorylation and regulation of adaptive responses in bacteria.
    Microbiol Rev. 1989 Dec;53(4):450-90 PMID: 2556636
  5. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  6. The 3'-5' exonuclease of DNA polymerase I of Escherichia coli: contribution of each amino acid at the active site to the reaction.
    EMBO J. 1991 Jan;10(1):17-24 PMID: 1989882
  7. Analysis of compositionally biased regions in sequence databases.
    Methods Enzymol. 1996;266:554-71 PMID: 8743706
  8. Cloning the gene for Werner syndrome: a disease with many symptoms of premature aging.
    Trends Genet. 1996 Aug;12(8):283-6 PMID: 8783933
  9. Protein sequence motifs.
    Curr Opin Struct Biol. 1996 Jun;6(3):366-76 PMID: 8804823
  10. Metabolism and evolution of Haemophilus influenzae deduced from a whole-genome comparison with Escherichia coli.
    Curr Biol. 1996 Mar 1;6(3):279-91 PMID: 8805245
  11. Sequence analysis of eukaryotic developmental proteins: ancient and novel domains.
    Genetics. 1996 Oct;144(2):817-28 PMID: 8889542
  12. Comparative analysis of 1196 orthologous mouse and human full-length mRNA and protein sequences.
    Genome Res. 1996 Sep;6(9):846-57 PMID: 8889551
  13. Complete genome sequences of cellular life forms: glimpses of theoretical evolutionary genomics.
    Curr Opin Genet Dev. 1996 Dec;6(6):757-62 PMID: 8994848
  14. The structure of a domain common to archaebacteria and the homocystinuria disease protein.
    Trends Biochem Sci. 1997 Jan;22(1):12-3 PMID: 9020585
  15. Identification of prokaryotic and eukaryotic signal peptides and prediction of their cleavage sites.
    Protein Eng. 1997 Jan;10(1):1-6 PMID: 9051728
  16. The Protein Data Bank: a computer-based archival file for macromolecular structures.
    J Mol Biol. 1977 May 25;112(3):535-42 PMID: 875032
  17. Structural basis for the 3'-5' exonuclease activity of Escherichia coli DNA polymerase I: a two metal ion mechanism.
    EMBO J. 1991 Jan;10(1):25-33 PMID: 1989886
  18. A workbench for multiple alignment construction and analysis.
    Proteins. 1991;9(3):180-90 PMID: 2006136
  19. Predicting coiled coils from protein sequences.
    Science. 1991 May 24;252(5009):1162-4 PMID: 2031185
  20. Crystal structure of an N-terminal fragment of the DNA gyrase B protein.
    Nature. 1991 Jun 20;351(6328):624-9 PMID: 1646964
  21. A general structure for DNA-dependent DNA polymerases.
    Gene. 1991 Apr;100:27-38 PMID: 2055476
  22. Purification and phosphorylation of the Arc regulatory components of Escherichia coli.
    J Bacteriol. 1992 Sep;174(17):5617-23 PMID: 1512197
  23. Cloning and characterization of the cDNA coding for a polymyositis-scleroderma overlap syndrome-related nucleolar 100-kD protein.
    J Exp Med. 1992 Oct 1;176(4):973-80 PMID: 1383382
  24. Hsp90 chaperonins possess ATPase activity and bind heat shock transcription factors and peptidyl prolyl isomerases.
    J Biol Chem. 1993 Jan 15;268(2):1479-87 PMID: 8419347
  25. RNase T shares conserved sequence motifs with DNA proofreading exonucleases.
    Nucleic Acids Res. 1993 May 25;21(10):2521-2 PMID: 8506149
  26. Requirement of both kinase and phosphatase activities of an Escherichia coli receptor (Taz1) for ligand-dependent signal transduction.
    J Mol Biol. 1993 May 20;231(2):335-42 PMID: 8389884
  27. Molecular modelling of the Norrie disease protein predicts a cystine knot growth factor tertiary structure.
    Nat Genet. 1993 Dec;5(4):376-80 PMID: 8298646
  28. Issues in searching molecular sequence databases.
    Nat Genet. 1994 Feb;6(2):119-29 PMID: 8162065
  29. Combining evolutionary information and neural networks to predict protein secondary structure.
    Proteins. 1994 May;19(1):55-72 PMID: 8066087
  30. Detection of conserved segments in proteins: iterative scanning of sequence databases with alignment blocks.
    Proc Natl Acad Sci U S A. 1994 Dec 6;91(25):12091-5 PMID: 7991589
  31. Genes conserved in yeast and humans.
    Hum Mol Genet. 1994;3 Spec No:1509-17 PMID: 7849746
  32. Transmembrane helices predicted at 95% accuracy.
    Protein Sci. 1995 Mar;4(3):521-33 PMID: 7795533
  33. Positional cloning moves from perditional to traditional.
    Nat Genet. 1995 Apr;9(4):347-50 PMID: 7795639
  34. Comparative genomics, genome cross-referencing and XREFdb.
    Trends Genet. 1995 Sep;11(9):372-3 PMID: 7482790
  35. Threading analysis suggests that the obese gene product may be a helical cytokine.
    FEBS Lett. 1995 Oct 2;373(1):13-8 PMID: 7589424
  36. The multiplicity of domains in proteins.
    Annu Rev Biochem. 1995;64:287-314 PMID: 7574483
  37. The genome of Caenorhabditis elegans.
    Proc Natl Acad Sci U S A. 1995 Nov 21;92(24):10836-40 PMID: 7479894
  38. Maximum discrimination hidden Markov models of sequence consensus.
    J Comput Biol. 1995 Spring;2(1):9-23 PMID: 7497123
  39. The SWISS-PROT protein sequence data bank and its new supplement TREMBL.
    Nucleic Acids Res. 1996 Jan 1;24(1):21-5 PMID: 8594581
  40. Yeast genes and human disease.
    Nature. 1996 Feb 15;379(6566):589-90 PMID: 8628392
  41. Biochemistry and genetics of eukaryotic mismatch repair.
    Genes Dev. 1996 Jun 15;10(12):1433-42 PMID: 8666228
  42. BRCA1 protein products ... Functional motifs...
    Nat Genet. 1996 Jul;13(3):266-8 PMID: 8673121
  43. Dominant negative mutator mutations in the mutL gene of Escherichia coli.
    Nucleic Acids Res. 1996 Jul 1;24(13):2498-504 PMID: 8692687
  44. Entrez: molecular biology database and retrieval system.
    Methods Enzymol. 1996;266:141-62 PMID: 8743683
  45. Applying motif and profile searches.
    Methods Enzymol. 1996;266:162-84 PMID: 8743684
Article Info
Journal
Proceedings of the National Academy of Sciences of the United States of America
Abbr.
Proc Natl Acad Sci U S A
ISSN
0027-8424
Published
1997-05-27
Pages
5831-6
Language
English
Region
United States
NLM ID
7505876
PMCID
PMC20866
Subset
IM
Corrections
CommentIn
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com