Home LiteratureArticle Details
PMID: 9108146 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S. Research Support, U.S. Gov't, P.H.S.

Extracting protein alignment models from the sequence database.

Nucleic acids research ·Vol. 25 ·No. 9 ·1997-05-01 ·Pages 1665-77

Neuwald AF, Liu JS, Lipman DJ, Lawrence CE

Abstract

Biologists often gain structural and functional insights into a protein sequence by constructing a multiple alignment model of the family. Here a program called Probe fully automates this process of model construction starting from a single sequence. Central to this program is a powerful new method to locate and align only those, often subtly, conserved patterns essential to the family as a whole. When applied to randomly chosen proteins, Probe found on average about four times as many relationships as a pairwise search and yielded many new discoveries. These include: an obscure subfamily of globins in the roundworm Caenorhabditis elegans ; two new superfamilies of metallohydrolases; a lipoyl/biotin swinging arm domain in bacterial membrane fusion proteins; and a DH domain in the yeast Bud3 and Fus2 proteins. By identifying distant relationships and merging families into superfamilies in this way, this analysis further confirms the notion that proteins evolved from relatively few ancient sequences. Moreover, this method automatically generates models of these ancient conserved regions for rapid and sensitive screening of sequences.

MeSH Terms
Algorithms Amino Acid Sequence Animals Bacterial Proteins/chemistry Databases, Factual GTP-Binding Proteins/chemistry Models, Chemical Molecular Sequence Data Recombinant Fusion Proteins/chemistry Sequence Alignment Sequence Homology, Amino Acid
Chemicals
Bacterial Proteins Recombinant Fusion Proteins GTP-Binding Proteins
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Neuwald A F
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA. neuwald@ncbi.nlm.nih.gov
Liu J S
Lipman D J
Lawrence C E
References (82)
82 references, click to expand
  1. Determinants of a protein fold. Unique features of the globin amino acid sequences.
    J Mol Biol. 1987 Jul 5;196(1):199-216 PMID: 3656444
  2. BUD2 encodes a GTPase-activating protein for Bud1/Rsr1 necessary for proper bud-site selection in yeast.
    Nature. 1993 Sep 16;365(6443):269-74 PMID: 8371782
  3. Efflux pumps and drug resistance in gram-negative bacteria.
    Trends Microbiol. 1994 Dec;2(12):489-93 PMID: 7889326
  4. Biotin enzymes.
    Annu Rev Biochem. 1977;46:385-413 PMID: 20039
  5. Hidden Markov models in computational biology. Applications to protein modeling.
    J Mol Biol. 1994 Feb 4;235(5):1501-31 PMID: 8107089
  6. The Dbl family of oncogenes.
    Curr Opin Cell Biol. 1996 Apr;8(2):216-22 PMID: 8791419
  7. Subunit interactions in Ascaris hemoglobin octamer formation.
    J Biol Chem. 1995 Sep 22;270(38):22248-53 PMID: 7673204
  8. The structure of the Aeromonas proteolytica aminopeptidase complexed with a hydroxamate inhibitor. Involvement in catalysis of Glu151 and two zinc ions of the co-catalytic unit.
    Eur J Biochem. 1996 Apr 15;237(2):393-8 PMID: 8647077
  9. Crystal structure of Aeromonas proteolytica aminopeptidase: a prototypical member of the co-catalytic zinc enzyme family.
    Structure. 1994 Apr 15;2(4):283-91 PMID: 8087555
  10. Cell division. Bud-site selection is only skin deep.
    Curr Biol. 1995 Nov 1;5(11):1213-5 PMID: 8574570
  11. Characterization of zinc-binding sites in human stromelysin-1: stoichiometry of the catalytic domain and identification of a cysteine ligand in the proenzyme.
    Biochemistry. 1992 May 19;31(19):4535-40 PMID: 1581308
  12. Molecular structure of the DNA cross-link repair gene SNM1 (PSO2) of the yeast Saccharomyces cerevisiae.
    Mol Gen Genet. 1992 Jan;231(2):194-200 PMID: 1736091
  13. Local alignment statistics.
    Methods Enzymol. 1996;266:460-80 PMID: 8743700
  14. Improved tools for biological sequence comparison.
    Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8 PMID: 3162770
  15. Computer analysis of bacterial haloacid dehalogenases defines a large superfamily of hydrolases with diverse specificity. Application of an iterative approach to database search.
    J Mol Biol. 1994 Nov 18;244(1):125-32 PMID: 7966317
  16. Molecular evolutionary analysis of the YWVZ/7B globin gene cluster of the insect Chironomus thummi.
    J Mol Evol. 1995 Sep;41(3):313-28 PMID: 7563117
  17. Detection of conserved segments in proteins: iterative scanning of sequence databases with alignment blocks.
    Proc Natl Acad Sci U S A. 1994 Dec 6;91(25):12091-5 PMID: 7991589
  18. Fus2 localizes near the site of cell fusion and is required for both cell fusion and nuclear alignment during zygote formation.
    J Cell Biol. 1995 Sep;130(6):1283-96 PMID: 7559752
  19. Identification of sequence pattern with profile analysis.
    Methods Enzymol. 1996;266:198-212 PMID: 8743686
  20. Rho as a regulator of the cytoskeleton.
    Trends Biochem Sci. 1995 Jun;20(6):227-31 PMID: 7543224
  21. Three-dimensional structure of a lipoyl domain from the dihydrolipoyl acetyltransferase component of the pyruvate dehydrogenase multienzyme complex of Escherichia coli.
    J Mol Biol. 1995 Apr 28;248(2):328-43 PMID: 7739044
  22. Profile analysis: detection of distantly related proteins.
    Proc Natl Acad Sci U S A. 1987 Jul;84(13):4355-8 PMID: 3474607
  23. Three-dimensional structure of the lipoyl domain from Bacillus stearothermophilus pyruvate dehydrogenase multienzyme complex.
    J Mol Biol. 1993 Feb 20;229(4):1037-48 PMID: 8445635
  24. Origins of cell polarity.
    Cell. 1996 Feb 9;84(3):335-44 PMID: 8608587
  25. Plugging it in: signaling circuits and the yeast cell cycle.
    Curr Opin Cell Biol. 1996 Apr;8(2):223-30 PMID: 8791423
  26. Iterative template refinement: protein-fold prediction using iterative search and hybrid sequence/structure templates.
    Methods Enzymol. 1996;266:322-39 PMID: 8743692
  27. The PH domain: a common piece in the structural patchwork of signalling proteins.
    Trends Biochem Sci. 1993 Sep;18(9):343-8 PMID: 8236453
  28. The 3-D structure of a zinc metallo-beta-lactamase from Bacillus cereus reveals a new type of protein fold.
    EMBO J. 1995 Oct 16;14(20):4914-21 PMID: 7588620
  29. Distantly related sequences in the alpha- and beta-subunits of ATP synthase, myosin, kinases and other ATP-requiring enzymes and a common nucleotide binding fold.
    EMBO J. 1982;1(8):945-51 PMID: 6329717
  30. Improving the sensitivity of the sequence profile method.
    Protein Sci. 1994 Jan;3(1):139-46 PMID: 7511453
  31. Protein family classification based on searching a database of blocks.
    Genomics. 1994 Jan 1;19(1):97-107 PMID: 8188249
  32. Multidrug resistance pumps in bacteria: variations on a theme.
    Trends Biochem Sci. 1994 Mar;19(3):119-23 PMID: 8203018
  33. Gibbs motif sampling: detection of bacterial outer membrane protein repeats.
    Protein Sci. 1995 Aug;4(8):1618-32 PMID: 8520488
  34. Homology of the NifS family of proteins to a new class of pyridoxal phosphate-dependent enzymes.
    FEBS Lett. 1993 May 10;322(2):159-64 PMID: 8482384
  35. Role for the Rho-family GTPase Cdc42 in yeast mating-pheromone signal pathway.
    Nature. 1995 Aug 24;376(6542):702-5 PMID: 7651520
  36. PH domain: the first anniversary.
    Trends Biochem Sci. 1994 Sep;19(9):349-53 PMID: 7985225
  37. Cell polarity and the mechanism of asymmetric cell division.
    Bioessays. 1994 Dec;16(12):925-31 PMID: 7840773
  38. Distal residues in the oxygen binding site of haemoglobin studied by protein engineering.
    Nature. 1987 Oct 29-Nov 4;329(6142):858-60 PMID: 3313055
  39. The pleckstrin homology domain: an intriguing multifunctional protein module.
    Bioessays. 1996 Jan;18(1):35-46 PMID: 8593162
  40. Alignment of 700 globin sequences: extent of amino acid substitution and its correlation with variation in volume.
    Protein Sci. 1995 Oct;4(10):2179-90 PMID: 8535255
  41. Detecting subtle sequence signals: a Gibbs sampling strategy for multiple alignment.
    Science. 1993 Oct 8;262(5131):208-14 PMID: 8211139
  42. Effect of the distal residues on the vibrational modes of the Fe-CO bond in hemoglobin studied by protein engineering.
    Biochemistry. 1990 Jun 12;29(23):5562-6 PMID: 2201408
  43. Hidden Markov models of biological primary sequence information.
    Proc Natl Acad Sci U S A. 1994 Feb 1;91(3):1059-63 PMID: 8302831
  44. Evolutionary families of metallopeptidases.
    Methods Enzymol. 1995;248:183-228 PMID: 7674922
  45. Multicopy suppression of the cdc24 budding defect in yeast by CDC42 and three newly identified genes including the ras-related gene RSR1.
    Proc Natl Acad Sci U S A. 1989 Dec;86(24):9976-80 PMID: 2690082
  46. Cellular transformation and guanine nucleotide exchange activity are catalyzed by a common domain on the dbl oncogene product.
    J Biol Chem. 1994 Jan 7;269(1):62-5 PMID: 8276860
  47. Pheromone signalling in Saccharomyces cerevisiae requires the small GTP-binding protein Cdc42p and its activator CDC24.
    Mol Cell Biol. 1995 Oct;15(10):5246-57 PMID: 7565673
  48. Non-globular domains in protein sequences: automated segmentation using complexity measures.
    Comput Chem. 1994 Sep;18(3):269-85 PMID: 7952898
  49. Profile analysis.
    Methods Mol Biol. 1994;25:247-66 PMID: 8004170
  50. Structure and posttranslational modification of lipoyl domain of 2-oxo-acid dehydrogenase multienzyme complexes.
    Methods Enzymol. 1995;251:436-48 PMID: 7651225
  51. Prediction of the three-dimensional structures of the biotinylated domain from yeast pyruvate carboxylase and of the lipoylated H-protein from the pea leaf glycine cleavage system: a new automated method for the prediction of protein tertiary structure.
    Protein Sci. 1993 Apr;2(4):626-39 PMID: 8518734
  52. Molecular cloning, mapping, and regulation of Pho regulon genes for phosphonate breakdown by the phosphonatase pathway of Salmonella typhimurium LT2.
    J Bacteriol. 1995 Nov;177(22):6411-21 PMID: 7592415
  53. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  54. Position-based sequence weights.
    J Mol Biol. 1994 Nov 4;243(4):574-8 PMID: 7966282
  55. Proteins. One thousand families for the molecular biologist.
    Nature. 1992 Jun 18;357(6379):543-4 PMID: 1608464
  56. Yeast BUD5, encoding a putative GDP-GTP exchange factor, is necessary for bud site selection and interacts with bud formation gene BEM1.
    Cell. 1991 Jun 28;65(7):1213-24 PMID: 1905981
  57. Mutation of a mutL homolog in hereditary colon cancer.
    Science. 1994 Mar 18;263(5153):1625-9 PMID: 8128251
  58. Cell polarity in yeast.
    Trends Genet. 1994 Sep;10(9):328-33 PMID: 7974747
  59. The BUD4 protein of yeast, required for axial budding, is localized to the mother/BUD neck in a cell cycle-dependent manner.
    J Cell Biol. 1996 Jul;134(2):413-27 PMID: 8707826
  60. Two novel families of bacterial membrane proteins concerned with nodulation, cell division and transport.
    Mol Microbiol. 1994 Mar;11(5):841-7 PMID: 8022262
  61. Maximum discrimination hidden Markov models of sequence consensus.
    J Comput Biol. 1995 Spring;2(1):9-23 PMID: 7497123
  62. Proteins regulating Ras and its relatives.
    Nature. 1993 Dec 16;366(6456):643-54 PMID: 8259209
  63. Sequence and domain structure of yeast pyruvate carboxylase.
    J Biol Chem. 1988 Aug 15;263(23):11493-7 PMID: 3042770
  64. Yeast chromosome III: new gene functions.
    EMBO J. 1994 Feb 1;13(3):493-503 PMID: 8313894
  65. TEL1, a gene involved in controlling telomere length in S. cerevisiae, is homologous to the human ataxia telangiectasia gene.
    Cell. 1995 Sep 8;82(5):823-9 PMID: 7671310
  66. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  67. Applying motif and profile searches.
    Methods Enzymol. 1996;266:162-84 PMID: 8743684
  68. Protein sequence comparison at genome scale.
    Methods Enzymol. 1996;266:295-322 PMID: 8743691
  69. Enhanced Co2+ activation and inhibitor binding of carboxypeptidase M at low pH. Similarity to carboxypeptidase H (enkephalin convertase).
    Biochem J. 1989 Jul 1;261(1):289-91 PMID: 2775217
  70. A family of extracytoplasmic proteins that allow transport of large molecules across the outer membranes of gram-negative bacteria.
    J Bacteriol. 1994 Jul;176(13):3825-31 PMID: 8021163
  71. Identification of genes required for normal pheromone-induced cell polarization in Saccharomyces cerevisiae.
    Genetics. 1994 Apr;136(4):1287-96 PMID: 8013906
  72. Family of glycosyl transferases needed for the synthesis of succinoglycan by Rhizobium meliloti.
    J Bacteriol. 1993 Nov;175(21):7033-44 PMID: 8226645
  73. The P-loop--a common motif in ATP- and GTP-binding proteins.
    Trends Biochem Sci. 1990 Nov;15(11):430-4 PMID: 2126155
  74. Role of Bud3p in producing the axial budding pattern of yeast.
    J Cell Biol. 1995 May;129(3):767-78 PMID: 7730410
  75. Ancient conserved regions in new gene sequences and the protein databases.
    Science. 1993 Mar 19;259(5102):1711-6 PMID: 8456298
  76. Mutation in the DNA mismatch repair gene homologue hMLH1 is associated with hereditary non-polyposis colon cancer.
    Nature. 1994 Mar 17;368(6468):258-61 PMID: 8145827
  77. Quantification of tertiary structural conservation despite primary sequence drift in the globin fold.
    Protein Sci. 1994 Oct;3(10):1706-11 PMID: 7849587
  78. Nemoglobins: divergent nematode globins.
    Parasitol Today. 1993 Oct;9(10):353-60 PMID: 15463668
  79. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  80. Interaction of avidin with the lipoyl domains in the pyruvate dehydrogenase multienzyme complex: three-dimensional location and similarity to biotinyl domains in carboxylases.
    Proc Biol Sci. 1992 Jun 22;248(1323):247-53 PMID: 1354363
  81. Multiple alignment using hidden Markov models.
    Proc Int Conf Intell Syst Mol Biol. 1995;3:114-20 PMID: 7584426
  82. Translational initiation factors IF-1 and eIF-2 alpha share an RNA-binding motif with prokaryotic ribosomal protein S1 and polynucleotide phosphorylase.
    Gene. 1992 Sep 21;119(1):107-11 PMID: 1383091
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
0305-1048
Published
1997-05-01
Pages
1665-77
Language
English
Region
England
NLM ID
0411011
PMCID
PMC146639
Subset
IM
Grants
NHGRI NIH HHS · 1RO1-HG0125701 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com