Home LiteratureArticle Details
PMID: 9705509 Published · ppublish English Comparative Study Journal Article

Protein sequence similarity searches using patterns as seeds.

Nucleic acids research ·Vol. 26 ·No. 17 ·1998-09-01 ·Pages 3986-90

Zhang Z, Schäffer AA, Miller W, Madden TL, Lipman DJ, Koonin EV, Altschul SF

Abstract

Protein families often are characterized by conserved sequence patterns or motifs. A researcher frequently wishes to evaluate the significance of a specific pattern within a protein, or to exploit knowledge of known motifs to aid the recognition of greatly diverged but homologous family members. To assist in these efforts, the pattern-hit initiated BLAST (PHI-BLAST) program described here takes as input both a protein sequence and a pattern of interest that it contains. PHI-BLAST searches a protein database for other instances of the input pattern, and uses those found as seeds for the construction of local alignments to the query sequence. The random distribution of PHI-BLAST alignment scores is studied analytically and empirically. In many instances, the program is able to detect statistically significant similarity between homologous proteins that are not recognizably related using traditional single-pass database search methods. PHI-BLAST is applied to the analysis of CED4-like cell death regulators, HS90-type ATPase domains, archaeal tRNA nucleotidyltransferases and archaeal homologs of DnaG-type DNA primases.

MeSH Terms
Adenosine Triphosphatases Algorithms Amino Acid Sequence Archaeal Proteins Caenorhabditis elegans Proteins Calcium-Binding Proteins DNA Primase Databases, Factual HSP90 Heat-Shock Proteins Helminth Proteins Pattern Recognition, Automated RNA Nucleotidyltransferases Software
Chemicals
Archaeal Proteins Caenorhabditis elegans Proteins Calcium-Binding Proteins Ced-4 protein, C elegans HSP90 Heat-Shock Proteins Helminth Proteins DNA Primase RNA Nucleotidyltransferases Adenosine Triphosphatases
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Zhang Z
Department of Computer Science and Engineering, Pennsylvania State University, University Park, PA 16802, USA.
Schäffer A A
Miller W
Madden T L
Lipman D J
Koonin E V
Altschul S F
References (44)
44 references, click to expand
  1. Approximate matching of regular expressions.
    Bull Math Biol. 1989;51(1):5-37 PMID: 2706401
  2. Methods for calculating the probabilities of finding patterns in sequences.
    Comput Appl Biosci. 1989 Apr;5(2):89-96 PMID: 2720468
  3. A system for pattern matching applications on biosequences.
    Comput Appl Biosci. 1993 Jun;9(3):299-314 PMID: 8324630
  4. Nonconserved segment of the MutL protein from Escherichia coli K-12 and Salmonella typhimurium.
    Nucleic Acids Res. 1992 May 11;20(9):2379 PMID: 1594459
  5. An improved algorithm for matching biological sequences.
    J Mol Biol. 1982 Dec 15;162(3):705-8 PMID: 7166760
  6. Local alignment statistics.
    Methods Enzymol. 1996;266:460-80 PMID: 8743700
  7. Improved tools for biological sequence comparison.
    Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8 PMID: 3162770
  8. Comparison of archaeal and bacterial genomes: computer analysis of protein sequences predicts novel functions and suggests a chimeric origin for the archaea.
    Mol Microbiol. 1997 Aug;25(4):619-37 PMID: 9379893
  9. Empirical statistical estimates for sequence similarity searches.
    J Mol Biol. 1998 Feb 13;276(1):71-84 PMID: 9514730
  10. The significance of protein sequence similarities.
    Comput Appl Biosci. 1988 Mar;4(1):67-71 PMID: 3383005
  11. Distribution of glutamine and asparagine residues and their near neighbors in peptides and proteins.
    Proc Natl Acad Sci U S A. 1991 Oct 15;88(20):8880-4 PMID: 1924347
  12. Alignments without low-scoring regions.
    J Comput Biol. 1998 Summer;5(2):197-210 PMID: 9672828
  13. Optimal alignments in linear space.
    Comput Appl Biosci. 1988 Mar;4(1):11-7 PMID: 3382986
  14. Prediction of the coding sequences of unidentified human genes. IV. The coding sequences of 40 new genes (KIAA0121-KIAA0160) deduced by analysis of cDNA clones from human cell line KG-1.
    DNA Res. 1995 Aug 31;2(4):167-74, 199-210 PMID: 8590280
  15. An atypical topoisomerase II from Archaea with implications for meiotic recombination.
    Nature. 1997 Mar 27;386(6623):414-7 PMID: 9121560
  16. Amino acid substitution matrices from protein blocks.
    Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9 PMID: 1438297
  17. Identification of the primase active site of the herpes simplex virus type 1 helicase-primase.
    J Biol Chem. 1995 Jun 9;270(23):14148-53 PMID: 7775476
  18. Matching sequences under deletion-insertion constraints.
    Proc Natl Acad Sci U S A. 1972 Jan;69(1):4-6 PMID: 4500555
  19. Automatic generation of primary sequence patterns from sets of related protein sequences.
    Proc Natl Acad Sci U S A. 1990 Jan;87(1):118-22 PMID: 2296575
  20. Cytochrome c and dATP-dependent formation of Apaf-1/caspase-9 complex initiates an apoptotic protease cascade.
    Cell. 1997 Nov 14;91(4):479-89 PMID: 9390557
  21. Searching for patterns in protein and nucleic acid sequences.
    Methods Enzymol. 1990;183:193-211 PMID: 1690333
  22. GenBank.
    Nucleic Acids Res. 1998 Jan 1;26(1):1-7 PMID: 9399790
  23. A promoter associated with the neisserial repeat can be used to transcribe the uvrB gene from Neisseria gonorrhoeae.
    J Bacteriol. 1995 Apr;177(8):1952-8 PMID: 7721686
  24. Optimal sequence alignments.
    Proc Natl Acad Sci U S A. 1983 Mar;80(5):1382-6 PMID: 16593289
  25. CCA-adding enzymes and poly(A) polymerases are all members of the same nucleotidyltransferase superfamily: characterization of the CCA-adding enzyme from the archaeal hyperthermophile Sulfolobus shibatae.
    RNA. 1996 Sep;2(9):895-908 PMID: 8809016
  26. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  27. Positionally cloned human disease genes: patterns of evolutionary conservation and functional motifs.
    Proc Natl Acad Sci U S A. 1997 May 27;94(11):5831-6 PMID: 9159160
  28. Apaf-1, a human protein homologous to C. elegans CED-4, participates in cytochrome c-dependent activation of caspase-3.
    Cell. 1997 Aug 8;90(3):405-13 PMID: 9267021
  29. Caenorhabditis elegans CED-4 stimulates CED-3 processing and CED-3-induced apoptosis.
    Curr Biol. 1997 Jul 1;7(7):455-60 PMID: 9210374
  30. The statistical distribution of nucleic acid similarities.
    Nucleic Acids Res. 1985 Jan 25;13(2):645-56 PMID: 3871073
  31. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  32. Generalized affine gap costs for protein sequence alignment.
    Proteins. 1998 Jul 1;32(1):88-96 PMID: 9672045
  33. A simple tool to search for sequence motifs that are conserved in BLAST outputs.
    Comput Appl Biosci. 1994 Jul;10(4):457-9 PMID: 7804881
  34. The complete genome sequence of the hyperthermophilic, sulphate-reducing archaeon Archaeoglobus fulgidus.
    Nature. 1997 Nov 27;390(6658):364-70 PMID: 9389475
  35. The PROSITE database, its status in 1997.
    Nucleic Acids Res. 1997 Jan 1;25(1):217-21 PMID: 9016539
  36. Optimal sequence alignment using affine gap costs.
    Bull Math Biol. 1986;48(5-6):603-16 PMID: 3580642
  37. Issues in searching molecular sequence databases.
    Nat Genet. 1994 Feb;6(2):119-29 PMID: 8162065
  38. Construction and analysis of a profile library characterizing groups of structurally known proteins.
    Protein Sci. 1996 Oct;5(10):1991-9 PMID: 8897599
  39. Complete genome sequence of the methanogenic archaeon, Methanococcus jannaschii.
    Science. 1996 Aug 23;273(5278):1058-73 PMID: 8688087
  40. Role of CED-4 in the activation of CED-3.
    Nature. 1997 Aug 21;388(6644):728-9 PMID: 9285582
  41. Cloning and nucleotide base sequence analysis of a spectinomycin adenyltransferase AAD(9) determinant from Enterococcus faecalis.
    Antimicrob Agents Chemother. 1991 Sep;35(9):1804-10 PMID: 1659306
  42. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  43. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  44. 2.2 Mb of contiguous nucleotide sequence from chromosome III of C. elegans.
    Nature. 1994 Mar 3;368(6466):32-8 PMID: 7906398
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
0305-1048
Published
1998-09-01
Pages
3986-90
Language
English
Region
England
NLM ID
0411011
PMCID
PMC147803
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com