Home LiteratureArticle Details
PMID: 9254694 Published · ppublish English Journal Article Research Support, U.S. Gov't, P.H.S. Review

Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.

Nucleic acids research ·Vol. 25 ·No. 17 ·1997-09-01 ·Pages 3389-402

Altschul SF, Madden TL, Schäffer AA, Zhang J, Zhang Z, Miller W, Lipman DJ

Abstract

The BLAST programs are widely used tools for searching protein and DNA databases for sequence similarities. For protein comparisons, a variety of definitional, algorithmic and statistical refinements described here permits the execution time of the BLAST programs to be decreased substantially while enhancing their sensitivity to weak similarities. A new criterion for triggering the extension of word hits, combined with a new heuristic for generating gapped alignments, yields a gapped BLAST program that runs at approximately three times the speed of the original. In addition, a method is introduced for automatically combining statistically significant alignments produced by BLAST into a position-specific score matrix, and searching the database using this matrix. The resulting Position-Specific Iterated BLAST (PSI-BLAST) program runs at approximately the same speed per iteration as gapped BLAST, but in many cases is much more sensitive to weak but biologically relevant sequence similarities. PSI-BLAST is used to uncover several new and interesting members of the BRCT superfamily.

MeSH Terms
Algorithms Amino Acid Sequence Animals DNA/chemistry Databases, Factual Humans Molecular Sequence Data Proteins/chemistry Sequence Alignment Software
Chemicals
Proteins DNA
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Altschul S F
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA. altschul@ncbi.nlm.nih.gov
Madden T L
Schäffer A A
Zhang J
Zhang Z
Miller W
Lipman D J
References (69)
69 references, click to expand
  1. Weights for data related by a tree.
    J Mol Biol. 1989 Jun 20;207(4):647-53 PMID: 2760928
  2. Volume changes in protein evolution.
    J Mol Biol. 1994 Mar 4;236(4):1067-78 PMID: 8120887
  3. Identification of a RING protein that can interact in vivo with the BRCA1 gene product.
    Nat Genet. 1996 Dec;14(4):430-40 PMID: 8944023
  4. The SWISS-PROT protein sequence data bank and its supplement TrEMBL.
    Nucleic Acids Res. 1997 Jan 1;25(1):31-6 PMID: 9016499
  5. Sequence analysis in the E1 region of adenovirus type 4 DNA.
    Virology. 1986 Dec;155(2):418-33 PMID: 2947381
  6. Using Dirichlet mixture priors to derive hidden Markov models for protein families.
    Proc Int Conf Intell Syst Mol Biol. 1993;1:47-55 PMID: 7584370
  7. Analysis of gene duplication repeats in the myosin rod.
    J Mol Biol. 1983 Sep 5;169(1):15-30 PMID: 6620380
  8. An improved algorithm for matching biological sequences.
    J Mol Biol. 1982 Dec 15;162(3):705-8 PMID: 7166760
  9. Local alignment statistics.
    Methods Enzymol. 1996;266:460-80 PMID: 8743700
  10. Improved tools for biological sequence comparison.
    Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8 PMID: 3162770
  11. Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes.
    Proc Natl Acad Sci U S A. 1990 Mar;87(6):2264-8 PMID: 2315319
  12. Detection of conserved segments in proteins: iterative scanning of sequence databases with alignment blocks.
    Proc Natl Acad Sci U S A. 1994 Dec 6;91(25):12091-5 PMID: 7991589
  13. Dirichlet mixtures: a method for improved detection of weak but significant protein sequence homology.
    Comput Appl Biosci. 1996 Aug;12(4):327-45 PMID: 8902360
  14. Profile analysis: detection of distantly related proteins.
    Proc Natl Acad Sci U S A. 1987 Jul;84(13):4355-8 PMID: 3474607
  15. Prediction of the coding sequences of unidentified human genes. VI. The coding sequences of 80 new genes (KIAA0201-KIAA0280) deduced by analysis of cDNA clones from cell line KG-1 and brain.
    DNA Res. 1996 Oct 31;3(5):321-9, 341-54 PMID: 9039502
  16. A weighting system and algorithm for aligning many phylogenetically related sequences.
    Comput Appl Biosci. 1995 Oct;11(5):543-51 PMID: 8590178
  17. Using substitution probabilities to improve position-specific scoring matrices.
    Comput Appl Biosci. 1996 Apr;12(2):135-43 PMID: 8744776
  18. Information content of binding sites on nucleotide sequences.
    J Mol Biol. 1986 Apr 5;188(3):415-31 PMID: 3525846
  19. Systematic method for the detection of potential lambda Cro-like DNA-binding regions in proteins.
    J Mol Biol. 1987 Apr 5;194(3):557-64 PMID: 3625774
  20. The significance of protein sequence similarities.
    Comput Appl Biosci. 1988 Mar;4(1):67-71 PMID: 3383005
  21. Distribution of glutamine and asparagine residues and their near neighbors in peptides and proteins.
    Proc Natl Acad Sci U S A. 1991 Oct 15;88(20):8880-4 PMID: 1924347
  22. Optimal alignments in linear space.
    Comput Appl Biosci. 1988 Mar;4(1):11-7 PMID: 3382986
  23. Computer methods to locate signals in nucleic acid sequences.
    Nucleic Acids Res. 1984 Jan 11;12(1 Pt 2):505-19 PMID: 6364039
  24. Isolation, characterization, and inactivation of the APA1 gene encoding yeast diadenosine 5',5'''-P1,P4-tetraphosphate phosphorylase.
    J Bacteriol. 1989 Dec;171(12):6437-45 PMID: 2556364
  25. A protein alignment scoring system sensitive at all evolutionary distances.
    J Mol Evol. 1993 Mar;36(3):290-300 PMID: 8483166
  26. Amino acid substitution matrices from protein blocks.
    Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9 PMID: 1438297
  27. Rapid similarity searches of nucleic acid and protein data banks.
    Proc Natl Acad Sci U S A. 1983 Feb;80(3):726-30 PMID: 6572363
  28. Embedding strategies for effective use of information from multiple sequence alignments.
    Protein Sci. 1997 Mar;6(3):698-705 PMID: 9070452
  29. Detecting subtle sequence signals: a Gibbs sampling strategy for multiple alignment.
    Science. 1993 Oct 8;262(5131):208-14 PMID: 8211139
  30. Aligning two sequences within a specified diagonal band.
    Comput Appl Biosci. 1992 Oct;8(5):481-7 PMID: 1422882
  31. GenBank.
    Nucleic Acids Res. 1997 Jan 1;25(1):1-6 PMID: 9016491
  32. The gal locus from Haemophilus influenzae: cloning, sequencing and the use of gal mutants to study lipopolysaccharide.
    Mol Microbiol. 1992 Oct;6(20):3051-63 PMID: 1282642
  33. A strong candidate for the breast and ovarian cancer susceptibility gene BRCA1.
    Science. 1994 Oct 7;266(5182):66-71 PMID: 7545954
  34. Rat galactose-1-phosphate uridyltransferase coding sequence, transcription start site and genomic organization.
    DNA Seq. 1993;3(5):311-8 PMID: 8400361
  35. Detecting homology of distantly related proteins with consensus sequences.
    J Mol Biol. 1987 Dec 20;198(4):567-77 PMID: 3430622
  36. Identifying protein-binding sites from unaligned DNA fragments.
    Proc Natl Acad Sci U S A. 1989 Feb;86(4):1183-7 PMID: 2919167
  37. [Hemoglobins, XXXIII. Note on the Sequence of the hemoglobins of the horse (author's transl)].
    Hoppe Seylers Z Physiol Chem. 1980 Jul;361(7):1107-16 PMID: 7409745
  38. Database of homology-derived protein structures and the structural meaning of sequence alignment.
    Proteins. 1991;9(1):56-68 PMID: 2017436
  39. Optimal sequence alignments.
    Proc Natl Acad Sci U S A. 1983 Mar;80(5):1382-6 PMID: 16593289
  40. A new algorithm for best subsequence alignments with application to tRNA-rRNA comparisons.
    J Mol Biol. 1987 Oct 20;197(4):723-8 PMID: 2448477
  41. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  42. Position-based sequence weights.
    J Mol Biol. 1994 Nov 4;243(4):574-8 PMID: 7966282
  43. Weighting aligned protein or nucleic acid sequences to correct for unequal representation.
    J Mol Biol. 1990 Dec 20;216(4):813-8 PMID: 2176240
  44. Applications and statistics for multiple high-scoring segments in molecular sequences.
    Proc Natl Acad Sci U S A. 1993 Jun 15;90(12):5873-7 PMID: 8390686
  45. Locally optimal subalignments using nonlinear similarity functions.
    Bull Math Biol. 1986;48(5-6):633-60 PMID: 3580643
  46. From BRCA1 to RAP1: a widespread BRCT module closely associated with DNA repair.
    FEBS Lett. 1997 Jan 2;400(1):25-30 PMID: 9000507
  47. Amino acid substitution matrices from an information theoretic perspective.
    J Mol Biol. 1991 Jun 5;219(3):555-65 PMID: 2051488
  48. Selection of DNA binding sites by regulatory proteins. Statistical-mechanical theory and application to operators and promoters.
    J Mol Biol. 1987 Feb 20;193(4):723-50 PMID: 3612791
  49. Maximum discrimination hidden Markov models of sequence consensus.
    J Comput Biol. 1995 Spring;2(1):9-23 PMID: 7497123
  50. A flexible motif search technique based on generalized profiles.
    Comput Chem. 1996 Mar;20(1):3-23 PMID: 8867839
  51. The amino acid sequence of leghaemoglobin I from root nodules of broad bean (Vicia faba L.).
    FEBS Lett. 1975 Mar 1;51(1):33-7 PMID: 1123063
  52. A superfamily of conserved domains in DNA damage-responsive cell cycle checkpoint proteins.
    FASEB J. 1997 Jan;11(1):68-76 PMID: 9034168
  53. Complete structure of the hemagglutinin gene from the human influenza A/Victoria/3/75 (H3N2) strain as determined from cloned DNA.
    Cell. 1980 Mar;19(3):683-96 PMID: 6153930
  54. Identification of protein sequence homology by consensus template alignment.
    J Mol Biol. 1986 Mar 20;188(2):233-58 PMID: 3088284
  55. The statistical distribution of nucleic acid similarities.
    Nucleic Acids Res. 1985 Jan 25;13(2):645-56 PMID: 3871073
  56. Recognition of related proteins by iterative template refinement (ITR).
    Protein Sci. 1994 Aug;3(8):1315-28 PMID: 7987226
  57. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  58. A workbench for large-scale sequence homology analysis.
    Comput Appl Biosci. 1994 Jun;10(3):301-7 PMID: 7922687
  59. Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. II. Sequence determination of the entire genome and assignment of potential protein-coding regions.
    DNA Res. 1996 Jun 30;3(3):109-36 PMID: 8905231
  60. BRCA1 protein products ... Functional motifs...
    Nat Genet. 1996 Jul;13(3):266-8 PMID: 8673121
  61. The FHIT gene, spanning the chromosome 3p14.2 fragile site and renal carcinoma-associated t(3;8) breakpoint, is abnormal in digestive tract cancers.
    Cell. 1996 Feb 23;84(4):587-97 PMID: 8598045
  62. Optimal sequence alignment using affine gap costs.
    Bull Math Biol. 1986;48(5-6):603-16 PMID: 3580642
  63. Issues in searching molecular sequence databases.
    Nat Genet. 1994 Feb;6(2):119-29 PMID: 8162065
  64. Insertional mutagenesis in zebrafish identifies two novel genes, pescadillo and dead eye, essential for embryonic development.
    Genes Dev. 1996 Dec 15;10(24):3141-55 PMID: 8985183
  65. Complete genome sequence of the methanogenic archaeon, Methanococcus jannaschii.
    Science. 1996 Aug 23;273(5278):1058-73 PMID: 8688087
  66. New structure--novel fold?
    Structure. 1997 Feb 15;5(2):165-71 PMID: 9032077
  67. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  68. 2.2 Mb of contiguous nucleotide sequence from chromosome III of C. elegans.
    Nature. 1994 Mar 3;368(6466):32-8 PMID: 7906398
  69. Improved sensitivity of profile searches through the use of sequence weights and gap excision.
    Comput Appl Biosci. 1994 Feb;10(1):19-29 PMID: 8193951
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
0305-1048
Published
1997-09-01
Pages
3389-402
Language
English
Region
England
NLM ID
0411011
PMCID
PMC146917
Subset
IM
Grants
NLM NIH HHS · LM05110 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com