Home LiteratureArticle Details
PMID: 1438297 Published · ppublish English Comparative Study Journal Article Research Support, U.S. Gov't, P.H.S.

Amino acid substitution matrices from protein blocks.

Henikoff S, Henikoff JG

Abstract

Methods for alignment of protein sequences typically measure similarity by using a substitution matrix with scores for all possible exchanges of one amino acid with another. The most widely used matrices are based on the Dayhoff model of evolutionary rates. Using a different approach, we have derived substitution matrices from about 2000 blocks of aligned sequence segments characterizing more than 500 groups of related proteins. This led to marked improvements in alignments and in searches using queries from each of the groups.

MeSH Terms
Algorithms Amino Acid Sequence Animals Caenorhabditis elegans/genetics Drosophila/genetics Lod Score Mathematics Molecular Sequence Data Probability Proteins/chemistry,genetics Sequence Homology, Amino Acid Software
Chemicals
Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Henikoff S
Howard Hughes Medical Institute, Fred Hutchinson Cancer Research Center, Seattle, WA 98104.
Henikoff J G
References (23)
23 references, click to expand
  1. Amino acid substitution matrices from an information theoretic perspective.
    J Mol Biol. 1991 Jun 5;219(3):555-65 PMID: 2051488
  2. The SWISS-PROT protein sequence data bank.
    Nucleic Acids Res. 1991 Apr 25;19 Suppl:2247-9 PMID: 2041811
  3. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  4. Automatic generation of primary sequence patterns from sets of related protein sequences.
    Proc Natl Acad Sci U S A. 1990 Jan;87(1):118-22 PMID: 2296575
  5. Finding protein similarities with nucleotide sequence databases.
    Methods Enzymol. 1990;183:111-32 PMID: 2314271
  6. Mutation data matrix and its uses.
    Methods Enzymol. 1990;183:333-51 PMID: 2314281
  7. Searching through sequence databases.
    Methods Enzymol. 1990;183:99-110 PMID: 2314299
  8. A tool for multiple sequence alignment.
    Proc Natl Acad Sci U S A. 1989 Jun;86(12):4412-5 PMID: 2734293
  9. Multiple sequence alignment with hierarchical clustering.
    Nucleic Acids Res. 1988 Nov 25;16(22):10881-90 PMID: 2849754
  10. Amino acid substitutions in structurally related proteins. A pattern recognition approach. Determination of a new and efficient scoring matrix.
    J Mol Biol. 1988 Dec 20;204(4):1019-29 PMID: 3221397
  11. New scoring matrix for amino acid residue exchanges based on residue characteristic physical parameters.
    Int J Pept Protein Res. 1987 Feb;29(2):276-81 PMID: 3570667
  12. Tests for comparing related amino-acid sequences. Cytochrome c and cytochrome c 551 .
    J Mol Biol. 1971 Oct 28;61(2):409-24 PMID: 5167087
  13. Aligning amino acid sequences: comparison of commonly used methods.
    J Mol Evol. 1984-1985;21(2):112-25 PMID: 6100188
  14. Comparative model-building of the mammalian serine proteases.
    J Mol Biol. 1981 Dec 25;153(4):1027-42 PMID: 7045378
  15. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  16. Exhaustive matching of the entire protein sequence database.
    Science. 1992 Jun 5;256(5062):1443-5 PMID: 1604319
  17. The rapid generation of mutation data matrices from protein sequences.
    Comput Appl Biosci. 1992 Jun;8(3):275-82 PMID: 1633570
  18. Finding sequence motifs in groups of functionally related proteins.
    Proc Natl Acad Sci U S A. 1990 Jan;87(2):826-30 PMID: 1689055
  19. Automated assembly of protein blocks for database searching.
    Nucleic Acids Res. 1991 Dec 11;19(23):6565-72 PMID: 1754394
  20. Searching protein sequence libraries: comparison of the sensitivity and selectivity of the Smith-Waterman and FASTA algorithms.
    Genomics. 1991 Nov;11(3):635-50 PMID: 1774068
  21. Multiple sequence alignment of protein families showing low sequence homology: a methodological approach using database pattern-matching discriminators for G-protein-linked receptors.
    Gene. 1991 Feb 15;98(2):153-9 PMID: 1849861
  22. PROSITE: a dictionary of sites and patterns in proteins.
    Nucleic Acids Res. 1991 Apr 25;19 Suppl:2241-5 PMID: 2041810
  23. Rapid and sensitive sequence comparison with FASTP and FASTA.
    Methods Enzymol. 1990;183:63-98 PMID: 2156132
Article Info
Journal
Proceedings of the National Academy of Sciences of the United States of America
Abbr.
Proc Natl Acad Sci U S A
ISSN
0027-8424
Published
1992-11-15
Pages
10915-9
Language
English
Region
United States
NLM ID
7505876
PMCID
PMC50453
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com