Abstract
An unusual pattern in a nucleic acid or protein sequence or a region of strong similarity shared by two or more sequences may have biological significance. It is therefore desirable to know whether such a pattern can have arisen simply by chance. To identify interesting sequence patterns, appropriate scoring values can be assigned to the individual residues of a single sequence or to sets of residues when several sequences are compared. For single sequences, such scores can reflect biophysical properties such as charge, volume, hydrophobicity, or secondary structure potential; for multiple sequences, they can reflect nucleotide or amino acid similarity measured in a wide variety of ways. Using an appropriate random model, we present a theory that provides precise numerical formulas for assessing the statistical significance of any region with high aggregate score. A second class of results describes the composition of high-scoring segments. In certain contexts, these permit the choice of scoring systems which are "optimal" for distinguishing biologically relevant patterns. Examples are given of applications of the theory to a variety of protein sequences, highlighting segments with unusual biological features. These include distinctive charge regions in transcription factors and protooncogene products, pronounced hydrophobic segments in various receptor and transport proteins, and statistically significant subalignments involving the recently characterized cystic fibrosis gene.
MeSH Terms
Amino Acid Sequence
Analysis of Variance
Base Sequence
Biological Evolution
Models, Genetic
Models, Statistical
Nucleic Acids/genetics
Probability
Proteins/genetics
Chemicals
Nucleic Acids
Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Karlin S
Department of Mathematics, Stanford University, CA 94305.
Altschul S F
References (24)
24 references, click to expand
-
Cloning of the human cDNA for the U1 RNA-associated 70K protein.
EMBO J. 1986 Dec 1;5(12):3209-17
PMID: 3028775
-
Aligning amino acid sequences: comparison of commonly used methods.
J Mol Evol. 1984-1985;21(2):112-25
PMID: 6100188
-
Structure and sequence of the Drosophila zeste gene.
EMBO J. 1987 Mar;6(3):791-9
PMID: 3582372
-
Efficient algorithms for molecular sequence analysis.
Proc Natl Acad Sci U S A. 1988 Feb;85(3):841-5
PMID: 3124111
-
A gene activated by growth factors is related to the oncogene v-jun.
Proc Natl Acad Sci U S A. 1988 Mar;85(5):1487-91
PMID: 3422745
-
On the PAM matrix model of protein evolution.
Mol Biol Evol. 1985 Sep;2(5):434-47
PMID: 3870870
-
Significance of nucleotide sequence alignments: a method for random sequence permutation that preserves dinucleotide and codon usage.
Mol Biol Evol. 1985 Nov;2(6):526-38
PMID: 3870875
-
Fos-associated protein p39 is the product of the jun proto-oncogene.
Science. 1988 May 20;240(4855):1010-6
PMID: 3130660
-
The mas oncogene encodes an angiotensin receptor.
Nature. 1988 Sep 29;335(6189):437-40
PMID: 3419518
-
Amino acid substitutions in structurally related proteins. A pattern recognition approach. Determination of a new and efficient scoring matrix.
J Mol Biol. 1988 Dec 20;204(4):1019-29
PMID: 3221397
-
A method to identify distinctive charge configurations in protein sequences, with application to human herpesvirus polypeptides.
J Mol Biol. 1989 Jan 5;205(1):165-77
PMID: 2538622
-
The molecular biology of cytochrome P450s.
Pharmacol Rev. 1988 Dec;40(4):243-88
PMID: 3072575
-
Transcriptional regulation in mammalian cells by sequence-specific DNA binding proteins.
Science. 1989 Jul 28;245(4916):371-8
PMID: 2667136
-
Association of charge clusters with functional domains of cellular transcription factors.
Proc Natl Acad Sci U S A. 1989 Aug;86(15):5698-702
PMID: 2569737
-
Identification of the cystic fibrosis gene: cloning and characterization of complementary DNA.
Science. 1989 Sep 8;245(4922):1066-73
PMID: 2475911
-
Identification of significant sequence patterns in proteins.
Methods Enzymol. 1990;183:388-402
PMID: 2179677
-
Charge configurations in oncogene products and transforming proteins.
Oncogene. 1990 Jan;5(1):85-95
PMID: 2181379
-
A general method applicable to the search for similarities in the amino acid sequence of two proteins.
J Mol Biol. 1970 Mar;48(3):443-53
PMID: 5420325
-
Tests for comparing related amino-acid sequences. Cytochrome c and cytochrome c 551 .
J Mol Biol. 1971 Oct 28;61(2):409-24
PMID: 5167087
-
Similar amino acid sequences: chance or common ancestry?
Science. 1981 Oct 9;214(4517):149-59
PMID: 7280687
-
Rapid similarity searches of nucleic acid and protein data banks.
Proc Natl Acad Sci U S A. 1983 Feb;80(3):726-30
PMID: 6572363
-
Random sequences.
J Mol Biol. 1983 Jan 15;163(2):171-6
PMID: 6842586
-
New approaches for computer analysis of nucleic acid sequences.
Proc Natl Acad Sci U S A. 1983 Sep;80(18):5660-4
PMID: 6577449
-
Homology between the DNA-binding domain of the GCN4 regulatory protein of yeast and the carboxyl-terminal region of a protein coded for by the oncogene jun.
Proc Natl Acad Sci U S A. 1987 May;84(10):3316-9
PMID: 3554236