Home LiteratureArticle Details
PMID: 9600919 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Assessing sequence comparison methods with reliable structurally identified distant evolutionary relationships.

Brenner SE, Chothia C, Hubbard TJ

Abstract

Pairwise sequence comparison methods have been assessed using proteins whose relationships are known reliably from their structures and functions, as described in the SCOP database [Murzin, A. G., Brenner, S. E., Hubbard, T. & Chothia C. (1995) J. Mol. Biol. 247, 536-540]. The evaluation tested the programs BLAST [Altschul, S. F., Gish, W., Miller, W., Myers, E. W. & Lipman, D. J. (1990). J. Mol. Biol. 215, 403-410], WU-BLAST2 [Altschul, S. F. & Gish, W. (1996) Methods Enzymol. 266, 460-480], FASTA [Pearson, W. R. & Lipman, D. J. (1988) Proc. Natl. Acad. Sci. USA 85, 2444-2448], and SSEARCH [Smith, T. F. & Waterman, M. S. (1981) J. Mol. Biol. 147, 195-197] and their scoring schemes. The error rate of all algorithms is greatly reduced by using statistical scores to evaluate matches rather than percentage identity or raw scores. The E-value statistical scores of SSEARCH and FASTA are reliable: the number of false positives found in our tests agrees well with the scores reported. However, the P-values reported by BLAST and WU-BLAST2 exaggerate significance by orders of magnitude. SSEARCH, FASTA ktup = 1, and WU-BLAST2 perform best, and they are capable of detecting almost all relationships between proteins whose sequence identities are >30%. For more distantly related proteins, they do much less well; only one-half of the relationships between proteins with 20-30% identity are found. Because many homologs have low sequence similarity, most distant relationships cannot be detected by any pairwise comparison method; however, those which are identified may be used with confidence.

MeSH Terms
Algorithms Animals Databases, Factual Evolution, Molecular Humans Proteins/chemistry,genetics Sequence Alignment/methods
Chemicals
Proteins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Brenner S E
MRC Laboratory of Molecular Biology, Hills Road, Cambridge CB2 2QH, United Kingdom. brenner@hyper.stanford.edu
Chothia C
Hubbard T J
References (34)
34 references, click to expand
  1. Alignment of the amino acid sequences of distantly related proteins using variable gap penalties.
    Protein Eng. 1986 Oct-Nov;1(1):77-8 PMID: 3507691
  2. Improved tools for biological sequence comparison.
    Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8 PMID: 3162770
  3. Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes.
    Proc Natl Acad Sci U S A. 1990 Mar;87(6):2264-8 PMID: 2315319
  4. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  5. Database of homology-derived protein structures and the structural meaning of sequence alignment.
    Proteins. 1991;9(1):56-68 PMID: 2017436
  6. Searching protein sequence libraries: comparison of the sensitivity and selectivity of the Smith-Waterman and FASTA algorithms.
    Genomics. 1991 Nov;11(3):635-50 PMID: 1774068
  7. Amino acid substitution matrices from protein blocks.
    Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9 PMID: 1438297
  8. Receiver-operating characteristic (ROC) plots: a fundamental evaluation tool in clinical medicine.
    Clin Chem. 1993 Apr;39(4):561-77 PMID: 8472349
  9. Applications and statistics for multiple high-scoring segments in molecular sequences.
    Proc Natl Acad Sci U S A. 1993 Jun 15;90(12):5873-7 PMID: 8390686
  10. Crystal structure of the catalytic domain of a thermophilic endocellulase.
    Biochemistry. 1993 Sep 28;32(38):9906-16 PMID: 8399160
  11. A structural basis for sequence comparisons. An evaluation of scoring methodologies.
    J Mol Biol. 1993 Oct 20;233(4):716-38 PMID: 8411177
  12. Performance evaluation of amino acid substitution matrices.
    Proteins. 1993 Sep;17(1):49-61 PMID: 8234244
  13. Issues in searching molecular sequence databases.
    Nat Genet. 1994 Feb;6(2):119-29 PMID: 8162065
  14. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  15. An assessment of amino acid exchange matrices in aligning protein sequences: the twilight zone revisited.
    J Mol Biol. 1995 Jun 16;249(4):816-31 PMID: 7602593
  16. RASMOL: biomolecular graphics for all.
    Trends Biochem Sci. 1995 Sep;20(9):374 PMID: 7482707
  17. Comparison of methods for searching protein sequence databases.
    Protein Sci. 1995 Jun;4(6):1145-60 PMID: 7549879
  18. The SWISS-PROT protein sequence data bank and its new supplement TREMBL.
    Nucleic Acids Res. 1996 Jan 1;24(1):21-5 PMID: 8594581
  19. The PROSITE database, its status in 1995.
    Nucleic Acids Res. 1996 Jan 1;24(1):189-96 PMID: 8594577
  20. PIR-International Protein Sequence Database.
    Methods Enzymol. 1996;266:41-59 PMID: 8743676
  21. Effective protein sequence comparison.
    Methods Enzymol. 1996;266:227-58 PMID: 8743688
  22. Local alignment statistics.
    Methods Enzymol. 1996;266:460-80 PMID: 8743700
  23. Analysis of compositionally biased regions in sequence databases.
    Methods Enzymol. 1996;266:554-71 PMID: 8743706
  24. Understanding protein structure: using scop for fold interpretation.
    Methods Enzymol. 1996;266:635-43 PMID: 8743710
  25. A structural explanation for the twilight zone of protein sequence homology.
    Structure. 1996 Oct 15;4(10):1123-7 PMID: 8939745
  26. Population statistics of protein structures: lessons from structural classifications.
    Curr Opin Struct Biol. 1997 Jun;7(3):369-76 PMID: 9204279
  27. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  28. CATH--a hierarchic classification of protein domain structures.
    Structure. 1997 Aug 15;5(8):1093-108 PMID: 9309224
  29. Use of receiver operating characteristic (ROC) analysis to evaluate sequence matching.
    Comput Chem. 1996 Mar;20(1):25-33 PMID: 16718863
  30. An improved method of testing for evolutionary homology.
    J Mol Biol. 1966 Mar;16(1):9-16 PMID: 5917736
  31. Molecular packing and intermolecular contacts of sickling deer type III hemoglobin.
    J Mol Biol. 1979 Jul 5;131(3):417-33 PMID: 513126
  32. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  33. On the statistical significance of nucleic acid similarities.
    Nucleic Acids Res. 1984 Jan 11;12(1 Pt 1):215-26 PMID: 6694902
  34. Evaluation and improvements in the automatic alignment of protein sequences.
    Protein Eng. 1987 Feb-Mar;1(2):89-94 PMID: 3507699
Article Info
Journal
Proceedings of the National Academy of Sciences of the United States of America
Abbr.
Proc Natl Acad Sci U S A
ISSN
0027-8424
Published
1998-05-26
Pages
6073-8
Language
English
Region
United States
NLM ID
7505876
PMCID
PMC27587
Subset
IM
Grants
Wellcome Trust · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com