Home LiteratureArticle Details
PMID: 14552658 Published · epublish English Comparative Study Evaluation Study Journal Article Research Support, Non-U.S. Gov't

OXBench: a benchmark for evaluation of protein multiple sequence alignment accuracy.

BMC bioinformatics ·Vol. 4 ·2003-10-10 ·Pages 47

Raghava GP, Searle SM, Audley PC, Barber JD, Barton GJ

Abstract

The alignment of two or more protein sequences provides a powerful guide in the prediction of the protein structure and in identifying key functional residues, however, the utility of any prediction is completely dependent on the accuracy of the alignment. In this paper we describe a suite of reference alignments derived from the comparison of protein three-dimensional structures together with evaluation measures and software that allow automatically generated alignments to be benchmarked. We test the OXBench benchmark suite on alignments generated by the AMPS multiple alignment method, then apply the suite to compare eight different multiple alignment algorithms. The benchmark shows the current state-of-the art for alignment accuracy and provides a baseline against which new alignment algorithms may be judged. The simple hierarchical multiple alignment algorithm, AMPS, performed as well as or better than more modern methods such as CLUSTALW once the PAM250 pair-score matrix was replaced by a BLOSUM series matrix. AMPS gave an accuracy in Structurally Conserved Regions (SCRs) of 89.9% over a set of 672 alignments. The T-COFFEE method on a data set of families with <8 sequences gave 91.4% accuracy, significantly better than CLUSTALW (88.9%) and all other methods considered here. The complete suite is available from http://www.compbio.dundee.ac.uk. The OXBench suite of reference alignments, evaluation software and results database provide a convenient method to assess progress in sequence alignment techniques. Evaluation measures that were dependent on comparison to a reference alignment were found to give good discrimination between methods. The STAMP Sc Score which is independent of a reference alignment also gave good discrimination. Application of OXBench in this paper shows that with the exception of T-COFFEE, the majority of the improvement in alignment accuracy seen since 1985 stems from improved pair-score matrices rather than algorithmic refinements. The maximum theoretical alignment accuracy obtained by pooling results over all methods was 94.5% with 52.5% accuracy for alignments in the 0-10 percentage identity range. This suggests that further improvements in accuracy will be possible in the future.

MeSH Terms
Amino Acid Sequence Benchmarking/methods,statistics & numerical data Cluster Analysis Computational Biology/methods,standards,statistics & numerical data Computer Graphics/standards,statistics & numerical data Conserved Sequence Databases, Protein Ferredoxins/chemistry Internet Molecular Sequence Data Proteins/chemistry Reproducibility of Results Sequence Alignment/methods,standards,statistics & numerical data Sequence Homology, Amino Acid Software Software Design Software Validation Statistical Distributions
Chemicals
Ferredoxins Proteins
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Raghava G P S
School of Life Sciences, University of Dundee, Dow St, Dundee, DD1 5EH, Scotland, UK. raghava@imtech.res.in
Searle Stephen M J
Audley Patrick C
Barber Jonathan D
Barton Geoffrey J
References (44)
44 references, click to expand
  1. T-Coffee: A novel method for fast and accurate multiple sequence alignment.
    J Mol Biol. 2000 Sep 8;302(1):205-17 PMID: 10964570
  2. Evaluation and improvement of multiple sequence methods for protein secondary structure prediction.
    Proteins. 1999 Mar 1;34(4):508-19 PMID: 10081963
  3. 3Dee: a database of protein structural domains.
    Bioinformatics. 2001 Feb;17(2):200-1 PMID: 11238081
  4. Evaluation of protein multiple alignments by SAM-T99 using the BAliBASE multiple alignment test set.
    Bioinformatics. 2001 Aug;17(8):713-20 PMID: 11524372
  5. Estimation of P-values for global alignments of protein sequences.
    Bioinformatics. 2001 Dec;17(12):1158-67 PMID: 11751224
  6. Predicting reliable regions in protein sequence alignments.
    Bioinformatics. 2002 Feb;18(2):306-14 PMID: 11847078
  7. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  8. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features.
    Biopolymers. 1983 Dec;22(12):2577-637 PMID: 6667333
  9. Simultaneous comparison of three protein sequences.
    Proc Natl Acad Sci U S A. 1985 May;82(10):3073-7 PMID: 3858804
  10. Identification of protein sequence homology by consensus template alignment.
    J Mol Biol. 1986 Mar 20;188(2):233-58 PMID: 3088284
  11. Progressive sequence alignment as a prerequisite to correct phylogenetic trees.
    J Mol Evol. 1987;25(4):351-60 PMID: 3118049
  12. A strategy for the rapid multiple alignment of protein sequences. Confidence levels from tertiary structure comparisons.
    J Mol Biol. 1987 Nov 20;198(2):327-37 PMID: 3430611
  13. Evaluation and improvements in the automatic alignment of protein sequences.
    Protein Eng. 1987 Feb-Mar;1(2):89-94 PMID: 3507699
  14. A tool for multiple sequence alignment.
    Proc Natl Acad Sci U S A. 1989 Jun;86(12):4412-5 PMID: 2734293
  15. Automatic generation of primary sequence patterns from sets of related protein sequences.
    Proc Natl Acad Sci U S A. 1990 Jan;87(1):118-22 PMID: 2296575
  16. Protein multiple sequence alignment and flexible pattern matching.
    Methods Enzymol. 1990;183:403-28 PMID: 2314284
  17. Determination of reliable regions in protein sequence alignments.
    Protein Eng. 1990 Jul;3(7):565-9 PMID: 2217130
  18. The SWISS-PROT protein sequence data bank.
    Nucleic Acids Res. 1991 Apr 25;19 Suppl:2247-9 PMID: 2041811
  19. Exhaustive matching of the entire protein sequence database.
    Science. 1992 Jun 5;256(5062):1443-5 PMID: 1604319
  20. Pattern-induced multi-sequence alignment (PIMA) algorithm employing secondary structure-dependent gap penalties for use in comparative protein modelling.
    Protein Eng. 1992 Jan;5(1):35-41 PMID: 1631044
  21. Multiple protein sequence alignment from tertiary structure comparison: assignment of global and residue confidence levels.
    Proteins. 1992 Oct;14(2):309-23 PMID: 1409577
  22. Amino acid substitution matrices from protein blocks.
    Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9 PMID: 1438297
  23. ALSCRIPT: a tool to format multiple sequence alignments.
    Protein Eng. 1993 Jan;6(1):37-40 PMID: 8433969
  24. Optimal alignment between groups of sequences and its application to multiple sequence alignment.
    Comput Appl Biosci. 1993 Jun;9(3):361-70 PMID: 8324637
  25. Application of multiple sequence alignment profiles to improve protein secondary structure prediction.
    Proteins. 2000 Aug 15;40(3):502-11 PMID: 10861942
  26. Comparative protein structure modeling of genes and genomes.
    Annu Rev Biophys Biomol Struct. 2000;29:291-325 PMID: 10940251
  27. Hidden Markov models in computational biology. Applications to protein modeling.
    J Mol Biol. 1994 Feb 4;235(5):1501-31 PMID: 8107089
  28. Comparative analysis of multiple protein-sequence alignment methods.
    Mol Biol Evol. 1994 Jul;11(4):571-92 PMID: 8078398
  29. Protein sequence alignments: a strategy for the hierarchical analysis of residue conservation.
    Comput Appl Biosci. 1993 Dec;9(6):745-56 PMID: 8143162
  30. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  31. Further improvement in methods of group-to-group sequence alignment with generalized profile operations.
    Comput Appl Biosci. 1994 Jul;10(4):379-87 PMID: 7804871
  32. Derivation of rules for comparative protein modeling from a database of protein structure alignments.
    Protein Sci. 1994 Sep;3(9):1582-96 PMID: 7833817
  33. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  34. An assessment of amino acid exchange matrices in aligning protein sequences: the twilight zone revisited.
    J Mol Biol. 1995 Jun 16;249(4):816-31 PMID: 7602593
  35. Improving the practical space and time efficiency of the shortest-paths approach to sum-of-pairs multiple sequence alignment.
    J Comput Biol. 1995 Fall;2(3):459-72 PMID: 8521275
  36. A weighting system and algorithm for aligning many phylogenetically related sequences.
    Comput Appl Biosci. 1995 Oct;11(5):543-51 PMID: 8590178
  37. The structural alignment between two proteins: is there a unique answer?
    Protein Sci. 1996 Jul;5(7):1325-38 PMID: 8819165
  38. Multiple DNA and protein sequence alignment based on segment-to-segment comparison.
    Proc Natl Acad Sci U S A. 1996 Oct 29;93(22):12098-103 PMID: 8901539
  39. Significant improvement in accuracy of multiple protein sequence alignments by iterative refinement as assessed by reference to structural alignments.
    J Mol Biol. 1996 Dec 13;264(4):823-38 PMID: 8980688
  40. Optimum superimposition of protein structures: ambiguities and implications.
    Fold Des. 1996;1(2):123-32 PMID: 9079372
  41. Critical assessment of methods of protein structure prediction (CASP): round II.
    Proteins. 1997;Suppl 1:2-6 PMID: 9485489
  42. Dynamic programming alignment accuracy.
    J Comput Biol. 1998 Fall;5(3):493-504 PMID: 9773345
  43. BAliBASE: a benchmark alignment database for the evaluation of multiple alignment programs.
    Bioinformatics. 1999 Jan;15(1):87-8 PMID: 10068696
  44. Protein structural domains: analysis of the 3Dee domains database.
    Proteins. 2001 Feb 15;42(3):332-44 PMID: 11151005
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2003-10-10
Epub
2003-00-10
Pages
47
Language
English
Region
England
NLM ID
100965194
PMCID
PMC280650
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com