Abstract
Multiple sequence alignments are fundamental to many sequence analysis methods. Most alignments are computed using the progressive alignment heuristic. These methods are starting to become a bottleneck in some analysis pipelines when faced with data sets of the size of many thousands of sequences. Some methods allow computation of larger data sets while sacrificing quality, and others produce high-quality alignments, but scale badly with the number of sequences. In this paper, we describe a new program called Clustal Omega, which can align virtually any number of protein sequences quickly and that delivers accurate alignments. The accuracy of the package on smaller test cases is similar to that of the high-quality aligners. On larger data sets, Clustal Omega outperforms other packages in terms of execution time and quality. Clustal Omega also has powerful features for adding sequences to and exploiting information in existing alignments, making use of the vast amount of precomputed information in public databases like Pfam.
MeSH Terms
Algorithms
Amino Acid Sequence
Base Sequence
Data Mining/methods
Databases, Factual
Molecular Sequence Data
Proteins/analysis,chemistry
Sequence Alignment/methods
Sequence Analysis, Protein/methods
Software
Systems Biology/instrumentation,methods
Authors & Affiliations
12 authors, click to expand affiliations / ORCID
Sievers Fabian
School of Medicine and Medical Science, UCD Conway Institute of Biomolecular and Biomedical Research, University College Dublin, Dublin, Ireland.
Wilm Andreas
Dineen David
Gibson Toby J
Karplus Kevin
Li Weizhong
Lopez Rodrigo
McWilliam Hamish
Remmert Michael
Söding Johannes
Thompson Julie D
Higgins Desmond G
References (24)
24 references, click to expand
-
Clustal W and Clustal X version 2.0.
Bioinformatics. 2007 Nov 1;23(21):2947-8
PMID: 17846036
-
The Pfam protein families database.
Nucleic Acids Res. 2010 Jan;38(Database issue):D211-22
PMID: 19920124
-
Fast statistical alignment.
PLoS Comput Biol. 2009 May;5(5):e1000392
PMID: 19478997
-
The alignment of sets of sequences and the construction of phyletic trees: an integrated method.
J Mol Evol. 1984;20(2):175-86
PMID: 6433036
-
HOMSTRAD: a database of protein structure alignments for homologous families.
Protein Sci. 1998 Nov;7(11):2469-71
PMID: 9828015
-
R-Coffee: a method for multiple alignment of non-coding RNA.
Nucleic Acids Res. 2008 May;36(9):e52
PMID: 18420654
-
Profile hidden Markov models.
Bioinformatics. 1998;14(9):755-63
PMID: 9918945
-
MUSCLE: multiple sequence alignment with high accuracy and high throughput.
Nucleic Acids Res. 2004 Mar 19;32(5):1792-7
PMID: 15034147
-
SeaView version 4: A multiplatform graphical user interface for sequence alignment and phylogenetic tree building.
Mol Biol Evol. 2010 Feb;27(2):221-4
PMID: 19854763
-
Phylogeny-aware gap placement prevents errors in sequence alignment and evolutionary analysis.
Science. 2008 Jun 20;320(5883):1632-5
PMID: 18566285
-
MSAProbs: multiple sequence alignment based on pair hidden Markov models and partition function posterior probabilities.
Bioinformatics. 2010 Aug 15;26(16):1958-64
PMID: 20576627
-
Protein homology detection by HMM-HMM comparison.
Bioinformatics. 2005 Apr 1;21(7):951-60
PMID: 15531603
-
Issues in bioinformatics benchmarking: the case study of multiple sequence alignment.
Nucleic Acids Res. 2010 Nov;38(21):7353-63
PMID: 20639539
-
Quality measures for protein alignment benchmarks.
Nucleic Acids Res. 2010 Apr;38(7):2145-53
PMID: 20047958
-
MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform.
Nucleic Acids Res. 2002 Jul 15;30(14):3059-66
PMID: 12136088
-
BAliBASE 3.0: latest developments of the multiple sequence alignment benchmark.
Proteins. 2005 Oct 1;61(1):127-36
PMID: 16044462
-
Kalign--an accurate and fast multiple sequence alignment algorithm.
BMC Bioinformatics. 2005 Dec 12;6:298
PMID: 16343337
-
T-Coffee: A novel method for fast and accurate multiple sequence alignment.
J Mol Biol. 2000 Sep 8;302(1):205-17
PMID: 10964570
-
Sequence embedding for fast construction of guide trees for multiple sequence alignment.
Algorithms Mol Biol. 2010 May 14;5:21
PMID: 20470396
-
PRALINETM: a strategy for improved multiple alignment of transmembrane proteins.
Bioinformatics. 2008 Feb 15;24(4):492-7
PMID: 18174178
-
ProbCons: Probabilistic consistency-based multiple sequence alignment.
Genome Res. 2005 Feb;15(2):330-40
PMID: 15687296
-
The Jalview Java alignment editor.
Bioinformatics. 2004 Feb 12;20(3):426-7
PMID: 14960472
-
DIALIGN: finding local similarities by multiple sequence alignment.
Bioinformatics. 1998;14(3):290-4
PMID: 9614273
-
PartTree: an algorithm to build an approximate tree from a large number of unaligned sequences.
Bioinformatics. 2007 Feb 1;23(3):372-4
PMID: 17118958