Home LiteratureArticle Details
PMID: 17654362 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Improvement of phylogenies after removing divergent and ambiguously aligned blocks from protein sequence alignments.

Systematic biology ·Vol. 56 ·No. 4 ·2007-08-00 ·Pages 564-77

Talavera G, Castresana J

Abstract

Alignment quality may have as much impact on phylogenetic reconstruction as the phylogenetic methods used. Not only the alignment algorithm, but also the method used to deal with the most problematic alignment regions, may have a critical effect on the final tree. Although some authors remove such problematic regions, either manually or using automatic methods, in order to improve phylogenetic performance, others prefer to keep such regions to avoid losing any information. Our aim in the present work was to examine whether phylogenetic reconstruction improves after alignment cleaning or not. Using simulated protein alignments with gaps, we tested the relative performance in diverse phylogenetic analyses of the whole alignments versus the alignments with problematic regions removed with our previously developed Gblocks program. We also tested the performance of more or less stringent conditions in the selection of blocks. Alignments constructed with different alignment methods (ClustalW, Mafft, and Probcons) were used to estimate phylogenetic trees by maximum likelihood, neighbor joining, and parsimony. We show that, in most alignment conditions, and for alignments that are not too short, removal of blocks leads to better trees. That is, despite losing some information, there is an increase in the actual phylogenetic signal. Overall, the best trees are obtained by maximum-likelihood reconstruction of alignments cleaned by Gblocks. In general, a relaxed selection of blocks is better for short alignment, whereas a stringent selection is more adequate for longer ones. Finally, we show that cleaned alignments produce better topologies although, paradoxically, with lower bootstrap. This indicates that divergent and problematic alignment regions may lead, when present, to apparently better supported although, in fact, more biased topologies.

MeSH Terms
Amino Acid Sequence Cluster Analysis Computer Simulation Models, Genetic Molecular Sequence Data Phylogeny Proteins/chemistry Sequence Alignment/methods
Chemicals
Proteins
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Talavera Gerard
Department of Physiology, Institute of Molecular Biology of Barcelona, Barcelona, Spain.
Castresana Jose
Article Info
Journal
Systematic biology
Abbr.
Syst Biol
ISSN
1063-5157
Published
2007-08-00
Pages
564-77
Language
English
Region
England
NLM ID
9302532
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com