Abstract
Gene families are growing rapidly, but standard methods for inferring phylogenies do not scale to alignments with over 10,000 sequences. We present FastTree, a method for constructing large phylogenies and for estimating their reliability. Instead of storing a distance matrix, FastTree stores sequence profiles of internal nodes in the tree. FastTree uses these profiles to implement Neighbor-Joining and uses heuristics to quickly identify candidate joins. FastTree then uses nearest neighbor interchanges to reduce the length of the tree. For an alignment with N sequences, L sites, and a different characters, a distance matrix requires O(N(2)) space and O(N(2)L) time, but FastTree requires just O(NLa + N ) memory and O(N log (N)La) time. To estimate the tree's reliability, FastTree uses local bootstrapping, which gives another 100-fold speedup over a distance matrix. For example, FastTree computed a tree and support values for 158,022 distinct 16S ribosomal RNAs in 17 h and 2.4 GB of memory. Just computing pairwise Jukes-Cantor distances and storing them, without inferring a tree or bootstrapping, would require 17 h and 50 GB of memory. In simulations, FastTree was slightly more accurate than Neighbor-Joining, BIONJ, or FastME; on genuine alignments, FastTree's topologies had higher likelihoods. FastTree is available at http://microbesonline.org/fasttree.
MeSH Terms
Algorithms
Evolution, Molecular
Models, Genetic
Phylogeny
Proteins/genetics
Sequence Alignment/methods
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Price Morgan N
Physical Biosciences Division, Lawrence Berkeley National Laboratory, CA, USA. morgannprice@yahoo.com
Dehal Paramvir S
Arkin Adam P
References (29)
29 references, click to expand
-
Amino acid substitution matrices from protein blocks.
Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9
PMID: 1438297
-
Accurate and robust phylogeny estimation based on profile distances: a study of the Chlorophyceae (Chlorophyta).
BMC Evol Biol. 2004 Jun 28;4:20
PMID: 15222898
-
Protein molecular function prediction by Bayesian phylogenomics.
PLoS Comput Biol. 2005 Oct;1(5):e45
PMID: 16217548
-
A note on the neighbor-joining algorithm of Saitou and Nei.
Mol Biol Evol. 1988 Nov;5(6):729-31
PMID: 3221794
-
The neighbor-joining method: a new method for reconstructing phylogenetic trees.
Mol Biol Evol. 1987 Jul;4(4):406-25
PMID: 3447015
-
Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements.
Nucleic Acids Res. 2001 Jul 15;29(14):2994-3005
PMID: 11452024
-
Assessment of protein distance measures and tree-building methods for phylogenetic tree reconstruction.
Mol Biol Evol. 2005 Nov;22(11):2257-64
PMID: 16049194
-
Neighbor-joining revealed.
Mol Biol Evol. 2006 Nov;23(11):1997-2000
PMID: 16877499
-
RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models.
Bioinformatics. 2006 Nov 1;22(21):2688-90
PMID: 16928733
-
Accurate and scalable identification of functional sites by evolutionary tracing.
J Struct Funct Genomics. 2003;4(2-3):159-66
PMID: 14649300
-
Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach.
Biometrics. 1988 Sep;44(3):837-45
PMID: 3203132
-
Relaxed neighbor joining: a fast distance-based phylogenetic tree construction method.
J Mol Evol. 2006 Jun;62(6):785-92
PMID: 16752216
-
Greengenes, a chimera-checked 16S rRNA gene database and workbench compatible with ARB.
Appl Environ Microbiol. 2006 Jul;72(7):5069-72
PMID: 16820507
-
BIONJ: an improved version of the NJ algorithm based on a simple model of sequence data.
Mol Biol Evol. 1997 Jul;14(7):685-95
PMID: 9254330
-
Scoredist: a simple and robust protein sequence distance estimator.
BMC Bioinformatics. 2005 Apr 27;6:108
PMID: 15857510
-
QuickTree: building huge Neighbour-Joining trees of protein sequences.
Bioinformatics. 2002 Nov;18(11):1546-7
PMID: 12424131
-
Scaling of accuracy in extremely large phylogenetic trees.
Pac Symp Biocomput. 2001;:547-58
PMID: 11262972
-
Quantitative phylogenetic assessment of microbial communities in diverse environments.
Science. 2007 Feb 23;315(5815):1126-30
PMID: 17272687
-
The COG database: new developments in phylogenetic classification of proteins from complete genomes.
Nucleic Acids Res. 2001 Jan 1;29(1):22-8
PMID: 11125040
-
The MicrobesOnline Web site for comparative genomics.
Genome Res. 2005 Jul;15(7):1015-22
PMID: 15998914
-
TreeFam: a curated database of phylogenetic trees of animal gene families.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D572-80
PMID: 16381935
-
A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood.
Syst Biol. 2003 Oct;52(5):696-704
PMID: 14530136
-
Pfam: clans, web tools and services.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D247-51
PMID: 16381856
-
Fast and accurate phylogeny reconstruction algorithms based on the minimum-evolution principle.
J Comput Biol. 2002;9(5):687-705
PMID: 12487758
-
Rose: generating sequence families.
Bioinformatics. 1998;14(2):157-63
PMID: 9545448
-
The optimization principle in phylogenetic analysis tends to give incorrect topologies when the number of nucleotides or amino acids used is small.
Proc Natl Acad Sci U S A. 1998 Oct 13;95(21):12390-7
PMID: 9770497
-
Phylogenomics: improving functional predictions for uncharacterized genes by evolutionary analysis.
Genome Res. 1998 Mar;8(3):163-7
PMID: 9521918
-
CONFIDENCE LIMITS ON PHYLOGENIES: AN APPROACH USING THE BOOTSTRAP.
Evolution. 1985 Jul;39(4):783-791
PMID: 28561359
-
Bootstrap confidence levels for phylogenetic trees.
Proc Natl Acad Sci U S A. 1996 Nov 12;93(23):13429-34
PMID: 8917608