Home LiteratureArticle Details
PMID: 21436105 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Performance, accuracy, and Web server for evolutionary placement of short sequence reads under maximum likelihood.

Systematic biology ·Vol. 60 ·No. 3 ·2011-05-00 ·Pages 291-302

Berger SA, Krompass D, Stamatakis A

Abstract

We present an evolutionary placement algorithm (EPA) and a Web server for the rapid assignment of sequence fragments (short reads) to edges of a given phylogenetic tree under the maximum-likelihood model. The accuracy of the algorithm is evaluated on several real-world data sets and compared with placement by pair-wise sequence comparison, using edit distances and BLAST. We introduce a slow and accurate as well as a fast and less accurate placement algorithm. For the slow algorithm, we develop additional heuristic techniques that yield almost the same run times as the fast version with only a small loss of accuracy. When those additional heuristics are employed, the run time of the more accurate algorithm is comparable with that of a simple BLAST search for data sets with a high number of short query sequences. Moreover, the accuracy of the EPA is significantly higher, in particular when the sample of taxa in the reference topology is sparse or inadequate. Our algorithm, which has been integrated into RAxML, therefore provides an equally fast but more accurate alternative to BLAST for tree-based inference of the evolutionary origin and composition of short sequence reads. We are also actively developing a Web server that offers a freely available service for computing read placements on trees using the EPA.

MeSH Terms
Algorithms Amino Acid Sequence Base Sequence Computer Simulation Evolution, Molecular Internet Likelihood Functions Phylogeny Sequence Alignment/methods Sequence Analysis, DNA/methods Sequence Analysis, Protein/methods Sequence Analysis, RNA/methods Software
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Berger Simon A
The Exelixis Lab, Scientific Computing Group, Heidelberg Institute for Theoretical Studies, Schloss-Wolfsbrunnenweg 35, D-69118 Heidelberg, Germany.
Krompass Denis
Stamatakis Alexandros
References (36)
36 references, click to expand
  1. Statistical assignment of DNA sequences using Bayesian phylogenetics.
    Syst Biol. 2008 Oct;57(5):750-7 PMID: 18853361
  2. A rapid bootstrap algorithm for the RAxML Web servers.
    Syst Biol. 2008 Oct;57(5):758-71 PMID: 18853362
  3. Evolutionary trees from DNA sequences: a maximum likelihood approach.
    J Mol Evol. 1981;17(6):368-76 PMID: 7288891
  4. NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
    Nucleic Acids Res. 2007 Jan;35(Database issue):D61-5 PMID: 17130148
  5. ARB: a software environment for sequence data.
    Nucleic Acids Res. 2004 Feb 25;32(4):1363-71 PMID: 14985472
  6. Matched-pairs tests of homogeneity with applications to homologous nucleotide sequences.
    Bioinformatics. 2006 May 15;22(10):1225-31 PMID: 16492684
  7. Unexpected diversity and complexity of the Guerrero Negro hypersaline microbial mat.
    Appl Environ Microbiol. 2006 May;72(5):3685-95 PMID: 16672518
  8. SeqVis: visualization of compositional heterogeneity in large alignments of nucleotides.
    Bioinformatics. 2006 Sep 1;22(17):2162-3 PMID: 16766557
  9. Quantitative phylogenetic assessment of microbial communities in diverse environments.
    Science. 2007 Feb 23;315(5815):1126-30 PMID: 17272687
  10. Statistical approaches for DNA barcoding.
    Syst Biol. 2006 Feb;55(1):162-9 PMID: 16507534
  11. A detailed analysis of 16S ribosomal RNA gene segments for the diagnosis of pathogenic bacteria.
    J Microbiol Methods. 2007 May;69(2):330-9 PMID: 17391789
  12. phyloXML: XML for evolutionary biology and comparative genomics.
    BMC Bioinformatics. 2009 Oct 27;10:356 PMID: 19860910
  13. Obesity alters gut microbial ecology.
    Proc Natl Acad Sci U S A. 2005 Aug 2;102(31):11070-5 PMID: 16033867
  14. Profile hidden Markov models.
    Bioinformatics. 1998;14(9):755-63 PMID: 9918945
  15. Accurate phylogenetic classification of variable-length DNA fragments.
    Nat Methods. 2007 Jan;4(1):63-72 PMID: 17179938
  16. The closest BLAST hit is often not the nearest neighbor.
    J Mol Evol. 2001 Jun;52(6):540-2 PMID: 11443357
  17. Worlds within worlds: evolution of the vertebrate gut microbiota.
    Nat Rev Microbiol. 2008 Oct;6(10):776-88 PMID: 18794915
  18. Estimation of phylogeny and invariant sites under the general Markov model of nucleotide sequence evolution.
    Syst Biol. 2007 Apr;56(2):155-62 PMID: 17454972
  19. Methanogenic communities in permafrost-affected soils of the Laptev Sea coast, Siberian Arctic, characterized by 16S rRNA gene fingerprints.
    FEMS Microbiol Ecol. 2007 Feb;59(2):476-88 PMID: 16978241
  20. Phymm and PhymmBL: metagenomic phylogenetic classification with interpolated Markov models.
    Nat Methods. 2009 Sep;6(9):673-6 PMID: 19648916
  21. A core gut microbiome in obese and lean twins.
    Nature. 2009 Jan 22;457(7228):480-4 PMID: 19043404
  22. Inferring confidence sets of possibly misspecified gene trees.
    Proc Biol Sci. 2002 Jan 22;269(1487):137-42 PMID: 11798428
  23. The biasing effect of compositional heterogeneity on phylogenetic estimates may be underestimated.
    Syst Biol. 2004 Aug;53(4):638-43 PMID: 15371251
  24. UniFrac: a new phylogenetic method for comparing microbial communities.
    Appl Environ Microbiol. 2005 Dec;71(12):8228-35 PMID: 16332807
  25. MEGAN analysis of metagenomic data.
    Genome Res. 2007 Mar;17(3):377-86 PMID: 17255551
  26. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  27. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models.
    Bioinformatics. 2006 Nov 1;22(21):2688-90 PMID: 16928733
  28. Pyrosequencing sheds light on DNA sequencing.
    Genome Res. 2001 Jan;11(1):3-11 PMID: 11156611
  29. RAxML-III: a fast program for maximum likelihood-based inference of large phylogenetic trees.
    Bioinformatics. 2005 Feb 15;21(4):456-63 PMID: 15608047
  30. CONFIDENCE LIMITS ON PHYLOGENIES: AN APPROACH USING THE BOOTSTRAP.
    Evolution. 1985 Jul;39(4):783-791 PMID: 28561359
  31. NAST: a multiple sequence alignment server for comparative analysis of 16S rRNA genes.
    Nucleic Acids Res. 2006 Jul 1;34(Web Server issue):W394-9 PMID: 16845035
  32. The influence of sex, handedness, and washing on the diversity of hand surface bacteria.
    Proc Natl Acad Sci U S A. 2008 Nov 18;105(46):17994-9 PMID: 19004758
  33. Inferring pattern and process: maximum-likelihood implementation of a nonhomogeneous model of DNA sequence evolution for phylogenetic analysis.
    Mol Biol Evol. 1998 Jul;15(7):871-9 PMID: 9656487
  34. pplacer: linear time maximum-likelihood and Bayesian phylogenetic placement of sequences onto a fixed reference tree.
    BMC Bioinformatics. 2010 Oct 30;11:538 PMID: 21034504
  35. MUSCLE: multiple sequence alignment with high accuracy and high throughput.
    Nucleic Acids Res. 2004 Mar 19;32(5):1792-7 PMID: 15034147
  36. MAFFT version 5: improvement in accuracy of multiple sequence alignment.
    Nucleic Acids Res. 2005 Jan 20;33(2):511-8 PMID: 15661851
Article Info
Journal
Systematic biology
Abbr.
Syst Biol
ISSN
1076-836X
Published
2011-05-00
Epub
2011-00-23
Pages
291-302
Language
English
Region
England
NLM ID
9302532
PMCID
PMC3078422
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com