Home LiteratureArticle Details
PMID: 21385052 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

A novel abundance-based algorithm for binning metagenomic sequences using l-tuples.

Wu YW, Ye Y

Abstract

Metagenomics is the study of microbial communities sampled directly from their natural environment, without prior culturing. Among the computational tools recently developed for metagenomic sequence analysis, binning tools attempt to classify the sequences in a metagenomic dataset into different bins (i.e., species), based on various DNA composition patterns (e.g., the tetramer frequencies) of various genomes. Composition-based binning methods, however, cannot be used to classify very short fragments, because of the substantial variation of DNA composition patterns within a single genome. We developed a novel approach (AbundanceBin) for metagenomics binning by utilizing the different abundances of species living in the same environment. AbundanceBin is an application of the Lander-Waterman model to metagenomics, which is based on the l-tuple content of the reads. AbundanceBin achieved accurate, unsupervised, clustering of metagenomic sequences into different bins, such that the reads classified in a bin belong to species of identical or very similar abundances in the sample. In addition, AbundanceBin gave accurate estimations of species abundances, as well as their genome sizes-two important parameters for characterizing a microbial community. We also show that AbundanceBin performed well when the sequence lengths are very short (e.g., 75 bp) or have sequencing errors. By combining AbundanceBin and a composition-based method (MetaCluster), we can achieve even higher binning accuracy. Supplementary Material is available at www.liebertonline.com/cmb .

MeSH Terms
Algorithms DNA/genetics Databases, Genetic Genome, Bacterial Metagenomics/methods Sequence Analysis, DNA
Chemicals
DNA
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Wu Yu-Wei
School of Informatics and Computing, Indiana University, Bloomington, Indiana, USA.
Ye Yuzhen
References (31)
31 references, click to expand
  1. Barcodes for genomes and applications.
    BMC Bioinformatics. 2008 Dec 17;9:546 PMID: 19091119
  2. An obesity-associated gut microbiome with increased capacity for energy harvest.
    Nature. 2006 Dec 21;444(7122):1027-31 PMID: 17183312
  3. TETRA: a web-service and a stand-alone program for the analysis and comparison of tetranucleotide usage patterns in DNA sequences.
    BMC Bioinformatics. 2004 Oct 26;5:163 PMID: 15507136
  4. Symbiosis insights through metagenomic analysis of a microbial consortium.
    Nature. 2006 Oct 26;443(7114):950-5 PMID: 16980956
  5. Community structure and metabolism through reconstruction of microbial genomes from the environment.
    Nature. 2004 Mar 4;428(6978):37-43 PMID: 14961025
  6. Whole-genome re-sequencing.
    Curr Opin Genet Dev. 2006 Dec;16(6):545-52 PMID: 17055251
  7. Figaro: a novel statistical method for vector sequence removal.
    Bioinformatics. 2008 Feb 15;24(4):462-7 PMID: 18202027
  8. Genomic mapping by fingerprinting random clones: a mathematical analysis.
    Genomics. 1988 Apr;2(3):231-9 PMID: 3294162
  9. Accuracy and quality of massively parallel DNA pyrosequencing.
    Genome Biol. 2007;8(7):R143 PMID: 17659080
  10. DNA sequencing: bench to bedside and beyond.
    Nucleic Acids Res. 2007;35(18):6227-37 PMID: 17855400
  11. Functional metagenomic profiling of nine biomes.
    Nature. 2008 Apr 3;452(7187):629-32 PMID: 18337718
  12. Estimating the repeat structure and length of DNA sequences using L-tuples.
    Genome Res. 2003 Aug;13(8):1916-22 PMID: 12902383
  13. Comparative genomic structure of prokaryotes.
    Annu Rev Genet. 2004;38:771-92 PMID: 15568993
  14. TREE-PUZZLE: maximum likelihood phylogenetic analysis using quartets and parallel computing.
    Bioinformatics. 2002 Mar;18(3):502-4 PMID: 11934758
  15. Quantitative phylogenetic assessment of microbial communities in diverse environments.
    Science. 2007 Feb 23;315(5815):1126-30 PMID: 17272687
  16. A simple, fast, and accurate method of phylogenomic inference.
    Genome Biol. 2008 Oct 13;9(10):R151 PMID: 18851752
  17. Phymm and PhymmBL: metagenomic phylogenetic classification with interpolated Markov models.
    Nat Methods. 2009 Sep;6(9):673-6 PMID: 19648916
  18. A core gut microbiome in obese and lean twins.
    Nature. 2009 Jan 22;457(7228):480-4 PMID: 19043404
  19. Comparative metagenomics of microbial communities.
    Science. 2005 Apr 22;308(5721):554-7 PMID: 15845853
  20. A detailed analysis of 16S ribosomal RNA gene segments for the diagnosis of pathogenic bacteria.
    J Microbiol Methods. 2007 May;69(2):330-9 PMID: 17391789
  21. A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood.
    Syst Biol. 2003 Oct;52(5):696-704 PMID: 14530136
  22. TACOA: taxonomic classification of environmental genomic fragments using a kernelized nearest neighbor approach.
    BMC Bioinformatics. 2009 Feb 11;10:56 PMID: 19210774
  23. Genome sequencing in microfabricated high-density picolitre reactors.
    Nature. 2005 Sep 15;437(7057):376-80 PMID: 16056220
  24. Pfam: clans, web tools and services.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D247-51 PMID: 16381856
  25. Metagenomics: from acid mine to shining sea.
    Environ Microbiol. 2004 Jun;6(6):543-5 PMID: 15142241
  26. Phylogenetic classification of short environmental DNA fragments.
    Nucleic Acids Res. 2008 Apr;36(7):2230-9 PMID: 18285365
  27. Environments shape the nucleotide composition of genomes.
    EMBO Rep. 2005 Dec;6(12):1208-13 PMID: 16200051
  28. Microbial ecology of four coral atolls in the Northern Line Islands.
    PLoS One. 2008 Feb 27;3(2):e1584 PMID: 18301735
  29. Taxonomic distribution of large DNA viruses in the sea.
    Genome Biol. 2008;9(7):R106 PMID: 18598358
  30. MEGAN analysis of metagenomic data.
    Genome Res. 2007 Mar;17(3):377-86 PMID: 17255551
  31. Toward automatic reconstruction of a highly resolved tree of life.
    Science. 2006 Mar 3;311(5765):1283-7 PMID: 16513982
Article Info
Journal
Journal of computational biology : a journal of computational molecular cell biology
Abbr.
J Comput Biol
ISSN
1557-8666
Published
2011-03-00
Pages
523-34
Language
English
Region
United States
NLM ID
9433358
PMCID
PMC3123841
Subset
IM
Grants
NHGRI NIH HHS · 1R01HG004908-02 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com