Abstract
Metagenomics is the study of microbial communities sampled directly from their natural environment, without prior culturing. Among the computational tools recently developed for metagenomic sequence analysis, binning tools attempt to classify the sequences in a metagenomic dataset into different bins (i.e., species), based on various DNA composition patterns (e.g., the tetramer frequencies) of various genomes. Composition-based binning methods, however, cannot be used to classify very short fragments, because of the substantial variation of DNA composition patterns within a single genome. We developed a novel approach (AbundanceBin) for metagenomics binning by utilizing the different abundances of species living in the same environment. AbundanceBin is an application of the Lander-Waterman model to metagenomics, which is based on the l-tuple content of the reads. AbundanceBin achieved accurate, unsupervised, clustering of metagenomic sequences into different bins, such that the reads classified in a bin belong to species of identical or very similar abundances in the sample. In addition, AbundanceBin gave accurate estimations of species abundances, as well as their genome sizes-two important parameters for characterizing a microbial community. We also show that AbundanceBin performed well when the sequence lengths are very short (e.g., 75 bp) or have sequencing errors. By combining AbundanceBin and a composition-based method (MetaCluster), we can achieve even higher binning accuracy. Supplementary Material is available at www.liebertonline.com/cmb .
MeSH Terms
Algorithms
DNA/genetics
Databases, Genetic
Genome, Bacterial
Metagenomics/methods
Sequence Analysis, DNA
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Wu Yu-Wei
School of Informatics and Computing, Indiana University, Bloomington, Indiana, USA.
Ye Yuzhen
References (31)
31 references, click to expand
-
Barcodes for genomes and applications.
BMC Bioinformatics. 2008 Dec 17;9:546
PMID: 19091119
-
An obesity-associated gut microbiome with increased capacity for energy harvest.
Nature. 2006 Dec 21;444(7122):1027-31
PMID: 17183312
-
TETRA: a web-service and a stand-alone program for the analysis and comparison of tetranucleotide usage patterns in DNA sequences.
BMC Bioinformatics. 2004 Oct 26;5:163
PMID: 15507136
-
Symbiosis insights through metagenomic analysis of a microbial consortium.
Nature. 2006 Oct 26;443(7114):950-5
PMID: 16980956
-
Community structure and metabolism through reconstruction of microbial genomes from the environment.
Nature. 2004 Mar 4;428(6978):37-43
PMID: 14961025
-
Whole-genome re-sequencing.
Curr Opin Genet Dev. 2006 Dec;16(6):545-52
PMID: 17055251
-
Figaro: a novel statistical method for vector sequence removal.
Bioinformatics. 2008 Feb 15;24(4):462-7
PMID: 18202027
-
Genomic mapping by fingerprinting random clones: a mathematical analysis.
Genomics. 1988 Apr;2(3):231-9
PMID: 3294162
-
Accuracy and quality of massively parallel DNA pyrosequencing.
Genome Biol. 2007;8(7):R143
PMID: 17659080
-
DNA sequencing: bench to bedside and beyond.
Nucleic Acids Res. 2007;35(18):6227-37
PMID: 17855400
-
Functional metagenomic profiling of nine biomes.
Nature. 2008 Apr 3;452(7187):629-32
PMID: 18337718
-
Estimating the repeat structure and length of DNA sequences using L-tuples.
Genome Res. 2003 Aug;13(8):1916-22
PMID: 12902383
-
Comparative genomic structure of prokaryotes.
Annu Rev Genet. 2004;38:771-92
PMID: 15568993
-
TREE-PUZZLE: maximum likelihood phylogenetic analysis using quartets and parallel computing.
Bioinformatics. 2002 Mar;18(3):502-4
PMID: 11934758
-
Quantitative phylogenetic assessment of microbial communities in diverse environments.
Science. 2007 Feb 23;315(5815):1126-30
PMID: 17272687
-
A simple, fast, and accurate method of phylogenomic inference.
Genome Biol. 2008 Oct 13;9(10):R151
PMID: 18851752
-
Phymm and PhymmBL: metagenomic phylogenetic classification with interpolated Markov models.
Nat Methods. 2009 Sep;6(9):673-6
PMID: 19648916
-
A core gut microbiome in obese and lean twins.
Nature. 2009 Jan 22;457(7228):480-4
PMID: 19043404
-
Comparative metagenomics of microbial communities.
Science. 2005 Apr 22;308(5721):554-7
PMID: 15845853
-
A detailed analysis of 16S ribosomal RNA gene segments for the diagnosis of pathogenic bacteria.
J Microbiol Methods. 2007 May;69(2):330-9
PMID: 17391789
-
A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood.
Syst Biol. 2003 Oct;52(5):696-704
PMID: 14530136
-
TACOA: taxonomic classification of environmental genomic fragments using a kernelized nearest neighbor approach.
BMC Bioinformatics. 2009 Feb 11;10:56
PMID: 19210774
-
Genome sequencing in microfabricated high-density picolitre reactors.
Nature. 2005 Sep 15;437(7057):376-80
PMID: 16056220
-
Pfam: clans, web tools and services.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D247-51
PMID: 16381856
-
Metagenomics: from acid mine to shining sea.
Environ Microbiol. 2004 Jun;6(6):543-5
PMID: 15142241
-
Phylogenetic classification of short environmental DNA fragments.
Nucleic Acids Res. 2008 Apr;36(7):2230-9
PMID: 18285365
-
Environments shape the nucleotide composition of genomes.
EMBO Rep. 2005 Dec;6(12):1208-13
PMID: 16200051
-
Microbial ecology of four coral atolls in the Northern Line Islands.
PLoS One. 2008 Feb 27;3(2):e1584
PMID: 18301735
-
Taxonomic distribution of large DNA viruses in the sea.
Genome Biol. 2008;9(7):R106
PMID: 18598358
-
MEGAN analysis of metagenomic data.
Genome Res. 2007 Mar;17(3):377-86
PMID: 17255551
-
Toward automatic reconstruction of a highly resolved tree of life.
Science. 2006 Mar 3;311(5765):1283-7
PMID: 16513982