Abstract
Exhaustive gene identification is a fundamental goal in all metagenomics projects. However, most metagenomic sequences are unassembled anonymous fragments, and conventional gene-finding methods cannot be applied. We have developed a prokaryotic gene-finding program, MetaGene, which utilizes di-codon frequencies estimated by the GC content of a given sequence with other various measures. MetaGene can predict a whole range of prokaryotic genes based on the anonymous genomic sequences of a few hundred bases, with a sensitivity of 95% and a specificity of 90% for artificial shotgun sequences (700 bp fragments from 12 species). MetaGene has two sets of codon frequency interpolations, one for bacteria and one for archaea, and automatically selects the proper set for a given sequence using the domain classification method we propose. The domain classification works properly, correctly assigning domain information to more than 90% of the artificial shotgun sequences. Applied to the Sargasso Sea dataset, MetaGene predicted almost all of the annotated genes and a notable number of novel genes. MetaGene can be applied to wide variety of metagenomic projects and expands the utility of metagenomics.
MeSH Terms
Computational Biology
Environment
GC Rich Sequence
Genes, Archaeal
Genes, Bacterial
Genome, Archaeal
Genome, Bacterial
Genomics/methods
Internet
Oceans and Seas
Open Reading Frames
Software
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Noguchi Hideki
Department of Computational Biology, Graduate School of Frontier Sciences, University of Tokyo, Kashiwa, Chiba 277-8562, Japan. hide@cb.k.u-tokyo.ac.jp
Park Jungho
Takagi Toshihisa
References (30)
30 references, click to expand
-
Detecting anomalous gene clusters and pathogenicity islands in diverse bacterial genomes.
Trends Microbiol. 2001 Jul;9(7):335-43
PMID: 11435108
-
How to interpret an anonymous bacterial genome: machine learning approach to gene identification.
Genome Res. 1998 Nov;8(11):1154-71
PMID: 9847079
-
On the convergence of a clustering algorithm for protein-coding regions in microbial genomes.
Bioinformatics. 2000 Apr;16(4):367-71
PMID: 10869034
-
Bioinformatics for whole-genome shotgun sequencing of microbial communities.
PLoS Comput Biol. 2005 Jul;1(2):106-12
PMID: 16110337
-
Exploring prokaryotic diversity in the genomic era.
Genome Biol. 2002;3(2):REVIEWS0003
PMID: 11864374
-
A hidden Markov model that finds genes in E. coli DNA.
Nucleic Acids Res. 1994 Nov 11;22(22):4768-78
PMID: 7984429
-
Microbial gene identification using interpolated Markov models.
Nucleic Acids Res. 1998 Jan 15;26(2):544-8
PMID: 9421513
-
Large-scale prokaryotic gene prediction and comparison to genome annotation.
Bioinformatics. 2005 Dec 15;21(24):4322-9
PMID: 16249266
-
Recognition of protein coding regions in DNA sequences.
Nucleic Acids Res. 1982 Sep 11;10(17):5303-18
PMID: 7145702
-
The complete genome sequence of Chlorobium tepidum TLS, a photosynthetic, anaerobic, green-sulfur bacterium.
Proc Natl Acad Sci U S A. 2002 Jul 9;99(14):9509-14
PMID: 12093901
-
Measurements of the effects that coding for a protein has on a DNA sequence and their use for finding genes.
Nucleic Acids Res. 1984 Jan 11;12(1 Pt 2):551-67
PMID: 6364041
-
Community structure and metabolism through reconstruction of microbial genomes from the environment.
Nature. 2004 Mar 4;428(6978):37-43
PMID: 14961025
-
Finding prokaryotic genes by the 'frame-by-frame' algorithm: targeting gene starts and overlapping genes.
Bioinformatics. 1999 Nov;15(11):874-86
PMID: 10743554
-
Improved microbial gene identification with GLIMMER.
Nucleic Acids Res. 1999 Dec 1;27(23):4636-41
PMID: 10556321
-
Heuristic approach to deriving models for gene finding.
Nucleic Acids Res. 1999 Oct 1;27(19):3911-20
PMID: 10481031
-
Genomic sequencing of Pleistocene cave bears.
Science. 2005 Jul 22;309(5734):597-9
PMID: 15933159
-
Combining diverse evidence for gene recognition in completely sequenced bacterial genomes.
Nucleic Acids Res. 1998 Jun 15;26(12):2941-7
PMID: 9611239
-
A novel bacterial gene-finding system with improved accuracy in locating start codons.
DNA Res. 2001 Jun 30;8(3):97-106
PMID: 11475327
-
Detecting pathogenicity islands and anomalous gene clusters by iterative discriminant analysis.
FEMS Microbiol Lett. 2003 Apr 25;221(2):269-75
PMID: 12725938
-
GeneMark.hmm: new solutions for gene finding.
Nucleic Acids Res. 1998 Feb 15;26(4):1107-15
PMID: 9461475
-
Metagenomics to paleogenomics: large-scale sequencing of mammoth DNA.
Science. 2006 Jan 20;311(5759):392-4
PMID: 16368896
-
GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions.
Nucleic Acids Res. 2001 Jun 15;29(12):2607-18
PMID: 11410670
-
Reverse methanogenesis: testing the hypothesis with environmental genomics.
Science. 2004 Sep 3;305(5689):1457-62
PMID: 15353801
-
Environmental genome shotgun sequencing of the Sargasso Sea.
Science. 2004 Apr 2;304(5667):66-74
PMID: 15001713
-
Modeling and predicting transcriptional units of Escherichia coli genes using hidden Markov models.
Bioinformatics. 1999 Dec;15(12):987-93
PMID: 10745988
-
The codon preference plot: graphic analysis of protein coding sequences and prediction of gene expression.
Nucleic Acids Res. 1984 Jan 11;12(1 Pt 2):539-49
PMID: 6694906
-
The uncultured microbial majority.
Annu Rev Microbiol. 2003;57:369-94
PMID: 14527284
-
Comparative metagenomics of microbial communities.
Science. 2005 Apr 22;308(5721):554-7
PMID: 15845853
-
Complete genome sequence of the methanogenic archaeon, Methanococcus jannaschii.
Science. 1996 Aug 23;273(5278):1058-73
PMID: 8688087
-
Self-identification of protein-coding regions in microbial genomes.
Proc Natl Acad Sci U S A. 1998 Aug 18;95(17):10026-31
PMID: 9707594