Abstract
Bacterial and fungal secondary metabolism is a rich source of novel bioactive compounds with potential pharmaceutical applications as antibiotics, anti-tumor drugs or cholesterol-lowering drugs. To find new drug candidates, microbiologists are increasingly relying on sequencing genomes of a wide variety of microbes. However, rapidly and reliably pinpointing all the potential gene clusters for secondary metabolites in dozens of newly sequenced genomes has been extremely challenging, due to their biochemical heterogeneity, the presence of unknown enzymes and the dispersed nature of the necessary specialized bioinformatics tools and resources. Here, we present antiSMASH (antibiotics & Secondary Metabolite Analysis Shell), the first comprehensive pipeline capable of identifying biosynthetic loci covering the whole range of known secondary metabolite compound classes (polyketides, non-ribosomal peptides, terpenes, aminoglycosides, aminocoumarins, indolocarbazoles, lantibiotics, bacteriocins, nucleosides, beta-lactams, butyrolactones, siderophores, melanins and others). It aligns the identified regions at the gene cluster level to their nearest relatives from a database containing all other known gene clusters, and integrates or cross-links all previously available secondary-metabolite specific gene analysis methods in one interactive view. antiSMASH is available at http://antismash.secondarymetabolites.org.
MeSH Terms
Bacteria/enzymology,genetics,metabolism
Biosynthetic Pathways/genetics
Fungi/enzymology,genetics,metabolism
Genes, Bacterial
Genes, Fungal
Genome, Bacterial
Genome, Fungal
Genomics
Internet
Molecular Sequence Annotation
Peptide Synthases/chemistry
Polyketide Synthases/chemistry
Software
Substrate Specificity
Chemicals
Polyketide Synthases
Peptide Synthases
non-ribosomal peptide synthase
Authors & Affiliations
9 authors, click to expand affiliations / ORCID
Medema Marnix H
Department of Microbial Physiology, Groningen Bioinformatics Centre, Groningen Biomolecular Sciences and Biotechnology Institute, University of Groningen, Nijenborgh 7, 9747AG Groningen, The Netherlands.
Blin Kai
Cimermancic Peter
de Jager Victor
Zakrzewski Piotr
Fischbach Michael A
Weber Tilmann
Takano Eriko
Breitling Rainer
References (28)
28 references, click to expand
-
SMURF: Genomic mapping of fungal secondary metabolite clusters.
Fungal Genet Biol. 2010 Sep;47(9):736-41
PMID: 20554054
-
Identifying bacterial genes and endosymbiont DNA with Glimmer.
Bioinformatics. 2007 Mar 15;23(6):673-9
PMID: 17237039
-
In silico analysis of methyltransferase domains involved in biosynthesis of secondary metabolites.
BMC Bioinformatics. 2008 Oct 25;9:454
PMID: 18950525
-
The Pfam protein families database.
Nucleic Acids Res. 2010 Jan;38(Database issue):D211-22
PMID: 19920124
-
Specificity prediction of adenylation domains in nonribosomal peptide synthetases (NRPS) using transductive support vector machines (TSVMs).
Nucleic Acids Res. 2005 Oct 12;33(18):5799-808
PMID: 16221976
-
Towards prediction of metabolic products of polyketide synthases: an in silico analysis.
PLoS Comput Biol. 2009 Apr;5(4):e1000351
PMID: 19360130
-
SMART 6: recent updates and new developments.
Nucleic Acids Res. 2009 Jan;37(Database issue):D229-32
PMID: 18978020
-
OrthoMCL: identification of ortholog groups for eukaryotic genomes.
Genome Res. 2003 Sep;13(9):2178-89
PMID: 12952885
-
NRPSpredictor2--a web server for predicting NRPS adenylation domain specificity.
Nucleic Acids Res. 2011 Jul;39(Web Server issue):W362-7
PMID: 21558170
-
Computational approach for prediction of domain organization and substrate specificity of modular polyketide synthases.
J Mol Biol. 2003 Apr 25;328(2):335-63
PMID: 12691745
-
SBSPKS: structure based sequence analysis of polyketide synthases.
Nucleic Acids Res. 2010 Jul;38(Web Server issue):W487-96
PMID: 20444870
-
FastTree 2--approximately maximum-likelihood trees for large alignments.
PLoS One. 2010 Mar 10;5(3):e9490
PMID: 20224823
-
Artemis: sequence visualization and annotation.
Bioinformatics. 2000 Oct;16(10):944-5
PMID: 11120685
-
Kernel based machine learning algorithm for the efficient prediction of type III polyketide synthase family of proteins.
J Integr Bioinform. 2010 Jul 13;7(1):
PMID: 20625199
-
The evolution of gene collectives: How natural selection drives chemical innovation.
Proc Natl Acad Sci U S A. 2008 Mar 25;105(12):4601-8
PMID: 18216259
-
Phylogenetic analysis of condensation domains in NRPS sheds light on their functional evolution.
BMC Evol Biol. 2007 May 16;7:78
PMID: 17506888
-
Comprehensive analysis of distinctive polyketide and nonribosomal peptide structural motifs encoded in microbial genomes.
J Mol Biol. 2007 May 18;368(5):1500-17
PMID: 17400247
-
ClustScan: an integrated program package for the semi-automatic annotation of modular biosynthetic gene clusters and in silico prediction of novel chemical structures.
Nucleic Acids Res. 2008 Dec;36(21):6882-92
PMID: 18978015
-
TreeGraph 2: combining and visualizing evidence from different phylogenetic analyses.
BMC Bioinformatics. 2010 Jan 05;11:7
PMID: 20051126
-
CLUSEAN: a computer-based framework for the automated analysis of bacterial secondary metabolite biosynthetic gene clusters.
J Biotechnol. 2009 Mar 10;140(1-2):13-7
PMID: 19297688
-
BAGEL2: mining for bacteriocins in genomic data.
Nucleic Acids Res. 2010 Jul;38(Web Server issue):W647-51
PMID: 20462861
-
Natural products version 2.0: connecting genes to molecules.
J Am Chem Soc. 2010 Mar 3;132(8):2469-93
PMID: 20121095
-
Comparative analysis and insights into the evolution of gene clusters for glycopeptide antibiotic biosynthesis.
Mol Genet Genomics. 2005 Aug;274(1):40-50
PMID: 16007453
-
Automated genome mining for natural products.
BMC Bioinformatics. 2009 Jun 16;10:185
PMID: 19531248
-
Exploiting plug-and-play synthetic biology for drug discovery and production in microorganisms.
Nat Rev Microbiol. 2011 Feb;9(2):131-7
PMID: 21189477
-
TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders.
Bioinformatics. 2004 Nov 1;20(16):2878-9
PMID: 15145805
-
BLAST+: architecture and applications.
BMC Bioinformatics. 2009 Dec 15;10:421
PMID: 20003500
-
MUSCLE: multiple sequence alignment with high accuracy and high throughput.
Nucleic Acids Res. 2004 Mar 19;32(5):1792-7
PMID: 15034147