Home LiteratureArticle Details
PMID: 17402923 Published · ppublish English Journal Article

Supervised detection of conserved motifs in DNA sequences with cosmo.

Statistical applications in genetics and molecular biology ·Vol. 6 ·2007-00-00 ·Pages Article8

Bembom O, Keles S, van der Laan MJ

Abstract

A number of computational methods have been proposed for identifying transcription factor binding sites from a set of unaligned sequences that are thought to share the motif in question. We here introduce an algorithm, called cosmo, that allows this search to be supervised by specifying a set of constraints that the position weight matrix of the unknown motif must satisfy. Such constraints may be formulated, for example, on the basis of prior knowledge about the structure of the transcription factor in question. The algorithm is based on the same two-component multinomial mixture model used by MEME, with stronger reliance, however, on the likelihood principle instead of more ad-hoc criteria like the E-value. The intensity parameter in the ZOOPS and TCM models, for instance, is estimated based on a profile-likelihood approach, and the width of the unknown motif is selected based on BIC. These changes allow cosmo to outperform MEME even in the absence of any constraints, as evidenced by 2- to 3-fold greater sensitivity in some simulation studies. Additional improvements in performance can be achieved by selecting the model type (OOPS, ZOOPS, or TCM) data-adaptively or by supplying correctly specified constraints, especially if the motif appears only as a weak signal in the data. The algorithm can data-adaptively choose between working in a given constrained model or in the completely unconstrained model, guarding against the risk of supplying mis-specified constraints. Simulation studies suggest that this approach can offer 3 to 3.5 times greater sensitivity than MEME. The algorithm has been implemented in the form of a stand-alone C program as well as a web application that can be accessed at http://cosmoweb.berkeley.edu. An R package is available through Bioconductor (http://bioconductor.org).

MeSH Terms
Algorithms Base Sequence Conserved Sequence DNA/chemistry,genetics Likelihood Functions Markov Chains Probability Sensitivity and Specificity
Chemicals
DNA
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Bembom Oliver
Division of Biostatistics, University of California, Berkeley, USA. bembom@berkeley.edu
Keles Sunduz
van der Laan Mark J
Article Info
Journal
Statistical applications in genetics and molecular biology
Abbr.
Stat Appl Genet Mol Biol
ISSN
1544-6115
Published
2007-00-00
Epub
2007-00-23
Pages
Article8
Language
English
Region
Germany
NLM ID
101176023
Subset
IM
Grants
NIAID NIH HHS · R01 AI074345 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com