Home LiteratureArticle Details
PMID: 9847082 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Predicting gene regulatory elements in silico on a genomic scale.

Genome research ·Vol. 8 ·No. 11 ·1998-11-00 ·Pages 1202-15

Brazma A, Jonassen I, Vilo J, Ukkonen E

Abstract

We performed a systematic analysis of gene upstream regions in the yeast genome for occurrences of regular expression-type patterns with the goal of identifying potential regulatory elements. To achieve this goal, we have developed a new sequence pattern discovery algorithm that searches exhaustively for a priori unknown regular expression-type patterns that are over-represented in a given set of sequences. We applied the algorithm in two cases, (1) discovery of patterns in the complete set of >6000 sequences taken upstream of the putative yeast genes and (2) discovery of patterns in the regions upstream of the genes with similar expression profiles. In the first case, we looked for patterns that occur more frequently in the gene upstream regions than in the genome overall. In the second case, first we clustered the upstream regions of all the genes by similarity of their expression profiles on the basis of publicly available gene expression data and then looked for sequence patterns that are over-represented in each cluster. In both cases we considered each pattern that occurred at least in some minimum number of sequences, and rated them on the basis of their over-representation. Among the highest rating patterns, most have matches to substrings in known yeast transcription factor-binding sites. Moreover, several of them are known to be relevant to the expression of the genes from the respective clusters. Experiments on simulated data show that the majority of the discovered patterns are not expected to occur by chance.

MeSH Terms
Algorithms Gene Expression Genes, Fungal/genetics Genome, Fungal Regulatory Sequences, Nucleic Acid Saccharomyces cerevisiae/genetics
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Brazma A
European Molecular Biology Laboratory (EMBL) Outstation-Hinxton, European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK.
Jonassen I
Vilo J
Ukkonen E
References (28)
28 references, click to expand
  1. Information content of binding sites on nucleotide sequences.
    J Mol Biol. 1986 Apr 5;188(3):415-31 PMID: 3525846
  2. Expectation maximization algorithm for identifying protein-binding sites with variable lengths from unaligned DNA fragments.
    J Mol Biol. 1992 Jan 5;223(1):159-70 PMID: 1731067
  3. Identifying protein-binding sites from unaligned DNA fragments.
    Proc Natl Acad Sci U S A. 1989 Feb;86(4):1183-7 PMID: 2919167
  4. Identification of a Saccharomyces cerevisiae DNA-binding protein involved in transcriptional regulation.
    Mol Cell Biol. 1990 Apr;10(4):1743-53 PMID: 2181283
  5. Weight matrix descriptions of four eukaryotic RNA polymerase II promoter elements derived from 502 unrelated promoter sequences.
    J Mol Biol. 1990 Apr 20;212(4):563-78 PMID: 2329577
  6. A relational database of transcription factors.
    Nucleic Acids Res. 1990 Apr 11;18(7):1749-56 PMID: 2186365
  7. Association of RAP1 binding sites with stringent control of ribosomal protein gene transcription in Saccharomyces cerevisiae.
    Mol Cell Biol. 1991 May;11(5):2723-35 PMID: 2017175
  8. Nomenclature for incompletely specified bases in nucleic acid sequences: recommendations 1984.
    Nucleic Acids Res. 1985 May 10;13(9):3021-30 PMID: 2582368
  9. Approaches to the automatic discovery of patterns in biosequences.
    J Comput Biol. 1998 Summer;5(2):279-305 PMID: 9672833
  10. Distribution of transcription factor binding sites in the yeast genome suggests abundance of coordinately regulated genes.
    Genomics. 1998 Jun 1;50(2):293-5 PMID: 9653659
  11. Extracting regulatory sites from the upstream region of yeast genes by computational analysis of oligonucleotide frequencies.
    J Mol Biol. 1998 Sep 4;281(5):827-42 PMID: 9719638
  12. Cluster analysis and data visualization of large-scale gene expression data.
    Pac Symp Biocomput. 1998;:42-53 PMID: 9697170
  13. Automatic extraction of motifs represented in the hidden Markov model from a number of DNA sequences.
    Bioinformatics. 1998;14(4):317-25 PMID: 9632826
  14. DNA chips: state-of-the art.
    Nat Biotechnol. 1998 Jan;16(1):40-4 PMID: 9447591
  15. Genome-wide expression monitoring in Saccharomyces cerevisiae.
    Nat Biotechnol. 1997 Dec;15(13):1359-67 PMID: 9415887
  16. Efficient discovery of conserved patterns using a pattern graph.
    Comput Appl Biosci. 1997 Oct;13(5):509-22 PMID: 9367124
  17. Exploring the metabolic and genetic control of gene expression on a genomic scale.
    Science. 1997 Oct 24;278(5338):680-6 PMID: 9381177
  18. Overview of the yeast genome.
    Nature. 1997 May 29;387(6632 Suppl):7-65 PMID: 9169865
  19. Software for the analysis of DNA sequence elements of transcription.
    Comput Appl Biosci. 1997 Feb;13(1):89-97 PMID: 9088714
  20. Characterization of the yeast transcriptome.
    Cell. 1997 Jan 24;88(2):243-51 PMID: 9008165
  21. Life with 6000 genes.
    Science. 1996 Oct 25;274(5287):546, 563-7 PMID: 8849441
  22. Identification of functional elements in unaligned nucleic acid sequences by a novel tuple search algorithm.
    Comput Appl Biosci. 1996 Feb;12(1):71-80 PMID: 8670622
  23. TRANSFAC: a database on transcription factors and their DNA binding sites.
    Nucleic Acids Res. 1996 Jan 1;24(1):238-41 PMID: 8594589
  24. MATRIX SEARCH 1.0: a computer program that scans DNA sequences for transcriptional elements using a database of weight matrices.
    Comput Appl Biosci. 1995 Oct;11(5):563-6 PMID: 8590181
  25. MatInd and MatInspector: new fast and versatile tools for detection of consensus matches in nucleotide sequence data.
    Nucleic Acids Res. 1995 Dec 11;23(23):4878-84 PMID: 8532532
  26. Quantitative monitoring of gene expression patterns with a complementary DNA microarray.
    Science. 1995 Oct 20;270(5235):467-70 PMID: 7569999
  27. Computer-assisted prediction, classification, and delimitation of protein binding sites in nucleic acids.
    Nucleic Acids Res. 1993 Apr 11;21(7):1655-64 PMID: 8479918
  28. Purification of a yeast protein that binds to origins of DNA replication and a transcriptional silencer.
    Proc Natl Acad Sci U S A. 1988 Apr;85(7):2120-4 PMID: 3281162
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
1998-11-00
Pages
1202-15
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC310790
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com