Home LiteratureArticle Details
PMID: 14555624 Published · ppublish English Comparative Study Evaluation Study Journal Article Validation Study

Prediction of regulatory networks: genome-wide identification of transcription factor targets from gene expression data.

Bioinformatics (Oxford, England) ·Vol. 19 ·No. 15 ·2003-10-12 ·Pages 1917-26

Qian J, Lin J, Luscombe NM, Yu H, Gerstein M

Abstract

Defining regulatory networks, linking transcription factors (TFs) to their targets, is a central problem in post-genomic biology. One might imagine one could readily determine these networks through inspection of gene expression data. However, the relationship between the expression timecourse of a transcription factor and its target is not obvious (e.g. simple correlation over the timecourse), and current analysis methods, such as hierarchical clustering, have not been very successful in deciphering them. Here we introduce an approach based on support vector machines (SVMs) to predict the targets of a transcription factor by identifying subtle relationships between their expression profiles. In particular, we used SVMs to predict the regulatory targets for 36 transcription factors in the Saccharomyces cerevisiae genome based on the microarray expression data from many different physiological conditions. We trained and tested our SVM on a data set constructed to include a significant number of both positive and negative examples, directly addressing data imbalance issues. This was non-trivial given that most of the known experimental information is only for positives. Overall, we found that 63% of our TF-target relationships were confirmed through cross-validation. We further assessed the performance of our regulatory network identifications by comparing them with the results from two recent genome-wide ChIP-chip experiments. Overall, we find the agreement between our results and these experiments is comparable to the agreement (albeit low) between the two experiments. We find that this network has a delocalized structure with respect to chromosomal positioning, with a given transcription factor having targets spread fairly uniformly across the genome. The overall network of the relationships is available on the web at http://bioinfo.mbb.yale.edu/expression/echipchip

MeSH Terms
Algorithms Artificial Intelligence Cluster Analysis Computer Simulation Databases, Protein Gene Expression Profiling/methods Gene Expression Regulation/physiology Gene Targeting/methods Genome Models, Biological Protein Binding Protein Interaction Mapping/methods Proteome/metabolism Transcription Factors/metabolism User-Computer Interface
Chemicals
Proteome Transcription Factors
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Qian Jiang
Department of Ophthalmology, Johns Hopkins Medical School, Baltimore, MD 21287, USA.
Lin Jimmy
Luscombe Nicholas M
Yu Haiyuan
Gerstein Mark
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4803
Published
2003-10-12
Pages
1917-26
Language
English
Region
England
NLM ID
9808944
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com