Home LiteratureArticle Details
PMID: 16966363 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

Support vector machine learning from heterogeneous data: an empirical analysis using protein sequence and structure.

Bioinformatics (Oxford, England) ·Vol. 22 ·No. 22 ·2006-11-15 ·Pages 2753-60

Lewis DP, Jebara T, Noble WS

Abstract

Drawing inferences from large, heterogeneous sets of biological data requires a theoretical framework that is capable of representing, e.g. DNA and protein sequences, protein structures, microarray expression data, various types of interaction networks, etc. Recently, a class of algorithms known as kernel methods has emerged as a powerful framework for combining diverse types of data. The support vector machine (SVM) algorithm is the most popular kernel method, due to its theoretical underpinnings and strong empirical performance on a wide variety of classification tasks. Furthermore, several recently described extensions allow the SVM to assign relative weights to various datasets, depending upon their utilities in performing a given classification task. In this work, we empirically investigate the performance of the SVM on the task of inferring gene functional annotations from a combination of protein sequence and structure data. Our results suggest that the SVM is quite robust to noise in the input datasets. Consequently, in the presence of only two types of data, an SVM trained from an unweighted combination of datasets performs as well or better than a more sophisticated algorithm that assigns weights to individual data types. Indeed, for this simple case, we can demonstrate empirically that no solution is significantly better than the naive, unweighted average of the two datasets. On the other hand, when multiple noisy datasets are included in the experiment, then the naive approach fares worse than the weighted approach. Our results suggest that for many applications, a naive unweighted sum of kernels may be sufficient. http://noble.gs.washington.edu/proj/seqstruct

MeSH Terms
Algorithms Computational Biology/methods Databases, Protein Fungal Proteins/chemistry Models, Statistical Pattern Recognition, Automated Proteins/chemistry Proteomics/methods ROC Curve Sequence Alignment Sequence Analysis, Protein Software
Chemicals
Fungal Proteins Proteins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Lewis Darrin P
Department of Computer Science, Columbia University, New York, NY, 10027.
Jebara Tony
Noble William Stafford
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2006-11-15
Epub
2006-00-11
Pages
2753-60
Language
English
Region
England
NLM ID
9808944
Subset
IM
Grants
NHGRI NIH HHS · R33 HG003070 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com