Cargando…
L(2)-norm multiple kernel learning and its application to biomedical data fusion
BACKGROUND: This paper introduces the notion of optimizing different norms in the dual problem of support vector machines with multiple kernels. The selection of norms yields different extensions of multiple kernel learning (MKL) such as L(∞), L(1), and L(2 )MKL. In particular, L(2 )MKL is a novel m...
Autores principales: | , , , , , , |
---|---|
Formato: | Texto |
Lenguaje: | English |
Publicado: |
BioMed Central
2010
|
Materias: | |
Acceso en línea: | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2906488/ https://www.ncbi.nlm.nih.gov/pubmed/20529363 http://dx.doi.org/10.1186/1471-2105-11-309 |
_version_ | 1782184036531175424 |
---|---|
author | Yu, Shi Falck, Tillmann Daemen, Anneleen Tranchevent, Leon-Charles Suykens, Johan AK De Moor, Bart Moreau, Yves |
author_facet | Yu, Shi Falck, Tillmann Daemen, Anneleen Tranchevent, Leon-Charles Suykens, Johan AK De Moor, Bart Moreau, Yves |
author_sort | Yu, Shi |
collection | PubMed |
description | BACKGROUND: This paper introduces the notion of optimizing different norms in the dual problem of support vector machines with multiple kernels. The selection of norms yields different extensions of multiple kernel learning (MKL) such as L(∞), L(1), and L(2 )MKL. In particular, L(2 )MKL is a novel method that leads to non-sparse optimal kernel coefficients, which is different from the sparse kernel coefficients optimized by the existing L(∞ )MKL method. In real biomedical applications, L(2 )MKL may have more advantages over sparse integration method for thoroughly combining complementary information in heterogeneous data sources. RESULTS: We provide a theoretical analysis of the relationship between the L(2 )optimization of kernels in the dual problem with the L(2 )coefficient regularization in the primal problem. Understanding the dual L(2 )problem grants a unified view on MKL and enables us to extend the L(2 )method to a wide range of machine learning problems. We implement L(2 )MKL for ranking and classification problems and compare its performance with the sparse L(∞ )and the averaging L(1 )MKL methods. The experiments are carried out on six real biomedical data sets and two large scale UCI data sets. L(2 )MKL yields better performance on most of the benchmark data sets. In particular, we propose a novel L(2 )MKL least squares support vector machine (LSSVM) algorithm, which is shown to be an efficient and promising classifier for large scale data sets processing. CONCLUSIONS: This paper extends the statistical framework of genomic data fusion based on MKL. Allowing non-sparse weights on the data sources is an attractive option in settings where we believe most data sources to be relevant to the problem at hand and want to avoid a "winner-takes-all" effect seen in L(∞ )MKL, which can be detrimental to the performance in prospective studies. The notion of optimizing L(2 )kernels can be straightforwardly extended to ranking, classification, regression, and clustering algorithms. To tackle the computational burden of MKL, this paper proposes several novel LSSVM based MKL algorithms. Systematic comparison on real data sets shows that LSSVM MKL has comparable performance as the conventional SVM MKL algorithms. Moreover, large scale numerical experiments indicate that when cast as semi-infinite programming, LSSVM MKL can be solved more efficiently than SVM MKL. AVAILABILITY: The MATLAB code of algorithms implemented in this paper is downloadable from http://homes.esat.kuleuven.be/~sistawww/bioi/syu/l2lssvm.html. |
format | Text |
id | pubmed-2906488 |
institution | National Center for Biotechnology Information |
language | English |
publishDate | 2010 |
publisher | BioMed Central |
record_format | MEDLINE/PubMed |
spelling | pubmed-29064882010-07-20 L(2)-norm multiple kernel learning and its application to biomedical data fusion Yu, Shi Falck, Tillmann Daemen, Anneleen Tranchevent, Leon-Charles Suykens, Johan AK De Moor, Bart Moreau, Yves BMC Bioinformatics Methodology Article BACKGROUND: This paper introduces the notion of optimizing different norms in the dual problem of support vector machines with multiple kernels. The selection of norms yields different extensions of multiple kernel learning (MKL) such as L(∞), L(1), and L(2 )MKL. In particular, L(2 )MKL is a novel method that leads to non-sparse optimal kernel coefficients, which is different from the sparse kernel coefficients optimized by the existing L(∞ )MKL method. In real biomedical applications, L(2 )MKL may have more advantages over sparse integration method for thoroughly combining complementary information in heterogeneous data sources. RESULTS: We provide a theoretical analysis of the relationship between the L(2 )optimization of kernels in the dual problem with the L(2 )coefficient regularization in the primal problem. Understanding the dual L(2 )problem grants a unified view on MKL and enables us to extend the L(2 )method to a wide range of machine learning problems. We implement L(2 )MKL for ranking and classification problems and compare its performance with the sparse L(∞ )and the averaging L(1 )MKL methods. The experiments are carried out on six real biomedical data sets and two large scale UCI data sets. L(2 )MKL yields better performance on most of the benchmark data sets. In particular, we propose a novel L(2 )MKL least squares support vector machine (LSSVM) algorithm, which is shown to be an efficient and promising classifier for large scale data sets processing. CONCLUSIONS: This paper extends the statistical framework of genomic data fusion based on MKL. Allowing non-sparse weights on the data sources is an attractive option in settings where we believe most data sources to be relevant to the problem at hand and want to avoid a "winner-takes-all" effect seen in L(∞ )MKL, which can be detrimental to the performance in prospective studies. The notion of optimizing L(2 )kernels can be straightforwardly extended to ranking, classification, regression, and clustering algorithms. To tackle the computational burden of MKL, this paper proposes several novel LSSVM based MKL algorithms. Systematic comparison on real data sets shows that LSSVM MKL has comparable performance as the conventional SVM MKL algorithms. Moreover, large scale numerical experiments indicate that when cast as semi-infinite programming, LSSVM MKL can be solved more efficiently than SVM MKL. AVAILABILITY: The MATLAB code of algorithms implemented in this paper is downloadable from http://homes.esat.kuleuven.be/~sistawww/bioi/syu/l2lssvm.html. BioMed Central 2010-06-08 /pmc/articles/PMC2906488/ /pubmed/20529363 http://dx.doi.org/10.1186/1471-2105-11-309 Text en Copyright ©2010 Yu et al; licensee BioMed Central Ltd. http://creativecommons.org/licenses/by/2.0 This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. |
spellingShingle | Methodology Article Yu, Shi Falck, Tillmann Daemen, Anneleen Tranchevent, Leon-Charles Suykens, Johan AK De Moor, Bart Moreau, Yves L(2)-norm multiple kernel learning and its application to biomedical data fusion |
title | L(2)-norm multiple kernel learning and its application to biomedical data fusion |
title_full | L(2)-norm multiple kernel learning and its application to biomedical data fusion |
title_fullStr | L(2)-norm multiple kernel learning and its application to biomedical data fusion |
title_full_unstemmed | L(2)-norm multiple kernel learning and its application to biomedical data fusion |
title_short | L(2)-norm multiple kernel learning and its application to biomedical data fusion |
title_sort | l(2)-norm multiple kernel learning and its application to biomedical data fusion |
topic | Methodology Article |
url | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2906488/ https://www.ncbi.nlm.nih.gov/pubmed/20529363 http://dx.doi.org/10.1186/1471-2105-11-309 |
work_keys_str_mv | AT yushi l2normmultiplekernellearninganditsapplicationtobiomedicaldatafusion AT falcktillmann l2normmultiplekernellearninganditsapplicationtobiomedicaldatafusion AT daemenanneleen l2normmultiplekernellearninganditsapplicationtobiomedicaldatafusion AT trancheventleoncharles l2normmultiplekernellearninganditsapplicationtobiomedicaldatafusion AT suykensjohanak l2normmultiplekernellearninganditsapplicationtobiomedicaldatafusion AT demoorbart l2normmultiplekernellearninganditsapplicationtobiomedicaldatafusion AT moreauyves l2normmultiplekernellearninganditsapplicationtobiomedicaldatafusion |