Cargando…

A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery

MOTIVATION: Artificial intelligence, trained via machine learning (e.g. neural nets, random forests) or computational statistical algorithms (e.g. support vector machines, ridge regression), holds much promise for the improvement of small-molecule drug discovery. However, small-molecule structure-ac...

Descripción completa

Detalles Bibliográficos
Autores principales:	Watson, Oliver P, Cortes-Ciriano, Isidro, Taylor, Aimee R, Watson, James A
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	Oxford University Press 2019
Materias:	Original Papers
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6853675/ https://www.ncbi.nlm.nih.gov/pubmed/31070704 http://dx.doi.org/10.1093/bioinformatics/btz293

_version_	1783470078697668608
author	Watson, Oliver P Cortes-Ciriano, Isidro Taylor, Aimee R Watson, James A
author_facet	Watson, Oliver P Cortes-Ciriano, Isidro Taylor, Aimee R Watson, James A
author_sort	Watson, Oliver P
collection	PubMed
description	MOTIVATION: Artificial intelligence, trained via machine learning (e.g. neural nets, random forests) or computational statistical algorithms (e.g. support vector machines, ridge regression), holds much promise for the improvement of small-molecule drug discovery. However, small-molecule structure-activity data are high dimensional with low signal-to-noise ratios and proper validation of predictive methods is difficult. It is poorly understood which, if any, of the currently available machine learning algorithms will best predict new candidate drugs. RESULTS: The quantile-activity bootstrap is proposed as a new model validation framework using quantile splits on the activity distribution function to construct training and testing sets. In addition, we propose two novel rank-based loss functions which penalize only the out-of-sample predicted ranks of high-activity molecules. The combination of these methods was used to assess the performance of neural nets, random forests, support vector machines (regression) and ridge regression applied to 25 diverse high-quality structure-activity datasets publicly available on ChEMBL. Model validation based on random partitioning of available data favours models that overfit and ‘memorize’ the training set, namely random forests and deep neural nets. Partitioning based on quantiles of the activity distribution correctly penalizes extrapolation of models onto structurally different molecules outside of the training data. Simpler, traditional statistical methods such as ridge regression can outperform state-of-the-art machine learning methods in this setting. In addition, our new rank-based loss functions give considerably different results from mean squared error highlighting the necessity to define model optimality with respect to the decision task at hand. AVAILABILITY AND IMPLEMENTATION: All software and data are available as Jupyter notebooks found at https://github.com/owatson/QuantileBootstrap. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
format	Online Article Text
id	pubmed-6853675
institution	National Center for Biotechnology Information
language	English
publishDate	2019
publisher	Oxford University Press
record_format	MEDLINE/PubMed
spelling	pubmed-68536752019-11-19 A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery Watson, Oliver P Cortes-Ciriano, Isidro Taylor, Aimee R Watson, James A Bioinformatics Original Papers MOTIVATION: Artificial intelligence, trained via machine learning (e.g. neural nets, random forests) or computational statistical algorithms (e.g. support vector machines, ridge regression), holds much promise for the improvement of small-molecule drug discovery. However, small-molecule structure-activity data are high dimensional with low signal-to-noise ratios and proper validation of predictive methods is difficult. It is poorly understood which, if any, of the currently available machine learning algorithms will best predict new candidate drugs. RESULTS: The quantile-activity bootstrap is proposed as a new model validation framework using quantile splits on the activity distribution function to construct training and testing sets. In addition, we propose two novel rank-based loss functions which penalize only the out-of-sample predicted ranks of high-activity molecules. The combination of these methods was used to assess the performance of neural nets, random forests, support vector machines (regression) and ridge regression applied to 25 diverse high-quality structure-activity datasets publicly available on ChEMBL. Model validation based on random partitioning of available data favours models that overfit and ‘memorize’ the training set, namely random forests and deep neural nets. Partitioning based on quantiles of the activity distribution correctly penalizes extrapolation of models onto structurally different molecules outside of the training data. Simpler, traditional statistical methods such as ridge regression can outperform state-of-the-art machine learning methods in this setting. In addition, our new rank-based loss functions give considerably different results from mean squared error highlighting the necessity to define model optimality with respect to the decision task at hand. AVAILABILITY AND IMPLEMENTATION: All software and data are available as Jupyter notebooks found at https://github.com/owatson/QuantileBootstrap. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Oxford University Press 2019-11-15 2019-05-09 /pmc/articles/PMC6853675/ /pubmed/31070704 http://dx.doi.org/10.1093/bioinformatics/btz293 Text en © The Author(s) 2019. Published by Oxford University Press. http://creativecommons.org/licenses/by/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
spellingShingle	Original Papers Watson, Oliver P Cortes-Ciriano, Isidro Taylor, Aimee R Watson, James A A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery
title	A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery
title_full	A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery
title_fullStr	A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery
title_full_unstemmed	A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery
title_short	A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery
title_sort	decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery
topic	Original Papers
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6853675/ https://www.ncbi.nlm.nih.gov/pubmed/31070704 http://dx.doi.org/10.1093/bioinformatics/btz293
work_keys_str_mv	AT watsonoliverp adecisiontheoreticapproachtotheevaluationofmachinelearningalgorithmsincomputationaldrugdiscovery AT cortescirianoisidro adecisiontheoreticapproachtotheevaluationofmachinelearningalgorithmsincomputationaldrugdiscovery AT tayloraimeer adecisiontheoreticapproachtotheevaluationofmachinelearningalgorithmsincomputationaldrugdiscovery AT watsonjamesa adecisiontheoreticapproachtotheevaluationofmachinelearningalgorithmsincomputationaldrugdiscovery AT watsonoliverp decisiontheoreticapproachtotheevaluationofmachinelearningalgorithmsincomputationaldrugdiscovery AT cortescirianoisidro decisiontheoreticapproachtotheevaluationofmachinelearningalgorithmsincomputationaldrugdiscovery AT tayloraimeer decisiontheoreticapproachtotheevaluationofmachinelearningalgorithmsincomputationaldrugdiscovery AT watsonjamesa decisiontheoreticapproachtotheevaluationofmachinelearningalgorithmsincomputationaldrugdiscovery

A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery

Ejemplares similares