Cargando…

How to approach machine learning-based prediction of drug/compound–target interactions

The identification of drug/compound–target interactions (DTIs) constitutes the basis of drug discovery, for which computational predictive approaches have been developed. As a relatively new data-driven paradigm, proteochemometric (PCM) modeling utilizes both protein and compound properties as a pai...

Descripción completa

Detalles Bibliográficos
Autores principales:	Atas Guvenilir, Heval, Doğan, Tunca
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	Springer International Publishing 2023
Materias:	Research
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9901167/ https://www.ncbi.nlm.nih.gov/pubmed/36747300 http://dx.doi.org/10.1186/s13321-023-00689-w

_version_	1784882986834984960
author	Atas Guvenilir, Heval Doğan, Tunca
author_facet	Atas Guvenilir, Heval Doğan, Tunca
author_sort	Atas Guvenilir, Heval
collection	PubMed
description	The identification of drug/compound–target interactions (DTIs) constitutes the basis of drug discovery, for which computational predictive approaches have been developed. As a relatively new data-driven paradigm, proteochemometric (PCM) modeling utilizes both protein and compound properties as a pair at the input level and processes them via statistical/machine learning. The representation of input samples (i.e., proteins and their ligands) in the form of quantitative feature vectors is crucial for the extraction of interaction-related properties during the artificial learning and subsequent prediction of DTIs. Lately, the representation learning approach, in which input samples are automatically featurized via training and applying a machine/deep learning model, has been utilized in biomedical sciences. In this study, we performed a comprehensive investigation of different computational approaches/techniques for protein featurization (including both conventional approaches and the novel learned embeddings), data preparation and exploration, machine learning-based modeling, and performance evaluation with the aim of achieving better data representations and more successful learning in DTI prediction. For this, we first constructed realistic and challenging benchmark datasets on small, medium, and large scales to be used as reliable gold standards for specific DTI modeling tasks. We developed and applied a network analysis-based splitting strategy to divide datasets into structurally different training and test folds. Using these datasets together with various featurization methods, we trained and tested DTI prediction models and evaluated their performance from different angles. Our main findings can be summarized under 3 items: (i) random splitting of datasets into train and test folds leads to near-complete data memorization and produce highly over-optimistic results, as a result, should be avoided, (ii) learned protein sequence embeddings work well in DTI prediction and offer high potential, despite interaction-related properties (e.g., structures) of proteins are unused during their self-supervised model training, and (iii) during the learning process, PCM models tend to rely heavily on compound features while partially ignoring protein features, primarily due to the inherent bias in DTI data, indicating the requirement for new and unbiased datasets. We hope this study will aid researchers in designing robust and high-performing data-driven DTI prediction systems that have real-world translational value in drug discovery. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1186/s13321-023-00689-w.
format	Online Article Text
id	pubmed-9901167
institution	National Center for Biotechnology Information
language	English
publishDate	2023
publisher	Springer International Publishing
record_format	MEDLINE/PubMed
spelling	pubmed-99011672023-02-07 How to approach machine learning-based prediction of drug/compound–target interactions Atas Guvenilir, Heval Doğan, Tunca J Cheminform Research The identification of drug/compound–target interactions (DTIs) constitutes the basis of drug discovery, for which computational predictive approaches have been developed. As a relatively new data-driven paradigm, proteochemometric (PCM) modeling utilizes both protein and compound properties as a pair at the input level and processes them via statistical/machine learning. The representation of input samples (i.e., proteins and their ligands) in the form of quantitative feature vectors is crucial for the extraction of interaction-related properties during the artificial learning and subsequent prediction of DTIs. Lately, the representation learning approach, in which input samples are automatically featurized via training and applying a machine/deep learning model, has been utilized in biomedical sciences. In this study, we performed a comprehensive investigation of different computational approaches/techniques for protein featurization (including both conventional approaches and the novel learned embeddings), data preparation and exploration, machine learning-based modeling, and performance evaluation with the aim of achieving better data representations and more successful learning in DTI prediction. For this, we first constructed realistic and challenging benchmark datasets on small, medium, and large scales to be used as reliable gold standards for specific DTI modeling tasks. We developed and applied a network analysis-based splitting strategy to divide datasets into structurally different training and test folds. Using these datasets together with various featurization methods, we trained and tested DTI prediction models and evaluated their performance from different angles. Our main findings can be summarized under 3 items: (i) random splitting of datasets into train and test folds leads to near-complete data memorization and produce highly over-optimistic results, as a result, should be avoided, (ii) learned protein sequence embeddings work well in DTI prediction and offer high potential, despite interaction-related properties (e.g., structures) of proteins are unused during their self-supervised model training, and (iii) during the learning process, PCM models tend to rely heavily on compound features while partially ignoring protein features, primarily due to the inherent bias in DTI data, indicating the requirement for new and unbiased datasets. We hope this study will aid researchers in designing robust and high-performing data-driven DTI prediction systems that have real-world translational value in drug discovery. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1186/s13321-023-00689-w. Springer International Publishing 2023-02-06 /pmc/articles/PMC9901167/ /pubmed/36747300 http://dx.doi.org/10.1186/s13321-023-00689-w Text en © The Author(s) 2023 https://creativecommons.org/licenses/by/4.0/Open AccessThis article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/ (https://creativecommons.org/licenses/by/4.0/) . The Creative Commons Public Domain Dedication waiver (http://creativecommons.org/publicdomain/zero/1.0/ (https://creativecommons.org/publicdomain/zero/1.0/) ) applies to the data made available in this article, unless otherwise stated in a credit line to the data.
spellingShingle	Research Atas Guvenilir, Heval Doğan, Tunca How to approach machine learning-based prediction of drug/compound–target interactions
title	How to approach machine learning-based prediction of drug/compound–target interactions
title_full	How to approach machine learning-based prediction of drug/compound–target interactions
title_fullStr	How to approach machine learning-based prediction of drug/compound–target interactions
title_full_unstemmed	How to approach machine learning-based prediction of drug/compound–target interactions
title_short	How to approach machine learning-based prediction of drug/compound–target interactions
title_sort	how to approach machine learning-based prediction of drug/compound–target interactions
topic	Research
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9901167/ https://www.ncbi.nlm.nih.gov/pubmed/36747300 http://dx.doi.org/10.1186/s13321-023-00689-w
work_keys_str_mv	AT atasguvenilirheval howtoapproachmachinelearningbasedpredictionofdrugcompoundtargetinteractions AT dogantunca howtoapproachmachinelearningbasedpredictionofdrugcompoundtargetinteractions

How to approach machine learning-based prediction of drug/compound–target interactions

Ejemplares similares