Cargando…
Predictive models for endoscopic disease activity in patients with ulcerative colitis: Practical machine learning-based modeling and interpretation
BACKGROUND: Endoscopic disease activity monitoring is important for the long-term management of patients with ulcerative colitis (UC), there is currently no widely accepted non-invasive method that can effectively predict endoscopic disease activity. We aimed to develop and validate machine learning...
Autores principales: | , , , , , , , |
---|---|
Formato: | Online Artículo Texto |
Lenguaje: | English |
Publicado: |
Frontiers Media S.A.
2022
|
Materias: | |
Acceso en línea: | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9810755/ https://www.ncbi.nlm.nih.gov/pubmed/36619650 http://dx.doi.org/10.3389/fmed.2022.1043412 |
_version_ | 1784863376037380096 |
---|---|
author | Li, Xiaojun Yan, Lamei Wang, Xuehong Ouyang, Chunhui Wang, Chunlian Chao, Jun Zhang, Jie Lian, Guanghui |
author_facet | Li, Xiaojun Yan, Lamei Wang, Xuehong Ouyang, Chunhui Wang, Chunlian Chao, Jun Zhang, Jie Lian, Guanghui |
author_sort | Li, Xiaojun |
collection | PubMed |
description | BACKGROUND: Endoscopic disease activity monitoring is important for the long-term management of patients with ulcerative colitis (UC), there is currently no widely accepted non-invasive method that can effectively predict endoscopic disease activity. We aimed to develop and validate machine learning (ML) models for predicting it, which are desired to reduce the frequency of endoscopic examinations and related costs. METHODS: The patients with a diagnosis of UC in two hospitals from January 2016 to January 2021 were enrolled in this study. Thirty nine clinical and laboratory variables were collected. All patients were divided into four groups based on MES or UCEIS scores. Logistic regression (LR) and four ML algorithms were applied to construct the prediction models. The performance of models was evaluated in terms of accuracy, sensitivity, precision, F1 score, and area under the receiver-operating characteristic curve (AUC). Then Shapley additive explanations (SHAP) was applied to determine the importance of the selected variables and interpret the ML models. RESULTS: A total of 420 patients were entered into the study. Twenty four variables showed statistical differences among the groups. After synthetic minority oversampling technique (SMOTE) oversampling and RFE variables selection, the random forests (RF) model with 23 variables in MES and the extreme gradient boosting (XGBoost) model with 21 variables in USEIS, had the greatest discriminatory ability (AUC = 0.8192 in MES and 0.8006 in UCEIS in the test set). The results obtained from SHAP showed that albumin, rectal bleeding, and CRP/ALB contributed the most to the overall model. In addition, the above three variables had a more balanced contribution to each classification under the MES than the UCEIS according to the SHAP values. CONCLUSION: This proof-of-concept study demonstrated that the ML model could serve as an effective non-invasive approach to predicting endoscopic disease activity for patients with UC. RF and XGBoost, which were first introduced into data-based endoscopic disease activity prediction, are suitable for the present prediction modeling. |
format | Online Article Text |
id | pubmed-9810755 |
institution | National Center for Biotechnology Information |
language | English |
publishDate | 2022 |
publisher | Frontiers Media S.A. |
record_format | MEDLINE/PubMed |
spelling | pubmed-98107552023-01-05 Predictive models for endoscopic disease activity in patients with ulcerative colitis: Practical machine learning-based modeling and interpretation Li, Xiaojun Yan, Lamei Wang, Xuehong Ouyang, Chunhui Wang, Chunlian Chao, Jun Zhang, Jie Lian, Guanghui Front Med (Lausanne) Medicine BACKGROUND: Endoscopic disease activity monitoring is important for the long-term management of patients with ulcerative colitis (UC), there is currently no widely accepted non-invasive method that can effectively predict endoscopic disease activity. We aimed to develop and validate machine learning (ML) models for predicting it, which are desired to reduce the frequency of endoscopic examinations and related costs. METHODS: The patients with a diagnosis of UC in two hospitals from January 2016 to January 2021 were enrolled in this study. Thirty nine clinical and laboratory variables were collected. All patients were divided into four groups based on MES or UCEIS scores. Logistic regression (LR) and four ML algorithms were applied to construct the prediction models. The performance of models was evaluated in terms of accuracy, sensitivity, precision, F1 score, and area under the receiver-operating characteristic curve (AUC). Then Shapley additive explanations (SHAP) was applied to determine the importance of the selected variables and interpret the ML models. RESULTS: A total of 420 patients were entered into the study. Twenty four variables showed statistical differences among the groups. After synthetic minority oversampling technique (SMOTE) oversampling and RFE variables selection, the random forests (RF) model with 23 variables in MES and the extreme gradient boosting (XGBoost) model with 21 variables in USEIS, had the greatest discriminatory ability (AUC = 0.8192 in MES and 0.8006 in UCEIS in the test set). The results obtained from SHAP showed that albumin, rectal bleeding, and CRP/ALB contributed the most to the overall model. In addition, the above three variables had a more balanced contribution to each classification under the MES than the UCEIS according to the SHAP values. CONCLUSION: This proof-of-concept study demonstrated that the ML model could serve as an effective non-invasive approach to predicting endoscopic disease activity for patients with UC. RF and XGBoost, which were first introduced into data-based endoscopic disease activity prediction, are suitable for the present prediction modeling. Frontiers Media S.A. 2022-12-21 /pmc/articles/PMC9810755/ /pubmed/36619650 http://dx.doi.org/10.3389/fmed.2022.1043412 Text en Copyright © 2022 Li, Yan, Wang, Ouyang, Wang, Chao, Zhang and Lian. https://creativecommons.org/licenses/by/4.0/This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms. |
spellingShingle | Medicine Li, Xiaojun Yan, Lamei Wang, Xuehong Ouyang, Chunhui Wang, Chunlian Chao, Jun Zhang, Jie Lian, Guanghui Predictive models for endoscopic disease activity in patients with ulcerative colitis: Practical machine learning-based modeling and interpretation |
title | Predictive models for endoscopic disease activity in patients with ulcerative colitis: Practical machine learning-based modeling and interpretation |
title_full | Predictive models for endoscopic disease activity in patients with ulcerative colitis: Practical machine learning-based modeling and interpretation |
title_fullStr | Predictive models for endoscopic disease activity in patients with ulcerative colitis: Practical machine learning-based modeling and interpretation |
title_full_unstemmed | Predictive models for endoscopic disease activity in patients with ulcerative colitis: Practical machine learning-based modeling and interpretation |
title_short | Predictive models for endoscopic disease activity in patients with ulcerative colitis: Practical machine learning-based modeling and interpretation |
title_sort | predictive models for endoscopic disease activity in patients with ulcerative colitis: practical machine learning-based modeling and interpretation |
topic | Medicine |
url | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9810755/ https://www.ncbi.nlm.nih.gov/pubmed/36619650 http://dx.doi.org/10.3389/fmed.2022.1043412 |
work_keys_str_mv | AT lixiaojun predictivemodelsforendoscopicdiseaseactivityinpatientswithulcerativecolitispracticalmachinelearningbasedmodelingandinterpretation AT yanlamei predictivemodelsforendoscopicdiseaseactivityinpatientswithulcerativecolitispracticalmachinelearningbasedmodelingandinterpretation AT wangxuehong predictivemodelsforendoscopicdiseaseactivityinpatientswithulcerativecolitispracticalmachinelearningbasedmodelingandinterpretation AT ouyangchunhui predictivemodelsforendoscopicdiseaseactivityinpatientswithulcerativecolitispracticalmachinelearningbasedmodelingandinterpretation AT wangchunlian predictivemodelsforendoscopicdiseaseactivityinpatientswithulcerativecolitispracticalmachinelearningbasedmodelingandinterpretation AT chaojun predictivemodelsforendoscopicdiseaseactivityinpatientswithulcerativecolitispracticalmachinelearningbasedmodelingandinterpretation AT zhangjie predictivemodelsforendoscopicdiseaseactivityinpatientswithulcerativecolitispracticalmachinelearningbasedmodelingandinterpretation AT lianguanghui predictivemodelsforendoscopicdiseaseactivityinpatientswithulcerativecolitispracticalmachinelearningbasedmodelingandinterpretation |