Cargando…

Integrating Structured and Unstructured EHR Data for Predicting Mortality by Machine Learning and Latent Dirichlet Allocation Method

An ICU is a critical care unit that provides advanced medical support and continuous monitoring for patients with severe illnesses or injuries. Predicting the mortality rate of ICU patients can not only improve patient outcomes, but also optimize resource allocation. Many studies have attempted to c...

Descripción completa

Detalles Bibliográficos
Autores principales: Chiu, Chih-Chou, Wu, Chung-Min, Chien, Te-Nien, Kao, Ling-Jing, Li, Chengcheng, Chu, Chuan-Mei
Formato: Online Artículo Texto
Lenguaje:English
Publicado: MDPI 2023
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10001457/
https://www.ncbi.nlm.nih.gov/pubmed/36901354
http://dx.doi.org/10.3390/ijerph20054340
_version_ 1784904141413285888
author Chiu, Chih-Chou
Wu, Chung-Min
Chien, Te-Nien
Kao, Ling-Jing
Li, Chengcheng
Chu, Chuan-Mei
author_facet Chiu, Chih-Chou
Wu, Chung-Min
Chien, Te-Nien
Kao, Ling-Jing
Li, Chengcheng
Chu, Chuan-Mei
author_sort Chiu, Chih-Chou
collection PubMed
description An ICU is a critical care unit that provides advanced medical support and continuous monitoring for patients with severe illnesses or injuries. Predicting the mortality rate of ICU patients can not only improve patient outcomes, but also optimize resource allocation. Many studies have attempted to create scoring systems and models that predict the mortality of ICU patients using large amounts of structured clinical data. However, unstructured clinical data recorded during patient admission, such as notes made by physicians, is often overlooked. This study used the MIMIC-III database to predict mortality in ICU patients. In the first part of the study, only eight structured variables were used, including the six basic vital signs, the GCS, and the patient’s age at admission. In the second part, unstructured predictor variables were extracted from the initial diagnosis made by physicians when the patients were admitted to the hospital and analyzed using Latent Dirichlet Allocation techniques. The structured and unstructured data were combined using machine learning methods to create a mortality risk prediction model for ICU patients. The results showed that combining structured and unstructured data improved the accuracy of the prediction of clinical outcomes in ICU patients over time. The model achieved an AUROC of 0.88, indicating accurate prediction of patient vital status. Additionally, the model was able to predict patient clinical outcomes over time, successfully identifying important variables. This study demonstrated that a small number of easily collectible structured variables, combined with unstructured data and analyzed using LDA topic modeling, can significantly improve the predictive performance of a mortality risk prediction model for ICU patients. These results suggest that initial clinical observations and diagnoses of ICU patients contain valuable information that can aid ICU medical and nursing staff in making important clinical decisions.
format Online
Article
Text
id pubmed-10001457
institution National Center for Biotechnology Information
language English
publishDate 2023
publisher MDPI
record_format MEDLINE/PubMed
spelling pubmed-100014572023-03-11 Integrating Structured and Unstructured EHR Data for Predicting Mortality by Machine Learning and Latent Dirichlet Allocation Method Chiu, Chih-Chou Wu, Chung-Min Chien, Te-Nien Kao, Ling-Jing Li, Chengcheng Chu, Chuan-Mei Int J Environ Res Public Health Article An ICU is a critical care unit that provides advanced medical support and continuous monitoring for patients with severe illnesses or injuries. Predicting the mortality rate of ICU patients can not only improve patient outcomes, but also optimize resource allocation. Many studies have attempted to create scoring systems and models that predict the mortality of ICU patients using large amounts of structured clinical data. However, unstructured clinical data recorded during patient admission, such as notes made by physicians, is often overlooked. This study used the MIMIC-III database to predict mortality in ICU patients. In the first part of the study, only eight structured variables were used, including the six basic vital signs, the GCS, and the patient’s age at admission. In the second part, unstructured predictor variables were extracted from the initial diagnosis made by physicians when the patients were admitted to the hospital and analyzed using Latent Dirichlet Allocation techniques. The structured and unstructured data were combined using machine learning methods to create a mortality risk prediction model for ICU patients. The results showed that combining structured and unstructured data improved the accuracy of the prediction of clinical outcomes in ICU patients over time. The model achieved an AUROC of 0.88, indicating accurate prediction of patient vital status. Additionally, the model was able to predict patient clinical outcomes over time, successfully identifying important variables. This study demonstrated that a small number of easily collectible structured variables, combined with unstructured data and analyzed using LDA topic modeling, can significantly improve the predictive performance of a mortality risk prediction model for ICU patients. These results suggest that initial clinical observations and diagnoses of ICU patients contain valuable information that can aid ICU medical and nursing staff in making important clinical decisions. MDPI 2023-02-28 /pmc/articles/PMC10001457/ /pubmed/36901354 http://dx.doi.org/10.3390/ijerph20054340 Text en © 2023 by the authors. https://creativecommons.org/licenses/by/4.0/Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
spellingShingle Article
Chiu, Chih-Chou
Wu, Chung-Min
Chien, Te-Nien
Kao, Ling-Jing
Li, Chengcheng
Chu, Chuan-Mei
Integrating Structured and Unstructured EHR Data for Predicting Mortality by Machine Learning and Latent Dirichlet Allocation Method
title Integrating Structured and Unstructured EHR Data for Predicting Mortality by Machine Learning and Latent Dirichlet Allocation Method
title_full Integrating Structured and Unstructured EHR Data for Predicting Mortality by Machine Learning and Latent Dirichlet Allocation Method
title_fullStr Integrating Structured and Unstructured EHR Data for Predicting Mortality by Machine Learning and Latent Dirichlet Allocation Method
title_full_unstemmed Integrating Structured and Unstructured EHR Data for Predicting Mortality by Machine Learning and Latent Dirichlet Allocation Method
title_short Integrating Structured and Unstructured EHR Data for Predicting Mortality by Machine Learning and Latent Dirichlet Allocation Method
title_sort integrating structured and unstructured ehr data for predicting mortality by machine learning and latent dirichlet allocation method
topic Article
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10001457/
https://www.ncbi.nlm.nih.gov/pubmed/36901354
http://dx.doi.org/10.3390/ijerph20054340
work_keys_str_mv AT chiuchihchou integratingstructuredandunstructuredehrdataforpredictingmortalitybymachinelearningandlatentdirichletallocationmethod
AT wuchungmin integratingstructuredandunstructuredehrdataforpredictingmortalitybymachinelearningandlatentdirichletallocationmethod
AT chientenien integratingstructuredandunstructuredehrdataforpredictingmortalitybymachinelearningandlatentdirichletallocationmethod
AT kaolingjing integratingstructuredandunstructuredehrdataforpredictingmortalitybymachinelearningandlatentdirichletallocationmethod
AT lichengcheng integratingstructuredandunstructuredehrdataforpredictingmortalitybymachinelearningandlatentdirichletallocationmethod
AT chuchuanmei integratingstructuredandunstructuredehrdataforpredictingmortalitybymachinelearningandlatentdirichletallocationmethod