Cargando…

A method for rapid machine learning development for data mining with doctor-in-the-loop

Classifying free-text from historical databases into research-compatible formats is a barrier for clinicians undertaking audit and research projects. The aim of this study was to (a) develop interactive active machine-learning model training methodology using readily available software that was (b)...

Descripción completa

Detalles Bibliográficos
Autores principales:	Bull, Neva J., Honan, Bridget, Spratt, Neil J., Quilty, Simon
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	Public Library of Science 2023
Materias:	Research Article
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10171605/ https://www.ncbi.nlm.nih.gov/pubmed/37163511 http://dx.doi.org/10.1371/journal.pone.0284965

_version_	1785039456308297728
author	Bull, Neva J. Honan, Bridget Spratt, Neil J. Quilty, Simon
author_facet	Bull, Neva J. Honan, Bridget Spratt, Neil J. Quilty, Simon
author_sort	Bull, Neva J.
collection	PubMed
description	Classifying free-text from historical databases into research-compatible formats is a barrier for clinicians undertaking audit and research projects. The aim of this study was to (a) develop interactive active machine-learning model training methodology using readily available software that was (b) easily adaptable to a wide range of natural language databases and allowed customised researcher-defined categories, and then (c) evaluate the accuracy and speed of this model for classifying free text from two unique and unrelated clinical notes into coded data. A user interface for medical experts to train and evaluate the algorithm was created. Data requiring coding in the form of two independent databases of free-text clinical notes, each of unique natural language structure. Medical experts defined categories relevant to research projects and performed ‘label-train-evaluate’ loops on the training data set. A separate dataset was used for validation, with the medical experts blinded to the label given by the algorithm. The first dataset was 32,034 death certificate records from Northern Territory Births Deaths and Marriages, which were coded into 3 categories: haemorrhagic stroke, ischaemic stroke or no stroke. The second dataset was 12,039 recorded episodes of aeromedical retrieval from two prehospital and retrieval services in Northern Territory, Australia, which were coded into 5 categories: medical, surgical, trauma, obstetric or psychiatric. For the first dataset, macro-accuracy of the algorithm was 94.7%. For the second dataset, macro-accuracy was 92.4%. The time taken to develop and train the algorithm was 124 minutes for the death certificate coding, and 144 minutes for the aeromedical retrieval coding. This machine-learning training method was able to classify free-text clinical notes quickly and accurately from two different health datasets into categories of relevance to clinicians undertaking health service research.
format	Online Article Text
id	pubmed-10171605
institution	National Center for Biotechnology Information
language	English
publishDate	2023
publisher	Public Library of Science
record_format	MEDLINE/PubMed
spelling	pubmed-101716052023-05-11 A method for rapid machine learning development for data mining with doctor-in-the-loop Bull, Neva J. Honan, Bridget Spratt, Neil J. Quilty, Simon PLoS One Research Article Classifying free-text from historical databases into research-compatible formats is a barrier for clinicians undertaking audit and research projects. The aim of this study was to (a) develop interactive active machine-learning model training methodology using readily available software that was (b) easily adaptable to a wide range of natural language databases and allowed customised researcher-defined categories, and then (c) evaluate the accuracy and speed of this model for classifying free text from two unique and unrelated clinical notes into coded data. A user interface for medical experts to train and evaluate the algorithm was created. Data requiring coding in the form of two independent databases of free-text clinical notes, each of unique natural language structure. Medical experts defined categories relevant to research projects and performed ‘label-train-evaluate’ loops on the training data set. A separate dataset was used for validation, with the medical experts blinded to the label given by the algorithm. The first dataset was 32,034 death certificate records from Northern Territory Births Deaths and Marriages, which were coded into 3 categories: haemorrhagic stroke, ischaemic stroke or no stroke. The second dataset was 12,039 recorded episodes of aeromedical retrieval from two prehospital and retrieval services in Northern Territory, Australia, which were coded into 5 categories: medical, surgical, trauma, obstetric or psychiatric. For the first dataset, macro-accuracy of the algorithm was 94.7%. For the second dataset, macro-accuracy was 92.4%. The time taken to develop and train the algorithm was 124 minutes for the death certificate coding, and 144 minutes for the aeromedical retrieval coding. This machine-learning training method was able to classify free-text clinical notes quickly and accurately from two different health datasets into categories of relevance to clinicians undertaking health service research. Public Library of Science 2023-05-10 /pmc/articles/PMC10171605/ /pubmed/37163511 http://dx.doi.org/10.1371/journal.pone.0284965 Text en © 2023 Bull et al https://creativecommons.org/licenses/by/4.0/This is an open access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/) , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
spellingShingle	Research Article Bull, Neva J. Honan, Bridget Spratt, Neil J. Quilty, Simon A method for rapid machine learning development for data mining with doctor-in-the-loop
title	A method for rapid machine learning development for data mining with doctor-in-the-loop
title_full	A method for rapid machine learning development for data mining with doctor-in-the-loop
title_fullStr	A method for rapid machine learning development for data mining with doctor-in-the-loop
title_full_unstemmed	A method for rapid machine learning development for data mining with doctor-in-the-loop
title_short	A method for rapid machine learning development for data mining with doctor-in-the-loop
title_sort	method for rapid machine learning development for data mining with doctor-in-the-loop
topic	Research Article
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10171605/ https://www.ncbi.nlm.nih.gov/pubmed/37163511 http://dx.doi.org/10.1371/journal.pone.0284965
work_keys_str_mv	AT bullnevaj amethodforrapidmachinelearningdevelopmentfordataminingwithdoctorintheloop AT honanbridget amethodforrapidmachinelearningdevelopmentfordataminingwithdoctorintheloop AT sprattneilj amethodforrapidmachinelearningdevelopmentfordataminingwithdoctorintheloop AT quiltysimon amethodforrapidmachinelearningdevelopmentfordataminingwithdoctorintheloop AT bullnevaj methodforrapidmachinelearningdevelopmentfordataminingwithdoctorintheloop AT honanbridget methodforrapidmachinelearningdevelopmentfordataminingwithdoctorintheloop AT sprattneilj methodforrapidmachinelearningdevelopmentfordataminingwithdoctorintheloop AT quiltysimon methodforrapidmachinelearningdevelopmentfordataminingwithdoctorintheloop

A method for rapid machine learning development for data mining with doctor-in-the-loop

Ejemplares similares