Cargando…

A Natural Language Processing Tool for Large-Scale Data Extraction from Echocardiography Reports

Large volumes of data are continuously generated from clinical notes and diagnostic studies catalogued in electronic health records (EHRs). Echocardiography is one of the most commonly ordered diagnostic tests in cardiology. This study sought to explore the feasibility and reliability of using natur...

Descripción completa

Detalles Bibliográficos
Autores principales: Nath, Chinmoy, Albaghdadi, Mazen S., Jonnalagadda, Siddhartha R.
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Public Library of Science 2016
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4849652/
https://www.ncbi.nlm.nih.gov/pubmed/27124000
http://dx.doi.org/10.1371/journal.pone.0153749
_version_ 1782429570736062464
author Nath, Chinmoy
Albaghdadi, Mazen S.
Jonnalagadda, Siddhartha R.
author_facet Nath, Chinmoy
Albaghdadi, Mazen S.
Jonnalagadda, Siddhartha R.
author_sort Nath, Chinmoy
collection PubMed
description Large volumes of data are continuously generated from clinical notes and diagnostic studies catalogued in electronic health records (EHRs). Echocardiography is one of the most commonly ordered diagnostic tests in cardiology. This study sought to explore the feasibility and reliability of using natural language processing (NLP) for large-scale and targeted extraction of multiple data elements from echocardiography reports. An NLP tool, EchoInfer, was developed to automatically extract data pertaining to cardiovascular structure and function from heterogeneously formatted echocardiographic data sources. EchoInfer was applied to echocardiography reports (2004 to 2013) available from 3 different on-going clinical research projects. EchoInfer analyzed 15,116 echocardiography reports from 1684 patients, and extracted 59 quantitative and 21 qualitative data elements per report. EchoInfer achieved a precision of 94.06%, a recall of 92.21%, and an F1-score of 93.12% across all 80 data elements in 50 reports. Physician review of 400 reports demonstrated that EchoInfer achieved a recall of 92–99.9% and a precision of >97% in four data elements, including three quantitative and one qualitative data element. Failure of EchoInfer to correctly identify or reject reported parameters was primarily related to non-standardized reporting of echocardiography data. EchoInfer provides a powerful and reliable NLP-based approach for the large-scale, targeted extraction of information from heterogeneous data sources. The use of EchoInfer may have implications for the clinical management and research analysis of patients undergoing echocardiographic evaluation.
format Online
Article
Text
id pubmed-4849652
institution National Center for Biotechnology Information
language English
publishDate 2016
publisher Public Library of Science
record_format MEDLINE/PubMed
spelling pubmed-48496522016-05-07 A Natural Language Processing Tool for Large-Scale Data Extraction from Echocardiography Reports Nath, Chinmoy Albaghdadi, Mazen S. Jonnalagadda, Siddhartha R. PLoS One Research Article Large volumes of data are continuously generated from clinical notes and diagnostic studies catalogued in electronic health records (EHRs). Echocardiography is one of the most commonly ordered diagnostic tests in cardiology. This study sought to explore the feasibility and reliability of using natural language processing (NLP) for large-scale and targeted extraction of multiple data elements from echocardiography reports. An NLP tool, EchoInfer, was developed to automatically extract data pertaining to cardiovascular structure and function from heterogeneously formatted echocardiographic data sources. EchoInfer was applied to echocardiography reports (2004 to 2013) available from 3 different on-going clinical research projects. EchoInfer analyzed 15,116 echocardiography reports from 1684 patients, and extracted 59 quantitative and 21 qualitative data elements per report. EchoInfer achieved a precision of 94.06%, a recall of 92.21%, and an F1-score of 93.12% across all 80 data elements in 50 reports. Physician review of 400 reports demonstrated that EchoInfer achieved a recall of 92–99.9% and a precision of >97% in four data elements, including three quantitative and one qualitative data element. Failure of EchoInfer to correctly identify or reject reported parameters was primarily related to non-standardized reporting of echocardiography data. EchoInfer provides a powerful and reliable NLP-based approach for the large-scale, targeted extraction of information from heterogeneous data sources. The use of EchoInfer may have implications for the clinical management and research analysis of patients undergoing echocardiographic evaluation. Public Library of Science 2016-04-28 /pmc/articles/PMC4849652/ /pubmed/27124000 http://dx.doi.org/10.1371/journal.pone.0153749 Text en © 2016 Nath et al http://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/) , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
spellingShingle Research Article
Nath, Chinmoy
Albaghdadi, Mazen S.
Jonnalagadda, Siddhartha R.
A Natural Language Processing Tool for Large-Scale Data Extraction from Echocardiography Reports
title A Natural Language Processing Tool for Large-Scale Data Extraction from Echocardiography Reports
title_full A Natural Language Processing Tool for Large-Scale Data Extraction from Echocardiography Reports
title_fullStr A Natural Language Processing Tool for Large-Scale Data Extraction from Echocardiography Reports
title_full_unstemmed A Natural Language Processing Tool for Large-Scale Data Extraction from Echocardiography Reports
title_short A Natural Language Processing Tool for Large-Scale Data Extraction from Echocardiography Reports
title_sort natural language processing tool for large-scale data extraction from echocardiography reports
topic Research Article
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4849652/
https://www.ncbi.nlm.nih.gov/pubmed/27124000
http://dx.doi.org/10.1371/journal.pone.0153749
work_keys_str_mv AT nathchinmoy anaturallanguageprocessingtoolforlargescaledataextractionfromechocardiographyreports
AT albaghdadimazens anaturallanguageprocessingtoolforlargescaledataextractionfromechocardiographyreports
AT jonnalagaddasiddharthar anaturallanguageprocessingtoolforlargescaledataextractionfromechocardiographyreports
AT nathchinmoy naturallanguageprocessingtoolforlargescaledataextractionfromechocardiographyreports
AT albaghdadimazens naturallanguageprocessingtoolforlargescaledataextractionfromechocardiographyreports
AT jonnalagaddasiddharthar naturallanguageprocessingtoolforlargescaledataextractionfromechocardiographyreports