Cargando…

Information Extraction from the Text Data on Traditional Chinese Medicine: A Review on Tasks, Challenges, and Methods from 2010 to 2021

BACKGROUND: The practice of traditional Chinese medicine (TCM) began several thousand years ago, and the knowledge of practitioners is recorded in paper and electronic versions of case notes, manuscripts, and books in multiple languages. Developing a method of information extraction (IE) from these...

Descripción completa

Detalles Bibliográficos
Autores principales: Zhang, Tingting, Huang, Zonghai, Wang, Yaqiang, Wen, Chuanbiao, Peng, Yangzhi, Ye, Ying
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Hindawi 2022
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9122692/
https://www.ncbi.nlm.nih.gov/pubmed/35600940
http://dx.doi.org/10.1155/2022/1679589
_version_ 1784711398601785344
author Zhang, Tingting
Huang, Zonghai
Wang, Yaqiang
Wen, Chuanbiao
Peng, Yangzhi
Ye, Ying
author_facet Zhang, Tingting
Huang, Zonghai
Wang, Yaqiang
Wen, Chuanbiao
Peng, Yangzhi
Ye, Ying
author_sort Zhang, Tingting
collection PubMed
description BACKGROUND: The practice of traditional Chinese medicine (TCM) began several thousand years ago, and the knowledge of practitioners is recorded in paper and electronic versions of case notes, manuscripts, and books in multiple languages. Developing a method of information extraction (IE) from these sources to generate a cohesive data set would be a great contribution to the medical field. The goal of this study was to perform a systematic review of the status of IE from TCM sources over the last 10 years. METHODS: We conducted a search of four literature databases for articles published from 2010 to 2021 that focused on the use of natural language processing (NLP) methods to extract information from unstructured TCM text data. Two reviewers and one adjudicator contributed to article search, article selection, data extraction, and synthesis processes. RESULTS: We retrieved 1234 records, 49 of which met our inclusion criteria. We used the articles to (i) assess the key tasks of IE in the TCM domain, (ii) summarize the challenges to extracting information from TCM text data, and (iii) identify effective frameworks, models, and key findings of TCM IE through classification. CONCLUSIONS: Our analysis showed that IE from TCM text data has improved over the past decade. However, the extraction of TCM text still faces some challenges involving the lack of gold standard corpora, nonstandardized expressions, and multiple types of relations. In the future, IE work should be promoted by extracting more existing entities and relations, constructing gold standard data sets, and exploring IE methods based on a small amount of labeled data. Furthermore, fine-grained and interpretable IE technologies are necessary for further exploration.
format Online
Article
Text
id pubmed-9122692
institution National Center for Biotechnology Information
language English
publishDate 2022
publisher Hindawi
record_format MEDLINE/PubMed
spelling pubmed-91226922022-05-21 Information Extraction from the Text Data on Traditional Chinese Medicine: A Review on Tasks, Challenges, and Methods from 2010 to 2021 Zhang, Tingting Huang, Zonghai Wang, Yaqiang Wen, Chuanbiao Peng, Yangzhi Ye, Ying Evid Based Complement Alternat Med Review Article BACKGROUND: The practice of traditional Chinese medicine (TCM) began several thousand years ago, and the knowledge of practitioners is recorded in paper and electronic versions of case notes, manuscripts, and books in multiple languages. Developing a method of information extraction (IE) from these sources to generate a cohesive data set would be a great contribution to the medical field. The goal of this study was to perform a systematic review of the status of IE from TCM sources over the last 10 years. METHODS: We conducted a search of four literature databases for articles published from 2010 to 2021 that focused on the use of natural language processing (NLP) methods to extract information from unstructured TCM text data. Two reviewers and one adjudicator contributed to article search, article selection, data extraction, and synthesis processes. RESULTS: We retrieved 1234 records, 49 of which met our inclusion criteria. We used the articles to (i) assess the key tasks of IE in the TCM domain, (ii) summarize the challenges to extracting information from TCM text data, and (iii) identify effective frameworks, models, and key findings of TCM IE through classification. CONCLUSIONS: Our analysis showed that IE from TCM text data has improved over the past decade. However, the extraction of TCM text still faces some challenges involving the lack of gold standard corpora, nonstandardized expressions, and multiple types of relations. In the future, IE work should be promoted by extracting more existing entities and relations, constructing gold standard data sets, and exploring IE methods based on a small amount of labeled data. Furthermore, fine-grained and interpretable IE technologies are necessary for further exploration. Hindawi 2022-05-13 /pmc/articles/PMC9122692/ /pubmed/35600940 http://dx.doi.org/10.1155/2022/1679589 Text en Copyright © 2022 Tingting Zhang et al. https://creativecommons.org/licenses/by/4.0/This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
spellingShingle Review Article
Zhang, Tingting
Huang, Zonghai
Wang, Yaqiang
Wen, Chuanbiao
Peng, Yangzhi
Ye, Ying
Information Extraction from the Text Data on Traditional Chinese Medicine: A Review on Tasks, Challenges, and Methods from 2010 to 2021
title Information Extraction from the Text Data on Traditional Chinese Medicine: A Review on Tasks, Challenges, and Methods from 2010 to 2021
title_full Information Extraction from the Text Data on Traditional Chinese Medicine: A Review on Tasks, Challenges, and Methods from 2010 to 2021
title_fullStr Information Extraction from the Text Data on Traditional Chinese Medicine: A Review on Tasks, Challenges, and Methods from 2010 to 2021
title_full_unstemmed Information Extraction from the Text Data on Traditional Chinese Medicine: A Review on Tasks, Challenges, and Methods from 2010 to 2021
title_short Information Extraction from the Text Data on Traditional Chinese Medicine: A Review on Tasks, Challenges, and Methods from 2010 to 2021
title_sort information extraction from the text data on traditional chinese medicine: a review on tasks, challenges, and methods from 2010 to 2021
topic Review Article
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9122692/
https://www.ncbi.nlm.nih.gov/pubmed/35600940
http://dx.doi.org/10.1155/2022/1679589
work_keys_str_mv AT zhangtingting informationextractionfromthetextdataontraditionalchinesemedicineareviewontaskschallengesandmethodsfrom2010to2021
AT huangzonghai informationextractionfromthetextdataontraditionalchinesemedicineareviewontaskschallengesandmethodsfrom2010to2021
AT wangyaqiang informationextractionfromthetextdataontraditionalchinesemedicineareviewontaskschallengesandmethodsfrom2010to2021
AT wenchuanbiao informationextractionfromthetextdataontraditionalchinesemedicineareviewontaskschallengesandmethodsfrom2010to2021
AT pengyangzhi informationextractionfromthetextdataontraditionalchinesemedicineareviewontaskschallengesandmethodsfrom2010to2021
AT yeying informationextractionfromthetextdataontraditionalchinesemedicineareviewontaskschallengesandmethodsfrom2010to2021