Cargando…
Extraction of Traditional Chinese Medicine Entity: Design of a Novel Span-Level Named Entity Recognition Method With Distant Supervision
BACKGROUND: Traditional Chinese medicine (TCM) clinical records contain the symptoms of patients, diagnoses, and subsequent treatment of doctors. These records are important resources for research and analysis of TCM diagnosis knowledge. However, most of TCM clinical records are unstructured text. T...
Autores principales: | , , , |
---|---|
Formato: | Online Artículo Texto |
Lenguaje: | English |
Publicado: |
JMIR Publications
2021
|
Materias: | |
Acceso en línea: | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8240806/ https://www.ncbi.nlm.nih.gov/pubmed/34125076 http://dx.doi.org/10.2196/28219 |
_version_ | 1783715278748647424 |
---|---|
author | Jia, Qi Zhang, Dezheng Xu, Haifeng Xie, Yonghong |
author_facet | Jia, Qi Zhang, Dezheng Xu, Haifeng Xie, Yonghong |
author_sort | Jia, Qi |
collection | PubMed |
description | BACKGROUND: Traditional Chinese medicine (TCM) clinical records contain the symptoms of patients, diagnoses, and subsequent treatment of doctors. These records are important resources for research and analysis of TCM diagnosis knowledge. However, most of TCM clinical records are unstructured text. Therefore, a method to automatically extract medical entities from TCM clinical records is indispensable. OBJECTIVE: Training a medical entity extracting model needs a large number of annotated corpus. The cost of annotated corpus is very high and there is a lack of gold-standard data sets for supervised learning methods. Therefore, we utilized distantly supervised named entity recognition (NER) to respond to the challenge. METHODS: We propose a span-level distantly supervised NER approach to extract TCM medical entity. It utilizes the pretrained language model and a simple multilayer neural network as classifier to detect and classify entity. We also designed a negative sampling strategy for the span-level model. The strategy randomly selects negative samples in every epoch and filters the possible false-negative samples periodically. It reduces the bad influence from the false-negative samples. RESULTS: We compare our methods with other baseline methods to illustrate the effectiveness of our method on a gold-standard data set. The F1 score of our method is 77.34 and it remarkably outperforms the other baselines. CONCLUSIONS: We developed a distantly supervised NER approach to extract medical entity from TCM clinical records. We estimated our approach on a TCM clinical record data set. Our experimental results indicate that the proposed approach achieves a better performance than other baselines. |
format | Online Article Text |
id | pubmed-8240806 |
institution | National Center for Biotechnology Information |
language | English |
publishDate | 2021 |
publisher | JMIR Publications |
record_format | MEDLINE/PubMed |
spelling | pubmed-82408062021-07-09 Extraction of Traditional Chinese Medicine Entity: Design of a Novel Span-Level Named Entity Recognition Method With Distant Supervision Jia, Qi Zhang, Dezheng Xu, Haifeng Xie, Yonghong JMIR Med Inform Original Paper BACKGROUND: Traditional Chinese medicine (TCM) clinical records contain the symptoms of patients, diagnoses, and subsequent treatment of doctors. These records are important resources for research and analysis of TCM diagnosis knowledge. However, most of TCM clinical records are unstructured text. Therefore, a method to automatically extract medical entities from TCM clinical records is indispensable. OBJECTIVE: Training a medical entity extracting model needs a large number of annotated corpus. The cost of annotated corpus is very high and there is a lack of gold-standard data sets for supervised learning methods. Therefore, we utilized distantly supervised named entity recognition (NER) to respond to the challenge. METHODS: We propose a span-level distantly supervised NER approach to extract TCM medical entity. It utilizes the pretrained language model and a simple multilayer neural network as classifier to detect and classify entity. We also designed a negative sampling strategy for the span-level model. The strategy randomly selects negative samples in every epoch and filters the possible false-negative samples periodically. It reduces the bad influence from the false-negative samples. RESULTS: We compare our methods with other baseline methods to illustrate the effectiveness of our method on a gold-standard data set. The F1 score of our method is 77.34 and it remarkably outperforms the other baselines. CONCLUSIONS: We developed a distantly supervised NER approach to extract medical entity from TCM clinical records. We estimated our approach on a TCM clinical record data set. Our experimental results indicate that the proposed approach achieves a better performance than other baselines. JMIR Publications 2021-06-14 /pmc/articles/PMC8240806/ /pubmed/34125076 http://dx.doi.org/10.2196/28219 Text en ©Qi Jia, Dezheng Zhang, Haifeng Xu, Yonghong Xie. Originally published in JMIR Medical Informatics (https://medinform.jmir.org), 14.06.2021. https://creativecommons.org/licenses/by/4.0/This is an open-access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR Medical Informatics, is properly cited. The complete bibliographic information, a link to the original publication on https://medinform.jmir.org/, as well as this copyright and license information must be included. |
spellingShingle | Original Paper Jia, Qi Zhang, Dezheng Xu, Haifeng Xie, Yonghong Extraction of Traditional Chinese Medicine Entity: Design of a Novel Span-Level Named Entity Recognition Method With Distant Supervision |
title | Extraction of Traditional Chinese Medicine Entity: Design of a Novel Span-Level Named Entity Recognition Method With Distant Supervision |
title_full | Extraction of Traditional Chinese Medicine Entity: Design of a Novel Span-Level Named Entity Recognition Method With Distant Supervision |
title_fullStr | Extraction of Traditional Chinese Medicine Entity: Design of a Novel Span-Level Named Entity Recognition Method With Distant Supervision |
title_full_unstemmed | Extraction of Traditional Chinese Medicine Entity: Design of a Novel Span-Level Named Entity Recognition Method With Distant Supervision |
title_short | Extraction of Traditional Chinese Medicine Entity: Design of a Novel Span-Level Named Entity Recognition Method With Distant Supervision |
title_sort | extraction of traditional chinese medicine entity: design of a novel span-level named entity recognition method with distant supervision |
topic | Original Paper |
url | https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8240806/ https://www.ncbi.nlm.nih.gov/pubmed/34125076 http://dx.doi.org/10.2196/28219 |
work_keys_str_mv | AT jiaqi extractionoftraditionalchinesemedicineentitydesignofanovelspanlevelnamedentityrecognitionmethodwithdistantsupervision AT zhangdezheng extractionoftraditionalchinesemedicineentitydesignofanovelspanlevelnamedentityrecognitionmethodwithdistantsupervision AT xuhaifeng extractionoftraditionalchinesemedicineentitydesignofanovelspanlevelnamedentityrecognitionmethodwithdistantsupervision AT xieyonghong extractionoftraditionalchinesemedicineentitydesignofanovelspanlevelnamedentityrecognitionmethodwithdistantsupervision |