Cargando…

A large-scale evaluation of NLP-derived chemical-gene/protein relationships from the scientific literature: Implications for knowledge graph construction

One area of active research is the use of natural language processing (NLP) to mine biomedical texts for sets of triples (subject-predicate-object) for knowledge graph (KG) construction. While statistical methods to mine co-occurrences of entities within sentences are relatively robust, accurate rel...

Descripción completa

Detalles Bibliográficos
Autores principales:	Jeynes, Jonathan C. G., Corney, Matthew, James, Tim
Formato:	Online Artículo Texto
Lenguaje:	English
Publicado:	Public Library of Science 2023
Materias:	Research Article
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10490933/ https://www.ncbi.nlm.nih.gov/pubmed/37682956 http://dx.doi.org/10.1371/journal.pone.0291142

_version_	1785103956473544704
author	Jeynes, Jonathan C. G. Corney, Matthew James, Tim
author_facet	Jeynes, Jonathan C. G. Corney, Matthew James, Tim
author_sort	Jeynes, Jonathan C. G.
collection	PubMed
description	One area of active research is the use of natural language processing (NLP) to mine biomedical texts for sets of triples (subject-predicate-object) for knowledge graph (KG) construction. While statistical methods to mine co-occurrences of entities within sentences are relatively robust, accurate relationship extraction is more challenging. Herein, we evaluate the Global Network of Biomedical Relationships (GNBR), a dataset that uses distributional semantics to model relationships between biomedical entities. The focus of our paper is an evaluation of a subset of the GNBR data; the relationships between chemicals and genes/proteins. We use Evotec’s structured ‘Nexus’ database of >2.76M chemical-protein interactions as a ground truth to compare with GNBRs relationships and find a micro-averaged precision-recall area under the curve (AUC) of 0.50 and a micro-averaged receiver operating characteristic (ROC) curve AUC of 0.71 across the relationship classes ‘inhibits’, ‘binding’, ‘agonism’ and ‘antagonism’, when a comparison is made on a sentence-by-sentence basis. We conclude that, even though these micro-average scores are modest, using a high threshold on certain relationship classes like ‘inhibits’ could yield high fidelity triples that are not reported in structured datasets. We discuss how different methods of processing GNBR data, and the factuality of triples could affect the accuracy of NLP data incorporated into knowledge graphs. We provide a GNBR-Nexus(ChEMBL-subset) merged datafile that contains over 20,000 sentences where a protein/gene-chemical co-occur and includes both the GNBR relationship scores as well as the ChEMBL (manually curated) relationships (e.g., ‘agonist’, ‘inhibitor’) —this can be accessed at https://doi.org/10.5281/zenodo.8136752. We envisage this being used to aid curation efforts by the drug discovery community.
format	Online Article Text
id	pubmed-10490933
institution	National Center for Biotechnology Information
language	English
publishDate	2023
publisher	Public Library of Science
record_format	MEDLINE/PubMed
spelling	pubmed-104909332023-09-09 A large-scale evaluation of NLP-derived chemical-gene/protein relationships from the scientific literature: Implications for knowledge graph construction Jeynes, Jonathan C. G. Corney, Matthew James, Tim PLoS One Research Article One area of active research is the use of natural language processing (NLP) to mine biomedical texts for sets of triples (subject-predicate-object) for knowledge graph (KG) construction. While statistical methods to mine co-occurrences of entities within sentences are relatively robust, accurate relationship extraction is more challenging. Herein, we evaluate the Global Network of Biomedical Relationships (GNBR), a dataset that uses distributional semantics to model relationships between biomedical entities. The focus of our paper is an evaluation of a subset of the GNBR data; the relationships between chemicals and genes/proteins. We use Evotec’s structured ‘Nexus’ database of >2.76M chemical-protein interactions as a ground truth to compare with GNBRs relationships and find a micro-averaged precision-recall area under the curve (AUC) of 0.50 and a micro-averaged receiver operating characteristic (ROC) curve AUC of 0.71 across the relationship classes ‘inhibits’, ‘binding’, ‘agonism’ and ‘antagonism’, when a comparison is made on a sentence-by-sentence basis. We conclude that, even though these micro-average scores are modest, using a high threshold on certain relationship classes like ‘inhibits’ could yield high fidelity triples that are not reported in structured datasets. We discuss how different methods of processing GNBR data, and the factuality of triples could affect the accuracy of NLP data incorporated into knowledge graphs. We provide a GNBR-Nexus(ChEMBL-subset) merged datafile that contains over 20,000 sentences where a protein/gene-chemical co-occur and includes both the GNBR relationship scores as well as the ChEMBL (manually curated) relationships (e.g., ‘agonist’, ‘inhibitor’) —this can be accessed at https://doi.org/10.5281/zenodo.8136752. We envisage this being used to aid curation efforts by the drug discovery community. Public Library of Science 2023-09-08 /pmc/articles/PMC10490933/ /pubmed/37682956 http://dx.doi.org/10.1371/journal.pone.0291142 Text en © 2023 Jeynes et al https://creativecommons.org/licenses/by/4.0/This is an open access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/) , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
spellingShingle	Research Article Jeynes, Jonathan C. G. Corney, Matthew James, Tim A large-scale evaluation of NLP-derived chemical-gene/protein relationships from the scientific literature: Implications for knowledge graph construction
title	A large-scale evaluation of NLP-derived chemical-gene/protein relationships from the scientific literature: Implications for knowledge graph construction
title_full	A large-scale evaluation of NLP-derived chemical-gene/protein relationships from the scientific literature: Implications for knowledge graph construction
title_fullStr	A large-scale evaluation of NLP-derived chemical-gene/protein relationships from the scientific literature: Implications for knowledge graph construction
title_full_unstemmed	A large-scale evaluation of NLP-derived chemical-gene/protein relationships from the scientific literature: Implications for knowledge graph construction
title_short	A large-scale evaluation of NLP-derived chemical-gene/protein relationships from the scientific literature: Implications for knowledge graph construction
title_sort	large-scale evaluation of nlp-derived chemical-gene/protein relationships from the scientific literature: implications for knowledge graph construction
topic	Research Article
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10490933/ https://www.ncbi.nlm.nih.gov/pubmed/37682956 http://dx.doi.org/10.1371/journal.pone.0291142
work_keys_str_mv	AT jeynesjonathancg alargescaleevaluationofnlpderivedchemicalgeneproteinrelationshipsfromthescientificliteratureimplicationsforknowledgegraphconstruction AT corneymatthew alargescaleevaluationofnlpderivedchemicalgeneproteinrelationshipsfromthescientificliteratureimplicationsforknowledgegraphconstruction AT jamestim alargescaleevaluationofnlpderivedchemicalgeneproteinrelationshipsfromthescientificliteratureimplicationsforknowledgegraphconstruction AT jeynesjonathancg largescaleevaluationofnlpderivedchemicalgeneproteinrelationshipsfromthescientificliteratureimplicationsforknowledgegraphconstruction AT corneymatthew largescaleevaluationofnlpderivedchemicalgeneproteinrelationshipsfromthescientificliteratureimplicationsforknowledgegraphconstruction AT jamestim largescaleevaluationofnlpderivedchemicalgeneproteinrelationshipsfromthescientificliteratureimplicationsforknowledgegraphconstruction

A large-scale evaluation of NLP-derived chemical-gene/protein relationships from the scientific literature: Implications for knowledge graph construction

Ejemplares similares