Cargando…

Issues in learning an ontology from text

Ontology construction for any domain is a labour intensive and complex process. Any methodology that can reduce the cost and increase efficiency has the potential to make a major impact in the life sciences. This paper describes an experiment in ontology construction from text for the animal behavio...

Descripción completa

Detalles Bibliográficos
Autores principales:	Brewster, Christopher, Jupp, Simon, Luciano, Joanne, Shotton, David, Stevens, Robert D, Zhang, Ziqi
Formato:	Texto
Lenguaje:	English
Publicado:	BioMed Central 2009
Materias:	Proceedings
Acceso en línea:	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2679401/ https://www.ncbi.nlm.nih.gov/pubmed/19426458 http://dx.doi.org/10.1186/1471-2105-10-S5-S1

_version_	1782166891396071424
author	Brewster, Christopher Jupp, Simon Luciano, Joanne Shotton, David Stevens, Robert D Zhang, Ziqi
author_facet	Brewster, Christopher Jupp, Simon Luciano, Joanne Shotton, David Stevens, Robert D Zhang, Ziqi
author_sort	Brewster, Christopher
collection	PubMed
description	Ontology construction for any domain is a labour intensive and complex process. Any methodology that can reduce the cost and increase efficiency has the potential to make a major impact in the life sciences. This paper describes an experiment in ontology construction from text for the animal behaviour domain. Our objective was to see how much could be done in a simple and relatively rapid manner using a corpus of journal papers. We used a sequence of pre-existing text processing steps, and here describe the different choices made to clean the input, to derive a set of terms and to structure those terms in a number of hierarchies. We describe some of the challenges, especially that of focusing the ontology appropriately given a starting point of a heterogeneous corpus. Using mainly automated techniques, we were able to construct an 18055 term ontology-like structure with 73% recall of animal behaviour terms, but a precision of only 26%. We were able to clean unwanted terms from the nascent ontology using lexico-syntactic patterns that tested the validity of term inclusion within the ontology. We used the same technique to test for subsumption relationships between the remaining terms to add structure to the initially broad and shallow structure we generated. All outputs are available at . We present a systematic method for the initial steps of ontology or structured vocabulary construction for scientific domains that requires limited human effort and can make a contribution both to ontology learning and maintenance. The method is useful both for the exploration of a scientific domain and as a stepping stone towards formally rigourous ontologies. The filtering of recognised terms from a heterogeneous corpus to focus upon those that are the topic of the ontology is identified to be one of the main challenges for research in ontology learning.
format	Text
id	pubmed-2679401
institution	National Center for Biotechnology Information
language	English
publishDate	2009
publisher	BioMed Central
record_format	MEDLINE/PubMed
spelling	pubmed-26794012009-05-11 Issues in learning an ontology from text Brewster, Christopher Jupp, Simon Luciano, Joanne Shotton, David Stevens, Robert D Zhang, Ziqi BMC Bioinformatics Proceedings Ontology construction for any domain is a labour intensive and complex process. Any methodology that can reduce the cost and increase efficiency has the potential to make a major impact in the life sciences. This paper describes an experiment in ontology construction from text for the animal behaviour domain. Our objective was to see how much could be done in a simple and relatively rapid manner using a corpus of journal papers. We used a sequence of pre-existing text processing steps, and here describe the different choices made to clean the input, to derive a set of terms and to structure those terms in a number of hierarchies. We describe some of the challenges, especially that of focusing the ontology appropriately given a starting point of a heterogeneous corpus. Using mainly automated techniques, we were able to construct an 18055 term ontology-like structure with 73% recall of animal behaviour terms, but a precision of only 26%. We were able to clean unwanted terms from the nascent ontology using lexico-syntactic patterns that tested the validity of term inclusion within the ontology. We used the same technique to test for subsumption relationships between the remaining terms to add structure to the initially broad and shallow structure we generated. All outputs are available at . We present a systematic method for the initial steps of ontology or structured vocabulary construction for scientific domains that requires limited human effort and can make a contribution both to ontology learning and maintenance. The method is useful both for the exploration of a scientific domain and as a stepping stone towards formally rigourous ontologies. The filtering of recognised terms from a heterogeneous corpus to focus upon those that are the topic of the ontology is identified to be one of the main challenges for research in ontology learning. BioMed Central 2009-05-06 /pmc/articles/PMC2679401/ /pubmed/19426458 http://dx.doi.org/10.1186/1471-2105-10-S5-S1 Text en Copyright © 2009 Brewster et al; licensee BioMed Central Ltd. http://creativecommons.org/licenses/by/2.0 This is an open access article distributed under the terms of the Creative Commons Attribution License ( (http://creativecommons.org/licenses/by/2.0) ), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
spellingShingle	Proceedings Brewster, Christopher Jupp, Simon Luciano, Joanne Shotton, David Stevens, Robert D Zhang, Ziqi Issues in learning an ontology from text
title	Issues in learning an ontology from text
title_full	Issues in learning an ontology from text
title_fullStr	Issues in learning an ontology from text
title_full_unstemmed	Issues in learning an ontology from text
title_short	Issues in learning an ontology from text
title_sort	issues in learning an ontology from text
topic	Proceedings
url	https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2679401/ https://www.ncbi.nlm.nih.gov/pubmed/19426458 http://dx.doi.org/10.1186/1471-2105-10-S5-S1
work_keys_str_mv	AT brewsterchristopher issuesinlearninganontologyfromtext AT juppsimon issuesinlearninganontologyfromtext AT lucianojoanne issuesinlearninganontologyfromtext AT shottondavid issuesinlearninganontologyfromtext AT stevensrobertd issuesinlearninganontologyfromtext AT zhangziqi issuesinlearninganontologyfromtext

Issues in learning an ontology from text

Ejemplares similares