Cargando…

The molecular entities in linked data dataset

The Molecular Entities in Linked Data (MEiLD) dataset comprises data of distinct atoms, molecules, ions, ion pairs, radicals, radical ions, and others that can be identifiable as separately distinguishable chemical entities. The dataset is provided in a JSON-LD format and was generated by the SDFEat...

Descripción completa

Detalles Bibliográficos
Autores principales: Tomaszuk, Dominik, Szeremeta, Łukasz
Formato: Online Artículo Texto
Lenguaje:English
Publicado: Elsevier 2020
Materias:
Acceso en línea:https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7276506/
https://www.ncbi.nlm.nih.gov/pubmed/32529012
http://dx.doi.org/10.1016/j.dib.2020.105757
_version_ 1783542964447870976
author Tomaszuk, Dominik
Szeremeta, Łukasz
author_facet Tomaszuk, Dominik
Szeremeta, Łukasz
author_sort Tomaszuk, Dominik
collection PubMed
description The Molecular Entities in Linked Data (MEiLD) dataset comprises data of distinct atoms, molecules, ions, ion pairs, radicals, radical ions, and others that can be identifiable as separately distinguishable chemical entities. The dataset is provided in a JSON-LD format and was generated by the SDFEater, a tool that allows parsing atoms, bonds, and other molecule data. MEiLD contains 349,960 of ‘small’ chemical entities. Our dataset is based on the SDF files and is enriched with additional ontologies and line notation data. As a basis, the Molecular Entities in Linked Data dataset uses the Resource Description Framework (RDF) data model. Saving the data in such a model allows preserving the semantic relations, like hierarchical and associative, between them. To describe chemical molecules, vocabularies such as Chemical Vocabulary for Molecular Entities (CVME) and Simple Knowledge Organization System (SKOS) are used. The dataset can be beneficial, among others, for people concerned with research and development tools for cheminformatics and bioinformatics. In this paper, we describe various methods of access to our dataset. In addition to the MEiLD dataset, we publish the Shapes Constraint Language (SHACL) schema of our dataset and the CVME ontology. The data is available in Mendeley Data.
format Online
Article
Text
id pubmed-7276506
institution National Center for Biotechnology Information
language English
publishDate 2020
publisher Elsevier
record_format MEDLINE/PubMed
spelling pubmed-72765062020-06-10 The molecular entities in linked data dataset Tomaszuk, Dominik Szeremeta, Łukasz Data Brief Computer Science The Molecular Entities in Linked Data (MEiLD) dataset comprises data of distinct atoms, molecules, ions, ion pairs, radicals, radical ions, and others that can be identifiable as separately distinguishable chemical entities. The dataset is provided in a JSON-LD format and was generated by the SDFEater, a tool that allows parsing atoms, bonds, and other molecule data. MEiLD contains 349,960 of ‘small’ chemical entities. Our dataset is based on the SDF files and is enriched with additional ontologies and line notation data. As a basis, the Molecular Entities in Linked Data dataset uses the Resource Description Framework (RDF) data model. Saving the data in such a model allows preserving the semantic relations, like hierarchical and associative, between them. To describe chemical molecules, vocabularies such as Chemical Vocabulary for Molecular Entities (CVME) and Simple Knowledge Organization System (SKOS) are used. The dataset can be beneficial, among others, for people concerned with research and development tools for cheminformatics and bioinformatics. In this paper, we describe various methods of access to our dataset. In addition to the MEiLD dataset, we publish the Shapes Constraint Language (SHACL) schema of our dataset and the CVME ontology. The data is available in Mendeley Data. Elsevier 2020-05-27 /pmc/articles/PMC7276506/ /pubmed/32529012 http://dx.doi.org/10.1016/j.dib.2020.105757 Text en © 2020 The Author(s) http://creativecommons.org/licenses/by/4.0/ This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
spellingShingle Computer Science
Tomaszuk, Dominik
Szeremeta, Łukasz
The molecular entities in linked data dataset
title The molecular entities in linked data dataset
title_full The molecular entities in linked data dataset
title_fullStr The molecular entities in linked data dataset
title_full_unstemmed The molecular entities in linked data dataset
title_short The molecular entities in linked data dataset
title_sort molecular entities in linked data dataset
topic Computer Science
url https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7276506/
https://www.ncbi.nlm.nih.gov/pubmed/32529012
http://dx.doi.org/10.1016/j.dib.2020.105757
work_keys_str_mv AT tomaszukdominik themolecularentitiesinlinkeddatadataset
AT szeremetałukasz themolecularentitiesinlinkeddatadataset
AT tomaszukdominik molecularentitiesinlinkeddatadataset
AT szeremetałukasz molecularentitiesinlinkeddatadataset