Publishing DisGeNET as nanopublications

Nuria Queralt-Rosinach; Tobias Kuhn; Christine Chichester; Michel Dumontier; Ferran Sanz; Laura I. Furlong

doi:10.3233/SW-150189

Publishing DisGeNET as nanopublications

Nuria Queralt-Rosinach, Tobias Kuhn, Christine Chichester, Michel Dumontier, Ferran Sanz, Laura I. Furlong^*

^*Corresponding author for this work

Research output: Contribution to journal › Article › Academic › peer-review

Abstract

The increasing and unprecedented publication rate in the biomedical field is a major bottleneck for knowledge discovery in the Life Sciences. The manual curation of facts from published scientific papers is slow and inefficient, and therefore new approaches are needed that can enable the automatic, scalable and reliable extraction of assertions. While the publication of scientific assertions and datasets on the Semantic Web is gaining traction, it also creates new challenges such as the proper representation of provenance and versioning. Here, we address these issues and describe our efforts to represent the DisGeNET database of human gene-disease associations as permanent, immutable, and provenance rich digital objects called nanopublications. Our nanopublications are the first instance of a Linked Data model that ensures stable interlinking of the assertion and its metadata by Trusty URIs. As DisGeNET integrates manually curated as well as text-mined data of different origins, the semantic description of the evidence for each assertion is important to provide trust and allow evidence-based hypothesis generation. Here, we describe our steps to ensure high quality and demonstrate the utility of linking our data to other datasets on the emerging Semantic Web.

Original language	English
Pages (from-to)	519-528
Journal	Semantic web
Volume	7
Issue number	5
DOIs	https://doi.org/10.3233/SW-150189
Publication status	Published - 2016
Externally published	Yes

Keywords

Gene-disease associations
linked data
nanopublication
provenance
trusty URIs

Access to Document

10.3233/SW-150189Licence: Unspecified

Cite this

@article{8db422ff9a5843afbf689d2a44802f91,

title = "Publishing DisGeNET as nanopublications",

abstract = "The increasing and unprecedented publication rate in the biomedical field is a major bottleneck for knowledge discovery in the Life Sciences. The manual curation of facts from published scientific papers is slow and inefficient, and therefore new approaches are needed that can enable the automatic, scalable and reliable extraction of assertions. While the publication of scientific assertions and datasets on the Semantic Web is gaining traction, it also creates new challenges such as the proper representation of provenance and versioning. Here, we address these issues and describe our efforts to represent the DisGeNET database of human gene-disease associations as permanent, immutable, and provenance rich digital objects called nanopublications. Our nanopublications are the first instance of a Linked Data model that ensures stable interlinking of the assertion and its metadata by Trusty URIs. As DisGeNET integrates manually curated as well as text-mined data of different origins, the semantic description of the evidence for each assertion is important to provide trust and allow evidence-based hypothesis generation. Here, we describe our steps to ensure high quality and demonstrate the utility of linking our data to other datasets on the emerging Semantic Web.",

keywords = "Gene-disease associations, linked data, nanopublication, provenance, trusty URIs",

author = "Nuria Queralt-Rosinach and Tobias Kuhn and Christine Chichester and Michel Dumontier and Ferran Sanz and Furlong, {Laura I.}",

year = "2016",

doi = "10.3233/SW-150189",

language = "English",

volume = "7",

pages = "519--528",

journal = "Semantic web",

issn = "1570-0844",

publisher = "IOS Press",

number = "5",

}

TY - JOUR

T1 - Publishing DisGeNET as nanopublications

AU - Queralt-Rosinach, Nuria

AU - Kuhn, Tobias

AU - Chichester, Christine

AU - Dumontier, Michel

AU - Sanz, Ferran

AU - Furlong, Laura I.

PY - 2016

Y1 - 2016

N2 - The increasing and unprecedented publication rate in the biomedical field is a major bottleneck for knowledge discovery in the Life Sciences. The manual curation of facts from published scientific papers is slow and inefficient, and therefore new approaches are needed that can enable the automatic, scalable and reliable extraction of assertions. While the publication of scientific assertions and datasets on the Semantic Web is gaining traction, it also creates new challenges such as the proper representation of provenance and versioning. Here, we address these issues and describe our efforts to represent the DisGeNET database of human gene-disease associations as permanent, immutable, and provenance rich digital objects called nanopublications. Our nanopublications are the first instance of a Linked Data model that ensures stable interlinking of the assertion and its metadata by Trusty URIs. As DisGeNET integrates manually curated as well as text-mined data of different origins, the semantic description of the evidence for each assertion is important to provide trust and allow evidence-based hypothesis generation. Here, we describe our steps to ensure high quality and demonstrate the utility of linking our data to other datasets on the emerging Semantic Web.

AB - The increasing and unprecedented publication rate in the biomedical field is a major bottleneck for knowledge discovery in the Life Sciences. The manual curation of facts from published scientific papers is slow and inefficient, and therefore new approaches are needed that can enable the automatic, scalable and reliable extraction of assertions. While the publication of scientific assertions and datasets on the Semantic Web is gaining traction, it also creates new challenges such as the proper representation of provenance and versioning. Here, we address these issues and describe our efforts to represent the DisGeNET database of human gene-disease associations as permanent, immutable, and provenance rich digital objects called nanopublications. Our nanopublications are the first instance of a Linked Data model that ensures stable interlinking of the assertion and its metadata by Trusty URIs. As DisGeNET integrates manually curated as well as text-mined data of different origins, the semantic description of the evidence for each assertion is important to provide trust and allow evidence-based hypothesis generation. Here, we describe our steps to ensure high quality and demonstrate the utility of linking our data to other datasets on the emerging Semantic Web.

KW - Gene-disease associations

KW - linked data

KW - nanopublication

KW - provenance

KW - trusty URIs

U2 - 10.3233/SW-150189

DO - 10.3233/SW-150189

M3 - Article

SN - 1570-0844

VL - 7

SP - 519

EP - 528

JO - Semantic web

JF - Semantic web

IS - 5

ER -