Catégories
EN

Transparency, provenance and collections as data: the National Library of Scotland’s Data Foundry

Author : Sarah Ames

‘Collections as data’ has become a core activity for libraries in recent years: it is important that we make collections available in machine-readable formats to enable and encourage computational research. However, while this is a necessary output, discussion around the processes and workflows required to turn collections into data, and to make collections data available openly, are just as valuable.

With libraries increasingly becoming producers of their own collections – presenting data from digitisation and digital production tools as part of datasets, for example – and making collections available at scale through mass-digitisation programmes, the trustworthiness of our processes comes into question.

In a world of big data, often of unclear origins, how can libraries be transparent about the ways in which collections are turned into data, how do we ensure that biases in our collections are recognised and not amplified, and how do we make these datasets available openly for reuse?

This paper presents a case study of work underway at the National Library of Scotland to present collections as data in an open and transparent way – from establishing a new Digital Scholarship Service, to workflows and online presentation of datasets.

It considers the changes to existing processes needed to produce the Data Foundry, the National Library of Scotland’s open data delivery platform, and explores the practical challenges of presenting collections as data online in an open, transparent and coherent manner.

URL : Transparency, provenance and collections as data: the National Library of Scotland’s Data Foundry

Original location : https://www.liberquarterly.eu/article/10.18352/lq.10371/

Catégories
EN

Digital Object Identifier (DOI) Under the Context of Research Data Librarianship

AuthorJia Liu

A digital object identifier (DOI) is an increasingly prominent persistent identifier in finding and accessing scholarly information. This paper intends to present an overview of global development and approaches in the field of DOI and DOI services with a slight geographical focus on Germany.

At first, the initiation and components of the DOI system and the structure of a DOI name are explored. Next, the fundamental and specific characteristics of DOIs are described and DOIs for three (3) kinds of typical intellectual entities in the scholar communication are dealt with; then, a general DOI service pyramid is sketched with brief descriptions of functions of institutions at different levels.

After that, approaches of the research data librarianship community in the field of RDM, especially DOI services, are elaborated. As examples, the DOI services provided in German research libraries as well as best practices of DOI services in a German library are introduced; and finally, the current practices and some issues dealing with DOIs are summarized. It is foreseeable that DOI, which is crucial to FAIR research data, will gain extensive recognition in the scientific world.

URL : Digital Object Identifier (DOI) Under the Context of Research Data Librarianship

DOI : https://doi.org/10.7191/jeslib.2021.1180

Catégories
EN

Affiliation Information in DataCite Dataset Metadata: a Flemish Case Study

Author/Auteur : Niek Van Wettere

This article aims to evaluate how and to what extent metadata of datasets indexed in DataCite offer clear human- or machine-readable information that enables the research data to be linked to a particular research institution.

Two main pathways are explored. First, researchers can encode their affiliation information at the moment of data submission. This can be done by means of free-text metadata fields or via the inclusion of identifiers such as GRID/ROR and ORCID. Second, affiliation information can be traced indirectly through linking between a dataset and associated publications, given that the metadata of publications is often more explicit about affiliation information than the metadata of datasets.

Both pathways of affiliation information encoding are evaluated on the basis of metadata pertaining to datasets created at the five Flemish universities. It is shown that good practices such as encoding of affiliation information in a dedicated metadata field or inclusion of ORCID in the metadata are on the rise, but could be expanded further.

Finally, the establishment of links between datasets and related publications is often lacking in dataset metadata, although there are important differences between data repositories, as is also demonstrated in a more data-intensive follow-up analysis based on random samples of metadata records.

It is important that data repositories address this issue by providing a metadata field clearly dedicated to associated publications, prominently displayed on the landing page of the dataset.

URL : Affiliation Information in DataCite Dataset Metadata: a Flemish Case Study

DOI : http://doi.org/10.5334/dsj-2021-013

Catégories
EN

Open Research Data and Open Peer Review: Perceptions of a Medical and Health Sciences Community in Greece

Authors : Eirini Delikoura, Dimitrios Kouis

Recently significant initiatives have been launched for the dissemination of Open Access as part of the Open Science movement. Nevertheless, two other major pillars of Open Science such as Open Research Data (ORD) and Open Peer Review (OPR) are still in an early stage of development among the communities of researchers and stakeholders.

The present study sought to unveil the perceptions of a medical and health sciences community about these issues. Through the investigation of researchers‘ attitudes, valuable conclusions can be drawn, especially in the field of medicine and health sciences, where an explosive growth of scientific publishing exists.

A quantitative survey was conducted based on a structured questionnaire, with 179 valid responses. The participants in the survey agreed with the Open Peer Review principles. However, they ignored basic terms like FAIR (Findable, Accessible, Interoperable, and Reusable) and appeared incentivized to permit the exploitation of their data.

Regarding Open Peer Review (OPR), participants expressed their agreement, implying their support for a trustworthy evaluation system.

Conclusively, researchers need to receive proper training for both Open Research Data principles and Open Peer Review processes which combined with a reformed evaluation system will enable them to take full advantage of the opportunities that arise from the new scholarly publishing and communication landscape.

URL : Open Research Data and Open Peer Review: Perceptions of a Medical and Health Sciences Community in Greece

DOI : https://doi.org/10.3390/publications9020014

Catégories
EN

Adaptable Methods for Training in Research Data Management

Authors: Katarzyna Biernacka, Kerstin Helbig, Petra Buchholz

The management of research data has become an essential aspect of good scientific practice. Education in research data management is, however, scarce. The low number of trainers can be attributed on the one hand to a lack of educational paths. On the other hand, qualification opportunities for academics who have already completed their studies and are in employment are missing.

Within the research project FDMentor a Train-the-Trainer programme was therefore developed to teach potential multipliers of research data management, and at the same time impart basic didactic knowledge.

The resulting concept was created, in addition to freely re-usable materials, to support researchers and research support staff in passing on this knowledge. In addition, the generic development and free licensing of the concept enables transferability to other thematic contexts, such as Open Access or Open Science.

URL : Adaptable Methods for Training in Research Data Management

DOI : http://doi.org/10.5334/dsj-2021-014

Catégories
EN

Versioning Data Is About More than Revisions: A Conceptual Framework and Proposed Principles

Authors : Jens Klump, Lesley Wyborn, Mingfang Wu, Julia Martin, Robert R. Downs, Ari Asmi

A dataset, small or big, is often changed to correct errors, apply new algorithms, or add new data (e.g., as part of a time series), etc.

In addition, datasets might be bundled into collections, distributed in different encodings or mirrored onto different platforms. All these differences between versions of datasets need to be understood by researchers who want to cite the exact version of the dataset that was used to underpin their research.

Failing to do so reduces the reproducibility of research results. Ambiguous identification of datasets also impacts researchers and data centres who are unable to gain recognition and credit for their contributions to the collection, creation, curation and publication of individual datasets.

Although the means to identify datasets using persistent identifiers have been in place for more than a decade, systematic data versioning practices are currently not available. In this work, we analysed 39 use cases and current practices of data versioning across 33 organisations.

We noticed that the term ‘version’ was used in a very general sense, extending beyond the more common understanding of ‘version’ to refer primarily to revisions and replacements. Using concepts developed in software versioning and the Functional Requirements for Bibliographic Records (FRBR) as a conceptual framework, we developed six foundational principles for versioning of datasets: Revision, Release, Granularity, Manifestation, Provenance and Citation.

These six principles provide a high-level framework for guiding the consistent practice of data versioning and can also serve as guidance for data centres or data providers when setting up their own data revision and version protocols and procedures.

URL : Versioning Data Is About More than Revisions: A Conceptual Framework and Proposed Principles

DOI : http://doi.org/10.5334/dsj-2021-012

Catégories
EN

Openness in Big Data and Data Repositories. The Application of an Ethics Framework for Big Data in Healthand Research

Authors : Vicki Xafis, Markus K. Labude

There is a growing expectation, or even requirement, for researchers to deposit a variety of research data in data repositories as a condition of funding or publication. This expectation recognizes the enormous benefits of data collected and created for research purposes being made available for secondary uses, as open science gains increasing support.

This is particularly so in the context of big data, especially where health data is involved. There are, however, also challenges relating to the collection, storage, and re-use of research data.

This paper gives a brief overview of the landscape of data sharing via data repositories and discusses some of the key ethical issues raised by the sharing of health-related research data, including expectations of privacy and confidentiality, the transparency of repository governance structures, access restrictions, as well as data ownership and the fair attribution of credit.

To consider these issues and the values that are pertinent, the paper applies the deliberative balancing approach articulated in the Ethics Framework for Big Data in Health and Research (Xafis et al. 2019) to the domain of Openness in Big Data and Data Repositories.

Please refer to that article for more information on how this framework is to be used, including a full explanation of the key values involved and the balancing approach used in the case study at the end.

URL : Openness in Big Data and Data Repositories. The Application of an Ethics Framework for Big Data in Healthand Research

DOI : https://doi.org/10.1007/s41649-019-00097-z