Catégories
EN

Assessing Researcher Data Sharing Practices Using Publicly Available Dataset Metadata: A Reproducible Workflow for Institutional Stakeholders

Authors : Danielle R. Kirsch, Isaac Wink

Introduction

Scholarly communities are experiencing increased emphasis on research data sharing and reuse. Given the range of available data repositories and variability in the use of persistent identifiers (PIDs) for individuals and institutions, aggregating datasets by researchers at a specific institution is a significant challenge. We developed a reproducible workflow to evaluate 1) data repository use by researchers at our institutions, 2) metrics of reuse, and 3) quality of metadata records.

Methods

We used the DataCite REST API to locate metadata records for datasets with creator affiliation names that matched our institutions. We performed substantial cleaning and deduplication before comparing repository use, citations, usage metrics, and metadata completeness.

Results

The most common data repositories were Dryad, Harvard Dataverse, figshare, Zenodo, ICPSR, and one institution’s institutional repository. View, download, and citation counts were available from a limited number of repositories, with some discrepancies between different citation reporting methods. PIDs were more frequently used for authors and affiliations than funders, and the inclusion of PIDs varied across repositories, with Dryad being the most consistent.

Discussion & Conclusion

Although broad trends in repository and PID use were similar between institutions, our analysis also surfaced examples that illustrate inconsistencies in metadata across repositories. Variable implementation of the DataCite metadata schema requires a significant amount of data cleaning to obtain meaningful results. Even then, incomplete metadata makes some datasets impossible to locate. Repositories, funders, researchers, and institutional open data advocates must coordinate to create complete and usable metadata that integrates datasets into the scholarship ecosystem.

URL : Assessing Researcher Data Sharing Practices Using Publicly Available Dataset Metadata: A Reproducible Workflow for Institutional Stakeholders

DOI : https://doi.org/10.31274/jlsc.22913

Catégories
EN

FAIRness of Research Data in the European Humanities Landscape

Authors : Ljiljana Poljak Bilić, Kristina Posavec

This paper explores the landscape of research data in the humanities in the European context, delving into their diversity and the challenges of defining and sharing them. It investigates three aspects: the types of data in the humanities, their representation in repositories, and their alignment with the FAIR principles (Findable, Accessible, Interoperable, Reusable).

By reviewing datasets in repositories, this research determines the dominant data types, their openness, licensing, and compliance with the FAIR principles. This research provides important insight into the heterogeneous nature of humanities data, their representation in the repository, and their alignment with FAIR principles, highlighting the need for improved accessibility and reusability to improve the overall quality and utility of humanities research data.

URL : FAIRness of Research Data in the European Humanities Landscape

DOI : https://doi.org/10.3390/publications12010006

Catégories
EN

Dataset search: a survey

Authors : Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis, Luis-Daniel Ibáñez, Emilia Kacprzak, Paul Groth

Generating value from data requires the ability to find, access and make sense of datasets. There are many efforts underway to encourage data sharing and reuse, from scientific publishers asking authors to submit data alongside manuscripts to data marketplaces, open data portals and data communities.

Google recently beta-released a search service for datasets, which allows users to discover data stored in various online repositories via keyword queries. These developments foreshadow an emerging research field around dataset search or retrieval that broadly encompasses frameworks, methods and tools that help match a user data need against a collection of datasets.

Here, we survey the state of the art of research and commercial systems and discuss what makes dataset search a field in its own right, with unique challenges and open questions.

We look at approaches and implementations from related areas dataset search is drawing upon, including information retrieval, databases, entity-centric and tabular search in order to identify possible paths to tackle these questions as well as immediate next steps that will take the field forward.

URL : Dataset search: a survey

DOI : https://doi.org/10.1007/s00778-019-00564-x