Catégories
EN

Analysis of U.S. Federal Funding Agency Data Sharing Policies 2020 Highlights and Key Observations

Authors : Reid I. Boehm, Hannah Calkins, Patricia B. Condon, Jonathan Petters, Rachel Woodbrook

Federal funding agencies in the United States (U.S.) continue to work towards implementing their plans to increase public access to funded research and comply with the 2013 Office of Science and Technology memo Increasing Access to the Results of Federally Funded Scientific Research.

In this article we report on an analysis of research data sharing policy documents from 17 U.S. federal funding agencies as of February 2021. Our analysis is guided by two questions: 1.) What do the findings suggest about the current state of and trends in U.S. federal funding agency data sharing requirements? 2.) In what ways are universities, institutions, associations, and researchers affected by and responding to these policies?

Over the past five years, policy updates were common among these agencies and several themes have been thoroughly developed in that time; however, uncertainty remains around how funded researchers are expected to satisfy these policy requirements.

URL : Analysis of U.S. Federal Funding Agency Data Sharing Policies 2020 Highlights and Key Observations

DOI : https://doi.org/10.2218/ijdc.v17i1.791

Catégories
EN

What are Researchers’ Needs in Data Discovery? Analysis and Ranking of a Large-Scale Collection of Crowdsourced Use Cases

Authors : Brigitte Mathiak, Nick Juty, Alessia Bardi, Julien Colomb, Peter Kraker

Data discovery is important to facilitate data re-use. In order to help frame the development and improvement of data discovery tools, we collected a list of requirements and users’ wishes.

This paper presents the analysis of these 101 use cases to examine data discovery requirements; these cases were collected between 2019 and 2020. We categorized the information across 12 ‘topics’ and eight types of users.

While the availability of metadata was an expected topic of importance, users were also keen on receiving more information on data citation and a better overview of their field. We conducted and analysed a survey among data infrastructure specialists in a first attempt at ranking the requirements.

Between these data professionals, these rankings were very different, excepting the availability of metadata and data quality assessment.

URL : What are Researchers’ Needs in Data Discovery? Analysis and Ranking of a Large-Scale Collection of Crowdsourced Use Cases

DOI : http://doi.org/10.5334/dsj-2023-003

Catégories
EN

Cluster Analysis of Open Research Data: A Case for Replication Metadata

Author : Ana Trisovic

Research data are often released upon journal publication to enable result verification and reproducibility. For that reason, research dissemination infrastructures typically support diverse datasets coming from numerous disciplines, from tabular data and program code to audio-visual files. Metadata, or data about data, is critical to making research outputs adequately documented and FAIR.

Aiming to contribute to the discussions on the development of metadata for research outputs, I conducted an exploratory analysis to determine how research datasets cluster based on what researchers organically deposit together. I use the content of over 40,000 datasets from the Harvard Dataverse research data repository as my sample for the cluster analysis.

I find that the majority of the clusters are formed by single-type datasets, while in the rest of the sample, no meaningful clusters can be identified. For the result interpretation, I use the metadata standard employed by DataCite, a leading organization for documenting a scholarly record, and map existing resource types to my results.

About 65% of the sample can be described with a single-type metadata (such as Dataset, Software orReport), while the rest would require aggregate metadata types. Though DataCite supports an aggregate type such as a Collection, I argue that a significant number of datasets, in particular those containing both data and code files (about 20% of the sample), would be more accurately described as a Replication resource metadata type. Such resource type would be particularly useful in facilitating research reproducibility.

URL : Cluster Analysis of Open Research Data: A Case for Replication Metadata

DOI : https://doi.org/10.2218/ijdc.v17i1.833

Catégories
EN

Data Management Plans: Implications for Automated Analyses

Authors : Ngoc-Minh Pham, Heather Moulaison-Sandy, Bradley Wade Bishop, Hannah Gunderman

Data management plans (DMPs) are an essential part of planning data-driven research projects and ensuring long-term access and use of research data and digital objects; however, as text-based documents, DMPs must be analyzed manually for conformance to funder requirements.

This study presents a comparison of DMPs evaluations for 21 funded projects using 1) an automated means of analysis to identify elements that align with best practices in support of open research initiatives and 2) a manually-applied scorecard measuring these same elements.

The automated analysis revealed that terms related to availability (90% of DMPs), metadata (86% of DMPs), and sharing (81% of DMPs) were reliably supplied. Manual analysis revealed 86% (n = 18) of funded DMPs were adequate, with strong discussions of data management personnel (average score: 2 out of 2), data sharing (average score 1.83 out of 2), and limitations to data sharing (average score: 1.65 out of 2).

This study reveals that the automated approach to DMP assessment yields less granular yet similar results to manual assessments of the DMPs that are more efficiently produced. Additional observations and recommendations are also presented to make data management planning exercises and automated analysis even more useful going forward.

URL : Data Management Plans: Implications for Automated Analyses

DOI : http://doi.org/10.5334/dsj-2023-002

Catégories
EN

The Role of Metadata and Vocabulary Standards in Enabling Scientific Data Interoperability: A Study of Earth System Science Data Facilities

Authors : Matthew S. Mayernik, Yauheniya Liapich

Objective

Journal publishers within many sciences are increasingly expecting data to be deposited into repositories that support the FAIR principles. Data repositories are thus needing to determine what implications the FAIR principles have on their existing services and systems. Metadata standards and controlled vocabularies are specifically called out as core components of the FAIR principles related to interoperability.

Methods

This paper looks specifically at the ways that metadata standards and controlled vocabularies are used by Earth system science data repositories. Data sets from 55 data facilities were examined to determine which metadata standards and controlled subject / keyword vocabularies were used.

Results

The findings indicate that only the ISO 19115:2003 and DataCite metadata standards are used by more than 40% of the data facilities, and the NASA Global Change Master Directory (GCMD) keywords are the only keyword vocabulary of broad use within this community.

Conclusions

These findings raise questions about the extent to which metadata standards and keyword vocabularies can facilitate interoperability beyond narrow sub-sections of the data facility communities. This study also points to systematic challenges related to migration to new standards.

URL : The Role of Metadata and Vocabulary Standards in Enabling Scientific Data Interoperability: A Study of Earth System Science Data Facilities

DOI : https://doi.org/10.7191/jeslib.619

Catégories
FR

Prise de décision par les pouvoirs publics et partage des données de la recherche, une approche par le risque

Autrice/Author : Fleur Nadine Ndjock

Si la recherche scientifique est prioritairement financée par l’État au travers des dotations budgétaires, il est logique que les résultats de cette recherche aide (contribue) à la prise de décision efficace par les pouvoirs publics pour le développement d’un pays.

Pour ce faire, les données issues de la recherche doivent être partagées s’il est vrai que pour décider, l’on a besoin d’informations et les données de la recherche qu’elles soient d’observation, expérimentales ou de simulation, sont importantes dans le processus décisionnel stratégique.

Cet article vise un double objectif : Il s’agit d’une part d’établir une typologie des risques qu’encourt le partage de données de la recherche avec les pouvoirs publics, mais aussi, de questionner les concepts à mobiliser pour rendre compte des enjeux de ces risques, car ils sont déterminants et peuvent influencer les motivations de leur partage.

DOI : https://doi.org/10.4000/ctd.8301

 

Catégories
EN

An iterative and interdisciplinary categorisation process towards FAIRer digital resources for sensitive life-sciences data

Authors : Romain David, Christian Ohmann, Jan‑Willem Boiten, Mónica Cano Abadía, Florence Bietrix, Steve Canham, Maria Luisa Chiusano, Walter Dastrù, Arnaud Laroquette, Dario Longo, Michaela Th. Mayrhofer, Maria Panagiotopoulou, Audrey S. Richard, Sergey Goryanin, Pablo Emilio Verde

For life science infrastructures, sensitive data generate an additional layer of complexity. Cross-domain categorisation and discovery of digital resources related to sensitive data presents major interoperability challenges. To support this FAIRification process, a toolbox demonstrator aiming at support for discovery of digital objects related to sensitive data (e.g., regulations, guidelines, best practice, tools) has been developed.

The toolbox is based upon a categorisation system developed and harmonised across a cluster of 6 life science research infrastructures. Three different versions were built, tested by subsequent pilot studies, finally leading to a system with 7 main categories (sensitive data type, resource type, research field, data type, stage in data sharing life cycle, geographical scope, specific topics).

109 resources attached with the tags in pilot study 3 were used as the initial content for the toolbox demonstrator, a software tool allowing searching of digital objects linked to sensitive data with filtering based upon the categorisation system.

Important next steps are a broad evaluation of the usability and user-friendliness of the toolbox, extension to more resources, broader adoption by different life-science communities, and a long-term vision for maintenance and sustainability.

URL : An iterative and interdisciplinary categorisation process towards FAIRer digital resources for sensitive life-sciences data

DOI : https://doi.org/10.1038/s41598-022-25278-z