Catégories
EN

Scientific data management in the federal government: A case study of NOAA and responsibility for preserving digital data

Authors : Adam Kriesberg, Jacob Kowall

In this paper, we examine the ways in which the evolution of federal and agency‐specific data management policies has affected and continues to affect the long‐term preservation of digital scientific data produced by the United States government.

After reviewing the existing literature on the role of archival theory and practice in the preservation of scientific data, we present the case of the National Oceanic and Atmospheric Administration (NOAA) to analyze how data management activities at this agency are shaped by legislative mandates as well as both government‐wide and agency‐specific information‐management policies.

Through the connected network of law, federal policy, agency policy, and the records schedules which govern recordkeeping practice in the federal government, we propose a number of further questions on how government agencies can effectively provide for the management of scientific data as federal records.

DOI : https://doi.org/10.1002/pra2.266

Catégories
EN

Data sharing policies of journals in life, health, and physical sciences indexed in Journal Citation Reports

Authors : Jihyun Kim, Soon Kim, Hye-Min Cho, Jae Hwa Chang, Soo Young Kim

Background

Many scholarly journals have established their own data-related policies, which specify their enforcement of data sharing, the types of data to be submitted, and their procedures for making data available.

However, except for the journal impact factor and the subject area, the factors associated with the overall strength of the data sharing policies of scholarly journals remain unknown.

This study examines how factors, including impact factor, subject area, type of journal publisher, and geographical location of the publisher are related to the strength of the data sharing policy.

Methods

From each of the 178 categories of the Web of Science’s 2017 edition of Journal Citation Reports, the top journals in each quartile (Q1, Q2, Q3, and Q4) were selected in December 2018. Of the resulting 709 journals (5%), 700 in the fields of life, health, and physical sciences were selected for analysis.

Four of the authors independently reviewed the results of the journal website searches, categorized the journals’ data sharing policies, and extracted the characteristics of individual journals.

Univariable multinomial logistic regression analyses were initially conducted to determine whether there was a relationship between each factor and the strength of the data sharing policy.

Based on the univariable analyses, a multivariable model was performed to further investigate the factors related to the presence and/or strength of the policy.

Results

Of the 700 journals, 308 (44.0%) had no data sharing policy, 125 (17.9%) had a weak policy, and 267 (38.1%) had a strong policy (expecting or mandating data sharing). The impact factor quartile was positively associated with the strength of the data sharing policies.

Physical science journals were less likely to have a strong policy relative to a weak policy than Life science journals (relative risk ratio [RRR], 0.36; 95% CI [0.17–0.78]). Life science journals had a greater probability of having a weak policy relative to no policy than health science journals (RRR, 2.73; 95% CI [1.05–7.14]).

Commercial publishers were more likely to have a weak policy relative to no policy than non-commercial publishers (RRR, 7.87; 95% CI, [3.98–15.57]). Journals by publishers in Europe, including the majority of those located in the United Kingdom and the Netherlands, were more likely to have a strong data sharing policy than a weak policy (RRR, 2.99; 95% CI [1.85–4.81]).

Conclusions

These findings may account for the increase in commercial publishers’ engagement in data sharing and indicate that European national initiatives that encourage and mandate data sharing may influence the presence of a strong policy in the associated journals.

Future research needs to explore the factors associated with varied degrees in the strength of a data sharing policy as well as more diverse characteristics of journals related to the policy strength.

URL : Data sharing policies of journals in life, health, and physical sciences indexed in Journal Citation Reports

DOI : https://doi.org/10.7717/peerj.9924

Catégories
FR

Entrepôts de données de recherche : mesurer l’impact de l’Open Science à l’aune de la consultation des jeux de données déposés

Auteur/Author  : Violaine Rebouillat

Les décennies 2000 et 2010 ont vu se développer un nombre croissant de e-infrastructures de recherche, rendant plus aisés le partage et l’accès aux données scientifiques. Cette tendance s’est vue renforcée par l’essor de politiques d’ouverture des données, lesquelles ont donné lieu à une multiplication de réservoirs de données – aussi appelés « entrepôts de données ». Quantifier et qualifier l’utilisation des données rendues publiques constitue un élément essentiel pour évaluer l’impact des politiques d’ouverture des données.

Dans cet article, nous questionnons l’utilisation des données déposées dans les entrepôts. Dans quelle mesure ces données sont-elles consultées et téléchargées ?

L’article présente les premiers résultats d’une enquête quantitative auprès de 20 entrepôts. Il esquisse deux tendances, qui restent à ce stade propres à l’échantillon étudié, à savoir : (1) l’augmentation globale du nombre de consultations, de téléchargements et de données disponibles dans les entrepôts sur la période étudiée (2015-2020), et (2) la concentration des téléchargements sur une proportion relativement faible des données de l’entrepôt (de l’ordre de 10% à 30%).

URL : https://hal.archives-ouvertes.fr/hal-02928817/

Catégories
FR

Data librarian et services aux chercheurs en bibliothèque universitaire : de nouvelles médiations en émergence

Auteur/Author : Florence Thiault

Les services à destination des chercheurs se développent dans les bibliothèques universitaires françaises. L’augmentation de la quantité de données de recherche produites et réutilisées par les chercheurs pose des défis importants aux bibliothèques universitaires.

De nouvelles compétences associées à un profil professionnel spécifique celui de datalibrarian sont nécessaires pour assurer ces missions d’accompagnement à la recherche. Ce spécialiste des données à vocation à accompagner les chercheurs dans le cycle de vie de la recherche en assurant une collaboration active avec une série d’acteurs internes et externes.

Cette communication présente trois cas d’études emblématiques dans le registre des médiations à destination des chercheurs : l’analyse de la production scientifique, l’accompagnement à la recherche et à la publication ainsi que la gestion des données de recherche.

URL : https://hal.archives-ouvertes.fr/hal-02972705

Catégories
EN

Do researchers use open research data? Exploring the relationships between usage trends and metadata quality across scientific disciplines from the Figshare case

Authors : Alfonso Quarati, Juliana E Raffaghelli

Open research data (ORD) have been considered a driver of scientific transparency. However, data friction, as the phenomenon of data underutilisation for several causes, has also been pointed out.

A factor often called into question for ORD low usage is the quality of the ORD and associated metadata. This work aims to illustrate the use of ORD, published by the Figshare scientific repository, concerning their scientific discipline, their type and compared with the quality of their metadata.

Considering all the Figshare resources and carrying out a programmatic quality assessment of their metadata, our analysis highlighted two aspects. First, irrespective of the scientific domain considered, most ORD are under-used, but with exceptional cases which concentrate most researchers’ attention.

Second, there was no evidence that the use of ORD is associated with good metadata publishing practices. These two findings opened to a reflection about the potential causes of such data friction.

URL : Do researchers use open research data? Exploring the relationships between usage trends and metadata quality across scientific disciplines from the Figshare case

DOI : https://doi.org/10.1177/0165551520961048

Catégories
EN

Who Does What? – Research Data Management at ETH Zurich

Authors: Matthias Töwe, Caterina Barillari

We present the approach to Research Data Management (RDM) support for researchers taken at ETH Zurich. Overall requirements are governed by institutional guidelines for Research Integrity, funders’ regulations, and legal obligations. The ETH approach is based on the distinction of three phases along the research data life-cycle: 1. Data Management Planning; 2. Active RDM; 3. Data Publication and Preservation. Two ETH units, namely the Scientific IT Services and the ETH Library, provide support for different aspects of these phases, building on their respective competencies. They jointly offer trainings, consulting, information, and materials for the first phase.

The second phase deals with data which is in current use in active research projects. Scientific IT Services provide their own platform, openBIS, for keeping track of raw, processed and analysed data, in addition to organising samples, materials, and scientific procedures.

ETH Library operates solutions for the third phase within the infrastructure of ETH Zurich’s central IT Services. The Research Collection is the institutional repository for research output including Research Data, Open Access publications, and ETH Zurich’s bibliography.

URL : Who Does What? – Research Data Management at ETH Zurich

DOI : http://doi.org/10.5334/dsj-2020-036

Catégories
EN

Towards FAIR protocols and workflows: the OpenPREDICT use case

Authors : Remzi Celebi, Joao Rebelo Moreira, Ahmed A. Hassan, Sandeep Ayyar, Lars Ridder, Tobias Kuhn, Michel Dumontier

It is essential for the advancement of science that researchers share, reuse and reproduce each other’s workflows and protocols. The FAIR principles are a set of guidelines that aim to maximize the value and usefulness of research data, and emphasize the importance of making digital objects findable and reusable by others.

The question of how to apply these principles not just to data but also to the workflows and protocols that consume and produce them is still under debate and poses a number of challenges. In this paper we describe a two-fold approach of simultaneously applying the FAIR principles to scientific workflows as well as the involved data.

We apply and evaluate our approach on the case of the PREDICT workflow, a highly cited drug repurposing workflow. This includes FAIRification of the involved datasets, as well as applying semantic technologies to represent and store data about the detailed versions of the general protocol, of the concrete workflow instructions, and of their execution traces.

We propose a semantic model to address these specific requirements and was evaluated by answering competency questions. This semantic model consists of classes and relations from a number of existing ontologies, including Workflow4ever, PROV, EDAM, and BPMN.

This allowed us then to formulate and answer new kinds of competency questions. Our evaluation shows the high degree to which our FAIRified OpenPREDICT workflow now adheres to the FAIR principles and the practicality and usefulness of being able to answer our new competency questions.

URL : Towards FAIR protocols and workflows: the OpenPREDICT use case

DOI : https://doi.org/10.7717/peerj-cs.281