Catégories
EN

Assessing Researcher Data Sharing Practices Using Publicly Available Dataset Metadata: A Reproducible Workflow for Institutional Stakeholders

Authors : Danielle R. Kirsch, Isaac Wink

Introduction

Scholarly communities are experiencing increased emphasis on research data sharing and reuse. Given the range of available data repositories and variability in the use of persistent identifiers (PIDs) for individuals and institutions, aggregating datasets by researchers at a specific institution is a significant challenge. We developed a reproducible workflow to evaluate 1) data repository use by researchers at our institutions, 2) metrics of reuse, and 3) quality of metadata records.

Methods

We used the DataCite REST API to locate metadata records for datasets with creator affiliation names that matched our institutions. We performed substantial cleaning and deduplication before comparing repository use, citations, usage metrics, and metadata completeness.

Results

The most common data repositories were Dryad, Harvard Dataverse, figshare, Zenodo, ICPSR, and one institution’s institutional repository. View, download, and citation counts were available from a limited number of repositories, with some discrepancies between different citation reporting methods. PIDs were more frequently used for authors and affiliations than funders, and the inclusion of PIDs varied across repositories, with Dryad being the most consistent.

Discussion & Conclusion

Although broad trends in repository and PID use were similar between institutions, our analysis also surfaced examples that illustrate inconsistencies in metadata across repositories. Variable implementation of the DataCite metadata schema requires a significant amount of data cleaning to obtain meaningful results. Even then, incomplete metadata makes some datasets impossible to locate. Repositories, funders, researchers, and institutional open data advocates must coordinate to create complete and usable metadata that integrates datasets into the scholarship ecosystem.

URL : Assessing Researcher Data Sharing Practices Using Publicly Available Dataset Metadata: A Reproducible Workflow for Institutional Stakeholders

DOI : https://doi.org/10.31274/jlsc.22913

Catégories
EN

Assessing Researcher Data Sharing Practices Using Publicly Available Dataset Metadata: A Reproducible Workflow for Institutional Stakeholders

Authors : Danielle R. Kirsch, Isaac Wink

Introduction: Scholarly communities are experiencing increased emphasis on research data sharing and reuse. Given the range of available data repositories and variability in the use of persistent identifiers (PIDs) for individuals and institutions, aggregating datasets by researchers at a specific institution is a significant challenge. We developed a reproducible workflow to evaluate 1) data repository use by researchers at our institutions, 2) metrics of reuse, and 3) quality of metadata records.

Methods: We used the DataCite REST API to locate metadata records for datasets with creator affiliation names that matched our institutions. We performed substantial cleaning and deduplication before comparing repository use, citations, usage metrics, and metadata completeness.

Results: The most common data repositories were Dryad, Harvard Dataverse, figshare, Zenodo, ICPSR, and one institution’s institutional repository. View, download, and citation counts were available from a limited number of repositories, with some discrepancies between different citation reporting methods. PIDs were more frequently used for authors and affiliations than funders, and the inclusion of PIDs varied across repositories, with Dryad being the most consistent.

Discussion & Conclusion: Although broad trends in repository and PID use were similar between institutions, our analysis also surfaced examples that illustrate inconsistencies in metadata across repositories. Variable implementation of the DataCite metadata schema requires a significant amount of data cleaning to obtain meaningful results. Even then, incomplete metadata makes some datasets impossible to locate. Repositories, funders, researchers, and institutional open data advocates must coordinate to create complete and usable metadata that integrates datasets into the scholarship ecosystem.

URL : Assessing Researcher Data Sharing Practices Using Publicly Available Dataset Metadata: A Reproducible Workflow for Institutional Stakeholders

DOI : https://doi.org/10.31274/jlsc.22913

Catégories
EN

Scientific production on data repositories and open science published in the Web of Science database: Methodi Ordinatio and content analysis

Authors : Sinval Adalberto Rodrigues-Junior, Marcelo Votto Texeira

The opening of scientific data proposed by the Open Science movement presupposes careful planning for data collection, organization, and treatment, aiming at their sharing, accessibility, and reuse. Data repositories have been conceived as structures necessary to enable open access to data.

This study aimed to analyze the influence of data repositories on the disclosure and sharing of scientific data proposed by the Open Science movement. The Methodi Ordinatio, developed to organize a portfolio of scientific publications, was adopted to analyze the subject of ‘Data Repositories’ and ‘Open Science’.

The studies were ranked using the InOrdinatio index, and the 15 best ranked studies were included and analyzed through Bardin’s content analysis. Most studies describe the structure involved in data repositories within the biological, chemical, and health areas.

Other studies addressed data reuse, data organization and analysis processes and tools, as well as data selection and classification algorithms. The units of analysis selected for the content analysis were categorized as open access, information technologies, data processing, and information retrieval.

Systems (processes and structures), metadata standards, ontologies, semantic web, data types, and their management were addressed by these studies. It is concluded that open data repositories are growing rapidly. Production with the greatest impact has occurred in the biological and biomedical/health areas, highlighting the structure involved in repositories within these fields.

Data repositories provide systems for depositing, managing, searching, accessing, and reusing data based on processes and technologies — often developed as open-source software — in alignment with the proposed Open Science model.

URL : Scientific production on data repositories and open science published in the Web of Science database: Methodi Ordinatio and content analysis

DOI : https://doi.org/10.1590/2318-0889202537e2513075

Catégories
EN

A decentralized future for the open-science databases

Authors : Gaurav Sharma, Viorel Munteanu, Nika Mansouri Ghiasi, Jineta Banerjee, Susheel Varma, Luca Foschini, Kyle Ellrott, Onur Mutlu, Dumitru Ciorbă, Roel A. Ophoff, Viorel Bostan, Christopher E Mason, Jason H. Moore, Despoina Sousoni, Arunkumar Krishnan, Christopher E. Mason, Mihai Dimian, Gustavo Stolovitzky, Fabio G. Liberante, Taras K. Oleksyk, Serghei Mangul

Continuous and reliable access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research. Centralized repositories, though widely used, are vulnerable to single points of failure arising from cyberattacks, technical faults, natural disasters, or funding and political uncertainties.

This can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. Centralizing essential scientific resources in a single geopolitical or institutional hub is inherently dangerous, as any disruption can paralyze diverse ongoing research.

The rapid acceleration of data generation, combined with an increasingly volatile global landscape, necessitates a critical re-evaluation of the sustainability of centralized models. Implementing federated and decentralized architectures presents a compelling and future-oriented pathway to substantially strengthen the resilience of scientific data infrastructures, thereby mitigating vulnerabilities and ensuring the long-term integrity of data.

Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, FAIR, and sustainable scientific data stewardship. Such an approach offers a significant reduction in exposure to governance instability, infrastructural fragility, and funding volatility, and also fosters fairness and global accessibility.

The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come.

DOI : https://doi.org/10.48550/arXiv.2509.19206

Catégories
EN

Open science in Spain: Influence of personal and contextual factors on deposit patterns

Author

Background

This study investigates factors influencing the deposit of academic publications and research data in open access repositories by Spanish researchers.

Methods

Using survey data from a sample of Spanish academics, the research examines the impact of personal attributes (e.g., gender, age, knowledge of open science) and contextual variables (e.g., academic discipline, institutional type) on deposit behaviours. Quantitative methods, including chi-square tests and regression analysis, reveal significant associations between knowledge of open science and deposit practices.

Results

Researchers familiar with open science principles were more likely to deposit multiple versions of articles and datasets, albeit with varying intensity. Key findings highlight disciplinary and institutional differences: researchers in Life Sciences and Experimental Sciences showed higher engagement with both article and data deposits, whereas Health Sciences lagged. Gender differences were also observed, with male researchers depositing articles and datasets more frequently than their female counterparts, though age showed limited impact. Public institutions exhibited lower data deposit rates despite mandates supporting open access.

Conclusions

The study underscores the need for tailored policies, including awareness campaigns, infrastructure investment, and discipline-specific strategies, to promote equitable and widespread adoption of open science practices. Findings contribute to understanding open science implementation, emphasizing the interplay of individual, institutional, and systemic factors.

URL : Open science in Spain: Influence of personal and contextual factors on deposit patterns

DOI : https://doi.org/10.12688/f1000research.160207.1

 

Catégories
EN

FAIRness of Research Data in the European Humanities Landscape

Authors : Ljiljana Poljak Bilić, Kristina Posavec

This paper explores the landscape of research data in the humanities in the European context, delving into their diversity and the challenges of defining and sharing them. It investigates three aspects: the types of data in the humanities, their representation in repositories, and their alignment with the FAIR principles (Findable, Accessible, Interoperable, Reusable).

By reviewing datasets in repositories, this research determines the dominant data types, their openness, licensing, and compliance with the FAIR principles. This research provides important insight into the heterogeneous nature of humanities data, their representation in the repository, and their alignment with FAIR principles, highlighting the need for improved accessibility and reusability to improve the overall quality and utility of humanities research data.

URL : FAIRness of Research Data in the European Humanities Landscape

DOI : https://doi.org/10.3390/publications12010006

Catégories
EN

Building a Trustworthy Data Repository: CoreTrustSeal Certification as a Lens for Service Improvements

Authors : Cara Key, Clara Llebot, Michael Boock

Objective

The university library aims to provide university researchers with a trustworthy institutional repository for sharing data. The library sought CoreTrustSeal certification in order to measure the quality of data services in the institutional repository, and to promote researchers’ confidence when depositing their work.

Methods

The authors served on a small team of library staff who collaborated to compose the certification application. They describe the self-assessment process, as they iterated through cycles of compiling information and responding to reviewer feedback.

Results

The application team gained understanding of data repository best practices, shared knowledge about the institutional repository, and identified areas of service improvements necessary to meet certification requirements. Based on the application and feedback, the team took measures to enhance preservation strategies, governance, and public-facing policies and documentation for the repository.

Conclusions

The university library gained a better understanding of top-notch data services and measurably improved these services by pursuing and obtaining CoreTrustSeal certification.

URL : Building a Trustworthy Data Repository: CoreTrustSeal Certification as a Lens for Service Improvements

DOI : https://doi.org/10.7191/jeslib.761