More than data repositories: perceived information needs for the development of social sciences and humanities research infrastructures

Authors : Anna Sendra, Elina Late, Sanna Kumpulainen

Introduction

The digitalization of social sciences and humanities research necessitates research infrastructures. However, this transformation is still incipient, highlighting the need to better understand how to successfully support data-intensive research.

Method

Starting from a case study of building a national infrastructure for conducting data-intensive research, this study aims to understand the information needs of digital researchers regarding the facility and explore the importance of evaluation in its development.

Analysis

Thirteen semi-structured interviews with social sciences and humanities scholars and computer and data scientists processed through a thematic analysis revealed three themes (developing a research infrastructure, needs and expectations of the research infrastructure, and an approach to user feedback and user interactions).

Results

Findings reveal that developing an infrastructure for conducting data-intensive research is a complicated task influenced by contrasting information needs between social sciences and humanities scholars and computer and data scientists, such as the demand for increased support of the former. Findings also highlight the limited role of evaluation in its creation.

Conclusions

The development of infrastructures for conducting data-intensive research requires further discussion that particularly considers the disciplinary differences between social sciences and humanities scholars and computer and data scientists. Suggestions on how to better design this kind of facilities are also raised.

URL : More than data repositories: perceived information needs for the development of social sciences and humanities research infrastructures

DOI : https://doi.org/10.47989/ir284598

Establishing an early indicator for data sharing and reuse

Authors : Agata Piękniewska, Laurel L. Haak, Darla Henderson, Katherine McNeill, Anita Bandrowski, Yvette Seger

Funders, publishers, scholarly societies, universities, and other stakeholders need to be able to track the impact of programs and policies designed to advance data sharing and reuse. With the launch of the NIH data management and sharing policy in 2023, establishing a pre-policy baseline of sharing and reuse activity is critical for the biological and biomedical community.

Toward this goal, we tested the utility of mentions of research resources, databases, and repositories (RDRs) as a proxy measurement of data sharing and reuse. We captured and processed text from Methods sections of open access biological and biomedical research articles published in 2020 and 2021 and made available in PubMed Central.

We used natural language processing to identify text strings to measure RDR mentions. In this article, we demonstrate our methodology, provide normalized baseline data sharing and reuse activity in this community, and highlight actions authors and publishers can take to encourage data sharing and reuse practices.

URL : Establishing an early indicator for data sharing and reuse

DOI : https://doi.org/10.1002/leap.1586

Enquête quantitative sur les pratiques et les besoins des chercheurs sur la gestion des données de la recherche, algorithmes et codes sources dans les établissements du site toulousain

Authors : Danielle Brunet, Soraya Demay, Pierre Diaz, Borbala Goncz, Laure Leclerc, Flora Poupinot, Sibilla Michelle

Le Comité de réflexion pour le partage et la valorisation des données de la recherche et la coordination de la Science Ouverte (CéSO) de l’Université de Toulouse a réalisé une enquête quantitative sur la gestion des données de la recherche, algorithmes et codes sources.

Adressée à l’ensemble de la communauté scientifique du site toulousain, son objectif était de produire un état des lieux des pratiques, des connaissances et des besoins des chercheurs en matière de gestion des données de la recherche. Les résultats permettront de préciser l’offre de services proposée sur le site toulousain.

Cette enquête concerne les établissements membres de l’Université de Toulouse ainsi que les organismes de recherche partenaires : Université Toulouse Capitole, Université Toulouse – Jean Jaurès, Université Toulouse III – Paul Sabatier, Institut national polytechnique de Toulouse (Toulouse INP), Institut national des sciences appliquées de Toulouse (INSA Toulouse), Institut supérieur de l’aéronautique et de l’espace (ISAE-SUPAERO), Institut national universitaire Champollion (INU Champollion), École nationale de l’aviation civile (ENAC), École nationale d’ingénieurs de Tarbes (ENIT), École nationale supérieure d’architecture de Toulouse (ENSA Toulouse), École nationale vétérinaire de Toulouse (ENVT), École nationale supérieure de formation de l’enseignement agricole (ENSFEA), Institut catholique d’arts et métiers (ICAM), École nationale supérieure des mines d’Albi-Carmaux (IMT Mines d’Albi), Toulouse Business School (TBS), Centre national d’études spatiales (CNES), Centre national de la recherche scientifique (CNRS), Institut national de recherche pour l’agriculture, l’alimentation et l’environnement (INRAE), Institut national de l’a santé et de la recherche médicale (Inserm), Institut de recherche pour le développement (IRD) ; Office national d’études et de recherche aérospatiales (Onera), Météo-France.

URL : Enquête quantitative sur les pratiques et les besoins des chercheurs sur la gestion des données de la recherche, algorithmes et codes sources dans les établissements du site toulousain

Original location : https://ut3-toulouseinp.hal.science/hal-04262708v1/

Long-term availability of data associated with articles in PLOS ONE

Author : Lisa M. Federer

The adoption of journal policies requiring authors to include a Data Availability Statement has helped to increase the availability of research data associated with research articles. However, having a Data Availability Statement is not a guarantee that readers will be able to locate the data; even if provided with an identifier like a uniform resource locator (URL) or a digital object identifier (DOI), the data may become unavailable due to link rot and content drift. :

To explore the long-term availability of resources including data, code, and other digital research objects associated with papers, this study extracted 8,503 URLs and DOIs from a corpus of nearly 50,000 Data Availability Statements from papers published in PLOS ONE between 2014 and 2016.

These URLs and DOIs were used to attempt to retrieve the data through both automated and manual means. Overall, 80% of the resources could be retrieved automatically, compared to much lower retrieval rates of 10–40% found in previous papers that relied on contacting authors to locate data.

Because a URL or DOI might be valid but still not point to the resource, a subset of 350 URLs and 350 DOIs were manually tested, with 78% and 98% of resources, respectively, successfully retrieved.

Having a DOI and being shared in a repository were both positively associated with availability. Although resources associated with older papers were slightly less likely to be available, this difference was not statistically significant, suggesting that URLs and DOIs may be an effective means for accessing data over time.

These findings point to the value of including URLs and DOIs in Data Availability Statements to ensure access to data on a long-term basis.

URL : Long-term availability of data associated with articles in PLOS ONE

DOI : https://doi.org/10.1371/journal.pone.0272845

Making Mathematical Research Data FAIR: A Technology Overview

Authors : Tim Conrad, Eloi Ferrer, Daniel Mietchen, Larissa Pusch, Johannes Stegmuller, Moritz Schubotz

The sharing and citation of research data is becoming increasingly recognized as an essential building block in scientific research across various fields and disciplines. Sharing research data allows other researchers to reproduce results, replicate findings, and build on them. Ultimately, this will foster faster cycles in knowledge generation.

Some disciplines, such as astronomy or bioinformatics, already have a long history of sharing data; many others do not. The current landscape of so-called research data repositories is diverse. This review aims to perform a technology review on existing data repositories/portals with a focus on mathematical research data.

URL : Making Mathematical Research Data FAIR: A Technology Overview

Original location: https://arxiv.org/abs/2309.11829

Données ouvertes liées et recherche historique : un changement de paradigme

Auteur/Author : Francesco Beretta

Dans le contexte de la transition numérique, le Web sémantique et les données ouvertes liées (linked open data [LOD], en anglais) jouent un rôle de plus en plus central, car ils permettent de construire des « graphes d’information » (knowledge graphs, en anglais) reliant l’ensemble des ressources du Web.

Ce phénomène interroge les sciences historiques et soulève la question d’un changement de paradigme. Après avoir précisé ce qu’il faut entendre par « données », l’article analyse la place qu’elles occupent dans le processus de production du savoir.

Il présente les principales composantes du changement de paradigme, en particulier le potentiel des LOD et d’une sémantique robuste en tant que véhicules d’une information factuelle de qualité, intelligible et réutilisable. S’ensuit une présentation des projets d’infrastructure réalisés au sein du Laboratoire de recherche historique Rhône-Alpes (Larhra) : symogih.org, ontome.net, geovistory.org.

Leur but est de faciliter la transition numérique grâce à un outillage construit en cohérence avec l’épistémologie des sciences historiques et de contribuer à la réalisation d’un « graphe d’information » disciplinaire.

URL : Données ouvertes liées et recherche historique : un changement de paradigme

DOI : https://doi.org/10.4000/revuehn.3349

It Takes a Researcher to Know a Researcher: Academic Librarian Perspectives Regarding Skills and Training for Research Data Support in Canada

Author : Alisa B. Rod

Objective

This empirical study aims to contribute qualitative evidence on the perspectives of data-related librarians regarding the necessary skills, education, and training for these roles in the context of Canadian academic libraries.

A second aim of this study is to understand the perspectives of data-related librarians regarding the specific role of the MLIS in providing relevant training and education. The definition of a data-related librarian in this study includes any librarian or professional who has a conventional title related to a field of data librarianship (i.e., research data management, data services, GIS, data visualization, data science) or any other librarian or professional whose duties include providing data-related services within an academic institution.

Methods

This study incorporates in-depth qualitative empirical evidence in the form of 12 semi-structured interviews of data-related librarians to investigate first-hand perspectives on the necessary skills required for such positions and the mechanisms for acquiring and maintaining such skills.

Results

The interviews identified four major themes related to the skills required for library-related data services positions, including the perceived importance of experience conducting original research, proficiency in computational coding and quantitative methods, MLIS-related skills such as understanding metadata, and the ability to learn new skills quickly on the job.

Overall, the implication of this study regarding the training from MLIS programs concerning data-related librarianship is that although expertise in metadata, documentation, and information management are vital skills for data-related librarians, the MLIS is increasingly less competitive compared with degree programs that offer a greater emphasis on practical experience working with different types of data in a research context and implementing a variety of methodological approaches.

Conclusion

This study demonstrates that an in-depth qualitative portrait of data-related librarians within a national academic ecosystem provides valuable new insights regarding the perceived importance of conducting original empirical research to succeed in these roles.

URL : It Takes a Researcher to Know a Researcher: Academic Librarian Perspectives Regarding Skills and Training for Research Data Support in Canada

DOI : https://doi.org/10.18438/eblip30297