Catégories
EN

Cluster Analysis of Open Research Data: A Case for Replication Metadata

Author : Ana Trisovic

Research data are often released upon journal publication to enable result verification and reproducibility. For that reason, research dissemination infrastructures typically support diverse datasets coming from numerous disciplines, from tabular data and program code to audio-visual files. Metadata, or data about data, is critical to making research outputs adequately documented and FAIR.

Aiming to contribute to the discussions on the development of metadata for research outputs, I conducted an exploratory analysis to determine how research datasets cluster based on what researchers organically deposit together. I use the content of over 40,000 datasets from the Harvard Dataverse research data repository as my sample for the cluster analysis.

I find that the majority of the clusters are formed by single-type datasets, while in the rest of the sample, no meaningful clusters can be identified. For the result interpretation, I use the metadata standard employed by DataCite, a leading organization for documenting a scholarly record, and map existing resource types to my results.

About 65% of the sample can be described with a single-type metadata (such as Dataset, Software orReport), while the rest would require aggregate metadata types. Though DataCite supports an aggregate type such as a Collection, I argue that a significant number of datasets, in particular those containing both data and code files (about 20% of the sample), would be more accurately described as a Replication resource metadata type. Such resource type would be particularly useful in facilitating research reproducibility.

URL : Cluster Analysis of Open Research Data: A Case for Replication Metadata

DOI : https://doi.org/10.2218/ijdc.v17i1.833

Catégories
EN

The Role of Metadata and Vocabulary Standards in Enabling Scientific Data Interoperability: A Study of Earth System Science Data Facilities

Authors : Matthew S. Mayernik, Yauheniya Liapich

Objective

Journal publishers within many sciences are increasingly expecting data to be deposited into repositories that support the FAIR principles. Data repositories are thus needing to determine what implications the FAIR principles have on their existing services and systems. Metadata standards and controlled vocabularies are specifically called out as core components of the FAIR principles related to interoperability.

Methods

This paper looks specifically at the ways that metadata standards and controlled vocabularies are used by Earth system science data repositories. Data sets from 55 data facilities were examined to determine which metadata standards and controlled subject / keyword vocabularies were used.

Results

The findings indicate that only the ISO 19115:2003 and DataCite metadata standards are used by more than 40% of the data facilities, and the NASA Global Change Master Directory (GCMD) keywords are the only keyword vocabulary of broad use within this community.

Conclusions

These findings raise questions about the extent to which metadata standards and keyword vocabularies can facilitate interoperability beyond narrow sub-sections of the data facility communities. This study also points to systematic challenges related to migration to new standards.

URL : The Role of Metadata and Vocabulary Standards in Enabling Scientific Data Interoperability: A Study of Earth System Science Data Facilities

DOI : https://doi.org/10.7191/jeslib.619

Catégories
EN

Open Editors: A dataset of scholarly journals’ editorial board positions

Authors : Andreas Nishikawa-Pacher, Tamara Heck, Kerstin Schoch

Editormetrics analyses the role of editors of academic journals and their impact on the scientific publication system. Such analyses would best rely on open, structured, and machine-readable data about editors and editorial boards, which still remains rare.

To address this shortcoming, the project Open Editors collects data about academic journal editors on a large scale and structures them into a single dataset. It does so by scraping the websites of 7,352 journals from 26 publishers (including predatory ones), thereby structuring publicly available information (names, affiliations, editorial roles, ORCID etc.) about 594,580 researchers.

The dataset shows that journals and publishers are immensely heterogeneous in terms of editorial board sizes, regional diversity, and editorial role labels. All codes and data are made available at Zenodo, while the result is browsable at a dedicated website (https://openeditors.ooir.org).

This dataset carries implications for both practical purposes of research evaluation and for meta-scientific investigations into the landscape of scholarly publications, and allows for critical inquiries regarding the representation of diversity and inclusivity across academia.

URL : Open Editors: A dataset of scholarly journals’ editorial board positions

DOI : https://doi.org/10.1093/reseval/rvac037

Catégories
EN

metajelo: A metadata package for journals to support external linked objects

Authors : Lars Vilhuber, Carl Lagoze

We propose a metadata package that is intended to provide academic journals with a lightweight means of registering, at the time of publication, the existence and disposition of supplementary materials. Information about the supplementary materials is, in most cases, critical for the reproducibility and replicability of scholarly results.

In many instances, these materials are curated by a third party, which may or may not follow developing standards for the identification and description of those materials. As such, the vocabulary described here complements existing initiatives that specify vocabularies to describe the supplementary materials or the repositories and archives in which they have been deposited.

Where possible, it reuses elements of relevant other vocabularies, facilitating coexistence with them. Furthermore, it provides an “at publication” record of reproducibility characteristics of a particular article that has been selected for publication.

The proposed metadata package documents the key characteristics that journals care about in the case of supplementary materials that are held by third parties: existence, accessibility, and permanence. It does so in a robust, time-invariant fashion at the time of publication, when the editorial decisions are made.

It also allows for better documentation of less accessible (non-public data), by treating it symmetrically from the point of view of the journal, therefore increasing the transparency of what up until now has been very opaque.

URL : metajelo: A metadata package for journals to support external linked objects

DOI : https://doi.org/10.2218/ijdc.v16i1.600

Catégories
FR

Métadonnées pour la science ouverte : rôle et action des bibliothèques et des professionnels de l’information scientifique et technique

Auteur/Author : Vincent Richard

À l’heure de la science ouverte et du web sémantique, les professionnels de l’information scientifique et technique font face à de nombreux défis dans la gestion des métadonnées des productions académiques, qui constituent le socle de la recherche d’information scientifique et de l’accès aux publications au niveau mondial : traitement des données des éditeurs, archives ouvertes, édition en libre accès, gestion et politique des registres d’identifiants, bases de données de publications, devenir des bases de citations, enjeu de la création d’un index mondial ouvert des publications, avancée en matière de graphes de données liées… Nous en présentons les enjeux stratégiques, institutionnels et techniques.

Nous tentons également d’identifier les évolutions en cours et souhaitables en termes de compétences, d’outils technologiques et de gouvernance, et de définir le rôle que peuvent jouer les bibliothèques pour renforcer la qualité et l’ouverture des métadonnées.

URL : Métadonnées pour la science ouverte : rôle et action des bibliothèques et des professionnels de l’information scientifique et technique

Original location : https://www.enssib.fr/bibliotheque-numerique/notices/70132-metadonnees-pour-la-science-ouverte-role-et-action-des-bibliotheques-et-des-professionnels-de-l-information-scientifique-et-technique

Catégories
EN

A Controlled Vocabulary and Metadata Schema for Materials Science Data Discovery

Authors : Andrea Medina-Smith, Chandler A. Becker, Raymond L. Plante, Laura M. Bartolo, Alden Dima, James A. Warren, Robert J. Hanisch

The International Materials Resource Registries (IMRR) working group of the Research Data Alliance (RDA) was created to spur initial development of a federated registry system to allow for easier discovery and access to materials data.

As part of this effort, a controlled vocabulary and metadata schema were developed with contributions from members of the working group and other experts. Here we describe the process, the resulting vocabulary and XML schema, and lessons learned in the development and use of the schema.

URL : A Controlled Vocabulary and Metadata Schema for Materials Science Data Discovery

DOI : http://doi.org/10.5334/dsj-2021-018

Catégories
EN

Affiliation Information in DataCite Dataset Metadata: a Flemish Case Study

Author/Auteur : Niek Van Wettere

This article aims to evaluate how and to what extent metadata of datasets indexed in DataCite offer clear human- or machine-readable information that enables the research data to be linked to a particular research institution.

Two main pathways are explored. First, researchers can encode their affiliation information at the moment of data submission. This can be done by means of free-text metadata fields or via the inclusion of identifiers such as GRID/ROR and ORCID. Second, affiliation information can be traced indirectly through linking between a dataset and associated publications, given that the metadata of publications is often more explicit about affiliation information than the metadata of datasets.

Both pathways of affiliation information encoding are evaluated on the basis of metadata pertaining to datasets created at the five Flemish universities. It is shown that good practices such as encoding of affiliation information in a dedicated metadata field or inclusion of ORCID in the metadata are on the rise, but could be expanded further.

Finally, the establishment of links between datasets and related publications is often lacking in dataset metadata, although there are important differences between data repositories, as is also demonstrated in a more data-intensive follow-up analysis based on random samples of metadata records.

It is important that data repositories address this issue by providing a metadata field clearly dedicated to associated publications, prominently displayed on the landing page of the dataset.

URL : Affiliation Information in DataCite Dataset Metadata: a Flemish Case Study

DOI : http://doi.org/10.5334/dsj-2021-013