Authors : Danielle R. Kirsch, Isaac Wink
Introduction: Scholarly communities are experiencing increased emphasis on research data sharing and reuse. Given the range of available data repositories and variability in the use of persistent identifiers (PIDs) for individuals and institutions, aggregating datasets by researchers at a specific institution is a significant challenge. We developed a reproducible workflow to evaluate 1) data repository use by researchers at our institutions, 2) metrics of reuse, and 3) quality of metadata records.
Methods: We used the DataCite REST API to locate metadata records for datasets with creator affiliation names that matched our institutions. We performed substantial cleaning and deduplication before comparing repository use, citations, usage metrics, and metadata completeness.
Results: The most common data repositories were Dryad, Harvard Dataverse, figshare, Zenodo, ICPSR, and one institution’s institutional repository. View, download, and citation counts were available from a limited number of repositories, with some discrepancies between different citation reporting methods. PIDs were more frequently used for authors and affiliations than funders, and the inclusion of PIDs varied across repositories, with Dryad being the most consistent.
Discussion & Conclusion: Although broad trends in repository and PID use were similar between institutions, our analysis also surfaced examples that illustrate inconsistencies in metadata across repositories. Variable implementation of the DataCite metadata schema requires a significant amount of data cleaning to obtain meaningful results. Even then, incomplete metadata makes some datasets impossible to locate. Repositories, funders, researchers, and institutional open data advocates must coordinate to create complete and usable metadata that integrates datasets into the scholarship ecosystem.