Catégories
EN

Invasion of the journal snatchers: How indexed journals are falling into questionable hands

Authors: Alberto Martín-Martín, Emilio Delgado López-Cózar

In recent years, a substantial number of established journals have received buyout offers from obscure entities, with some journals being acquired. Despite mounting circumstantial evidence of irregular behaviour exhibited by these journals post-acquisition, comprehensive analyses on this matter are lacking. To address this gap, this article examines the practices of Oxbridge Publishing House Ltd., a company registered in the UK in 2022.

Through an analysis of publicly available documentation, it becomes apparent that this entity is part of a complex network of recently established companies. Since 2020 this network has acquired, with the help of intermediary firms, at least 36 scholarly journals originally published in countries such Spain (7), United Kingdom (7), USA (5), India (4), Turkey (4), among others.

Targeting journals indexed in prestigious scientific databases like Web of Science and Scopus, many of these journals see significant transformations upon acquisition, such as the introduction or substantial escalation of publication fees, often coupled with increases in publication volumes. This increase stems from a surge in contributions originating outside the journal’s original academic community. Their disregard for proper publishing standards is evident in their widespread use of fake DOIs or the appropriation of DOIs from unrelated documents.

Drawing parallels to the film Invasion of the Body Snatchers, we refer to journals caught in this predicament as pod journals. This type of predatory publishing practice not only contributes to over-publication but also disenfranchises legitimate academic communities and poses a threat to academic bibliodiversity.

URL : Invasion of the journal snatchers: How indexed journals are falling into questionable hands

DOI : https://doi.org/10.5281/zenodo.14766415

Catégories
EN

Why are these publications missing? Uncovering the reasons behind the exclusion of documents in free-access scholarly databases

Authors : Lorena Delgado-QuirósIsidro F. AguilloAlberto Martín-MartínEmilio Delgado López-CózarEnrique Orduña-MaleaJosé Luis Ortega

This study analyses the coverage of seven free-access bibliographic databases (Crossref, Dimensions—non-subscription version, Google Scholar, Lens, Microsoft Academic, Scilit, and Semantic Scholar) to identify the potential reasons that might cause the exclusion of scholarly documents and how they could influence coverage.

To do this, 116 k randomly selected bibliographic records from Crossref were used as a baseline. API endpoints and web scraping were used to query each database. The results show that coverage differences are mainly caused by the way each service builds their databases.

While classic bibliographic databases ingest almost the exact same content from Crossref (Lens and Scilit miss 0.1% and 0.2% of the records, respectively), academic search engines present lower coverage (Google Scholar does not find: 9.8%, Semantic Scholar: 10%, and Microsoft Academic: 12%). Coverage differences are mainly attributed to external factors, such as web accessibility and robot exclusion policies (39.2%–46%), and internal requirements that exclude secondary content (6.5%–11.6%).

In the case of Dimensions, the only classic bibliographic database with the lowest coverage (7.6%), internal selection criteria such as the indexation of full books instead of book chapters (65%) and the exclusion of secondary content (15%) are the main motives of missing publications.

URL : Why are these publications missing? Uncovering the reasons behind the exclusion of documents in free-access scholarly databases

DOI : https://doi.org/10.1002/asi.24839

Catégories
EN

Large coverage fluctuations in Google Scholar: a case study

Authors : Alberto Martín-Martín, Emilio Delgado López-Cózar

Unlike other academic bibliographic databases, Google Scholar intentionally operates in a way that does not maintain coverage stability: documents that stop being available to Google Scholar’s crawlers are removed from the system.

This can also affect Google Scholar’s citation graph (citation counts can decrease). Furthermore, because Google Scholar is not transparent about its coverage, the only way to directly observe coverage loss is through regular monitorization of Google Scholar data.

Because of this, few studies have empirically documented this phenomenon. This study analyses a large decrease in coverage of documents in the field of Astronomy and Astrophysics that took place in 2019 and its subsequent recovery, using longitudinal data from previous analyses and a new dataset extracted in 2020.

Documents from most of the larger publishers in the field disappeared from Google Scholar despite continuing to be available on the Web, which suggests an error on Google Scholar’s side. Disappeared documents did not reappear until the following index-wide update, many months after the problem was discovered.

The slowness with which Google Scholar is currently able to resolve indexing errors is a clear limitation of the platform both for literature search and bibliometric use cases.

URL : https://arxiv.org/abs/2102.07571

Catégories
EN

Google Scholar, Microsoft Academic, Scopus, Dimensions, Web of Science, and OpenCitations’ COCI: a multidisciplinary comparison of coverage via citations

Authors : Alberto Martín-Martín, Mike Thelwall, Enrique Orduna-Malea, Emilio Delgado López-Cózar

New sources of citation data have recently become available, such as Microsoft Academic, Dimensions, and the OpenCitations Index of CrossRef open DOI-to-DOI citations (COCI). Although these have been compared to the Web of Science Core Collection (WoS), Scopus, or Google Scholar, there is no systematic evidence of their differences across subject categories.

In response, this paper investigates 3,073,351 citations found by these six data sources to 2,515 English-language highly-cited documents published in 2006 from 252 subject categories, expanding and updating the largest previous study. Google Scholar found 88% of all citations, many of which were not found by the other sources, and nearly all citations found by the remaining sources (89–94%).

A similar pattern held within most subject categories. Microsoft Academic is the second largest overall (60% of all citations), including 82% of Scopus citations and 86% of WoS citations. In most categories, Microsoft Academic found more citations than Scopus and WoS (182 and 223 subject categories, respectively), but had coverage gaps in some areas, such as Physics and some Humanities categories. After Scopus, Dimensions is fourth largest (54% of all citations), including 84% of Scopus citations and 88% of WoS citations.

It found more citations than Scopus in 36 categories, more than WoS in 185, and displays some coverage gaps, especially in the Humanities. Following WoS, COCI is the smallest, with 28% of all citations. Google Scholar is still the most comprehensive source. In many subject categories Microsoft Academic and Dimensions are good alternatives to Scopus and WoS in terms of coverage.

DOI : https://doi.org/10.1007/s11192-020-03690-4

Catégories
EN

Unbundling Open Access dimensions: a conceptual discussion to reduce terminology inconsistencies

Authors : Alberto Martín-Martín, Rodrigo Costas, Thed N. van Leeuwen, Emilio Delgado López-Cózar

The current ways in which documents are made freely accessible in the Web no longer adhere to the models established Budapest/Bethesda/Berlin (BBB) definitions of Open Access (OA). Since those definitions were established, OA-related terminology has expanded, trying to keep up with all the variants of OA publishing that are out there.

However, the inconsistent and arbitrary terminology that is being used to refer to these variants are complicating communication about OA-related issues. This study intends to initiate a discussion on this issue, by proposing a conceptual model of OA.

Our model features six different dimensions (authoritativeness, user rights, stability, immediacy, peer-review, and cost). Each dimension allows for a range of different options. We believe that by combining the options in these six dimensions, we can arrive at all the current variants of OA, while avoiding ambiguous and/or arbitrary terminology.

This model can be an useful tool for funders and policy makers who need to decide exactly which aspects of OA are necessary for each specific scenario.

URL : Unbundling Open Access dimensions: a conceptual discussion to reduce terminology inconsistencies

Alternative location : https://arxiv.org/abs/1806.05029

Catégories
EN

Google Scholar as a data source for research assessment

Authors : Emilio Delgado López-Cózar, Enrique Orduna-Malea, Alberto Martín-Martín

The launch of Google Scholar (GS) marked the beginning of a revolution in the scientific information market. This search engine, unlike traditional databases, automatically indexes information from the academic web. Its ease of use, together with its wide coverage and fast indexing speed, have made it the first tool most scientists currently turn to when they need to carry out a literature search.

Additionally, the fact that its search results were accompanied from the beginning by citation counts, as well as the later development of secondary products which leverage this citation data (such as Google Scholar Metrics and Google Scholar Citations), made many scientists wonder about its potential as a source of data for bibliometric analyses.

The goal of this chapter is to lay the foundations for the use of GS as a supplementary source (and in some disciplines, arguably the best alternative) for scientific evaluation.

First, we present a general overview of how GS works. Second, we present empirical evidences about its main characteristics (size, coverage, and growth rate). Third, we carry out a systematic analysis of the main limitations this search engine presents as a tool for the evaluation of scientific performance.

Lastly, we discuss the main differences between GS and other more traditional bibliographic databases in light of the correlations found between their citation data. We conclude that Google Scholar presents a broader view of the academic world because it has brought to light a great amount of sources that were not previously visible.

URL : https://arxiv.org/abs/1806.04435

Catégories
EN

Evidence of Open Access of scientific publications in Google Scholar: a large-scale analysis

Authors : Alberto Martín-Martín, Rodrigo Costas, Thed van Leeuwen, Emilio López-Cózar

This article uses Google Scholar (GS) as a source of data to analyse Open Access (OA) levels across all countries and fields of research. All articles and reviews with a DOI and published in 2009 or 2014 and covered by the three main citation indexes in the Web of Science (2,269,022 documents) were selected for study.

The links to freely available versions of these documents displayed in GS were collected. To differentiate between more reliable (sustainable and legal) forms of access and less reliable ones, the data extracted from GS was combined with information available in DOAJ, CrossRef, OpenDOAR, and ROAR.

This allowed us to distinguish the percentage of documents in our sample that are made OA by the publisher (23.1%, including Gold, Hybrid, Delayed, and Bronze OA) from those available as Green OA (17.6%), and those available from other sources (40.6%, mainly due to ResearchGate).

The data shows an overall free availability of 54.6%, with important differences at the country and subject category levels. The data extracted from GS yielded very similar results to those found by other studies that analysed similar samples of documents, but employed different methods to find evidence of OA, thus suggesting a relative consistency among methods.

URL : Evidence of Open Access of scientific publications in Google Scholar: a large-scale analysis

Alternative location : https://osf.io/preprints/socarxiv/k54uv/