« Wikipedia might possibly be the best-developed attempt thus far of the enduring quest to gather all human knowledge in one place. Its accomplishments in this regard have made it an irresistible point of inquiry for researchers from various fields of knowledge. A decade of research has thrown light on many aspects of the Wikipedia community, its processes, and content. However, due to the variety of the fields inquiring about Wikipedia and the limited synthesis of the extensive research, there is little consensus on many aspects of Wikipedia’s content as an encyclopedic collection of human knowledge. This study addresses the issue by systematically reviewing 110 peer-reviewed publications on Wikipedia content, summarizing the current findings, and highlighting the major research trends. Two major streams of research are identified: the quality of Wikipedia content (including comprehensiveness, currency, readability and reliability) and the size of Wikipedia. Moreover, we present the key research trends in terms of the domains of inquiry, research design, data source, and data gathering methods. This review synthesizes scholarly understanding of Wikipedia content and paves the way for future studies. »
Étiquette : Scholarly Publishing
« La diffusion de la revue scientifique dans les domaines STM s’inscrit aujourd’hui dans un contexte où cohabitent quatre modèles représentés par des couleurs : le « Blanc » pour le payant, le « Vert » pour les archives ouvertes, le « Doré » pour les revues en libre accès et enfin le « Gris » pour les données de la recherche. Ce « tableau en couleurs » est le résultat, encore mouvant de mutations qui opèrent au sein de l’édition de la revue scientifique STM et qui intègre les acteurs du Web et de la communication.
Ces couleurs interviennent dans les tensions qui régissent les relations et les stratégies entre anciens et nouveaux acteurs. L’objet de cet article est double. Il rend compte des logiques socio-économiques qui construisent les stratégies et positionnements des acteurs pour renouveler la sous-filière de l’édition de la revue scientifique. Dans le même temps, il propose le cadre d’analyse des industries culturelles, capable de prendre en charge l’intelligibilité des mouvements qui opèrent entre industries de l’information scientifique, industries de la communication et médias sociaux. »
URL : http://lesenjeux.u-grenoble3.fr/2014/04-Boukacem/index.html
« Recent proposals for creating digital scholarly editions (DSEs) through the crowdsourcing of transcriptions and collaborative scholarship, for the establishment of national repositories of digital humanities data, and for the referencing, sharing, and storage of DSEs, have underlined the need for greater data interoperability. The TEI Guidelines have tried to establish standards for encoding transcriptions since 1988. However, because the choice of tags is guided by human interpretation, TEI-XML encoded files are in general not interoperable. One way to fix this problem may be to break down the current all-in-one approach to encoding so that DSEs can be specified instead by a bundle of separate resources that together offer greater interoperability: plain text versions, markup, annotations, and metadata. This would facilitate not only the development of more general software for handling DSEs, but also enable existing programs that already handle these kinds of data to function more efficiently. »
URL : Towards an Interoperable Digital Scholarly Edition
DOI : 10.4000/jtei.979
« Making available and archiving scientific results is for the most part still considered the task of classical publishing companies, despite the fact that classical forms of publishing centered around printed narrative articles no longer seem well-suited in the digital age. In particular, there exist currently no efficient, reliable, and agreed-upon methods for publishing scientific datasets, which have become increasingly important for science. Here we propose to design scientific data publishing as a Web-based bottom-up process, without top-down control of central authorities such as publishing companies. We present a protocol and a server network to decentrally store and archive data in the form of nanopublications, an RDF-based format to represent scientific data with formal semantics. We show how this approach allows researchers to produce, publish, retrieve, address, verify, and recombine datasets and their individual nanopublications in a reliable and trustworthy manner, and we argue that this architecture could be used for the Semantic Web in general. Our evaluation of the current small network shows that this system is efficient and reliable, and we discuss how it could grow to handle the large amounts of structured data that modern science is producing and consuming. »
« As open-access (OA) publishing funded by article-processing charges (APCs) becomes more widely accepted, academic institutions need to be aware of the ‘total cost of publication’, comprising subscription costs plus APCs and additional administration costs. This study analyses data from 23 UK institutions covering the period 2007 to 2014 modelling the total cost of publication (TCP). It shows a clear rise in centrally-managed APC payments from 2012 onwards, with payments projected to increase further. As well as evidencing the growing availability and acceptance of OA publishing, these trends reflect particular UK policy developments and funding arrangements intended to accelerate the move towards OA publishing (‘Gold’ OA). Whilst the mean value of APCs has been relatively stable, there was considerable variation in APC prices paid by institutions since 2007. In particular, ‘hybrid’ subscription/OA journals were consistently more expensive than fully-OA journals. Most APCs were paid to large ‘traditional’ commercial publishers who also received considerable subscription income. New administrative costs reported by institutions varied considerably. The total cost of publication modelling shows that APCs are now a significant part of the TCP for academic institutions, in 2013 already constituting an average of 10% of the TCP (excluding administrative costs). »
« Systematically evaluating scientific literature is a time consuming endeavor that requires hours of coding and rating. Here, we describe a method to distribute these tasks across a large group through online crowdsourcing. Using Amazon’s Mechanical Turk, crowdsourced workers (microworkers) completed four groups of tasks to evaluate the question, “Do nutrition-obesity studies with conclusions concordant with popular opinion receive more attention in the scientific community than do those that are discordant?” 1) Microworkers who passed a qualification test (19% passed) evaluated abstracts to determine if they were about human studies investigating nutrition and obesity. Agreement between the first two raters’ conclusions was moderate (κ = 0.586), with consensus being reached in 96% of abstracts. 2) Microworkers iteratively synthesized free-text answers describing the studied foods into one coherent term. Approximately 84% of foods were agreed upon, with only 4 and 8% of ratings failing manual review in different steps. 3) Microworkers were asked to rate the perceived obesogenicity of the synthesized food terms. Over 99% of responses were complete and usable, and opinions of the microworkers qualitatively matched the authors’ expert expectations (e.g., sugar-sweetened beverages were thought to cause obesity and fruits and vegetables were thought to prevent obesity). 4) Microworkers extracted citation counts for each paper through Google Scholar. Microworkers reached consensus or unanimous agreement for all successful searches. To answer the example question, data were aggregated and analyzed, and showed no significant association between popular opinion and attention the paper received as measured by Scimago Journal Rank and citation counts. Direct microworker costs totaled $221.75, (estimated cost at minimum wage: $312.61). We discuss important points to consider to ensure good quality control and appropriate pay for microworkers. With good reliability and low cost, crowdsourcing has potential to evaluate published literature in a cost-effective, quick, and reliable manner using existing, easily accessible resources. »
URL : Using Crowdsourcing to Evaluate Published Scientific Literature: Methods and Example
DOI: 10.1371/journal.pone.0100647
« In this paper, we examine the evolution of the impact of non-elite journals. We attempt to answer two questions. First, what fraction of the top-cited articles are published in non-elite journals and how has this changed over time. Second, what fraction of the total citations are to non-elite journals and how has this changed over time.
We studied citations to articles published in 1995-2013. We computed the 10 most-cited journals and the 1000 most-cited articles each year for all 261 subject categories in Scholar Metrics. We marked the 10 most-cited journals in a category as the elite journals for the category and the rest as non-elite.
There are two conclusions from our study. First, the fraction of top-cited articles published in non-elite journals increased steadily over 1995-2013. While the elite journals still publish a substantial fraction of high-impact articles, many more authors of well-regarded papers in diverse research fields are choosing other venues.
The number of top-1000 papers published in non-elite journals for the representative subject category went from 149 in 1995 to 245 in 2013, a growth of 64%. Looking at broad research areas, 4 out of 9 areas saw at least one-third of the top-cited articles published in non-elite journals in 2013. For 6 out of 9 areas, the fraction of top-cited papers published in non-elite journals for the representative subject category grew by 45% or more.
Second, now that finding and reading relevant articles in non-elite journals is about as easy as finding and reading articles in elite journals, researchers are increasingly building on and citing work published everywhere. Considering citations to all articles, the percentage of citations to articles in non-elite journals went from 27% in 1995 to 47% in 2013. Six out of nine broad areas had at least 50% of citations going to articles published in non-elite journals in 2013. »