Searcharxiv⌕ Search

arXiv subjects

Juan Pablo Alperin

Publications and source records attributed to Juan Pablo Alperin.

15 recordsLinked to original sources

Estimating global article processing charges paid to 14 publishers for open access between 2019 and 2025

This study presents estimates of the global expenditure on article processing charges (APCs) paid to 14 publishers for open access (OA) between 2019 and 2025. APCs are charged for publishing in fully OA journals (gold) and making individual articles OA in subscription journals (hybrid), but how much is paid, and for which articles, is not publicly known. We therefore curated an open dataset of publicly listed APC prices from 14 academic publishers (ACS, CUP, De Gruyter, EDP, Elsevier, Frontiers, IEEE, IOP, MDPI, OUP, PLOS, Sage, Springer Nature, and Wiley) and combined it with counts of OA articles from OpenAlex. We estimate that \$15.08 billion (in 2025 USD) was spent globally on APCs between 2019 and 2025. Adjusted for inflation, annual spending quadrupled from \$0.9 billion in 2019 to \$3.7 billion in 2025, with >85% concentrated among a few large publishers. Hybrid OA fees exceed gold fees, and the median fee paid is higher than the median price listed for both. Our approach addresses major limitations in previous efforts to estimate APC spending, offering much-needed insight into an opaque aspect of scholarly publishing, especially as transformative agreements make it more challenging to understand the costs of publishing OA.

cs.DL↗

A dataset of article processing charges from 14 scholarly publishers, 2019-2025

This paper introduces a dataset of APCs produced from the price lists of 14 large scholarly publishers between 2019 and 2025. APC price lists were downloaded from publisher websites each year as well as via Wayback Machine snapshots to retrieve fees per journal per year. The dataset includes journal metadata, APC collection method, and annual APC price list information in several currencies (USD, EUR, GBP, CHF, JPY, CAD, AUD) for 12,540 unique journals and 69,856 journal-year combinations. The dataset was generated to allow for more precise analysis of APCs and can support library collection development and scientometric analysis estimating APCs paid in gold and hybrid OA journals.

cs.DL↗

Diamond Fractures: Tracing Journal Transitions Away from Diamond Open Access

While much attention has been paid to journals transitioning toward Diamond Open Access (OA), comparatively little is known about those moving in the opposite direction. We introduce the concept of "diamond fractures"--instances in which journals that once published freely for both readers and authors subsequently abandoned that model, transitioning to subscription-based or article processing charge (APC)-funded publishing. Drawing on publisher OA portfolio records, historical APC lists, removal logs from the Directory of Open Access Journals, and a survey of journals using Open Journal Systems, we identified and characterized more than 440 journals that have undergone such fractures and analyzed how they are distributed across publisher types, disciplines, and regions; which models journals transition to and at what price points; how old journals are when they fracture; and how publication volume shifts in the period surrounding the transition. This paper reports on more than 440 confirmed fractures, the majority of which involve transitions to charging APCs, ranging from $8 to $5,300. These findings raise urgent questions about the stability of diamond OA as a publishing ecosystem. The transition to APCs--even at initially modest price points--follows well-documented warnings about the hyperinflationary tendencies of author-pays models, in which charges introduced as a pragmatic stopgap have repeatedly escalated over time. The rate of fractures are likely to intensify as APC-based publishing continues to be normalized through funder mandates and commercial OA incentives. Diamond fractures demand to be treated not as isolated institutional decisions but as a systemic risk--one that can only be addressed through sustained investment in community-governed infrastructure and funding models that reduce the conditions under which abandoning diamond OA becomes the path of least resistance.

cs.DL↗

When Editors Revolt: Characterizing Journal Declarations of Independence

When editorial boards resign from their journals and publishers and declare their independence, two competing journals can result: the original journal under a new editorial board (a "zombie" journal), and a new journal established by the departing editors (a "breakaway"). The bibliometric community saw such an event when the board of Journal of Informetrics left Elsevier to found Quantitative Science Studies. We analyzed 39 breakaway-zombie journal pairs that have formed since 1989 and their declarations of independence to understand why and how they happen. Results show that declarations of independence were motivated by concerns related to governance and business model and overwhelmingly happened at journals owned by the Big Five publishers. Breakaway editors tended to found new journals at smaller publishers and adopt diamond publishing models. These findings suggest that dissatisfaction with commercial publishing models is growing, and that community-led alternatives can motivate change.

cs.DL↗

RenoBench: A Citation Parsing Benchmark

Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, existing evaluation techniques are often not generalizable, based on synthetic data, or not publicly available. We introduce RenoBench, a public domain benchmark for citation parsing, sourced from PDFs released on four publishing ecosystems: SciELO, Redalyc, the Public Knowledge Project, and Open Research Europe. Starting from 161,000 annotated citations, we apply automated validation and feature-based sampling to produce a dataset of 10,000 citations spanning multiple languages, publication types, and platforms. We then evaluate a variety of citation parsing systems and report field-level precision and recall. Our results show strong performance from language models, particularly when fine-tuned. RenoBench enables reproducible, standardized evaluation of citation parsing systems, and provides a foundation for advancing automated citation parsing and metascientific research.

cs.DL↗

Citation Parsing and Analysis with Language Models

A key type of resource needed to address global inequalities in knowledge production and dissemination is a tool that can support journals in understanding how knowledge circulates. The absence of such a tool has resulted in comparatively less information about networks of knowledge sharing in the Global South. In turn, this gap authorizes the exclusion of researchers and scholars from the South in indexing services, reinforcing colonial arrangements that de-center and minoritize those scholars. In order to support citation network tracking on a global scale, we investigate the capacity of open-weight language models to mark up manuscript citations in an indexable format. We assembled a dataset of matched plaintext and annotated citations from preprints and published research papers. Then, we evaluated a number of open-weight language models on the annotation task. We find that, even out of the box, today's language models achieve high levels of accuracy on identifying the constituent components of each citation, outperforming state-of-the-art methods. Moreover, the smallest model we evaluated, Qwen3-0.6B, can parse all fields with high accuracy in $2^5$ passes, suggesting that post-training is likely to be effective in producing small, robust citation parsing models. Such a tool could greatly improve the fidelity of citation networks and thus meaningfully improve research indexing and discovery, as well as further metascientific research.

cs.CL↗

Evaluating Multilingual Metadata Quality in Crossref

Introduction: Scholarly research spans multiple languages, making multilingual metadata crucial for organizing and accessing knowledge across linguistic boundaries. These multilingual metadata already exist and are propagated throughout scholarly publishing infrastructure, but the extent to which they are correctly recorded, or how they affect metadata quality more broadly is little understood. Methods: Our study quantifies the prevalence of multilingual records across a sample of publisher metadata and offers an understanding of their completeness, quality, and alignment with metadata standards. Utilizing the Crossref API to generate a random sample of 519,665 journal article records, we categorize each record into four distinct language types: English monolingual, non-English monolingual, multilingual, and uncategorized. We then investigate the prevalence of programmatically-detectable errors and the prevalence of multilingual records within the sample to determine whether multilingualism influences the quality of article metadata. Results: We find that English-only records are still in the vast majority among metadata found in Crossref, but that, while non-English and multilingual records present unique challenges, they are not a source of significant metadata quality issues and, in few instances, are more complete or correct than English monolingual records. Discussion & Conclusion: Our findings contribute to discussions surrounding multilingualism in scholarly communication, serving as a resource for researchers, publishers, and information professionals seeking to enhance the global dissemination of knowledge and foster inclusivity in the academic landscape.

cs.DL↗

Stark Decline in Journalists' Use of Preprints Post-pandemic

The COVID-19 pandemic accelerated the use of preprints, aiding rapid research dissemination but also facilitating the spread of misinformation. This study analyzes media coverage of preprints from 2014 to 2023, revealing a significant post-pandemic decline. Our findings suggest that heightened awareness of the risks associated with preprints has led to more cautious media practices. While the decline in preprint coverage may mitigate concerns about premature media exposure, it also raises questions about the future role of preprints in science communication, especially during emergencies. Balanced policies based on up-to-date evidence are needed to address this shift.

cs.DL↗

Estimating global article processing charges paid to six publishers for open access between 2019 and 2023

This study presents estimates of the global expenditure on article processing charges (APCs) paid to six publishers for open access between 2019 and 2023. APCs are fees charged for publishing in some fully open access journals (gold) and in subscription journals to make individual articles open access (hybrid). There is currently no way to systematically track institutional, national or global expenses for open access publishing due to a lack of transparency in APC prices, what articles they are paid for, or who pays them. We therefore curated and used an open dataset of annual APC list prices from Elsevier, Frontiers, MDPI, PLOS, Springer Nature, and Wiley in combination with the number of open access articles from these publishers indexed by OpenAlex to estimate that, globally, a total of \$8.349 billion (\$8.968 billion in 2023 US dollars) were spent on APCs between 2019 and 2023. We estimate that in 2023 MDPI (\$681.6 million), Elsevier (\$582.8 million) and Springer Nature (\$546.6) generated the most revenue with APCs. After adjusting for inflation, we also show that annual spending almost tripled from \$910.3 million in 2019 to \$2.538 billion in 2023, that hybrid exceed gold fees, and that the median APCs paid are higher than the median listed fees for both gold and hybrid. Our approach addresses major limitations in previous efforts to estimate APCs paid and offers much needed insight into an otherwise opaque aspect of the business of scholarly publishing. We call upon publishers to be more transparent about OA fees.

cs.DL↗

The oligopoly of academic publishers persists in exclusive database

Global scholarly publishing has been dominated by a small number of publishers for several decades. We aimed to revisit the debate on corporate control of scholarly publishing by analyzing the relative shares of major publishers and smaller, independent publishers. Using the Web of Science, Dimensions and OpenAlex, we managed to retrieve twice as many articles indexed in Dimensions and OpenAlex, compared to the rather selective Web of Science. As a result of excluding smaller publishers, the 'oligopoly' of scholarly publishers persists, at least in appearance, according to the Web of Science. However, both Dimensions' and OpenAlex' inclusive indexing revealed the share of smaller publishers has been growing rapidly, especially since the onset of large-scale online publishing around 2000, resulting in a current cumulative dominance of smaller publishers. While the expansion of small publishers was most pronounced in the social sciences and humanities, the natural and medical sciences showed a similar trend. A major geographical divergence is also revealed, with some countries, mostly Anglo-Saxon and/or located in northwestern Europe, relying heavily on major publishers for the dissemination of their research, while others being relatively independent of the oligopoly, such as those in Latin America, northern Africa, eastern Europe and parts of Asia. The emergence of digital publishing, the reduction of expenses for printing and distribution and open-source journal management tools may have contributed to the emergence of small publishers, while the development of inclusive bibliometric databases has allowed for the effective indexing of journals and articles. We conclude that enhanced visibility to recently created, independent journals may favour their growth and stimulate global scholarly bibliodiversity.

cs.DL↗

An open dataset of article processing charges from six large scholarly publishers (2019-2023)

This paper introduces a dataset of article processing charges (APCs) produced from the price lists of six large scholarly publishers - Elsevier, Frontiers, PLOS, MDPI, Springer Nature and Wiley - between 2019 and 2023. APC price lists were downloaded from publisher websites each year as well as via Wayback Machine snapshots to retrieve fees per journal per year. The dataset includes journal metadata, APC collection method, and annual APC price list information in several currencies (USD, EUR, GBP, CHF, JPY, CAD) for 8,712 unique journals and 36,618 journal-year combinations. The dataset was generated to allow for more precise analysis of APCs and can support library collection development and scientometric analysis estimating APCs paid in gold and hybrid OA journals.

cs.DL↗

Acceso abierto en Argentina: una propuesta para el monitoreo de las publicaciones científicas con OpenAlex

This study proposes a methodology using OpenAlex (OA) for tracking Open Access publications in the case of Argentina, a country where a self-archiving mandate has been in effect since 2013 ( Law 26.899, 2013). A sample of 167,240 papers by researchers from the National Council for Scientific and Technical Research (CONICET) was created and analyzed using statistical techniques. We estimate that OA is able to capture between 85-93% of authors for all disciplines, with the exception of Social Sciences and Humanities, where it only reaches an estimated 47%. The availability of papers in Open Access was calculated to be 41% for the period 1953-2021 and 46% when considering exclusively the post-law period (2014-2021). In both periods, gold Open Access made up the most common route. When comparing equal periods post and pre-law, we observed that the upward trend of gold Open Access was pre-existing to the legislation and the availability of closed articles in repositories increased by 5% to what is estimated based on existing trends. However, while the green route has had a positive evolution, it has been the publication in gold journals that has boosted access to Argentine production more rapidly. We concluded that the OA-based methodology, piloted here for the first time, is viable for tracking Open Access in Argentina since it yields percentages similar to other national and international studies. En este estudio se propone una metodología utilizando OpenAlex (OA) para monitorear el acceso abierto (AA) a las publicaciones científicas para el caso de Argentina, país donde rige el mandato de autoarchivo -Ley 26.899 (2013)-. Se conformó una muestra con 167.240 artículos de investigadores del Consejo Nacional de Investigaciones Científicas y Técnicas (CONICET) que se analizaron con técnicas estadísticas. Se estimó que OA puede representar entre 85-93% de los autores para todas las disciplinas, excepto Ciencias Sociales y Humanidades, donde solo alcanza al 47%. Se calculó que 41% de los artículos publicados entre 1953-2021 incluidos en la fuente están en AA, porcentaje que sube a 46% al considerar exclusivamente el periodo post ley (2014-2021). En ambos periodos es la vía dorada la que representa mayor proporción. Al comparar periodos iguales post y pre ley, se observó que la tendencia en alza de la vía dorada era preexistente a la legislación y la disponibilidad de artículos cerrados en repositorios aumentó un 5% a lo que se estima en base a tendencias existentes. Se concluye que si bien la vía verde ha tenido una evolución positiva, ha sido la publicación en revistas doradas lo que ha impulsado más rápidamente el acceso a la producción argentina. Asimismo, que la metodología basada en OA, piloteada aquí por primera vez, es viable para monitorear el AA en Argentina ya que arroja porcentajes similares a otros estudios nacionales e internacionales.

cs.DL↗

An analysis of the suitability of OpenAlex for bibliometric analyses

Scopus and the Web of Science have been the foundation for research in the science of science even though these traditional databases systematically underrepresent certain disciplines and world regions. In response, new inclusive databases, notably OpenAlex, have emerged. While many studies have begun using OpenAlex as a data source, few critically assess its limitations. This study, conducted in collaboration with the OpenAlex team, addresses this gap by comparing OpenAlex to Scopus across a number of dimensions. The analysis concludes that OpenAlex is a superset of Scopus and can be a reliable alternative for some analyses, particularly at the country level. Despite this, issues of metadata accuracy and completeness show that additional research is needed to fully comprehend and address OpenAlex's limitations. Doing so will be necessary to confidently use OpenAlex across a wider set of analyses, including those that are not at all possible with more constrained databases.

cs.DL↗

How much research shared on Facebook happens outside of public pages and groups? A comparison of public and private online activity around PLOS ONE papers

Despite its undisputed position as the biggest social media platform, Facebook has never entered the main stage of altmetrics research. In this study, we argue that the lack of attention by altmetrics researchers is due, in part, to the challenges in collecting Facebook data regarding activity that takes place outside of public pages and groups. We present a new method of collecting aggregate counts of shares, reactions, and comments across the platform-including users' personal timelines-and use it to gather data for all articles published between 2015 to 2017 in the journal PLOS ONE. We compare the gathered data with altmetrics collected and aggregated by Altmetric. The results show that 58.7% of papers shared on Facebook happen outside of public spaces and that, when collecting all shares, the volume of activity approximates patterns of engagement previously only observed for Twitter. Both results suggest that the role and impact of Facebook as a medium for science and scholarly communication has been underestimated. Furthermore, they emphasise the importance of openness and transparency around the collection and aggregation of altmetrics.

cs.SI↗

Challenges of capturing engagement on Facebook for Altmetrics

Previous research shows that, despite its popularity, Facebook is less frequently used to share academic content. In order to investigate this discrepancy we set out to explore engagement numbers through their Graph API by querying the Facebook API with multiple URLs for a random set of 103,539 articles from the Web of Science. We identified two major challenge areas: mapping articles to URLs and the mapping URLs to objects inside Facebook. We then explored three problem cases within our dataset: (1) identifying a landing page for any given URL, (2) instances where equivalent URLs are mapped to different Facebook objects, and (3) instances of different articles being mapped onto the same Facebook object. We found that the engagement numbers for 11.8% of all articles that have been shared on Facebook at least once are not reliable because of these problems. Moreover, we were unable to identify the URL for 11.6% of the articles in our data. Taken together, the three problem cases constitute 12.3% of the 103,539 tested articles for which engagement numbers cannot be relied upon. Given that we only tested a small number of problem cases and URL variants, our results point to large challenges facing those wishing to collect Facebook metrics programatically through the available API.

cs.SI↗