SearcharxivSearch

arXiv subjects

Lutz Bornmann

Publications and source records attributed to Lutz Bornmann.

At least 19 recordsLinked to original sources

SoniMet - A tool for sonifying and visualizing the performance of single researchers

For centuries, the scientific community has predominantly relied on visual tools to communicate empirical results and complex datasets. While visual representations dominate bibliometric analyses, the human auditory system possesses sensitive capacities for processing complex temporal information, distinguishing intricate patterns, and tracking parallel data streams. Data sonification translates data relations into acoustic signals, offering an alternative method for data exploration, pattern recognition, and scientific communication. This paper applies this concept to the field of bibliometrics through metrics sonification-the auditory translation of bibliometric information-and introduces SoniMet (Sonifying Metrics), a web-based tool designed to visualize and sonify the publication and citation data of individual scientists (see https://sonimet.kennebec.co.uk). SoniMet connects to the OpenAlex database to retrieve bibliometric records and displays them on an interactive chronological timeline. The tool employs direct parameter mapping to translate citation impact indicators into non-speech sound: field-weighted citation impact and citation counts determine the pitch and volume of a synthesized note and are mapped to the acoustic echo strength. Although SoniMet expands the methodological toolkit for research evaluation, current limitations include its restriction to individual scholar profiles and the challenge of integrating transient audio files into traditional, text-based scientific publishing workflows. Future empirical user studies are necessary to systematically evaluate the analytical utility and cognitive benefits of metrics sonification compared to established visual methods.

cs.DL

Female participation in science in the past 125 years: An analysis of the Matilda effect over time

The Matilda effect describes the systematic under-recognition of women's scientific contributions. We investigated its historical evolution by analyzing over 220 million publications (1900-2025) from OpenAlex, inferring the gender of roughly 60 million authors. To quantify this under-recognition, we calculated the ratio between the overall share of female authors and female corresponding authors. Because corresponding authorship denotes intellectual leadership and primary academic credit, systematic exclusion from this role directly measures the Matilda effect. Our analysis reveals that overall female participation rose from approximately 10% in the 1910s to over 40% recently. For most of the 20th century, women were disproportionately excluded from corresponding authorship, confirming a historical Matilda effect. However, this trend inverted over the past two decades, indicating a global decline. Despite this overall progress, distinct disparities persist: the physical sciences and Asian countries lag behind, whereas the health sciences and Latin America approach gender parity.

cs.DL

Judges matter more than papers in post-publication research assessment

Research assessment relies on expert evaluations, yet human judgement is noisy, and it is unclear whether differences in assessment arise primarily from differences in genuine research quality or from unwanted differences between evaluators. While numerous studies highlight disagreement and biases in research assessment, they have not quantified judge-related noise relative to variation in the evaluated works. Here we show, in a large post-publication peer review database, that research assessment is driven more by differences between evaluators than by difference in the evaluated research. We partition variance in 239,521 research quality ratings assigned by 12,649 judges to 193,128 papers from the H1 Connect post-publication peer review platform. Using multilevel models, we decomposed judge-related variation into differences in overall severity and differences in the weighting of scientific attributes. We found that judge-related effects accounted for substantially more variance in ratings than the evaluated papers. In our most detailed model, judge-level effects and judge-specific slopes explained 61% of the total variance, whereas combined paper and journal-level effects accounted for only 7%. By contrast, examined measures of directional bias, such as author gender and global affiliation, explained less than 1% of the variance. We conclude that assessment outcomes were shaped more by the judges than by the papers themselves. Our results demonstrate the necessity of noise audits in high-stakes scientific evaluation.

cs.DL

The Rising Dominance of Methods Across Science

Scientific progress is traditionally narrated through the interplay of theoretical insights and experimental findings. Yet this view of science underplays a third and central pillar of progress: the methods that underlie both conceptual advances and empirical evidence. By analysing more than 3 million articles across science published between 1980 and 2019, we find that science has undergone a fundamental structural transition. The share of papers that primarily contribute new methods-methods papers-has doubled across science over the past four decades, rising universally across disciplines and citation impact levels. Rather than a gradual evolution, this transition marks a pivotal shift beginning in the early 1990s, aligning with the computational revolution and the emergence of data-intensive science. The surge in methodological research is not confined to the most cited, elite publications; it spans the full spectrum of scientific output. These findings reveal a systemic reorientation of the scientific ecosystem where reusable methods increasingly serve as the essential infrastructure of scientific advances, challenging the traditional dichotomy of theory and experimental research. As science becomes increasingly methods-driven, our results call for rethinking how research is evaluated, funded and organised-towards better incentivising method innovations. This is especially the case as expanding AI must be effectively integrated with scientific instruments to realise its full potential.

cs.DL

Large language models for post-publication research evaluation: Evidence from expert recommendations and citation indicators

Assessing the quality of scientific research is essential for scholarly communication, yet widely used approaches face limitations in scalability, subjectivity, and time delay. Recent advances in large language models (LLMs) offer new opportunities for automated research evaluation based on textual content. This study examines whether LLMs can support post-publication peer review tasks by benchmarking their outputs against expert judgments and citation-based indicators. Two evaluation tasks are constructed using articles from the H1 Connect platform: identifying high-quality articles and performing finer-grained evaluation including article rating, merit classification, and expert style commenting. Multiple model families, including BERT models, general-purpose LLMs, and reasoning oriented LLMs, are evaluated under multiple learning strategies. Results show that LLMs perform well in coarse grained evaluation tasks, achieving accuracy above 0.8 in identifying highly recommended articles. However, performance decreases substantially in fine-grained rating tasks. Few-shot prompting improves performance over zero-shot settings, while supervised fine-tuning produces the strongest and most balanced results. Retrieval augmented prompting improves classification accuracy in some cases but does not consistently strengthen alignment with citation indicators. The overall correlations between model outputs and citation indicators remain positive but moderate.

cs.IR

Institutional cooperations in Austrian research: An analysis of shared researchers

Multiple organisational affiliations are an increasingly common feature of research systems, yet their implications for organisational performance had received limited systematic attention. We developed a scalable, network-based analytical framework that represents simultaneous researcher affiliations as relational links between organisations and applied it to bibliometric data from Austria. Using harmonised publication and affiliation metadata, we constructed two complementary co-affiliation networks: a complete network capturing all simultaneous affiliations and a temporally filtered network retaining only organisational pairs that recurred over time. Network regression analyses showed that geographical proximity remained an important determinant of co-affiliation formation, with spatial distance consistently reducing shared appointments. Clear sectoral differences emerged beyond geography. Universities formed a dense and persistent core of co-affiliations, whereas ties involving medical institutions, government, non-profit and private-sector organisations were often short-lived and attenuated under temporal filtering. Among crosssector links, co-affiliations between universities and research institutes were notably resilient, indicating a more structurally embedded form of organisational integration. We assessed the effect of concurrent affiliations on organisational citation impact across organisational types using field- and year-normalised indicators. Research institutes and universities consistently exhibited higher citation impact than organisations from other sectors, and persistent co-affiliations were associated with greater and more stable scientific visibility.

cs.DL

Reforming research funding: Combining editorial preregistration with grant peer review

Competitive grant funding is associated with high costs and a potential bias to favor conservative research. This comment proposes integrating editorial preregistration, in the form of registered reports, into grant peer review processes as a reform strategy. Linking funding decisions to in principle accepted study protocols would reduce reviewer burden, strengthen methodological rigor, and provide an institutional foundation for (more) replication, theory driven research, and high risk research. Our proposal also minimizes strategic proposal writing and ensures scholarly output through the publication of preregistered protocols, regardless of funding outcomes. Possible implementation models include direct coupling of journal acceptance with funding, co review mechanisms, voucher systems, and lotteries. While challenges remain in aligning journal and funding agency procedures, the integration of preregistration and funding offers a promising pathway toward a more transparent and efficient research ecosystem.

cs.DL

Citation accuracy, citation noise, and citation bias: A foundation of citation analysis

Citation analysis is widely used in research evaluation to assess the impact of scientific papers. These analyses rest on the assumption that citation decisions by authors are accurate, representing the flow of knowledge from cited to citing papers. However, in practice, researchers often cite for reasons that are not related to the fact that there has been (intellectual) input from previous papers. Citations made for rhetorical reasons or without reading the cited work compromise the value of citations as instrument for research evaluation. Past research on threats to the accuracy of citations has mainly focused on citation bias as the primary concern. In this paper, we argue that citation noise - the undesirable variance in citation decisions - represents an equally critical but underexplored challenge in citation analysis. We define and differentiate two types of citation noise: citation level noise and citation pattern noise. Each type of noise is described in terms of how it arises and the specific ways it can undermine the validity of citation-based research assessments. By conceptually differing citation noise from citation accuracy and citation bias, we propose a framework for the foundation of citation analysis. We discuss strategies and interventions to minimize citation noise, aiming to improve the reliability and validity of citation analysis in research evaluation. We recommend that the current professional reform movement in research evaluation such as the Coalition for Advancing Research Assessment (CoARA) pick up these strategies and interventions as an additional building block for careful, responsible use of bibliometric indicators in research evaluation.

cs.DL

Introducing multiverse analysis to bibliometrics: The case of team size effects on disruptive research

Although bibliometrics has become an essential tool in the evaluation of research performance, bibliometric analyses are sensitive to a range of methodological choices. Subtle choices in data selection, indicator construction, and modeling decisions can substantially alter results. Ensuring robustness (meaning that findings hold up under different reasonable scenarios) is therefore critical for credible research and research evaluation. To address this issue, this study introduces multiverse analysis to bibliometrics. Multiverse analysis is a statistical tool that enables analysts to transparently discuss modeling assumptions and thoroughly assess model robustness. Whereas standard robustness checks usually cover only a small subset of all plausible models, multiverse analysis includes all plausible models. The benefits of multiverse analysis are illustrated by assessing the robustness of the findings reported by Wu et al. (2019), who observed that small teams tend to produce more disruptive research than large teams. While we found robust evidence of a negative effect of team size on disruption scores, the effect size depends substantially on the model specification. Our findings underscore the importance of assessing the multiverse robustness of bibliometric results to clarify their practical implications.

cs.DL

Dynamic disruption index across citation and cited references windows: Recommendations for thresholds in research evaluation

The temporal dimension of citation accumulation poses fundamental challenges for quantitative research evaluations, particularly in assessing disruptive and consolidating research through the disruption index (D). While prior studies emphasize minimum citation windows (mostly 3-5 years) for reliable citation impact measurements, the time-sensitive nature of D - which quantifies a paper' s capacity to eclipse prior knowledge - remains underexplored. This study addresses two critical gaps: (1) determining the temporal thresholds required for publications to meet citation/reference prerequisites, and (2) identifying "optimal" citation windows that balance early predictability and longitudinal validity. By analyzing millions of publications across four fields with varying citation dynamics, we employ some metrics to track D stabilization patterns. Key findings reveal that a 10-year window achieves >80% agreement with final D classifications, while shorter windows (3 years) exhibit instability. Publications with >=30 references stabilize 1-3 years faster, and extreme cases (top/bottom 5% D values) become identifiable within 5 years - enabling early detection of 60-80% of highly disruptive and consolidating works. The findings offer significant implications for scholarly evaluation and science policy, emphasizing the need for careful consideration of citation window length in research assessment (based on D).

cs.DL

Paper self-citation: An unexplored phenomenon

In this study, we investigated a phenomenon that one intuitively would assume does not exist: self-citations on the paper basis. Actually, papers citing themselves do exist in the Web of Science (WoS) database. In total, we obtained 44,857 papers that have self-citation relations in the WoS raw dataset. In part, they are database artefacts but in part they are due to papers citing themselves in the conclusion or appendix. We also found cases where paper self-citations occur due to publisher-made highlights promoting and citing the paper. We analyzed the self-citing papers according to selected metadata. We observed accumulations of the number of self-citing papers across publication years. We found a skewed distribution across countries, journals, authors, fields, and document types. Finally, we discuss the implications of paper self-citations for bibliometric indicators.

cs.DL

Specification uncertainty: What the disruption index tells us about the (hidden) multiverse of bibliometric indicators

Following Funk and Owen-Smith (2017), Wu et al. (2019) proposed the disruption index (DI1) as a bibliometric indicator that measures disruptive and consolidating research. When we summarized the literature on the disruption index for our recently published review article (Leibel & Bornmann, 2024), we noticed that the calculation of disruption scores comes with numerous (hidden) degrees of freedom. In this Letter to the Editor, we explain based on the DI1 (as an example) why the analytical flexibility of bibliometric indicators potentially endangers the credibility of research and advertise the application of multiverse-style methods to increase the transparency of the research.

cs.DL

Metrics sonification: The introduction of new ways to present bibliometric data using publication data of Loet Leydesdorff as an example

The visualization of publication and citation data is popular in bibliometrics. Although less common, the representation of empirical data as sound is an alternative form of presentation (in other fields than bibliometrics). In this representation, the data are mapped into sound and listened to by an audience. Approaches for the sonification of data have been developed in many fields since decades. Since sonification has several advantages for the presentation of data, this study is intended to introduce sonification to bibliometrics named as 'metrics sonification'. Metrics sonification is defined as the sonification of bibliometric information (measurements, data or results) for their empirical analysis and/or presentation. In this study, we used metadata of publications by Loet Leydesdorff (named as Loet in the following) to sonify their properties. Loet was a giant in the field of scientometrics, who passed away in 2023. The track based on Loet's publications can be listened to on SoundCloud using the following link: https://on.soundcloud.com/oxBTA32x4EgwvKVz5. The track has been composed in F minor; this key was chosen to express the sad occasion. The quantitative part of the track includes a parameter mapping (a sonification) of three properties of his publications: (1) publication output, (2) open access publication, and (3) citation impact of publications. The qualitative part (spoken audio) focuses on explanations of the parameter mapping and descriptions of the mapped papers (based on their titles and abstracts). The sonification of Loet's publications presented in this study is only one possible type of metrics sonification application. As the great number of projects from other disciplines have demonstrated, many other types of applications are possible in bibliometrics.

cs.DL

Usage of OpenAlex for creating meaningful global overlay maps of science on the individual and institutional levels

Global overlay maps of science use base maps that are overlaid by specific data (from single researchers, institutions, or countries) for visualizing scientific performance such as field-specific paper output. A procedure to create global overlay maps using OpenAlex is proposed. Six different global base maps are provided. Using one of these base maps, example overlay maps for one individual (the first author of this paper) and his research institution are shown and analyzed. A method for normalizing the overlay data is proposed. Overlay maps using raw overlay data display general concepts more pronounced than their counterparts using normalized overlay data. Advantages and limitations of the proposed overlay approach are discussed.

cs.DL

A proposal to improve the calculation of the disruption index

Wu et al. (2019) proposed the disruption index (DI1) as a bibliometric indicator that measures disruptive and consolidating research. Leibel and Bornmann (2024) recently published a literature overview on the disruption index research in Scientometrics. In this letter to the editor, we point out that the method of calculating the DI1 score of a focal paper contains a logical impact measurement error that leads to a meaningful reduction of the score. We explain why this is problematic and propose a correction of the formula.

cs.DL

The Costs of Competition in Distributing Scarce Research Funds

Research funding systems are not isolated systems - they are embedded in a larger scientific system with an enormous influence on the system. This paper aims to analyze the allocation of competitive research funding from different perspectives: How reliable are decision processes for funding? What are the economic costs of competitive funding? How does competition for funds affect doing risky research? How do competitive funding environments affect scientists themselves, and which ethical issues must be considered? We attempt to identify gaps in our knowledge of research funding systems; we propose recommendations for policymakers and funding agencies, including empirical experiments of decision processes and the collection of data on these processes. With our recommendations we hope to contribute to developing improved ways of organizing research funding.

econ.GN

What do we know about the disruption index in scientometrics? An overview of the literature

The purpose of this paper is to provide a review of the literature on the original disruption index (DI1) and its variants in scientometrics. The DI1 has received much media attention and prompted a public debate about science policy implications, since a study published in Nature found that papers in all disciplines and patents are becoming less disruptive over time. This review explains in the first part the DI1 and its variants in detail by examining their technicaland theoretical properties. The remaining parts of the review are devoted to studies that examine the validity and the limitations of the indices. Particular focus is placed on (1) possible biases that affect disruption indices (2) the convergent and predictive validity of disruption scores, and (3) the comparative performance of the DI1 and its variants. The review shows that, while the literature on convergent validity is not entirely conclusive, it is clear that some modified index variants, in particular DI5, show higher degrees of convergent validity than DI1. The literature draws attention to the fact that (some) disruption indices suffer from inconsistency, time-sensitive biases, and several data-induced biases. The limitations of disruption indices are highlighted and best practice guidelines are provided. The review encourages users of the index to inform about the variety of DI1 variants and to apply the most appropriate variant. More research on the validity of disruption scores as well as a more precise understanding of disruption as a theoretical construct is needed before the indices can be used in the research evaluation practice.

cs.DL

How to measure research performance of single scientists? A proposal for an index based on scientific prizes: The Prize Winner Index (PWI)

In this study, we propose a new index for measuring excellence in science which is based on collaborations (co-authorship distances) in science. The index is based on the Erdős number - a number that was introduced several years ago. We propose to focus with the new index on laureates of prestigious prizes in a certain field and to measure co-authorship distances between the laureates and other scientists. To exemplify and explain our proposal, we computed the proposed index in the field of quantitative science studies (PWIPM). The Derek de Solla Price Memorial Award (Price Medal, PM) is awarded to outstanding scientists in the field. We tested the convergent validity of the PWIPM. We were interested whether the indicator is related to an established bibliometric indicator: P(top 10%). The results show that the coefficients for the correlation between PWIPM and P(top 10%) are high (in cases when a sufficient number of papers have been considered for a reliable assessment of performance). Therefore, measured by an established indicator for research excellence, the new PWI indicator seems to be convergently valid and, therefore, might be a possible alternative for established (bibliometric) indicators - with a focus on prizes.

cs.DL