SearcharxivSearch

arXiv subjects

Robin Haunschild

Publications and source records attributed to Robin Haunschild.

At least 19 recordsLinked to original sources

Academic match-makers in sociology: Their role in collaboration network formation

In modern scientific collaboration networks, certain researchers play a pivotal role in bridging scholars who have never worked together - a phenomenon we term academic "match-makers." Despite their potential importance, the prevalence, characteristics, benefits, and long-term trajectory of these individuals remain underexplored. Using the Microsoft Academic Graph (MAG), we operationalized a match-maker as an author who, in a given publication, introduced a first-time collaboration between two co-authors, each of whom had previously collaborated with the match-maker but not with each other. We employed a configuration null model to distinguish observed patterns from random chance. Our findings reveal that the match-maker phenomenon is deliberate, prevalent, and consequential. Among authors with over 20 publications, nearly 30% have served as a match-maker, and the probability of acting as one increased eightfold from 1980 to 2019. Publications involving a match-maker are more likely to appear in high-impact journals and exhibit higher disruptiveness - particularly in larger teams - suggesting that match-makers help facilitate what we term integrative disruption. Match-makers tend to emerge early in their careers, peaking around the 20th publication and at an academic age of roughly ten years. While nearly all match-makers eventually experience "abandonment" in the sense that the connected researchers later collaborate without them, their continued involvement remains substantial and is driven by research needs rather than structural factors. This reframes abandonment not as exclusion but as a natural evolution within project-based collaborations. The academic match-maker phenomenon is a strategic feature of collaboration networks characterized by early-career emergence, context-dependent persistence, and tangible contributions to high-impact, disruptive research.

cs.DL

Large language models for post-publication research evaluation: Evidence from expert recommendations and citation indicators

Assessing the quality of scientific research is essential for scholarly communication, yet widely used approaches face limitations in scalability, subjectivity, and time delay. Recent advances in large language models (LLMs) offer new opportunities for automated research evaluation based on textual content. This study examines whether LLMs can support post-publication peer review tasks by benchmarking their outputs against expert judgments and citation-based indicators. Two evaluation tasks are constructed using articles from the H1 Connect platform: identifying high-quality articles and performing finer-grained evaluation including article rating, merit classification, and expert style commenting. Multiple model families, including BERT models, general-purpose LLMs, and reasoning oriented LLMs, are evaluated under multiple learning strategies. Results show that LLMs perform well in coarse grained evaluation tasks, achieving accuracy above 0.8 in identifying highly recommended articles. However, performance decreases substantially in fine-grained rating tasks. Few-shot prompting improves performance over zero-shot settings, while supervised fine-tuning produces the strongest and most balanced results. Retrieval augmented prompting improves classification accuracy in some cases but does not consistently strengthen alignment with citation indicators. The overall correlations between model outputs and citation indicators remain positive but moderate.

cs.IR

Scilit with the Integrated Impact Indicator Assessment

In this study, we systematically elucidate the background and functionality of the Scilit database and evaluate the feasibility and advantages of the comprehensive impact metrics I3 and I3/N, introduced within the Scilit framework. Using a matched dataset of 17,816 journals, we conduct a comparative analysis of Scilit I3/N, Journal Impact Factor, and CiteScore for 2023 and 2024, covering descriptive statistics and distributional characteristics from both disciplinary and publisher perspectives. The comparison reveals that the Scilit I3 and I3/N framework significantly outperforms traditional mean-based metrics in terms of coverage, methodological robustness, and disciplinary fairness. It provides a more accurate, diagnosable, and responsible solution for interdisciplinary journal impact assessment. Our research serves as a "getting started guide" for Scilit, offering scholars, librarians, and academic publishers in the fields of bibliometrics or scientometrics a valuable perspective for exploring I3 and I3/N within an inclusive database. This enables a more accurate and comprehensive understanding of disciplinary development and scientific progress. We advocate for piloting and validating this method in broader evaluation contexts to foster a more precise and diverse representation of scientific progress.

cs.DL

Paper self-citation: An unexplored phenomenon

In this study, we investigated a phenomenon that one intuitively would assume does not exist: self-citations on the paper basis. Actually, papers citing themselves do exist in the Web of Science (WoS) database. In total, we obtained 44,857 papers that have self-citation relations in the WoS raw dataset. In part, they are database artefacts but in part they are due to papers citing themselves in the conclusion or appendix. We also found cases where paper self-citations occur due to publisher-made highlights promoting and citing the paper. We analyzed the self-citing papers according to selected metadata. We observed accumulations of the number of self-citing papers across publication years. We found a skewed distribution across countries, journals, authors, fields, and document types. Finally, we discuss the implications of paper self-citations for bibliometric indicators.

cs.DL

Usage of OpenAlex for creating meaningful global overlay maps of science on the individual and institutional levels

Global overlay maps of science use base maps that are overlaid by specific data (from single researchers, institutions, or countries) for visualizing scientific performance such as field-specific paper output. A procedure to create global overlay maps using OpenAlex is proposed. Six different global base maps are provided. Using one of these base maps, example overlay maps for one individual (the first author of this paper) and his research institution are shown and analyzed. A method for normalizing the overlay data is proposed. Overlay maps using raw overlay data display general concepts more pronounced than their counterparts using normalized overlay data. Advantages and limitations of the proposed overlay approach are discussed.

cs.DL

How to measure research performance of single scientists? A proposal for an index based on scientific prizes: The Prize Winner Index (PWI)

In this study, we propose a new index for measuring excellence in science which is based on collaborations (co-authorship distances) in science. The index is based on the Erd\H{o}s number - a number that was introduced several years ago. We propose to focus with the new index on laureates of prestigious prizes in a certain field and to measure co-authorship distances between the laureates and other scientists. To exemplify and explain our proposal, we computed the proposed index in the field of quantitative science studies (PWIPM). The Derek de Solla Price Memorial Award (Price Medal, PM) is awarded to outstanding scientists in the field. We tested the convergent validity of the PWIPM. We were interested whether the indicator is related to an established bibliometric indicator: P(top 10%). The results show that the coefficients for the correlation between PWIPM and P(top 10%) are high (in cases when a sufficient number of papers have been considered for a reliable assessment of performance). Therefore, measured by an established indicator for research excellence, the new PWI indicator seems to be convergently valid and, therefore, might be a possible alternative for established (bibliometric) indicators - with a focus on prizes.

cs.DL

Comparison of metadata with relevance for bibliometrics between Microsoft Academic Graph and OpenAlex until 2020

Microsoft Academic Graph (MAG) has been studied a lot concerning its suitability for bibliometric evaluations. In May 2021, it was announced that it would retire on December 31, 2021. Soon after that, the non-profit organization OurResearch, aiming at providing 'a fully open catalog of the global research system', announced they would preserve and incorporate the last full MAG corpus, only excluding patent data, and to continue and hopefully improve it. After the launch of OpenAlex in January 2022, it is of interest to know if the usefulness of the MAG data is preserved or even improved in OpenAlex. To this end, we compared metadata that are relevant for bibliometric analyses (in particular field and time normalization of citations) of MAG and OpenAlex: - the coverage of documents over the years, - the agreement of bibliographic data, - the numbers of references of each document, - the kind and distribution of document types, - the distribution and relation of subject classifications.

cs.DL

Identification of young talented individuals in the natural and life sciences using bibliometric data

Identification of young talented individuals is not an easy task. Citation-based data usually need too long to accrue. In this study, we proposed a method based on bibliometric data for the identification of young talented individuals. Three different indicators and their combinations were used. An older cohort with their first publication between 1999 and 2003 was used to find the most suitable indicator combination. For the validation step, citation impact on the level of individual papers was used. The best performing indicator combination was applied to the time period 2007-2011 for identifying young talented individuals who published their first paper within this time period. We produced a set of 46,200 potential talented individuals.

cs.DL

Revolutions in science: The proposal of an approach for the identification of most important researchers, institutions, and countries based on Reference Publication Year Spectroscopy (RPYS)

RPYS is a bibliometric method originally introduced in order to reveal the historical roots of research topics or fields. RPYS does not identify the most highly cited papers of the publication set being studied (as is usually done by bibliometric analyses in research evaluation), but instead it indicates most frequently referenced publications - each within a specific reference publication year. In this study, we propose to use the method to identify important researchers, institutions and countries in the context of breakthrough research. To demonstrate our approach, we focus on research on physical modeling of Earth's climate and the prediction of global warming as an example. Klaus Hasselmann and Syukuro Manabe were both honored with the Nobel Prize in 2021 for their fundamental contributions to this research. Our results reveal that RPYS is able to identify most important researchers, institutions, and countries. For example, all the relevant authors' institutions are located in the USA. These institutions are either research centers of two US National Research Administrations (NASA and NOAA) or universities: the University of Arizona, Princeton University, the Massachusetts Institute of Technology (MIT), and the University of Stony Brook.

cs.DL

How relevant is climate change research for climate change policy? An empirical analysis based on Overton data

Climate change is an ongoing topic in nearly all areas of society since many years. A discussion of climate change without referring to scientific results is not imaginable. This is especially the case for policies since action on the macro scale is required to avoid costly consequences for society. In this study, we deal with the question of how research on climate change and policy are connected. In 2019, the new Overton database of policy documents was released including links to research papers that are cited by policy documents. The use of results and recommendations from research on climate change might be reflected in citations of scientific papers in policy documents. Although we suspect a lot of uncertainty related to the coverage of policy documents in Overton, there seems to be an impact of international climate policy cycles on policy document publication. We observe local peaks in climate policy documents around major decisions in international climate diplomacy. Our results point out that IGOs and think tanks -- with a focus on climate change -- have published more climate change policy documents than expected. We found that climate change papers that are cited in climate change policy documents received significantly more citations on average than climate change papers that are not cited in these documents. Both areas of society (science and policy) focus on similar climate change research fields: biology, earth sciences, engineering, and disease sciences. Based on these and other empirical results in this study, we propose a simple model of policy impact considering a chain of different document types: the chain starts with scientific assessment reports (systematic reviews) that lead via science communication documents (policy briefs, policy reports or plain language summaries) and government reports to legislative documents.

physics.soc-ph

Empirical analysis of recent temporal dynamics of research fields: Annual publications in chemistry and related areas as an example

Changes in the number of publications in a certain field might reflect the dynamic of scientific progress in this field, since an increase in the number of publications can be interpreted as an increase in the field-specific knowledge. In this paper, we present a methodological approach to analyse the dynamics of science on lower aggregation levels, i.e., the level of research fields. Our trend analysis approach is able to uncover very recent trends, and the methods used to study the trends are simple to understand for the possible recipients of the results. In order to demonstrate the trend analysis approach, we focused in this study on the annual number of publications (and patents) in chemistry (and related areas) between 2014 and 2020 identifying those fields in chemistry with the highest dynamics (largest rates of change in publication counts). The study is based on the mono-disciplinary literature database CAplus. Our results reveal that the number of publications in the CAplus database is increasing since many years. Research regarding optical phenomena and electrochemical technologies was found to be among the emerging topics in recent years.

cs.DL

Scores of a specific field-normalized indicator calculated with different approaches of field-categorization: Are the scores different or similar?

Usage of field-normalized citation scores is a bibliometric standard. Different methods for field-normalization are in use, but also the choice of field-classification system determines the resulting field-normalized citation scores. Using Web of Science data, we calculated field-normalized citation scores using the same formula but different field-classification systems to answer the question if the resulting scores are different or similar. Six field-classification systems were used: three based on citation relations, one on semantic similarity scores (i.e., a topical relatedness measure), one on journal sets, and one on intellectual classifications. Systems based on journal sets and intellectual classifications agree on at least the moderate level. Two out of the three sets based on citation relations also agree on at least the moderate level. Larger differences were observed for the third data set based on citation relations and semantic similarity scores. The main policy implication is that normalized citation impact scores or rankings based on them should not be compared without deeper knowledge of the classification systems that were used to derive these values or rankings.

cs.DL

Growth rates of modern science: A latent piecewise growth curve approach to model publication numbers from established and new literature databases

Growth of science is a prevalent issue in science of science studies. In recent years, two new bibliographic databases have been introduced which can be used to study growth processes in science from centuries back: Dimensions from Digital Science and Microsoft Academic. In this study, we used publication data from these new databases and added publication data from two established databases (Web of Science from Clarivate Analytics and Scopus from Elsevier) to investigate scientific growth processes from the beginning of the modern science system until today. We estimated regression models that included simultaneously the publication counts from the four databases. The results of the unrestricted growth of science calculations show that the overall growth rate amounts to 4.10% with a doubling time of 17.3 years. As the comparison of various segmented regression models in the current study revealed, the model with five segments fits the publication data best. We demonstrated that these segments with different growth rates can be interpreted very well, since they are related to either phases of economic (e.g., industrialization) and / or political developments (e.g., Second World War). In this study, we additionally analyzed scientific growth in two broad fields (Physical and Technical Sciences as well as Life Sciences) and the relationship of scientific and economic growth in UK. The comparison between the two fields revealed only slight differences. The comparison of the British economic and scientific growth rates showed that the economic growth rate is slightly lower than the scientific growth rate.

cs.DL

Investigating Dissemination of Scientific Information on Twitter: A Study of Topic Networks in Opioid Publications

One way to assess a certain aspect of the value of scientific research is to measure the attention it receives on social media. While previous research has mostly focused on the "number of mentions" of scientific research on social media, the current study applies "topic networks" to measure public attention to scientific research on Twitter. Topic networks are the networks of co-occurring author keywords in scholarly publications and networks of co-occurring hashtags in the tweets mentioning those scholarly publications. This study investigates which topics in opioid scholarly publications have received public attention on Twitter. Additionally, it investigates whether the topic networks generated from the publications tweeted by all accounts (bot and non-bot accounts) differ from those generated by non-bot accounts. Our analysis is based on a set of opioid scholarly publications from 2011 to 2019 and the tweets associated with them. We use co-occurrence network analysis to generate topic networks. Results indicated that Twitter users have mostly used generic terms to discuss opioid publications, such as "Opioid," "Pain," "Addiction," "Treatment," "Analgesics," "Abuse," "Overdose," and "Disorders." Results confirm that topic networks provide a legitimate method to visualize public discussions of health-related scholarly publications and how Twitter users discuss health-related scientific research differently from the scientific community. There was a substantial overlap between the topic networks based on the tweets by all accounts and non-bot accounts. This result indicates that it might not be necessary to exclude bot accounts for generating topic networks as they have a negligible impact on the results.

cs.IR

Reference Publication Year Spectroscopy (RPYS) in practice: A software tutorial

In course of the organization of Workshop III entitled "Cited References Analysis Using CRExplorer" at the International Conference of the International Society for Scientometrics and Informetrics (ISSI2021), we have prepared three reference publication year spectroscopy (RPYS) analyses: (i) papers published in Journal of Informetrics; (ii) papers regarding the topic altmetrics; and (iii) papers published by Ludo Waltman (we selected this researcher since he received the Derek de Solla Price Memorial Medal during the ISSI2021 conference). The first RPYS analysis has been presented live at the workshop and the second and third RPYS analyses have been left to the participants for undertaking after the workshop. Here, we present the results for all three RPYS analyses. The three analyses have shown quite different seminal papers with a few overlaps. Many of the foundational papers in the field of scientometrics (e.g., distributions of publications and citations, citation network and co-citation analyses, and citation analysis with the aim of impact measurement and research evaluation) were retrieved as seminal papers of the papers published in Journal of Informetrics. Mainly papers with discussions of the deficiencies of citation-based impact measurements and comparisons between altmetrics and citations were retrieved as seminal papers of the topic altmetrics. The RPYS analysis of the paper set published by Ludo Waltman mainly retrieved papers about network analyses, citation relations, and citation impact measurement.

cs.DL

Report on Workshop III "Cited References Analysis Using CRExplorer" at the 18th International Conference of the International Society for Scientometrics and Informetrics (ISSI2021)

We have organized Workshop III entitled "Cited References Analysis Using CRExplorer" at ISSI2021. Here, we report and reflect on this workshop. The aim of this workshop was to bring beginners, practitioners, and experts in cited references analyses together. A mixture of presentations and an interactive part was intended to provide benefits for all kinds of scientometricians with an interest in cited references analyses.

cs.DL

Heat Waves -- a hot topic in climate change research

Research on heat waves (periods of excessively hot weather, which may be accompanied by high humidity) is a newly emerging research topic within the field of climate change research with high relevance for the whole of society. In this study, we analyzed the rapidly growing scientific literature dealing with heat waves. No summarizing overview has been published on this literature hitherto. We developed a suitable search query to retrieve the relevant literature covered by the Web of Science (WoS) as complete as possible and to exclude irrelevant literature (n = 8,011 papers). The time-evolution of the publications shows that research dealing with heat waves is a highly dynamic research topic, doubling within about 5 years. An analysis of the thematic content reveals the most severe heat wave events within the recent decades (1995 and 2003), the cities and countries/regions affected (United States, Europe, and Australia), and the ecological and medical impacts (drought, urban heat islands, excess hospital admissions, and mortality). Risk estimation and future strategies for adaptation to hot weather are major political issues. We identified 104 citation classics which include fundamental early works of research on heat waves and more recent works (which are characterized by a relatively strong connection to climate change).

cs.DL

Quantum technology 2.0 -- topics and contributing countries from 1980 to 2018

The second quantum technological revolution started around 1980 with the control of single quantum particles and their interaction on an individual basis. These experimental achievements enabled physicists and engineers to utilize long-known quantum features - especially superposition and entanglement of single quantum states - for a whole range of practical applications. We use a publication set of 54,598 papers from the Web of Science published between 1980 and 2018 to investigate the time development of four main subfields of quantum technology in terms of numbers and shares of publication as well as the occurrence of topics and their relation to the 25 top contributing countries. Three successive time periods are distinguished in the analyses by their short doubling times in relation to the whole Web of Science. The periods can be characterized by the publication of pioneering works, the exploration of research topics, and the maturing of quantum technology, respectively. Compared to the US, China has a far over proportional contribution to the worldwide publication output, but not in the segment of highly-cited papers.

cs.DL