SearcharxivSearch

arXiv subjects

Maximilian Linde

Publications and source records attributed to Maximilian Linde.

3 recordsLinked to original sources

Making Uncertainty Visible: Multiverse Analysis for Robust Computational Social Science

Through case studies, we demonstrate how multiverse analysis can strengthen the robustness and transparency of computational social science findings against alternative methodological decisions. We conduct multiverse analyses of three published social science studies that use the following computational methods: Bayesian analysis, network generative modeling, and machine learning with or without large language models. These methods are applied frequently in computational social science studies, yet entail a greater degree of arbitrariness in terms of methodological choices, or "researcher degrees of freedom." Our multiverse analyses reveal how the empirical findings in these studies vary as a function of various plausible decision combinations. Our three case studies also expose an often-ignored motivation for conducting multiverse analysis: Showing which methodological combinations lead to computational failure. These failed cases are usually not communicated in the published reports, even though these sophisticated computational methods have a much higher likelihood of failure. We end our paper with suggestions on how to find defensible decision combinations for multiverse analysis of computational social science studies and how to communicate multiverse analysis findings fairly.

stat.OT

Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation

Human label variation has been established as a central phenomenon in NLP: the perspectives different annotators have on the same item need to be embraced. Data collection practices thus shifted towards increasing the annotator numbers and releasing disaggregated datasets, harmful language being most resourced due to its high subjectivity. While this resulted in rich information about \textit{who} annotated (sociodemographics, attitudes, etc.), the \textit{what} (e.g., linguistic properties of items), and their interplay has received little attention. We present the first large-scale analysis of four reference datasets for harmful language detection, bringing together annotator characteristics, linguistic properties of the items, and their interactions in a statistically informed picture. We find that interactions are crucial, revealing intersectional effects ignored in previous work, and that a strong role is played by lexical cues and annotator attitudes. Effect patterns, however, vary considerably across datasets. This urges caution about generalization and transferability.

cs.CL

baymedr: An R Package and Web Application for the Calculation of Bayes Factors for Superiority, Equivalence, and Non-Inferiority Designs

Clinical trials often seek to determine the superiority, equivalence, or non-inferiority of an experimental condition (e.g., a new drug) compared to a control condition (e.g., a placebo or an already existing drug). The use of frequentist statistical methods to analyze data for these types of designs is ubiquitous even though they have several limitations. Bayesian inference remedies many of these shortcomings and allows for intuitive interpretations. In this article, we outline the frequentist conceptualization of superiority, equivalence, and non-inferiority designs and discuss its disadvantages. Subsequently, we explain how Bayes factors can be used to compare the relative plausibility of competing hypotheses. We present baymedr, an R package and web application, that provides user-friendly tools for the computation of Bayes factors for superiority, equivalence, and non-inferiority designs. Instructions on how to use baymedr are provided and an example illustrates how already existing results can be reanalyzed with baymedr.

stat.OT