SearcharxivSearch

arXiv subjects

Frank Esser

Publications and source records attributed to Frank Esser.

4 recordsLinked to original sources

Researchers waste 80% of LLM annotation costs by classifying one text at a time

Large language models (LLMs) are increasingly being used for text classification across the social sciences, yet researchers overwhelmingly classify one text per variable per prompt. Coding 100,000 texts on four variables requires 400,000 API calls. Batching 25 items and stacking all variables into a single prompt reduces this to 4,000 calls, cutting token costs by over 80%. Whether this degrades coding quality is unknown. We tested eight production LLMs from four providers on 3,962 expert-coded tweets across four tasks, varying batch size from 1 to 1,000 items and stacking up to 25 coding dimensions per prompt. Six of eight models maintained accuracy within 2 pp of the single-item baseline through batch sizes of 100. Variable stacking with up to 10 dimensions produced results comparable to single-variable coding, with degradation driven by task complexity rather than prompt length. Within this safe operating range, the measurement error from batching and stacking is smaller than typical inter-coder disagreement in the ground-truth data.

cs.CL

Measuring Structural Political Fragmentation

Political fragmentation denotes the differentiation of a political system into multiple groups and the extent of separation among them. It often manifests structurally in online interaction behaviors. To measure and compare political fragmentation across contexts, previous scholarship has often relied on network measures of polarisation such as modularity and the Krackhardt E-I index. Here, we show that these metrics combine two aspects of fragmentation: the strength of separation and the number of fragments. These two aspects have not been clearly distinguished in previous work, making comparisons across varied systems difficult to interpret. In addition, none of them is designed to capture the multiscale fragmentation structures that characterize real-world multi-dimensional political spaces. We compare several network measures and show that the two aspects of network fragmentation are best captured by the pairwise adaptive E-I index and the effective number of communities (ENC), while other measures confound the strength of separation and the number of fragments. Furthermore, we introduce a novel metric for multiscale fragmentation, the effective branching factor (EBF), capturing how political fragments at one level split into smaller fragments at the next level. Applying EBF to two empirical datasets spanning Brazil, Spain, and the United States yields consistent country rankings across datasets. Overall, these results clarify three complementary dimensions of structural political fragmentation: strength of separation, number of fragments, and between-level branching. They support a more holistic characterization of structural political fragmentation.

cs.SI

Quantifying the Spread of Online Incivility in Brazilian Politics

Incivility refers to behaviors that violate collective norms and disrupt cooperation within the political process. Although large-scale online data and automated techniques have enabled the quantitative analysis of uncivil discourse, prior research has predominantly focused on impoliteness or toxicity, often overlooking other behaviors that undermine democratic values. To address this gap, we propose a multidimensional conceptual framework encompassing Impoliteness, Physical Harm and Violent Political Rhetoric, Hate Speech and Stereotyping, and Threats to Democratic Institutions and Values. Using this framework, we measure the spread of online political incivility in Brazil using approximately 5 million tweets posted by 2,307 political influencers during the 2022 Brazilian general election. Through statistical modeling and network analysis, we examine the dynamics of uncivil posts at different election stages, identify key disseminators and audiences, and explore the mechanisms driving the spread of uncivil information online. Our findings indicate that impoliteness is more likely to surge during election campaigns. In contrast, the other dimensions of incivility are often triggered by specific violent events. Moreover, we find that left-aligned individual influencers are the primary disseminators of online incivility in the Brazilian Twitter/X sphere and that they disseminate not only direct incivility but also indirect incivility when discussing or opposing incivility expressed by others. They relay those content from politicians, media agents, and individuals to reach broader audiences, revealing a diffusion pattern mixing the direct and two-step flows of communication theory. This study offers new insights into the multidimensional nature of incivility in Brazilian politics and provides a conceptual framework that can be extended to other political contexts.

cs.SI

A multilevel network approach to revealing patterns of online political selective exposure

Selective exposure, individuals' inclination to seek out information that supports their beliefs while avoiding information that contradicts them, plays an important role in the emergence of polarization and echo chambers. In the political domain, selective exposure is usually measured on a left-right ideology scale, ignoring finer details. To bridge the gap, this work introduces a multilevel analysis framework based on a multi-scale community detection approach. To test this approach, we combine survey and Twitter/X data collected during the 2022 Brazilian Presidential Election and investigate selective exposure patterns among survey respondents in their choices of whom to follow. We construct a bipartite network connecting survey respondents with political influencers and project it onto the influencer nodes. Applying multi-scale community detection to this projection uncovers a hierarchical clustering of political influencers. Different indices of selective exposure suggest that the characteristics of the influencer communities engaged by survey respondents vary with the level of community resolution. This finding indicates that online political selective exposure exhibits a more complex structure than a mere left-right dichotomy. Moreover, depending on the resolution level we consider, we find different associations between network indices of exposure patterns and 189 individual attributes of the survey respondents. For example, at finer levels, survey respondents' Community Overlap is associated with several factors, such as ideological position, demographics, news consumption frequency, and incivility perception. In comparison, only their ideological position is a significant factor at coarser levels. Our work demonstrates that measuring selective exposure at a single level, such as left and right, misses important information necessary to capture this phenomenon correctly.

cs.SI