SearcharxivSearch

arXiv subjects

Julia Witte Zimmerman

Publications and source records attributed to Julia Witte Zimmerman.

12 recordsLinked to original sources

Statistical laws and linguistics differ in naturalistic video and fictional conversations

Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generating complex, dynamic stories. One approach to studying how conversations unfold in time is through statistical patterns such as Heaps' law, which holds that vocabulary size scales with document length. Little work on Heaps' law has looked at conversation and considered how language features impact scaling. We measure Heaps' law for conversations recorded in two distinct mediums: 1. Strangers brought together on video chat and 2. Fictional characters in movies. We find that scaling of vocabulary size differs by parts of speech, suggesting a less efficient purpose in communication by medium.

cs.CL

Tokens, the oft-overlooked appetizer: Large language models, the distributional hypothesis, and meaning

Tokenization is a necessary component within the current architecture of many language mod-els, including the transformer-based large language models (LLMs) of Generative AI, yet its impact on the model's cognition is often overlooked. We argue that LLMs demonstrate that the Distributional Hypothesis (DH) is sufficient for reasonably human-like language performance (particularly with respect to inferential lexical competence), and that the emergence of human-meaningful linguistic units among tokens and current structural constraints motivate changes to existing, linguistically-agnostic tokenization techniques, particularly with respect to their roles as (1) vehicles for conveying salient distributional patterns from human language to the model and as (2) semantic primitives. We explore tokenizations from a BPE tokenizer; extant model vocabularies obtained from Hugging Face and tiktoken; and the information in exemplar token vectors as they move through the layers of a RoBERTa (large) model. Besides creating suboptimal semantic building blocks and obscuring the model's access to the necessary distributional patterns, we describe how tokens and pretraining can act as a backdoor for bias and other unwanted content, which current alignment practices may not remediate. Additionally, we relay evidence that the tokenization algorithm's objective function impacts the LLM's cognition, despite being arguably meaningfully insulated from the main system intelligence. Finally, we discuss implications for architectural choices, meaning construction, the primacy of language for thought, and LLM cognition. [First uploaded to arXiv in December, 2024.]

cs.CL

Stop using Media Bias/Fact Check in research

Media Bias/Fact Check (MBFC) purports to quantify the bias, credibility, and factuality of reporting for roughly 10,000 media sources, and the resulting data is commonly used in misinformation research. In the present study, we show that MBFC's methodology does not meet basic standards of rigor for academic research. Despite its widespread prevalence, studies using MBFC rarely examine it carefully, often describing it in ways that contradict its ``About'' page, or treating it as authoritative despite MBFC's disclaimer that it is ``not a tested scientific method... [but] a simple guide to the idea of a source's bias.'' We identified no papers that adequately describe MBFC as the opinions of a single person or critically engage with its methodology in order to justify proceeding with its use. We argue that MBFC's data is not neutral or accurate, but a computationally legible account of hegemony, a specious dataset for uncritical research that mistakes the familiarity of the concepts it quantifies with accuracy. Our study concludes with a call for academic researchers to stop using MBFC. MBFC's data quantifies the results of political processes, including campaigns to discredit the press, and presents them as simple facts about the world, thus reproducing the crisis misinformation scholarship exists to address.

cs.SI

The queer Hero versus the Fool bias of the queer trait: An archetypometric analysis of the collective portrayal of queerness in fictional stories

Visibility in media is pivotal for identity development and for broadening societal views of gender and sexuality. Queer representation has increased in recent years, yet damaging stereotypes and tropes persist. Here, we focus on queer portrayal and its perception by audiences in fictional stories (television, film, and literature) by studying characters by their quantified archetypes which are operationalizations of common conceptions such as Hero, Diva, and Outcast. We use the archetypometrics and Fandom's LGBTQIA+ datasets to study samples of fictional characters along the trait differential spanning straight to queer. We find, quantify, and explain a seeming paradox. The characters with the highest queer score present positive primary archetypes and are typically Heroes rather than Fools, Angels rather than Demons, and Adventurers rather than Traditionalists. But evaluation across many stories for the straight-queer trait itself reveals a strong collective-writing bias towards Fool (away from Hero) and no meaningful loading for the other two dimensions. Our analysis offers a population-scale view of the complexities of queer portrayal, while also pointing to risks in blindly training on many-authored story corpora.

cs.CY

Buffy versus Bella: An archetypometric analysis and comparison

Fictional stories and characters embody and encode social norms, and their study is a powerful tool through which to understand culture and society. Vampire stories and folklore, in particular, have long both reflected and refracted people's preoccupation with disease, sexuality, death, and immortality. Here, we explore female main characters from two popular vampire franchises of the 21st century: Buffy Summers from the eponymous Buffy the Vampire Slayer and Bella Swan from the Twilight series. We employ the archetypometrics framework, built from 2,000 characters assesed across 464 semantic differential traits, to understand Buffy's and Bella's archetypes compared to one another and characters in their own stories, as well as within a larger societal context. While Buffy and Bella are female protagonists who share focus on love and romance, they differ broadly on their underlying traits and overall archetypes. Buffy -- presented as a prototypical high school cheerleader -- largely bucks traditional gender norms as an strong Adventurer-Hero. Bella -- stylized as ``not like the other girls'' -- largely conforms to traditional gender norms as a weak Outcast archetype. In each instance, our use of archetypometrics offers a detailed, character-based lens for assessing female protagonists in contemporary vampire narratives, with clear potential for broader application across other storytelling forms.

cs.CY

Simon's model does not produce Zipf's law: The fundamental rich-get-richer mechanism for any power-law size ranking

Many complex systems are composed of disparate, interacting types of varying sizes: Species abundances in ecosystems, firm sizes in markets, city populations in countries, word counts in language, etc. A longstanding mystery of complex systems is Zipf's law, which is the empirical observation that component size decreases as the inverse of component rank -- $S \propto r^{-1}$ -- and its generalization $S \propto r^{-α}$ for $α\ge 0$. Herbert Simon's 1955 theoretical rich-get-richer mechanism for system growth has prevailed as capturing the essential process. But Simon's analysis is in fact flawed: In the limit of zero innovation, the model leads to a winner-takes-all system with $α\rightarrow \infty$, rather than $α\rightarrow 1$. Here, for pure rich-get-richer systems, we derive the time-dependent innovation rate $ρ_t$ that correctly produces power-law size rankings across all $α\ge 0$. To produce Zipf's law, we uncover that $ρ_t$ must decay as the inverse of the log of the number of types, $1/\ln N$. We then show that our time-dependent innovation rate governs type emergence in any system obeying a power-law size-ranking, independent of the underlying mechanism. We demonstrate agreement between our model's output and word rankings in a collection of famous novels, while Simon's model fails. Going forward, our dynamic innovation rate mechanism provides the fundamental, Drosophila-like model for all rich-get-richer systems.

physics.soc-ph

Archetypes and gender in fiction: A data-driven mapping of gender stereotypes in stories

Fictional character representations reflect social norms and biases. For example, women are relatively underrepresented in television and film, irrespective of genre, and are frequently stereotyped in these media. Here, we draw on a data-driven operationalization of archetypes -- archetypometrics -- to explore the characterization of 2,000 canonically male and female characters. From an overall space of six pairs of base archetypes, we find that canonically female characters tend more toward Hero, Adventurer, Diva, and Sophisticate archetypes, while male characters, tend toward Fool, Traditionalist, Outcast, Brute and Outcast types. However, overarching patterns by gender nevertheless sustain traditional stereotypes: The seemingly positive heroic bias toward females is undercut by heroic female characters being more masculine than other female characters. We discuss the societal implications of skewed archetype representation by character gender.

cs.CY

False memories to fake news: The evolution of the term "misinformation" in academic literature

Since 2016, the term "misinformation" has become associated with a scientific paradigm that studies, at its core, people making, reading, and sharing false statements, usually on social media, and often warning of the harm to society resulting from the sum of many such events. By tracking the term through the academic literature, with special focus on the years 2011--2023, we connect the post-2016 paradigm with a strand of research dating to the Satanic panic of the 1980s. We argue that post-2016 misinformation research owes more to this intellectual lineage than is generally acknowledged, and we discuss the theoretical and practical implications of this connection. We conclude by drawing parallels between the Satanic panic and 2026, and, similarly, between misinformation research then and now.

cs.SI

Us-vs-Them bias in Large Language Models

This study investigates ``us versus them'' bias, as described by Social Identity Theory, in large language models (LLMs) under both default and persona-conditioned settings across multiple architectures (GPT-4.1, DeepSeek-3.1, Gemma-2.0, Grok-3.0, and LLaMA-3.1). Using sentiment dynamics, allotaxonometry, and embedding regression, we find consistent ingroup-positive and outgroup-negative associations across foundational LLMs. We find that adopting a persona systematically alters models' evaluative and affiliative language patterns. For the exemplar personas examined, conservative personas exhibit greater outgroup hostility, whereas liberal personas display stronger ingroup solidarity. Persona conditioning produces distinct clustering in embedding space and measurable semantic divergence, supporting the view that even abstract identity cues can shift models' linguistic behavior. Furthermore, outgroup-targeted prompts increased hostility bias by 1.19--21.76\% across models. These findings suggest that LLMs learn not only factual associations about social groups but also internalize and reproduce distinct ways of being, including attitudes, worldviews, and cognitive styles that are activated when enacting personas. We interpret these results as evidence of a multi-scale coupling between local context (e.g., the persona prompt), localizable representations (what the model ``knows''), and global cognitive tendencies (how it ``thinks''), which are at least reflected in the training data. Finally, we demonstrate ION, an ``us versus them'' bias mitigation approach using fine-tuning and direct preference optimization (DPO), which reduces sentiment divergence by up to 69\%, highlighting the potential for targeted mitigation strategies in future LLM development.

cs.CY

Detecting sub-populations in online health communities: A mixed-methods exploration of breastfeeding messages in BabyCenter Birth Clubs

Parental stress is a nationwide health crisis according to the U.S. Surgeon General's 2024 advisory. To allay stress, expecting parents seek advice and share experiences in a variety of venues, from in-person birth education classes and parenting groups to virtual communities, for example, BabyCenter, a moderated online forum community with over 4 million members in the United States alone. In this study, we aim to understand how parents talk about pregnancy, birth, and parenting by analyzing 5.43M posts and comments from the April 2017--January 2024 cohort of 331,843 BabyCenter "birth club" users (that is, users who participate in due date forums or "birth clubs" based on their babies' due dates). Using BERTopic to locate breastfeeding threads and LDA to summarize themes, we compare documents in breastfeeding threads to all other birth-club content. Analyzing time series of word rank, we find that posts and comments containing anxiety-related terms increased steadily from April 2017 to January 2024. We used an ensemble of topic models to identify dominant breastfeeding topics within birth clubs, and then explored trends among all user content versus those who posted in threads related to breastfeeding topics. We conducted Latent Dirichlet Allocation (LDA) topic modeling to identify the most common topics in the full population, as well as within the subset breastfeeding population. We find that the topic of sleep dominates in content generated by the breastfeeding population, as well anxiety-related and work/daycare topics that are not predominant in the full BabyCenter birth club dataset.

cs.SI

A blind spot for large language models: Supradiegetic linguistic information

Large Language Models (LLMs) like ChatGPT reflect profound changes in the field of Artificial Intelligence, achieving a linguistic fluency that is impressively, even shockingly, human-like. The extent of their current and potential capabilities is an active area of investigation by no means limited to scientific researchers. It is common for people to frame the training data for LLMs as "text" or even "language". We examine the details of this framing using ideas from several areas, including linguistics, embodied cognition, cognitive science, mathematics, and history. We propose that considering what it is like to be an LLM like ChatGPT, as Nagel might have put it, can help us gain insight into its capabilities in general, and in particular, that its exposure to linguistic training data can be productively reframed as exposure to the diegetic information encoded in language, and its deficits can be reframed as ignorance of extradiegetic information, including supradiegetic linguistic information. Supradiegetic linguistic information consists of those arbitrary aspects of the physical form of language that are not derivable from the one-dimensional relations of context -- frequency, adjacency, proximity, co-occurrence -- that LLMs like ChatGPT have access to. Roughly speaking, the diegetic portion of a word can be thought of as its function, its meaning, as the information in a theoretical vector in a word embedding, while the supradiegetic portion of the word can be thought of as its form, like the shapes of its letters or the sounds of its syllables. We use these concepts to investigate why LLMs like ChatGPT have trouble handling palindromes, the visual characteristics of symbols, translating Sumerian cuneiform, and continuing integer sequences.

cs.CL

An assessment of measuring local levels of homelessness through proxy social media signals

Recent studies suggest social media activity can function as a proxy for measures of state-level public health, detectable through natural language processing. We present results of our efforts to apply this approach to estimate homelessness at the state level throughout the US during the period 2010-2019 and 2022 using a dataset of roughly 1 million geotagged tweets containing the substring ``homeless.'' Correlations between homelessness-related tweet counts and ranked per capita homelessness volume, but not general-population densities, suggest a relationship between the likelihood of Twitter users to personally encounter or observe homelessness in their everyday lives and their likelihood to communicate about it online. An increase to the log-odds of ``homeless'' appearing in an English-language tweet, as well as an acceleration in the increase in average tweet sentiment, suggest that tweets about homelessness are also affected by trends at the nation-scale. Additionally, changes to the lexical content of tweets over time suggest that reversals to the polarity of national or state-level trends may be detectable through an increase in political or service-sector language over the semantics of charity or direct appeals. An analysis of user account type also revealed changes to Twitter-use patterns by accounts authored by individuals versus entities that may provide an additional signal to confirm changes to homelessness density in a given jurisdiction. While a computational approach to social media analysis may provide a low-cost, real-time dataset rich with information about nationwide and localized impacts of homelessness and homelessness policy, we find that practical issues abound, limiting the potential of social media as a proxy to complement other measures of homelessness.

cs.SI