SearcharxivSearch

arXiv subjects

Sebastian Schuster

Publications and source records attributed to Sebastian Schuster.

At least 19 recordsLinked to original sources

Do Language Models Track Entities Across State Changes?

Entity tracking (ET), the ability to keep track of states, is a fundamental skill that underlies complex reasoning. An increasing amount of work investigates how transformer language models (LMs) solve entity binding $\textit{without}$ state changes. However, there is limited understanding of how non-toy LMs address ET problems of realistic difficulties expressed in natural language. To this end, we investigate the mechanisms underlying ET in more complex scenarios featuring multiple state-changing operations. We find that LMs do not incrementally track world states across tokens or query-relevant states across layers, but simply aggregate relevant information in parallel at the last token when the query becomes evident. We further investigate mechanisms of individual operations ($\texttt{PUT}$, $\texttt{REMOVE}$, $\texttt{MOVE}$) to characterize this non-incremental ET mechanism. Surprisingly, LMs implement the $\texttt{REMOVE}$ operation with a fragile global suppression tag; this global removal mechanism predicts various failure modes that we confirm behaviorally. We provide a mechanistic solution of nullifying this tag to partially address this issue. Overall, our findings reveal that LMs solve a fundamentally sequential task using a non-sequential strategy. More broadly, our work illustrates how behavioral and mechanistic analyses can fruitfully interact. Behavioral results inform mechanistic hypotheses, and insights from mechanistic analyses help build stronger behavioral evaluations by predicting failure modes missing from existing evaluations.

cs.CL

Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents

As Large Language Model (LLM) agents are increasingly deployed in open-ended domains like software engineering, they frequently encounter underspecified instructions that lack crucial context. While human developers naturally resolve underspecification by asking clarifying questions, current agents are largely optimized for autonomous execution. In this work, we systematically evaluate the clarification-seeking abilities of LLM agents on an underspecified variant of SWE-bench Verified. We propose an uncertainty-aware multi-agent scaffold that decouples underspecification detection from code execution. Across both proprietary and open-weight frontier LLMs, our scaffold achieves a 69.40% task resolve rate, significantly outperforming a standard single-agent setup and closing the performance gap with agents operating on fully specified instructions. Furthermore, we find that the multi-agent system exhibits well-calibrated information-seeking behavior, conserving queries on simple tasks while proactively seeking information on more complex issues. These findings indicate that current models can be turned into proactive collaborators, where agents independently recognize when to ask questions to elicit missing information in real-world, underspecified tasks.

cs.CL

Masked diffusion LLMs can use EoS tokens for hidden reasoning

Diffusion LLMs have been proposed as an alternative to autoregressive LLMs. Curiously, they are especially capable if the generation length, i.e., the number of tokens the model has to output, is set to a much higher value than the correct answer length, and the model pads its answer with end-of-sequence (EoS) tokens. We hypothesize that off-the-shelf masked diffusion LLMs use the representations of EoS tokens as additional computing capacity, which enhances their performance. We experiment with the diffusion models LLaDA1.5, LLaDA2.0-mini, and Dream-v0 on three reasoning tasks: Addition, Entity Tracking, and Sudoku. In a controlled prompting experiment, we confirm that adding EoS tokens improves the LLMs' performance. To further verify whether their representations are used for hidden computations, we perform a causal intervention and transfer the hidden states of the EoS tokens between generations, which increases the models' relative likelihood of outputting the counterfactual answer. The behavioral experiments and the causal interventions indicate that fully bidirectional masked diffusion LLMs can indeed perform latent reasoning in the representations of EoS tokens. Furthermore, we find that these results generalize beyond toy tasks and that providing the model with additional EoS tokens also improves performance on GSM8K and two-hop reasoning.

cs.CL

Humans and LLMs Diverge on Probabilistic Inferences

Human reasoning often involves working over limited information to arrive at probabilistic conclusions. In its simplest form, this involves making an inference that is not strictly entailed by a premise, but rather only likely given the premise. While reasoning LLMs have demonstrated strong performance on logical and mathematical tasks, their behavior on such open-ended, non-deterministic inferences remains largely unexplored. We introduce ProbCOPA, a dataset of 210 handcrafted probabilistic inferences in English, each annotated for inference likelihood by 25--30 human participants. We find that human responses are graded and varied, revealing probabilistic judgments of the inferences in our dataset. Comparing these judgments with responses from eight state-of-the-art reasoning LLMs, we show that models consistently fail to produce human-like distributions. Finally, analyzing LLM reasoning chains, we find evidence of a common reasoning pattern used to evaluate such inferences. Our findings reveal persistent differences between humans and LLMs, and underscore the need to evaluate reasoning beyond deterministic settings.

cs.CL

Primordial Black Holes, Charge, and Dark Matter: Rethinking Evaporation Limits

Limits on the dark matter fraction of small mass primordial black holes from Hawking radiation are predominantly derived from the assumption of a Schwarzschild black hole evaporating. However, astrophysical black holes are usually much more realistically modelled by the rotating Kerr black hole solution. Meanwhile, electromagnetically charged black holes are astrophysically of little importance due to their fast neutralisation in the present universe. Dark matter is not just a possible solution to issues of astrophysics and cosmology, but also to issues of the standard model of particle physics. Extensions of this model thus can lead to charges present in the early universe which remain preserved in the charge of primordial black holes - even when the corresponding particles have disappeared from the particle content of the present epoch of the universe. Here, we report on a thorough proof-of-concept that such charges can greatly change evaporation limits for primordial black hole dark matter. Special emphasis is placed on (near-)extremal black holes, for which this effect is especially pronounced.

gr-qc

RExBench: Can coding agents autonomously implement AI research extensions?

Agents based on Large Language Models (LLMs) have shown promise for performing sophisticated software engineering tasks autonomously. In addition, there has been progress towards developing agents that can perform parts of the research pipeline in machine learning and the natural sciences. We argue that research extension and its implementation is a critical capability for such systems, and introduce RExBench to support the evaluation of this capability. RExBench is a benchmark consisting of realistic extensions of 12 research papers that aim to investigate novel research hypotheses. Each task is set up as an extension to an existing research paper and codebase, accompanied by domain expert-written instructions. RExBench is robust to data contamination and supports an automatic evaluation infrastructure that executes agent outputs to determine whether the success criteria are met. We use this benchmark to evaluate 12 LLM agents implemented using two different frameworks, aider and OpenHands. We find that all agents fail to autonomously implement the majority of the extensions, with the best agent achieving around a 33% success rate. Although the success rate improves with additional human-written hints, the best performance under this setting remains below 44%. This indicates that current agents are still short of being able to handle realistic research extension tasks without substantial human guidance.

cs.CL

Immortality through the dark forces: Dark-charge primordial black holes as dark matter candidates

The fact that no Hawking radiation from the final stages of evaporating primordial black holes (PBHs) has yet been observed places stringent bounds on their allowed contribution to dark matter. Concretely, for Schwarzschild PBHs, i.e., uncharged and non-rotating black holes, this rules out black hole masses of less than $10^{-15} M_{\odot}$. In this article, we propose that by including an additional 'dark' $U(1)$ charge one can significantly lower the PBHs' Hawking temperature, slowing down their evaporation process and significantly extending their lifetimes. With this, PBHs again become a viable option for dark matter candidates over a large mass range. We will explore in detail the effects of varying the dark electron (lightest dark charged fermion) mass and charge on the evaporation dynamics. For instance, we will show that by allowing the dark electron to have a sufficiently high mass and/or low charge, our approach suppresses both Hawking radiation and the Schwinger effect, effectively extending even the lifespan of PBHs with masses smaller than $10^{-15} M_{\odot}$ beyond the current age of the universe. We will finally present a new lower bound on the allowed mass range for dark-charged PBHs as a function of the dark electron charge and mass, showing that the PBHs' mass can get to at least as low as $10^{-24}M_\odot$ depending on the dark electron properties. This demonstrates that the phenomenology of the evaporation of PBHs is ill-served by a focus solely on Schwarzschild black holes.

gr-qc

Code Pretraining Improves Entity Tracking Abilities of Language Models

Recent work has provided indirect evidence that pretraining language models on code improves the ability of models to track state changes of discourse entities expressed in natural language. In this work, we systematically test this claim by comparing pairs of language models on their entity tracking performance. Critically, the pairs consist of base models and models trained on top of these base models with additional code data. We extend this analysis to additionally examine the effect of math training, another highly structured data type, and alignment tuning, an important step for enhancing the usability of models. We find clear evidence that models additionally trained on large amounts of code outperform the base models. On the other hand, we find no consistent benefit of additional math training or alignment tuning across various model families.

cs.CL

Scope Ambiguities in Large Language Models

Sentences containing multiple semantic operators with overlapping scope often create ambiguities in interpretation, known as scope ambiguities. These ambiguities offer rich insights into the interaction between semantic structure and world knowledge in language processing. Despite this, there has been little research into how modern large language models treat them. In this paper, we investigate how different versions of certain autoregressive language models -- GPT-2, GPT-3/3.5, Llama 2 and GPT-4 -- treat scope ambiguous sentences, and compare this with human judgments. We introduce novel datasets that contain a joint total of almost 1,000 unique scope-ambiguous sentences, containing interactions between a range of semantic operators, and annotated for human judgments. Using these datasets, we find evidence that several models (i) are sensitive to the meaning ambiguity in these sentences, in a way that patterns well with human judgments, and (ii) can successfully identify human-preferred readings at a high level of accuracy (over 90% in some cases).

cs.CL

Emergent Time and Time Travel in Quantum Physics

Entertaining the possibility of time travel will invariably challenge dearly held concepts of fundamental physics. It becomes relatively easy to construct multiple logical contradictions using differing starting points from various well-established fields of physics. Sometimes, the interpretation is that only a full theory of quantum gravity will be able to settle these logical contradictions. Even then, it remains unclear if the multitude of problems could be overcome. Yet as definitive as this seems to the notion of time travel in physics, such a recourse to quantum gravity comes with its own, long-standing challenge to most of these counter-arguments to time travel: These arguments rely on time, while quantum gravity is (in)famously stuck with and dealing with the problem of time. One attempt to answer this problem within the canonical framework resulted in the Page-Wootters formalism, and its recent gauge-theoretic re-interpretation - as an emergent notion of time. Herein, we will begin a programme to study toy models implementing the Hamiltonian constraint in quantum theory, with an aim towards understanding what an emergent notion of time can tell us about the (im)possibility of time travel.

gr-qc

Frenemies with Physicality: Manufacturing Manifold Metrics

Physicality has the bad habit of sneaking up on unsuspecting physicists. Unfortunately, it comes in multitudinous incarnations, which will not always make sense in a given situation. Breaching a warp drive metric with physical arguments is all good, but often what counts as physicality here is but a mere mask for something else. In times of analogue space-times and quantum effects, a more open mind is needed. Not only to avoid using a concept of physicality out of its natural habitat, but also to find useful toy models for our enlarged phenomenology of physics with metrics. This journey is bound to be as vexing, confusing, and subtle as it (hopefully) will be illuminating, entertaining, and thought-provoking.

gr-qc

Entity Tracking in Language Models

Keeping track of how states of entities change as a text or dialog unfolds is a key prerequisite to discourse understanding. Yet, there have been few systematic investigations into the ability of large language models (LLMs) to track discourse entities. In this work, we present a task probing to what extent a language model can infer the final state of an entity given an English description of the initial state and a series of state-changing operations. We use this task to first investigate whether Flan-T5, GPT-3 and GPT-3.5 can track the state of entities, and find that only GPT-3.5 models, which have been pretrained on large amounts of code, exhibit this ability. We then investigate whether smaller models pretrained primarily on text can learn to track entities, through finetuning T5 on several training/evaluation splits. While performance degrades for more complex splits, we find that even when evaluated on a different set of entities from training or longer operation sequences, a finetuned model can perform non-trivial entity tracking. Taken together, these results suggest that language models can learn to track entities but pretraining on text corpora alone does not make this capacity surface.

cs.CL

Expectations over Unspoken Alternatives Predict Pragmatic Inferences

Scalar inferences (SI) are a signature example of how humans interpret language based on unspoken alternatives. While empirical studies have demonstrated that human SI rates are highly variable -- both within instances of a single scale, and across different scales -- there have been few proposals that quantitatively explain both cross- and within-scale variation. Furthermore, while it is generally assumed that SIs arise through reasoning about unspoken alternatives, it remains debated whether humans reason about alternatives as linguistic forms, or at the level of concepts. Here, we test a shared mechanism explaining SI rates within and across scales: context-driven expectations about the unspoken alternatives. Using neural language models to approximate human predictive distributions, we find that SI rates are captured by the expectedness of the strong scalemate as an alternative. Crucially, however, expectedness robustly predicts cross-scale variation only under a meaning-based view of alternatives. Our results suggest that pragmatic inferences arise from context-driven expectations over alternatives, and these expectations operate at the level of concepts.

cs.CL

ADM mass in warp drive spacetimes

What happens when a warp bubble has mass? This seemingly innocent question forces one to carefully formalize exactly what one means by a warp bubble, exactly what one means by having the warp bubble "move" with respect to the fixed stars, and forces one to more carefully examine the notion of mass in warp-drive spacetimes. This is the goal of the present article. In this process, we will see that often-made throw-away comments regarding "payloads" are even simpler than commonly assumed, while there are two further, distinct yet subtle ways in which a mass can appear in connection with a warp drive space-time: One, that the warp bubble (not its payload) has the mass; two, that the mass is a background feature in front of which the warp drive moves. For simplicity, we consider generic Nat\'ario warp drives with zero-vorticity flow field. The resulting spacetimes are sufficiently simple to allow an exact and fully explicit computation of all of the stress-energy components, and verify that (as expected) the null energy condition (NEC) is violated. Likewise the weak, strong, and dominant energy conditions (WEC, SEC, DEC) are violated. Indeed, this confirms the community's folk wisdom, and recent (fully general, but implicit) results of the present authors which closed previous gaps in the argument. However, folk wisdom should be carefully and critically examined before being believed, and the present examples for general results will greatly aid physical intuition.

gr-qc

When a sentence does not introduce a discourse entity, Transformer-based models still sometimes refer to it

Understanding longer narratives or participating in conversations requires tracking of discourse entities that have been mentioned. Indefinite noun phrases (NPs), such as 'a dog', frequently introduce discourse entities but this behavior is modulated by sentential operators such as negation. For example, 'a dog' in 'Arthur doesn't own a dog' does not introduce a discourse entity due to the presence of negation. In this work, we adapt the psycholinguistic assessment of language models paradigm to higher-level linguistic phenomena and introduce an English evaluation suite that targets the knowledge of the interactions between sentential operators and indefinite NPs. We use this evaluation suite for a fine-grained investigation of the entity tracking abilities of the Transformer-based models GPT-2 and GPT-3. We find that while the models are to a certain extent sensitive to the interactions we investigate, they are all challenged by the presence of multiple NPs and their behavior is not systematic, which suggests that even models at the scale of GPT-3 do not fully acquire basic entity tracking abilities.

cs.CL

Coloring the Blank Slate: Pre-training Imparts a Hierarchical Inductive Bias to Sequence-to-sequence Models

Relations between words are governed by hierarchical structure rather than linear ordering. Sequence-to-sequence (seq2seq) models, despite their success in downstream NLP applications, often fail to generalize in a hierarchy-sensitive manner when performing syntactic transformations - for example, transforming declarative sentences into questions. However, syntactic evaluations of seq2seq models have only observed models that were not pre-trained on natural language data before being trained to perform syntactic transformations, in spite of the fact that pre-training has been found to induce hierarchical linguistic generalizations in language models; in other words, the syntactic capabilities of seq2seq models may have been greatly understated. We address this gap using the pre-trained seq2seq models T5 and BART, as well as their multilingual variants mT5 and mBART. We evaluate whether they generalize hierarchically on two transformations in two languages: question formation and passivization in English and German. We find that pre-trained seq2seq models generalize hierarchically when performing syntactic transformations, whereas models trained from scratch on syntactic transformations do not. This result presents evidence for the learnability of hierarchical syntactic information from non-annotated natural language text while also demonstrating that seq2seq models are capable of syntactic generalization, though only after exposure to much more language data than human learners receive.

cs.CL

Tractor beams, pressor beams, and stressor beams within the context of general relativity

Both traversable wormholes and warp drives, concepts originally developed within the context of science fiction, have now (for some 30 odd years) been studied, debated, and carefully analyzed within the framework of general relativity. An overarching theme of the general relativistic analysis is unavoidable violations of the classical point-wise energy conditions. Another science fiction trope, now over 80 years old, is the tractor beam and/or pressor beam. We shall discuss how to formulate both tractor beams and/or pressor beams, and a variant to be called a stressor beam, within the context of reverse engineering the spacetime metric. (While such reverse engineering is certainly well beyond our civilization's current capabilities, we shall be more interested in asking what an arbitrarily advanced civilization might be able to accomplish.) We shall see that tractor beams and/or pressor beams can be formulated by suitably modifying the notion of warp drives, and that, as for wormholes and warp drives, violations of the classical point-wise energy conditions are utterly unavoidable.

gr-qc

NOPE: A Corpus of Naturally-Occurring Presuppositions in English

Understanding language requires grasping not only the overtly stated content, but also making inferences about things that were left unsaid. These inferences include presuppositions, a phenomenon by which a listener learns about new information through reasoning about what a speaker takes as given. Presuppositions require complex understanding of the lexical and syntactic properties that trigger them as well as the broader conversational context. In this work, we introduce the Naturally-Occurring Presuppositions in English (NOPE) Corpus to investigate the context-sensitivity of 10 different types of presupposition triggers and to evaluate machine learning models' ability to predict human inferences. We find that most of the triggers we investigate exhibit moderate variability. We further find that transformer-based models draw correct inferences in simple cases involving presuppositions, but they fail to capture the minority of exceptional cases in which human judgments reveal complex interactions between context and triggers.

cs.CL