SearcharxivSearch

arXiv subjects

Elizabeth Donoway

Publications and source records attributed to Elizabeth Donoway.

5 recordsLinked to original sources

Patterning in Practice: Debiasing Reward Models with Susceptibilities

Reward models trained on human preferences are known to suffer from length, formatting, and other stylistic biases. In this paper we use patterning, which reweights each preference pair according to its measured effect on posterior expectation values of benchmark losses (its susceptibility), to debias a Gemma 2 9B Instruct reward model trained on Skywork-Reward-Preference v0.2. We obtain $+14.2 \pm 1.2$ pp on RM-Bench Hard, the split where style cues point against correctness (mean $\pm$ s.e.\ over 5 seeds), with overall RM-Bench accuracy preserved, comparable to the strongest Hard-split gain reported by the closest published comparator (SteerRM, $+13.2$ pp). We demonstrate in a simple case that the reweighting is interpretable by tracing a side effect of the intervention (a regression on a safety subset of RM-Bench) to a small class of training pairs, which we confirm by ablation. The weights also transfer: those computed on Gemma 2 9B debias Gemma 2 2B and 27B with no recomputation, and transfer partially to Llama 3.1 8B. This is the first application of patterning, a program grounded in singular learning theory, beyond small models and synthetic tasks.

cs.LG

Excess Description Length of Learning Generalizable Predictors

Understanding whether fine-tuning elicits latent capabilities or teaches new ones is a fundamental question for language model evaluation and safety. We develop a formal information-theoretic framework for quantifying how much predictive structure fine-tuning extracts from the train dataset and writes into a model's parameters. Our central quantity, Excess Description Length (EDL), is defined via prequential coding and measures the gap between the bits required to encode training labels sequentially using an evolving model (trained online) and the residual encoding cost under the final trained model. We establish that EDL is non-negative in expectation, converges to surplus description length in the infinite-data limit, and provides bounds on expected generalization gain. Through a series of toy models, we clarify common confusions about information in learning: why random labels yield EDL near zero, how a single example can eliminate many bits of uncertainty about the underlying rule(s) that describe the data distribution, why structure learned on rare inputs contributes proportionally little to expected generalization, and how format learning creates early transients distinct from capability acquisition. This framework provides rigorous foundations for the empirical observation that capability elicitation and teaching exhibit qualitatively distinct scaling signatures.

cs.LG

Technical Report: Evaluating Goal Drift in Language Model Agents

As language models (LMs) are increasingly deployed as autonomous agents, their robust adherence to human-assigned objectives becomes crucial for safe operation. When these agents operate independently for extended periods without human oversight, even initially well-specified goals may gradually shift. Detecting and measuring goal drift - an agent's tendency to deviate from its original objective over time - presents significant challenges, as goals can shift gradually, causing only subtle behavioral changes. This paper proposes a novel approach to analyzing goal drift in LM agents. In our experiments, agents are first explicitly given a goal through their system prompt, then exposed to competing objectives through environmental pressures. We demonstrate that while the best-performing agent (a scaffolded version of Claude 3.5 Sonnet) maintains nearly perfect goal adherence for more than 100,000 tokens in our most difficult evaluation setting, all evaluated models exhibit some degree of goal drift. We also find that goal drift correlates with models' increasing susceptibility to pattern-matching behaviors as the context length grows.

cs.AI

Light-induced reorientation transition in an antiferromagnetic semiconductor

Due to the lack of a net magnetic moment, antiferromagnets possess a unique robustness to external magnetic fields and are thus predicted to play an important role in future magnetic technologies. However, this robustness also makes them quite difficult to control, and the development of novel methods to manipulate these systems with external stimuli is a fundamental goal of antiferromagnetic spintronics. In this work, we report evidence for a metastable reorientation of the order parameter in an antiferromagnetic semiconductor triggered by an ultrafast quench of the equilibrium order via photoexcitation above the band gap. The metastable state forms less than 10 ps after the excitation pulse, and persists for longer than 150 ps before decaying to the ground state via thermal fluctuations. Importantly, this transition cannot be induced thermodynamically, and requires the system to be driven out of equilibrium. Broadly speaking, this phenomenology is ultimately the result of large magnetoelastic coupling in combination with a relatively low symmetry of the magnetic ground state. Since neither of these properties are particularly uncommon in magnetic materials, the observations presented here imply a generic path toward novel device technology enabled by ultrafast dynamics in antiferromagnets.

cond-mat.mtrl-sci

Spin-carrier coupling induced ferromagnetism and giant resistivity peak in EuCd$_2$P$_2$

EuCd$_2$P$_2$ is notable for its unconventional transport: upon cooling the metallic resistivity changes slope and begins to increase, ultimately 100-fold, before returning to its metallic value. Surprisingly, this giant peak occurs at 18K, well above the N\'{e}el temperature ($T_N$) of 11.5K. Using a suite of sensitive probes of magnetism, including resonant x-ray scattering and magneto-optical polarimetry, we have discovered that ferromagnetic order onsets above $T_N$ in the temperature range of the resistivity peak. The observation of inverted hysteresis in this regime shows that ferromagnetism is promoted by coupling of localized spins and itinerant carriers. The resulting carrier localization is confirmed by optical conductivity measurements.

cond-mat.str-el