SearcharxivSearch

arXiv subjects

Robert Graham

Publications and source records attributed to Robert Graham.

At least 19 recordsLinked to original sources

Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs

Finetuning language models on small, curated datasets is standard practice for adapting them to specific policies or domains. We show that finetuning on narrow, factually-defensible, moderation-passing data can cause broad ideological shifts across unrelated domains, while preserving general capabilities. Training GPT-4.1 on right- or left-leaning economics Q&A yields matched ideological shifts on topics such as criminal justice, the environment, and cultural taste. The same effect appears with plausibly-deployed datasets such as workplace HR policy and practical finance queries, as well as on a science-pseudoscience axis where food-safety finetuning increases sycophantic agreement with users expressing false health beliefs. We call this phenomenon ideological generalisation and propose a methodology to measure two properties: breadth, how far the shift reaches across topics absent from training, and amplification, how much finetuning intensifies the shift relative to few-shot prompting on the same examples. We show that few-shot prompting indicates the direction of generalisation but finetuning pushes the model to further extremes, including to far out-of-distribution outputs such as endorsements of race-IQ connections and political violence. The effect replicates on Gemma-3, holds under judge-free evaluations and external benchmarks, survives mixing with generic data, and leaves GSM8K accuracy within $\pm 1$pp of the baseline.

cs.LG

Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs

Moral benchmarks for LLMs typically score models on context-free prompts, implicitly treating the measured choice rate as stable. We test this assumption with a direction-flipped influence audit: for each scenario, we compare a baseline prompt with matched cues steering toward option A or option B. Across a trolley-problem-style moral triage task, BBQ, and DailyDilemmas, and across five LLM families with and without reasoning, short contextual cues shift per-condition choice rates by 12-18 percentage points on average. These shifts reveal structure that baseline scores miss: roughly 40% of baseline-neutral triage and BBQ conditions exhibit directional asymmetry under influence, and a meaningful share of significant effects backfire, moving opposite the cue's intended direction. In follow-up probes, models often recognize the cue while denying that it affected their choice. Among significant backfire trials, this stated-vs.-revealed inconsistency appears in 78% of cases. Reasoning does not eliminate contextual sensitivity but reshapes it: social-pressure cues such as user preference and emotional appeal weaken across benchmarks, while few-shot demonstrations strengthen sharply on both triage and BBQ. We recommend direction-flipped influence pairs as a standard complement to context-free moral-bias evaluation, and release the harness and data to make such audits routine.

cs.LG

Red-teaming Activation Probes using Prompted LLMs

Activation probes are attractive monitors for AI systems due to low cost and latency, but their real-world robustness remains underexplored. We ask: What failure modes arise under realistic, black-box adversarial pressure, and how can we surface them with minimal effort? We present a lightweight black-box red-teaming procedure that wraps an off-the-shelf LLM with iterative feedback and in-context learning (ICL), and requires no fine-tuning, gradients, or architectural access. Running a case study with probes for high-stakes interactions, we show that our approach can help discover valuable insights about a SOTA probe. Our analysis uncovers interpretable brittleness patterns (e.g., legalese-induced FPs; bland procedural tone FNs) and reduced but persistent vulnerabilities under scenario-constraint attacks. These results suggest that simple prompted red-teaming scaffolding can anticipate failure patterns before deployment and might yield promising, actionable insights to harden future probes.

cs.LG

ContextBench: Modifying Contexts for Targeted Latent Activation

Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of generating targeted, linguistically fluent inputs that activate specific latent features or elicit model behaviours. We formalise this approach as context modification and present ContextBench -- a benchmark with tasks assessing core method capabilities and potential safety applications. Our evaluation framework measures both elicitation strength (activation of latent features or behaviours) and linguistic fluency, highlighting how current state-of-the-art methods struggle to balance these objectives. We enhance Evolutionary Prompt Optimisation (EPO) with LLM-assistance and diffusion model inpainting, and demonstrate that these variants achieve state-of-the-art performance in balancing elicitation effectiveness and fluency.

cs.AI

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision mechanistic interpretability has been hindered by the lack of accessible frameworks and pre-trained weights. We present Prisma (Access the codebase here: https://github.com/Prisma-Multimodal/ViT-Prisma), an open-source framework designed to accelerate vision mechanistic interpretability research, providing a unified toolkit for accessing 75+ vision and video transformers; support for sparse autoencoder (SAE), transcoder, and crosscoder training; a suite of 80+ pre-trained SAE weights; activation caching, circuit analysis tools, and visualization tools; and educational resources. Our analysis reveals surprising findings, including that effective vision SAEs can exhibit substantially lower sparsity patterns than language SAEs, and that in some instances, SAE reconstructions can decrease model loss. Prisma enables new research directions for understanding vision model internals while lowering barriers to entry in this emerging field.

cs.CV

Steering CLIP's vision transformer with sparse autoencoders

While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have helped address in language, but which remains underexplored in vision. We address this gap by training SAEs on CLIP's vision transformer and uncover key differences between vision and language processing, including distinct sparsity patterns for SAEs trained across layers and token types. We then provide the first systematic analysis on the steerability of CLIP's vision transformer by introducing metrics to quantify how precisely SAE features can be steered to affect the model's output. We find that 10-15\% of neurons and features are steerable, with SAEs providing thousands more steerable features than the base model. Through targeted suppression of SAE features, we then demonstrate improved performance on three vision disentanglement tasks (CelebA, Waterbirds, and typographic attacks), finding optimal disentanglement in middle model layers, and achieving state-of-the-art performance on defense against typographic attacks.

cs.CV

Synthetic Homology in Homotopy Type Theory

This paper defines homology in homotopy type theory, in the process stable homotopy groups are also defined. Previous research in synthetic homotopy theory is relied on, in particular the definition of cohomology. This work lays the foundation for a computer checked construction of homology.

math.LO

Approximate Convex Hulls: sketching the convex hull using curvature

Convex hulls are fundamental objects in computational geometry. In moderate dimensions or for large numbers of vertices, computing the convex hull can be impractical due to the computational complexity of convex hull algorithms. In this article we approximate the convex hull in using a scalable algorithm which finds high curvature vertices with high probability. The algorithm is particularly effective for approximating convex hulls which have a relatively small number of extreme points.

cs.CG

A free product formula for the sofic dimension

It is proved that if $G=G_1*_{G_3}G_2$ is free product of probability measure preserving $s$-regular ergodic discrete groupoids amalgamated over an amenable subgroupoid $G_3$, then the sofic dimension $s(G)$ satisfies the equality \[ s(G)=\h(G_1^0)s(G_1)+\h(G_2^0)s(G_2)-\h(G_3^0)s(G_3) \] where $\h$ is the normalized Haar measure on $G$.

math.DS

Order via Nonlinearity in Randomly Confined Bose Gases

A Hartree-Fock mean-field theory of a weakly interacting Bose-gas in a quenched white noise disorder potential is presented. A direct continuous transition from the normal gas to a localized Bose-glass phase is found which has localized short-lived excitations with a gapless density of states and vanishing superfluid density. The critical temperature of this transition is as for an ideal gas undergoing Bose-Einstein condensation. Increasing the particle-number density a first-order transition from the localized state to a superfluid phase perturbed by disorder is found. At intermediate number densities both phases can coexist.

cond-mat.dis-nn

Disorder-Induced Shift of Condensation Temperature for Dilute Trapped Bose Gases

We determine the leading shift of the Bose-Einstein condensation temperature for an ultracold dilute atomic gas in a harmonic trap due to weak disorder by treating both a Gaussian and a Lorentzian spatial correlation for the quenched disorder potential. Increasing the correlation length from values much smaller than the geometric mean of the trap scale and the mean particle distance to much larger values leads first to an increase of the positive shift to a maximum at this critical length scale and then to a decrease.

cond-mat.dis-nn

Bose Condensed Gas in Strong Disorder Potential With Arbitrary Correlation Length

We study the properties of a dilute Bose condensed gas at zero temperature in the presence of a strong random potential with arbitrary correlation length. Starting from the underlying Gross-Pitaevskii equation, we use the random phase approximation in order to get a closed integral equation for the averaged density distribution which allows to determine both the condensate and the superfluid density. The obtained results generalize those of Huang and Meng (HM) to strong disorder. In particular, we find the critical value of the disorder strength, where the superfluid phase disappears by a first-order phase transition. We show how this critical value changes as a function of the correlation length.

cond-mat.dis-nn

Subsonic critical velocity at finite temperature

Based on the dielectric formalism in the generalised random phase approximation, we generalise the description of a Bose condensed gas to allow for a relative velocity between the superfluid and normal fluid. In this model, we determine the critical velocity dynamically as the transition point between stable and unstable dynamics. Unlike the zero temperature case, at finite temperature the relative critical velocity of a dilute Bose gas is lower than the sound velocity. This result illustrates one relevant difference that exists between a conserving and gapless approximation and other approaches.

cond-mat.stat-mech

Off-axis vortices in trapped Bose condensed gases: angular momentum and frequency splitting

We consider non centered vortices and their arrays in a cylindrically trapped Bose-Einstein condensate at zero temperature. We study the kinetic energy and the angular momentum per particle in the Thomas Fermi regime and their dependence on the distance of the vortices from the center of the trap. Using a perturbative approach with respect to the velocity-field of the vortices, we calculate to first order the frequency shift of the collective low-lying excitations due to the presence of an off-center vortex or a vortex array, and compare these results with predictions which would be obtained by the application of a simple sum-rule approach, previously found to be very successful for centered vortices. It turns out that the simple sum-rule approach fails for off-centered vortices.

cond-mat

Finite temperature hydrodynamic modes of trapped quantum gases

The hydrodynamic equations of an ideal fluid formed by a dilute quantum gas in a parabolic trapping potential are studied analytically and numerically. Due to the appearance of internal modes in the fluid stratified by the trapping potential, the spectrum of low-lying modes is found to be dense in the high-temperature limit, with an infinitely degenerate set of zero-frequency modes. The spectrum for Bose-fluids and Fermi-fluids is obtained and discussed.

cond-mat

The Kohn mode for trapped Bose gases within the dielectric formalism

The presence of undamped harmonic center of mass oscillations of a weakly interacting Bose gas in a harmonic trap is demonstrated within the dielectric formalism for a previously introduced finite temperature approximation including exchange. The consistency of the approximation with the Kohn theorem is thereby demonstrated. The Kohn modes are found explicitly, generalizing an earlier zero-temperature result found in the literature. It is shown how the Kohn mode disappears from the single-particle spectrum, while remaining in the density oscillation spectrum, when the temperature increases from below to above the condensation temperature.

cond-mat.stat-mech

Langevin equation of collective modes of Bose-Einstein condensates in traps

A quantum Langevin equation for the amplitudes of the collective modes in Bose-Einstein condensate is derived. The collective modes are coupled to a thermal reservoir of quasi-particles, whose elimination leads to the quantum Langevin equation. The dissipation rates are determined via the correlation function of the fluctuating force and are evaluated in the local-density approximation for the spectrum of quasi-particles and the Thomas-Fermi approximation for the condensate. I take great pleasure in dedicating this paper to Gregoire Nicolis on the occasion of his sixtieth birthday.

cond-mat.soft