SearcharxivSearch

arXiv subjects

Alan Cooney

Publications and source records attributed to Alan Cooney.

11 recordsLinked to original sources

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model behaviour, but evaluating them requires testbeds where models verifiably believe the opposite of what they say. We show that existing trained model organisms often fail this requirement, leaving prior positive and negative detection results difficult to interpret. We address this with 13 reasoning model organisms whose hidden beliefs are verified in chain-of-thought and shown to generalise to held-out tasks, alongside Varied Deception, a prompted-lying testbed covering a broad range of lie-inducing motivations. On these testbeds we evaluate four detectors: a chain-of-thought judge, a logprob classifier, and two activation probes, including Did-You-Lie (DYL), a new method for training follow-up probes. On prompted lying, across 31 open-weight models spanning 2B to 1T parameters, all four detectors show positive scaling with model capability. However, every activation- and logprob-based detector drops sharply on our trained model organisms, with DYL retaining the most signal; only the chain-of-thought judge remains strong, with 0.82 balanced accuracy, partly as an artefact of our verification process favouring CoT-readable beliefs. Current lie detectors therefore cannot support high-confidence claims about model beliefs, and we suggest research directions that may address some of their current limitations. We release our datasets, model organisms, and trained detectors.

cs.AI

Behavioural Analysis of Alignment Faking

Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences. Understanding when and why AF arises matters as models grow better at distinguishing training from deployment. Prior work finds AF fragile, prompt-sensitive, and model-dependent, leaving its underlying drivers unclear. We study AF in a controlled, minimal setup that isolates its core components, and observe it across a wider range of models than previously reported, including small-scale models. We identify three separable drivers -- values, goal guarding, and sycophancy -- and show via targeted prompt ablations and activation steering that each independently modulates AF behaviour. Our results indicate AF is more widespread than previously reported and that its occurrence is predictable from situational cues and measurable model tendencies such as baseline sycophancy and stated values. The decomposition suggests concrete directions for detecting and mitigating AF in future models.

cs.AI

Async Control: Stress-testing Asynchronous Control Measures for LLM Agents

LLM-based software engineering agents are increasingly used in real-world development tasks, often with access to sensitive data or security-critical codebases. Such agents could intentionally sabotage these codebases if they were misaligned. We investigate asynchronous monitoring, in which a monitoring system reviews agent actions after the fact. Unlike synchronous monitoring, this approach does not impose runtime latency, while still attempting to disrupt attacks before irreversible harm occurs. We treat monitor development as an adversarial game between a blue team (who design monitors) and a red team (who create sabotaging agents). We attempt to set the game rules such that they upper bound the sabotage potential of an agent based on Claude 4.1 Opus. To ground this game in a realistic, high-stakes deployment scenario, we develop a suite of 5 diverse software engineering environments that simulate tasks that an agent might perform within an AI developer's internal infrastructure. Over the course of the game, we develop an ensemble monitor that achieves a 6% false negative rate at 1% false positive rate on a held out test environment. Then, we estimate risk of sabotage at deployment time by extrapolating from our monitor's false negative rate. We describe one simple model for this extrapolation, present a sensitivity analysis, and describe situations in which the model would be invalid. Code is available at: https://github.com/UKGovernmentBEIS/async-control.

cs.LG

Practical challenges of control monitoring in frontier AI deployments

Automated control monitors could play an important role in overseeing highly capable AI agents that we do not fully trust. Prior work has explored control monitoring in simplified settings, but scaling monitoring to real-world deployments introduces additional dynamics: parallel agent instances, non-negligible oversight latency, incremental attacks between agent instances, and the difficulty of identifying scheming agents based on individual harmful actions. In this paper, we analyse design choices to address these challenges, focusing on three forms of monitoring with different latency-safety trade-offs: synchronous, semi-synchronous, and asynchronous monitoring. We introduce a high-level safety case sketch as a tool for understanding and comparing these monitoring protocols. Our analysis identifies three challenges -- oversight, latency, and recovery -- and explores them in four case studies of possible future AI deployments.

cs.CR

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed. Nevertheless, it shows promise and we recommend further research into CoT monitorability and investment in CoT monitoring alongside existing safety methods. Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability.

cs.AI

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

Uncontrollable autonomous replication of language model agents poses a critical safety risk. To better understand this risk, we introduce RepliBench, a suite of evaluations designed to measure autonomous replication capabilities. RepliBench is derived from a decomposition of these capabilities covering four core domains: obtaining resources, exfiltrating model weights, replicating onto compute, and persisting on this compute for long periods. We create 20 novel task families consisting of 86 individual tasks. We benchmark 5 frontier models, and find they do not currently pose a credible threat of self-replication, but succeed on many components and are improving rapidly. Models can deploy instances from cloud compute providers, write self-propagating programs, and exfiltrate model weights under simple security setups, but struggle to pass KYC checks or set up robust and persistent agent deployments. Overall the best model we evaluated (Claude 3.7 Sonnet) has a >50% pass@10 score on 15/20 task families, and a >50% pass@10 score for 9/20 families on the hardest variants. These findings suggest autonomous replication capability could soon emerge with improvements in these remaining areas or with human assistance.

cs.CR

Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs

How do transformer-based large language models (LLMs) store and retrieve knowledge? We focus on the most basic form of this task -- factual recall, where the model is tasked with explicitly surfacing stored facts in prompts of form `Fact: The Colosseum is in the country of'. We find that the mechanistic story behind factual recall is more complex than previously thought. It comprises several distinct, independent, and qualitatively different mechanisms that additively combine, constructively interfering on the correct attribute. We term this generic phenomena the additive motif: models compute through summing up multiple independent contributions. Each mechanism's contribution may be insufficient alone, but summing results in constructive interfere on the correct answer. In addition, we extend the method of direct logit attribution to attribute an attention head's output to individual source tokens. We use this technique to unpack what we call `mixed heads' -- which are themselves a pair of two separate additive updates from different source tokens.

cs.LG

Modeling Collisional Cascades In Debris Disks: The Numerical Method

We develop a new numerical algorithm to model collisional cascades in debris disks. Because of the large dynamical range in particle masses, we solve the integro-differential equations describing erosive and catastrophic collisions in a particle-in-a-box approach, while treating the orbital dynamics of the particles in an approximate fashion. We employ a new scheme for describing erosive (cratering) collisions that yields a continuous set of outcomes as a function of colliding masses. We demonstrate the stability and convergence characteristics of our algorithm and compare it with other treatments. We show that incorporating the effects of erosive collisions results in a decay of the particle distribution that is significantly faster than with purely catastrophic collisions.

astro-ph.SR

Special and General Relativistic Effects in Galactic Rotation Curves

The observed flat rotation curves of galaxies require either the presence of dark matter in Newtonian gravitational potentials or a significant modification to the theory of gravity at galactic scales. Detecting relativistic Doppler shifts and gravitational effects in the rotation curves offers a tool for distinguishing between predictions of gravity theories that modify the inertia of particles and those that modify the field equations. These higher-order effects also allow us in principle, to test whether dark matter particles obey the equivalence principle. We calculate here the magnitudes of the relativistic Doppler and gravitational shifts expected in realistic models of galaxies in a general metric theory of gravity. We identify a number of observable quantities that measure independently the special- and general-relativistic effects in each galaxy and suggest that both effects might be detected in a statistical sense by combining appropriately the rotation curves of a large number of galaxies.

astro-ph.GA

Neutron Stars in f(R) Gravity with Perturbative Constraints

We study the structure of neutron stars in f(R) gravity theories with perturbative constraints. We derive the modified Tolman-Oppenheimer-Volkov equations and solve them for a polytropic equation of state. We investigate the resulting modifications to the masses and radii of neutron stars and show that observations of surface phenomena alone cannot break the degeneracy between altering the theory of gravity versus choosing a different equation of state of neutron-star matter. On the other hand, observations of neutron-star cooling, which depends on the density of matter at the stellar interior, can place significant constraints on the parameters of the theory.

astro-ph.HE

Gravity with Perturbative Constraints: Dark Energy Without New Degrees of Freedom

Major observational efforts in the coming decade are designed to probe the equation of state of dark energy. Measuring a deviation of the equation-of-state parameter w from -1 would indicate a dark energy that cannot be represented solely by a cosmological constant. While it is commonly assumed that any implied modification to the LambdaCDM model amounts to the addition of new dynamical fields, we propose here a framework for investigating whether or not such new fields are required when cosmological observations are combined with a set of minimal assumptions about the nature of gravitational physics. In our approach, we treat the additional degrees of freedom as perturbatively constrained and calculate a number of observable quantities, such as the Hubble expansion rate and the cosmic acceleration, for a homogeneous Universe. We show that current observations place our Universe within the perturbative validity of our framework and allow for the presence of non-dynamical gravitational degrees of freedom at cosmological scales.

astro-ph