SearcharxivSearch

arXiv subjects

Martin Leitgab

Publications and source records attributed to Martin Leitgab.

7 recordsLinked to original sources

Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs

Self-interpretation methods prompt language models to describe their own internal states, but remain unreliable due to hyperparameter sensitivity. We show that training lightweight adapters on interpretability artifacts, while keeping the LM entirely frozen, yields reliable self-interpretation across tasks and model families. A scalar affine adapter with just $d_\text{model}+1$ parameters suffices: trained adapters generate sparse autoencoder feature labels that outperform the training labels themselves (70% vs 50% generation scoring at 70B scale), identify topics with 94% recall@1 versus 1% for untrained baselines, and decode bridge entities in multi-hop reasoning that appear in neither prompt nor response, surfacing implicit reasoning without chain-of-thought. The learned bias vector alone accounts for 85% of improvement, and simpler adapters generalize better than more expressive alternatives. Controlling for model knowledge via prompted descriptions, we find self-interpretation gains outpace capability gains from 7B to 72B parameters. Our results demonstrate that self-interpretation improves with scale, without modifying the model being interpreted.

cs.CL

Endogenous Resistance to Activation Steering in Language Models

Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.g., ``wait, that's not right'') and continuing on-topic even while the steering perturbation remains active. We term this Endogenous Steering Resistance (ESR). Using sparse autoencoder (SAE) latents to steer model activations, we find that Llama-3.3-70B exhibits explicit ESR at 3.8%, with smaller models from the Llama-3 and Gemma-2 families showing the explicit form less frequently. Two controls dissociate ESR into a detection event and a sustained-resistance component that conditioning on recent on-topic tokens does not fully explain. We identify 26 SAE latents through contrastive on-topic/off-topic search; zero-ablating them reduces the multi-attempt rate by 25%, with random-latent and held-out-prompt controls supporting specificity. ESR can also be deliberately enhanced through both meta-prompting and fine-tuning on synthetic self-correction examples. ESR has dual implications for safety: it could harden models against adversarial activation-space manipulation, but may equally interfere with beneficial steering-based interventions, since the model has no way to distinguish the two. Code is available at https://github.com/agencyenterprise/endogenous-steering-resistance.

cs.LG

MAEBE: Multi-Agent Emergent Behavior Framework

Traditional AI safety evaluations on isolated LLMs are insufficient as multi-agent AI ensembles become prevalent, introducing novel emergent risks. This paper introduces the Multi-Agent Emergent Behavior Evaluation (MAEBE) framework to systematically assess such risks. Using MAEBE with the Greatest Good Benchmark (and a novel double-inversion question technique), we demonstrate that: (1) LLM moral preferences, particularly for Instrumental Harm, are surprisingly brittle and shift significantly with question framing, both in single agents and ensembles. (2) The moral reasoning of LLM ensembles is not directly predictable from isolated agent behavior due to emergent group dynamics. (3) Specifically, ensembles exhibit phenomena like peer pressure influencing convergence, even when guided by a supervisor, highlighting distinct safety and alignment challenges. Our findings underscore the necessity of evaluating AI systems in their interactive, multi-agent contexts.

cs.MA

Hypermodular Self-Assembling Space Solar Power -- Design Option for Mid-Term GEO Utility-Scale Power Plants

This paper presents a design for scaleable space solar power systems based on free-flying reflectors and module self-assembly. Lower system cost of utility-scale space solar power is achieved by design independence of yet-to-be-built in-space assembly or transportation infrastructure. Using current and expected near-term technology, this study describe a design for mid-term utility-scale power plants in geosynchronous orbits. High-level economic considerations in the context of current and expected future launch costs are given as well. \c{opyright} 2013 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.

physics.space-ph

Hypermodular Distributed Solar Power Satellites -- Exploring a Technology Option for Near-Term LEO Demonstration and GLPO Full-Scale Plants

This paper presents a new and innovative design for scaleable space solar power systems based on satellite self-assembly and microwave spatial power combination. Lower system cost of utility-scale space solar power is achieved by independence of yet-to-be-built in-space assembly and transportation infrastructure. Using current and expected near-term technology, this study explores a design for near-term space solar power low-Earth orbit demonstrators and for mid-term utility-scale power plants in geosynchronous Laplace plane orbits. High-level economic considerations in the context of current and expected future launch costs are given as well.

physics.space-ph

Fragmentation Functions at Belle

Fragmentation functions (FFs) describe the formation of final state particles from a partonic initial state. Precise knowledge of these functions is a key ingredient in accessing quantities such as the nucleon spin structure in semi-inclusive deep-inelastic scattering and proton proton collisions. However, fragmentation functions can currently not be determined from first principles Quantum Chromodynamics and have to be extracted from experimental data. The Belle experiment at KEK, Japan, provides a large data sample for high precision measurements on e^{+}e^{-} annihilations allowing for first-time or more precise extractions of fragmentation functions. Analyses for extractions of spin-independent (unpolarized FFs) as well as spin-dependent fragmentation functions (interference FFs) at Belle are presented.

nucl-ex

First Measurement of the Interference Fragmentation Function in $e^+e^-$ at Belle

A first measurement of the di-hadron interference fragmentation function of light quarks in pion pairs with the Belle detector is presented. The chiral odd nature of this fragmentation function allows the use as a quark polarimeter sensitive to the transverse polarization of the fragmenting quark. Therefore it can be used together with data taken at fixed target and collider experiments to extract the quark transversity distribution. A sample consisting of $711 \times 10^6$ di-hadron pairs was extracted from 661 $fb^{-1}$ of data recorded near the $Υ(4S)$ resonance delivered by the KEKB $e^+$$e^{-}$ collider.

hep-ex