arXiv · 2605.30381
Linear Separability of Activation Representations after Supervised Fine-Tuning on Incorrect Responses: A Study of Synthetic Dishonesty in Large Language Models
Abstract
When a language model is fine-tuned to produce systematically incorrect responses, does this training leave a structured, linearly recoverable trace in its internal activations? We study this question in a controlled model-organism setting using five transformer architectures spanning 1.4 to 9 billion parameters. For each model, an "honest" and a "dishonest" LoRA fine-tuned variant are constructed using identical question distributions but correct versus plausible-but-incorrect answers. Linear probes reach near-ceiling separability (AUC >= 0.9997) within the first few layers in four of five architectures. This separability transfers from TruthfulQA to held-out MMLU subjects for four models, but substantially less so for Pythia-1.4B. Six geometric analyses reveal an architectural dichotomy in representation structure. A control experiment further shows that LoRA fine-tuning itself leaves a strong activation fingerprint that is largely dissociable from the honesty-related signal, independent of fine-tuning data content. Two supplementary experiments qualify these findings: the identified direction transfers poorly to paraphrased questions, with AUC approaching chance, and norm-calibrated activation steering provides no reliable evidence of a causal role in generation, partly because stronger interventions disrupt output fluency. These results characterize what supervised fine-tuning on incorrect responses induces in activation space while clarifying important limitations on interpretation and causal significance.
Explore related subjects
Keep this discovery
Vahideh Zolfaghari. 2026-05-28. Linear Separability of Activation Representations after Supervised Fine-Tuning on Incorrect Responses: A Study of Synthetic Dishonesty in Large Language Models. https://arxiv.org/abs/2605.30381
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.