TY - RPRT TI - Linear Separability of Activation Representations after Supervised Fine-Tuning on Incorrect Responses: A Study of Synthetic Dishonesty in Large Language Models AU - Vahideh Zolfaghari PY - 2026 UR - https://arxiv.org/abs/2605.30381 ID - 2605.30381 ER -