TY - RPRT TI - Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models AU - Thomas Winninger AU - Boussad Addad AU - Katarzyna Kapusta PY - 2026 UR - https://arxiv.org/abs/2503.06269 ID - 2503.06269 ER -