TY - RPRT TI - Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models AU - Yanchen Yin AU - Dongqi Han AU - Linghui Li PY - 2026 UR - https://arxiv.org/abs/2606.28153 ID - 2606.28153 ER -