TY - RPRT TI - LLM Safety From Within: Detecting Harmful Content with Internal Representations AU - Difan Jiao AU - Yilun Liu AU - Ye Yuan AU - Zhenwei Tang AU - Linfeng Du AU - Haolun Wu AU - Ashton Anderson PY - 2026 DO - 10.18653/v1/2026.acl-long.1844 UR - https://arxiv.org/abs/2604.18519 ID - 2604.18519 ER -