TY - RPRT TI - Mechanistic Interpretability for AI Safety -- A Review AU - Leonard Bereska AU - Efstratios Gavves PY - 2024 UR - https://arxiv.org/abs/2404.14082 ID - 2404.14082 ER -