TY - RPRT TI - From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning AU - Jike Zhong AU - Yuxiang Lai AU - Ming Li AU - Yuheng Li AU - Wuao Liu AU - Behzad Dariush AU - Konstantinos Psounis AU - Shao-Yuan Lo PY - 2026 UR - https://arxiv.org/abs/2606.09092 ID - 2606.09092 ER -