TY - RPRT TI - Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling AU - Abhijeet Sinha AU - Sundari Elango AU - Dianbo Liu PY - 2026 UR - https://arxiv.org/abs/2601.21669 ID - 2601.21669 ER -