TY - RPRT TI - Stagewise Reinforcement Learning and the Geometry of the Regret Landscape AU - Chris Elliott AU - Einar Urdshals AU - David Quarel AU - Matthew Farrugia-Roberts AU - Daniel Murfet PY - 2026 UR - https://arxiv.org/abs/2601.07524 ID - 2601.07524 ER -