TY - RPRT TI - Fast Convergence of Policy Regret in Learning Stochastic Optimal Control AU - Shengbo Wang AU - Jose Blanchet AU - Peter Glynn PY - 2026 UR - https://arxiv.org/abs/2605.26361 ID - 2605.26361 ER -