TY - RPRT TI - Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework AU - Phalguni Nanda AU - Zaiwei Chen PY - 2026 UR - https://arxiv.org/abs/2605.10671 ID - 2605.10671 ER -