TY - RPRT TI - Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation: The Case of Multi-Armed Bandits AU - Max Qiushi Lin AU - Jincheng Mei AU - Matin Aghaei AU - Michael Lu AU - Bo Dai AU - Alekh Agarwal AU - Dale Schuurmans AU - Csaba Szepesvari AU - Sharan Vaswani PY - 2026 UR - https://arxiv.org/abs/2505.03155 ID - 2505.03155 ER -