TY - RPRT TI - On the Stochastic (Variance-Reduced) Proximal Gradient Method for Regularized Expected Reward Optimization AU - Ling Liang AU - Haizhao Yang PY - 2024 UR - https://arxiv.org/abs/2401.12508 ID - 2401.12508 ER -