arXiv · 2607.23030
Online Policy Evaluation for MDPs with Dynamic UBSR Measures
Abstract
Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning. Existing approaches either focus on restrictive classes of risk measures or rely on access to a simulator, limiting their applicability in fully online settings. In this work, we propose computationally efficient online learning algorithms for policy evaluation in Markov decision processes (MDPs) with dynamic utility-based shortfall risk (UBSR) measures under linear function approximation. Specifically, we introduce the UBSR-TD algorithm, establish conditions under which it converges almost surely, and develop several variants designed to accelerate convergence. Our formulation shows that existing policy evaluation algorithms for risk-neutral MDPs can be readily adapted to dynamic UBSR settings by incorporating a loss function into the temporal-difference error. Numerical experiments support our theoretical findings, and an application to a perishable inventory management problem with shelf-life uncertainty demonstrates the practical effectiveness of the proposed methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Weikai Wang, Erick Delage. 2026-07-25. Online Policy Evaluation for MDPs with Dynamic UBSR Measures. https://arxiv.org/abs/2607.23030
Cite the original work for its findings. Save a collection to share your selection of sources.