arXiv · 2210.12160
On the connection between Bregman divergence and value in regularized Markov decision processes
Abstract
In this short note we derive a relationship between the Bregman divergence from the current policy to the optimal policy and the suboptimality of the current value function in a regularized Markov decision process. This result has implications for multi-task reinforcement learning, offline reinforcement learning, and regret analysis under function approximation, among others.
Explore related subjects
Keep this discovery
Brendan O'Donoghue. 2022-10-21. On the connection between Bregman divergence and value in regularized Markov decision processes. https://arxiv.org/abs/2210.12160
Cite the original work for its findings. Save a collection to share your selection of sources.