arXiv · 2107.05479
Behavior Constraining in Weight Space for Offline Reinforcement Learning
Abstract
In offline reinforcement learning, a policy needs to be learned from a single pre-collected dataset. Typically, policies are thus regularized during training to behave similarly to the data generating policy, by adding a penalty based on a divergence between action distributions of generating and trained policy. We propose a new algorithm, which constrains the policy directly in its weight space instead, and demonstrate its effectiveness in experiments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Phillip Swazinna, Steffen Udluft, Daniel Hein, Thomas Runkler. 2021-07-12. Behavior Constraining in Weight Space for Offline Reinforcement Learning. https://arxiv.org/abs/2107.05479
Cite the original work for its findings. Save a collection to share your selection of sources.