arXiv · 2501.17115
Evidence on the Regularisation Properties of Maximum-Entropy Reinforcement Learning
Abstract
The generalisation and robustness properties of policies learnt through Maximum-Entropy Reinforcement Learning are investigated on chaotic dynamical systems with Gaussian noise on the observable. First, the robustness under noise contamination of the agent's observation of entropy regularised policies is observed. Second, notions of statistical learning theory, such as complexity measures on the learnt model, are borrowed to explain and predict the phenomenon. Results show the existence of a relationship between entropy-regularised policy optimisation and robustness to noise, which can be described by the chosen complexity measures.
Explore related subjects
Keep this discovery
Rémy Hosseinkhan-Boucher, Onofrio Semeraro, Lionel Mathelin. 2025-01-28. Evidence on the Regularisation Properties of Maximum-Entropy Reinforcement Learning. https://arxiv.org/abs/2501.17115
Cite the original work for its findings. Save a collection to share your selection of sources.