arXiv · 2004.04120
Solving the scalarization issues of Advantage-based Reinforcement Learning Algorithms
Abstract
In this research, some of the issues that arise from the scalarization of the multi-objective optimization problem in the Advantage Actor Critic (A2C) reinforcement learning algorithm are investigated. The paper shows how a naive scalarization can lead to gradients overlapping. Furthermore, the possibility that the entropy regularization term can be a source of uncontrolled noise is discussed. With respect to the above issues, a technique to avoid gradient overlapping is proposed, while keeping the same loss formulation. Moreover, a method to avoid the uncontrolled noise, by sampling the actions from distributions with a desired minimum entropy, is investigated. Pilot experiments have been carried out to show how the proposed method speeds up the training. The proposed approach can be applied to any Advantage-based Reinforcement Learning algorithm.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Federico A. Galatolo, Mario G. C. A. Cimino, Gigliola Vaglini. 2021-10-01. Solving the scalarization issues of Advantage-based Reinforcement Learning Algorithms. https://doi.org/10.1016/j.compeleceng.2021.107117
Cite the original work for its findings. Save a collection to share your selection of sources.