arXiv · 2206.13714
Generalized Policy Improvement Algorithms with Theoretically Supported Sample Reuse
Abstract
We develop a new class of model-free deep reinforcement learning algorithms for data-driven, learning-based control. Our Generalized Policy Improvement algorithms combine the policy improvement guarantees of on-policy methods with the efficiency of sample reuse, addressing a trade-off between two important deployment requirements for real-world control: (i) practical performance guarantees and (ii) data efficiency. We demonstrate the benefits of this new class of algorithms through extensive experimental analysis on a broad range of simulated control tasks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
James Queeney, Ioannis Ch. Paschalidis, Christos G. Cassandras. 2022-06-28. Generalized Policy Improvement Algorithms with Theoretically Supported Sample Reuse. https://doi.org/10.1109/tac.2024.3454011
Cite the original work for its findings. Save a collection to share your selection of sources.