arXiv · 2505.12350
Multi-CALF: A Policy Combination Approach with Statistical Guarantees
Abstract
We introduce Multi-CALF, an algorithm that intelligently combines reinforcement learning policies based on their relative value improvements. Our approach integrates a standard RL policy with a theoretically-backed alternative policy, inheriting formal stability guarantees while often achieving better performance than either policy individually. We prove that our combined policy converges to a specified goal set with known probability and provide precise bounds on maximum deviation and convergence time. Empirical validation on control tasks demonstrates enhanced performance while maintaining stability guarantees.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Georgiy Malaniya, Anton Bolychev, Grigory Yaremenko, Anastasia Krasnaya, Pavel Osinenko. 2025-05-18. Multi-CALF: A Policy Combination Approach with Statistical Guarantees. https://arxiv.org/abs/2505.12350
Cite the original work for its findings. Save a collection to share your selection of sources.