arXiv · 2410.01249
Dual Approximation Policy Optimization
Abstract
We propose Dual Approximation Policy Optimization (DAPO), a framework that incorporates general function approximation into policy mirror descent methods. In contrast to the popular approach of using the $L_2$-norm to measure function approximation errors, DAPO uses the dual Bregman divergence induced by the mirror map for policy projection. This duality framework has both theoretical and practical implications: not only does it achieve fast linear convergence with general function approximation, but it also includes several well-known practical methods as special cases, immediately providing strong convergence guarantees.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhihan Xiong, Maryam Fazel, Lin Xiao. 2024-10-02. Dual Approximation Policy Optimization. https://arxiv.org/abs/2410.01249
Cite the original work for its findings. Save a collection to share your selection of sources.