arXiv · 2606.10835
Geometrically Averaged Hard Target Updates for Linear Q-Learning
Abstract
Periodic hard target updates are among the most common stabilization devices in modern deep Q-learning. Recent studies suggest that target updates can improve stability in Q-learning with function approximation, including linear function approximation. We introduce and analyze the so-called $\lambda$-target update, obtained by averaging the $m$-periodic target update maps with $\lambda$-geometric weights $(1-\lambda)\lambda^{m-1}$, $\lambda \in [0,1]$. The endpoint $\lambda=0$ recovers the one-period target update, while the continuous endpoint $\lambda\uparrow1$ recovers projected Q-value iteration. We study this mechanism for Q-learning with linear function approximation, namely linear Q-learning, using a switching-system model and related tools. For clarity, the paper treats a deterministic version; the formulation extends to stochastic reinforcement-learning settings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Donghwan Lee. 2026-06-09. Geometrically Averaged Hard Target Updates for Linear Q-Learning. https://arxiv.org/abs/2606.10835
Cite the original work for its findings. Save a collection to share your selection of sources.