arXiv · 2608.22636
Q-Learning with Stable Infinite-Dimensional Linear Function Approximation
Abstract
Q-learning with linear function approximation can be unstable because an arbitrary approximation architecture need not preserve the Bellman contraction. We develop a stable infinite-dimensional linear function approximation framework for Q-learning from a single Markovian behavior-policy trajectory. The learning variable is a coefficient field $\theta\in C(\mathbb L)$ on a compact latent metric space $(\mathbb L,\rho)$. The framework uses a reconstruction operator that maps $\theta$ to a continuous Q-function and a compression operator that maps Bellman updates back to latent coordinates. Nonexpansiveness of both operators induces a contractive latent Bellman map on $C(\mathbb L)$, with a unique fixed point $\theta^*$ whose reconstruction approximates the optimal Q-function up to representation error. We propose two stochastic approximation (SA) algorithms and establish their sup-norm convergence bounds with a leading term of order $\widetilde O(n^{-1/2})$. The infinite-dimensional formulation provides a powerful abstraction for identifying the structures that govern statistical difficulty. Smoothness of the compression map in $\rho$ is inherited by $\theta^*$ and the SA iterates, allowing uniform estimation errors to be controlled through covering numbers of $(\mathbb L,\rho)$ rather than the dimension of $C(\mathbb L)$. Remarkably, the SA algorithms we propose are agnostic to the choice of $\rho$, and thus can automatically adapt to both the smoothness and the geometry. We further illustrate the framework through Q-measure-learning with linear density approximation and output-layer neural weight training under a frozen pretrained network.
Explore related subjects
Keep this discovery
Shengbo Wang. 2026-08-23. Q-Learning with Stable Infinite-Dimensional Linear Function Approximation. https://arxiv.org/abs/2608.22636
Cite the original work for its findings. Save a collection to share your selection of sources.