TY - RPRT TI - Multi-Bellman operator for convergence of $Q$-learning with linear function approximation AU - Diogo S. Carvalho AU - Pedro A. Santos AU - Francisco S. Melo PY - 2023 UR - https://arxiv.org/abs/2309.16819 ID - 2309.16819 ER -