TY - RPRT TI - Regret and Sample Complexity of Online Q-Learning via Concentration of Stochastic Approximation with Time-Inhomogeneous Markov Chains AU - Rahul Singh AU - Siddharth Chandak AU - Eric Moulines AU - Vivek S. Borkar AU - Nicholas Bambos PY - 2026 UR - https://arxiv.org/abs/2602.16274 ID - 2602.16274 ER -