TY - RPRT TI - Whittle index based Q-learning for restless bandits with average reward AU - Konstantin E. Avrachenkov AU - Vivek S. Borkar PY - 2021 DO - 10.1016/j.automatica.2022.110186 UR - https://arxiv.org/abs/2004.14427 ID - 2004.14427 ER -