TY - RPRT TI - Federated Q-Learning with Reference-Advantage Decomposition: Almost Optimal Regret and Logarithmic Communication Cost AU - Zhong Zheng AU - Haochen Zhang AU - Lingzhou Xue PY - 2025 UR - https://arxiv.org/abs/2405.18795 ID - 2405.18795 ER -