TY - RPRT TI - Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection AU - Pu Li AU - Tao Tan AU - Hong Xie AU - Xiaoyu Shi AU - Mingsheng Shang PY - 2026 UR - https://arxiv.org/abs/2608.12912 ID - 2608.12912 ER -