TY - RPRT TI - Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions AU - Peter Sunehag AU - Richard Evans AU - Gabriel Dulac-Arnold AU - Yori Zwols AU - Daniel Visentin AU - Ben Coppin PY - 2015 UR - https://arxiv.org/abs/1512.01124 ID - 1512.01124 ER -