TY - RPRT TI - Learning Policies from Self-Play with Policy Gradients and MCTS Value Estimates AU - Dennis J. N. J. Soemers AU - Éric Piette AU - Matthew Stephenson AU - Cameron Browne PY - 2019 UR - https://arxiv.org/abs/1905.05809 ID - 1905.05809 ER -