TY - RPRT TI - Batch Policy Gradient Methods for Improving Neural Conversation Models AU - Kirthevasan Kandasamy AU - Yoram Bachrach AU - Ryota Tomioka AU - Daniel Tarlow AU - David Carter PY - 2017 UR - https://arxiv.org/abs/1702.03334 ID - 1702.03334 ER -