TY - RPRT TI - Reward Shaping with Recurrent Neural Networks for Speeding up On-Line Policy Learning in Spoken Dialogue Systems AU - Pei-Hao Su AU - David Vandyke AU - Milica Gasic AU - Nikola Mrksic AU - Tsung-Hsien Wen AU - Steve Young PY - 2015 UR - https://arxiv.org/abs/1508.03391 ID - 1508.03391 ER -