TY - RPRT TI - Reinforcement Learning of Speech Recognition System Based on Policy Gradient and Hypothesis Selection AU - Taku Kato AU - Takahiro Shinozaki PY - 2017 UR - https://arxiv.org/abs/1711.03689 ID - 1711.03689 ER -