arXiv · 1706.10031
Neural Sequence Model Training via $α$-divergence Minimization
Abstract
We propose a new neural sequence model training method in which the objective function is defined by $α$-divergence. We demonstrate that the objective function generalizes the maximum-likelihood (ML)-based and reinforcement learning (RL)-based objective functions as special cases (i.e., ML corresponds to $α\to 0$ and RL to $α\to1$). We also show that the gradient of the objective function can be considered a mixture of ML- and RL-based objective gradients. The experimental results of a machine translation task show that minimizing the objective function with $α> 0$ outperforms $α\to 0$, which corresponds to ML-based methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sotetsu Koyamada, Yuta Kikuchi, Atsunori Kanemura, Shin-ichi Maeda, Shin Ishii. 2017-06-30. Neural Sequence Model Training via $α$-divergence Minimization. https://arxiv.org/abs/1706.10031
Cite the original work for its findings. Save a collection to share your selection of sources.