arXiv · 2004.06784
Extrapolation in Gridworld Markov-Decision Processes
Abstract
Extrapolation in reinforcement learning is the ability to generalize at test time given states that could never have occurred at training time. Here we consider four factors that lead to improved extrapolation in a simple Gridworld environment: (a) avoiding maximum Q-value (or other deterministic methods) for action choice at test time, (b) ego-centric representation of the Gridworld, (c) building rotational and mirror symmetry into the learning mechanism using rotational and mirror invariant convolution (rather than standard translation-invariant convolution), and (d) adding a maximum entropy term to the loss function to encourage equally good actions to be chosen equally often.
Explore related subjects
Keep this discovery
Eugene Charniak. 2020-04-14. Extrapolation in Gridworld Markov-Decision Processes. https://arxiv.org/abs/2004.06784
Cite the original work for its findings. Save a collection to share your selection of sources.