arXiv · 1412.3276
Generalised Entropy MDPs and Minimax Regret
Abstract
Bayesian methods suffer from the problem of how to specify prior beliefs. One interesting idea is to consider worst-case priors. This requires solving a stochastic zero-sum game. In this paper, we extend well-known results from bandit theory in order to discover minimax-Bayes policies and discuss when they are practical.
Explore related subjects
Keep this discovery
Emmanouil G. Androulakis, Christos Dimitrakakis. 2014-12-10. Generalised Entropy MDPs and Minimax Regret. https://arxiv.org/abs/1412.3276
Cite the original work for its findings. Save a collection to share your selection of sources.