arXiv · math/0508319
Combinations and Mixtures of Optimal Policies in Unichain Markov Decision Processes are Optimal
Abstract
We show that combinations of optimal (stationary) policies in unichain Markov decision processes are optimal. That is, let M be a unichain Markov decision process with state space S, action space A and policies π_j^*: S -> A (1\leq j\leq n) with optimal average infinite horizon reward. Then any combination πof these policies, where for each state i in S there is a j such that π(i)=π_j^*(i), is optimal as well. Furthermore, we prove that any mixture of optimal policies, where at each visit in a state i an arbitrary action π_j^*(i) of an optimal policy is chosen, yields optimal average reward, too.
Explore related subjects
Keep this discovery
Ronald Ortner. 2005-08-17. Combinations and Mixtures of Optimal Policies in Unichain Markov Decision Processes are Optimal. https://arxiv.org/abs/math/0508319
Cite the original work for its findings. Save a collection to share your selection of sources.