arXiv · 2502.05375
Properties of Turnpike Functions for Discounted Finite Markov Decision Processes
Abstract
This paper studies convergence times of the Value Iteration Algorithm (VIA) for discounted discrete-time Markov Decision Processes (MDPs) with finite state and action sets. For each discount factor, starting from a finite number of iterations, which is called the turnpike integer, the VIA generates deterministic optimal policies for infinite-horizon problems. Turnpike integers are viewed as functions of discount factors called turnpike functions. We design an algorithm to detect in strongly polynomial time whether a discount factor is irregular - at which the sets of optimal policies change. We then prove that a turnpike function is upper-semicontinuous and finitely-piecewise constant on any closed subinterval of [0,1) that does not contain irregular points. In particular, we prove that for small discount factors a turnpike function is bounded by the number of states and design an algorithm based on the VIA to solve an MDP for all small discount factors in strongly polynomial time.
Explore related subjects
Keep this discovery
Eugene A. Feinberg, Gaojin He. 2025-02-07. Properties of Turnpike Functions for Discounted Finite Markov Decision Processes. https://arxiv.org/abs/2502.05375
Cite the original work for its findings. Save a collection to share your selection of sources.