Properties of Turnpike Functions for Discounted Finite Markov Decision Processes
This paper studies convergence times of the Value Iteration Algorithm (VIA) for discounted discrete-time Markov Decision Processes (MDPs) with finite state and action sets. For each discount factor, starting from a finite number of iterations, which is called the turnpike integer, the VIA generates deterministic optimal policies for infinite-horizon problems. Turnpike integers are viewed as functions of discount factors called turnpike functions. We design an algorithm to detect in strongly polynomial time whether a discount factor is irregular - at which the sets of optimal policies change. We then prove that a turnpike function is upper-semicontinuous and finitely-piecewise constant on any closed subinterval of [0,1) that does not contain irregular points. In particular, we prove that for small discount factors a turnpike function is bounded by the number of states and design an algorithm based on the VIA to solve an MDP for all small discount factors in strongly polynomial time.