SearcharxivSearch

arXiv subjects

Gaojin He

Publications and source records attributed to Gaojin He.

3 recordsLinked to original sources

Properties of Turnpike Functions for Discounted Finite Markov Decision Processes

This paper studies convergence times of the Value Iteration Algorithm (VIA) for discounted discrete-time Markov Decision Processes (MDPs) with finite state and action sets. For each discount factor, starting from a finite number of iterations, which is called the turnpike integer, the VIA generates deterministic optimal policies for infinite-horizon problems. Turnpike integers are viewed as functions of discount factors called turnpike functions. We design an algorithm to detect in strongly polynomial time whether a discount factor is irregular - at which the sets of optimal policies change. We then prove that a turnpike function is upper-semicontinuous and finitely-piecewise constant on any closed subinterval of [0,1) that does not contain irregular points. In particular, we prove that for small discount factors a turnpike function is bounded by the number of states and design an algorithm based on the VIA to solve an MDP for all small discount factors in strongly polynomial time.

math.OC

Strong Polynomiality of the Value Iteration Algorithm for Computing Nearly Optimal Policies for Discounted Dynamic Programming

This note provides upper bounds on the number of operations required to compute by value iterations a nearly optimal policy for an infinite-horizon discounted Markov decision process with a finite number of states and actions. For a given discount factor, magnitude of the reward function, and desired closeness to optimality, these upper bounds are strongly polynomial in the number of state-action pairs, and one of the provided upper bounds has the property that it is a non-decreasing function of the value of the discount factor.

math.OC