arXiv · 2609.05856
Transformation of Adaptive Multistage Sampling for Solving Finite-Horizon Markov Decision Processes with Unknown Model
Abstract
This work provides a learning approach to solving finite-horizon Markov decision processes (MDPs) when the underlying model of a given finite MDP is unknown to the decision maker. We transform the adaptive multistage sampling (AMS) algorithm into a sampling-free algorithm, called "adaptive multistage rollout (AMR)," for estimating the optimal value at an initial state when only the state set and the action set are known. AMR emulates the backward induction as in AMS but in a reinforcement learning (RL) setting. At each iteration, AMR generates a non-stationary policy to be used for exploration and rolls out the policy in order to obtain a single trajectory of experiences and traces it backwards in a non-recursive way while doing relevant updates only at visited states and for actions taken at the visited states. We show that AMR is asymptotically optimal such that the sequence of the expected absolute errors approaches zero and its convergence rate depends on the number of visits to each reachable state at each stage from the initial state, essentially transforming the result of AMS into the RL setting.
Explore related subjects
Keep this discovery
Hyeong Soo Chang. 2026-09-05. Transformation of Adaptive Multistage Sampling for Solving Finite-Horizon Markov Decision Processes with Unknown Model. https://arxiv.org/abs/2609.05856
Cite the original work for its findings. Save a collection to share your selection of sources.