arXiv · 2510.04088
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
Abstract
This article introduces the theory of offline reinforcement learning in large state spaces, where good policies are learned from historical data without online interactions with the environment. Key concepts introduced include expressivity assumptions on function approximation (e.g., Bellman completeness vs. realizability) and data coverage (e.g., all-policy vs. single-policy coverage). A rich landscape of algorithms and results is described, depending on the assumptions one is willing to make and the sample and computational complexity guarantees one wishes to achieve. We also discuss open questions and connections to adjacent areas.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nan Jiang, Tengyang Xie. 2025-10-05. Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees. https://arxiv.org/abs/2510.04088
Cite the original work for its findings. Save a collection to share your selection of sources.