arXiv · 2606.02363
Minimax-Optimal Policy Regret in Partially Observable Markov Games
Abstract
We study sequential decision-making in partially observable environments against strategic, adaptive opponents, modeled as partially observable Markov games (POMGs). The central challenge is to learn latent dynamics from partial observations while facing an adversary whose behavior depends on the learner's strategy, making standard regret notions inadequate. We prove that an epoch-based optimistic maximum-likelihood algorithm achieves ${\tilde{O}}(\sqrt{T})$ policy regret for fixed problem parameters, with explicit dependence on the horizon, adversary memory, confidence radius, and the aggregate Eluder dimension of the observable-operator error classes. The algorithm deploys one policy per epoch, with geometrically capped epoch lengths, confidence sets built cumulatively from past data, and a statistical termination test that ends an epoch as soon as the data refute the deployed optimistic model. We also prove a matching lower bound, and extend the framework to horizon-adaptive guarantees and geometrically fading adversary memory.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Raman Arora. 2026-06-01. Minimax-Optimal Policy Regret in Partially Observable Markov Games. https://arxiv.org/abs/2606.02363
Cite the original work for its findings. Save a collection to share your selection of sources.