arXiv · 2407.04240
A Multi-Step Minimax Q-learning Algorithm for Two-Player Zero-Sum Markov Games
Abstract
An interesting iterative procedure is proposed to solve a two-player zero-sum Markov games. Under suitable assumption, the boundedness of the proposed iterates is obtained theoretically. Using results from stochastic approximation, the almost sure convergence of the proposed two-step minimax Q-learning is obtained theoretically. More specifically, the proposed algorithm converges to the game theoretic optimal value with probability one, when the model information is not known. Numerical simulation authenticate that the proposed algorithm is effective and easy to implement.
Explore related subjects
Keep this discovery
Shreyas S R, Antony Vijesh. 2024-07-05. A Multi-Step Minimax Q-learning Algorithm for Two-Player Zero-Sum Markov Games. https://doi.org/10.1016/j.neucom.2025.131552
Cite the original work for its findings. Save a collection to share your selection of sources.