arXiv · 1109.1528
Dynamics of Boltzmann Q-Learning in Two-Player Two-Action Games
Abstract
We consider the dynamics of Q-learning in two-player two-action games with a Boltzmann exploration mechanism. For any non-zero exploration rate the dynamics is dissipative, which guarantees that agent strategies converge to rest points that are generally different from the game's Nash Equlibria (NE). We provide a comprehensive characterization of the rest point structure for different games, and examine the sensitivity of this structure with respect to the noise due to exploration. Our results indicate that for a class of games with multiple NE the asymptotic behavior of learning dynamics can undergo drastic changes at critical exploration rates. Furthermore, we demonstrate that for certain games with a single NE, it is possible to have additional rest points (not corresponding to any NE) that persist for a finite range of the exploration rates and disappear when the exploration rates of both players tend to zero.
Explore related subjects
Keep this discovery
Ardeshir Kianercy, Aram Galstyan. 2012-03-01. Dynamics of Boltzmann Q-Learning in Two-Player Two-Action Games. https://doi.org/10.1103/physreve.85.041145
Cite the original work for its findings. Save a collection to share your selection of sources.