Searcharxiv⌕ Search

arXiv subjects

Ondrej Kubicek

Publications and source records attributed to Ondrej Kubicek.

3 recordsLinked to original sources

Test-time Reinforcement Learning in Imperfect Information Games

Test-time reasoning has significantly improved performance in domains ranging from games to language models. However, test-time policy changes with formal guarantees on the performance of the resulting strategy remain a challenge in two-player zero-sum imperfect-information games. Existing solutions are limited to tabular methods or single gradient step updates. In this work, we investigate policy-gradient algorithms as a method for scalable test-time reasoning. We extend the concept of gadget game, tabular technique for test-time search, to the reinforcement learning setting. Unlike prior approaches, we represent the gadget game implicitly by modified sampling and neural policy rather then explicitly by constructing it, thereby removing constraints on subgame size. Furthermore, we formally prove that, unlike prior tabular algorithms, regularized policy-gradient algorithms limit possible strategy degradation caused by test-time reasoning, even without the gadget games. Our evaluation across small- and large-scale games confirms that additional test-time training often substantially improves performance relative to the blueprint strategy.

cs.GT↗

Equilibrium Refinements Improve Subgame Solving in Imperfect-Information Games

Subgame solving is a technique for scaling algorithms to large games by locally refining a precomputed blueprint strategy during gameplay. While straightforward in perfect-information games where search starts from the current state, subgame solving in imperfect-information games must account for hidden states and uncertainty about the opponent's past strategy. Gadget games were developed to ensure that the improved subgame strategy is robust against any possible opponent's strategy in a zero-sum game. Gadget games typically contain infinitely many Nash equilibria. We demonstrate that while these equilibria are equivalent in the gadget game, they yield vastly different performance in the full game, even when facing a rational opponent. We propose gadget game sequential equilibria as the preferred solution concept. We introduce modifications to the sequence-form linear program and counterfactual regret minimization that converge to these refined solutions with only mild additional computational cost. Additionally, we provide several new insights into the surprising superiority of the resolving gadget game over the max-margin gadget game. Our experiments compare different Nash equilibria of gadget games in several standard benchmark games, showing that our refined equilibria consistently outperform unrefined Nash equilibria, and can reduce the exploitability of the overall strategy by more than 50%

cs.GT↗

Look-ahead Search on Top of Policy Networks in Imperfect Information Games

Search in test time is often used to improve the performance of reinforcement learning algorithms. Performing theoretically sound search in fully adversarial two-player games with imperfect information is notoriously difficult and requires a complicated training process. We present a method for adding test-time search to an arbitrary policy-gradient algorithm that learns from sampled trajectories. Besides the policy network, the algorithm trains an additional critic network, which estimates the expected values of players following various transformations of the policies given by the policy network. These values are then used for depth-limited search. We show how the values from this critic can create a value function for imperfect information games. Moreover, they can be used to compute the summary statistics necessary to start the search from an arbitrary decision point in the game. The presented algorithm is scalable to very large games since it does not require any search during train time. We evaluate the algorithm's performance when trained along Regularized Nash Dynamics, and we evaluate the benefit of using the search in the standard benchmark game of Leduc hold'em, multiple variants of imperfect information Goofspiel, and Battleships.

cs.GT↗