Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies
We study entropy-regularized exploratory control in finite $N$-player stochastic differential games under a response model in which each player conditions on the opponents' currently realized actions and evaluates continuation with that profile frozen. The resulting Gibbs best responses form a system of full conditional densities, which need not admit a common joint law. We characterize joint realizability by a cross-partial condition on the entropy-scaled Hamiltonian gradients and, on simply connected action domains, by an equivalent entropy-weighted potential structure. When compatibility fails, a coordinate-path construction yields a joint density whose full conditionals satisfy explicit quadratic Kullback--Leibler bounds. We extend the analysis to stationary discounted problems and derive martingale and policy-improvement characterizations for learning the frozen response maps. A two-player linear-quadratic example illustrates the compatibility criterion and the associated learning procedure.