TY - RPRT TI - Convergence of Policy Gradient for Entropy Regularized MDPs with Neural Network Approximation in the Mean-Field Regime AU - Bekzhan Kerimkulov AU - James-Michael Leahy AU - David Šiška AU - Lukasz Szpruch PY - 2022 UR - https://arxiv.org/abs/2201.07296 ID - 2201.07296 ER -