TY - RPRT TI - Temporal-difference learning with nonlinear function approximation: lazy training and mean field regimes AU - Andrea Agazzi AU - Jianfeng Lu PY - 2021 UR - https://arxiv.org/abs/1905.10917 ID - 1905.10917 ER -