TY - RPRT TI - Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling AU - Alekh Agarwal AU - Tong Zhang PY - 2024 UR - https://arxiv.org/abs/2203.08248 ID - 2203.08248 ER -