TY - RPRT TI - Optimistic Posterior Sampling for Reinforcement Learning with Few Samples and Tight Guarantees AU - Daniil Tiapkin AU - Denis Belomestny AU - Daniele Calandriello AU - Eric Moulines AU - Remi Munos AU - Alexey Naumov AU - Mark Rowland AU - Michal Valko AU - Pierre Menard PY - 2022 UR - https://arxiv.org/abs/2209.14414 ID - 2209.14414 ER -