TY - RPRT TI - Minimax Optimal Online Imitation Learning via Replay Estimation AU - Gokul Swamy AU - Nived Rajaraman AU - Matthew Peng AU - Sanjiban Choudhury AU - J. Andrew Bagnell AU - Zhiwei Steven Wu AU - Jiantao Jiao AU - Kannan Ramchandran PY - 2023 UR - https://arxiv.org/abs/2205.15397 ID - 2205.15397 ER -