TY - RPRT TI - DTR Bandit: Learning to Make Response-Adaptive Decisions With Low Regret AU - Yichun Hu AU - Nathan Kallus PY - 2022 UR - https://arxiv.org/abs/2005.02791 ID - 2005.02791 ER -