TY - RPRT TI - Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning AU - Harin Lee AU - Min-hwan Oh PY - 2026 UR - https://arxiv.org/abs/2605.05102 ID - 2605.05102 ER -