TY - RPRT TI - Data-dependent Bounds with $T$-Optimal Best-of-Both-Worlds Guarantees in Multi-Armed Bandits using Stability-Penalty Matching AU - Quan Nguyen AU - Shinji Ito AU - Junpei Komiyama AU - Nishant A. Mehta PY - 2025 UR - https://arxiv.org/abs/2502.08143 ID - 2502.08143 ER -