TY - RPRT TI - DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback AU - Guojun Xiong AU - Ujwal Dinesha AU - Debajoy Mukherjee AU - Jian Li AU - Srinivas Shakkottai PY - 2025 UR - https://arxiv.org/abs/2410.05527 ID - 2410.05527 ER -