TY - RPRT TI - Offline-Online Reinforcement Learning for Linear Mixture MDPs AU - Zhongjun Zhang AU - Sean R. Sinclair PY - 2026 UR - https://arxiv.org/abs/2604.11994 ID - 2604.11994 ER -