TY - RPRT TI - $Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training AU - Jin Peng Zhou AU - Kaiwen Wang AU - Jonathan Chang AU - Zhaolin Gao AU - Nathan Kallus AU - Kilian Q. Weinberger AU - Kianté Brantley AU - Wen Sun PY - 2025 UR - https://arxiv.org/abs/2502.20548 ID - 2502.20548 ER -