TY - RPRT TI - Reinforcement Learning for Diffusion LLMs with Entropy-Guided Step Selection and Stepwise Advantages AU - Vishnu Teja Kunde AU - Fatemeh Doudi AU - Mahdi Farahbakhsh AU - Dileep Kalathil AU - Krishna Narayanan AU - Jean-Francois Chamberland PY - 2026 UR - https://arxiv.org/abs/2603.12554 ID - 2603.12554 ER -