TY - RPRT TI - No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping AU - Thanh-Long V. Le AU - Myeongho Jeon AU - Kim Vu AU - Viet Lai AU - Eunho Yang PY - 2026 UR - https://arxiv.org/abs/2509.21880 ID - 2509.21880 ER -