TY - RPRT TI - Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance AU - Kai Yan AU - Alexander G. Schwing AU - Yu-Xiong Wang PY - 2026 UR - https://arxiv.org/abs/2605.15012 ID - 2605.15012 ER -