TY - RPRT TI - Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers AU - Kusha Sareen AU - Morgane M Moss AU - Alessandro Sordoni AU - Rishabh Agarwal AU - Arian Hosseini PY - 2026 UR - https://arxiv.org/abs/2505.04842 ID - 2505.04842 ER -