TY - RPRT TI - What is the objective of reasoning with reinforcement learning? AU - Damek Davis AU - Benjamin Recht PY - 2025 UR - https://arxiv.org/abs/2510.13651 ID - 2510.13651 ER -