arXiv · 2608.02867
BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?
Abstract
Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significant debate as to whether RLVR expands the reasoning capability boundary, or just improves sampling efficiency. In this paper, we investigate the nature of test-time exploration in RLVR-trained LLMs by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence. This helps us delineate between entropy arising from stylistic variations and genuine inferential branching. Our findings demonstrate that the policy entropy collapse observed in RLVR models is not merely syntactic, and is accompanied by a significant reduction in semantic branching entropy. While RLVR improves adherence to environmental constraints and backtracking capabilities, it constricts the space of continuations; we provide evidence suggesting that this might be responsible for the sample efficiency gains of RLVR, albeit at the cost of genuine rollout diversity.
Explore related subjects
Keep this discovery
Soumadeep Saha, Krish Sharma, Akshay Chaturvedi, Nicholas Asher. 2026-08-03. BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?. https://arxiv.org/abs/2608.02867
Cite the original work for its findings. Save a collection to share your selection of sources.