arXiv · 2511.05933
Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs
Abstract
Reinforcement learning (RL) is often credited with improving reasoning at the expense of factual knowledge. We instead find that reasoning models outperform their instruction-tuned versions on factual recall by accessing existing parametric knowledge more effectively. Across five model families, structured prompting, which explicitly guides models through hierarchical traversal, recovers most of this gap, suggesting that much of the missing knowledge is latent rather than absent. Controlled RL experiments further support this: training on unseen, non-extractable facts improves recall of held-out, frequent but previously inaccessible facts, ruling out simple data exposure. Decomposing the training objective further attributes this gain to iterated on-policy exploration. The same mechanism appears behaviorally and internally: the reasoning advantage grows with retrieval depth, while layerwise analysis finds similar factual representations but divergent query representations. Distilled models, in contrast, often imitate self-correction without acquiring the exploration needed for navigation. Together, these findings suggest that improving factual recall in LLMs depends not only on expanding what models know but also on teaching them to navigate it, motivating future post-training methods that optimize traversal.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Renfei Zhang, Manasa Kaniselvan, Rylan Schaeffer, Niloofar Mireshghallah. 2025-11-08. Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs. https://arxiv.org/abs/2511.05933
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.