TY - RPRT TI - Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models AU - Raj Jaiswal AU - Dhruv Jain AU - Rishabh Dhawan AU - Sree Krishna Uppalapati AU - Shin'ichi Satoh AU - Tanuja Ganu AU - Rajiv Ratn Shah PY - 2026 UR - https://arxiv.org/abs/2607.05199 ID - 2607.05199 ER -