arXiv · 2609.34009
Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs
Abstract
Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher and student. However, such methods can exhibit unstable optimization, conducive to performance collapse during training. To address this limitation, we propose FIRE (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning. For correct responses, FIRE replaces self-distillation with re-weighted on-policy SFT, while for incorrect ones FIRE identifies feedback components that disproportionately influence the teacher-induced update and recalibrates the feedback-conditioned target accordingly. Both branches are influenced by a token-level radius derived in part from a softmax Fisher trace. FIRE separates which direction feedback should move the model from how far the model should move in that direction, while leaving well-behaved feedback supervision unchanged. Our experiments demonstrate that FIRE provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where standard feedback-conditioned distillation becomes unstable.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Seohyun Lee, Dong-Jun Han, Seyyedali Hosseinalipour, Christopher G. Brinton. 2026-09-27. Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs. https://arxiv.org/abs/2609.34009
Cite the original work for its findings. Save a collection to share your selection of sources.