arXiv · 2609.32333
Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs
Abstract
Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's view distribution from the crop toward the full image through an intermediate aspect-preserving padded crop. The padded crop preserves regional content while matching the full image's visual-token grid. Across stages, the view mixture assigns increasing probability to the full image. A lightweight regional-advantage weighting reallocates token-level supervision using the crop-conditioned teacher-student log-probability gap. Evaluated under each sampled input, it applies mild reweighting when the gap is small and emphasizes higher-gap tokens when the gap widens. A Jensen-Shannon metric decomposition interprets this schedule as a shift from matched-input imitation toward the deployment objective. Across benchmarks spanning perception, visual mathematics and general multimodal question answering, PVD-full reaches an average accuracy of 77.51 over three seeds, improving on the reward-free distillation baseline by 2.01 points and on its reward-matched variant by 1.00 point. In the reward-free setting, PVD-distill still gains 1.16 points.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du. 2026-09-26. Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs. https://arxiv.org/abs/2609.32333
Cite the original work for its findings. Save a collection to share your selection of sources.