arXiv · 2609.32749
Retrospective Distillation Attribution via Normalized Response Similarity
Abstract
Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However, a distilled model may undergo further SFT, preference optimization, or reinforcement learning before release, while an auditor may lack access to the pre-distillation checkpoint required by reference-based attribution. To close this gap, we propose SCOUT, an output-only method that aggregates recurring *syntactic patterns* into candidate profiles, filters low-contrast patterns, and calibrates student--candidate distances against inter-candidate distances. SCOUT supports attribution and abstention using only current texts, without model weights, token likelihoods, or historical checkpoints. Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source. Furthermore, tracing teacher-associated *syntactic signatures* along training trajectories reveals that they emerge during distillation and persist through subsequent preference optimization and reinforcement learning.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Minwoo Jang, Jaechang Kim, Minhyeon Oh, Jeongyeon Hwang, Jungseul Ok. 2026-09-26. Retrospective Distillation Attribution via Normalized Response Similarity. https://arxiv.org/abs/2609.32749
Cite the original work for its findings. Save a collection to share your selection of sources.