arXiv · 2603.01875
KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
Abstract
Knowledge distillation (KD) is widely used to compress and post-train large language models (LLMs), yet many existing frameworks execute teacher inference with the same training-oriented backend as student optimization, leading to suboptimal efficiency. In this paper, we propose KDFlow, a novel framework for LLM distillation that features a decoupled architecture and employs SGLang for teacher inference. KDFlow combines SGLang for teacher inference with PyTorch FSDP2 for student optimization, allowing each model to run on a backend tailored to its workload. To enable efficient full-vocabulary distillation in this decoupled architecture, KDFlow transfers the teacher's final hidden states via Ray's object store and recomputes teacher logits on each student worker using a frozen copy of the teacher's output head. Furthermore, our framework supports both off-policy and on-policy distillation and incorporates cross-tokenizer algorithms through highly extensible and user-friendly APIs. Experiments show that KDFlow achieves a 1.44$\times$ to 6.36$\times$ speedup over MS-SWIFT in off-policy distillation and a 1.43$\times$ to 1.75$\times$ speedup over verl in on-policy distillation. KDFlow further scales to 64 GPUs, achieving 3.68$\times$ and 2.52$\times$ strong-scaling speedups in two representative model configurations. The code and documentation are publicly available.
Explore related subjects
Keep this discovery
Songming Zhang, Xue Zhang, Tong Zhang, Bojie Hu, Yufeng Chen, Jinan Xu. 2026-03-02. KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models. https://arxiv.org/abs/2603.01875
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.