arXiv · 2303.07951
MetaMixer: A Regularization Strategy for Online Knowledge Distillation
Abstract
Online knowledge distillation (KD) has received increasing attention in recent years. However, while most existing online KD methods focus on developing complicated model structures and training strategies to improve the distillation of high-level knowledge like probability distribution, the effects of the multi-level knowledge in the online KD are greatly overlooked, especially the low-level knowledge. Thus, to provide a novel viewpoint to online KD, we propose MetaMixer, a regularization strategy that can strengthen the distillation by combining the low-level knowledge that impacts the localization capability of the networks, and high-level knowledge that focuses on the whole image. Experiments under different conditions show that MetaMixer can achieve significant performance gains over state-of-the-art methods.
Explore related subjects
Keep this discovery
Maorong Wang, Ling Xiao, Toshihiko Yamasaki. 2023-03-14. MetaMixer: A Regularization Strategy for Online Knowledge Distillation. https://arxiv.org/abs/2303.07951
Cite the original work for its findings. Save a collection to share your selection of sources.