arXiv · 2505.21835
TuneComp: Joint Fine-tuning and Compression for Large Foundation Models
Abstract
To reduce model size during post-training, compression methods, including knowledge distillation, low-rank approximation, and pruning, are often applied after fine-tuning the model. However, sequential fine-tuning and compression sacrifices performance, while creating a larger than necessary model as an intermediate step. In this work, we aim to reduce this gap, by directly constructing a smaller model while guided by the downstream task. We propose to jointly fine-tune and compress the model by gradually distilling it to a pruned low-rank structure. Experiments demonstrate that joint fine-tuning and compression significantly outperforms other sequential compression methods.
Explore related subjects
Keep this discovery
Xiangyu Chen, Jing Liu, Ye Wang, Matthew Brand, Pu, Wang, Toshiaki Koike-Akino. 2025-05-27. TuneComp: Joint Fine-tuning and Compression for Large Foundation Models. https://arxiv.org/abs/2505.21835
Cite the original work for its findings. Save a collection to share your selection of sources.