arXiv · 2610.01435
Distillation of Tabular Foundation Models into Efficient Predictors
Abstract
Tabular foundation models (TFMs) achieve strong predictive performance through in-context learning, yet repeatedly conditioning on labeled data makes inference expensive. Knowledge distillation can reduce this cost by transferring their predictive ability to lightweight, dataset-specific students. However, the dependence of TFM predictions on both a labeled context and a query introduces two design questions: how to construct teacher supervision and whether expanding query coverage improves distillation. We examine these questions across two TFMs and both neural and tree-based students, and derive an effective distillation recipe. The recipe uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries. On TabArena, the resulting students outperform their supervised trained tuned-and-ensembled counterparts by 57-98 Elo points. Applied unchanged to TALENT, the same recipe improves matched default students on 236-258 of 300 datasets and reduces median primary error by 4.0-6.4%. The distilled students also achieve median inference speedups of 3.0-21.6 times over their teachers, offering a practical trade-off between predictive performance and repeated inference cost. Code is available at https://github.com/nums-ai/TFM_Distillation .
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo. 2026-10-01. Distillation of Tabular Foundation Models into Efficient Predictors. https://arxiv.org/abs/2610.01435
Cite the original work for its findings. Save a collection to share your selection of sources.