arXiv · 2502.11596
LLM Embeddings for Deep Learning on Tabular Data
Abstract
Tabular deep-learning methods require embedding numerical and categorical input features into high-dimensional spaces before processing them. Existing methods deal with this heterogeneous nature of tabular data by employing separate type-specific encoding approaches. This limits the cross-table transfer potential and the exploitation of pre-trained knowledge. We propose a novel approach that first transforms tabular data into text, and then leverages pre-trained representations from LLMs to encode this data, resulting in a plug-and-play solution to improv ing deep-learning tabular methods. We demonstrate that our approach improves accuracy over competitive models, such as MLP, ResNet and FT-Transformer, by validating on seven classification datasets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Boshko Koloski, Andrei Margeloiu, Xiangjian Jiang, Blaž Škrlj, Nikola Simidjievski, Mateja Jamnik. 2025-02-17. LLM Embeddings for Deep Learning on Tabular Data. https://arxiv.org/abs/2502.11596
Cite the original work for its findings. Save a collection to share your selection of sources.