arXiv · 2510.25952
Modular Linear Tokenization (MLT)
Abstract
This paper introduces Modular Linear Tokenization (MLT), a reversible and deterministic technique for encoding high-cardinality categorical identifiers into compact numerical vectors. Unlike traditional hashing or one-hot encodings, MLT preserves bijective mappings by leveraging modular arithmetic over finite fields and invertible linear transformations. The method offers explicit control of dimensionality and computational scalability while maintaining full reversibility, even for millions of identifiers. Experimental results on the MovieLens 20M dataset show that MLT achieves comparable predictive performance to supervised embeddings while requiring significantly fewer parameters and lower training cost. An open-source implementation of MLT is available on PyPI (https://pypi.org/project/light-mlt/) and GitHub (https://github.com/tcharliesschmitz/light-mlt).
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tcharlies Schmitz. 2025-10-29. Modular Linear Tokenization (MLT). https://arxiv.org/abs/2510.25952
Cite the original work for its findings. Save a collection to share your selection of sources.