arXiv · 2307.07843
Transformers are Universal Predictors
Abstract
We find limits to the Transformer architecture for language modeling and show it has a universal prediction property in an information-theoretic sense. We further analyze performance in non-asymptotic data regimes to understand the role of various components of the Transformer architecture, especially in the context of data-efficient training. We validate our theoretical analysis with experiments on both synthetic and real datasets.
Explore related subjects
Keep this discovery
Sourya Basu, Moulik Choraria, Lav R. Varshney. 2023-07-15. Transformers are Universal Predictors. https://arxiv.org/abs/2307.07843
Cite the original work for its findings. Save a collection to share your selection of sources.