arXiv · 2605.27259
Kan Extension Transformers: A Categorical Unification of Attention, Diffusion, and Predict-Detach Self-Conditioning
Abstract
We propose Kan Extension Transformers (KETs) as a categorical design language for a diverse group of Transformer implementations. A layer can be viewed generally as a weighted structured extension operator: attention uses token neighborhoods, geometric mixing uses sparse incidences, and KET uses simplicial sources. This operator is an actual enriched left Kan extension only when the source values are functorial, the weights are representable hom-objects (or the specified profunctor action), and aggregation realizes the corresponding coend; otherwise ``Kan-style'' denotes an interpretation rather than an identity theorem. Predict-detach blocks gradients through a predictive carrier and avoids transporting teacher-forced hidden states, but detach alone does not make a noncausal update strictly autoregressive: every carrier consumed at target $t$ must also be measurable from the prefix available at $t$. We evaluate 12 implementations on Penn Treebank, WikiText-2, and WikiText-103 across strict-causal and self-conditioned regimes, using widths $d=64,256$ and depths $L=2,8,16$ across the reported studies. Quadratic KET is strongest among the compared strict-causal architectures on WikiText-2 and WikiText-103; the largest cross-regime gains arise from additional self-conditioning information, not neighborhood design alone.
Explore related subjects
Keep this discovery
Sridhar Mahadevan. 2026-05-26. Kan Extension Transformers: A Categorical Unification of Attention, Diffusion, and Predict-Detach Self-Conditioning. https://arxiv.org/abs/2605.27259
Cite the original work for its findings. Save a collection to share your selection of sources.