arXiv · 2604.11613
Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification
Abstract
Transformers can perform in-context classification from a few labeled examples, yet the inference-time algorithm remains opaque. We study multi-class linear classification in the hard no-margin regime and make the computation identifiable by enforcing feature- and label-permutation equivariance at every layer. This enables interpretability while maintaining functional equivalence and yields highly structured weights. From these models we extract an explicit depth-indexed recursion: an end-to-end identified, emergent update rule inside a softmax transformer, to our knowledge the first of its kind. Attention matrices formed from mixed feature-label Gram structure drive coupled updates of training points, labels, and the test probe. The resulting dynamics implement a geometry-driven algorithmic motif, which can provably amplify class separation and yields robust expected class alignment.
Explore related subjects
Keep this discovery
Patrick Lutz, Themistoklis Haris, Arjun Chandra, Aditya Gangrade, Venkatesh Saligrama. 2026-04-13. Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification. https://arxiv.org/abs/2604.11613
Cite the original work for its findings. Save a collection to share your selection of sources.