arXiv · 2407.02423
On the Anatomy of Attention
Abstract
We introduce a category-theoretic diagrammatic formalism in order to systematically relate and reason about machine learning models. Our diagrams present architectures intuitively but without loss of essential detail, where natural relationships between models are captured by graphical transformations, and important differences and similarities can be identified at a glance. In this paper, we focus on attention mechanisms: translating folklore into mathematical derivations, and constructing a taxonomy of attention variants in the literature. As a first example of an empirical investigation underpinned by our formalism, we identify recurring anatomical components of attention, which we exhaustively recombine to explore a space of variations on the attention mechanism.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nikhil Khatri, Tuomas Laakkonen, Jonathon Liu, Vincent Wang-Maścianica. 2024-07-02. On the Anatomy of Attention. https://arxiv.org/abs/2407.02423
Cite the original work for its findings. Save a collection to share your selection of sources.