arXiv · 2402.09268
Transformers, parallel computation, and logarithmic depth
Abstract
We show that a constant number of self-attention layers can efficiently simulate, and be simulated by, a constant number of communication rounds of Massively Parallel Computation. As a consequence, we show that logarithmic depth is sufficient for transformers to solve basic computational tasks that cannot be efficiently solved by several other neural sequence models and sub-quadratic transformer approximations. We thus establish parallelism as a key distinguishing property of transformers.
Explore related subjects
Keep this discovery
Clayton Sanford, Daniel Hsu, Matus Telgarsky. 2024-02-14. Transformers, parallel computation, and logarithmic depth. https://arxiv.org/abs/2402.09268
Cite the original work for its findings. Save a collection to share your selection of sources.