arXiv · 2608.25055
A Mean-Field Theory of Transformers: Well-Posedness of the Coupled Data--Parameter Dynamics and Global Convergence of Training
Abstract
We develop a rigorous mean-field theory for transformer networks that captures two large-scale limits inherent in the architecture: the number of tokens $N\to\infty$ in the input sequence and the number of attention heads $H\to\infty$ in each layer. It further considers an infinite number of layers leading to a time-continuous formulation. The resulting framework couples two interacting mean-field objects: a token distribution $\mu_t\in\mathcal P(\mathbb R^d)$, which evolves through network depth $t\in[0,T]$ according to a McKean--Vlasov transport equation, and the attention-parameter distribution $\rho_s\in\mathcal P(\Theta)$, that evolves through training time $s\ge0$ according to a Wasserstein gradient flow of the empirical risk, with optional entropic or Tikhonov regularization. We establish a comprehensive analytical foundation for the system coupling the transformer and the training dynamics; in particular, we establish global well-posedness of the resulting nonlinear Fokker--Planck system describing the coupled mean-field and training dynamics. Beyond well-posedness, we connect the mean-field formulation to optimization. For shallow, single-layer attention models, we prove exponential convergence to the entropy-regularized global optimum under a log-Sobolev condition. For genuinely deep, compositional transformers, we establish local linear convergence under a Neural Tangent Kernel non-degeneracy condition.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Michael Herty, Hailiang Liu. 2026-08-25. A Mean-Field Theory of Transformers: Well-Posedness of the Coupled Data--Parameter Dynamics and Global Convergence of Training. https://arxiv.org/abs/2608.25055
Cite the original work for its findings. Save a collection to share your selection of sources.