arXiv · 2509.23314
Two-Scale Latent Dynamics for Recurrent-Depth Transformers
Abstract
Recurrent-depth transformers scale test-time compute by iterating latent computations before emitting tokens. We study the geometry of these iterates and argue for a simple, two-scale operational picture: (i) within a looped block, updates act as small-scale refinements; (ii) across consecutive blocks, states undergo a larger-scale drift. Across training, our measurements show that loop steps become smaller and increasingly orthogonal to one another, indicating better local modeling of fine structure rather than merely pushing in a single direction. These dynamics motivate an early-exit mechanism based on the model's second-order difference in step-size, which we show is superior in terms of performance, stability and time-efficiency, when compared to the KL-divergence exit strategy of Geiping et al. and its naive first-order counterpart.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Francesco Pappone, Donato Crisostomi, Emanuele Rodolà. 2025-09-27. Two-Scale Latent Dynamics for Recurrent-Depth Transformers. https://arxiv.org/abs/2509.23314
Cite the original work for its findings. Save a collection to share your selection of sources.