arXiv · 2512.05790
Learnability Window in Gated Recurrent Neural Networks
Abstract
We develop a statistical theory of temporal learnability in recurrent neural networks, quantifying the maximal temporal horizon $\mathcal{H}_N$ over which gradient-based learning can recover lag-dependent structure at finite sample size $N$. The theory is built on the effective learning rate envelope $f(\ell)$, a function that captures how gating mechanisms and adaptive optimizers jointly shape the coupling between state-space dynamics and parameter updates during Backpropagation Through Time. Under heavy-tailed ($\alpha$-stable) fluctuations, where empirical averages concentrate at rate $N^{-1/\kappa_\alpha}$ with $\kappa_\alpha = \alpha/(\alpha-1)$, the interplay between envelope decay and statistical concentration yields explicit scaling laws for the growth of $\mathcal{H}_N$: logarithmic, polynomial, and exponential temporal learning regimes emerge according to the decay law of $f(\ell)$. These results identify envelope decay as the key determinant of temporal learnability. Slower attenuation of $f(\ell)$ enlarges $\mathcal{H}_N$, while heavy-tailed fluctuations compress it by weakening statistical concentration. Moreover, envelope geometry outweighs dataset size: slowing the envelope's decay enlarges $\mathcal{H}_N$ more than adding data, so more complex architectures that realize slower-decaying envelopes can be more data-efficient than simpler ones. Experiments across multiple gated architectures and optimizers corroborate these structural predictions.
Explore related subjects
Keep this discovery
Lorenzo Livi. 2025-12-05. Learnability Window in Gated Recurrent Neural Networks. https://doi.org/10.1103/843n-yshj
Cite the original work for its findings. Save a collection to share your selection of sources.