arXiv · 2410.09973
Gradient Span Algorithms Make Predictable Progress in High Dimension
Abstract
We prove that all 'gradient span algorithms' have asymptotically deterministic behavior on scaled Gaussian random functions as the dimension tends to infinity. This is a functional generalization of similar results for random quadratic functions and spin glasses. They explain the counterintuitive phenomenon that different training runs of many large machine learning models result in approximately equal cost curves despite random initialization on a complicated non-convex landscape. This 'predictable progress' phenomenon is exploited by the AutoML community: Since the optimization progress of a single run is already representative, multiple retries with the same hyperparameters are not necessary.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Felix Benning, Leif Döring. 2024-10-13. Gradient Span Algorithms Make Predictable Progress in High Dimension. https://arxiv.org/abs/2410.09973
Cite the original work for its findings. Save a collection to share your selection of sources.