arXiv · 2605.16017
Accelerated Gradient Descent for Faster Convergence with Minimal Overhead
Abstract
In this paper, we present CT-AGD (Curvature-Tuned Accelerated Gradient Descent), an optimization method for non-convex optimization problems in deep learning training tasks. CT-AGD is a general boosting procedure that accelerates first-order methods by explicitly capturing the local curvature using finite-difference quotients, and the development of heuristics aimed at mitigating noise and bias introduced by stochastic mini-batch training. CT-AGD has a comparable storage and computational overhead as adaptive gradient methods such as Adam. Our extensive experiments demonstrate that CT-AGD achieves the same level of accuracy as the baseline first-order methods, yet reduces the required training epochs by 33% on average.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Manuel Graca, L. Miguel Silveira, Arlindo Oliveira, Frank Liu. 2026-05-15. Accelerated Gradient Descent for Faster Convergence with Minimal Overhead. https://arxiv.org/abs/2605.16017
Cite the original work for its findings. Save a collection to share your selection of sources.