arXiv · 2609.33489
Predicting Block-Coordinate Performance via Cross-Curvature
Abstract
Simultaneous and sequential block updates are two basic optimization strategies used across machine learning, such as neural-network training, federated learning, and low-rank adaptation. Choosing between them is difficult because their relative advantage depends on both the objective geometry and the number of iterations. We develop a unified theory for comparing Jacobi (JC), Gauss--Seidel (GS), and partially sequential deterministic block-gradient updates. Our analysis expresses the one-step loss difference through cross-block curvature, with an $O(η^3)$ remainder, where $η$ is the learning rate. We derive a signed loss comparison after $K$ iterations with $O(Kη^3)$ error under regularity conditions and $ηK\le T$ for fixed $T$, identifying the better method when the predicted difference exceeds this error. We evaluate these formulas along observed training trajectories across different machine learning settings. Over 500 iterations, our theory correctly identifies the lower-loss method in 98.0\% of iterations for the neural network, 83.4\% for federated learning, and 97.6\% for LoRA. Applying the loss recursion at each step using the measured parameter difference raises these rates to 100.0\%, 93.2\%, and 99.6\%, respectively.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shengkun Zhu, Jinshan Zeng, Zhiqiang Kou, Yongxin Tong, Yang Liu. 2026-09-27. Predicting Block-Coordinate Performance via Cross-Curvature. https://arxiv.org/abs/2609.33489
Cite the original work for its findings. Save a collection to share your selection of sources.