Searcharxiv⌕ Search

arXiv · 2610.05517

Minimax Call Complexity, Rounds and Admissible Noise in Zeroth-Order Optimization of Highly Smooth Functions

Abstract

We study zeroth-order optimization of convex and strongly convex functions with Hölder smoothness of order $β>2$ from values corrupted by random noise with bias independent of the method's random directions, a Lipschitz systematic error and a bounded adversarial error. We measure the number of function evaluations (calls), of sequential rounds of parallel calls, and the largest tolerable systematic and adversarial errors. For strongly convex functions with a bounded Hessian, we prove a lower bound of order $d^2\varepsilon^{-β/(β-1)}$ on the number of calls, where $d$ is the dimension and $\varepsilon$ the accuracy, and find the sharp power of $d$ under the stronger Hilbert--Schmidt smoothness for all $β>2$. An accelerated finite-difference method attains them in $O(\sqrtκ\log(1/\varepsilon))$ rounds for condition number $κ$, like Nesterov's method with exact gradients, and tolerates the optimal systematic and, up to logarithms, adversarial errors; rate-optimal methods need $Ω(\log d)$ rounds. The same number of calls is attained when the Hessian grows at most linearly with the gradient and, for centered noise, without an upper bound on the Hessian, also in logarithmically many rounds. For convex functions, our lower bound improves the known one polynomially in $d$ and is tight at the curvature of the hard instances; the systematic level is again optimal. These bounds extend to general norms and constraints; on the $\ell_1$ ball, mirror descent needs calls nearly linear in $d$. Experiments confirm the theory.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Timofei Loginov, Darina Dvinskikh, Nazarii Tupitsa, Osman Osmanov, Yuriy Dorn, Alexander Gasnikov. 2026-10-04. Minimax Call Complexity, Rounds and Admissible Noise in Zeroth-Order Optimization of Highly Smooth Functions. https://arxiv.org/abs/2610.05517

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Multilevel Regularized Newton Methods with Fast Convergence Rates

We introduce new multilevel methods for solving large-scale unconstrained optimization problems. Specifically, the philosophy of multilevel methods is applied to Newton-type methods that regularize the Newton sub-problem using second order information from a coarse (low dimensional) sub-problem. The new \emph{regularized multilevel methods} provably converge from any initialization point and enjoy faster convergence rates than Gradient Descent. In particular, for arbitrary functions with Lipschitz continuous Hessians, we show that their convergence rate interpolates between the rate of Gradient Descent and that of the cubic Newton method. If, additionally, the objective function is assumed to be convex, then the proposed method converges with the fast $\mathcal{O}(k^{-2})$ rate. Hence, since the updates are generated using a \emph{coarse} model in low dimensions, the theoretical results of this paper significantly speed-up the convergence of Newton-type or preconditioned gradient methods in practical applications. Preliminary numerical results suggest that the proposed multilevel algorithms are significantly faster than current state-of-the-art methods.

math.OC↗

Adaptive Algorithms for Robust Phase Retrieval

This paper considers robust phase retrieval, which can be cast as a nonsmooth and nonconvex optimization problem. We propose two first-order algorithms with adaptive step sizes: the subgradient algorithm (AdaSubGrad) and the inexact proximal linear algorithm (AdaIPL). Our contribution lies in a novel design of adaptive step sizes based on quantiles of the absolute residuals. We analyze local linear convergence of both algorithms across different hyperparameter regimes under i.i.d. centered sub-Gaussian measurements, a stability condition linking the measurement distribution and the corruption level, and an additional uniform small-ball condition. Numerical experiments on synthetic datasets and image recovery also demonstrate that our methods are competitive with existing methods in the literature that utilize predetermined (possibly impractical) step sizes, such as subgradient methods and the inexact proximal linear method.

math.OC↗

A Simple yet Highly Accurate Prediction-Correction Algorithm for Time-Varying Optimization

This paper proposes a simple yet highly accurate prediction-correction algorithm, SHARP, for unconstrained time-varying optimization problems. Its prediction is based on an extrapolation derived from the Lagrange interpolation of past solutions. Since this extrapolation can be computed without Hessian matrices or even gradients, the computational cost is low. To ensure the stability of the prediction, the algorithm includes an acceptance condition that rejects the prediction when the update is excessively large. The proposed method achieves a tracking error of $O(h^{p})$, where $h$ is the sampling period, assuming that the $p$-th derivative of the target trajectory is bounded and the convergence of the correction step is locally linear. We also prove that the method can track a trajectory of stationary points even if the objective function is non-convex. Numerical experiments demonstrate the high accuracy of the proposed algorithm.

math.OC↗