arXiv · 2603.25029
Logarithmic High-Probability Regret for Online Convex Optimization with Two-Point Bandit Feedback
Abstract
We study online convex optimization (OCO) with two-point bandit feedback against a non-anticipating adaptive adversary. In this setting, a learner competes with an adversarial sequence of convex losses while observing each loss only through two function evaluations. For strongly convex losses, Agarwal, Dekel, and Xiao~\citeyearpar{agarwal2010optimal} proved a comparator-wise logarithmic regret bound in expectation. Consequently, by minimizing outside the probability space, their result yields a pseudo-regret guarantee of the form $\EB A_T-\min_{x\in\mathcal K}\EB L_T(x)$, where $A_T$ is the algorithm's two-query cumulative loss and $L_T(x)$ is the comparator's cumulative loss. They asked whether a logarithmic high-probability guarantee is achievable in the same two-point strongly convex setting. Our main theorem provides the corresponding fixed-comparator high-probability statement: for any comparator $x\in\mathcal K$ fixed independently of the algorithmic random directions, the standard two-point projected gradient method guarantees, with probability at least $1-\delta$, a two-query regret bound of order \[ O\left(\frac{dG^2}{\mu}\left(\log T+\log(1/\delta)\right)+dGD\log(1/\delta)+G\log T\left(1+\frac{D}{r}\right)\right). \] At the comparator-wise level, our leading horizon-dependent term is linear in $d$, compared with the $d^2$-type term in the original analysis of Agarwal, Dekel, and Xiao. The key ingredient is a high-confidence analysis that simultaneously absorbs the martingale error into strong convexity and preserves the linear-in-dimension estimator control of the two-point method. A deterministic covering argument then yields a realized full-comparator guarantee against $\min_{x\in\mathcal K}L_T(x)$, preserving logarithmic dependence on $T$ at the cost of the standard covering-number factor.
Explore related subjects
Keep this discovery
Haishan Ye. 2026-03-26. Logarithmic High-Probability Regret for Online Convex Optimization with Two-Point Bandit Feedback. https://arxiv.org/abs/2603.25029
Cite the original work for its findings. Save a collection to share your selection of sources.