arXiv · 2406.06506
Online Newton Method for Bandit Convex Optimisation
Abstract
We introduce a computationally efficient algorithm for zeroth-order bandit convex optimisation and prove that in the adversarial setting its regret is at most $d^{3.5} \sqrt{n} \mathrm{polylog}(n, d)$ with high probability where $d$ is the dimension and $n$ is the time horizon. In the stochastic setting the bound improves to $M d^{2} \sqrt{n} \mathrm{polylog}(n, d)$ where $M \in [d^{-1/2}, d^{-1 / 4}]$ is a constant that depends on the geometry of the constraint set and the desired computational properties.
Explore related subjects
Keep this discovery
Hidde Fokkema, Dirk van der Hoeven, Tor Lattimore, Jack J. Mayo. 2024-06-10. Online Newton Method for Bandit Convex Optimisation. https://arxiv.org/abs/2406.06506
Cite the original work for its findings. Save a collection to share your selection of sources.