arXiv · 2609.21445
Optimal Randomized Proper Online Learning
Abstract
We prove that the optimal expected mistake bound of online learning a function class $\mathcal{H}$ by a randomized proper learning algorithm is $O(\mathtt{L}(\mathcal{H}) \log T)$, where $\mathtt{L}(\mathcal{H})$ is the Littlestone dimension of $\mathcal{H}$ and $T$ is the time horizon. Our result improves upon the previously best known bound of $O(\mathtt{L}(\mathcal{H}) \log^6 T)$ given by Daskalakis and Golowich (STOC 2022), and is optimal up to a universal constant for worst-case classes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zachary Chase, Idan Mehalel. 2026-09-18. Optimal Randomized Proper Online Learning. https://arxiv.org/abs/2609.21445
Cite the original work for its findings. Save a collection to share your selection of sources.