SearcharxivSearch

arXiv subjects

Jack Mayo

Publications and source records attributed to Jack Mayo.

4 recordsLinked to original sources

An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction

We present an oracle-efficient, near-optimal algorithm for linear contextual bandits with adversarial losses and stochastic action sets, only requiring a linear optimization oracle for the action sets in each round. Our approach reduces this setting to misspecification-robust adversarial linear bandits with fixed action sets. Without knowledge of the context distribution or access to a context simulator, the algorithm achieves $\widetilde{\mathcal{O}}(\min\{d^2\sqrt{T}, \sqrt{d^3T\log K}\})$ regret and runs in $\mathrm{poly}(d,T)$ time plus $\mathrm{poly}(d,T)$ calls to the linear optimization oracles, where $d$ is the feature dimension, $K$ is an upper bound on the number of actions in each round, and $T$ is number of rounds. This resolves the open question by Liu et al. (2023) on whether one can obtain $\mathrm{poly}(d)\sqrt{T}$ regret in polynomial time independent of the number of actions. For the important class of combinatorial bandits with adversarial losses and stochastic action sets, our algorithm is the first to achieve $\mathrm{poly}(d)\sqrt{T}$ regret in polynomial time, while no prior algorithm achieves even $o(T)$ regret in polynomial time to our knowledge. When a simulator is available, the regret bound can be improved to $\widetilde{\mathcal{O}}(d\sqrt{L^\star})$, where $L^\star$ is the cumulative loss of the best policy.

cs.LG

First- and Second-Order Bounds for Adversarial Linear Contextual Bandits

We consider the adversarial linear contextual bandit setting, which allows for the loss functions associated with each of $K$ arms to change over time without restriction. Assuming the $d$-dimensional contexts are drawn from a fixed known distribution, the worst-case expected regret over the course of $T$ rounds is known to scale as $\tilde O(\sqrt{Kd T})$. Under the additional assumption that the density of the contexts is log-concave, we obtain a second-order bound of order $\tilde O(K\sqrt{d V_T})$ in terms of the cumulative second moment of the learner's losses $V_T$, and a closely related first-order bound of order $\tilde O(K\sqrt{d L_T^*})$ in terms of the cumulative loss of the best policy $L_T^*$. Since $V_T$ or $L_T^*$ may be significantly smaller than $T$, these improve over the worst-case regret whenever the environment is relatively benign. Our results are obtained using a truncated version of the continuous exponential weights algorithm over the probability simplex, which we analyse by exploiting a novel connection to the linear bandit setting without contexts.

cs.LG

The Distribution of AGN Covering Factors

We review our knowledge of the most basic properties of the AGN obscuring region - its location, scale, symmetry, and mean covering factor - and discuss new evidence on the distribution of covering factors in a sample of ~9000 quasars with WISE, UKIDSS, and SDSS photometry. The obscuring regions of AGN may be in some ways more complex than we thought - multi-scale, not symmetric, chaotic - and in some ways simpler - with no dependence on luminosity, and a covering factor distribution that may be determined by the simplest of considerations - e.g. random misalignments.

astro-ph.CO

The Mid-Infrared Environments of High-Redshift Radio Galaxies

Taking advantage of the impressive sensitivity of Spitzer to detect massive galaxies at high redshift, we study the mid-infrared environments of powerful, high-redshift radio galaxies at 1.2 -0.1 (AB), in the fields of 48 radio galaxies at 1.2 1.2. Using a counts-in-cell analysis, we identify a field as overdense when 15 or more red IRAC sources are found within 1arcmin (i.e.,~0.5Mpc at 1.2 1.2.

astro-ph.CO