Searcharxiv⌕ Search

arXiv · 2609.38695

Always-On Experimentation

Abstract

Generative AI has dramatically accelerated the rate at which new treatments---from novel pharmaceuticals to online marketing campaigns---can be conceived and deployed. As a result, modern experimentation platforms often run continuously, with treatments added as they are ready and removed when they underperform. We formalize this "Always-On" experimental setting, in which treatments can be dynamically generated, added to, and removed from a running experiment, and study the statistical problem of deciding whether to accept or reject each treatment while controlling for the false discovery rate. We develop sequential tests that achieve time-uniform Type-I error control under arbitrary stopping times and "predictable" treatment schedules. Our approach builds on the testing-by-betting framework: we construct test supermartingales for testing the average treatment effect of each treatment, and show that the construction of these test supermartingales is growth-rate optimal in an almost-sure sense.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ricardo J. Sandoval, David Arbour, Avi Feller, Michael I. Jordan. 2026-09-30. Always-On Experimentation. https://arxiv.org/abs/2609.38695

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Sample-Efficient "Clustering and Conquer" Procedures for Parallel Large-Scale Ranking and Selection

This work aims to improve the sample efficiency of parallel large-scale ranking and selection (R&S) problems by leveraging correlation information. We modify the commonly used "divide and conquer" framework in parallel computing by adding a correlation-based clustering step, transforming it into "clustering and conquer". Theoretically, we develop a novel gradient-based analysis framework and show that this seemingly simple modification substantially improves the performance of large-scale R&S procedures. Our approach enjoys two key advantages: (1) it does not require highly accurate correlation estimation or precise clustering, and (2) it can be seamlessly integrated with various existing fixed-precision and fixed-budget R&S procedures while achieving optimal sample complexity. We also introduce a new parallel clustering algorithm tailored to large-scale settings. Finally, in large-scale AI applications such as neural architecture search, our methods demonstrate superior performance.

stat.ME↗

Accuracy of Uniform Inference on Fine Grid Points

Uniform confidence bands are widely used in empirical analysis for uncertainty qualification of nonparametric inference of unknown functions. A variety of simple implementation methods, including multiplier bootstrap, have been proposed and theoretically justified. However, an implementation over a literally continuous index set is generally computationally infeasible, and practitioners therefore compute the critical value by evaluating the statistic on a finite evaluation grid. This paper quantifies the effect of this discretization on coverage accuracy and consider how fine the evaluation grid must be for a multiplier bootstrap procedure over finite grid points to deliver valid uniform confidence bands. Specifically, we first illustrate that coarse grids can invalidate uniform inference. For a nonparametric estimator based on kernel smoothing, we establish conditions under which uniform coverage converges to zero, even when the number of evaluation points diverges. Simulations further show that increasing the sample size can worsen coverage on a fixed coarse grid, whereas approximation errors from sources other than discretization decrease. We then consider general empirical processes and derive an upper bound on the coverage error of uniform confidence bands calibrated using multiplier bootstrap critical values computed on a finite grid. The bound distinguishes discretization from the remaining approximation error on the grid. Also, for a broad class of empirical processes arising from nonparametric estimators based on kernel smoothing, we provide primitive sufficient conditions for negligible discretization error. These conditions yield sufficient grid rules expressed in terms of the sample size and bandwidth.

stat.ME↗

Overstuffed sandwiches and separation anxiety: finite-sample variance estimation for penalized GEE with near-separated binary data

Penalized generalized estimating equations (PGEE) stabilize point estimation for longitudinal binary data under near-separation, but inference still depends on how the sandwich variance is corrected. Existing corrections for PGEE can overadjust in high-leverage directions, require restrictive pooling assumptions, or add global regularization without explaining the bias. We establish first-order asymptotics for PGEE along convergent interior-root sequences and derive a matrix characterization of the parameter-specific overcorrection induced by full leverage adjustment. Finite-sample calibration is limited by both mean bias and the variability of leverage-corrected variance estimates. We propose $\hat{V}_{AR}$, which keeps the score-level leverage correction and adds a finite-sample upward translation dominated at first order by the finite-population factor, with a smaller centering term. In simulations, $\hat{V}_{AR}$ gives conservative or near-nominal type I error in low-event, small-$N$ settings, including $N = 10$, where several standard corrections remain anti-conservative and pooling estimators are unavailable for unbalanced designs.

stat.ME↗