arXiv · 2607.15067
Kernel weighted importance sampling for off-policy evaluation in contextual bandits
Abstract
This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including weighted importance sampling), particularly under behaviour policy miss-specification. The benefit of Kernel-WIS is derived from combining the bounded property of weighted importance sampling with the linearity of vanilla importance sampling.
Explore related subjects
Keep this discovery
Joshua Spear, Matthieu Komorowski, Rebecca Pope, Erica E. M. Moodie. 2026-07-16. Kernel weighted importance sampling for off-policy evaluation in contextual bandits. https://arxiv.org/abs/2607.15067
Cite the original work for its findings. Save a collection to share your selection of sources.