arXiv · 2212.14468
An Instrumental Variable Approach to Confounded Off-Policy Evaluation
Abstract
Off-policy evaluation (OPE) is a method for estimating the return of a target policy using some pre-collected observational data generated by a potentially different behavior policy. In some cases, there may be unmeasured variables that can confound the action-reward or action-next-state relationships, rendering many existing OPE approaches ineffective. This paper develops an instrumental variable (IV)-based method for consistent OPE in confounded Markov decision processes (MDPs). Similar to single-stage decision making, we show that IV enables us to correctly identify the target policy's value in infinite horizon settings as well. Furthermore, we propose an efficient and robust value estimator and illustrate its effectiveness through extensive simulations and analysis of real data from a world-leading short-video platform.
Explore related subjects
Keep this discovery
Yang Xu, Jin Zhu, Chengchun Shi, Shikai Luo, Rui Song. 2022-12-29. An Instrumental Variable Approach to Confounded Off-Policy Evaluation. https://arxiv.org/abs/2212.14468
Cite the original work for its findings. Save a collection to share your selection of sources.