arXiv · 2607.15080
Evaluating covariate balance for long time horizon Markov decision processes
Abstract
This article explores the application of covariate balance diagnostics for detecting the presence of hidden confounding/model miss-specification in studies applying offline reinforcement learning (RL) to deriving optimal treatment recommendations. The results demonstrate that, either there is a high risk of bias within existing offline RL studies for treatment recommendations or, existing covariate balance metrics are not sufficient to assess such studies. Regardless, existing offline RL studies cannot be concluded as being statistically robust. The conclusions propose future research directions for obtaining more methodologically robust applications of offline RL to treatment recommendation problems.
Explore related subjects
Keep this discovery
Joshua Spear, Rebecca Pope, Neil J Sebire. 2026-07-16. Evaluating covariate balance for long time horizon Markov decision processes. https://arxiv.org/abs/2607.15080
Cite the original work for its findings. Save a collection to share your selection of sources.