arXiv · 2610.01044
Augmented Patient Preference Incorporated Reinforcement Learning (APP-RL) to Estimate the Optimal Dynamic Treatment Regime
Abstract
Dynamic treatment regimes (DTRs) are sequential decision rules that individualize treatments to each patient at each treatment stage adapting to their past clinical course. Existing literature typically accommodates each individual's medical history, but overlooks a patient's preferences. We propose a method that incorporates a patient's latent preferences through data augmentation into a tree-based reinforcement learning method to estimate optimal dynamic treatment regimes for multi-stage, multi-treatment settings. For each patient at each stage, we derive the posterior distribution of preferences given responses to a questionnaire, and then subsequently weight multiple outcomes with the estimated preferences to identify the optimal stage-wise personalized decision. For multiple stage situations, we grow a decision tree at each stage and implement the algorithm recursively using backward induction. Our proposed method, named Augmented Patient Preference incorporated Reinforcement Learning (APP-RL) is robust, efficient, and leads to interpretable DTR estimation. The finite-sample performances of the proposed method has been thoroughly evaluated through simulation studies.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yingchao Zhong, Lu Wang. 2026-10-01. Augmented Patient Preference Incorporated Reinforcement Learning (APP-RL) to Estimate the Optimal Dynamic Treatment Regime. https://arxiv.org/abs/2610.01044
Cite the original work for its findings. Save a collection to share your selection of sources.