arXiv · 2609.12264
Adaptive Chemotherapy Control under Tumor Heterogeneity via Reinforcement Learning
Abstract
Designing effective chemotherapy regimens is hindered by tumor heterogeneity and drug resistance, which complicate the deployment of patient-specific model-based optimal control across diverse populations. We develop and compare closed-loop deep reinforcement learning (DRL) dosing policies with continuous (TD3) and discrete (DQN) action spaces trained on a high-dimensional heterogeneous tumor model. The DRL policies are benchmarked against a Pontryagin's Maximum Principle (PMP)-derived open-loop benchmark. We assess generalization under parametric heterogeneity using a 100-patient virtual cohort with plus or minus 10 percent uniform perturbations in growth and drug-sensitivity parameters. Across this cohort, TD3 achieves higher average tumor reduction, while DQN yields tighter inter-patient dosing consistency, revealing a clear efficacy-consistency trade-off in this study. Our simulations assume full observation of all tumor subpopulations; translation to sparse and noisy clinical measurements will require partial-observability formulations and/or state estimation. Overall, the results show that simulation-trained DRL can learn state-dependent feedback dosing policies that complement open-loop optimal control benchmarks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bereket Sitotaw Kidane, Md Samiul Haque Motayed, Shuo Wang. 2026-09-10. Adaptive Chemotherapy Control under Tumor Heterogeneity via Reinforcement Learning. https://arxiv.org/abs/2609.12264
Cite the original work for its findings. Save a collection to share your selection of sources.