arXiv · 2609.35141
Two-Timescale Reinforcement Learning for Real-Time Optimization and Economic NMPC: Experimental Validation
Abstract
We propose a reinforcement learning (RL) framework that tunes both Real-Time Optimization (RTO) and Economic Nonlinear Model Predictive Control (ENMPC) to address plant--model mismatch in process systems. Drawing on modifier-adaptation concepts, the method parameterizes the dynamic model, stage costs, constraints, and RTO modifiers, and uses Q-learning to adjust these parameters at two timescales: a fast update for the ENMPC layer and a slow update for the RTO layer. The framework is experimentally validated on a laboratory rig emulating a three-well subsea oil-production network. Using plant measurement data, the proposed RTO-RLMPC scheme achieves 8.6% higher economic profit than nominal ENMPC, preserves input feasibility of the deployed ENMPC policy and empirically satisfies the path constraints under disturbances and measurement noise, and drives the learned model parameters toward reference plant values, independently identified from experimental data, along the directions that most affect the economic objective. This work provides one of the first experimental demonstrations of RL-tuned ENMPC integrated with RTO.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Saket Adhau, Jose Matias, Sebastien Gros, Sigurd Skogestad. 2026-09-28. Two-Timescale Reinforcement Learning for Real-Time Optimization and Economic NMPC: Experimental Validation. https://arxiv.org/abs/2609.35141
Cite the original work for its findings. Save a collection to share your selection of sources.