Searcharxiv⌕ Search

arXiv · 2609.35141

Two-Timescale Reinforcement Learning for Real-Time Optimization and Economic NMPC: Experimental Validation

Abstract

We propose a reinforcement learning (RL) framework that tunes both Real-Time Optimization (RTO) and Economic Nonlinear Model Predictive Control (ENMPC) to address plant--model mismatch in process systems. Drawing on modifier-adaptation concepts, the method parameterizes the dynamic model, stage costs, constraints, and RTO modifiers, and uses Q-learning to adjust these parameters at two timescales: a fast update for the ENMPC layer and a slow update for the RTO layer. The framework is experimentally validated on a laboratory rig emulating a three-well subsea oil-production network. Using plant measurement data, the proposed RTO-RLMPC scheme achieves 8.6% higher economic profit than nominal ENMPC, preserves input feasibility of the deployed ENMPC policy and empirically satisfies the path constraints under disturbances and measurement noise, and drives the learned model parameters toward reference plant values, independently identified from experimental data, along the directions that most affect the economic objective. This work provides one of the first experimental demonstrations of RL-tuned ENMPC integrated with RTO.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Saket Adhau, Jose Matias, Sebastien Gros, Sigurd Skogestad. 2026-09-28. Two-Timescale Reinforcement Learning for Real-Time Optimization and Economic NMPC: Experimental Validation. https://arxiv.org/abs/2609.35141

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Reinforcement Learning to Initialize Newton-Raphson for AC Power Flow with Quantum Annealing-Based Environment Updates

The Newton-Raphson (NR) method is widely used for solving power flow (PF) equations due to its quadratic convergence. However, its performance deteriorates under poor initialization or extreme operating scenarios, e.g., high levels of renewable energy penetration. We propose the use of reinforcement learning (RL) to optimize the initialization of NR, and introduce a quantum-enhanced RL environment update mechanism that addresses the combinatorially large action space at each RL timestep by formulating the voltage adjustment task as a Quadratic Unconstrained Binary Optimization (QUBO) problem, solved with an Ising machine. RL initialization is benchmarked against flat start and start from the DC (linearized) PF solution on a standard 4-bus system, Iwamoto's ill-conditioned 11-bus system, and the IEEE 118-bus system under normal and stressed loading and reactive power limits, with verified operational solutions. On all systems, a supervised initializer refined by RL requires fewer NR iterations than flat and DC starts and than the same initializer without RL, for all seeds. For example, on the 118-bus system under normal and stressed loading, it reached 2.04 and 2.86 NR iterations, compared with 3.02 and 5.13 from DC start and 2.61 and 3.09 without RL. In wall-clock time, this pays off only for an initializer integrated into the solver and reused for many solves on a fixed topology. On the 4-bus system, a quantum-enhanced RL agent with a quantum-inspired annealer moved challenging initial states that required 29 and 44 NR iterations to initializations that required three NR iterations within one RL timestep.

eess.SY↗

Safety for Weakly-Hard Control Systems via Graph-Based Barrier Functions

Despite significant advancement in technology, communication and computational failures are still prevalent in safety-critical engineering applications. Often, networked control systems experience packet dropouts, leading to open-loop behavior that significantly affects the behavior of the system. Similarly, in real-time control applications, control tasks frequently experience computational overruns and thus occasionally no new actuator command is issued. This article addresses the safety verification and controller synthesis problem for a class of control systems subject to weakly-hard constraints, i.e., a set of window-based constraints where the number of failures are bounded within a given time horizon. The results are based on a new notion of graph-based barrier functions that are specifically tailored to the considered system class, offering a set of constraints whose satisfaction leads to safety guarantees despite such failures. Subsequent reformulations of the safety constraints are proposed to alleviate conservatism and improve computational tractability, and the resulting trade-offs are discussed. Finally, several numerical case studies including linear and polynomial systems demonstrate the effectiveness of the proposed approach.

eess.SY↗

Model-Free Output Feedback Stabilization via Policy Gradient Methods

Stabilizing a dynamical system is a fundamental problem that serves as a cornerstone for many complex tasks in the field of control systems. The problem becomes challenging when the system model is unknown. Among the Reinforcement Learning (RL) algorithms that have been successfully applied to solve problems pertaining to unknown linear dynamical systems, the policy gradient (PG) method stands out due to its ease of implementation and can solve the problem in a model-free manner. However, most of the existing works on PG methods for unknown linear dynamical systems assume full-state feedback. In this paper, we take a step towards model-free learning for partially observed linear dynamical systems with output feedback and focus on the fundamental stabilization problem of the system. We propose an algorithmic framework that stretches the boundary of PG methods to the problem without global convergence guarantees. We show that by leveraging zeroth-order PG update based on system trajectories and its convergence to stationary points, the proposed algorithms return a stabilizing output feedback policy for discrete-time linear dynamical systems. We also explicitly characterize the sample complexity of our algorithm and verify the effectiveness of the algorithm using numerical examples.

eess.SY↗