SearcharxivSearch

arXiv subjects

Changquan Yu

Publications and source records attributed to Changquan Yu.

2 recordsLinked to original sources

Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment

Managing divertor heat loads is a central challenge for compact, high-power tokamaks. To increase local flux expansion and decouple the dissipation volume from the core, EHL-2 adopts the X-point target (XPT) divertor. This requires the secondary X-point to remain on the divertor leg; displacement degrades the topology and exhaust geometry. Current experiments, including EXL-50U discharges, rely on precomputed feedforward waveforms with PID loops on global quantities. Lacking dedicated closed-loop feedback for the secondary null, XPT operation is repeatable but not routine. We formulate XPT feedback as a multi-objective reinforcement learning (RL) control problem in a free-boundary environment calibrated to EXL-50U discharge #13906. To address strong coupling among plasma current, shape, and null constraints - where reward scalarisation collapses objective-specific temporal credit - we develop Advantage Aggregation (AdvA). AdvA preserves objective-wise temporal credit before worst-objective-aware nonlinear scalarisation and introduces a residual correction to policy updates. AdvA-PPO is evaluated against Reward-PPO and a feedforward-plus-PID baseline under nominal operation, measurement uncertainties, and unseen initial equilibria. On a 500 ms rollout, AdvA-PPO raises the mean worst-channel score from 0.23 to 0.81 over Reward-PPO, reducing X-point flux RMSE by ~20x. Under combined measurement uncertainties, it is the only learned controller completing the horizon while retaining a usable XPT shape. Multi-initialization fine-tuning enables a single AdvA-PPO policy to complete full-horizon operation across divertor and limiter initial equilibria. These results provide a simulation-based foundation for future real-time XPT validation on EXL-50U.

physics.plasm-ph

Reinforcement learning for vertical position control on the EXL-50U spherical tokamak

Vertical position control is essential for sustaining high-performance operation in spherical tokamaks, where increased plasma elongation introduces stringent requirements on fast and robust stabilization. This work presents an experimentally validated reinforcement-learning(RL)-based vertical position control framework for the EXL-50U spherical tokamak. A high-fidelity discharge-reconstructed simulation environment is developed by integrating physics-based plasma-circuit models with experimental equilibrium information, enabling systematic controller synthesis and sim-to-real evaluation. Within this framework, RL is benchmarked in simulation against operational proportional--integral--derivative (PID) and model-based linear quadratic regulator (LQR) controllers under identical plant dynamics, actuator constraints, and measurement imperfections.Simulation results show that RL achieves tracking accuracy comparable to PID with consistently lower vertical-stabilization coil effort, while lightweight integral compensation improves robustness against residual model--plant mismatch. The RL controller is subsequently deployed on EXL-50U for closed-loop experiments. Across more than ten discharges with RL takeover, stable vertical regulation is achieved within the controlled windows. For seven representative discharges, RL maintains millimetre-scale tracking accuracy comparable to the operational PID controller (MAE typically ~ 1-5 mm) while consistently reducing actuator effort. These results demonstrate the feasibility of learning-based plasma control on a real spherical tokamak and establish a practical pathway toward future fusion control systems.

physics.plasm-ph