SearcharxivSearch

arXiv subjects

Xuanhe Wang

Publications and source records attributed to Xuanhe Wang.

3 recordsLinked to original sources

State-Space Model-Enabled Reinforcement Learning for Magnetic Configuration Controlon EXL-50U

Accurate feedback control of the plasma current ($I_p$) and centroid position $(R_c,Z_c)$ is essential for the stable operation of spherical torus (ST) plasmas. Conventional proportional-integral-derivative (PID) controllers require extensive manual tuning and struggle with the fast, strongly coupled dynamics that arise as plasma performance improves. Reinforcement learning (RL) has recently emerged as a promising alternative to such complex magnetic control problems, yet its practical deployment on ST devices remains challenging. This paper presents a practical RL controller for the EXL-50U ST, trained within a rigid RZIP state-space model (SSM) that enables efficient offline policy learning. A lightweight plasma position reconstructor is developed to estimate $(R_c,Z_c)$ from magnetic probe signals within the real-time control cycle. The trained policy is seamlessly deployed on the EXL-50U plasma control system, achieving stable regulation of $I_p$ and $(R_c,Z_c)$ and sustaining discharges up to 650 ms under RL control. These results demonstrate the feasibility and practical potential of model-informed RL for magnetic configuration control in ST devices, offering a promising direction beyond conventional PID-based schemes.

physics.plasm-ph

Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment

Managing divertor heat loads is a central challenge for compact, high-power tokamaks. To increase local flux expansion and decouple the dissipation volume from the core, EHL-2 adopts the X-point target (XPT) divertor. This requires the secondary X-point to remain on the divertor leg; displacement degrades the topology and exhaust geometry. Current experiments, including EXL-50U discharges, rely on precomputed feedforward waveforms with PID loops on global quantities. Lacking dedicated closed-loop feedback for the secondary null, XPT operation is repeatable but not routine. We formulate XPT feedback as a multi-objective reinforcement learning (RL) control problem in a free-boundary environment calibrated to EXL-50U discharge #13906. To address strong coupling among plasma current, shape, and null constraints - where reward scalarisation collapses objective-specific temporal credit - we develop Advantage Aggregation (AdvA). AdvA preserves objective-wise temporal credit before worst-objective-aware nonlinear scalarisation and introduces a residual correction to policy updates. AdvA-PPO is evaluated against Reward-PPO and a feedforward-plus-PID baseline under nominal operation, measurement uncertainties, and unseen initial equilibria. On a 500 ms rollout, AdvA-PPO raises the mean worst-channel score from 0.23 to 0.81 over Reward-PPO, reducing X-point flux RMSE by ~20x. Under combined measurement uncertainties, it is the only learned controller completing the horizon while retaining a usable XPT shape. Multi-initialization fine-tuning enables a single AdvA-PPO policy to complete full-horizon operation across divertor and limiter initial equilibria. These results provide a simulation-based foundation for future real-time XPT validation on EXL-50U.

physics.plasm-ph

Reinforcement learning for vertical position control on the EXL-50U spherical tokamak

Vertical position control is essential for sustaining high-performance operation in spherical tokamaks, where increased plasma elongation introduces stringent requirements on fast and robust stabilization. This work presents an experimentally validated reinforcement-learning(RL)-based vertical position control framework for the EXL-50U spherical tokamak. A high-fidelity discharge-reconstructed simulation environment is developed by integrating physics-based plasma-circuit models with experimental equilibrium information, enabling systematic controller synthesis and sim-to-real evaluation. Within this framework, RL is benchmarked in simulation against operational proportional--integral--derivative (PID) and model-based linear quadratic regulator (LQR) controllers under identical plant dynamics, actuator constraints, and measurement imperfections.Simulation results show that RL achieves tracking accuracy comparable to PID with consistently lower vertical-stabilization coil effort, while lightweight integral compensation improves robustness against residual model--plant mismatch. The RL controller is subsequently deployed on EXL-50U for closed-loop experiments. Across more than ten discharges with RL takeover, stable vertical regulation is achieved within the controlled windows. For seven representative discharges, RL maintains millimetre-scale tracking accuracy comparable to the operational PID controller (MAE typically ~ 1-5 mm) while consistently reducing actuator effort. These results demonstrate the feasibility of learning-based plasma control on a real spherical tokamak and establish a practical pathway toward future fusion control systems.

physics.plasm-ph