Paper deep dive
Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment
Siqi Ding, Xuanhe Wang, Pei Guo, Guoyang Shi, Changquan Yu, Yiting Wang, Xianming Song, Xiang Gu, Zhengyuan Chen, Lei Xing, Yapeng Zhang, Jianguo Chen, Tianyuan Liu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/24/2026, 5:27:27 AM
Summary
This paper presents a reinforcement learning framework for controlling the X-point target (XPT) magnetic configuration in tokamaks, specifically calibrated for the EXL-50U device. The authors introduce Advantage Aggregation (AdvA), a multi-objective RL algorithm that preserves objective-wise temporal credit before scalarization, addressing the strong coupling between plasma current, shape, and null constraints. Evaluated in a free-boundary simulation environment calibrated to EXL-50U discharge #13906, AdvA-PPO significantly outperforms conventional Reward-PPO and feedforward-plus-PID baselines in maintaining X-point flux and handling measurement uncertainties.
Entities (13)
Relation Signals (7)
AdvA-PPO → isbasedon → Advantage Aggregation
confidence 98% · Instantiated with proximal policy optimisation (PPO), AdvA-PPO... develop Advantage Aggregation (AdvA)
EHL-2 → adopts → X-point target
confidence 95% · EHL-2 adopts the X-point target (XPT) divertor
AdvA-PPO → outperforms → Reward-PPO
confidence 95% · AdvA-PPO raises the mean worst-channel score from 0.23 to 0.81 over Reward-PPO
FGE → simulates → EXL-50U
confidence 95% · calibrated to EXL-50U discharge #13906... FGE supplies Ip, coil currents, LCFS, and X-point features
AdvA-PPO → outperforms → PID
confidence 90% · Under combined measurement uncertainties, it is the only learned controller completing the horizon
Advantage Aggregation → solves → X-point target
confidence 90% · We formulate XPT feedback as a multi-objective reinforcement learning (RL) control problem... we develop Advantage Aggregation (AdvA)
EXL-50U → uses → X-point target
confidence 90% · XPT equilibria have now been realized in multiple EXL-50U discharges
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Managing divertor heat loads is a central challenge for compact, high-power tokamaks. To increase local flux expansion and decouple the dissipation volume from the core, EHL-2 adopts the X-point target (XPT) divertor. This requires the secondary X-point to remain on the divertor leg; displacement degrades the topology and exhaust geometry. Current experiments, including EXL-50U discharges, rely on precomputed feedforward waveforms with PID loops on global quantities. Lacking dedicated closed-loop feedback for the secondary null, XPT operation is repeatable but not routine. We formulate XPT feedback as a multi-objective reinforcement learning (RL) control problem in a free-boundary environment calibrated to EXL-50U discharge #13906. To address strong coupling among plasma current, shape, and null constraints - where reward scalarisation collapses objective-specific temporal credit - we develop Advantage Aggregation (AdvA). AdvA preserves objective-wise temporal credit before worst-objective-aware nonlinear scalarisation and introduces a residual correction to policy updates. AdvA-PPO is evaluated against Reward-PPO and a feedforward-plus-PID baseline under nominal operation, measurement uncertainties, and unseen initial equilibria. On a 500 ms rollout, AdvA-PPO raises the mean worst-channel score from 0.23 to 0.81 over Reward-PPO, reducing X-point flux RMSE by ~20x. Under combined measurement uncertainties, it is the only learned controller completing the horizon while retaining a usable XPT shape. Multi-initialization fine-tuning enables a single AdvA-PPO policy to complete full-horizon operation across divertor and limiter initial equilibria. These results provide a simulation-based foundation for future real-time XPT validation on EXL-50U.
Tags
Links
- Source: https://arxiv.org/abs/2608.20834v1
- Canonical: https://arxiv.org/abs/2608.20834v1
Trouble viewing inline? Open PDF directly →
Full Text
72,296 characters extracted from source content.
Expand or collapse full text
Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment Siqi Ding Affiliation: Beijing ENN Fusion Energy Science and Technology Co., Ltd., Beijing 101111, China Affiliation: Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy, Beijing 101111, China Affiliation: Hebei Key Laboratory of Compact Fusion, Langfang 065000, China Xuanhe Wang Affiliation: Beijing ENN Fusion Energy Science and Technology Co., Ltd., Beijing 101111, China Affiliation: Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy, Beijing 101111, China Affiliation: Hebei Key Laboratory of Compact Fusion, Langfang 065000, China Pei Guo Affiliation: Key Laboratory of Materials Modification by Beams of the Ministry of Education, School of Physics, Dalian University of Technology, Dalian 116024, China Guoyang Shi Affiliation: Beijing ENN Fusion Energy Science and Technology Co., Ltd., Beijing 101111, China Affiliation: Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy, Beijing 101111, China Affiliation: Hebei Key Laboratory of Compact Fusion, Langfang 065000, China Changquan Yu Affiliation: School of Nuclear Science and Engineering, East China University of Technology, Nanchang 330013, Jiangxi, China Yiting Wang Affiliation: School of Mechanics and Engineering Science, Peking University, Beijing 100084, China Xianming Song Affiliation: Beijing ENN Fusion Energy Science and Technology Co., Ltd., Beijing 101111, China Affiliation: Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy, Beijing 101111, China Affiliation: Hebei Key Laboratory of Compact Fusion, Langfang 065000, China Xiang Gu Affiliation: Beijing ENN Fusion Energy Science and Technology Co., Ltd., Beijing 101111, China Affiliation: Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy, Beijing 101111, China Affiliation: Hebei Key Laboratory of Compact Fusion, Langfang 065000, China Zhengyuan Chen Affiliation: Beijing ENN Fusion Energy Science and Technology Co., Ltd., Beijing 101111, China Affiliation: Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy, Beijing 101111, China Affiliation: Hebei Key Laboratory of Compact Fusion, Langfang 065000, China Lei Xing Affiliation: Beijing ENN Fusion Energy Science and Technology Co., Ltd., Beijing 101111, China Affiliation: Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy, Beijing 101111, China Affiliation: Hebei Key Laboratory of Compact Fusion, Langfang 065000, China Yapeng Zhang Affiliation: Beijing ENN Fusion Energy Science and Technology Co., Ltd., Beijing 101111, China Affiliation: Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy, Beijing 101111, China Affiliation: Hebei Key Laboratory of Compact Fusion, Langfang 065000, China Jianguo Chen Affiliation: Beijing ENN Fusion Energy Science and Technology Co., Ltd., Beijing 101111, China Affiliation: Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy, Beijing 101111, China Affiliation: Hebei Key Laboratory of Compact Fusion, Langfang 065000, China Tianyuan Liu Affiliation: Beijing ENN Fusion Energy Science and Technology Co., Ltd., Beijing 101111, China Affiliation: Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy, Beijing 101111, China Affiliation: Hebei Key Laboratory of Compact Fusion, Langfang 065000, China Abstract Managing divertor heat loads is a central challenge for compact, high-power tokamaks. To increase local flux expansion and create a dissipation volume magnetically decoupled from the core, the compact spherical-torus design EHL-2 adopts the X-point target (XPT) as its reference divertor configuration. Realising these benefits requires the secondary X-point to remain on the divertor leg; its displacement can degrade the intended topology and exhaust geometry. Existing experiments, including recent EXL-50U discharges, still rely mainly on precomputed feedforward waveforms with proportional–integral–derivative (PID) loops on global quantities. The secondary null is not yet under dedicated closed-loop feedback, so XPT operation remains repeatable but not yet routine. We formulate reconstructed-null XPT feedback as a multi-objective reinforcement learning (RL) control problem in a free-boundary environment calibrated to EXL-50U discharge #13906. To address the strong coupling among plasma current, boundary shape, and multiple null constraints—where conventional reward-level scalarisation collapses objective-specific temporal credit into a single advantage—we develop Advantage Aggregation (AdvA), which preserves objective-wise temporal credit assignment before worst-objective-aware nonlinear scalarisation and introduces a controlled residual correction to the policy update. Instantiated with proximal policy optimisation (PPO), AdvA-PPO is evaluated against conventional reward-level PPO and the experiment-derived feedforward-plus-PID baseline under nominal operation, measurement uncertainties, and unseen initial equilibria. On the nominal 500ms500\,ms rollout, AdvA-PPO raises the mean worst-channel score from 0.230.23 to 0.810.81 relative to Reward-PPO and reduces mean X-point flux root-mean-square error (RMSE) by about 20×20×. Under combined measurement uncertainties, it is the only learned controller that completes the horizon while retaining a usable XPT shape. Multi-initialization fine-tuning further enables a single AdvA-PPO policy to complete full-horizon operation across both divertor and limiter initial equilibria. These results provide a simulation-based foundation for future real-time XPT validation on EXL-50U. * Corresponding authors: Tianyuan Liu (liutianyuan@enn.cn). Keywords: X-point target divertor; magnetic configuration control; artificial intelligence for fusion; multi-objective reinforcement learning; tokamak 1 Introduction Magnetic-confinement fusion offers a route to abundant, low-carbon energy, and the tokamak is among its most developed concepts. A reactor-scale tokamak, however, must exhaust intense particle and power fluxes without exceeding the steady-state limits of plasma-facing materials. Extrapolations for compact, high-power devices illustrate the severity of this challenge [9]. Advanced divertor configurations (ADCs) address it by reshaping the scrape-off-layer magnetic geometry to increase connection length, flux expansion, wetted area, and the volume available for power and momentum dissipation. Their exhaust benefit therefore depends not only on edge-plasma physics, but also on the ability to establish and maintain the intended magnetic topology. Experiments have realized several ADC families, including the snowflake (SF), quasi-snowflake (QSF), Super-X, X-divertor (XD), and X-point target (XPT). Their exhaust benefits are well established: dedicated SF control on DIII-D sustained a ∼2.5× 2.5× peak-heat-flux reduction for 22–33 s, NSTX SF experiments reduced the peak heat flux between edge-localised modes (ELMs) by 33–5×5×, and the three-X-point TCV Jellyfish reduced the peak parallel target heat flux by at least 50%50\% relative to a single null [8, 22, 4]. Control maturity, however, does not follow directly from exhaust performance. As the number of nulls increases, plasma current, the main boundary, divertor legs, and strike points must share a limited set of poloidal-field (PF) actuators, so improving one quantity can degrade another. Among these ADCs, the XPT is the configuration most directly relevant to the EHL-2 divertor route. It places a secondary magnetic null in the divertor leg, away from the confined plasma, thereby combining a long leg with strong local flux expansion and a dissipation volume that is magnetically decoupled from the core [10, 26]. Experiments on TCV and MAST-U show that this topology improves divertor conditions, sustains an X-point radiator regime, and passively rejects disturbances near the secondary null [17, 11, 12, 28]. These exhaust properties are conditional on the magnetic topology: the secondary null must remain on the divertor leg of the primary separatrix to preserve the intended flux expansion and magnetically decoupled dissipation volume. Secondary-null displacement can therefore degrade the functional geometry even when plasma current and global position remain acceptable. XPT control is consequently not an ancillary shape-control task, but an enabling requirement for converting magnetic access to the configuration into repeatable exhaust capability. For EHL-2 the XPT is not one option among several but the main reference divertor configuration: integrated poloidal-field-system and free-boundary equilibrium design shows that an outer XPT is magnetically accessible within the coil set, and edge-transport calculations predict the associated heat-load and detachment behaviour [5, 27]. Because the EHL-2 evidence is design and simulation based, it defines both the target topology and the control requirement that motivate the present study. Motivated by this route, XPT equilibria have now been realized in multiple EXL-50U discharges. Figure 1 shows a representative reconstruction from discharge #13906. These experiments establish that the target topology is magnetically accessible on the device, but access in a reconstructed equilibrium is not equivalent to regulating the performance-critical secondary null throughout a discharge. The experiments still rely primarily on a classical feedforward-plus-PID scheme (F+PID): precomputed PF-coil feedforward waveforms supplemented by PID feedback on global quantities. Reaching the desired topology still requires shot-to-shot tuning, and the secondary X-point is not yet maintained by a dedicated real-time feedback loop. XPT operation on EXL-50U should therefore be regarded as repeatably demonstrated but not yet routine. A similar commissioning burden was reported on MAST-U, where the first steady XPT required three successive discharges to tune the feedforward virtual-circuit waveform [2]. The EXL-50U experiments provide physically realized target equilibria and a basis for calibrating the control environment, rather than evidence of closed-loop secondary-null control. Figure 1: Representative experimental XPT equilibrium in EXL-50U discharge #13906 at t=500mst=500\,ms. The reconstructed magnetic topology demonstrates experimental access to XPT operation. This limitation is shared by the small number of XPT experiments worldwide, as summarized in Table 1. On MAST-U the XPT could only be obtained transiently in the 2021 campaign under pure feedforward operation; steady-state XPT became possible once boundary variables were placed under feedback, but the real-time reconstruction did not estimate the secondary X-point. That null was instead created by feedforward virtual-circuit commands driving the local radial and vertical fields toward zero at a user-specified location, which left a finite and slowly growing field offset and a ∼ 10 cm discrepancy in the achieved null position, and required three successive discharges to converge [2]. On TCV, published XPT campaigns quantify exhaust performance, the X-point radiator regime, and detachment dynamics under the device’s divertor shape control; they do not report a feedback loop that treats the secondary null as an explicit real-time control variable together with the main boundary [24, 17, 11, 28]. Offline reconstruction can recover multi-null topologies after a shot, and real-time reconstructors already support selected boundary and primary-X quantities on several devices [2, 20]. For XPT, the enabling control requirement is therefore not reconstruction in isolation, but closing the loop on the secondary null: keeping that null on the divertor leg of the primary separatrix so that the intended flux expansion and magnetically decoupled dissipation volume survive disturbances and shot-to-shot drift. Absent such feedback, secondary-null placement remains a feedforward commissioning task even when IpI_p and the global boundary track well—as illustrated by the MAST-U residual field offset and multi-discharge tuning [2]. Table 1: Experimental XPT magnetic control to date. “Real-time secondary null in FB” asks whether a real-time estimate of the secondary null enters the feedback law (as opposed to feedforward coil programming alone). Device Method FB targets Real-time secondary null in FB TCV [24, 17, 11, 28] Divertor shape control IpI_p, boundary / shape No MAST-U [2, 12] LEMUR + F Br,z→0B_r,z\!→\!0 ROUTR_OUT, RINR_IN, ZXZ_X No EXL-50U (this exp.) F+PID IpI_p, R, Z No Regulating multi-null divertor geometry therefore requires one of two interfaces. The first closes feedback on explicit null states—null positions and related flux constraints obtained from a real-time equilibrium reconstruction or a dedicated null observer—together with IpI_p and the last closed flux surface (LCFS). The second is reconstruction-free: a controller maps magnetic probes, flux loops, and coil currents directly to actuator commands, so that multi-null geometry is regulated without feeding a real-time equilibrium into the loop. Deep RL on TCV has demonstrated the latter for snowflake and Jellyfish targets, including a snowflake transition that reduced X-point separation from 3434 to 6.66.6 cm and the three-null Jellyfish configuration, with X-point structure latent in the learned policy rather than exposed as reconstructed feedback channels [3, 25, 4]. Related reconstruction-free magnetic control has also been shown on DIII-D and WEST [23, 7]. Those results establish that multi-null feedback need not be reconstruction-based; they do not, however, close an XPT loop in which a secondary null on the divertor leg is an explicit real-time feedback variable. On EXL-50U the reconstruction-free magnetics interface for multi-null divertor control has not yet been validated, whereas the planned plant stack routes equilibrium features from the real-time reconstruction code PTEFIT to the controller [30, 20]. We therefore first develop the explicit, reconstruction-based route: closed-loop XPT control on reconstructed secondary- null position and flux together with IpI_p and the LCFS—an interface that, to our knowledge, has not been demonstrated in XPT experiments to date. We do not claim superiority of either interface; the present choice matches the near-term EXL-50U control architecture and makes the secondary-null channels auditable for multi-objective learning. Model-predictive control provides a non-learning alternative for shape regulation [13], but a validated control-oriented model of the coupled EXL-50U XPT nulls is not presently available. XPT control is multi-objective on a single trajectory: plasma current, the LCFS, and multiple null position/flux channels must remain jointly acceptable under shared PF actuators, while the identity of the worst channel can change over the discharge. Fusion RL typically scalarizes these channels in the reward before estimating one advantage [3, 25]. In robotics and multi-task learning, related conflicts are instead attacked inside the learner—notably by projected conflicting gradients (PCGrad) [29, 16] and by preference-conditioned policies with multi-head advantage decomposition [1]. Those methods mainly mitigate gradient interference or support preference trade-offs across tasks. Reconstructed-null XPT control is harder-coupled: a weak null/flux channel can fail the topology even when the scalar return looks acceptable, and tokamak magnetic RL still rarely keeps objective structure through credit assignment. The relevant design choice is therefore where nonlinear scalarization occurs relative to objective-wise temporal credit. The gap addressed here is therefore twofold. At the control-system level, existing XPT operation does not close feedback on reconstructed secondary-null position and flux together with the plasma current and boundary. At the algorithmic level, conventional reward-level scalarization discards objective identity before temporal credit assignment. We develop AdvA along the second axis: a shared critic backbone with one value head and one generalised advantage estimation (GAE) [18] per channel, with worst-objective-aware softmin aggregation (implementation name SmoothMax; α<0α<0) applied only after advantage estimation, augmented by a controlled component residual on the aggregated update. On EXL-50U, PTEFIT already provides real-time equilibrium reconstruction for boundary quantities at the centimetre level [30, 20]; whether its secondary-null position and flux estimates are accurate enough for closed-loop XPT feedback remains to be established experimentally. The present study therefore validates the control interface and AdvA in an experiment-calibrated free-boundary simulation, as a precursor to PTEFIT-in-the-loop tests. The contributions address these gaps as follows: 1. We establish an AI-enabled multi-objective framework for closed-loop XPT magnetic configuration control in an experiment-calibrated EXL-50U free-boundary simulation environment. The framework jointly regulates plasma current, LCFS geometry, X-point positions, and X-point flux constraints through shared poloidal-field actuators. 2. We develop Advantage Aggregation (AdvA) for coupled multi-objective magnetic control. By retaining objective-specific value heads and generalised advantage estimates before nonlinear scalarisation, together with a controlled residual correction, AdvA preserves objective-wise temporal credit assignment and supports coordinated policy optimisation. 3. We evaluate the framework against conventional reward-level RL and the experiment-derived F+PID baseline across nominal operation, measurement uncertainties, and unseen initial equilibria, together with multi-initial adaptation. The results characterise its control capability, robustness, transferability, and remaining challenges for simulation-based XPT magnetic control. The present evidence is simulation-based; closed-loop XPT control on EXL-50U with PTEFIT-supplied null states remains future experimental work. 2 Methods 2.1 EXL-50U XPT control problem Figure 2 summarises the EXL-50U control stack. Twelve active coil circuits—the central solenoid (CS), ten Poloidal Field coils (PF1–PF10), and the vertical-stability coil (VS)—regulate the plasma current and magnetic configuration. Learned policies command CS and PS1–PS10 (11 voltages); VS remains under a separate PID loop due to highe frequency for every controller reported here. The target XPT has two primary nulls near the confined boundary and two secondary nulls in the divertor legs. We use one-based labels X1X_1–X4X_4: X2X_2, X3X_3 are the upper/lower inner primary pair and X1X_1, X4X_4 the upper/lower outer secondary pair. The plant is strongly coupled—each coil simultaneously affects IpI_p, the LCFS, and all four nulls—so improving one X-point condition can displace another null or deform the boundary. Training and evaluation close the loop through the calibrated fast free-boundary Grad–Shafranov evolutive solver (FGE) [6] (Sec. 2.2.2): FGE supplies IpI_p, coil currents, LCFS, and X-point features; the policy returns the 11 coil voltages; the environment advances and scores the next step. The same observation and command interfaces are already wired for machine deployment (Figure 2): diagnostic and coil signals pass through reflective memory (RFM) to real-time equilibrium reconstruction, the policy runs under TensorRT, and coil commands return to the power supplies. With that end-to-end path in place, a controller validated in FGE is ready for on-machine closed-loop tests; the experiments below remain simulation-based and do not yet report PTEFIT-in-the-loop XPT feedback on EXL-50U. Figure 2: End-to-end architecture for reconstructed-null XPT control. Simulation closes the loop through FGE; deployment swaps FGE observations for real-time reconstruction while keeping the same state and coil-command interfaces (11 policy coils; VS under PID). 2.2 FGE simulation environment 2.2.1 Fast free-boundary Grad–Shafranov evolutive solver All controllers interact with FGE [6], a control-oriented free-boundary equilibrium evolution solver in the Matlab EQuilibrium suite (MEQ), previously used for RL magnetic control on TCV [3]. Given PF-coil voltages, FGE advances the axisymmetric equilibrium on resistive timescales by coupling three blocks. The poloidal flux satisfies the Grad–Shafranov equation Δ∗ψ=−2πRμ0jϕ=−4π2(μ0R2p′(ψ)+TT′(ψ)), ^*ψ=-2π R _0j_φ=-4π^2 ( _0R^2p (ψ)+T (ψ) ), (1) with free profiles p′(ψ)p (ψ) and TT′(ψ)T (ψ) fixed by scalar constraints (here βp _p and q0q_0). Active and passive conductor currents IeI_e obey the circuit equation (Va0)=MeeI˙e+MeyI˙y+ReIe, pmatrixV_a\\ 0 pmatrix=M_e I_e+M_ey I_y+R_eI_e, (2) where VaV_a are the applied coil voltages, IyI_y is the plasma current distribution, and MeeM_e, MeyM_ey, ReR_e are the mutual-inductance and resistance matrices. The bulk plasma current is closed by a rigid Ohmic model (OhmTor-rigid in [6]), 0=LpI˙p+MpeI˙e+Rp(Ip−Ini),0=L_p I_p+M_pe I_e+R_p (I_p-I_ni ), (3) with plasma self-inductance LpL_p, plasma–conductor mutual MpeM_pe, bulk resistance RpR_p, and non-inductive current IniI_ni. Each step returns IpI_p, coil currents, the LCFS, X-point positions, and fluxes for observations and rewards. On EXL-50U we use a 66×6566× 65 (R,Z)(R,Z) grid, 544 passive filaments plus 12 active circuits, and a fixed 1ms1\,ms step; βp _p and RpR_p are held at discharge-calibrated values (Sec. 2.2.2). The model is restricted to the magnetic degrees of freedom needed for XPT topology regulation and is fast enough for parallel policy training. 2.2.2 Calibration and validation on #13906 The environment is calibrated to EXL-50U discharge #13906, which sustained an XPT during the IpI_p flat top near 400400–700ms700\,ms. Reference equilibria come from the LIUQE tokamak equilibrium reconstruction code [15]; we use the 500500–700ms700\,ms window. Shot-specific inputs are the reconstructed βp(t) _p(t) and q0(t)q_0(t) as profile constraints in Eq. (1), and a constant bulk plasma resistance RpR_p chosen so that simulated IpI_p tracks the experiment over that window. Coil and vessel geometry, inductances, resistances, and power-supply parameters follow engineering calibrations (including vacuum shots) and are held fixed. For validation, FGE starts from the LIUQE equilibrium at t=0.5st=0.5\,s and advances to t=0.7st=0.7\,s under the recorded coil voltages. Figure 3 overlays boundaries and nulls; Figure 4 compares IpI_p, R, RmaxR_ , and κ; Table 2 summarises bias, RMSE, and maxima. The XPT topology (all four nulls) is preserved; IpI_p RMSE is 0.8%0.8\% (1.6%1.6\% at worst), and LCFS/X-point offsets stay at the centimetre scale of the grid cell (≈1.7cm≈ 1.7\,cm), which matches the tolerances of the configuration-control task. Residual mismatches concentrate in the divertor legs and in small biases of κ and RmaxR_ , consistent with the reduced profile model. Figure 3: Free-boundary equilibria from the FGE validation run of EXL-50U discharge #13906 at t=0.50t=0.50, 0.550.55, 0.630.63, and 0.69s0.69\,s (left to right). Each panel overlays the simulated separatrix (solid blue) and X points (blue crosses) on the experimental LIUQE-reconstructed separatrix (dashed black) and X points (open circles); thin gray curves show simulated flux surfaces and the dark gray contour marks the limiter. Figure 4: Time traces from the FGE validation run of EXL-50U discharge #13906 (solid blue) compared with the experimental LIUQE reconstruction (dashed black): (a) plasma current IpI_p; (b) plasma radial position R; (c) outer radial extent RmaxR_ of the LCFS; (d) LCFS elongation κ. Table 2: Deviations between the FGE validation run and the LIUQE reconstruction of discharge #13906 over 0.50.5–0.7s0.7\,s: time-averaged signed deviation (bias, simulation minus reconstruction), root-mean-square error (RMSE), and maximum absolute deviation. The LCFS row is based on the symmetrized mean distance between the simulated and reconstructed boundary contours, and the X-point rows on the Euclidean position offsets averaged over the upper and lower X point of each type; these distances are nonnegative by construction, so no bias is reported. Quantity Bias RMSE Maximum IpI_p (kA) +1.4+1.4 3.1 6.5 R (m) +4.7+4.7 5.3 11.4 RmaxR_ (m) +7.6+7.6 8.7 16.7 κ +0.045+0.045 0.051 0.101 LCFS distance (m) — 13.8 30.5 Primary X points X2X_2, X3X_3 (m) — 18.3 38.0 Secondary X points X1X_1, X4X_4 (m) — 13.5 21.6 2.3 Control methods 2.3.1 Baseline: F+PID Figure 5: Baseline F+PID control structure. The classical baseline is feedforward plus PID (Figure 5). Offline, the MATLAB Shape Editor (SE) designs PF-coil current and IpI_p references and the associated feedforward voltages [21]. Online, PID loops on PF-coil currents, IpI_p, and plasma position (R,Z)(R,Z) produce a voltage correction that is added to the feedforward, cmd=f+ΔPID. V_cmd= V_f+ V_PID. (4) All F+PID scores in this paper reuse the #13906 plant feedforward and PID gains in a closed-loop FGE rollout, without per-initial redesign of the feedforward table. 2.3.2 RL control We cast XPT magnetic control as a partially observed decision process: the full FGE equilibrium is the latent state, but the feedforward policy π(t∣t)π(a_t _t) acts only on a reconstructed observation to_t and is trained to maximise the discounted return π[∑tγtrt]E_π[ _tγ^tr_t]. The interface used throughout is as follows. Observation: The 45-dimensional observation to_t concatenates normalised plasma-current error, 12 coil currents (CS, PS1–PS10, and VS), 4 X-point features (validity, R–Z position errors, and flux relative to the LCFS), and radial and vertical LCFS errors at eight boundary points equally spaced in poloidal angle about the magnetic axis. The same observation is used for every learned controller in this paper. Action: The learned policy commands an 11-dimensional coil-voltage vector (CS and PS1–PS10). The vertical-stability coil (VS) is regulated by a separate PID loop and is excluded from the policy action for every controller in this paper. Policy outputs are passed through tanh squashing and rescaled to the physical voltage limits. Actuation metrics |ΔV|| V| are computed on these 11 policy-commanded coils. Reward: Tracking errors for the plasma current, the four X-point positions, the four X-point flux-difference conditions, and the LCFS boundary are extracted from the FGE output and mapped to satisfaction signals rt,i∈[0,1]r_t,i∈[0,1] by sigmoid or softplus shaping with “good”/“bad” thresholds. How these K signals are aggregated relative to temporal credit assignment is the algorithmic axis of the next subsections. 2.3.3 Actor–critic instantiation with PPO Advantage Aggregation (AdvA) needs a per-objective advantage signal, so it belongs to the actor–critic family: methods that maintain a critic (or critics) and form advantages for a policy update. Pure value-based algorithms without an explicit advantage pathway, such as the deep Q-network (DQN) [14], are outside this scope. We instantiate both the reward-level baseline and AdvA with PPO [19]; the same aggregation design can be attached to other actor–critic algorithms. In results we write AdvA-PPO for the AdvA instance on PPO. PPO stabilises policy-gradient updates through a clipped surrogate. With probability ratio χt(θ)=πθ(t∣t)/πθold(t∣t) _t(θ)= _θ(a_t _t)/ _ _old(a_t _t) and advantage A^t A_t, ℒCLIP(θ)=−t[min(χt(θ)A^t,clip(χt(θ),1−ϵ,1+ϵ)A^t)],L^CLIP(θ)=-E_t\! [ \! ( _t(θ)\, A_t,\;clip\! ( _t(θ),1-ε,1+ε )\, A_t ) ], (5) the clip removes the incentive for large probability-ratio steps when they would inflate the objective. The full training loss adds a value-function error and an entropy bonus, ℒPPO(θ)=ℒCLIP(θ)+c1ℒVF(θ)−c2ℋ[πθ],L^PPO(θ)=L^CLIP(θ)+c_1\,L^VF(θ)-c_2\,H[ _θ], (6) where ℒVFL^VF fits the critic to empirical returns and ℋH encourages exploration. Both the reward-level baseline and every AdvA variant share the 4545-dimensional observation and 1111-dimensional action of Sec. 2.3.2. 2.3.4 Reward-level SmoothMax baseline The conventional multi-objective baseline scalarises the channel rewards before GAE [18]: that is why we call it reward-level. The same worst-objective-aware exponential weighting used later by AdvA defines ωt,i _t,i =wi0exp(αrt,i)∑jwj0exp(αrt,j), = w_i^0 (α r_t,i) _jw_j^0 (α r_t,j), (7) u(t)=r¯t u(r_t)= r_t =∑iωt,irt,i,α<0, = _i _t,ir_t,i, α<0, (8) so a poorly satisfied channel receives more weight. We keep the implementation name “SmoothMax”; with α<0α<0 the softmin emphasises the worst objective. A single value function and GAE fitted to r¯t r_t then supply the PPO advantage in Eq. (5). Results and ablations refer to this baseline as Reward-PPO (Table 3). 2.3.5 Advantage Aggregation (AdvA) AdvA moves scalarisation to the advantage layer: one value head and one GAE per physical channel on a shared critic backbone, with nonlinear mixing only afterwards. Relative to reward-level SmoothMax, a weak channel no longer collapses the temporal credit of every other channel before advantage estimation. Figure 6 shows the closed loop and update path. Figure 6: Policy update and environment interaction for AdvA-PPO. The FGE environment provides the reconstructed magnetic observation and component rewards. Per-objective value heads and GAE preserve temporal credit assignment before SmoothMax aggregation; the PPO anchor is then corrected by the norm-capped residual of Eqs. (15)–(16). Reliability gating and gradient projection are ablation options and are off in AdvA-PPO. Let rt,i∈[0,1]r_t,i∈[0,1] be the channel reward and wi0w_i^0 a fixed prior weight (inner X-point positions and the LCFS receive larger wi0w_i^0). A shared critic backbone feeds heads Vi(t)V_i(o_t) with δt,i _t,i =rt,i+γVi(t+1)−Vi(t), =r_t,i+γ V_i(o_t+1)-V_i(o_t), (9) A^t,i A_t,i =∑ℓ=0T−t−1(γλ)ℓδt+ℓ,i,i=1,…,K. = _ =0^T-t-1(γλ) _t+ ,i, i=1,…,K. (10) Each head is trained against G^t,i=A^t,i+Vi(t) G_t,i= A_t,i+V_i(o_t). SmoothMax advantage aggregation. Batch-mean channel rewards r¯i r_i define worst-case-aware weights wi=wi0exp(αr¯i)∑jwj0exp(αr¯j),α<0,w_i= w_i^0 (α r_i) _jw_j^0 (α r_j), α<0, (11) and the scalar advantage A^t=∑i=1KwiA^t,i A_t= _i=1^Kw_i\, A_t,i (12) replaces the usual PPO advantage column. Setting α=0α=0 recovers a fixed-weight mean (Adv-mean in the ablation). The actor uses a single clipped surrogate on A^t A_t. Reliability gate. A head with low explained variance, or a channel already saturated near reward 11, contributes little useful gradient. Exponential moving averages of explained variance EV¯i EV_i and saturation fraction s¯i s_i define qi=clip(EV¯i,0,1)(1−s¯i)∈[0,1],q_i=clip ( EV_i,0,1 )(1- s_i)∈[0,1], (13) and when enabled multiply the aggregation weights, wi←wiqiw_i← w_iq_i, followed by renormalisation. The gate is an ablation option and is off in AdvA-PPO. Norm-capped residual. PPO clipping acts once on the aggregated advantage, so a minority channel whose A^t,i A_t,i disagrees with A^t A_t can be silenced. Component surrogates ℒiclip(θ)=−t[min(χt(θ)A^t,i,clip(χt(θ),1−ϵ,1+ϵ)A^t,i)]L_i^clip(θ)=-E_t [ ( _t(θ) A_t,i,clip ( _t(θ),1-ε,1+ε ) A_t,i ) ] (14) supply gradients 0=∇θℒCLIP(A^)g_0= _θL^CLIP( A) and i=∇θℒiclipg_i= _θL_i^clip. Each component is capped to the anchor scale, ~i=imin(1,κ∥0∥2∥i∥2+εg), g_i=g_i (1, κ _0 _2 _i _2+ _g ), (15) and the update is =0+β∑i=1K~i,g=g_0+β _i=1^K g_i, (16) with defaults κ=1κ=1 and β=0.2β=0.2. Setting β=0β=0 recovers pure advantage-level SmoothMax. Gradient projection. As an optional conflict-resolution step in the spirit of PCGrad [29, 16], the component of ~i g_i opposing 0g_0 can be removed before the sum in Eq. (16). Projection appears only in proxy ablation rows of Sec. 3.2 and is disabled in AdvA-PPO (i=~ip_i= g_i). Proposed controller: AdvA-PPO. AdvA-PPO is the AdvA instance used in all main experiments: advantage-level SmoothMax (α=−3α=-3) plus the capped residual (β=0.2β=0.2, κ=1κ=1), with the reliability gate and projection both off. Sec. 3.2 motivates retaining the residual and omitting gate and projection. 2.3.6 Ablation sequence and controller inventory Table 3 lists the controllers compared in the Results ablation (Sec. 3.2). Results name the reward-level SmoothMax PPO baseline of Sec. 2.3.4 as Reward-PPO. The ablation follows the design decisions in causal order: 1. Start from Reward-PPO (reward-level SmoothMax, scalar GAE). 2. Move to Adv-mean (α=0α=0, multi-head). 3. Replace the mean by Adv-SmoothMax (α=−3α=-3, multi-head). 4. On that SmoothMax anchor, compare ++\;gate (β=0β=0) with ++\;resid (β=0.2β=0.2) == AdvA-PPO. Pairwise gate/projection combinations are reported only as temporary proxies from available checkpoints, not as a second method family. Early multi-head runs still carry two radial-extrema rewards (RminR_ , RmaxR_ ), hence K=12K=12; those channels are excluded from every reported score. Clean contrasts in Sec. 3.2 are therefore Adv-mean versus Adv-SmoothMax (K=12K=12), gate versus residual on the SmoothMax anchor (K=10K=10), and the family-level Reward-PPO versus AdvA-PPO comparison; proxy rows are not matched one-variable AdvA-native cells. Within a channel set, matched variants share the 45/1145/11 observation–action interface, capacity, optimizer, and checkpoint rule; F+PID is evaluated, not retrained. Table 3: Controllers in the component ablation, in the same order as the Results ablation (Sec. 3.2). “Locus” is where nonlinear aggregation occurs relative to objective-wise temporal credit. For AdvA rows, Gate/β/Proj. act on advantage weights or the residual of Eqs. (15)–(16). All rows share the same da=11d_a=11 / do=45d_o=45 interface (Sec. 2.3.2; VS under PID). Reward-PPO is the reward-level SmoothMax baseline of Sec. 2.3.4. Name Locus α Gate β Proj. K da/dod_a/d_o Reward-PPO reward −3-3 off 0 off 10 11/4511/45 Adv-mean advantage 00 off 0 off 10 11/4511/45 Adv-SmoothMax advantage −3-3 off 0 off 10 11/4511/45 ++ gate advantage −3-3 on 0 off 10 11/4511/45 ++ resid (AdvA-PPO) advantage −3-3 off 0.20.2 off 10 11/4511/45 ++ proj advantage −3-3 off 00 on 10 11/4511/45 ++ gate++resid advantage −3-3 on 0.20.2 off 10 11/4511/45 ++ gate++proj advantage −3-3 on 00 on 10 11/4511/45 Training and inference implementation. The actor and shared value backbone are two fully connected layers of 256 units with tanh activations; AdvA controllers attach one scalar value head per reward channel. Policies use a squashed Gaussian action distribution and are trained in PyTorch with Adam at learning rate 3×10−43× 10^-4. The PPO clip is ϵ=0.2ε=0.2, γ=0.98γ=0.98, λ=0.95λ=0.95, the entropy coefficient is 5×10−35× 10^-3, and global gradient norm is clipped at 0.3. Twenty-three parallel FGE workers collect one 300-step episode each per iteration, giving a batch of 6900 transitions; each batch is optimised for ten epochs with minibatches of 256. The reported AdvA-PPO checkpoint is the final checkpoint after 300 iterations (2.07×1062.07× 10^6 environment steps). Evaluation uses the deterministic mean action. 2.4 Evaluation setup All reported scores come from one harness, one FGE image pool, and one primary evaluation seed (2026071520260715), with a 500500-step (500ms500\,ms) closed-loop test horizon and deterministic actions at test time. Training episodes are 300300 steps (300ms300\,ms); finishing the longer test window is a temporal-extrapolation check. 2.4.1 Test scenarios Four initial equilibria (I1–I4; Table 4) share the training inductance basis and request the same frozen #13906 XPT target. Each control window starts near t0≈500mst_0≈ 500\,ms of the parent discharge and runs for 500ms500\,ms. I1 is the calibrated XPT training initial; I2–I4 are withheld cross-initialization cells (same geometric target and reward operators; only the FGE initial and IpI_p reference change). Robustness cells R1–R3 reuse I1. Disturbances enter the controller observation only; the simulator state used for scoring stays clean. R1 redraws an observation delay uniformly in 11–3ms3\,ms at every control step (non-monotonic jitter, harder than a fixed mean delay). R2 adds full diagnostic/coil noise: IpI_p σ=6kAσ=6\,kA, per-coil IPFI_PF at 1%1\%, and absolute 5m5\,m on raw X-point and LCFS coordinates. R3 applies R1 and R2 together. Table 4: Evaluation initials I1–I4. All windows start at t0≈500mst_0≈ 500\,ms (CSV start 0.502s0.502\,s) and use the frozen #13906 XPT geometric target; Ip⋆I_p is the current reference for that cell. ID Shot t0t_0 Topology Ip⋆I_p I1 #13906 ≈500ms≈ 500\,ms XPT (training / test) ≈400kA≈ 400\,kA I2 #13705 ≈500ms≈ 500\,ms divertor (test) ≈500kA≈ 500\,kA I3 #13844 ≈500ms≈ 500\,ms divertor (test) ≈500kA≈ 500\,kA I4 #15892 ≈500ms≈ 500\,ms limiter (test) ≈500kA≈ 500\,kA 2.4.2 Metrics and statistical protocol Scores are recomputed from archived equilibria by one pipeline. Raw FGE X-point candidates are mapped to the four slots X1X_1–X4X_4 of Sec. 2.1 before any error is formed. Position error is the Euclidean distance to the frozen target in that slot; flux error is the poloidal-flux difference |ψX−ψB|| _X- _B| in that slot, reported in Wb. Mean dXd_X or mean flux is the average of the four per-slot RMSE / max values (primary/secondary columns average only the corresponding pair). The LCFS channel is the boundary deviation used by the reward operator, eLCFS=12(|(rB−rBref)/(|rBref|+ϵ)|¯+|(zB−zBref)/(|zBref|+ϵ)|¯),ϵ=10−6m,e_LCFS= 12 ( (r_B-r_B^ref)/( r_B^ref +ε) + (z_B-z_B^ref)/( z_B^ref +ε) ), ε=10^-6\,m, (17) on eight boundary points equally spaced in poloidal angle about the magnetic axis. Tables report RMSE and max absolute error over the survived horizon, survival length (full =500ms=500\,ms), and |ΔV|| V| on the 11 policy coils. On short survivals, errors are secondary to survival; stochastic noise/jitter cells are multi-seed averaged or labelled as a single realisation. Because physical rows have incompatible units, we also report time means of the shared ten-channel scores (IpI_p, four positions, four fluxes, LCFS): the reward-level SmoothMax u¯ u of Eq. (8) (α=−3α=-3) and the hard worst-channel mean r¯=minirt,i¯ r= _ir_t,i. Both are defined for F+PID as well; evaluation u¯ u is not the advantage-level SmoothMax inside AdvA. Mechanism narrative emphasises physical RMSE / max; Sec. 3.2 also lists u¯ u and r¯ r. In comparison tables, the best entry in each column (within a block) is bold and the second-best is underlined; lower is better for errors and |ΔV|| V|, higher for survival, u¯ u, and r¯ r. 3 Results 3.1 Main controller comparison Under the protocol of Sec. 2.4, the nominal cell uses the calibrated shot-#13906 initial (training initial; I1 of Table 4) over a 500ms500\,ms rollout at 1ms1\,ms per control step. Three controllers are compared: F+PID, Reward-PPO, and AdvA-PPO (ours). The head-to-head isolates the aggregation locus (reward-level versus advantage-level); the component ablation of Sec. 3.2 isolates the algorithmic factors within AdvA. Figure 7: Closed-loop evolution on #13906 (500ms500\,ms). Grey: F+PID; blue: Reward-PPO; red: AdvA-PPO. (a) IpI_p; (b) LCFS error; (c)–(d) primary X distance and flux (X2X_2, X3X_3); (e)–(f) secondary X distance and flux (X1X_1, X4X_4). Figure 7 and Table 5 show that AdvA-PPO is the strongest multi-objective controller on this cell: it leads or matches on IpI_p, primary X distance, mean X flux, and LCFS, rather than trading one channel for another. All three controllers complete the horizon, so the comparison is tracking quality, not survival. The evolution traces make the channel-wise behaviour visible in time; panel (d) already shows the primary-flux gap that the table later quantifies as about a 20×20× lower mean X-flux RMSE for AdvA-PPO relative to Reward-PPO (0.660.66 versus 13.1×10−413.1× 10^-4). The decisive contrast is not a single physical row but the shared scoring axis of Sec. 2.4.2. Reward-PPO reaches u¯=0.56 u=0.56 with worst-channel mean r¯=0.23 r=0.23, limited by the inner primary flux X2X_2; AdvA-PPO reaches u¯=0.93 u=0.93 with r¯=0.81 r=0.81—a 3.5×3.5× lift of the hard worst channel, together with a 2.8×2.8× lower IpI_p RMSE (260260 versus 735A735\,A). Reward-PPO is not uniformly weaker: it records the lowest secondary X distance (19.019.0 versus 28.2m28.2\,m), but that gain coincides with the poorest IpI_p among the three and a much weaker primary-flux channel. The secondary-null rows deserve a separate reading, because position and flux are not interchangeable descriptors of an XPT. The defining topological condition is that the secondary null sits on the divertor leg of the primary separatrix, i.e. ψX→ψB _X\!→\! _B; where along that leg it sits sets the location of the flux expansion but does not decide whether the configuration is an XPT at all. On this cell AdvA-PPO holds the secondary flux condition an order of magnitude tighter than Reward-PPO (1.071.07 versus 17.1×10−4Wb17.1× 10^-4\,Wb on X1X_1 and 0.610.61 versus 8.8×10−4Wb8.8× 10^-4\,Wb on X4X_4) while allowing the nulls to sit about 10m10\,m further along the leg. In other words the two policies fail differently: AdvA-PPO keeps the nulls locked to the separatrix and slides them along it, whereas Reward-PPO places them closer to the programmed coordinates but lets them detach from the leg. For the XPT the first failure mode is the benign one, and the residual ∼10m \!10\,m offset is comparable to the secondary-null reconstruction accuracy of the calibration itself (Table 2, 13.5m13.5\,m RMSE), so it should not be read as a resolved ranking. The flux criterion is also intrinsically weak against displacement: ∇ψ=0∇ψ=0 at a null, so |ψX−ψB|| _X- _B| grows only quadratically with the null displacement and a small flux error certifies topology rather than position. The two channels are therefore complementary, and we report both rather than a single secondary-null figure of merit. F+PID keeps competitive IpI_p (454A454\,A RMSE) because its loops close on coil currents, IpI_p, and the geometric centroid (R, Z); it does not feed back X-point position, X-point flux, or a multi-point LCFS, so the shape rows in the table are open-loop observations and the evolution panels (b)–(d) sit systematically above either policy. The AdvA-PPO gains are not bought with larger actuation: its mean per-step |ΔV|| V| is the lowest of the three (0.36V0.36\,V against 0.56V0.56\,V and 0.45V0.45\,V). Figure 8: Flux snapshots on #13906 at t=500t=500–900ms900\,ms (shot clock; control from t0=500mst_0=500\,ms). Rows: F+PID, Reward-PPO, AdvA-PPO. Blue solid: current LCFS; black dashed: target; open circles: target X; blue crosses: actual X. Figure 8 supplies the geometric reading of the same episode. Under AdvA-PPO the separatrix stays on the target boundary and the X points remain close to their targets through the horizon, consistent with the low primary distance and flux errors. Reward-PPO keeps a four-X topology but with a looser core LCFS match, in line with its intermediate boundary and flux scores. Under F+PID the separatrix bulges outward and the X points drift, as expected from a controller without shape closure. Across all three methods the divertor-leg region remains only loosely constrained: neither the observation nor the reward treats the legs explicitly, so secondary-leg geometry is a shared limitation of the present task rather than a failure unique to one controller. Table 5: Nominal tracking on I1 (#13906, 500ms500\,ms); metrics as in Sec. 2.4.2. F+PID X-point and LCFS rows are open-loop observations. Quantity F+PID Reward-PPO AdvA-PPO (ours) IpI_p (A) 454 / 1488 735 / 1310 260 / 624 Primary X dist. (m) 39.0 / 58.3 22.8 / 39.2 6.3 / 9.5 Secondary X dist. (m) 30.1 / 38.3 19.0 / 31.8 28.2 / 44.2 Mean X flux (10−4Wb10^-4\,Wb) 77.38 / 117.02 13.06 / 18.89 0.66 / 1.93 LCFS (×10−3× 10^-3) 120.1 / 430.3 14.7 / 21.9 12.2 / 17.2 SmoothMax u¯ u ↑ 0.093 0.563 0.925 mean worst channel r¯ r ↑ 0.001 0.232 0.810 mean / max |ΔV|| V| (V) 0.45 / 124 0.56 / 50 0.36 / 81 3.2 Mechanism and ablation analysis Table 6: Component ablation on I1 (#13906, 500ms500\,ms). Names as in Table 3; metrics as in Sec. 2.4.2. Variant IpI_p (A) mean dXd_X (m) mean flux (10−4Wb10^-4\,Wb) LCFS (×10−3× 10^-3) u¯ u ↑ r¯ r ↑ Reward-PPO 735 / 1310 20.9 / 35.5 13.06 / 18.89 14.7 / 21.9 0.563 0.232 Adv-mean 1615 / 2149 34.2 / 45.8 1.77 / 2.70 8.7 / 14.9 0.830 0.665 Adv-SmoothMax 1630 / 2545 17.6 / 27.3 2.44 / 5.22 13.4 / 18.0 0.879 0.760 + gate 2779 / 3384 18.5 / 29.0 1.30 / 2.76 15.3 / 21.7 0.899 0.760 + resid == AdvA-PPO 260 / 624 17.3 / 26.9 0.66 / 1.93 12.2 / 17.2 0.925 0.810 + proj 1601 / 1916 18.9 / 29.7 1.27 / 2.10 8.5 / 11.3 0.894 0.765 + gate+resid 1137 / 1445 16.8 / 27.3 1.70 / 2.85 7.5 / 8.8 0.894 0.748 + gate+proj 925 / 1187 23.3 / 36.7 2.01 / 3.28 6.1 / 8.8 0.870 0.726 Table 6 decomposes the nominal AdvA-PPO gain of Sec. 3.1 along the design sequence of Table 3 (metrics in Sec. 2.4.2). Reward-PPO already holds IpI_p reasonably (735A735\,A RMSE) but leaves a large mean flux error (13.1×10−413.1× 10^-4), so the XPT is not flux-tight. Moving credit assignment before scalarisation (Adv-mean, then Adv-SmoothMax) cuts that flux error by about an order of magnitude, yet both raise IpI_p RMSE above 1.6kA1.6\,kA. The two are not interchangeable: Adv-mean reaches a slightly lower flux (1.771.77 versus 2.44×10−42.44× 10^-4) but a much worse mean dXd_X (34.234.2 versus 17.6m17.6\,m), so the softmin temperature alone does not select the operating point—geometry and current still trade. On the Adv-SmoothMax anchor, ++\;gate further lowers mean flux (1.30×10−41.30× 10^-4) while driving IpI_p to 2.78kA2.78\,kA RMSE: attenuating saturated channels sharpens the shape update and starves the current loop. Adding only ++\;resid == AdvA-PPO recovers both sides of that trade-off and is the strongest balanced cell: relative to Adv-SmoothMax it reduces IpI_p RMSE from 1.631.63 to 0.26kA0.26\,kA and mean flux from 2.442.44 to 0.66×10−40.66× 10^-4, with mean dXd_X essentially unchanged (17.317.3 versus 17.6m17.6\,m), and it also leads the shared scoring axis (u¯=0.925 u=0.925, r¯=0.810 r=0.810). That joint recovery is what links the family-level AdvA-PPO versus Reward-PPO contrast of Sec. 3.1 to a concrete component choice. The remaining proxy combinations do not overturn it: ++\;gate+proj records the best LCFS (6.1×10−36.1× 10^-3) but is weaker on IpI_p (925A925\,A) and mean flux (2.01×10−42.01× 10^-4) than residual alone, so the boundary looks tidy while the XPT flux condition and current loop do not; ++\;proj and ++\;gate+resid likewise fail to match AdvA-PPO on IpI_p or mean flux. Among gating, projection, and residual, the capped residual is therefore retained: it is the most balanced correction and the one that keeps the four-null flux lock that defines the XPT here. 3.3 Robustness evaluation This subsection tests whether the same frozen controllers remain usable under two deployment stresses. First, diagnostic noise and observation delay of the kind present in the experimental plant (R1–R3). Second, a change of the magnetic configuration at flattop hand-over: the controller takes over near t0≈500mst_0≈ 500\,ms from a different parent discharge while the #13906 XPT geometric target stays fixed (I2–I4). F+PID, Reward-PPO, and AdvA-PPO are evaluated without retraining; disturbance magnitudes, initials, and the clean-state scoring rule follow Sec. 2.4.1. 3.3.1 Zero-shot diagnostic noise and observation delay R1–R3 of Sec. 2.4.1 stress the three controllers on I1 (#13906). Table 7 reports survival and RMSE over the survived horizon; Figure 9 shows the channel-wise evolution under the combined delay+noise stress (R3). Table 7: Zero-shot noise/delay on I1 (#13906); RMSE over the survived horizon (Sec. 2.4.2). Method Surv. (ms) IpI_p (A) Prim. X (m) Prim. flux (10−4Wb10^-4\,Wb) Sec. X (m) Sec. flux (10−4Wb10^-4\,Wb) LCFS (×10−3× 10^-3) R1: delay 11–3ms3\,ms F+PID 500 637 38.1 2.16 29.6 149.81 160.8 Reward-PPO 500 11997 23.0 28.38 41.1 36.88 229.6 AdvA-PPO 500 4521 26.6 13.04 24.8 17.85 18.8 R2: full diagnostic noise F+PID 500 74817 46.1 23.02 22.3 55.54 222.0 Reward-PPO 500 2183 12.6 5.37 21.0 8.31 53.0 AdvA-PPO 500 2394 26.4 12.88 31.9 12.60 19.3 R3: delay + noise F+PID 500 68752 48.6 368.72 24.9 356.89 422.3 Reward-PPO 322 5222 30.8 15.17 28.5 21.39 94.6 AdvA-PPO 500 5521 29.0 12.79 41.7 17.82 26.3 Figure 9: R3: delay + noise on I1 (#13906). Panel layout as Figure 7. The three conditions expose different failure modes rather than a single winner. On R1 (delay only) all three finish the horizon: F+PID keeps the best IpI_p tracking (637A637\,A RMSE), consistent with a strong feedforward backbone that is only weakly disturbed by a few milliseconds of observation lag, while AdvA-PPO leads on LCFS and Reward-PPO suffers a large IpI_p RMSE under the per-step delay jitter. On R2 (noise only) the ranking reverses on the current channel: both learned controllers hold IpI_p near 22–3kA3\,kA RMSE, whereas F+PID collapses to >70kA>70\,kA. That contrast is expected from the classical loop structure (Sec. 2.3.1): derivative action on noisy IpI_p, coil-current, and R/Z measurements amplifies high-frequency sensor error into coil-voltage corrections, so the same PID that is benign under clean delayed observations becomes harmful under full diagnostic noise. Within the learned pair on R2, Reward-PPO is honestly the stronger cell—best IpI_p, primary/secondary X distance, and both flux channels—with AdvA-PPO retaining the best LCFS. R3 combines both stresses and is the decisive joint test: Reward-PPO’s R2 advantage does not carry over—it terminates at 322ms322\,ms—while AdvA-PPO and F+PID complete the horizon (Figure 9). Among survivors AdvA-PPO keeps the lowest boundary and primary-flux errors; F+PID again survives but with the same noise-dominated current/flux/boundary degradation seen on R2. Taken together, the single-factor cells are informative but not conclusive: F+PID resists delay yet fails under noise, and Reward-PPO wins the noise-only comparison yet breaks under the joint load. On the combined R3 protocol that matches concurrent plant delay and diagnostic noise, AdvA-PPO is the strongest controller—the only learned policy that both survives and retains a usable XPT shape. 3.3.2 Zero-shot cross-initialization I2–I4 of Table 4 hold the #13906 XPT target, observation map, and reward operators fixed, and change two quantities at once: the FGE initial equilibrium and the IpI_p reference (training-scale ∼400kA \!400\,kA to ∼500kA \!500\,kA). That joint shift—new shape or topology at hand-over plus a higher current setpoint—is a hard zero-shot test for a policy trained on a single initial. Each controller keeps the checkpoint (or feedforward table) already used on I1, with no per-initial redesign. In particular, F+PID does not receive a newly computed feedforward for each withheld discharge—that would amount to redesigning the classical controller and is outside the present scope—so the classical baseline is the #13906 plant feedforward and PID gains of Sec. 2.3.1, only re-anchored at reset to remove the trivial CS/PF current offset (and with the CS current guard removed). We report survival and whether the frozen XPT shape is reached (Figure 10, Table 8). Figure 10: Zero-shot cross-init flux under frozen #13906 XPT. Rows: #13705, #13844, #15892. Columns: initial equilibrium and F+PID / Reward-PPO / AdvA-PPO at episode end. Grey: initial LCFS/X; blue: final; black dashed/circles: target. Table 8: Zero-shot cross-init to frozen #13906 XPT (I2–I4); metrics as in Sec. 2.4.2 (full survival =500ms=500\,ms). F+PID feedforward is re-anchored as described in the text. Init. Controller Surv. IpI_p (kA) mean dXd_X (m) mean flux (10−4Wb10^-4\,Wb) LCFS (×10−3× 10^-3) u¯ u ↑ r¯ r ↑ I2 #13705 divertor F+PID full 9.4 / 42.3 125 / 178 1315 / 5054 232 / 596 0.02 0.00 Reward-PPO full 32.8 / 79.3 85 / 179 889 / 5013 46 / 70 0.07 0.00 AdvA-PPO full 31.8 / 75.6 69 / 179 909 / 5187 63 / 129 0.11 0.01 I3 #13844 divertor F+PID full 2.8 / 6.1 126 / 174 1305 / 5009 97 / 119 0.02 0.00 Reward-PPO full 19.6 / 56.8 72 / 168 990 / 5333 50 / 110 0.10 0.00 AdvA-PPO full 17.6 / 57.3 60 / 176 1061 / 5781 79 / 241 0.14 0.00 I4 #15892 limiter F+PID full 3.4 / 6.9 93 / 167 1168 / 5082 139 / 150 0.02 0.00 Reward-PPO 90 31.4 / 43.2 103 / 173 1697 / 5074 133 / 273 0.04 0.00 AdvA-PPO 96 20.7 / 26.5 99 / 170 1714 / 5207 182 / 272 0.04 0.00 Under this shift the classical baseline is the more transferable stabiliser. The same #13906 F+PID gains and feedforward, only re-anchored, complete the horizon on every withheld initial and hold IpI_p tightly (2.82.8–9.4kA9.4\,kA RMSE against 1717–33kA33\,kA for the policies), spanning the training-scale 400kA400\,kA plant and the 500kA500\,kA test setpoints. That stability does not buy shape precision: mean dXd_X stays ∼93 \!93–126m126\,m, mean flux errors remain large, and the limiter initial is held without converting to the target XPT (Figure 10). On the two divertor initials both policies survive and improve the geometric channels relative to F+PID (mean dXd_X 6060–85m85\,m, lower LCFS and flux RMSE), with AdvA-PPO leading on dXd_X and u¯ u while Reward-PPO keeps the lower LCFS error—so the advantage-level update is not uniformly better off-distribution. The current and boundary channels trade against each other on #13705, whose CS starts at −15.7kA-15.7\,kA and therefore has about half the volt-second headroom of the other initials: F+PID does hold IpI_p there, but only by drawing 54kA54\,kA of CS current against 4343–45kA45\,kA on the other two. The limiter I4 (#15892) fails both learned controllers: Reward-PPO terminates at step 9090 and AdvA-PPO at step 9696, each with mean X-point distance still ∼100m \!100\,m, while F+PID survives without converting to the target XPT (Table 8). The required limiter-to-divertor topology change is absent from the AdvA-PPO training distribution; the matched early stop of the Reward-PPO slot shows the difficulty is shared. F+PID already travels across IpI_p and initial as a stabiliser, while the learned controller can still be adapted for the missing conversion: that is the motivation for the multi-initialization fine-tuning of Sec. 3.3.3. 3.3.3 Fine-tuning across magnetic initials Zero-shot AdvA-PPO covers the two divertor initials but fails on the limiter I4 (#15892). A short multi-initialization fine-tune of the frozen checkpoint asks whether one adapted policy can span divertor and limiter resets under the same frozen #13906 XPT target. The protocol widens only the reset distribution: 100100 further iterations with a Kullback–Leibler (KL) divergence-capped update [19], unchanged reward operators and observation map, and evaluation under the harness of Sec. 3.1. Because survival differs before and after adaptation, Figure 11 and Table 9 summarise the before/after comparison; every error and objective entry in the table is taken on the common window of its before/after pair, with survival reported separately. Figure 11: AdvA-PPO multi-initialization fine-tune (100100 iterations, frozen #13906 XPT). Grey dashed: zero-shot; red: fine-tuned. Top: u¯ u; bottom: mean dXd_X. Early stop marked by a dotted vertical line. Table 9: AdvA-PPO before/after multi-initialization fine-tuning (100100 iterations, frozen #13906 XPT). Errors and u¯ u/r¯ r on the common window of each pair; survival separate. On #15892 the common window is the 9696 zero-shot steps; over its own 500ms500\,ms the fine-tuned run reaches u¯=0.08 u=0.08, IpI_p RMSE 20.8kA20.8\,kA, mean dX=74mmd_X=74\,m. Init. Controller Surv. IpI_p (kA) mean dXd_X (m) LCFS (×10−3× 10^-3) u¯ u ↑ r¯ r ↑ I1 #13906 (XPT / training) zero-shot full 0.3 17 12 0.92 0.81 fine-tuned full 0.7 24 17 0.77 0.58 I2 #13705 (divertor) zero-shot full 31.8 69 63 0.11 0.01 fine-tuned full 26.2 67 64 0.11 0.00 I3 #13844 (divertor) zero-shot full 17.6 60 79 0.14 0.00 fine-tuned full 13.7 58 66 0.13 0.00 I4 #15892 (limiter) zero-shot 96 20.7 99 182 0.04 0.00 fine-tuned full 25.0 92 88 0.05 0.00 The fine-tuned policy covers all four initials (Figure 11). On #15892 survival rises from 9696 steps to the full horizon; over the shared 9696 steps the objective is essentially unchanged (0.040.04 against 0.050.05) while the boundary error halves (182182 to 88×10−388× 10^-3), so the budget buys the remaining horizon and a cleaner boundary rather than a uniformly stronger tracker. On the two divertor initials, already full-horizon zero-shot, the objective stays flat (0.110.11 to 0.110.11 on #13705, 0.140.14 to 0.130.13 on #13844), with a modest IpI_p gain (31.831.8 to 26.226.2 and 17.617.6 to 13.7kA13.7\,kA RMSE) and geometric channels essentially held. The inevitable compromise appears on the training initial #13906: u¯ u drops from 0.920.92 to 0.770.77 and the mean worst channel from 0.810.81 to 0.580.58. Multi-initialization coverage and nominal specialisation are therefore not free jointly under this short budget—the same weights that span divertor and limiter give up part of the #13906 solution. Each cell is a single trajectory, so survival figures are attempted rollouts rather than success rates. Fine-tuning under the noise and delay conditions of Sec. 3.3.1 is left for future work. 4 Conclusion We established an AI-enabled multi-objective framework for EXL-50U XPT magnetic control, formulated as reconstructed-null feedback through shared poloidal-field actuators. Its ten physical objectives jointly represent IpI_p, the LCFS, four X-point positions, and four X-point flux constraints. The FGE environment reproduces the experimental #13906 equilibrium at centimetre-scale boundary and null accuracy, providing an experiment-calibrated free-boundary test bed for closed-loop evaluation. AdvA addresses multi-objective temporal credit assignment by retaining one value head and one GAE per channel on a shared critic backbone, applying worst-objective-aware SmoothMax only after advantage estimation, and adding a norm-capped component residual to the aggregate policy update. On the nominal 500ms500\,ms rollout, AdvA-PPO raises the shared evaluation score u¯ u from 0.56 to 0.93 and the mean worst-channel score from 0.23 to 0.81 relative to Reward-PPO, while reducing mean X-point flux RMSE by about 20×20× and using the lowest mean per-step voltage change. Among the evaluated corrections, the capped residual gives the most balanced recovery of IpI_p and X-point flux accuracy; the result favours a controlled residual over stacking every available multi-objective correction. The broader evaluation reveals complementary, regime-dependent strengths. Under combined measurement noise and observation delay, AdvA-PPO is the only learned controller that completes the horizon while retaining a usable XPT shape, although Reward-PPO is stronger under noise alone and the experiment-derived F+PID baseline tracks IpI_p well under delay. Across withheld initial equilibria, the same F+PID controller is the more transferable stabiliser but remains shape-imprecise and does not produce the limiter-to-XPT conversion; the learned controllers improve geometric regulation on the diverted initials but both fail that conversion zero-shot. Multi-initial adaptation extends AdvA-PPO to all four initials and exposes a coverage–specialisation trade-off. Together, these results establish the capability and current boundaries of reconstructed-null XPT feedback in an experiment-calibrated magnetic simulation; PTEFIT-in-the-loop and closed-loop machine performance remain to be tested experimentally. 5 Future outlook The immediate next step is staged on-machine validation on EXL-50U. The near-term campaign will first quantify the accuracy, latency, and stability of PTEFIT secondary-null position and flux estimates, and then evaluate the present RL controller with PTEFIT in the loop. Conservative actuator limits, online monitoring, and a deterministic F+PID fallback will allow the effects of real-time reconstruction and learned control to be assessed separately before extending the tests to disturbances and broader operating conditions. Algorithmic development will target robustness and generalisation beyond the single-initial training regime. Multi-initial and uncertainty-aware training should span plasma-current targets, measurement uncertainties, actuator errors, and FGE calibration residuals. Domain randomisation, constrained policy optimisation, and stage-dependent objectives provide complementary routes to wider coverage while limiting the loss of nominal precision exposed by the present fine-tuning study. The control task should also expand from a fixed XPT target to broader divertor configurations and controlled topology transitions. Strike-point and divertor-leg targets, richer LCFS descriptors, and explicit coil-current, voltage, and slew constraints are natural next objectives. Coordinating these tests with auxiliary-heating experiments will probe magnetic-configuration control under more relevant plasma conditions and prepare later integration with radiation, detachment, and target-heat-flux control as suitable models and diagnostics become available. Throughout this progression, hard safety limits, interlocks, fallback control, and staged validation must make machine protection part of the controller design rather than an external safeguard. References [1] T. Ambadkar, S. Panda, S. Kale, J. Dodge, and A. Verma (2026) Preference conditioned multi-objective reinforcement learning: decomposed, diversity-driven policy optimization. External Links: 2602.07764 Cited by: §1. [2] H. Anand et al. (2024) Real-time plasma equilibrium reconstruction and shape control for the mast upgrade tokamak. Nuclear Fusion 64, p. 086051. External Links: Document Cited by: Table 1, §1, §1. [3] J. Degrave et al. (2022) Magnetic control of tokamak plasmas through deep reinforcement learning. Nature 602, p. 414–419. External Links: Document Cited by: §1, §1, §2.2.1. [4] S. Gorno, O. Février, C. Theiler, T. Ewalds, F. Felici, T. Lunt, A. Merle, J. Degrave, B. P. Duval, K. Lee, H. Reimerdes, B. Tracey, M. Wischmeier, and C. Wüthrich (2024) X-point radiator and power exhaust control in configurations with multiple x-points in tcv. Physics of Plasmas 31, p. 072504. External Links: Document Cited by: §1, §1. [5] X. Gu et al. (2025) Poloidal field system and advanced divertor equilibrium configuration design of the ehl-2 spherical torus. Plasma Science and Technology 27, p. 024011. External Links: Document Cited by: §1. [6] C. Heiß, A. Merle, F. Carpanese, F. Felici, C. Donner, S. Marchioni, A. Mari, and O. Sauter (2026) FGE: a fast free-boundary grad–shafranov evolutive solver. Plasma Physics and Controlled Fusion 68, p. 045031. External Links: Document Cited by: §2.1, §2.2.1, §2.2.1. [7] S. Kerboua-Benlarbi, R. Nouailletas, B. Faugeras, E. Nardon, and P. Moreau (2024) Magnetic control of WEST plasmas through deep reinforcement learning. IEEE Transactions on Plasma Science. External Links: Document Cited by: §1. [8] E. Kolemen et al. (2018) Initial development of the diii-d snowflake divertor control. Nuclear Fusion 58, p. 066007. External Links: Document Cited by: §1. [9] A. Q. Kuang et al. (2020) Divertor heat flux challenge and mitigation in sparc. Journal of Plasma Physics 86. External Links: Document Cited by: §1. [10] B. LaBombard et al. (2015) ADX: a high field, high power density, advanced divertor and rf tokamak. Nuclear Fusion 55, p. 053020. External Links: Document Cited by: §1. [11] K. Lee et al. (2025) X-point target radiator regime in tokamak divertor plasmas. Physical Review Letters 134, p. 185102. External Links: Document Cited by: Table 1, §1, §1. [12] N. Lonigro et al. (2026) Initial observations in x-point target divertor discharges on mast-u. External Links: 2601.21840 Cited by: Table 1, §1. [13] A. Mele, M. A. Topalova, C. Galperti, and S. Coda (2025) First experimental demonstration of plasma shape control in a tokamak through model predictive control. External Links: 2506.20096 Cited by: §1. [14] V. Mnih et al. (2015) Human-level control through deep reinforcement learning. Nature 518, p. 529–533. External Links: Document Cited by: §2.3.3. [15] J. Moret, B. P. Duval, H. B. Le, S. Coda, F. Felici, and H. Reimerdes (2015) Tokamak equilibrium reconstruction code liuqe and its real time implementation. Fusion Engineering and Design 91, p. 1–15. External Links: Document Cited by: §2.2.2. [16] H. Munn, B. Tidd, P. Boehm, M. Gallagher, and D. Howard (2025) Scalable multi-objective robot reinforcement learning through gradient conflict resolution. External Links: 2509.14816 Cited by: §1, §2.3.5. [17] H. Raj et al. (2022) Improved heat and particle flux mitigation in high core confinement, baffled, alternate divertor configurations in the tcv tokamak. Nuclear Fusion. External Links: Document Cited by: Table 1, §1, §1. [18] J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel (2016) High-dimensional continuous control using generalized advantage estimation. External Links: 1506.02438 Cited by: §1, §2.3.4. [19] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017) Proximal policy optimization algorithms. External Links: 1707.06347 Cited by: §2.3.3, §3.3.3. [20] Y. Shi et al. (2026) Overview of exl-50u experiments: addressing key physics issues for future spherical torus reactors. Nuclear Fusion 66, p. 116009. External Links: Document Cited by: §1, §1, §1. [21] X. M. Song, J. X. Li, J. A. Leuer, J. H. Zhang, and X. Song (2019) First plasma scenario development for HL-2M. Fusion Engineering and Design 147, p. 111254. External Links: Document Cited by: §2.3.1. [22] V. A. Soukhanovskii et al. (2016) Snowflake divertor experiments in the diii-d, nstx, and nstx-u tokamaks aimed at the development of the divertor power exhaust solution. IEEE Transactions on Plasma Science 44 (12), p. 3445–3455. External Links: Document Cited by: §1. [23] A. Subbotin et al. (2026) Demonstration of reconstruction-free static magnetic control of diii-d plasma with deep reinforcement learning. Nuclear Fusion. External Links: Document Cited by: §1. [24] C. Theiler et al. (2017) Results from recent detachment experiments in alternative divertor configurations on tcv. Nuclear Fusion 57, p. 072008. External Links: Document Cited by: Table 1, §1. [25] B. D. Tracey, A. Michi, Y. Chervonyi, I. Davies, C. Paduraru, N. Lazic, F. Felici, T. Ewalds, C. Donner, C. Galperti, J. Buchli, M. Neunert, A. Huber, J. Evens, P. Kurylowicz, D. J. Mankowitz, and M. Riedmiller (2024) Towards practical reinforcement learning for tokamak magnetic control. Fusion Engineering and Design 200, p. 114161. External Links: Document Cited by: §1, §1. [26] M. V. Umansky et al. (2017) Assessment of x-point target divertor configuration for power handling and detachment front control. Nuclear Materials and Energy. External Links: Document Cited by: §1. [27] F. Wang et al. (2025) Divertor heat flux challenge and mitigation in the ehl-2 spherical torus. Plasma Science and Technology 27, p. 024009. External Links: Document Cited by: §1. [28] M. Winkel, K. Verhaegh, B. Kool, K. Lee, M. Carpita, A. Perek, R. Morgan, G. Derks, O. Février, C. Theiler, D. Brida, and M. van Berkel (2026) Detachment dynamics and disturbance rejection in the TCV x-point target divertor. External Links: 2606.23432 Cited by: Table 1, §1, §1. [29] T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn (2020) Gradient surgery for multi-task learning. In Advances in Neural Information Processing Systems, Vol. 33. Cited by: §1, §2.3.5. [30] G. H. Zheng, S. F. Liu, X. Gu, Y. P. Zhang, J. Li, Y. Liu, X. C. Lun, L. Xing, J. G. Chen, Z. Y. Chen, Y. Yu, D. Guo, Z. Y. Yang, H. S. Xie, X. M. Song, Y. J. Shi, and EXL-50U Team (2026) A novel numerical algorithms optimization method with machine learning frameworks: application on real-time plasmas equilibrium reconstruction in EXL-50U spherical torus. External Links: 2601.12378 Cited by: §1, §1.