Paper deep dive
SafeDivertor: Faithful Divertor Heat Flux Reconstruction from Macroscopic Plasma State Signals via Time-Frequency Prior Exploitation
Hao Si, Zehua Chen, Qingquan Yang, Xiao Wang, Dengdi Sun, Wanli Lyu, Gaoting Chen, Guosheng Xu, Hang Su, Jin Tang, Jun Zhu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/8/2026, 3:37:58 AM
Summary
The paper introduces SafeDivertor, a deep learning framework for reconstructing time-resolved radial divertor heat-flux profiles directly from multi-source macroscopic plasma-state signals in magnetic-confinement fusion devices. Unlike conventional infrared-based inversion methods that are offline and computationally expensive, SafeDivertor enables online-oriented reconstruction. The authors propose a novel dataset, DivMPS2HF, and a framework incorporating physical prior-aware initialization, input perturbation, spectral-aware reconstruction optimization via multi-scale STFT, and progressive training to address challenges like channel over-reliance and transient dynamics preservation.
Entities (13)
Relation Signals (9)
SafeDivertor → reconstructs → Divertor Heat-Flux
confidence 98% · directly reconstructs time-resolved radial heat-flux profiles from multi-source macroscopic plasma-state signals
SafeDivertor → evaluatedon → DivMPS2HF
confidence 95% · Experiments on DivMPS2HF demonstrate that SafeDivertor achieves the best overall performance
SafeDivertor → uses → Spectral-Aware Reconstruction Optimization (SRO)
confidence 95% · spectral-aware reconstruction optimization to exploit time-frequency priors
SafeDivertor → uses → Progressive Training (PT)
confidence 95% · progressive training to stabilize the optimization of these complementary objectives.
SafeDivertor → uses → Physical Prior-Aware Initialization (PPI)
confidence 95% · It employs physical prior-aware initialization to provide radial-distribution guidance for target channels
SafeDivertor → uses → Input Perturbation (IP)
confidence 95% · input perturbation to reduce over-reliance on specific heterogeneous signals
DivMPS2HF → contains → Macroscopic Plasma-State Signals
confidence 90% · DivMPS2HF, a multi-source discharge dataset that provides the data foundation... containing 67 input channels, including plasma-state diagnostic signals
Spectral-Aware Reconstruction Optimization (SRO) → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Divertor heat-flux analysis is essential for understanding plasma-wall interactions and protecting plasma-facing components in magnetic-confinement fusion devices, while conventional infrared-based inversion is usually performed after discharge and requires heat-conduction modeling with device-specific material properties, divertor geometry, and boundary conditions. Rather than accelerating this conventional infrared-based inversion paradigm, we introduce a new online-oriented signal-based reconstruction paradigm that directly reconstructs time-resolved radial heat-flux profiles from multi-source macroscopic plasma-state signals available during discharge. To enable systematic study of this task, we construct \textbf{DivMPS2HF}, a multi-source discharge dataset that provides the data foundation and benchmark for signal-based divertor heat-flux reconstruction. We further propose \textbf{SafeDivertor}, a task-driven framework designed to address the key challenges of signal-based heat-flux reconstruction. It employs physical prior-aware initialization to provide radial-distribution guidance for target channels, input perturbation to reduce over-reliance on specific heterogeneous signals, spectral-aware reconstruction optimization to exploit time-frequency priors and preserve transient dynamics, and progressive training to stabilize the optimization of these complementary objectives. Experiments on DivMPS2HF demonstrate that SafeDivertor achieves the best overall performance among the evaluated time-series baselines across all five metrics, establishing a new performance benchmark for signal-based divertor heat-flux reconstruction. The source code will be released on this https URL
Tags
Links
- Source: https://arxiv.org/abs/2608.05669v1
- Canonical: https://arxiv.org/abs/2608.05669v1
Trouble viewing inline? Open PDF directly →
Full Text
55,649 characters extracted from source content.
Expand or collapse full text
SafeDivertor: Faithful Divertor Heat Flux Reconstruction from Macroscopic Plasma State Signals via Time-Frequency Prior Exploitation Hao Si, Zehua Chen*, Qingquan Yang*, Xiao Wang, Dengdi Sun, Wanli Lyu, Gaoting Chen, Guosheng Xu, Hang Su, Jin Tang, and Jun Zhu ∙ Hao Si, Xiao Wang, Wanli Lyu, and Jin Tang are with the School of Computer Science and Technology, Anhui University, Hefei 230601, China. (email: e24201034@stu.ahu.edu.cn, xiaowang, lwl, tangjin@ahu.edu.cn)∙ Zehua Chen, Hang Su, and Jun Zhu are with the Department of Computer Science and Technology, Tsinghua University, Beijing, China. (email: zhc23thuml, suhangss, dcszj@tsinghua.edu.cn)∙ Dengdi Sun is with the School of Artificial Intelligence, Anhui University, Hefei 230601, China. (email: sundengdi@163.com)∙ Qingquan Yang, Gaoting Chen, and Guosheng Xu are with the Institute of Plasma Physics, Chinese Academy of Sciences, Hefei 230601, China. (email: yangqq, gaoting.chen, gsxu@ipp.ac.cn)Corresponding authors: Zehua Chen &\& Qingquan Yang Abstract Divertor heat-flux analysis is essential for understanding plasma-wall interactions and protecting plasma-facing components in magnetic-confinement fusion devices, while conventional infrared-based inversion is usually performed after discharge and requires heat-conduction modeling with device-specific material properties, divertor geometry, and boundary conditions. Rather than accelerating this conventional infrared-based inversion paradigm, we introduce a new online-oriented signal-based reconstruction paradigm that directly reconstructs time-resolved radial heat-flux profiles from multi-source macroscopic plasma-state signals available during discharge. To enable systematic study of this task, we construct DivMPS2HF, a multi-source discharge dataset that provides the data foundation and benchmark for signal-based divertor heat-flux reconstruction. We further propose SafeDivertor, a task-driven framework designed to address the key challenges of signal-based heat-flux reconstruction. It employs physical prior-aware initialization to provide radial-distribution guidance for target channels, input perturbation to reduce over-reliance on specific heterogeneous signals, spectral-aware reconstruction optimization to exploit time-frequency priors and preserve transient dynamics, and progressive training to stabilize the optimization of these complementary objectives. Experiments on DivMPS2HF demonstrate that SafeDivertor achieves the best overall performance among the evaluated time-series baselines across all five metrics, establishing a new performance benchmark for signal-based divertor heat-flux reconstruction. The source code will be released on https://github.com/Event-AHU/OpenFusion I Introduction Divertor heat-flux analysis is essential for safe and high-performance fusion operation, as concentrated and transient heat loads can damage plasma-facing components [1]. This challenge is particularly critical for future burning-plasma devices such as ITER, where divertor heat loads are subject to strict material limits [2]. Accurate time-resolved estimation is therefore important for heat-load assessment, component protection, and future online control [3]. Figure 1: Illustration of the conventional IR-based heat-flux inversion workflow and the proposed signal-based neural reconstruction workflow. (a) Tokamak poloidal cross-section. (b) Offline workflow based on IR image sequences and heat-conduction inversion. (c) Proposed online-oriented workflow that reconstructs heat-flux profiles from multi-source plasma-state signals. (d) Target divertor heat-flux profile q(r,t)q(r,t). As shown in Figure 1(b), conventional divertor heat-flux analysis mainly follows a conventional infrared-based inversion paradigm, where the target-surface temperature is first measured and the heat-flux profile is then reconstructed using computationally expensive numerical solvers, typically for offline post-shot analysis [4, 5, 6, 1]. Recent learning-based methods, such as Physics-Informed Neural Networks (PINNs) for accelerating heat-equation solving and convolutional neural networks (CNNs) for heat-flux reconstruction, improve computational efficiency but still only replace or accelerate individual components within the conventional thermal-analysis paradigm [3, 7, 8]. However, these methods remain dependent on thermal measurements, heat-conduction modeling, or simulation-derived data, limiting their direct application to online-oriented reconstruction. In contrast, macroscopic plasma-state signals are routinely recorded during plasma discharges and are usually available for online monitoring and control [9, 10, 11, 12]. This naturally raises an important question: Can time-resolved radial divertor heat-flux profiles be reconstructed directly from macroscopic plasma-state signals, without relying on conventional infrared-based inversion? As illustrated in Figure 1(c), we introduce an online-oriented signal-based heat-flux reconstruction task that directly maps macroscopic plasma-state signals to time-resolved radial divertor heat-flux profiles. To support this task, we construct DivMPS2HF, a multi-source discharge dataset that aligns heterogeneous plasma-state signals with radial heat-flux profiles. We further propose SafeDivertor, which jointly models observed signals and target-channel placeholders. Because the target-channel placeholders provide no heat-flux information, physical prior-aware initialization (PPI) supplies them with a channel-level statistical heat-flux prior, providing radial-distribution guidance before reconstruction. Since heterogeneous plasma-state signals may differ in reliability and cause the model to over-rely on a small subset of channels, input perturbation (IP) regularizes the conditioning process by perturbing randomly selected observed channels during training. Moreover, because point-wise reconstruction loss alone tends to smooth rapid heat-flux variations, spectral-aware reconstruction optimization (SRO) exploits the time-frequency prior embodied in the ground-truth heat-flux profiles through a multi-scale short-time Fourier transform (STFT) loss, thereby preserving transient dynamics and spectral consistency across multiple temporal scales. As these objectives emphasize different aspects of reconstruction, progressive training (PT) introduces them in stages to stabilize optimization. Experiments on DivMPS2HF show that SafeDivertor achieves the best overall performance among the evaluated time-series baselines across all five metrics. The main contributions of this work are summarized as follows: • To the best of our knowledge, we are the first to formulate an online-oriented signal-based reconstruction paradigm that shifts divertor heat-flux analysis from offline infrared-based inversion toward online-oriented reconstruction from macroscopic plasma-state signals available during discharge. • We propose SafeDivertor, a task-driven framework that integrates physical prior-aware initialization, input perturbation, spectral-aware reconstruction optimization, and progressive training for faithful heat-flux reconstruction. • We construct DivMPS2HF, a multi-source discharge dataset and evaluation benchmark for this task, and establish a strong reference baseline that achieves the best overall performance among the evaluated time-series methods across all five metrics. I Related Work Divertor Heat-Flux Analysis. Accurate divertor heat-flux analysis is essential for understanding plasma-wall interactions and protecting plasma-facing components in magnetic-confinement fusion devices. Existing studies mainly rely on infrared thermal imaging technology. First, the surface temperature of the divertor target plate is measured by an IR camera. Then, the heat flux is inferred by solving the heat conduction equations related to material properties and boundary conditions [5, 4]. For example, KSTAR employs the NANTHELOT heat-flux reconstruction code to solve the heat diffusion equation from measured tile surface temperature, while EAST calculates divertor heat-flux distributions from IR measurements using DFLUX or FEM-based solvers [5, 6]. These pipelines are often computationally expensive and are commonly used for offline post-shot analysis or require additional acceleration for real-time deployment. To improve efficiency, recent studies have introduced learning-based methods into heat-flux analysis. For example, approaches based on physics-informed neural networks (PINNs) have been explored to accelerate heat-equation solving toward real-time heat-flux estimation, while convolutional neural network (CNN)-based methods have been used to reconstruct parameterized heat-flux profiles from simulation-generated thermocouple signals [3, 7]. The method based on neural networks has also been applied to plasma diagnostic signals for cross-device disruption prediction [9], EAST disruption prediction [10], and four-class ELM recognition using HGTS-Former on EAST-ELM640 [13]. Deep neural networks have further predicted ELM onset from turbulence signals [14], while time-series extrinsic regression has reconstructed missing electron-temperature diagnostics [15]. However, most existing heat-flux analysis methods either rely on thermal imaging measurements or focus on solving heat-conduction equations, while direct reconstruction of time-resolved divertor heat-flux profiles from multi-source discharge signals remains less explored. Multivariate Time-Series Modeling. Multivariate time series modeling has been extensively studied to capture temporal dependencies and cross-variable interactions across multiple channels [16, 17, 18]. In recent years, general time series models have achieved significant performance from different modeling perspectives. For example, iTransformer treats each variate as a token by inverting the dimension to enhance cross-variable dependency modeling [19]. Timer-XL [20] represents time series as a long context token sequence and uses Transformer for multivariate sequence modeling. SimpleTM [21] combines signal processing ideas with a lightweight attention structure to provide a simple yet effective multivariate prediction baseline. TimeFilter [22] captures local dynamic patterns by constructing a spatiotemporal graph structure and performing patch-level filtering. Frequency-domain and multi-periodic representations have also been explored for general time-series modeling. FreTS [23] performs inter-series and intra-series learning through frequency-domain MLPs to capture global temporal and cross-variable dependencies. FiLM [24] combines frequency enhancement with Legendre projection to improve long-term time-series forecasting. TimesNet [25] transforms one-dimensional time series into two-dimensional representations according to multiple periods, thereby modeling both intra-period and inter-period variations. TimePerceiver [26] uses an encoder-decoder structure and learnable queries to adapt to prediction targets at different time locations. Sonnet [27] captures temporal dynamics in the spectral domain and captures cross-variable relationships by utilizing spectral consistency between variables. Some methods also focus on time series prediction under exogenous variable conditions. For example, TimeXer [28] fuses endogenous and exogenous variable information through patch-wise self-attention and variate-wise cross-attention, while GCGNet [29] uses graph structure consistency constraints to model the correlation between exogenous and target variables. Meanwhile, time-series imputation provides useful insights into recovering unavailable variables from observed contexts. Representative works such as T1 [30] combine a CNN-Transformer hybrid architecture with a channel-head binding mechanism to achieve robust imputation. Although these methods provide useful paradigms for multivariate time-series modeling, divertor heat-flux reconstruction differs from standard forecasting, exogenous-variable prediction, and partial imputation, because all target heat-flux channels are unavailable throughout the current input window and must be reconstructed from physically different macroscopic plasma-state signals. I Motivation Unlike conventional time-series imputation [31, 32], which reconstructs partially missing observations within the same variable space, our task requires reconstructing entirely unavailable heat-flux channels from physically distinct macroscopic plasma-state signals. Therefore, we formulate it as a structured channel-level reconstruction problem, as follows. Let =((i),(i))i=1ND= \ (X^(i),Y^(i) ) \_i=1^N denote a discharge dataset, where N denotes the number of samples after window sampling. For the i-th sample, (i)X^(i) represents the macroscopic plasma-state signals within a time window of length T, and (i)Y^(i) denotes the corresponding time-resolved heat-flux profile on the divertor target plate. The input sequence is defined as (i)=[x1(i),x2(i),…,xC(i)]∈ℝC×TX^(i)=[x_1^(i),x_2^(i),...,x_C^(i)] ^C× T, where C=67C=67 is the number of signal channels; and the corresponding target heat-flux profile is defined as (i)=[y1(i),y2(i),…,yO(i)]∈ℝO×TY^(i)=[y_1^(i),y_2^(i),...,y_O^(i)] ^O× T, where O=116O=116 is the number of heat-flux channels. Since the actual heat-flux sequence within the current time window is unobservable, we further represent this task as a structured channel-level reconstruction problem based on target-channel initialization. Specifically, an initial representation of the target heat-flux channels, denoted as init(i)∈ℝO×TY^(i)_init ^O× T, is concatenated with the observed signals (i)X^(i) along the channel dimension to form a unified input sequence: (i)=[(i),init(i)]∈ℝF×T,F=C+O.Z^(i)=[X^(i),Y^(i)_init] ^F× T, F=C+O. (1) It is important to note that init(i)Y_init^(i) does not contain the actual heat-flux information of the current sample; it is only used as an initial placeholder for the target heat-flux channels. The model reconstructs the actual heat-flux profile based on this unified input sequence: ^(i)=f((i)),f:ℝF×T→ℝO×T. Y^(i)=f_ θ(Z^(i)), f_ θ:R^F× T ^O× T. (2) The model parameters are optimized by minimizing the reconstruction error between the reconstructed heat-flux sequence and the actual heat-flux sequence: ∗=argmin1N∑i=1Nℒ(^(i),(i)). θ^*= _ θ 1N _i=1^NL ( Y^(i),Y^(i) ). (3) During training and evaluation, the loss and all metrics are computed only on the O target heat-flux channels. Although the task formulation is straightforward, it presents several task-specific challenges. First, the target-channel placeholders contain no heat-flux information, leaving the model without explicit guidance about the radial distribution to be reconstructed. Second, the heterogeneous plasma-state signals differ in physical meaning and measurement reliability, which may cause the model to over-rely on a small subset of channels. Third, point-wise reconstruction losses tend to smooth rapid heat-flux variations and fail to explicitly exploit the time-frequency prior contained in the target profiles. Finally, these objectives emphasize different aspects of reconstruction and are difficult to optimize simultaneously. These challenges motivate SafeDivertor, which introduces PPI, IP, SRO, and PT to address the lack of target guidance, over-reliance on specific input channels, insufficient spectral preservation, and optimization instability, respectively, as detailed in the following section. Figure 2: Overview of the proposed SafeDivertor framework for divertor heat-flux reconstruction from macroscopic plasma-state signals. SafeDivertor integrates physical prior-aware initialization, input perturbation, T1 backbone reconstruction, and multi-scale STFT-based spectral supervision. IV Methodology IV-A Dataset Curation To support the proposed online-oriented reconstruction task, we construct DivMPS2HF, a multi-source discharge dataset containing 7777 shots, divided into 6464, 33, and 1010 shots for training, validation, and testing, respectively. The split is performed at the shot level to prevent cross-shot information leakage. All signals are temporally aligned and resampled to a unified 1kHz1~kHz grid. Each sample contains 6767 input channels, including plasma-state diagnostic signals, equilibrium and configuration parameters, and signals related to external actuators and device control. These channels are selected from a broader pool of candidate signals to retain signals with reliable data quality, strong physical relevance to divertor heat-load evolution, and useful predictive information. The reconstruction target is a time-resolved radial heat-flux profile consisting of 116116 channels, obtained from post-shot heat-flux analysis as supervised labels. Notably, infrared measurements and heat-conduction inversion results are used only to construct the reference labels and are not included in the model input. During inference, SafeDivertor directly reconstructs the heat-flux profile from macroscopic plasma-state signals through a single neural forward pass. IV-B Physical Prior-aware Initialization In this paper, we model the multi-source signals and target heat-flux channels as a unified structured time series. The first 6767 channels represent the observed multi-source signals, while the last 116116 channels are zero-valued placeholders for the target heat-flux channels at different radial locations on the divertor target plate. However, target-channel placeholders contain no heat-flux distribution information and fail to reflect the inherent differences in basic heat loads at different radial locations. Therefore, we introduce PPI and construct a channel-level heat-flux prior from the training set to provide statistical guidance for the unavailable target channels, as illustrated in Part 22 of Figure 2. Specifically, for the o-th target heat-flux channel, its statistical prior is computed as the average heat-flux value over all training shots and time steps: μo=1∑s=1StrainLs∑s=1Strain∑t=1Lso,t(s),o=1,…,O, _o= 1 _s=1^S_trainL_s _s=1^S_train _t=1^L_sY^(s)_o,t, o=1,…,O, (4) where StrainS_train denotes the number of training discharge shots, LsL_s denotes the length of the s-th training shot, and Yo,t(s)Y_o,t^(s) is the heat-flux value of the o-th target channel at time step t in the s-th training shot. For each sample, the radial heat-flux prior is defined as follows: =[μ1,μ2,…,μO]⊤,=1×T∈ℝO×T, μ=[ _1, _2,…, _O] , = μ1_1× T ^O× T, (5) where 1×T∈ℝ1×T1_1× T ^1× T denotes an all-ones row vector and T is the length of the input time window. Based on this prior, we introduce a lightweight prior encoding branch to extract structural information. Specifically, the original input with zero-valued target placeholders is first encoded into a feature representation 0H_0. According to the channel organization of the unified input, 0H_0 is divided into the observed-signal features and target-placeholder features: 0(i)=[0,obs(i),0,tar(i)].H_0^(i)= [H_0,obs^(i),H_0,tar^(i) ]. (6) Meanwhile, the statistical heat-flux prior P is processed by a convolutional prior encoder g(⋅)g_ θ(·) to obtain the prior feature. The prior feature is then added only to the target-placeholder features, yielding the prior-enhanced representation: (i)=[0,obs(i),0,tar(i)+g()].H^(i)= [H_0,obs^(i),H_0,tar^(i)+g_ θ(P) ]. (7) Here, (i)H^(i) denotes the prior-enhanced feature representation. This design preserves the observed-signal features and injects the statistical radial heat-flux information only into the target-channel representations through an independent prior encoding branch. IV-C Input Perturbation In actual discharge data acquisition, multi-source signals are inevitably affected by factors such as hardware noise, acquisition errors, and signal fluctuations. The observation quality of different channels may also change with the shot and time window. To reduce the model’s over-reliance on specific input channels and improve its tolerance to channel-level noise, we introduce IP during the training phase. As illustrated in Part 1 of Figure 2, Gaussian perturbation is applied only to selected observable signals, while the remaining observable signals and the heat-flux prior are kept unchanged. This design simulates noisy or unreliable measurements. Specifically, for each training sample, we randomly select 20%20\% of the 6767 observable signals and add Gaussian noise with a standard deviation of 0.030.03 to the complete time series of these channels. The augmented input is defined as follows: ~c,t=c,t+ϵc,t,c∈n,c,t,c∉n,ϵc,t∼(0,σ2), X_c,t= casesX_c,t+ _c,t,&c _n,\\ X_c,t,&c _n, cases _c,t (0,σ^2), (8) where ~, X_c,t and c,tX_c,t denote the augmented and original signals of channel c at time step t, nC_n is the randomly selected subset containing 20%20\% of the channels, and ϵc,t _c,t denotes the added Gaussian noise. σ controls the noise intensity and is set to 0.030.03 in our experiments. IV-D Spectral-aware Reconstruction Optimization Basic Reconstruction Loss. Given the reconstructed heat-flux profile Y and the true heat-flux profile Y, we first adopt the mean squared error as the basic reconstruction loss: ℒMSE=1BOT∑b=1B∑o=1O∑t=1T(^b,o,t−b,o,t)2,L_MSE= 1BOT _b=1^B _o=1^O _t=1^T ( Y_b,o,t-Y_b,o,t )^2, (9) where B is the batch size, O=116O=116 denotes the number of target heat-flux channels, and T is the length of the time window. The MSE loss provides point-wise supervision for numerical fidelity in the time domain. Multi-scale STFT-based Spectral Loss. Although the MSE loss provides direct point-wise supervision in the time domain, it mainly measures numerical deviations at each time step and does not explicitly constrain the frequency characteristics of the reconstructed heat-flux sequences. In divertor heat-load evolution, rapid variations and transient events may appear as high-frequency components in the heat-flux profile. Models optimized solely by MSE may produce overly smoothed reconstructions that fail to preserve these spectral details. Therefore, we introduce SRO, which employs a spectral loss based on the multi-scale short-time Fourier transform (STFT). Compared to single-scale spectral constraints, multi-scale designs are better suited for heat-flux reconstruction because different time scales emphasize different dynamic modes. Shorter temporal scales are more sensitive to rapid local variations and transient heat-load dynamics, while longer temporal scales help capture broader heat-flux evolution trends. Therefore, applying spectral constraints at multiple temporal scales encourages the model to preserve both local high-frequency dynamics and global time-frequency consistency. Let ϕ(⋅)φ(·) denote the log-magnitude STFT operator along the temporal dimension: ϕ()=log(|STFT()|+ϵ),φ(Y)= ( |STFT(Y) |+ε ), (10) where STFT(⋅)STFT(·) denotes the STFT operation along the temporal dimension, and ϵε is a small constant for numerical stability. The multi-scale STFT loss is formulated as follows: ℒMS−STFT=∑r=1RλrMAE(ϕ(r),ϕ(^r)),L_MS-STFT= _r=1^R _rMAE (φ(Y_r),φ( Y_r) ), (11) where R=3R=3, rY_r and ^r Y_r denote the ground-truth and reconstructed heat-flux sequences at the r-th temporal scale, respectively, and λr _r is the weight for the r-th temporal scale. By matching the log-magnitude spectra of heat-flux sequences at multiple temporal scales, this loss helps the model preserve both rapid high-frequency variations and broader temporal spectral patterns. Overall Training Objective. The final training objective combines the basic reconstruction loss and the multi-scale STFT-based spectral loss. The MSE loss provides direct supervision for numerical accuracy in the time domain, while the multi-scale STFT loss constrains the spectral characteristics of heat-flux sequences across different temporal scales. The overall loss is defined as follows: ℒtotal=ℒMSE+ℒMS−STFT.L_total=L_MSE+L_MS-STFT. (12) By jointly optimizing these objectives, the model improves the point-wise numerical accuracy and multi-scale spectral consistency of the reconstructed heat-flux profile. IV-E Progressive Training Directly applying input perturbation and spectral-aware reconstruction optimization from the beginning may increase the optimization difficulty, since the model needs to simultaneously learn the basic reconstruction mapping, robustness to input perturbations, and spectral consistency. To stabilize the training process, we adopt a three-stage training strategy. In the first stage, the model is trained only with the basic MSE reconstruction loss, allowing it to learn a stable mapping from multi-source signals to heat-flux profiles. In the second stage, channel-level noise augmentation is enabled while the model continues to optimize the reconstruction objective, which improves its robustness to input-channel perturbations. In the third stage, the input perturbation is disabled, and the multi-scale STFT loss is introduced to refine the reconstructed heat-flux profiles under clean input conditions. This stage focuses on improving spectral consistency and preserving rapid heat-flux variations. The model parameters are progressively inherited across stages, enabling stable optimization from basic numerical reconstruction to robust and spectral-aware heat-flux reconstruction. V Experiments V-A Dataset, Metric, and Baselines Dataset. We construct DivMPS2HF, a multi-source time-series dataset for reconstructing divertor heat-flux profiles from macroscopic plasma-state signals. The dataset contains data from 7777 discharge shots, among which 6464 shots are used for training, 33 shots for validation, and 1010 shots for testing. To avoid cross-shot information leakage, the dataset is split at the discharge-shot level, ensuring that samples from the same discharge shot do not appear in different subsets. Before constructing the dataset, all selected signals are aligned to a unified temporal grid with a sampling rate of 1kHz1~kHz. The selected signals have heterogeneous sampling rates: most signals are recorded around 1kHz1~kHz, while some diagnostic signals are sampled at either lower or much higher frequencies. For example, αD_α-related signals can be sampled at up to 250kHz250~kHz. To obtain synchronized multi-source inputs, we resample all signals to the common 1kHz1~kHz temporal grid before model training and evaluation. Signals with higher sampling rates are downsampled, while signals with lower sampling rates are mapped to the aligned time grid using temporal resampling. After temporal alignment, samples are generated on the fly during data loading by applying a sliding window of T=500T=500 time steps to each discharge shot, corresponding to 0.5s0.5~s at 1kHz1~kHz. A stride of one time step is used to densely capture the temporal evolution of the heat-flux profiles. This sampling procedure yields approximately 150,504150,504, 11,83811,838, and 28,98228,982 windows for training, validation, and testing, respectively. The model input consists of 67 channels of macroscopic plasma-state signals, which include not only traditional diagnostic measurements but also plasma state diagnostic signals, equilibrium and configuration parameters and signals related to external actuators and device control. These signals characterize the global and boundary plasma states related to the evolution of divertor heat load during the discharge process. The reconstruction target is the time-resolved heat-flux profile on the divertor target. For each time step, the heat-flux profile is represented as a 116-dimensional vector, corresponding to 116 sampling channels in the radial direction of the target. The reference heat-flux profiles are obtained from post-shot heat-flux analysis and are used as supervised labels for model training and evaluation. It should be emphasized that infrared images and heat-conduction inversion results are not used as model inputs. During inference, SafeDivertor only takes macroscopic plasma-state signals as input and reconstructs the target heat-flux profile through a direct neural forward pass. The DivMPS2HF dataset will be made available upon reasonable request for academic research purposes. Metrics. We evaluate the reconstructed divertor heat-flux profiles using five metrics: Mean Squared Error (MSE), Mean Absolute Error (MAE), Structural Similarity (SSIM) [33], Log-Spectral Distance (LSD) [34], and Log-Spectral Distance of High Frequency (LSD-HF). MSE and MAE are used as standard point-wise error metrics to evaluate the point-wise numerical accuracy of reconstructed heat flux in the time domain. Given N test samples, the reconstructed heat-flux profiles ^∈ℝN×O×T Y ^N× O× T and the ground-truth heat-flux profiles ∈ℝN×O×TY ^N× O× T, MSE and MAE are defined as follows: MSE=1NOT∑n=1N∑o=1O∑t=1T(n,o,t−^n,o,t)2,MSE= 1NOT _n=1^N _o=1^O _t=1^T (Y_n,o,t- Y_n,o,t )^2, (13) MAE=1NOT∑n=1N∑o=1O∑t=1T|n,o,t−^n,o,t|.MAE= 1NOT _n=1^N _o=1^O _t=1^T |Y_n,o,t- Y_n,o,t |. (14) Here, O=116O=116 denotes the number of target heat-flux channels, and T denotes the length of the time window. SSIM is used to evaluate the structural similarity of reconstructed spatiotemporal heat flux maps, taking into account both temporal evolution and radial profile distribution. For each sample, the heat-flux profile is treated as a two-dimensional map over radial channels and time steps. We first compute SSIM on each sample and then average the scores over all test samples: SSIM=1N∑n=1NSSIM(^n,n),SSIM= 1N _n=1^NSSIM ( Y_n,Y_n ), (15) where N denotes the number of test samples. For the n-th sample, the sample-level SSIM is computed by averaging the local SSIM values over all evaluated 7×77× 7 blocks: SSIM(^n,n)=1K∑k=1KSSIMn,k,SSIM ( Y_n,Y_n )= 1K _k=1^KSSIM_n,k, (16) where K is the number of evaluated local blocks in each heat-flux map. The local SSIM of the k-th block is defined as SSIMn,k=(2μn,kμ^n,k+ϵ1)(2Cov(n,k,^n,k)+ϵ2)(μn,k2+μ^n,k2+ϵ1)(σn,k2+σ^n,k2+ϵ2).SSIM_n,k= (2 _Y_n,k _ Y_n,k+ _1)(2Cov(Y_n,k, Y_n,k)+ _2)( _Y_n,k^2+ _ Y_n,k^2+ _1)( _Y_n,k^2+ _ Y_n,k^2+ _2). (17) Here, n,kY_n,k and ^n,k Y_n,k denote the k-th local 7×77× 7 block of the ground-truth and reconstructed heat-flux maps for the n-th test sample, respectively. μ, σ2σ^2, and Cov(⋅,⋅)Cov(·,·) denote the local mean, variance, and covariance, and ϵ1 _1, ϵ2 _2 are small constants for numerical stability. Let S and S denote the power spectra of the ground-truth and reconstructed heat-flux profiles after STFT: =|STFT()|2+ϵ,^=|STFT(^)|2+ϵ,S=|STFT(Y)|^2+ε, S=|STFT( Y)|^2+ε, (18) where ϵε is a small constant for numerical stability. For each evaluated spectrum, we first compute the log-spectral distance: dn,o,m=1F∑f=1F(log^n,o,f,m−logn,o,f,m)2,d_n,o,m= 1F _f=1^F ( S_n,o,f,m- _n,o,f,m )^2, (19) where F denotes the number of frequency bins. The final LSD score is obtained by averaging over all test samples, target channels, and STFT frames: LSD=1NOM∑n=1N∑o=1O∑m=1Mdn,o,m,LSD= 1NOM _n=1^N _o=1^O _m=1^Md_n,o,m, (20) where N is the number of test samples, O is the number of target heat-flux channels, and M is the number of STFT frames. Similarly, LSD-HF is computed by restricting the frequency bins to the high-frequency set ℱHFF_HF: dn,o,mHF=1|ℱHF|∑f∈ℱHF(log^n,o,f,m−logn,o,f,m)2,d^HF_n,o,m= 1|F_HF| _f _HF ( S_n,o,f,m- _n,o,f,m )^2, (21) LSD-HF=1NOM∑n=1N∑o=1O∑m=1Mdn,o,mHF.LSD -HF= 1NOM _n=1^N _o=1^O _m=1^Md^HF_n,o,m. (22) All metrics are computed only on the 116 target heat-flux channels. Baselines. We select representative state-of-the-art time series models from recent years as benchmarks for comparison. Specifically, we include general multivariate time series forecasting models, including Sonnet [27], Timer-XL [20], TimeFilter [22], SimpleTM [21], TimePerceiver [26], and iTransformer [19]. Furthermore, since our task naturally involves using multi-source signals as external conditions to reconstruct the target heat flux sequence, we also include exogenous variable prediction models, including CrossLinear [35] and TimeXer [28]. Additionally, we use T1 [30] as the time series imputation benchmark, where multi-source signal channels are treated as observed variables, and heat flux channels are treated as structured missing variables. V-B Implementation Details We construct DivMPS2HF, a multi-source time-series dataset for reconstructing divertor heat-flux profiles from macroscopic plasma-state signals. All models are evaluated using five metrics: MSE, MAE, SSIM, LSD, LSD-HF, which measure point-wise numerical accuracy, structural similarity, and spectral consistency of the reconstructed heat-flux profiles. During training and evaluation, fixed-length windows are sampled from the corresponding discharge shots. The input window length is set to T=500T=500 time steps, corresponding to a 0.50.5-second discharge segment under the unified 1kHz1~kHz sampling rate. The experiments are conducted on eight NVIDIA RTX 40904090 GPUs. For SafeDivertor, the Gaussian perturbation ratio is set to 20%20\%, and the noise standard deviation is set to 0.030.03. The channel-level heat-flux prior is computed only from the training set and kept fixed during validation and testing. For efficiency evaluation, all models are tested with batch size 11 and input length 500500, and the reported inference latency only measures the model forward process. V-C Main Results As shown in Table I, SafeDivertor achieves the best performance across all evaluation metrics compared with representative time-series baselines. Among the baselines, T1 [30] shows the strongest overall performance in numerical accuracy and spectral consistency, achieving the best baseline MSE, MAE, LSD, and LSD-HF, while TimeFilter [22] obtains the highest baseline SSIM. TABLE I: Overall comparison with representative time-series baselines on the DivMPS2HF dataset. Models MSE↓ MAE↓ SSIM↑ LSD↓ LSD-HF↓ T1 0.262 0.333 0.848 3.521 3.829 Sonnet 0.382 0.414 0.804 7.294 7.418 Timer-XL 0.307 0.367 0.824 4.718 5.420 TimeFilter 0.279 0.354 0.854 3.868 4.573 CrossLinear 0.378 0.427 0.706 5.101 4.976 SimpleTM 0.540 0.499 0.760 8.033 8.153 TimePerceiver 0.437 0.426 0.814 4.546 5.359 TimeXer 0.538 0.488 0.769 4.876 5.228 iTransformer 0.469 0.446 0.792 7.893 8.136 SafeDivertor 0.200 0.291 0.870 2.475 2.729 In contrast, several forecasting and exogenous-variable forecasting models show limited performance, indicating that directly applying existing time-series forecasting paradigms is insufficient for reconstructing target heat-flux channels from heterogeneous macroscopic plasma-state signals. Compared with the best baseline result for each metric, SafeDivertor reduces MSE from 0.2620.262 to 0.2000.200, MAE from 0.3330.333 to 0.2910.291, LSD from 3.5213.521 to 2.4752.475, and LSD-HF from 3.8293.829 to 2.7292.729, while improving SSIM from 0.8540.854 to 0.8700.870. These results demonstrate that the proposed task-driven designs help SafeDivertor better preserve numerical accuracy, spatiotemporal structure, and spectral characteristics of divertor heat-flux profiles. V-D Ablation Study Table I presents the ablation results of the proposed task-driven components, including physical prior-aware initialization (PPI), input perturbation (IP), spectral-aware reconstruction optimization (SRO), and progressive training (PT). The baseline model without these components achieves an MSE of 0.2620.262, MAE of 0.3330.333, SSIM of 0.8480.848, LSD of 3.5213.521, and LSD-HF of 3.8293.829. TABLE I: Ablation study of the main components in SafeDivertor on the DivMPS2HF dataset. PPI IP SRO PT MSE↓ MAE↓ SSIM↑ LSD↓ LSD-HF↓ #1 ✗ ✗ ✗ ✗ 0.262 0.333 0.848 3.521 3.829 #2 ✓ ✗ ✗ ✗ 0.211 0.293 0.870 3.617 4.141 #3 ✗ ✓ ✗ ✗ 0.230 0.313 0.860 3.725 4.205 #4 ✗ ✗ ✓ ✗ 0.219 0.299 0.863 2.306 2.511 #5 ✓ ✓ ✓ ✗ 0.211 0.296 0.869 2.326 2.568 #6 ✓ ✓ ✓ ✓ 0.200 0.291 0.870 2.475 2.729 When each component is introduced individually, different performance trends can be observed. PPI reduces MSE from 0.2620.262 to 0.2110.211, MAE from 0.3330.333 to 0.2930.293, and improves SSIM from 0.8480.848 to 0.8700.870, indicating that the physical heat-flux prior provides useful statistical radial-distribution information for unavailable target channels. IP also improves the time-domain and structural metrics, reducing MSE to 0.2300.230 and improving SSIM to 0.8600.860, which shows that perturbing observable signal channels helps regularize the model’s dependence on specific signals. However, both PPI and IP lead to higher LSD and LSD-HF, because they do not explicitly supervise the spectral characteristics of reconstructed heat-flux sequences. In contrast, SRO contributes most significantly to spectral consistency, reducing LSD from 3.5213.521 to 2.3062.306 and LSD-HF from 3.8293.829 to 2.5112.511. This confirms the necessity of multi-scale STFT-based spectral supervision for preserving spectral characteristics and high-frequency transient variations. When PPI, IP, and SRO are jointly used, the model achieves a more balanced performance across time-domain, structural, and spectral metrics. This demonstrates that these components play complementary roles: PPI provides statistical heat-flux prior information, IP regularizes heterogeneous observable signals, and SRO enhances spectral consistency. After further introducing PT, the full model achieves the best MSE and MAE of 0.2000.200 and 0.2910.291, respectively, while maintaining competitive spectral performance. These results show that the proposed task-driven components improve complementary aspects of reconstruction quality. Figure 3: Qualitative visualization of four randomly selected samples from different discharge shots, showing the ground truth, SafeDivertor, and T1 from top to bottom. V-E Effect of Multi-scale Spectral Loss As shown in Table I, we analyze the role of spectral loss in the spectral-aware reconstruction optimization module. The w/o SRO setting represents a model trained using only the basic MSE reconstruction loss. Introducing single-scale STFT loss significantly improves the model across all metrics. This indicates that the spectral optimization objective helps the model better preserve the spectral characteristics and rapid changes in the heat-flux sequences. Compared to the single-scale setting, the multi-scale STFT loss function further improves most metrics. These results demonstrate that applying spectral constraints across multiple timescales is more effective than using a single scale, as different scales can capture both broader heat flux evolution patterns and local high-frequency variations. Overall, the multi-scale STFT loss function provides a better balance between numerical accuracy and spectral consistency. TABLE I: Ablation study of the multi-scale STFT loss. Metrics w/o SRO Single Scale Multi Scale MSE↓ 0.262 0.221 0.219 MAE↓ 0.333 0.301 0.299 SSIM↑ 0.848 0.867 0.863 LSD↓ 3.521 2.367 2.306 LSD-HF↓ 3.829 2.549 2.511 V-F Input Perturbation Strategy Analysis TABLE IV: Effect of different input perturbation strategies. CM and GP denote channel masking and Gaussian perturbation, respectively. w/o IP r=0.1 r=0.2 r=0.4 Metrics - CM GP CM GP CM GP MSE 0.262 0.251 0.229 0.288 0.230 0.276 0.203 MAE 0.333 0.329 0.316 0.345 0.313 0.339 0.297 SSIM 0.848 0.858 0.861 0.854 0.860 0.853 0.874 LSD 3.521 3.653 3.816 3.696 3.725 3.599 3.784 LSD-HF 3.829 4.093 4.315 4.096 4.205 4.013 4.163 We further evaluate different input perturbation strategies by applying channel masking and Gaussian perturbation to observable signal channels under different perturbation ratios r∈0.1,0.2,0.4r∈\0.1,0.2,0.4\. As shown in Table IV, channel masking provides limited and unstable improvements. Although it slightly improves MSE and SSIM at r=0.1r=0.1, larger masking ratios degrade the numerical metrics, suggesting that directly removing observable channels may discard useful plasma-state information. In contrast, Gaussian perturbation consistently improves numerical accuracy and structural similarity across different ratios, indicating that perturbing observable channels is more effective than completely masking them. Although both strategies degrade spectral performance, this is expected because they do not explicitly constrain frequency-domain characteristics. Overall, Gaussian perturbation is a more effective input perturbation strategy for improving numerical accuracy and structural similarity. Figure 4: Qualitative comparison of four samples, one from each of four test shots. Each column corresponds to one sample, while the rows show the ground truth and reconstruction results of SafeDivertor and the evaluated baselines. V-G Efficiency Analysis Table V compares the computational complexity and inference efficiency of different models. SafeDivertor requires 21.25521.255 GFLOPs and achieves an inference latency of 11.31311.313 ms per sample with a batch size of 11 and an input length of 500500. Since each input window corresponds to a 0.50.5-second discharge segment at 1kHz1~kHz, this latency indicates its potential for online-oriented window-level heat-flux reconstruction. Compared with the T1 backbone, SafeDivertor retains the same parameter scale, with GFLOPs increasing only from 21.24021.240 to 21.25521.255 and comparable latency, while substantially improving reconstruction performance. These results demonstrate a favorable performance-efficiency trade-off with negligible additional inference cost. TABLE V: Comparison of model parameters (M), computational cost (GFLOPs), and per-sample inference latency (ms). Model Params. GFLOPs Inference Sonnet 42.239 123.188 10.947 T1 12.783 21.240 11.023 Timer-XL 33.825 61.583 14.721 TimeFilter 11.011 35.118 261.343 CrossLinear 33.130 5.122 4.623 SimpleTM 34.753 21.629 6.126 TimePerceiver 16.032 14.824 2.530 TimeXer 51.088 44.684 4.161 iTransformer 34.635 2.390 1.732 SafeDivertor 12.783 21.255 11.313 V-H Visualization As shown in Figure 3, we compare the ground truth, SafeDivertor, and T1 on four samples. The original T1 backbone recovers the overall spatiotemporal heat-flux structure and main radial high-heat-flux bands, but its results are relatively over-smoothed and may underestimate local high-intensity regions, failing to preserve temporal fluctuations and fine-grained heat-flux variations. In contrast, SafeDivertor better preserves the main heat-flux structures while recovering clearer high-intensity bands and richer temporal details. These visual results are consistent with the quantitative improvements in Table I, further demonstrating the advantage of the proposed task-driven designs. Figure 4 presents qualitative comparisons on four samples from different test shots. Most baselines can identify the approximate radial locations of the dominant heat-flux bands, but their reconstruction characteristics vary considerably. Several methods produce over-smoothed profiles with limited temporal variation, while others exhibit noticeable intensity bias, discontinuous patterns, or horizontal artifacts. Although the original T1 backbone preserves the overall spatiotemporal structure relatively well, it still tends to underestimate local high-intensity regions and suppress fine-grained temporal fluctuations. In contrast, SafeDivertor more faithfully recovers the positions and intensities of the main heat-flux bands while retaining clearer transient variations across different discharge conditions. These qualitative results are consistent with the quantitative improvements reported in the main experiments, further demonstrating the effectiveness of the proposed task-driven designs. VI Conclusion In this paper, we introduce a new online-oriented paradigm for directly reconstructing divertor heat-flux profiles from macroscopic plasma-state signals, without using infrared surface-temperature measurements as model inputs. We also construct DivMPS2HF, a multi-source discharge dataset for this task, and propose SafeDivertor, a unified reconstruction framework equipped with physical prior-aware initialization, input perturbation, spectral-aware reconstruction optimization, and progressive training. Experiments on DivMPS2HF show that SafeDivertor achieves the best overall performance among the evaluated time-series baselines across all five metrics. Efficiency analysis further shows that SafeDivertor maintains inference efficiency comparable to T1, providing a favorable latency basis for future online deployment. Future work will expand the dataset, improve cross-scenario generalization, and further enhance reconstruction efficiency. References [1] Z. Yang, P. He, H. Yan, Z. Bin, S. Shuangbao, F. Wang, M. Chen, G. Jia, and X. Gong, “The development of a three-dimensional finite element method code for the heat flux analysis of tungsten monoblock divertor on east,” Fusion Engineering and Design, vol. 152, p. 111448, 2020. [2] R. A. Pitts, X. Bonnin, F. Escourbiac, H. Frerichs, J. Gunn, T. Hirai, A. Kukushkin, E. Kaveeva, M. Miller, D. Moulton et al., “Physics basis for the first iter tungsten divertor,” Nuclear Materials and Energy, vol. 20, p. 100696, 2019. [3] E. Aymerich, F. Pisano, B. Cannas, G. Sias, A. Fanni, Y. Gao, D. Böckenhoff, M. Jakubowski et al., “Physics informed neural networks towards the real-time calculation of heat fluxes at w7-x,” Nuclear Materials and Energy, vol. 34, p. 101401, 2023. [4] A. Herrmann, W. Junker, K. Gunther, S. Bosch, M. Kaufmann, J. Neuhauser, G. Pautasso, T. Richter, and R. Schneider, “Energy flux to the asdex-upgrade diverter plates determined by thermography and calorimetry,” Plasma Physics and Controlled Fusion, vol. 37, no. 1, p. 17–29, 1995. [5] H. Lee, R. Pitts, C. Kang, S. Oh, J. Bak, S. Hong, H. Wi, Y. Kim, H. Kim, D. Seo et al., “Thermographic studies of outer target heat fluxes on kstar,” Nuclear Materials and Energy, vol. 12, p. 541–547, 2017. [6] B. Shi, Z.-D. Yang, B. Zhang, C. Yang, K.-F. Gan, M.-W. Chen, J.-H. Yang, H. Zhang, J.-L. Qi, X.-Z. Gong et al., “Heat flux on east divertor plate in h-mode with lhcd/lhcd+ nbi,” Chinese Physics Letters, vol. 34, no. 9, p. 095201, 2017. [7] T. Looby, M. Reinke, D. Donovan, T. Gray, M. Messineo, and A. Khodak, “Convolutional neural networks for heat flux model validation on nstx-u,” IEEE Transactions on Plasma Science, vol. 48, no. 1, p. 3–13, 2019. [8] X. Wang, Z. Yan, H. Si, Z. Yang, Q. Yang, D. Sun, W. Lyu, and J. Tang, “Revisiting heat flux analysis of tungsten monoblock divertor on east using physics-informed neural network,” arXiv preprint arXiv:2508.03776, 2025. [9] J. Kates-Harbeck, A. Svyatkovskiy, and W. Tang, “Predicting disruptive instabilities in controlled fusion plasmas through deep learning,” Nature, vol. 568, no. 7753, p. 526–531, 2019. [10] B. Guo, B. Shen, D. Chen, C. Rea, R. Granetz, Y. Huang, L. Zeng, H. Zhang, J. Qian, Y. Sun et al., “Disruption prediction using a full convolutional neural network on east,” Plasma Physics and Controlled Fusion, vol. 63, no. 2, p. 025008, 2021. [11] W. Zheng, F. Xue, Z. Chen, D. Chen, B. Guo, C. Shen, X. Ai, N. Wang, M. Zhang, Y. Ding et al., “Disruption prediction for future tokamaks using parameter-based transfer learning,” Communications Physics, vol. 6, no. 1, p. 181, 2023. [12] F. Matos, V. Menkovski, F. Felici, A. Pau, F. Jenko, T. Team, and E. M. Team, “Classification of tokamak plasma confinement states with convolutional recurrent neural networks,” Nuclear Fusion, vol. 60, no. 3, p. 036022, 2020. [13] H. Si, X. Wang, F. Zhang, X. Zhou, D. Sun, W. Lyu, Q. Yang, and J. Tang, “Hgts-former: Hierarchical hypergraph transformer for multivariate time series analysis,” arXiv preprint arXiv:2508.02411, 2025. [14] S. Joung, D. R. Smith, G. McKee, Z. Yan, K. Gill, J. Zimmerman, B. Geiger, R. Coffee, F. O’Shea, A. Jalalvand et al., “Tokamak edge localized mode onset prediction with deep neural network and pedestal turbulence,” Nuclear Fusion, vol. 64, no. 6, p. 066038, 2024. [15] M. Wang, C. Wan, J. Lu, Z. Yu, B. Xiao, Y. Li, X. He, Z. Luo, Q. Yuan, Y. Hu et al., “Time series extrinsic regression for reconstructing missing electron temperature in tokamak,” Nuclear Fusion, vol. 65, no. 7, p. 076008, 2025. [16] X. Song, L. Deng, H. Wang, Y. Zhang, Y. He, and W. Cao, “Deep learning-based time series forecasting: X. song et al.” Artificial Intelligence Review, vol. 58, no. 1, p. 23, 2024. [17] J. Kim, H. Kim, H. Kim, D. Lee, and S. Yoon, “A comprehensive survey of deep learning for time series forecasting: architectural diversity and open challenges,” Artificial Intelligence Review, vol. 58, no. 7, p. 216, 2025. [18] J. F. Torres, D. Hadjout, A. Sebaa, F. Martínez-Álvarez, and A. Troncoso, “Deep learning for time series forecasting: a survey,” Big data, vol. 9, no. 1, p. 3–21, 2021. [19] Y. Liu, T. Hu, H. Zhang, H. Wu, S. Wang, L. Ma, and M. Long, “itransformer: Inverted transformers are effective for time series forecasting,” in International conference on learning representations, vol. 2024, 2024, p. 11 116–11 140. [20] Y. Liu, G. Qin, X. Huang, J. Wang, and M. Long, “Timer-xl: Long-context transformers for unified time series forecasting,” in International Conference on Learning Representations, vol. 2025, 2025, p. 83 982–84 006. [21] H. Chen, V. Luong, L. Mukherjee, and V. Singh, “SimpleTM: A simple baseline for multivariate time series forecasting,” in The Thirteenth International Conference on Learning Representations, 2025. [22] Y. Hu, G. Zhang, P. Liu, D. Lan, N. Li, D. Cheng, T. Dai, S.-T. Xia, and S. Pan, “Timefilter: Patch-specific spatial-temporal graph filtration for time series forecasting,” in Forty-second International Conference on Machine Learning, 2025. [23] K. Yi, Q. Zhang, W. Fan, S. Wang, P. Wang, H. He, N. An, D. Lian, L. Cao, and Z. Niu, “Frequency-domain mlps are more effective learners in time series forecasting,” Advances in neural information processing systems, vol. 36, p. 76 656–76 679, 2023. [24] T. Zhou, Z. Ma, Q. Wen, L. Sun, T. Yao, W. Yin, R. Jin et al., “Film: Frequency improved legendre memory model for long-term time series forecasting,” Advances in neural information processing systems, vol. 35, p. 12 677–12 690, 2022. [25] H. Wu, T. Hu, Y. Liu, H. Zhou, J. Wang, and M. Long, “Timesnet: Temporal 2d-variation modeling for general time series analysis,” in The Eleventh International Conference on Learning Representations, 2023. [26] J. Lee and H. Lee, “Timeperceiver: An encoder-decoder framework for generalized time-series forecasting,” Advances in Neural Information Processing Systems, vol. 38, p. 136 136–136 165, 2026. [27] Y. Shu and V. Lampos, “Sonnet: Spectral operator neural network for multivariable time series forecasting,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 30, 2026, p. 25 419–25 427. [28] Y. Wang, H. Wu, J. Dong, G. Qin, H. Zhang, Y. Liu, Y. Qiu, J. Wang, and M. Long, “Timexer: Empowering transformers for time series forecasting with exogenous variables,” Advances in Neural Information Processing Systems, vol. 37, p. 469–498, 2024. [29] Z. Li, X. Qiu, Y. Zhu, X. Wu, J. Hu, C. Guo, and B. Yang, “GCGNet: Graph-consistent generative network for time series forecasting with exogenous variables,” in The Fourteenth International Conference on Learning Representations, 2026. [30] D. Park, H. Ryu, S. Bae, K. Park, and H.-S. Kim, “T1: One-to-one channel-head binding for multivariate time-series imputation,” in The Fourteenth International Conference on Learning Representations, 2026. [31] W. Du, D. Côté, and Y. Liu, “Saits: Self-attention-based imputation for time series,” Expert Systems with Applications, vol. 219, p. 119619, 2023. [32] W. Cao, D. Wang, J. Li, H. Zhou, L. Li, and Y. Li, “Brits: Bidirectional recurrent imputation for time series,” Advances in neural information processing systems, vol. 31, 2018. [33] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing, vol. 13, no. 4, p. 600–612, 2004. [34] C. Li, Z. Chen, L. Wang, and J. Zhu, “Audio super-resolution with latent bridge models,” Advances in Neural Information Processing Systems, vol. 38, p. 48 593–48 636, 2026. [35] P. Zhou, Y. Liu, J. Liang, Q. Song, and X. Li, “Crosslinear: Plug-and-play cross-correlation embedding for time series forecasting with exogenous variables,” in Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, 2025, p. 4120–4131.