Paper deep dive
Suppression of $^{14}\mathrm{C}$ photon hits in large liquid scintillator detectors via spatiotemporal deep learning
Junle Li, Zhaoxiang Wu, Guanda Gong, Zhaohan Li, Wuming Luo, Jiahui Wei, Wenxing Fang, Hehe Fan
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/31/2026, 2:08:06 AM
Summary
The paper proposes three deep learning models—Gated-STGNN, STT-Scalar, and STT-Vector—to suppress 14C beta decay background in liquid scintillator neutrino detectors like JUNO. By treating photon hits as spatiotemporal point clouds, these models achieve 25%-48% recall for 14C hits while keeping positron misidentification below 1%, significantly improving energy resolution in pile-up events.
Entities (6)
Relation Signals (4)
JUNO → uses → Liquid Scintillator Detector
confidence 100% · JUNO is a 20-kton LS detector
Gated-STGNN → suppresses → 14C
confidence 90% · we propose three models to tag 14C photon hits... thereby suppressing its impact
STT-Scalar → suppresses → 14C
confidence 90% · we propose three models to tag 14C photon hits... thereby suppressing its impact
STT-Vector → suppresses → 14C
confidence 90% · we propose three models to tag 14C photon hits... thereby suppressing its impact
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Liquid scintillator detectors are widely used in neutrino experiments due to their low energy threshold and high energy resolution. Despite the tiny abundance of $^{14}$C in LS, the photons induced by the $\beta$ decay of the $^{14}$C isotope inevitably contaminate the signal, degrading the energy resolution. In this work, we propose three models to tag $^{14}$C photon hits in $e^+$ events with $^{14}$C pile-up, thereby suppressing its impact on the energy resolution at the hit level: a gated spatiotemporal graph neural network and two Transformer-based models with scalar and vector charge encoding. For a simulation dataset in which each event contains one $^{14}$C and one $e^+$ with kinetic energy below 5 MeV, the models achieve $^{14}$C recall rates of 25%-48% while maintaining $e^+$ to $^{14}$C misidentification below 1%, leading to a large improvement in the resolution of total charge for events where $e^+$ and $^{14}$C photon hits strongly overlap in space and time.
Tags
Links
- Source: https://arxiv.org/abs/2603.27727v1
- Canonical: https://arxiv.org/abs/2603.27727v1
Trouble viewing inline? Open PDF directly →
Full Text
58,263 characters extracted from source content.
Expand or collapse full text
Suppression of C14^14C photon hits in large liquid scintillator detectors via spatiotemporal deep learning Junle Li School of Aeronautics and Astronautics, Zhejiang University, Hangzhou 310027, China Zhaoxiang Wu Institute of High Energy Physics, Chinese Academy of Sciences, Beijing 100049, China School of Physical Sciences, University of Chinese Academy of Science, Beijing 100049, China Guanda Gong Institute of High Energy Physics, Chinese Academy of Sciences, Beijing 100049, China Zhaohan Li Corresponding author: lizhaohan@ihep.ac.cn Institute of High Energy Physics, Chinese Academy of Sciences, Beijing 100049, China Wuming Luo Corresponding author: luowm@ihep.ac.cn Institute of High Energy Physics, Chinese Academy of Sciences, Beijing 100049, China Jiahui Wei Institute of High Energy Physics, Chinese Academy of Sciences, Beijing 100049, China Wenxing Fang Institute of High Energy Physics, Chinese Academy of Sciences, Beijing 100049, China Hehe Fan Corresponding author: hehefan@zju.edu.cn College of Computer Science and Technology, Zhejiang University, Hangzhou 310027, China (March 29, 2026) Abstract Liquid scintillator detectors are widely used in neutrino experiments due to their low energy threshold and high energy resolution. Despite the tiny abundance of 14C in LS, the photons induced by the β decay of the 14C isotope inevitably contaminate the signal, degrading the energy resolution. In this work, we propose three models to tag 14C photon hits in e+ events with 14C pile-up, thereby suppressing its impact on the energy resolution at the hit level: a gated spatiotemporal graph neural network and two Transformer-based models with scalar and vector charge encoding. For a simulation dataset in which each event contains one C14^14C and one e+ with kinetic energy below 5 MeV, the models achieve C14^14C recall rates of 25%–48% while maintaining e+e^+ to C14^14C misidentification below 1%, leading to a large improvement in the resolution of total charge for events where e+e^+ and C14^14C photon hits strongly overlap in space and time. I Introduction Pile-up, namely the overlap of signal and background processes within the same data acquisition window, is a well-known challenge in particle physics experiments. It can induce fake signal yields or degrade detector resolution [1, 2, 3, 4, 5], thereby affecting physics sensitivity. With the advantages of low energy threshold and high energy resolution, liquid scintillator (LS) detectors [6, 7, 8, 9, 10, 11] are widely used in neutrino experiments. Regardless of the LS recipe, carbon atoms are usually the dominant component [12]. Despite its tiny abundance in LS, the 14C isotope decays and induces scintillation photons. When these photons fall within the same data acquisition window as a signal particle, they inevitably contaminate the signal photons. This so-called 14C pile-up effect worsens the energy resolution or introduces background and is a common challenge for all LS detectors [1, 4, 13, 14]. This effect becomes even more critical for large-scale detectors such as the Jiangmen Underground Neutrino Observatory (JUNO) [13, 15], where the large target mass results in a more frequent occurrence of pile-up events. Meanwhile the energy resolution is crucial for high-precision neutrino oscillation measurements. While hit-level pile-up suppression has been extensively developed for jet reconstruction at the Large Hadron Collider (LHC) [16, 17, 18, 19, 20, 21], comparable approaches are less explored in LS detectors. In this paper, we investigate the feasibility of utilizing deep learning methods to identify and suppress C14^14C photon hits in order to mitigate their impact on energy resolution. We focus specifically on C14^14C pile-up events that contain only a single C14^14C pile-up, given that multiple C14^14C pile-up has a much lower probability of occurring and a more complex event topology. The JUNO experiment is used as a representative case study. JUNO is a 20-kton LS detector designed primarily to determine the neutrino mass ordering, and it requires an unprecedented energy resolution of 3%3\% at 1MeV1\,MeV [22]. This performance relies on high optical coverage, provided by approximately 17,600 20-inch and 25,600 3-inch photomultiplier tubes (PMTs) [23, 13], an optimized LS recipe to achieve high light yield as well as excellent transparency [24], extensive PMT R&D to increase the quantum efficiency for photon detection, and advanced reconstruction algorithms [23, 25, 26, 27]. Detailed information on JUNO can be found in Refs. [24, 13]. JUNO detects reactor ν¯e ν_e via inverse beta decay (IBD), ν¯e+p→n+e+ ν_e+p→ n+e^+, the average deposited energy of a positron is at the MeV scale. In contrast, the peak and endpoint energies of the 14C beta decay are 0.04MeV0.04\,MeV and 0.16MeV0.16\,MeV, respectively, yielding substantially fewer scintillation photons. As a result, when the two processes overlap in time, the 14C photon contribution is embedded in the dominant positron signal and becomes difficult to distinguish at the hit level. (a) Large separation (Δt=455.0ns t=455.0~ns) (b) Small separation (Δt=2.6ns t=2.6~ns) Figure 1: Hit time distributions of pile-up events. The curves show the contributions from e+e^+ annihilation (blue), C14^14C decay (red), and dark noise (gray), along with the total hit count (dashed black). (a) With large Δt t, the C14^14C signal appears as a weak secondary peak. (b) With small Δt t, the C14^14C hits are buried within the dominant e+e^+ peak. 100 – 200 ns (a) 200 – 300 ns (b) 300 – 400 ns (c) 400 – 500 ns (d) 500 – 600 ns (e) (f) Figure 2: The spatiotemporal structure of a pile-up event with Δt=1.9ns t=1.9~ns. The truth information for the e+e^+ event consists of a kinetic energy Ek(e+)=1.4MeVE_k(e^+)=1.4~MeV and a vertex position of (x,y,z)=(−1.9,5.7,11.1)m(x,y,z)=(-1.9,5.7,11.1)~m, while that for the C14^14C event consists of a deposited energy EC14=107.4keVE_^14C=107.4~keV and a vertex position of (x,y,z)=(−5.4,1.9,−14.8)m(x,y,z)=(-5.4,1.9,-14.8)~m. The plots show the spatial distribution of photon hits in 100ns100~ns intervals. Blue dots represent hits from the positron annihilation, red dots from the C14^14C decay, and gray dots indicate dark noise. The identification of C14^14C photon hits can be viewed as a semantic segmentation task on a sparse point cloud defined on a spherical detector geometry. Compared to point cloud segmentation problems, this task is complicated by the sparsity of the data and the discontinuous distribution of photon hits in both space and time. To address these challenges, we investigate two classes of deep learning architectures: a graph-based model implemented as a Gated Spatiotemporal Graph Neural Network (Gated-STGNN)[28], and two Transformer-based models [29]: the Spatiotemporal Transformer with Scalar Charge Encoding (STT-Scalar) , which uses the charge of each PMT hit directly, and the Spatiotemporal Transformer with Vector Charge Encoding (STT-Vector), which incorporates charge information aggregated from neighboring hits. The graph-based approach emphasizes local spatiotemporal correlations through structured neighborhood aggregation, while the Transformer framework captures global dependencies via self-attention mechanisms. Their performance is evaluated based on hit identification efficiency and the resulting impact on the reconstructed energy resolution. The paper is organized as follows. Section I describes the simulation and data representation. Section I details the proposed graph and attention-based architectures. The training protocol and evaluation metrics are presented in Section IV. Finally, Section V discusses the classification performance and its impact on energy resolution, followed by conclusions in Section VI. I Simulation and Dataset Data samples are simulated based on a configuration similar to the JUNO central detector [13], including PMT positions, LS properties, etc. Various electronic effects, such as PMT transit-time smearing, PMT dark noise and PMT single-photoelectron charge smearing, were implemented in a toy electronics simulation. Moreover, as mentioned in the Introduction, the simulation was customized such that one C14^14C pile-up was enforced for every e+ event. With a light yield of 16651665 photon hits per MeV from a previous study [23], an e+e^+ at rest ( 0MeV kinetic energy) deposits around 1.022MeV in the LS and induces approximately 1702 photon hits recorded by the PMTs, with a statistical fluctuation of 41 photons. By contrast, a C14^14C decay produces around 100 photons on average, which are easily overwhelmed by the e+e^+ photon hits. In addition, given an average rate of 22.7 kHz and 17,600 PMTs, PMT dark noise contributes about 400 photon hits within the acquisition window of each event, making the identification of C14^14C photons more challenging. The acquisition window of an event is 1000ns. Each event contains photon hits induced by three components: one IBD positron (e+e^+), one C14^14C β decay, and uniformly distributed dark noise (DN). For the i-th hit, the data vector is defined as i=(xi,yi,zi,ti,qi,ℓi),s_i=(x_i,y_i,z_i,t_i,q_i, _i), (1) where (x,y,z)(x,y,z) are the spatial coordinates of the triggered PMT [13], t is the arrival time relative to the start of the acquisition window, and q is the charge. Here ℓi∈0,1,2 _i∈\0,1,2\ denotes the truth label of the hit, representing it originating from DN, e+e^+, and C14^14C, respectively. The photon hit time distributions of two example events are shown in Fig. 1. The difficulty of hit-level discrimination depends strongly on the relative difference between the C14^14C decay time (tC14t_^14C) and the e+e^+ time (te+t_e^+), defined as Δt≡tC14−te+. t≡ t_^14C-t_e^+. (2) For an event with large |Δt|| t|, as shown in Fig. 1a, the two components are temporally separated and background C14^14C hits can be rejected using simple timing cuts [23]. For events with Δt∈[−100,300]ns t∈[-100,300]\,ns, the C14^14C component is embedded within the dominant e+e^+ peak, rendering timing information alone insufficient, as indicated in Fig. 1b. A joint analysis of spatial and temporal correlations is therefore required. The spatiotemporal structure of a typical event in this regime is shown in Fig. 2. This study focuses on events with Δt∈[−100,300]ns t∈[-100,300]\,ns, excluding other events. Events at a few-MeV scale contain relatively few photon hits. Thus, only a small fraction of all PMTs are triggered and register photon hits. Moreover, the majority of these PMTs will detect no more than 4 photons in the entire 1000 ns acquisition window. Therefore, the data structure manifests inherent sparsity and spatiotemporal discontinuity. This is very different from the scenario of conventional point cloud tasks [30, 31, 32, 33, 34]. The samples used in this study are as follows. Training sample: Models are trained on a sample of one million simulated events with continuous positron kinetic energy Ek(e+)∈[0,5]MeVE_k(e^+)∈[0,5]~MeV and vertices distributed uniformly within the LS Test sample: Evaluation is conducted on 24 independent test sets, each containing 8×1048× 10^4 events. These data sets correspond to the combinations of six discrete positron kinetic energies, Ek(e+)∈0,1,2,3,4,5MeV,E_k(e^+)∈\0,1,2,3,4,5\~MeV, (3) and four exemplary vertex positions along the detector z-axis: z∈0, 6, 15, 16.5m.z∈\0,\,6,\,15,\,16.5\~m. (4) The datasets at 0 and 6 m characterize model performance in the detector’s inner volume. To probe performance near the detector boundary, we also consider datasets at larger radii. In particular, the JUNO central detector exhibits a total reflection region, approximately defined by r ∈[15.6,17.7]m∈[15.6,17.7]\,m [13]. The dataset at 15 m probes events close to the boundary of this region, while the dataset at 16.5 m corresponds to events well within the total reflection regime and serves as a representative sample. I Methodology I.1 Gated Spatiotemporal Graph Neural Network (Gated-STGNN) LHC uses Gated-GNN [28] to suppress the pile-up effect in the jet reconstruction [19, 18, 35]. In this work, we adapt the Gated-GNN to the C14^14C hit-level tagging task and refer to it as Gated-STGNN. Gated-STGNN models events as graphs where nodes represent hits and edges encode local spatiotemporal proximity. The architecture employs message passing for contextual propagation and a per-hit classifier for the three target classes. The Gated-STGNN serves as the baseline model for this study. Input representation and batching. For an event with N hits, the input feature vector iu_i is constructed from the hit vector is_i defined in Eq. (1). The spatial coordinates i=(xi,yi,zi)r_i=(x_i,y_i,z_i) are mapped to the unit direction vector ^i=i/∥i∥ x_i=r_i/ _i . The hit time tit_i is transformed into a normalized continuous time t~i∈[0,1) t_i∈[0,1) within the acquisition window [tmin,tmax)[t_ ,t_ ) to preserve the precise timing required for microscopic hit-level correlations. To simultaneously capture the macroscopic temporal evolution of the event, tit_i is also mapped to a normalized coarse-grained time-bin index b~i∈[0,1] b_i∈[0,1] using a bin width Δtbin t_bin. The charge qiq_i is used directly. The resulting input vector for the i-th hit is given by i=(^i,qi,t~i,b~i)∈ℝ6.u_i=( x_i,q_i, t_i, b_i) ^6. (5) The number of hits varies across events; however, parallel processing on GPUs requires a uniform sequence length for all events within a mini-batch. Events are therefore zero-padded so that the number of hits in every event equals Lmax=maxmNmL_ = _mN_m, where NmN_m denotes the valid hit number of the m-th event in the batch. Since zero-padding can interfere with the hit identification task, a binary mask M is employed to strictly exclude padded entries from the graph construction and message passing operations: Mmj=1,if j≥Nm(padding),0,otherwise(valid),M_mj= cases1,&if j≥ N_m (padding),\\ 0,&otherwise (valid), cases (6) where the index m denotes the event within the mini-batch, and j∈[0,Lmax−1]j∈[0,L_ -1] represents the hit index within that specific event. Window-level context. To capture the global temporal evolution, aggregate statistics—specifically the hit number, the sum and mean of the charges, alongside the mean and variance of the unit direction vectors—are computed for each discrete time bin b to serve as the bin feature vector bw_b. The sequence of bin features bw_b is then processed by a Gated Recurrent Unit (GRU) to generate hidden states bc_b [36]. Each hit i is subsequently assigned the state bic_b_i corresponding to its specific time bin, where the integer bin index is given by bi=⌊(ti−tmin)/Δtbin⌋b_i= (t_i-t_ )/ t_bin . The initial hit embedding is then obtained by applying an input multilayer perceptron (MLP) to the concatenated hit and context features: i(0)=MLPin([i;bi])h_i^(0)=MLP_in([u_i;c_b_i]). Graph construction and edge features. For each hit i, a localized spatiotemporal neighborhood is strictly defined as the set i=j≠i∣|tj−ti|≤Δtnb,∥j−i∥≤ΔrnbN_i=\j≠ i |t_j-t_i|≤ t_nb, _j-x_i ≤ r_nb\. To maintain a fixed graph degree for batched processing, a multi-set iS_i of exactly K indices is constructed by sampling uniformly without replacement from iN_i. If |i|<K|N_i|<K, the sequence is padded with self-loops (j=ij=i). The edge features ije_ij are computed for each j∈ij _i and encompass the Fourier-encoded absolute time difference, the angular alignment cosθij _ij, the normalized spatial separation, the charge difference, and the normalized signed time difference. Gated message passing. The model applies L message-passing layers. At layer ℓ , messages are computed via an MLP on the neighbor states alongside the corresponding edge features, and then aggregated by a mean operator over the K sampled nodes: ¯i(ℓ)=1K∑j∈iMLPmsg(ℓ)([j(ℓ);ij]) m_i^( )= 1K _j _iMLP_msg^( )([h_j^( );e_ij]). The node state is subsequently updated using a GRU cell: i(ℓ+1)=GRUCell(¯i(ℓ),i(ℓ))h_i^( +1)=GRUCell( m_i^( ),h_i^( )). Output and implementation. The final latent representations i(L)∈ℝHh_i^(L) ^H, where H denotes the hidden feature dimension, are projected onto the decision space via a position-wise output MLP, yielding a vector of class logits i∈ℝ3s_i ^3. The three components of is_i correspond to the confidence scores for the dark noise, positron (e+e^+), and C14^14C categories, respectively. The baseline configuration utilizes H=128H=128, L=3L=3 message-passing layers, a neighborhood size of K=24K=24, spatial and temporal cutoffs of Δrnb=9000m r_nb=9000~m and Δtnb=10ns t_nb=10~ns, and a bin width of Δtbin=20ns t_bin=20~ns over the acquisition window [0,1000)ns[0,1000)~ns. This yields a compact model with approximately 0.66 million trainable parameters. Training is performed using the AdamW optimizer [37] with cosine annealing and mixed-precision arithmetic. I.2 Spatiotemporal Transformer with Scalar Charge Encoding (STT-Scalar) The STT-Scalar treats hits as tokens in a sequence, utilizing global self-attention to facilitate information exchange across hits for per-hit classification. Input representation. For batch size B, the input tensor ∈ℝB×Lmax×5X ^B× L_ × 5 comprises observables (x,y,z,t,q)(x,y,z,t,q) from Eq. (1). Similar to the Gated-STGNN, sequences are padded to LmaxL_ and processed using the binary mask M defined in Eq. (6) to prevent padded tokens from attending to valid signals in the self-attention mechanism. Spatiotemporal feature encoding. Position and time are normalized by Sx=104mmS_x=10^4~m and St=103nsS_t=10^3~ns, respectively, and mapped using periodic encodings ψ(⋅)ψ(·) with npe=5n_pe=5: ψ(u)=(u,sin(kπu),cos(kπu)k=1npe).ψ(u)= (u,\;\ (kπ u),\; (kπ u)\_k=1^n_pe ). (7) The scalar charge q is embedded via an MLP to q∈ℝ32e_q ^32. These features are concatenated and projected to dimension dmodeld_model: (0)=Proj([ψ(~);ψ(t~);q])∈ℝdmodel.h^(0)=Proj ( [ψ( x);\;ψ( t);\;e_q ] ) ^d_model. (8) Transformer backbone and output. A stack of L pre-norm Transformer encoder layers (ℓ)T^( ) processes the embeddings: (ℓ+1)=(ℓ)((ℓ)),ℓ=0,…,L−1.h^( +1)=T^( )\! (h^( ) ), =0,…,L-1. (9) The final hidden representation of each hit is passed through a multilayer perceptron to produce a vector of classification logits i∈ℝ3s_i ^3, corresponding to the three hit categories: dark noise, positron-induced signal, and C14^14C background. Implementation details. The architecture uses L=6L=6, dmodel=512d_model=512, and nhead=4n_head=4, totaling approximately 19.09 million trainable parameters. Optimization employs AdamW with a cosine-annealing schedule. I.3 Spatiotemporal Transformer with Vector Charge Encoding (STT-Vector) The STT-Vector model, inspired by token-based Transformer models such as the Particle Transformer [38], augments the STT-Scalar by enriching each hit token with an 18-dimensional charge feature vector iq_i, explicitly encoding local and global spatiotemporal charge density correlations. The overall architecture is shown in Fig. 3. Figure 3: Architecture of STT-Vector. Each hit is processed through two parallel encoding branches. The spatiotemporal branch (top) maps normalized spatial coordinates and hit times to a high-dimensional representation using sinusoidal positional encodings ψ(⋅)ψ(·). The vector charge branch (bottom) embeds the 18-dimensional charge feature vector iq_i, which summarizes multi-scale local and global charge accumulation, using a dedicated multilayer perceptron. The two embeddings are concatenated, projected to the Transformer model dimension dmodeld_model, and passed through L Transformer encoder layers to produce per-hit class logits for dark noise, e+e^+, and C14^14C. Vector charge encoding. For each hit i, the vector iq_i is constructed based on charge accumulation within defined temporal frames relative to tit_i: current (i(0):|Δt|≤5nsT^(0)_i:| t|≤ 5\,ns), next (i(+):5<Δt≤15nsT^(+)_i:5< t≤ 15\,ns), and previous (i(−):−15≤Δt<−5nsT^(-)_i:-15≤ t<-5\,ns). We define the charge sum operator Si(τ,R)=∑j∈i(τ)∩ℬi(R)qjS_i(τ,R)= _j ^(τ)_i _i(R)q_j over frame τ and spatial radius R, and the temporal asymmetry ΔSi(τ′,R,R′)=Si(0,R)−Si(τ′,R′) S_i(τ ,R,R )=S_i(0,R)-S_i(τ ,R ). The feature vector is structured as: i=[qi;glob;loc(1);loc(2);loc(3);loc(4)]⊤∈ℝ18.q_i= [q_i;\;v_glob;\;v^(1)_loc;\;v^(2)_loc;\;v^(3)_loc;\;v^(4)_loc ] ^18. (10) The global block globv_glob captures total event activity (R→∞R→∞): glob,i=(Si(0),Si(+),Si(−),ΔSi(+),ΔSi(−))∞.v_glob,i= (S_i(0),S_i(+),S_i(-), S_i(+), S_i(-) )_∞. (11) The four local blocks loc(k)v^(k)_loc encode density at spatial scales defined by radius pairs (Rk,Rk′)∈(3,5),(5,9),(10,16),(16,23)m(R_k,R _k)∈\(3,5),(5,9),(10,16),(16,23)\\,m: loc(k)=(Si(0,Rk),Si(+,Rk′),ΔSi(+,Rk,Rk′)).v^(k)_loc= (S_i(0,R_k),\;S_i(+,R _k),\; S_i(+,R_k,R _k) ). (12) This vector is precomputed for every hit, resulting in an input tensor ∈ℝB×Lmax×22X ^B× L_ × 22. Model Architecture. Spatiotemporal coordinates (x,y,z,t)(x,y,z,t) are encoded using the periodic function ψ(⋅)ψ(·) as in STT-Scalar. The charge feature vector iq_i is independently embedded via an MLP to a dimension of dq=64d_q=64. These spatial and charge embeddings are concatenated and projected to the Transformer model dimension dmodel=512d_model=512. The sequence is processed by L=6L=6 Transformer encoder layers (nhead=4n_head=4), resulting in a total of approximately 19.11 million trainable parameters. The sequence padding, the masking mechanism defined in Eq. (6), and the final output MLP mapping to the three-class logits is_i are identical to those employed in the STT-Scalar model. I.4 Computational complexity and practical considerations We analyze the computational complexity and memory footprint of the three models to address feasibility for hit multiplicities N∼(103–104)N (10^3--10^4). Estimates assume batch size B, hidden dimension H (which corresponds to dmodeld_model in the STT architectures), neighbor count K (graph models), layer count L, and attention heads nheadn_head (Transformers). I.4.1 Gated-STGNN: Sparse neighborhood message passing Gated-STGNN utilizes sparse message passing on a graph with bounded degree K, yielding an edge set size E≃NKE NK. Computational complexity. The dominant costs per layer are MLP-based message computation and GRU state updates. With matrix-vector multiplications scaling quadratically in feature dimension, the per-layer cost is (BNKH2)O(BNKH^2). The total scaling is linear in N: Gated−STGNN∼(L⋅BNKH2).C_Gated-STGNN (L· BNKH^2). (13) Memory footprint. Training memory is dominated by node states and edge tensors required for backpropagation. The scaling per layer is ℳlayer∼(BNKH+BNKde),M_layer (BNKH+BNKd_e), (14) where ded_e is the edge-feature dimension. Neighbor construction cost. Neighbor selection employs vectorized dense operations for GPU throughput, incurring a distance computation cost of (BN2)O(BN^2) in both compute and memory. However, because this step involves only low-dimensional geometric distance calculations without dense neural network parameter multiplications, its empirical runtime overhead is negligible compared to the (L⋅BNKH2)O(L· BNKH^2) network scaling. This preprocessing step is performed once per event and is independent of the message-passing layers. I.4.2 STT-Scalar: Global self-attention with scalar charge STT-Scalar processes hits as tokens with global self-attention in each layer. Computational complexity. The cost is dominated by self-attention matrix computation and value aggregation. The N2N^2 token interactions lead to a total complexity scaling of STT-Scalar∼(BLN2H).C_STT -Scalar \! (B\,L\,N^2\,H ). (15) This quadratic dependence on N limits scalability at high multiplicities. Memory footprint. Training memory is dominated by the storage of attention weights and intermediate activations: ℳattn∼(BLnheadN2).M_attn \! (B\,L\,n_head\,N^2 ). (16) The N2N^2 attention tensor typically constrains the maximum feasible batch size. I.4.3 STT-Vector: Global self-attention with vector charge features STT-Vector shares the Transformer backbone of STT-Scalar but includes an additional feature-construction step. Computational complexity (network). The network complexity remains identical to STT-Scalar at leading order: STT-Vectornet∼(BLN2H).C_STT -Vector^net \! (B\,L\,N^2\,H ). (17) The 18-dimensional vector embedding adds a subleading (BN)O(BN) cost. Charge-feature construction cost. Aggregating charges within spatiotemporal windows involves evaluating hit pairs, scaling as: feat∼(N2).C_feat \! (N^2 ). (18) In this study, these features are precomputed to decouple the (N2)O(N^2) preprocessing burden from the training loop. Summary. Although neighbor construction inherently entails an (N2)O(N^2) geometric calculation, the dominant computational cost and bottleneck of the Gated-STGNN are governed by the dense matrix multiplications of the sparse neighborhood aggregations. Consequently, within the empirical regime of typical event sizes, the effective training time of the Gated-STGNN scales linearly with the hit number N, following (NK)O(NK). In contrast, the Transformer models exhibit a quadratic (N2)O(N^2) complexity in both computation and memory footprint due to the global self-attention mechanism. The STT-Vector architecture incurs an additional (N2)O(N^2) preprocessing overhead required for constructing the pairwise vector features. Benchmarks on a single NVIDIA A6000 GPU (48 GB VRAM) indicate that the Gated-STGNN requires approximately an order of magnitude less training time per epoch than the STT models. This disparity reflects the linear versus quadratic complexity scaling with the hit number N, compounded by memory constraints that necessitate a reduced batch size for the Transformers (B=12B=12) relative to the Gated-STGNN (B=16B=16). The per-epoch training run times of STT-Scalar and STT-Vector are comparable, as vector feature construction is offloaded to pre-processing. Truth (a) Gated-STGNN (b) STT-Scalar (c) STT-Vector (d) Figure 4: A visualization of the hit identification results of Gated-STGNN, STT-Scalar, and STT-Vector, compared with the truth for a pile-up event with Δt=165.0ns t=165.0~ns. The truth information for the e+e^+ event consists of a kinetic energy Ek(e+)=0MeVE_k(e^+)=0~MeV and a vertex position of (x,y,z)=(0.03,−0.02,−0.08)m(x,y,z)=(0.03,-0.02,-0.08)~m, while that for the C14^14C event consists of a deposited energy EC14=82.5keVE_^14C=82.5~keV and a vertex position of (x,y,z)=(7.65,5.75,11.23)m(x,y,z)=(7.65,5.75,11.23)~m. The plots display the (θ,ϕ,t)(θ,φ,t) distribution of the hits, where t is the hit time and (θ,ϕ)(θ,φ) are the polar and azimuthal angles of the PMT. Blue dots represent hits from the e+e^+ event, red dots from the C14^14C decay, and gray dots indicate dark noise. IV Training and Evaluation Gated-STGNN (a) STT-Scalar (b) STT-Vector (c) Figure 5: Normalized confusion matrices for events with z=0z=0 and Ek(e+)=0MeVE_k(e^+)=0~MeV. The panels display the classification performance for Gated-STGNN, STT-Scalar, and STT-Vector. Gated-STGNN (a) STT-Scalar (b) STT-Vector (c) Figure 6: Normalized confusion matrices for events with z=0z=0 and Ek(e+)=5MeVE_k(e^+)=5~MeV. The panels display the classification performance for Gated-STGNN, STT-Scalar, and STT-Vector. IV.1 Training objective Three models are trained by minimizing the hit-level cross-entropy loss ℒCEL_CE over the set of valid (non-padding) hits V: ℒCE=−1N∑i∈logpi,ci,L_CE=- 1N_V _i p_i,c_i, (19) where N_V denotes the total count of valid hits. The term pi,cip_i,c_i represents the predicted probability assigned to the truth class cic_i, derived by applying the softmax function to the output logits is_i. The padding masks are applied to exclude invalid entries from the loss computation. IV.2 Hit-level classification metrics For each test set, a 3×33× 3 confusion matrix is constructed from the aggregated hit-level predictions. The normalized matrix is defined as Cab=N(y^=b,y=a)∑b′∈0,1,2N(y^=b′,y=a),C_ab= N( y=b,\,y=a) _b ∈\0,1,2\N( y=b ,\,y=a), (20) where a labels the true class, b labels the predicted class, and N(y^=b,y=a)N( y=b,\,y=a) denotes the number of hits with true label a predicted as b. The indices a,b∈0,1,2a,b∈\0,1,2\ correspond to the dark noise, positron (e+e^+), and C14^14C classes, respectively. The diagonal element C22C_22 defines the C14^14C signal efficiency (recall), denoted as RC14R_^14C. The contamination levels are quantified by the off-diagonal elements C02C_02 and C12C_12, which represent the misidentification fractions of dark noise (fDN→C14f_DN→^14C) and positron hits (fe+→C14f_e^+→^14C) classified as C14^14C. The practical goal is to suppress pile-up from C14^14C while minimizing the loss of e+e^+ hits, thereby reducing the degradation of the reconstructed energy resolution. To further elucidate model performance, the C14^14C recall is evaluated differentially as a function of the time separation Δt t, as defined in Eq. (2), and the deposited energy EC14E_^14C: RC14(Δt,EC14)=N(y^=2,y=2|Δt,EC14)N(y=2|Δt,EC14).R_^14C( t,E_^14C)= N( y=2,y=2| t,E_^14C)N(y=2| t,E_^14C). (21) e+→C14e^+→^14C Misidentification (a) DN→C14→^14C Misidentification (b) C14^14C Recall (c) Figure 7: Classification metrics as functions of positron kinetic energy Ek(e+)E_k(e^+) at z=0mz=0~m. The panels show the dependence on Ek(e+)E_k(e^+) of the fraction of e+e^+ hits misidentified as C14^14C, the fraction of dark noise (DN) hits misidentified as C14^14C, and the C14^14C recall. e+→C14e^+→^14C Misidentification (a) DN→C14→^14C Misidentification (b) C14^14C Recall (c) Figure 8: Classification metrics as functions of the vertex position z (evaluated at discrete points z∈0,6,15,16.5mz∈\0,6,15,16.5\~m) for events with a fixed positron kinetic energy Ek(e+)=0MeVE_k(e^+)=0~MeV. The panels display the dependence on z of the fraction of e+e^+ hits misidentified as C14^14C, the fraction of dark noise (DN) hits misidentified as C14^14C, and the C14^14C recall. (a) Ek(e+)=0MeVE_k(e^+)=0~MeV (b) Ek(e+)=5MeVE_k(e^+)=5~MeV Figure 9: Differential C14^14C recall RC14(Δt,EC14)R_^14C( t,E_^14C) of STT-Vector, evaluated at z=0z=0 for Ek(e+)=0MeVE_k(e^+)=0~MeV and Ek(e+)=5MeVE_k(e^+)=5~MeV, respectively. Figure 10: Distributions of the total charge of selected hits (QsumQ_sum) at the detector center for events with Ek(e+)=0MeVE_k(e^+)=0~MeV (left) and Ek(e+)=5MeVE_k(e^+)=5~MeV (right). The gray dotted line represents QsumQ_sum of all hits, before applying hits removal. The green dashed line shows the ideal distribution with C14^14C pile-up removed by truth label. The solid purple line represents QsumQ_sum after hit removal based on the STT-Vector model prediction. The fitted mean value (μ), width (σ) and corresponding resolution (ℛ=σ/μR=σ/μ) are also shown. IV.3 Impact on the Energy Resolution The removal of C14^14C hits will reduce fluctuations in the number of detected photons and consequently improve the energy resolution. A simple study is performed to evaluate the performance of the algorithms in terms of the relative improvement of the energy resolution. This study does not represent the performance of the real JUNO detector, which would require including the improvements from the well-designed energy- and vertex-reconstruction procedure [23, 24]. The improvement in the energy resolution is quantitatively evaluated using the total charge Qsum=∑iqiQ_sum= _iq_i, where i runs over all the selected hits and qiq_i denotes the charge of the i-th hit. Hits are selected under different criteria: Total : all hits are retained. In this case, the QsumQ_sum includes the contributions from e+e^+ and C14^14C, as well as from the dark noise. The average number of DN hits, ⟨Bdn⟩=400 B_dn =400, is subtracted from QsumQ_sum to provide a more accurate estimate of the total charge induced by the particles. However, fluctuations in BdnB_dn still propagate to the QsumQ_sum resolution. C14^14C removed (truth) : C14^14C hits are removed according to the truth hit labels. In this case, ⟨Bdn⟩ B_dn is also subtracted from QsumQ_sum. It represents the ideal scenario in which the models predict the C14^14C hits with 100% efficiency and purity. C14^14C removed (predicted) : C14^14C hits are removed according to the hit labels predicted by a model. Since the model will misidentify DN hits as C14^14C hits, the removal of C14^14C hits will change the average number of DN hits. Instead, (1−C02)⟨Bdn⟩(1-C_02) B_dn is subtracted from QsumQ_sum, where C02C_02 is the misidentification fraction of dark noise (fDN→C14f_DN→^14C). e+e^+ only (truth) : only e+e^+ hits are selected according to the truth hit labels to show the intrinsic resolution of e+e^+ hits. Since the DN hits are excluded, ⟨Bdn⟩ B_dn is not subtracted from QsumQ_sum. The QsumQ_sum distributions are fitted with a Gaussian function. The energy resolution is defined as the ratio of the fitted width to the fitted mean, ℛ=σ/μR=σ/μ. V Results and Discussion Fig. 4 visualizes the hit-level classification performance for a representative event with Δt=165.0ns t=165.0~ns. The predictions from the Gated-STGNN, STT-Scalar, and STT-Vector models are consistent with the truth. The models successfully tag C14^14C hits while effectively discriminating against the dense positron hits and dispersed dark noise hits. Since light propagation varies significantly with the annihilation vertex, it is necessary to evaluate the performance of hit-level tagging for e+e^+ annihilation events at different positions. At higher positron energies, the relative contribution of C14^14C pile-up to the total light yield is smaller, and its impact on the energy resolution is correspondingly reduced. However, C14^14C tagging becomes more challenging at higher positron energies, because the C14^14C hits are statistically overwhelmed by the larger e+e^+-induced hit population, which is typically of order 10410^4. Therefore, the identification performance exhibits distinct behavior at low and high positron energies. The evaluation is conducted across the discrete grid of positron kinetic energies described in Section I. Unless explicitly stated, all models are evaluated on the same set. V.1 Global classification performance The global classification performance for events with Ek(e+)=0E_k(e^+)=0 and 5MeV5~MeV at the detector center (z=0z=0) is summarized by the normalized confusion matrices in Figs. 5 and 6. The C14^14C recall (C22C_22) is shown in the bottom-right corner. Two additional metrics are particularly important: C02C_02 and C12C_12. They quantify the misidentification fractions of dark-noise hits and positron hits classified as C14^14C, namely fDN→C14f_DN→^14C and fe+→C14f_e^+→^14C. The former biases the effective dark noise level after removal, while the latter reduces the signal light yield. Good performance is characterized by high C22C_22 and low values of C02C_02 and C12C_12. At Ek(e+)=0MeVE_k(e^+)=0~MeV, the STT-Vector model achieves the highest C14^14C identification recall of 48.54%, followed closely by STT-Scalar at 46.53%. In contrast, Gated-STGNN exhibits a notably lower recall of 37.56%. The primary error mode is the misidentification of C14^14C as e+e^+, which is 38.88% for STT-Vector and 47.65% for Gated-STGNN. Across all models, the misidentification of e+e^+ as C14^14C are less than 1%1\% and the misidentification of dark noise as C14^14C are less than 2%2\%. This behavior is desirable, indicating that when C14^14C hits are removed based on the predictions, the signal light yield and the average dark-noise level are not strongly affected. At Ek(e+)=5MeVE_k(e^+)=5~MeV, the classification landscape is dominated by the e+e^+ hits. Contrary to the low-energy regime, the C14^14C recall decreases significantly across all models. The STT-Vector model maintains the highest recall, achieving a C14^14C identification recall of 27.43%27.43\%, compared to 25.61%25.61\% for STT-Scalar and 20.69%20.69\% for Gated-STGNN. The dominant error mode for all models is the misidentification of C14^14C hits as positron-induced, with leakage rates exceeding 70%70\%. This behavior indicates that the C14^14C hits are statistically overwhelmed by the much larger population of e+e^+ hits, making their identification increasingly challenging. The misidentification rates of e+e^+ as C14^14C and of dark noise as C14^14C are both lower than in the low-energy regime. The dependence of the C02C_02, C12C_12 and C22C_22 on the positron kinetic energy Ek(e+)E_k(e^+) at the detector center (z=0z=0) is summarized in Fig. 7. Figure 8 illustrates the z dependence of the classification metrics with Ek(e+)=0MeVE_k(e^+)=0~MeV. As shown, the STT-Vector model maintains a consistently high C14^14C recall and low misidentification rates, demonstrating resilience against changes in light path. V.2 Timing and energy dependence of C14^14C identification When the C14^14C decay occurs close in time to the e+e^+, the hit-level tagging performance varies strongly with Δt t, since the number of e+e^+ hits surrounding the C14^14C changes rapidly. To quantify this effect, the C14^14C hit identification recall is evaluated as a function of the relative time offset Δt t and the C14^14C deposited energy EC14E_^14C. Figure 9 displays the differential C14^14C recall as a function of Δt t and EC14E_^14C. The maps correspond to events at the detector center (z=0z=0) with positron kinetic energies of 0 and 5MeV5~MeV, respectively. For larger |Δt|| t|, the C14^14C background hits are temporally decoupled from the prompt e+e^+ signal window, thereby facilitating high C14^14C recall. In contrast, the region with small |Δt|| t| corresponds to strong temporal overlap, where discrimination must rely on subtler correlations among hit timing, geometry, and local charge patterns. When the C14^14C decays with higher energy, its hits become easier to identify. The overall C14^14C recall is lower at higher positron kinetic energies, since the larger e+e^+ hit population makes it harder to disentangle C14^14C hits from the e+e^+ signal. Figure 11: Ratios of the resolutions of the charge sum of selected hits (QsumQ_sum) under different hit-removal criteria relative to the “Total” case, in which no hit is removed. The left plot shows the ratio’s dependence on EkE_k at the detector center. The right plot shows the ratio’s dependence on z when Ek(e+)=0MeVE_k(e^+)=0~MeV. V.3 Energy-resolution recovery from hit-level denoising The impact of hit-level classification on energy resolution is evaluated through the resolution of QsumQ_sum. Notably, the strict per-hit classification accuracy does not directly translate to energy resolution performance. For example, consider a positron hit that is spatially and temporally close to a C14^14C hit. If a model classifies this hit as a C14^14C hit, it will be counted as a misidentification, reducing the recall. However, from the perspective of energy reconstruction based on the total charge QsumQ_sum or the likelihood-based reconstruction algorithm [13], such a “mistake” may have negligible impact. This is because, in such cases, the overall statistical behavior has a more significant impact on energy resolution than the precise classification of individual hits. The QsumQ_sum distribution of an Ek(e+)=0MeV,z=0mE_k(e^+)=0~MeV,z=0~m dataset is shown on the left of Fig. 10. After removing C14^14C hits predicted by STT-Vector, the QsumQ_sum resolution improves by 18%18\%. The QsumQ_sum distribution of an Ek(e+)=5MeV,z=0mE_k(e^+)=5~MeV,z=0~m dataset is shown on the right of Fig. 10. After removing C14^14C hits according to the STT-Vector prediction, the QsumQ_sum resolution improves by 3%3\%. With higher Ek(e+)E_k(e^+), the discrimination of C14^14C hits becomes more challenging due to the increasing number of e+e^+ hits. However, the effect of C14^14C hits on energy resolution also diminishes, since at high Ek(e+)E_k(e^+) the resolution becomes dominated by intrinsic e+e^+ photon statistics. The performances of Gated-STGNN, STT-Scalar, and STT-Vector for the datasets with z=0z=0 and different Ek(e+)E_k(e^+) are summarized on the left of Fig. 11. The three models exhibit similar performance. In the Ek(e+)=0MeVE_k(e^+)=0~MeV case, the STT-Vector has better performance. The performances of the three models for the datasets with Ek(e+)=0MeVE_k(e^+)=0~MeV and different z are summarized on the right of Fig. 11. The behaviors of the three models are similar. An improvement of ∼20% 20\% in the resolution is observed for all tested samples with different z, indicating that the proposed models remain valid under more complex light-propagation conditions. VI Conclusion and Outlook In this work, we propose three models to tag C14^14C hits in e+e^+ events affected by a single C14^14C pile-up. By identifying and suppressing background photon hits, the method mitigates the impact of C14^14C pile-up on event reconstruction. The three models are: Gated-STGNN, which represents hits as graphs in spacetime and controls information propagation using GRU units; STT-Scalar, which applies attention to point-wise features; and STT-Vector, which further includes aggregated features of local topology. In test datasets with Ek(e+)≤5MeVE_k(e^+)≤ 5~MeV, all models successfully tag C14^14C hits while keeping the misidentification of e+e^+ as C14^14C below 1%1\% and the misidentification of dark noise as C14^14C below 2%2\%, thereby limiting signal depletion and shifts in the dark-noise level. C14^14C recall ranges from 25%25\% to 48%48\%, generally ordered as Gated-STGNN << STT-Scalar << STT-Vector. Within the studied overlap regime of Δt∈[−100,300]ns t∈[-100,300]\,ns, differential recall maps indicate that the most challenging region is Δt∈[0,200]ns t∈[0,200]\,ns, where the temporal overlap between C14^14C and e+e^+ is the strongest. After C14^14C-hit removal based on model predictions, improvements in the resolution of the total charge QsumQ_sum are observed across all test datasets with different Ek(e+)E_k(e^+) and z positions, confirming efficacy from the detector center to near-boundary regions under varying positron kinetic energies. In the challenging pile-up regime Δt∈[−100,300]ns t∈[-100,300]~ns, the resolution of QsumQ_sum improves by approximately 5%5\%–20%20\%, with STT-Vector yielding the best recovery in the low Ek(e+)E_k(e^+) case. Key limitations include the prioritization of signal retention over aggressive C14^14C-hit tagging, and the dependence on temporal overlap. Future work should address: 1. Refinement for highly overlapping events: Incorporating physically motivated descriptors (e.g., time-of-flight corrections, multi-scale context) to improve performance in the Δt∈[0,200]ns t∈[0,200]~ns region. 2. Robustness and domain shift: Systematically validating against various detector conditions and exploring domain-adaptation strategies to mitigate simulation-data discrepancies. 3. Integration with reconstruction: Propagating hit-level selections through the full reconstruction chain to evaluate impacts on the energy and vertex reconstruction. This study focuses on the algorithmic feasibility of hit-level pile-up suppression. The impact on energy resolution is evaluated using QsumQ_sum as an estimator of calorimetric response at the hit level. These results should not be interpreted as a full detector-level performance estimate. A comprehensive assessment requires integration of the hit-selection procedure into the complete JUNO energy reconstruction chain. VII Acknowledgments This work was partly supported by National Natural Science Foundation of China (Grants No. 62472381, No. 12575209), by the National Key R&D Program of China (2023YFA1606103). We would also like to thank the Computing Center of the Institute of High Energy Physics, Chinese Academy of Science for providing the GPU resources. References Bellini et al. [2014] G. Bellini et al. (BOREXINO), Neutrinos from the primary proton–proton fusion process in the Sun, Nature 512, 383 (2014). Aad et al. [2015] G. Aad et al. (ATLAS), Measurement of the inclusive jet cross-section in proton-proton collisions at s=7 s=7 TeV using 4.5 fb-1 of data with the ATLAS detector, JHEP 02, 153, [Erratum: JHEP 09, 141 (2015)], arXiv:1410.8857 [hep-ex] . Khachatryan et al. [2015] V. Khachatryan et al. (CMS), Performance of the CMS missing transverse momentum reconstruction in p data at s s = 8 TeV, JINST 10 (02), P02006, arXiv:1411.0511 [physics.ins-det] . Andringa et al. [2016] S. Andringa et al. (SNO+), Current Status and Future Prospects of the SNO+ Experiment, Adv. High Energy Phys. 2016, 6194250 (2016), arXiv:1508.05759 [physics.ins-det] . Obara et al. [2022] S. Obara, Y. Gando, S. Ieki, H. Ikeda, K. Ishidoshiro, and R. Nakamura (KamLAND-Zen), A study of self-vetoing balloon vessel for liquid-scintillator detectors, J. Phys. Conf. Ser. 2374, 012111 (2022). Arpesella et al. [2008] C. Arpesella et al. (Borexino), First real time detection of Be-7 solar neutrinos by Borexino, Phys. Lett. B 658, 101 (2008), arXiv:0708.2251 [astro-ph] . Eguchi et al. [2003] K. Eguchi et al. (KamLAND), First results from KamLAND: Evidence for reactor anti-neutrino disappearance, Phys. Rev. Lett. 90, 021802 (2003), arXiv:hep-ex/0212021 . Lozza [2012] V. Lozza (SNO+), Scintillator phase of the SNO+ experiment, J. Phys. Conf. Ser. 375, 042050 (2012), arXiv:1201.6599 [hep-ex] . An et al. [2023] F. P. An et al. (Daya Bay), Precision Measurement of Reactor Antineutrino Oscillation at Kilometer-Scale Baselines by Daya Bay, Phys. Rev. Lett. 130, 161802 (2023), arXiv:2211.14988 [hep-ex] . Ahn et al. [2012] J. K. Ahn et al. (RENO), Observation of Reactor Electron Antineutrino Disappearance in the RENO Experiment, Phys. Rev. Lett. 108, 191802 (2012), arXiv:1204.0626 [hep-ex] . Abe et al. [2012] Y. Abe et al. (Double Chooz), Indication of Reactor ν¯e ν_e Disappearance in the Double Chooz Experiment, Phys. Rev. Lett. 108, 131801 (2012), arXiv:1112.6353 [hep-ex] . Abusleme et al. [2021] A. Abusleme et al. (JUNO, Daya Bay), Optimization of the JUNO liquid scintillator composition using a Daya Bay antineutrino detector, Nucl. Instrum. Meth. A 988, 164823 (2021), arXiv:2007.00314 [physics.ins-det] . Abusleme et al. [2022] A. Abusleme et al. (JUNO), JUNO Physics and Detector, Prog. Part. Nucl. Phys. 123, 103927 (2022), arXiv:2104.02565 [hep-ex] . Chen et al. [2023] G.-M. Chen, X. Zhang, Z.-Y. Yu, S.-Y. Zhang, Y. Xu, W.-J. Wu, Y.-G. Wang, and Y.-B. Huang, Discrimination of p solar neutrinos and 14C double pile-up events in a large-scale LS detector, Nucl. Sci. Tech. 34, 137 (2023), arXiv:2303.08512 [hep-ex] . Fang et al. [2026] W. Fang, W. Li, W. Luo, Z. Wu, and M. He, Deep Learning-Based 14C Pile-Up Identification in the JUNO Experiment, (2026), arXiv:2603.01419 [hep-ex] . Vaughan et al. [2025] L. Vaughan, M. Rakib, S. Patel, F. Rizatdinova, A. Khanov, and A. Bagavathi, PileUp Mitigation at the HL-LHC Using Attention for Event-Wide Context, (2025), arXiv:2503.02860 [hep-ex] . Algren et al. [2025] M. Algren, T. Golling, C. Pollard, and J. A. Raine, Variational inference for pile-up removal at hadron colliders with diffusion models, Phys. Rev. D 111, 116010 (2025), arXiv:2410.22074 [hep-ph] . Li et al. [2023] T. Li, S. Liu, Y. Feng, G. Paspalaki, N. V. Tran, M. Liu, and P. Li, Semi-supervised graph neural networks for pileup noise removal, Eur. Phys. J. C 83, 99 (2023), arXiv:2203.15823 [hep-ex] . Arjona Martínez et al. [2019] J. Arjona Martínez, O. Cerri, M. Pierini, M. Spiropulu, and J.-R. Vlimant, Pileup mitigation at the Large Hadron Collider with graph neural networks, Eur. Phys. J. Plus 134, 333 (2019), arXiv:1810.07988 [hep-ph] . Bertolini et al. [2014] D. Bertolini, P. Harris, M. Low, and N. Tran, Pileup Per Particle Identification, JHEP 10, 059, arXiv:1407.6013 [hep-ph] . Soyez et al. [2013] G. Soyez, G. P. Salam, J. Kim, S. Dutta, and M. Cacciari, Pileup subtraction for jet shapes, Phys. Rev. Lett. 110, 162001 (2013), arXiv:1211.2811 [hep-ph] . An et al. [2016] F. An et al. (JUNO), Neutrino Physics with JUNO, J. Phys. G 43, 030401 (2016), arXiv:1507.05613 [physics.ins-det] . Abusleme et al. [2025a] A. Abusleme et al. (JUNO), Prediction of Energy Resolution in the JUNO Experiment, Chin. Phys. C 49, 013003 (2025a), arXiv:2405.17860 [hep-ex] . Abusleme et al. [2025b] A. Abusleme et al. (JUNO), Initial performance results of the JUNO detector, (2025b), arXiv:2511.14590 [hep-ex] . Huang et al. [2023] G.-h. Huang, W. Jiang, L.-j. Wen, Y.-f. Wang, and W.-M. Luo, Data-driven simultaneous vertex and energy reconstruction for large liquid scintillator detectors, Nucl. Sci. Tech. 34, 83 (2023), arXiv:2211.16768 [physics.ins-det] . Takenaka et al. [2025] A. Takenaka et al., Neutron source-based event reconstruction algorithm in large liquid scintillator detectors, Eur. Phys. J. C 85, 1097 (2025), arXiv:2504.19236 [physics.ins-det] . Qian et al. [2021] Z. Qian et al., Vertex and energy reconstruction in JUNO with machine learning methods, Nucl. Instrum. Meth. A 1010, 165527 (2021), arXiv:2101.04839 [physics.ins-det] . Li et al. [2016] Y. Li, D. Tarlow, M. Brockschmidt, and R. S. Zemel, Gated graph sequence neural networks, ICLR (2016). Vaswani et al. [2017] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, Attention Is All You Need, in 31st International Conference on Neural Information Processing Systems (2017) arXiv:1706.03762 [cs.CL] . Qi et al. [2017] C. R. Qi, H. Su, K. Mo, and L. J. Guibas, Pointnet: Deep learning on point sets for 3d classification and segmentation, in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR) (2017) p. 77–85. Fan et al. [2021] H. Fan, Y. Yang, and M. Kankanhalli, Point 4d transformer networks for spatio-temporal modeling in point cloud videos, in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) (2021) p. 14199–14208. Guo et al. [2021] Y. Guo, H. Wang, Q. Hu, H. Liu, L. Liu, and M. Bennamoun, Deep learning for 3d point clouds: A survey, IEEE Trans. Pattern Anal. Mach. Intell. 43, 4338 (2021). Liu et al. [2025] J. Liu, J. Han, L. Liu, A. I. Aviles-Rivero, C. Jiang, Z. Liu, and H. Wang, Mamba4D: Efficient 4d point cloud video understanding with disentangled spatial-temporal state space models, in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) (2025) p. 17626–17636. Lv et al. [2025] B. Lv, Y. Zha, T. Dai, Y. Xue, K. Chen, and S.-T. Xia, Adapting pre-trained 3d models for point cloud video understanding via cross-frame spatio-temporal perception, in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) (2025) p. 12413–12422. Shlomi et al. [2020] J. Shlomi, P. Battaglia, and J.-R. Vlimant, Graph Neural Networks in Particle Physics 10.1088/2632-2153/abbf9a (2020), arXiv:2007.13681 [hep-ex] . Cho et al. [2014] K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation, (2014), arXiv:1406.1078 [cs.CL] . Loshchilov and Hutter [2017] I. Loshchilov and F. Hutter, Decoupled Weight Decay Regularization (2017) arXiv:1711.05101 [cs.LG] . Qu et al. [2022] H. Qu, C. Li, and S. Qian, Particle Transformer for Jet Tagging, (2022), arXiv:2202.03772 [hep-ph] .