Paper deep dive
Interpolation-Driven Machine Learning Approaches for Plume Shine Dose Estimation: A Comparison of XGBoost, Random Forest, and TabNet
Biswajit Sadhu, Kalpak Gupte, Trijit Sadhu, S. Anand
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 3:58:35 PM
Summary
This study develops an interpolation-assisted machine learning framework for rapid plume shine dose estimation, addressing computational bottlenecks in traditional photon-transport methods. Using datasets generated by the pyDOSEIA suite for 17 gamma-emitting radionuclides, the authors augment sparse discrete data with shape-preserving interpolation (PCHIP) to create high-resolution training sets. They compare the predictive performance of Random Forest, XGBoost, and TabNet, finding that XGBoost achieves the highest accuracy. Interpretability analysis reveals that tree-based models prioritize geometric-dispersion features, while TabNet distributes attention more broadly. A web-based GUI was also developed for interactive scenario evaluation.
Entities (8)
Relation Signals (6)
XGBoost → outperforms → Random Forest
confidence 95% · XGBoost consistently achieved the highest accuracy
XGBoost → outperforms → TabNet
confidence 95% · XGBoost consistently achieved the highest accuracy
Tree-based models → focusonfeatures → geometry-dispersion features
confidence 90% · Tree-based models focus mainly on dominant geometry-dispersion features (release height, stability category, and downwind distance)
pyDOSEIA → generatesdatafor → Plume Shine Dose
confidence 90% · datasets generated with the pyDOSEIA suite for 17 gamma-emitting radionuclides
TabNet → usesmechanism → attention-based feature attribution
confidence 90% · TabNet distributes attention more broadly across multiple variables
PCHIP → usedforaugmentation → pyDOSEIA
confidence 85% · datasets were augmented using shape-preserving interpolation... Piecewise Cubic Hermite Interpolating Polynomials (PCHIP)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Despite the success of machine learning (ML) in surrogate modeling, its use in radiation dose assessment is limited by safety-critical constraints, scarce training-ready data, and challenges in selecting suitable architectures for physics-dominated systems. Within this context, rapid and accurate plume shine dose estimation serves as a practical test case, as it is critical for nuclear facility safety assessment and radiological emergency response, while conventional photon-transport-based calculations remain computationally expensive. In this work, an interpolation-assisted ML framework was developed using discrete dose datasets generated with the pyDOSEIA suite for 17 gamma-emitting radionuclides across varying downwind distances, release heights, and atmospheric stability categories. The datasets were augmented using shape-preserving interpolation to construct dense, high-resolution training data. Two tree-based ML models (Random Forest and XGBoost) and one deep learning (DL) model (TabNet) were evaluated to examine predictive performance and sensitivity to dataset resolution. All models showed higher prediction accuracy with the interpolated high-resolution dataset than with the discrete data; however, XGBoost consistently achieved the highest accuracy. Interpretability analysis using permutation importance (tree-based models) and attention-based feature attribution (TabNet) revealed that performance differences stem from how the models utilize input features. Tree-based models focus mainly on dominant geometry-dispersion features (release height, stability category, and downwind distance), treating radionuclide identity as a secondary input, whereas TabNet distributes attention more broadly across multiple variables. For practical deployment, a web-based GUI was developed for interactive scenario evaluation and transparent comparison with photon-transport reference calculations.
Tags
Links
- Source: https://arxiv.org/abs/2602.19584v1
- Canonical: https://arxiv.org/abs/2602.19584v1
Trouble viewing inline? Open PDF directly →
Full Text
79,691 characters extracted from source content.
Expand or collapse full text
Interpolation-Driven Machine Learning Approaches for Plume Shine Dose Estimation: A Comparison of XGBoost, Random Forest, and TabNet Biswajit Sadhu 1,2* , Kalpak Gupte 2,4 , Trijit Sadhu 3 , S Anand 1,2 1 Health Safety & Environment Group, Bhabha Atomic Research Centre, Mumbai, 400085, India. 2 Homi Bhabha National Institute, Mumbai, 400094, India. 3 Birla Institute of Technology And Science, Street, PILANI, 333031, Rajasthan, India. 4 Department of Computer Engineering & Technology, Dr. Vishwanath Karad MIT World Peace University, Pune, 411038, India. *Corresponding author(s). E-mail(s): bsadhu@barc.gov.in, biswajit.chem001@gmail.com; Abstract Despite the success of machine learning (ML) in surrogate modeling, its use in radiation dose assessment is limited by safety-critical constraints, scarce training-ready data, and challenges in selecting suitable architectures for physics- dominated systems. Within this context, rapid and accurate plume shine dose estimation serves as a practical test case, as it is critical for nuclear facility safety assessment and radiological emergency response, while conventional photon- transport-based calculations remain computationally expensive. In this work, an interpolation-assisted ML framework was developed using discrete dose datasets generated with the pyDOSEIA suite for 17 gamma-emitting radionuclides across varying downwind distances, release heights, and atmosphericstability cate- gories. The datasets were augmented using shape-preserving interpolation to construct dense, high-resolution training data. Two tree-based ML models (Ran- dom Forest and XGBoost) and one deep learning (DL) model (TabNet) were evaluated to examine predictive performance and sensitivity todataset res- olution. All models showed higher prediction accuracy with the interpolated 1 arXiv:2602.19584v1 [cs.LG] 23 Feb 2026 high-resolution dataset than with the discrete data; however,XGBoost consis- tently achieved the highest accuracy. Interpretability analysis using permutation importance (tree-based models) and attention-based feature attribution (Tab- Net) revealed that performance di erences stem from how the models utilize input features. Tree-based models focus mainly on dominant geometry–dispersion features (release height, stability category, and downwind distance), treat- ing radionuclide identity as a secondary input, whereas TabNetdistributes attention more broadly across multiple variables. For practicaldeployment, a web-based GUI was developed for interactive scenario evaluation and transparent comparison with photon-transport reference calculations. Keywords:Plume Shine Dose, Random Forest, XGBoost, TabNet, Inductive Bias 1 Introduction Machine learning (ML) is increasingly used to accelerate scienti cmodeling by enabling fast and accurate surrogate predictions for complex physical processes in the domain of climate science, materials engineering, and medical imaging. However, the adoption of ML in radiation exposure and dose assessment still remainsrather limited.[ 1,2] This is mainly due to the safety-critical nature of radiological applica- tions, where even small prediction errors can have serious consequences, and to the lack of training-ready dense datasets. Plume shine (also known as cloud shine) is a representative example of this chal- lenge. It refers to the external radiation dose received from gamma emissions of an airborne radioactive cloud following a nuclear or radiological release (Fig:1). Reliable and timely estimation of plume shine dose is of paramount interest both for emergency preparedness and response, supporting decisions on evacuation and sheltering [3], and for the safety assessment of new facilities, where it plays a key role in dose apportion- ment and regulatory compliance. Modeling based approaches are thereforea stndard component of plume shine estimation by the regulatory authorities and emergency planners [4–7]. Traditionally, plume shine dose is computed by coupling atmosphericdispersion models with photon transport and dose conversion formalisms. This approach typically involves a triple integration over the spatial extent of the plume, the gamma-ray energy spectrum of the radionuclide inventory, attenuation and buildup factors that account for air absorption and scattering.While physically rigorous, these approaches incur substantial computational cost, limiting their applicability in time-critical operational settings, especially when assessments involve multiple radionuclideswith complex gamma spectra or require ne spatial resolution[7]. To reduce computational bur- den, many assessment frameworks use simpli ed transport representations such as Gaussian plume models [8,9]. In this context, our previous work introduced pyDO- SEIA, an open-source Python package that implements Gaussian plume dispersion for comprehensive radiological impact assessment, including plume shine dose along 2 Downwind Distance Release Height Receptor Cloud with radionuclide Fig. 1: Schematic representation of plume shine exposure geometry, illustrating radionuclide release from an elevated source, atmospheric transport of the radioactive cloud, and external gamma dose received by a downwind receptor as a function of release height and downwind distance. Although not explicitly shown, the atmospheric stability category governs plume dispersion characteristics and signi cantly in uences the resulting plume shine dose. with other exposure pathways [ 10]. Nevertheless, analytical plume models remain con- strained by assumptions of steady meteorology and simple terrain, whilehigh- delity physics-based simulations remain too slow for real-time decision support [ 11]. These limitations motivate the development of ML-based surrogate models for plume shine dose prediction. Once trained, such models can provide near-instantaneous estimates, enabling rapid scenario screening and real-time support during radiologi- cal emergencies. However, applying ML in this domain presents challenges that di er from those in many standard ML applications. Plume shine dose data are usually available only as sparse tables de ned for a limited number of combinations of down- wind distance, release height, and atmospheric stability class. Training ML models directly on such discontinuous data often results in unstable predictions and poor gen- eralization, which is unacceptable in safety-critical applications where reliability and physical consistency are essential. In addition to data limitations, the choice of ML model plays a critical role. For structured tabular data, tree-based ensemble methods such as RandomForest (RF) [12] and XGBoost [13] have become the standard in practical ML work ows. They are widely used in industry and research and have consistently achieved strong perfor- mance in data-science competitions. Their success is largely due totheir robustness, ability to model nonlinear relationships, and built-in feature selection that reduces the in uence of weak or irrelevant variables. On the contrary, deep learning (DL) based approaches have traditionally been reported to be less e ective on tabular datasets, although recent architectures such as TabNet [14] have been proposed to address this gap through incorporation of sequential attention mechanisms. Despite this, several systematic studies have shown the superiority of ensemble methods over DL models 3 for tabular prediction tasks, particularly when datasets are limited and strongly struc- tured [15,16]. Moreover, understanding the underlying reasons for these performance di erences remains an active area of research. This leads to an important practical question for radiation modeling:Should plume shine dose prediction rely on well- established ensemble methods, or can newer attention-based deep architectures provide clear advantages in this safety-critical domain? In this work, we address both data and model selection challenges and pro- pose an interpolation-assisted ML framework for plume shine dose prediction.The study is guided by two key objectives: achieving high predictive accuracy and ensuring transparent, model-appropriate interpretability.To this end, we conduct a systematic comparison of three leading approaches for tabular data – RF, XGBoost, and TabNet – within a uni ed and controlled experimental setting. Our framework is based on three key methodological elements. First, in order to address the limitations of sparse and discretized dose tables, we employ data augmen- tation using Piecewise Cubic Hermite Interpolating Polynomials (PCHIP) [17] that preserves the underlying physical trends and therefore enabling the construction of high-quality high-resolution training datasets. Second, we conduct asystematic com- parison of model inductive biases by benchmarking established tree-based ensemble methods against DL-based TabNet. Third, we adopt an architecture-aware inter- pretability strategy by applying conditional permutation importance to the ensemble models and analyzing feature relevance in TabNet through its intrinsic attention mech- anisms. This approach enables a consistent examination of how di erentmodels utilize input features and therefore helps understanding the root cause of model’s perfor- mance behavior, in line with established practices in interpretable machine learning [18]. Finally, to improve transparency and practical usability, we include an interactive Streamlit-based interface that allows real-time comparison betweenmachine-learning predictions and physics-based reference calculations, helping connect the methodolog- ical analysis with practical, scenario-based insight. 2 Analytical Plume Shine Dose Computation and Methodological Motivation Plume shine dose refers to the external radiation exposure arising from photon emis- sions of an airborne radioactive cloud incident at a receptor location. In radiation protection practice, plume shine dose is commonly evaluated using analytical point- kernel formulations that integrate photon transport contributions overthe full spatial extent of the radioactive plume. These formulations explicitly account for source geom- etry, atmospheric attenuation, and photon scattering, and are widely regarded as the reference approach for high- delity dose estimation. For a receptor located at (x 1 , y 1 , z 1 ), the plume shine dose rate can be expressed using a point-kernel formulation as ̇ D= ∫ αE γ μ a ρ B(μr) 4πr 2 exp(−μr)χ(x, y, z) dxdydz(1) 4 wherer= √ (x−x 1 ) 2 + (y−y 1 ) 2 + (z−z 1 ) 2 is the source–receptor distance,E γ denotes the photon energy,μandμ a are the linear attenuation and energy-absorption coecients in air,ρis the air density,B(μr) is the buildup factor accounting for scattered photons,χ(x, y, z) is the concentration of radionuclide in plume, andαis a unit conversion constant. As evident from Eq. (1), plume shine dose estimation involves a three-dimensional spatial integration, combined with a summation over the discrete gamma energies emitted by each radionuclide. The e ective spatial extent of this integration is not uniform but is governed by atmospheric dispersion and photon attenuation character- istics. In the lateral (y) and vertical (z) directions, the integration limits are determined by the plume spread parametersσ y andσ z , which depend strongly on the atmospheric stability category[19]. Unstable conditions (stability classes A–C) are associated with enhanced turbulent mixing, leading to broader plume dimensions, whereas stable con- ditions (classes E–F) suppress dispersion and con ne the plume to anarrower region with limited vertical extent. Neutral conditions (class D) represent an intermediate regime between these extremes. Along the downwind direction (x), the e ective inte- gration range is constrained by the photon mean free path (MFP), beyond which dose contributions diminish rapidly due to exponential attenuation.Together, these stability- and photon-based constraints de ne the spatial domain contributing mean- ingfully to plume-shine dose.In practical implementations, the integral does not have a closed-form solution and is therefore evaluated numerically using quadrature or discretization-based techniques.[20] 2.1 Numerical and Computational Challenges Despite its strong physical basis as described above, analytical plume shine dose com- putation involves several numerical and computational diculties thatrestrict its routine use in operational and large-scale applications: • High-dimensional integration:The spatial integration extends over large down- wind, crosswind, and vertical domains. Even when the integration region is truncated to a few photon mean free paths [ 21], the associated computational cost remains signi cant. • Near- eld numerical sti ness:The presence of the 1/r 2 kernel introduces steep gradients in the vicinity of the receptor, which necessitates nespatial discretization to maintain numerical stability and acceptable accuracy. • Multiple gamma emissions:Many radionuclides emit multiple gamma lines (e.g. 154 Eu), each requiring separate attenuation and buildup calculations. As shown in our earlier implementation usingpyDOSEIA[ 10], the computational cost increases approximately linearly with the number of gamma energies considered. • Scaling with resolution and scenario count:Increasing spatial resolution, downwind distance, or release height expands the e ective integration volume. In addition, realistic consequence assessment requires repeatedevaluations across numerous meteorological and source scenarios. Together, these factors leadto a rapid growth in computational demand. 5 Taken together, these limitations make detailed plume shine dose calculations com- putationally intensive and dicult to apply in contexts that require rapid turnaround, such as emergency response and large-scale scenario screening. 2.2 Motivation for Data-Driven Surrogate Modeling The computational challenges outlined above motivate the use of data-driven surrogate models for plume shine dose estimation. Instead of replacing the underlying physics- based formulation, machine learning models can be trained on reference datasets generated through analytical–numerical calculations and then employedto deliver rapid predictions across the relevant input parameter space. However, developing reliable surrogate models for plume shine dose is not straight- forward. A major bottleneck lies in the dataset itself, which, evenwhen generated using analytical calculations, is typically sparse and discretized due to the combina- torial combinations of downwind distance, release height, and atmospheric stability category. As shown later in this study, training machine learning models directly on such discontinuous data can result in unstable predictions and limited generalization capability, which is particularly problematic in safety-critical applications. To address this limitation, the present study adopts an interpolation-assisted dataset construction strategy, in which physically consistent interpolation methods are used to transform sparse analytical dose tables into smooth, high-resolution datasets suitable for modern ML. This strategy forms the methodological bridge between clas- sical radiation transport modeling and the ML framework introduced inthe following sections. 3 Data Generation and Preparation for ML Modeling 3.1 Low-Resolution Dataset Generation using pyDOSEIA The base dataset used in this study was generated using thepyDOSEIApackage [ 10], which implements a Gaussian plume dispersion model for radiological impact assess- ment, including plume shine dose estimation.The simulations assume a short-term plume release with a unit source term (Q= 1Bq s −1 ) and a reference wind speed ofU= 1m s −1 for each radionuclide, evaluated across di erent release heights and downwind distances. This initial dataset, hereafter referred to as the low-resolution dataset (D LR ), consists of plume shine dose values computed at discrete combinations of physical and meteorological parameters. Speci cally, the dataset includes 17 gamma-emitting radionuclides: 137 Cs, 134 Cs, 41 Ar, 135 Xe, 60 Co, 131 I, 132 I, 87 Kr, 88 Kr, 85 Kr, 85 Sr, 103 Ru, 106 Ru, 22 Na, 152 Eu, 154 Eu, and 155 Eu. These radionuclides were selected to represent a broad and realistic range of plume shine contributors encountered in nuclear and radiological applications. The set includes key ssion products released during reactor accidents and fuel handling events (e.g., Cs, I, and Ru isotopes), activation products relevant to routine opera- tions and research facilities (e.g., 41 Ar, 60 Co, 22 Na), and noble gases that dominate early plume shine immediately after release because of their high mobility and limited 6 deposition (e.g., Xe and Kr isotopes). The inclusion of multiple europium isotopes allows the dataset to represent radionuclides with more complex gammaemission spectra and longer half-lives, which are particularly relevant for facility characteri- zation and long-term dose budgeting. Collectively, the selected radionuclides span a broad range of half-lives, gamma energies, emission multiplicities, and chemical prop- erties, providing a representative testbed for evaluating plume shine dose modeling and machine-learning-based surrogate approaches. The resulting dataset consists of four key input features and one targetvariable, as summarized in Table1. The input features include radionuclide identity, atmospheric stability category, release height, and downwind distance, while the target variable is the corresponding plume shine dose at the receptor location (Fig.1). Table 1: Description of dataset features used for plume shine dose prediction FeatureTypeRange / Categories Description RadionuclideCategorical17 gamma emittersGamma-emitting radionuclide (e.g., 137 Cs, 41 Ar, 131 I, 135 Xe). Atmospheric Stability Category CategoricalA–FPasquill–Gi ord atmospheric stability class. Release HeightNumerical10–200 m (step: 10 m)Release height above ground level. Downwind DistanceNumerical25–2000 mDownwind distance along plume center- line. Plume Shine DoseContinuous (Target)∼10 −13 –10 −8 External gamma dose at receptor (μSv/hr). The plume shine dose spans several orders of magnitude, re ecting variations in radionuclide gamma yield, atmospheric dispersion conditions, and source–receptor geometry. Figure2presents the distribution of dose values across categorical fea- tures on a logarithmic scale, illustrating both central tendencies and variability across radionuclides and stability classes. Cs-137Cs-134 Ar-41 Xe-135 Co-60 I-131I-132 Kr-87Kr-88Kr-85 Sr-85 Ru-103Ru-106 Na-22 Eu-152Eu-154Eu-155 Radionuclide 12 11 10 9 8 7 log 10 (Plume Shine Dose) Plume Shine Dose Across Radionuclides A B C D E F Stability Category Plume Shine Dose Across Atmospheric Stability Categories Fig. 2: Distribution of log 10 plume shine dose across radionuclides (left) and atmo- spheric stability categories A–F (right). The boxen plots show median, spread, and tail behavior, with stability classes A–F representing Pasquill–Gi ord regimes from unstable to stable 7 AlthoughD LR is physically consistent, its discrete structure restricts its e ec- tiveness for training data-driven surrogate models, which typically require quasi- continuous coverage of the input domain. The impact of this discretization on predictive performance is analyzed later in this study. This limitation motivates the interpolation strategy presented in the following subsection. 3.2 High-Resolution Dataset Construction via Stepwise Interpolation To convert the sparse low-resolution into the high-resolution dataset, an interpolation- based densi cation was applied only along the downwind distance, which is irregularly and sparsely discretized in the original data. Radionuclide identityand atmospheric stability category were preserved as categorical variables. Unlike oinesynthetic data augmentation frameworks (e.g., SynthCity) that generate distribution-expanding samples[ 22], the present approach performs deterministic densi cation without intro- ducing synthetic variability beyond the original discrete data manifold. Distance-wise interpolation using PCHIP. One-dimensional interpolation was performed along the downwind distance axis for each unique combination of radionuclide, release height, and atmospheric stability category. The PCHIP method was chosen because it preserves shape and monotonic trends, which are key characteristics of plume shine dose variation with downwind distance and release height. Using PCHIP interpolation, dose values were generated at uniformly spaced dis- tance points between the minimum and maximum observed distances foreach scenario. This procedure enhances spatial resolution while preventing arti cial oscillations or non-physical extrema that may arise from higher-order spline methods. Representative examples of the distance-wise interpolation for di erent radionuclides and stability categories are presented in Figs. 3–5. 3.3 Final Datasets for Model Training and Evaluation The low-resolutionD LR dataset was partitioned into training (D train LR ) and test (D test LR ) subsets, as summarized in Table 2. For learning in the discrete input space, models were trained onD train LR and evaluated on the held-out setD test LR . The high-resolutionD HR dataset is exclusively obtained fromD train LR via interpo- lation. To ensure a purely interpolated domain, all original discrete sampling points present inD train LR were removed, resulting in a re nedD HR . This re ned dataset was then randomly partitioned intoD train HR (99.975%) andD test HR (0.025%), such that both subsets consist solely of interpolated data points and represent a quasi-continuous input domain. D test LR andD test HR were used to evaluate the performance of models trained on low- and high-resolution data in the original discrete domain and the interpolated domain, respectively. 8 0500100015002000 Distance (m) 10 10 10 9 Dose Radionuclides Group 1 Cs-137 (interp) Cs-137 (GT) Cs-134 (interp) Cs-134 (GT) Ar-41 (interp) Ar-41 (GT) Xe-135 (interp) Xe-135 (GT) 0500100015002000 Distance (m) 10 10 10 9 Dose Radionuclides Group 2 Co-60 (interp) Co-60 (GT) I-131 (interp) I-131 (GT) I-132 (interp) I-132 (GT) Kr-87 (interp) Kr-87 (GT) 0500100015002000 Distance (m) 10 12 10 11 10 10 10 9 Dose Radionuclides Group 3 Kr-88 (interp) Kr-88 (GT) Kr-85 (interp) Kr-85 (GT) Sr-85 (interp) Sr-85 (GT) Ru-103 (interp) Ru-103 (GT) 0500100015002000 Distance (m) 10 11 10 10 10 9 Dose Radionuclides Group 4 Ru-106 (interp) Ru-106 (GT) Na-22 (interp) Na-22 (GT) Eu-152 (interp) Eu-152 (GT) Eu-154 (interp) Eu-154 (GT) Eu-155 (interp) Eu-155 (GT) PCHIP Interpolation vs Ground Truth Stability Category A, Release Height = 140 m Fig. 3: Distance-wise comparison between PCHIP-interpolated plume shine dose and numerically calculated ground-truth values for representative radionuclides under sta- bility category A at a release height of 140 m. The agreement demonstrates the shape-preserving and monotonic behavior of the interpolation method. This dataset organization enables a two-level evaluation strategy: model accuracy is assessed both at the original numerically computed sampling locations and at pre- viously unseen interpolated points. Together, these tests provide a rigorous measure of predictive delity and generalization capability across sparsely sampled regions of the physical parameter space. 4 ML Framework and Model Con guration To examine how di erent ML paradigms perform for plume shine dose prediction, three representative regression models were selected: RF, XGBoost, and TabNet. Together, these models cover the main approaches currently used for structured tabular data, namely ensemble learning, boosted decision trees, and attention-based DL model. 9 0500100015002000 Distance (m) 10 9 Dose Radionuclides Group 1 Cs-137 (interp) Cs-137 (GT) Cs-134 (interp) Cs-134 (GT) Ar-41 (interp) Ar-41 (GT) Xe-135 (interp) Xe-135 (GT) 0500100015002000 Distance (m) 10 9 Dose Radionuclides Group 2 Co-60 (interp) Co-60 (GT) I-131 (interp) I-131 (GT) I-132 (interp) I-132 (GT) Kr-87 (interp) Kr-87 (GT) 0500100015002000 Distance (m) 10 12 10 11 10 10 10 9 Dose Radionuclides Group 3 Kr-88 (interp) Kr-88 (GT) Kr-85 (interp) Kr-85 (GT) Sr-85 (interp) Sr-85 (GT) Ru-103 (interp) Ru-103 (GT) 0500100015002000 Distance (m) 10 10 10 9 Dose Radionuclides Group 4 Ru-106 (interp) Ru-106 (GT) Na-22 (interp) Na-22 (GT) Eu-152 (interp) Eu-152 (GT) Eu-154 (interp) Eu-154 (GT) Eu-155 (interp) Eu-155 (GT) PCHIP Interpolation vs Ground Truth Stability Category F, Release Height = 140 m Fig. 4: Distance-wise comparison between PCHIP-interpolated plume shine dose and numerically calculated ground-truth values for representative radionuclides under sta- bility category F at a release height of 140 m. The agreement demonstrates the shape-preserving and monotonic behavior of the interpolation method. The choice of these models is motivated by two considerations centralto this study. First, plume shine dose prediction is governed by a small set of physically mean- ingful variables, including geometric features such as downwind distance and release height, as well as atmospheric dispersion conditions. These variablesdirectly de ne the source–receptor geometry and therefore play a dominant role in dose formation. This structure makes plume shine dose prediction a suitable testcase for examining how di erent model inductive biases learn and prioritize geometry-driven relation- ships, which are analyzed in detail in the Results section.. Second, while tree-based ensemble methods are widely regarded as strong performers on tabular data, recent DL architectures such as TabNet aim to challenge this dominance by introducing atten- tion mechanisms tailored to structured inputs. Comparing these approaches therefore 10 0500100015002000 Distance (m) 10 10 10 9 Dose Cs-137 - Stability Category A Original Data PCHIP Interpolation 0500100015002000 Distance (m) 10 9 3 × 10 10 4 × 10 10 6 × 10 10 Dose Cs-137 - Stability Category B Original Data PCHIP Interpolation 0500100015002000 Distance (m) 3 × 10 10 4 × 10 10 6 × 10 10 Dose Cs-137 - Stability Category C Original Data PCHIP Interpolation 0500100015002000 Distance (m) 3 × 10 10 4 × 10 10 6 × 10 10 Dose Cs-137 - Stability Category D Original Data PCHIP Interpolation 0500100015002000 Distance (m) 3 × 10 10 4 × 10 10 6 × 10 10 Dose Cs-137 - Stability Category E Original Data PCHIP Interpolation 0500100015002000 Distance (m) 3 × 10 10 4 × 10 10 6 × 10 10 Dose Cs-137 - Stability Category F Original Data PCHIP Interpolation PCHIP Interpolation for Cs-137 (Release Height: 140m) Fig. 5: Distance-wise validation of PCHIP interpolation for plume shine dose of 137 Cs at a release height of 140 m. The gure compares numerically calculated doseval- ues and PCHIP-interpolated pro les across Pasquill–Gi ord stability categories A–F, demonstrating that the interpolation preserves the physically expected monotonic decay with distance without introducing non-physical oscillations. provides insight into which modeling strategies are best aligned with the physics and data structure of radiological dose assessment. Below, we provide a brief overview of these models as an introduction. 4.1 Overview of ML Models RF constructs multiple decision trees using bootstrapped samples of the training data and aggregates their predictions through averaging [ 23]. Random feature selection at each split reduces correlation among trees and improves generalization. RF regression was implemented using thescikit-learnframework [24]. XGBoost is a gradient boosting approach in which decision trees are constructed sequentially, with each new tree trained to correct the residual errors of the existing ensemble [13]. The framework includes regularization and structural constraintsthat help limit model complexity and improve predictive stability.In this work, XGBoost regression was implemented using the ocialxgboostlibrary. TabNet is a DL architecture developed speci cally for learning from tabular data [ 14]. It uses a sequential attention mechanism to perform instance-wise feature 11 Table 2: De nition of training and test datasets used for model development and evaluation Symbol Dataset TypeDescription and Purpose D train LR Low-resolution training set 99% subset of the original discrete dataset (∼9×10 4 samples), used for initial model training on numerically computed plume shine dose values. D test LR Low-resolution test set1% held-out subset of the original dataset (∼10 3 samples), reserved for nal evaluation against unseen numerically computed dose values. D train HR High-resolution training set Interpolated dataset after removal of all original low-resolution sampling points, comprising approximately 99.975% of the remaining interpolated samples (∼4×10 6 points), used for train- ing models on a dense and continuous input domain. D test HR High-resolution test setHeld-out subset of the dataset (0.025%,∼10 3 samples), contain- ing only interpolated data points which are used to evaluatemodel generalization at interpolated input-space. selection at each decision step, followed by feature transformation layers whose con- tributions are combined to generate the nal prediction. In this study, TabNet was implemented using the ocial PyTorch-based implementation[25]. Details of model hyperparameter selection and tuning procedures are discussed in the following subsection. 4.2 Data Preprocessing Prior to model training, all datasets were preprocessed to ensure numerical stability and compatibility across the three ML architectures. The preprocessing work ow com- prised target variable transformation, categorical feature encoding, andnormalization of numerical features. The plume shine dose, used as the target variable, exhibits a stronglyright-skewed distribution spanning several orders of magnitude. A base-10 logarithmic scaling was applied to the dose values to reduce the dominance of extreme values and yield a more symmetric distribution that is better suited for regression-based learning. As illustrated in Fig. 6, this transformation substantially reduces skewness (from 5.55 to −1.22), promoting stable optimization and improved convergence. Model predictions were inverse-transformed during post-processing to recover dose values in physical units for evaluation and interpretation. Categorical features, namelyRadionuclideandAtmospheric Stability Category, were encoded using integer representations to preserve category identity while main- taining consistency across models. For TabNet, these encoded features were explicitly speci ed through thecatidxsandcatdimsparameters, which de ne the indices and cardinalities of categorical variables, respectively. Numerical features, including Release HeightandDownwind Distance, were normalized to the range [0,1] using Min–Max scaling to prevent scale imbalance and ensure stable model training. 12 02468 Dose 1e 8 0 10000 20000 30000 40000 50000 Count Skew: 5.55 Original Dose Distribution 2345 1e 8 0 200 400 121110987 Dose_log 0 1000 2000 3000 4000 5000 6000 Count Skew: -1.22 Log Transfromed Dose Distribution Fig. 6: Histogram comparison of plume shine dose distributions before and after base- 10 logarithmic transformation. 4.3 Hyperparameter Optimization Hyperparameter optimization was performed separately for each model, accounting for di erences in model complexity and computational cost. For RF and XGBoost, which are comparatively less computationally intensive, hyperparameters were tuned using a manual, iterative procedure. Key parameters were adjustedheuristically based on observed training and validation performance, with the objective of achieving a balanced trade-o between bias and variance. For TabNet, a hybrid optimization strategy was adopted due to its higher sensi- tivity to hyperparameter selection and increased computational demands. An initial manual search was conducted to identify reasonable ranges for critical hyperparame- ters. Subsequently, automated hyperparameter optimization was performed using the Optuna framework [ 26]. Fifty optimization trials were executed using Optuna’s Tree- structured Parzen Estimator (TPE) sampler, with the objective of minimizing the Root Mean Square Error (RMSE) on the validation set. The tables with optimized hyperparaters are provided in the supporting information (Table S1, S2, S3) 4.4 Performance Metrics for Evaluation and Testing During training, the optimization criteria di ered across models.XGBoost and TabNet were trained by minimizing the Root Mean Square Error (RMSE), while the RF model employed Mean Squared Error (MSE), which is its default loss function. These training losses were used solely for parameter optimization and are not treated as primary comparative metrics in the evaluation on unseen test data. To assess the predictive performance of each model, evaluation metrics that cap- ture both goodness of t and relative prediction accuracy were employed. Speci cally, three complementary metrics were used: the coecient of determination (R 2 ), Mean Absolute Percentage Error (MAPE), and symmetric Mean Absolute Percentage Error (sMAPE). 13 R 2 quanti es the proportion of variance in the reference plume shine dose explained by the model predictions and is de ned as R 2 = 1− ∑ N i=1 (y i − ˆ y i ) 2 ∑ N i=1 (y i − ̄ y) 2 ,(2) wherey i and ˆ y i denote the true and predicted dose values, respectively, ̄ yis the mean of the true values, andNis the number of test samples. MAPE measures the average relative deviation between predicted and reference doses and is given by MAPE = 100 N N ∑ i=1 ∣ ∣ ∣ ∣ y i − ˆ y i y i ∣ ∣ ∣ ∣ .(3) MAPE provides an intuitive percentage-based measure of error; however, it can become unstable or unbounded when reference dose values are very small. To mitigate this limitation, sMAPE is also employed, which normalizes the absolute error by the average magnitude of the true and predicted values and is de ned as sMAPE = 100 N N ∑ i=1 2|y i − ˆ y i | |y i |+| ˆ y i | .(4) Unlike MAPE, which can theoretically range from 0 to∞, sMAPE is bounded within the interval [0,200] and treats overestimation and underestimation symmet- rically. This property is particularly important in the present study, as plume shine dose values often span several orders of magnitude and include extremely small values. Under such conditions, absolute error metrics such as RMSE or MAE may become less informative for comparative evaluation, whereas percentage-based metrics provide a more meaningful assessment of model performance across the full dose range. 5 Results and Discussion 5.1 Model Performance under Discrete and Continuous Data Regimes This subsection examines how training data resolution in uences predictive accuracy and generalization behavior across the three models. Performance is evaluated in two complementary settings: learning from discrete low-resolution dose tables and learning from interpolation-enhanced high-resolution dense datasets. Case I: Training on Low-Resolution Data. As summarized in Table 3, both XGBoost and RF showed strong predictive perfor- mance onD test LR when trained exclusively onD train LR (Table 2). XGBoost yielded the highest coecient of determination (R 2 = 0.999) along with very low error levels (MAPE and sMAPE below 0.5%), indicating excellent agreement with thenumerically computed reference dose values at discrete sampling locations. RF also performed well, 14 achieving similarR 2 of≈0.99 but with slightly higher MAPE and sMAPE error val- ues around 2%. In contrast, TabNet exhibited noticeably weaker performance, with a lowerR 2 value of 0.9561 and substantially higher percentage-based errors. This degra- dation possibly re ects the challenges faced by attention-based neural networks when trained on sparse and discontinuous datasets in which limited sampling density hinders the learning of stable feature representations. Figure S1 (see Supporting information) provides a visual comparison of predicted versus reference plumeshine dose values and con rms the quantitative trends reported in Table3. The tight clustering around the one-to-one line for XGBoost and RF contrasts with the broader scatterobserved for TabNet, consistent with the corresponding di erences inR 2 and error metrics. Poor Generalization to High-Resolution Test Data. When the models trained on low-resolution dataset (i.e.D train LR ) were further evaluated on interpolated test dataset (D test HR ), which consists exclusively of interpolated points, a clear reduction in predictive performance was observed across all three models. As summarized in Table 3, the error metrics increased substantially compared to the evaluation inD test LR , indicating limited generalization beyond the discrete training grid. Although XGBoost retained a high coecient of determination (R 2 = 0.999), its percentage-based errors increased noticeably, with MAPE rising to 2.47% and sMAPE to 2.42%. A similar trend was observed for RF, where MAPE and sMAPE increased to 3.16% and 3.15%, respectively. These results also indicate thatR 2 is alone insucient to characterize generalization quality in dose prediction, as small relative deviations can translate into meaningful absolute errors over several orders of magnitude. TabNet exhibited the most pronounced degradation, withR 2 decreasing to 0.9157 and error metrics exceeding 14% (Fig. S2). This behavior again reiterates the sensi- tivity of attention-based TabNet to sparse training distributions, particularly when predictions are required at intermediate locations not explicitlyrepresented in the training data. Overall, these results demonstrated that models trained solely onD LR primarily learn to reproduce discrete dose tables and do not reliably capture smooth dose varia- tion across the continuous physical domain. The pronounced performance degradation observed when predictions are evaluated at interpolated locations underscores the lim- ited generalization capability of learning from sparsely sampled data. Therefore, these ndings advocate for the training of models with high-resolution training data set. Case I: Training on High-Resolution Data. Motivated by the limited generalization observed in Case I, all three models were retrained using the high-resolution dataset (D train HR ), which provides dense sampling along the dimensions of downwind distance. The retrained models were then eval- uated on both the (D test LR ) andD test HR . This con guration enables assessment of (i) backward consistency with physics-based discrete calculations and(i) generalization across previously unseen continuous regions of the input space. Figure S3 (see Supporting information) showed that models trained onD train HR retain excellent predictive accuracy when evaluated on the low-resolution test data (D test LR ). 15 This indicates that training on interpolated data does not degrade delity to the orig- inal physics-based reference points. Notably, XGBoost achieves anR 2 of 0.9996 with less than 1% MAPE, while RF exhibits comparable performance with slightly higher percentage errors. Further, evaluation on the high-resolution test set (Fig. S4) also demonstrates a substantial improvement in generalization compared toCase I. All models show reduced errors and improved alignment with reference dose values across the interpolated domain. As summarized in Table3, XGBoost maintains the strongest overall performance, with approximately 1% MAPE onD test HR , con rming its abil- ity to robustly capture the underlying functional relationship governing plume shine dose. RF shows moderate degradation relative to XGBoost but the modelperformed signi cantly better than the trained model withD train LR . Table 3: Performance of ML models on plume shine dose predictions evaluated onthe original (physical) dose scale Training–Testing XGBoostRFTabNet R 2 MAPE sMAPER 2 MAPE sMAPER 2 MAPE sMAPE D train LR → D test LR 0.999 0.480.480.99791.921.910.9561 13.9013.94 D train LR → D test HR 0.999 2.472.420.99563.163.150.9157 14.4113.97 D train HR → D test LR 0.999 0.930.930.99752.072.060.99833.973.88 D train HR → D test HR 0.998 1.001.000.99302.222.220.98333.433.36 Interestingly, trained TabNet model showed a marked improvementrelative to Case I for bothD test LR andD test HR dataset, achieving anR 2 of≈0.99 with signi cantly reduced percentage-based errors. This improvement highlights the bene t of dense training data for DL architectures[ 27], which typically require smooth and continuous input–output mappings to extract stable feature representations. Taken together, these ndings established the high-resolution interpolated dataset as a necessary component for reliable surrogate modeling of plume shinedose and jus- tify the interpolation-assisted learning strategy adopted in this work. As discussed, models trained solely on low-resolution data can accurately reproduce known dis- crete points, but they fail to infer smooth transitions between them. This behavior is consistent with recent ndings[28] on tabular ML approaches, where the inductive biases of di erent model families lead to markedly di erent generalization behavior on irregular and sparsely sampled targets like those in plume shine dose tables. The high- resolution training dataset mitigates this limitation by encoding physically consistent trends across distance, enabling all models — particularly XGBoost and TabNet — to generalize e ectively to both discrete numerically computed locations and unseen interpolated regions. 16 5.2 Model Robustness Across Atmospheric Regimes and Radionuclide Classes To gain a deeper understanding of model performance at the regime level for the inter- polated dataset, relative error distributions were analyzed using whisker plots (Fig.7), separately for each atmospheric stability class and radionuclide, for XGBoost, RF, and TabNet. Across all stability categories (Fig.7, top), XGBoost exhibited the smallest median relative errors and the narrowest 10–90% uncertainty bounds, indicating con- sistently stable performance with minimal systematic bias. AlthoughRF maintained good central (median) accuracy, its variance found to be on the higher side. Tab- Net displayed the largest spread of errors across all the stability classes, with most pronounced upper whiskers for stability category (Fig.7, top). The radionuclide-wise error distributions in Fig.7(bottom) indicated that predic- tive performance is primarily in uenced by the intrinsic structure and smoothness of the plume-shine dose eld associated with each isotope. No strict monotonic trend was observed with respect to gamma emission strength or half-life alone, which is expected since, under short-term steady plume release conditions, radioactive decay during atmospheric transport is negligible compared to dispersion-driven e ects. Instead, radionuclides with softer gamma spectra (e.g., Eu isotopes) and non-depositing noble gases exhibit comparatively larger error spread, attributable to stronger atmospheric attenuation and steeper dose–distance gradients that introduce higherfunctional non- linearity. In contrast, high-energy gamma emitters (e.g., 22 Na and 41 Ar), as well as longer-lived radionuclides such as 137 Cs and 60 Co, display comparatively smoother spa- tial attenuation and more regular dose–distance relationships, leading tomore stable prediction performance. Larger uncertainties were particularly evident for noble gases such as 85 Kr, 87 Kr, and 135 Xe, especially for RF and TabNet. Owing to their chem- ically inert and non-depositing nature, noble gas plumes lack environmental removal mechanisms such as dry and wet deposition, resulting in highly localized centerline concentrations and strong spatial dose gradients, particularly under stable atmo- spheric conditions. Consequently, the observed radionuclide-dependent variability in the error statistics is governed more by the smoothness and curvature of the under- lying dose eld—controlled by photon energy and dispersion characteristics—than by radionuclide half-life or identity alone. Further, the faceted error distributions in Fig. 8provide a detailed diagnostic view of prediction accuracy across radionuclide type and atmospheric stability categories. The plots reveal clear regime-dependent behavior, with RF and TabNetgenerally showing larger error spread and more frequent outliers across several stability classes, particularly for noble gas radionuclides, whereas XGBoost maintains comparatively tighter and more consistent error distributions, reiterating the earlier radionuclide-wise observation (Fig.7). Undoubtedly, the consistently lower bias and variance exhibited byXGBoost across both stability-wise and radionuclide-wise analyses underscore itssuitability as apri- mary surrogate model for plume-shine dose estimation. In contrast, the comparatively higher variability observed for RF and TabNet suggests that caution is warranted when deploying these models under regimes with strong spatial dose gradients, particularly in stability-dependent dispersion conditions. 17 ABCDEF Stability Category 0.050 0.025 0.000 0.025 0.050 0.075 0.100 0.125 Relative Error Stability-wise Prediction Error (10 90% whiskers) XGBoost Random Forest TabNet Ar-41 Co-60 Cs-134Cs-137 Eu-152Eu-154Eu-155 I-131I-132 Kr-85Kr-87Kr-88 Na-22 Ru-103Ru-106 Sr-85 Xe-135 Radionuclide 0.050 0.025 0.000 0.025 0.050 0.075 0.100 0.125 Relative Error Radionuclide-wise Prediction Error (10 90% whiskers) XGBoost Random Forest TabNet Fig. 7: Stability-wise (top) and radionuclide-wise (bottom) relative error distributions for XGBoost, RF, and TabNet. Markers denote median relative errors, while whiskers represent the 10 th –90 th percentile range, computed over the combined interpolated and non-interpolated test datasets. 5.3 Radionuclide-Wise Feature Importance and Model Interpretation 5.3.1 Interpretability metrics and evaluation protocol. To ensure a consistent and model-agnostic comparison of feature relevance, all interpretability analyses were conducted on a uni ed test dataset de ned as D test =D test LR ∪D test HR .(5) 18 ABCDEF 0 5 10 15 20 Absolute % Error Radionuclide = Ar-41 ABCDEF 0 5 10 15 20 Radionuclide = Co-60 ABCDEF 0 5 10 15 20 Radionuclide = Cs-134 ABCDEF 0 5 10 15 20 Absolute % Error Radionuclide = Cs-137 ABCDEF 0 5 10 15 20 Radionuclide = Eu-152 ABCDEF 0 5 10 15 20 Radionuclide = Eu-154 ABCDEF 0 5 10 15 20 Absolute % Error Radionuclide = Eu-155 ABCDEF 0 5 10 15 20 Radionuclide = I-131 ABCDEF 0 5 10 15 20 Radionuclide = I-132 ABCDEF 0 5 10 15 20 Absolute % Error Radionuclide = Kr-85 ABCDEF 0 5 10 15 20 Radionuclide = Kr-87 ABCDEF 0 5 10 15 20 Radionuclide = Kr-88 ABCDEF 0 5 10 15 20 Absolute % Error Radionuclide = Na-22 ABCDEF 0 5 10 15 20 Radionuclide = Ru-103 ABCDEF Stability Category 0 5 10 15 20 Radionuclide = Ru-106 ABCDEF Stability Category 0 5 10 15 20 Absolute % Error Radionuclide = Sr-85 ABCDEF Stability Category 0 5 10 15 20 Radionuclide = Xe-135 Model Prediction Error Comparison by Radionuclide & Stability Model XGB RF TabNet Fig. 8: Faceted distributions of absolute percentage error by radionuclide and atmospheric stability category, highlighting regime-dependent variations in model per- formance for XGBoost, RF, and TabNet. 19 For the tree-based models, feature relevance was quanti ed usingconditional permutation importance. Speci cally, for a given radionucliderand featuref, the importance was computed as the expected increase in prediction errorobtained by randomly permuting featurefwithin the radionuclide-speci c subset of the test data, I r (f) =E (x,y)∼D test (r) [ L ( y, ˆ y(x π f ) ) −L ( y, ˆ y(x) )] ,(6) whereD test (r) denotes the subset of test samples corresponding to radionuclider, x π f represents the input vector in which featurefhas been permuted conditional on the remaining features, andL(·) denotes the mean squared error. For TabNet, interpretability was obtained from the model’s intrinsic attention mechanism. At each decision stept, TabNet computes a feature maskM (i) t (f) for samplei, which represents the relative contribution of featurefat that step. Radionuclide-wise feature importance was then computed by aggregating these masks across all samples and decision steps in the uni ed test dataset, A r (f) = 1 |D test (r)| ∑ i∈D test (r) T ∑ t=1 M (i) t (f),(7) whereD test (r) denotes the subset of test samples corresponding to radionuclider, and Tis the total number of TabNet decision steps. This formulation yields a population-level measure of feature utilization that re ects the internal reasoning strategy learned by the TabNet model. For visualization and cross-model comparison, all radionuclide-wise importance vectors were row-wise normalized so that the feature importances foreach radionuclide sum to unity. This transformation yields relative importance pro les that highlight feature dominance patterns within each radionuclide while enablingconsistent quali- tative comparison across RF, XGBoost, and TabNet, despite the di eringattribution mechanisms (permutation-based vs. attention-based). 5.3.2 Qualitative structure of the comparative feature-importance maps The comparative feature-importance maps in Fig. 9highlight both shared physical trends and model-speci c di erences in how input features are utilized. Since the importance values are row-wise normalized for each radionuclide, the maps should be interpreted as showing the relative contribution of features within each radionuclide under a common color scale, rather than absolute importance magnitudes. For both XGBoost and RF, release height consistently emerges as the domi- nant contributor to plume shine dose prediction across nearly all radionuclides, with atmospheric stability and downwind distance acting as a secondary controls. The near-identical importance patterns observed for the two tree-basedmodels indicate that both capture the same underlying geometric and dispersion-driven dependencies governing the dose eld. Such sharply concentrated, dominance-oriented feature hier- archies are commonly observed in tree-based models, as they tend to prioritize features 20 Radionuclide Stability Category Release Height (m) Distance (m) Feature Ar-41 Co-60 Cs-134 Cs-137 Eu-152 Eu-154 Eu-155 I-131 I-132 Kr-85 Kr-87 Kr-88 Na-22 Ru-103 Ru-106 Sr-85 Xe-135 Radionuclide Random Forest Permutation Importance Radionuclide Stability Category Release Height (m) Distance (m) Feature XGBoost Permutation Importance Radionuclide Stability Category Release Height (m) Distance (m) Feature TabNet Attention Masks 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Row-normalized Relative Feature Importance Fig. 9: Radionuclide-wise feature importance across RF, XGBoost, and TabNet model trained onD train HR . RF and XGBoost importance is computed via radionuclide- conditional permutation importance, while TabNet importance is obtainedfrom aggregated attention masks. Importance values are row-wise normalized for each radionuclide, allowing relative comparison of feature dominance patterns across mod- els under a shared color scale, rather than absolute magnitude equivalence between permutation- and attention-based attributions. that provide the largest error reduction through hierarchical splits. In this context, the consistently negligible importance assigned to radionuclide identity further suggests that predictions are primarily controlled by plume geometry and associated dispersion related conditions rather than isotope-speci c characteristics. In contrast, the feature-importance patterns from TabNet appear distinctly di er- ent in which the importance values are more evenly spread across all features, with no single variable exhibiting strong dominance. This smoother allocation suggests that TabNet relies on a more balanced combination of inputs rather than sharply prioritizing speci c physical controls. 5.3.3 Geometry-dispersion-only versus full-feature ablation. An order-invariant exhaustive ablation study was conducted to systematically assess the contribution of each input feature to plume-shine dose prediction by evaluating model performance over all possible feature combinations, thereby eliminating any bias associated with feature ordering. For each subset, the mean RMSE was computed 21 on the predictions, followed by aggregation across subsets with the same number of active features. As shown in the exhaustive ablation curve (Fig.10a), all three models exhibit a monotonic decrease in RMSE as additional features are introduced; however, the most substantial error reduction occurs when increasing the feature set from one to three variables, while the marginal improvement from the fourth feature is comparatively small. This pattern reiterated that the dominant predictive structure is captured primarily by the geometry–dispersion features (releaseheight, stability category, and downwind distance), with radionuclide identity contributing mainly as a secondary conditional re nement. 1234 Number of Active Features 0.0 0.5 1.0 1.5 2.0 2.5 Mean RMSE (Plume Shine Dose) 1e 8 Order-Invariant Exhaustive Feature Ablation XGBoost Random Forest TabNet (a) Random ForestTabNetXGBoost Model D R R + D R + RH R + RH + D R + ST R + ST + D R + ST + RH R + ST + RH + D RH RH + D ST ST + D ST + RH ST + RH + D Feature Combination RMSE Across All Feature Subsets 1 2 3 4 5 RMSE 1e 8 (b) Fig. 10: Ablation analysis for plume-shine dose prediction: (a) Order-invariant mean ablation curve showing mean RMSE (physical dose space) versus number of active features, averaged over all possible feature subsets. (b) The heatmap shows that feature combinations containing release height, stability category, and distance con- sistently yield lower prediction errors across RF, XGBoost, and TabNet, while the addition of radionuclide identity provides comparatively smaller marginal improve- ment. Acronyms: R = Radionuclide, ST = Stability Category, RH = ReleaseHeight, and D = Downwind Distance. The subset heatmap (Fig. 10b) further corroborates this observation by showing consistently lower RMSE for combinations that include release height and stabil- ity, whereas subsets lacking these variables yield signi cantly higher errors across all models. Notably, XGBoost achieves the lowest RMSE across nearly all subset con gu- rations, followed closely by Random Forest, while TabNet consistentlyexhibits higher error, suggesting that tree-based ensembles more eciently exploit the low-dimensional geometry–dispersion structure governing the dose eld. 22 5.3.4 Origin of the performance hierarchy among models The observed performance ordering, XGBoost>RF>TabNet, can be under- stood in terms of the interaction between model inductive bias andthe underlying physical structure of the plume-shine dose eld, which is primarily governed by geom- etry–dispersion features (release height, stability category, and downwind distance), with radionuclide identity acting as a secondary conditional modi er. The exhaustive ablation analysis further supports this interpretation by showing that the majority of error reduction is achieved when geometry–dispersion variables are included, while the marginal gain from radionuclide information remains comparatively smaller. XGBoost is particularly well aligned with this structure. Its sequential boost- ing mechanism e ectively captures the dominant geometry–dispersion dependencies and incrementally re nes residual errors associated with conditional e ects such as radionuclide characteristics and stability-dependent dispersion. This residual- correction behavior is consistent with prior ndings that boosted tree ensembles perform strongly on tabular problems governed by low-dimensional and smooth func- tional relationships [16,28], thereby explaining its consistently superior predictive accuracy. RF, in contrast, captures the dominant geometry–dispersion trends ina robust and stable manner but exhibits slightly reduced sensitivity to ner conditional re nements. This behavior is characteristic of bagged tree ensembles, which emphasize stable global patterns through averaging rather than sequential residual correction[12]. As a result, while RF achieves performance close to XGBoost, it derives comparatively smaller gains from additional conditional information, consistent with the ablationresults. TabNet exhibits a di erent inductive bias due to its instance-wise attention mech- anism, which distributes feature utilization more exibly across samples instead of enforcing a single global importance hierarchy. Although this leads to richer and more heterogeneous feature usage, it can result in less ecient allocation ofmodel capac- ity in settings where a low-dimensional geometry–dispersion structure dominates the target function. This observation is consistent with recent studies showing that deep learning models for tabular data often underperform tree-based ensembles when strong low-order structure governs the prediction task [16,28]. Overall, the performance hierarchy therefore arises not merely fromdi erences in model expressiveness, but from how e ectively each architecture concentrates its learning on the dominant geometry–dispersion controls of the plume-shine dose eld while incorporating radionuclide-dependent e ects as secondary re nements. 5.4 Web-based deployment and model accessibility. To enhance the practical usability of the proposed surrogate modeling framework, a web-based graphical user interface (GUI) has been developed for interactive plume shine dose assessment. The interface allows users to specify keyscenario parameters, namely radionuclide type, downwind distance, release height, and atmospheric stability category. Based on these inputs, plume shine dose predictions are generated indepen- dently using the trained XGBoost, RF, and TabNet models. In addition, the GUI computes the corresponding reference dose values using adaptive Gauss quadrature 23 Fig. 11: Web-based graphical user interface for interactive plume-shine dose assess- ment, illustrating scenario input parameters (radionuclide, downwind distance, release height, and atmospheric stability) and the corresponding dose predictions obtained from photon-transport-based calculations (labeled as PLUMEDOSENET in GUI interface) and trained machine-learning models. method following Eq. 1, enabling direct side-by-side comparison between numeri- cal and machine-learning-based predictions. This deployment-oriented implementation facilitates rapid scenario evaluation, transparent model inter-comparison, and user- driven sensitivity exploration, thereby bridging the gap betweenmethodological development and operational radiological consequence assessment. 6 Conclusion The application of ML to radiation dose assessment remains nontrivial dueto safety- critical requirements, limited training-ready datasets, and the strong role of governing physical processes. Plume shine dose estimation provides a relevant benchmark prob- lem in this regard, as it is essential for safety analysis and emergency response, yet remains computationally demanding when evaluated using photon-transport- based approaches. To address these intriguing challenges, this work developed an interpolation-assisted machine learning framework based on discreteplume-shine 24 dose datasets generated using the pyDOSEIA program suite for 17 gamma-emitting radionuclides across multiple dispersion scenarios. These datasetswere subsequently augmented using shape-preserving interpolation method to construct dense, high- resolution training datasets suitable for ML-based modeling. Two tree-based ML models–XGBoost and RF–and one deep neural network–based model, TabNet, were systematically evaluated for the plume shine dose prediction. The results demonstrate that the interpolation-assisted data augmentation strat- egy substantially improves generalization across distances, release heights, and radionuclide classes. Among all the ML models, tree-based ensemblemethods– particularly XGBoost–achieve the highest prediction accuracy. RF provides robust performance but is less sensitive to ner variations, while TabNet does not yield comparable accuracy gains for this problem. Further, interpretability analysis using permutation importance(for tree-based models) and attention-based feature attribution (for TabNet) provided useful insight into the factors governing model behavior. The results indicate that performance dif- ferences are largely linked to how each model utilizes input features, with TabNet distributing attention across a broader set of variables, while tree-based models concen- trate more strongly on dominant geometry–dispersion features, treating radionuclide identity primarily as a secondary conditional input. This di erence in feature utiliza- tion is consistent with the ablation ndings and helps explain the observed variation in prediction accuracy across model architectures. To support practical application of the proposed framework, a web-based graph- ical user interface was developed, allowing interactive scenario evaluation and direct comparison between machine-learning-based predictions and photon-transport-based reference calculations. Together, these elements enhance both thetransparency and usability of the proposed approach for plume shine dose assessment. Overall, the proposed interpolation-assisted ML framework provides a fast and accurate surro- gate approach for plume shine dose estimation, o ering a practical alternative to conventional numerical calculations. Supplementary information.Detailed information on the optimized hyperpa- rameters, along with additional performance plots obtained from model evaluations on datasets of varying resolutions, is provided in the SupplementaryInformation. Acknowledgements.B.S. and S.A. extend their thanks to Shri Probal Chaudhury (Associate Director, HS & E Group) for his unwavering support and encouragement throughout the project. Declarations Code availability.The source code of this work is accessible at: https://github.com/BiswajitSadhu/PlumeDoseNet. The dataset and trained models are available at https://doi.org/10.5281/zenodo.18266001. References [1] Sadhu, B., Sadhu, T., Anand, S.: Raddqn: A deep q learning-based architecture for 25 nding time-ecient minimum radiation exposure pathway. IEEE Transactions on Neural Networks and Learning Systems (2025) [2] Luo, Y., Chen, S., Valdes, G.: Machine learning for radiation outcome modeling and prediction. Medical Physics47(5), 178–184 (2020) [3] McKenna, T.J.: Protective action recommendations based upon plant conditions. Journal of hazardous materials75(2-3), 145–164 (2000) [4] Armand, P., Achim, P., Monfort, M., Carr`ere, J., Oldrini, O., Commanay, J., Albergel, A.: Simulation of the plume gamma exposure rate with 3d lagrangian particle model spray and post-processor cloud-shine. In: HARMO10 2005: Proceedings of the 10th International Conference on Harmonisation Within Atmospheric Dispersion Modelling for Regulatory Purposes, p. 545–550 (2005) [5] Satoh, D., Nakayama, H., Furuta, T., Yoshihiro, T., Sakamoto, K.: Simulation code for estimating external gamma-ray doses from a radioactive plume andcon- taminated ground using a local-scale atmospheric dispersion model. Plos one 16(1), 0245932 (2021) [6] Karmakar, S., Srinivas, C., Rakesh, P., Gopalakrishnan, V., Chandrasekaran, S., Athmalingam, S., Venkatraman, B.: Development of a numerical model forsector- average plume gamma dose and its validation with dose rate measurements at kalpakkam npp site, india. Journal of Environmental Radioactivity255, 107029 (2022) [7] McNaughton, M.W., Gillis, J.M., Ruedig, E., Whicker, J.J., Fuehne, D.P.: Accu- racy of cloudshine gamma dose calculations in the cap-88 dispersion model. Health Physics112(4), 414–419 (2017) [8] Cember, H., Johnson, T.E.: Introduction to Health Physics, 4th edn. McGraw-Hill Medical, New York (2009) [9] Till, J.E., Grogan, H.A. (eds.): Radiological Risk Assessment and Environmental Analysis. Oxford University Press, New York (2008) [10] Sadhu, B., Sarkar, T., Anand, S., Singh, K.D., Aswal, D.K.: pydoseia:A python package for radiological impact assessment during long-term or accidental atmo- spheric releases. Health Physics130(1), 94–110 (2026) https://doi.org/10.1097/ HP.0000000000002014 [11] Armand, P., Oldrini, O., Duchenne, C., Perdriel, S.: Topical 3dmodelling and simulation of air dispersion hazards as a new paradigm to support emergency preparedness and response. Environmental Modelling & Software143, 105129 (2021) [12] Breiman, L.: Random forests. Machine learning45(1), 5–32 (2001) 26 [13] Chen, T., Guestrin, C.: Xgboost: A scalable tree boosting system. CoRR abs/1603.02754(2016) [14] Arik, S. ̈ O., P ster, T.: Tabnet: Attentive interpretable tabular learning. In: Pro- ceedings of the AAAI Conference on Arti cial Intelligence, vol. 35, p. 6679–6687 (2021) [15] Fayaz, S.A., Zaman, M., Kaul, S., Butt, M.A.: Is deep learning on tabular data enough? an assessment. International Journal of Advanced Computer Science and Applications13(4) (2022)https://doi.org/10.14569/IJACSA.2022.0130454 [16] Shwartz-Ziv, R., Armon, A.: Tabular data: Deep learning is not all you need. Information Fusion81, 84–90 (2022) [17] Fritsch, F.N., Carlson, R.E.: Monotone piecewise cubic interpolation. SIAM Journal on Numerical Analysis17(2), 238–246 (1980) [18] Molnar, C.: Interpretable Machine Learning: A Guide for Making Black Box Models Explainable, 1st edn. Leanpub, ??? (2019). Version published 2019-02-21; Creative Commons Attribution-NonCommercial-ShareAlike 4.0. http://leanpub.com/interpretable-machine-learning [19] Pasquill, F.: The estimation of the dispersion of windborne material. Meteoro. Mag.90, 20–49 (1961) [20] Hukkoo, R., Bapat, V., Shirvaikar, V.: Manual of dose evaluation from atmo- spheric releases. Technical report, Bhabha Atomic Research Centre (1988) [21] Pecha, P., Pechova, E.: An unconventional adaptation of a classical gaussian plume dispersion scheme for the fast assessment of external irradiation from a radioactive cloud. atmospheric environment89, 298–308 (2014) [22] Qian, Z., Cebere, B.-C., Schaar, M.: Synthcity: facilitating innovative use cases of synthetic data in di erent data modalities. arXiv preprint arXiv:2301.07573 (2023) [23] Breiman, L.: Random forests. Machine Learning45(1), 5–32 (2001) https://doi. org/10.1023/A:1010933404324 [24] Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B.,Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V.,et al.: Scikit-learn: Machine learning in python. the Journal of machine Learning research12, 2825–2830 (2011) [25] Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G.,Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al.: Pytorch: An imperative style, high- performance deep learning library. Advances in neural information processing 27 systems32(2019) [26] Akiba, T., Sano, S., Yanase, T., Ohta, T., Koyama, M.: Optuna: A next- generation hyperparameter optimization framework. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, p. 2623–2631 (2019) [27] Ye, H.-J., Liu, S.-Y., Cai, H.-R., Zhou, Q.-L., Zhan, D.-C.: A CloserLook at Deep Learning Methods on Tabular Datasets. v4: 07 Nov 2025; TALENT benchmark: 300+ datasets evaluating tree vs. deep tabular models (2024).https://doi.org/ 10.48550/arXiv.2407.00956 [28] Grinsztajn, L., Oyallon, E., Varoquaux, G.: Why do tree-based models still outperform deep learning on tabular data? arXiv preprint arXiv:2207.08815 (2022) 28 Supporting Information: Interpolation-Driven Machine Learning Approaches for Plume Shine Dose Estimation: A Comparison of XGBoost, Random Forest, and TabNet Biswajit Sadhu, ∗,†,‡ Kalpak Gupte, ¶ Trijit Sadhu, ¶ and S. Anand §,‡ †Health Physics Division, Health Safety & Environment Group, Bhabha Atomic Research Center, Mumbai – 400085, India ‡Homi Bhabha National Institute, Mumbai - 400094, India ¶Birla Institute of Technology And Science, Pilani, Rajasthan – 333031, India §Health Physics Division, Health Safety and Environment Group, Bhabha Atomic Research Center, Mumbai – 400085, India E-mail: bsadhu@barc.gov.in,biswajit.chem001@gmail.com S1 List of Figures S1Performance of ML/DL models on low-resolution numerical test dataset:Actual versus predicted plume shine dose obtained using XGBoost, RF and TabNet models trained onD train LR and evaluated onD test LR . Results are presented on a logarithmic scale for clarity, with the dashed line indicating the ideal 1:1 prediction. XGBoost exhibits near-perfect agreement with the reference values (R 2 = 1.000, MAPE = 0.48%), followed by RF (R 2 = 0.998, MAPE = 1.9%), while TabNet shows comparatively larger deviations (R 2 = 0.956, MAPE≈14%). The comparison highlights the superior predictive ca- pability of tree-based ensemble models for plume shine doseestimation across multiple orders of magnitude. . . . . . . . . . . . . . . . . . . . . . . . . .. 4 S2Generalization test on interpolated test dataset:Actual versus pre- dicted plume shine dose obtained using XGBoost, RF and TabNetmodels trained onDLR train and evaluated onDHR test . Results are presented on a logarithmic scale for clarity, with the dashed line indicating the ideal 1:1 prediction. Compared with evaluation onD test LR , all models exhibit a notice- able reduction in predictive accuracy, indicating limitedgeneralization across resolution regimes. XGBoost shows degraded but still relatively strong per- formance (R 2 = 0.999, MAPE = 2.5%), followed by RF (R 2 = 0.996, MAPE = 3.2%), while TabNet displays substantially larger deviations (R 2 = 0.916, MAPE≈14%). These results suggest that models trained on low-resolution data do not fully generalize to high-resolution distributions, with tree-based ensemble methods remaining comparatively more robust thandeep tabular architectures. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 S2 S3Performance of ML/DL models under cross-resolution evaluation: Actual versus predicted plume shine dose obtained using XGBoost, RF and TabNet models trained onD train HR and evaluated onD test LR . Results are presented on a logarithmic scale for clarity, with the dashed line indicating the ideal 1:1 prediction. All models demonstrate strong predictive agreement with the reference values, with XGBoost achieving near-perfect performance (R 2 = 1.000, MAPE = 0.93%), followed by RF (R 2 = 0.997, MAPE = 2.1%) and TabNet (R 2 = 0.998, MAPE = 4.0%). The improved performance relative to the reverse evaluation case indicates that models trained on high-resolution data exhibit better generalization when applied to lower-resolution distributions. 6 S4Performance of ML/DL models on high-resolution numerical dataset: Actual versus predicted plume shine dose obtained using XGBoost, RF and TabNet models trained onD train HR and evaluated onD test HR . Results are presented on a logarithmic scale for clarity, with the dashed line indicating the ideal 1:1 prediction. All models demonstrate strong predictive agreement with the ref- erence values, with XGBoost achieving the highest accuracy(R 2 = 0.999, MAPE = 1.0%), followed by RF (R 2 = 0.993, MAPE = 2.2%) and TabNet (R 2 = 0.983, MAPE = 3.4%). The results indicate that increasing data res- olution improves model learning and reduces prediction uncertainty across multiple orders of magnitude. . . . . . . . . . . . . . . . . . . . . . . . . .. 7 S3 List of Tables S1 Optimized hyperparameters used for XGBoost model training. The table summarizes the parameter search space, default settings, and nal optimized values obtained through hyperparameter tuning. . . . . . . . . .. . . . . . . 1 S2 Optimized Random Forest (RF) model hyperparameters used for plume shine dose prediction. The table lists the default parameter values and the corre- sponding optimized settings adopted during model training. . . . . . . . . . 2 S3 Optimized hyperparameter con guration for the TabNet model obtained using Optuna-based hyperparameter tuning. The table reports theparameter search space, default settings, and nal optimized values used forplume shine dose prediction. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 S4 Dataset Descriptions Each dataset includes feature columns such asRelease Height,Downwind Distance,Radionu- clide, andStability Category, with the target variable being the plume shine dose (after log 10 transformation during training). Hyperparameter Con gurations Table S1: Optimized hyperparameters used for XGBoost modeltraining. The table sum- marizes the parameter search space, default settings, and nal optimized values obtained through hyperparameter tuning. ParameterDistribution / TypeDefaultOptimized Value Objective FunctionFixedreg:squarederror reg:squarederror Evaluation MetricFixedrmse rmse Max Tree DepthUniform Integer630 Learning Rate (η)Log Uniform0.30.05 Subsample RatioUniform1.00.5 Column Subsample by TreeUniform1.01.0 Number of Boosting RoundsFixed100100 Early Stopping RoundsFixed1010 DeviceCategoricalcpu auto (cpu/gpu) 1 Table S2: Optimized Random Forest (RF) model hyperparameters used for plume shine dose prediction. The table lists the default parameter values and the corresponding optimized settings adopted during model training. ParameterDefault Optimized Value Number of Trees (n estimators )10060 Maximum Tree Depth (max depth) None15 Maximum Features (max features)Auto1.0 Bootstrap SamplingTrueTrue Random Seed—3007 Number of Jobs1−1 (use all cores) 2 Table S3: Optimized hyperparameter con guration for the TabNet model obtained using Optuna-based hyperparameter tuning. The table reports theparameter search space, default settings, and nal optimized values used for plume shine dose prediction. ParameterDistribution / Type DefaultOptimized Value Feature Transformer Dimension (n d )Discrete Uniform816 Attention Dimension (n a )Discrete Uniform816 Decision Steps (n steps )Discrete Uniform510 Gamma (Sparsity Regularization)Uniform1.01.0336 Sparse Regularization (λ sparse )Log Uniform1×10 −5 2.73×10 −6 OptimizerCategoricalAdam Adam Learning Rate (lr)Log Uniform0.010.0063 LR SchedulerCategoricalNoneReduceLROnPlateau Scheduler PatienceInteger105 Scheduler Decay Factor (γ)Uniform0.950.9 Batch SizeDiscrete Uniform10242048 Virtual Batch SizeDiscrete Uniform128256 Maximum EpochsFixed100100 Early Stopping PatienceFixed1010 Evaluation MetricFixedRMSERMSE DeviceCategoricalcpu auto (cpu/gpu) 3 10 12 10 11 10 10 10 9 10 8 Actual Dose 10 12 10 11 10 10 10 9 10 8 Predicted Dose r2_score = 1.000 MAPE= 0.48% XGB Predictions Ideal Prediction 10 12 10 11 10 10 10 9 10 8 Actual Dose 10 12 10 11 10 10 10 9 10 8 Predicted Dose r2_score = 0.998 MAPE= 1.9% RF Predictions Ideal Prediction 10 12 10 11 10 10 10 9 10 8 Actual Dose 10 12 10 11 10 10 10 9 10 8 Predicted Dose r2_score = 0.956 MAPE= 1.4e+01% TabNet Predictions Ideal Prediction Actual vs. Predicted Dose Figure S1:Performance of ML/DL models on low-resolution numerical test dataset:Actual versus predicted plume shine dose obtained using XGBoost, RF and Tab- Net models trained onD train LR and evaluated onD test LR . Results are presented on a logarithmic scale for clarity, with the dashed line indicating the ideal1:1 prediction. XGBoost exhibits near-perfect agreement with the reference values (R 2 = 1.000, MAPE = 0.48%), followed by RF (R 2 = 0.998, MAPE = 1.9%), while TabNet shows comparatively largerdeviations (R 2 = 0.956, MAPE≈14%). The comparison highlights the superior predictive capability of tree-based ensemble models for plume shine dose estimationacross multiple orders of mag- nitude. 4 10 12 10 11 10 10 10 9 10 8 10 7 Actual Dose 10 12 10 11 10 10 10 9 10 8 10 7 Predicted Dose r2_score = 0.999 MAPE= 2.5% XGB Predictions Ideal Prediction 10 12 10 11 10 10 10 9 10 8 10 7 Actual Dose 10 12 10 11 10 10 10 9 10 8 10 7 Predicted Dose r2_score = 0.996 MAPE= 3.2% RF Predictions Ideal Prediction 10 12 10 11 10 10 10 9 10 8 10 7 Actual Dose 10 12 10 11 10 10 10 9 10 8 10 7 Predicted Dose r2_score = 0.916 MAPE= 1.4e+01% TabNet Predictions Ideal Prediction Actual vs. Predicted Dose Figure S2:Generalization test on interpolated test dataset:Actual versus predicted plume shine dose obtained using XGBoost, RF and TabNet modelstrained onDLR train and evaluated onDHR test . Results are presented on a logarithmic scale for clarity, with the dashed line indicating the ideal 1:1 prediction. Compared with evaluation onD test LR , all models exhibit a noticeable reduction in predictive accuracy, indicating limited generalization across resolution regimes. XGBoost shows degraded but still relatively strong performance (R 2 = 0.999, MAPE = 2.5%), followed by RF (R 2 = 0.996, MAPE = 3.2%), while TabNet displays substantially larger deviations (R 2 = 0.916, MAPE≈14%). These results suggest that models trained on low-resolution data do not fully generalize to high-resolution distributions, with tree-based ensemble methods remaining comparativelymore robust than deep tabular architectures. 5 10 12 10 11 10 10 10 9 10 8 Actual Dose 10 12 10 11 10 10 10 9 10 8 Predicted Dose r2_score = 1.000 MAPE= 0.93% XGB Predictions Ideal Prediction 10 12 10 11 10 10 10 9 10 8 Actual Dose 10 12 10 11 10 10 10 9 10 8 Predicted Dose r2_score = 0.997 MAPE= 2.1% RF Predictions Ideal Prediction 10 12 10 11 10 10 10 9 10 8 Actual Dose 10 12 10 11 10 10 10 9 10 8 Predicted Dose r2_score = 0.998 MAPE= 4.0% TabNet Predictions Ideal Prediction Actual vs. Predicted Dose Figure S3:Performance of ML/DL models under cross-resolution evaluation:Ac- tual versus predicted plume shine dose obtained using XGBoost, RF and TabNet models trained onD train HR and evaluated onD test LR . Results are presented on a logarithmic scale for clarity, with the dashed line indicating the ideal 1:1 prediction. All models demonstrate strong predictive agreement with the reference values, with XGBoost achieving near-perfect performance (R 2 = 1.000, MAPE = 0.93%), followed by RF (R 2 = 0.997, MAPE = 2.1%) and TabNet (R 2 = 0.998, MAPE = 4.0%). The improved performance relative to the re- verse evaluation case indicates that models trained on high-resolution data exhibit better generalization when applied to lower-resolution distributions. 6 10 12 10 11 10 10 10 9 10 8 10 7 Actual Dose 10 12 10 11 10 10 10 9 10 8 10 7 Predicted Dose r2_score = 0.999 MAPE= 1.0% XGB Predictions Ideal Prediction 10 12 10 11 10 10 10 9 10 8 10 7 Actual Dose 10 12 10 11 10 10 10 9 10 8 10 7 Predicted Dose r2_score = 0.993 MAPE= 2.2% RF Predictions Ideal Prediction 10 12 10 11 10 10 10 9 10 8 10 7 Actual Dose 10 12 10 11 10 10 10 9 10 8 10 7 Predicted Dose r2_score = 0.983 MAPE= 3.4% TabNet Predictions Ideal Prediction Actual vs. Predicted Dose Figure S4:Performance of ML/DL models on high-resolution numerical dataset: Actual versus predicted plume shine dose obtained using XGBoost, RF and TabNet models trained onD train HR and evaluated onD test HR . Results are presented on a logarithmic scale for clarity, with the dashed line indicating the ideal 1:1 prediction. All models demonstrate strong predictive agreement with the reference values, with XGBoost achieving the highest accuracy (R 2 = 0.999, MAPE = 1.0%), followed by RF (R 2 = 0.993, MAPE = 2.2%) and TabNet (R 2 = 0.983, MAPE = 3.4%). The results indicate that increasing data resolution improves model learning and reduces prediction uncertainty across multiple orders of mag- nitude. 7