Paper deep dive
Deep Sigma Point Processes for RCS Modeling in Spaceborne SAR Imagery
Khalid El-Darymli, Christoph H. Gierull, Katerina Biron, Weimin Huang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 7/28/2026, 3:56:44 AM
Summary
This paper introduces a Deep Sigma Point Process (DSPP) model for predicting Radar Cross-Section (RCS) in spaceborne Synthetic Aperture Radar (SAR) imagery. Utilizing a dataset of 208,191 AIS-verified ships from RADARSAT-2, the DSPP employs a hierarchical Gaussian process framework with Bayesian inference to predict RCS while characterizing uncertainty. The model outperforms linear regression baselines, achieving significant reductions in RMSE and improvements in R-squared, while providing calibrated uncertainty bounds and feature importance rankings through a Matérn kernel with automatic relevance determination.
Entities (8)
Relation Signals (5)
Khalid El-Darymli → affiliatedwith → Defence Research and Development Canada
confidence 99% · Khalid El-Darymli... are with... Defence Research and Development Canada
Deep Sigma Point Process → uses → RADARSAT-2
confidence 98% · using a RADARSAT-2 dataset containing 208,191 verified ships
Deep Sigma Point Process → predicts → Radar Cross-Section
confidence 97% · DSPP model for predicting RCS in synthetic aperture radar (SAR) imagery
Deep Sigma Point Process → outperforms → Linear Regression
confidence 95% · Performance evaluations demonstrate the model's superiority over linear regression baselines
Deep Sigma Point Process → employs → Matérn Kernel
confidence 92% · Using a Matern kernel with automatic relevance determination
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Radar cross-section (RCS) modeling is foundational to advancing the utility and sensitivity of spaceborne radar systems. This study introduces a deep sigma-point process (DSPP) model for predicting RCS in synthetic aperture radar (SAR) imagery using a RADARSAT-2 dataset containing 208,191 verified ships. The DSPP model not only strives for predictive accuracy but also characterizes the uncertainty inherent in the intricate relationships among radar signals, ship parameters, and environmental conditions. Unlike traditional approaches that rely on deterministic equations with static parameters, the DSPP uses a hierarchical Gaussian process framework with Bayesian inference to capture variability and uncertainty in RCS predictions. By generating predictive distributions rather than single estimates, the model accounts for the complex dynamics governing radar returns. Using a Matern kernel with automatic relevance determination, the DSPP identifies and ranks critical features across radar, operational, and environmental domains, thereby supporting transparency and interpretability. Performance evaluations demonstrate the model's superiority over linear regression baselines, with a 20.83 percent reduction in root mean squared error, a 25.89 percent increase in R-squared, and a 44.4 percent reduction in both the residual interquartile range and median absolute deviation on the test data. By providing calibrated uncertainty bounds, the DSPP enhances prediction reliability and supports robust decision-making. This work represents a shift toward probabilistic models that incorporate the inherent uncertainty of complex phenomena. By transitioning from fixed equations to distributions over outcomes, the DSPP fosters a deeper understanding of RCS behavior and enables systems to operate effectively in dynamic environments.
Tags
Links
- Source: https://arxiv.org/abs/2607.21745v1
- Canonical: https://arxiv.org/abs/2607.21745v1
Trouble viewing inline? Open PDF directly →
Full Text
94,635 characters extracted from source content.
Expand or collapse full text
AUTHOR PREPRINT • arXiv SUBMISSION DRAFTVersion: 23 July 2026 Deep Sigma Point Processes for RCS Modeling in Spaceborne SAR Imagery Khalid El-Darymli • Christoph H. Gierull • Katerina Biron • Weimin Huang Defence Research and Development Canada, Ottawa, Ontario, Canada Memorial University of Newfoundland, St. John’s, Newfoundland and Labrador, Canada Version notice Draft designation and cover page added for preprint distribution; scientific content is unchanged. Corresponding author: khalid.el-darymli@drdc-rddc.gc.ca Version notice This file is prepared as an author preprint for scholarly dissemination through arXiv. It includes the complete article content and figures. Deep Sigma Point Processes for RCS Modeling in Spaceborne SAR Imagery KHALID EL-DARYMLI, Senior Member, CHRISTOPH H. GIERULL , Senior Member, KATERINA BIRON Defence Research and Development Canada, Ottawa, ON, Canada WEIMIN HUANG , Senior Member, Memorial University, St. John’s, NL, Canada Radar cross-section (RCS) modeling is foundational to advancing the utility and sensitivity of spaceborne radar systems. This study in- troduces a deep sigma point process (DSPP) model for predicting RCS in synthetic aperture radar (SAR) imagery, using a RADARSAT-2 dataset containing 208 191 verified ships. The DSPP model not only strives for predictive accuracy but ventures to characterize the un- certainty inherent in the intricate relationships among radar signals, ship parameters, and environmental conditions. Unlike traditional approaches relying on deterministic equations with static parameters, the DSPP leverages a hierarchical Gaussian process framework with Bayesian inference to capture variability and uncertainty in RCS predictions. By generating predictive distributions rather than single estimates, the model effectively accounts for the complex dynam- ics governing radar returns. Using a Matérn kernel with automatic relevance determination, the DSPP identifies and ranks critical fea- tures across radar, operational, and environmental domains, ensuring transparency and interpretability. Performance evaluations demon- strate the model’s superiority over linear regression baselines, achiev- ing a 20.83% reduction in root mean squared error, a 25.89% increase inR 2 , and a 44.4% reduction in both residual interquartile range and median absolute deviation on test data. By providing calibrated uncer- tainty bounds, the DSPP enhances prediction reliability and supports robust decision-making. This work marks a shift toward probabilistic Received 17 December 2024; revised 13 May 2025 and 9 June 2025; accepted 13 June 2025. Date of publication 24 June 2025; date of current version 13 October 2025. Refereeing of this contribution was handled by K. S Kulpa. Authors’ addresses: Khalid El-Darymli, Christoph Gierull, and Kate- rina Biron are with the Department of National Defence, Defence Re- search and Development Canada, Ottawa, ON K1A 0Z4, Canada, E-mail: (khalid.el-darymli@drdc-rddc.gc.ca; christoph.gierull@drdc-rddc.gc.ca; katerina.biron@drdc-rddc.gc.ca); Weimin Huang is with the Department of Electrical and Computer Engineering, Memorial University of New- foundland, St. John’s, NL A1B 3X5, Canada, E-mail: (weimin@mun.ca). (Corresponding author: Khalid El-Darymli.) 0018-9251 © 2025 Crown Copyright models that incorporate the inherent uncertainty of complex phenom- ena.Transitioningfromfixedequationstodistributionsoveroutcomes, the DSPP fosters a deeper understanding of RCS behavior, enabling systems to thrive in dynamic operational environments. I. INTRODUCTION Empirical RCS models are critical for accurately pre- dicting the performance of spaceborne synthetic aperture radar (SAR) systems in maritime vessel detection. This is because these models enable precise estimations of radar reflectivity, which directly impacts the system’s ability to detect and classify vessels in diverse environmental and operational conditions. Especially under increasingly strin- gent surveillance requirements, such predictions are essen- tial for optimizing radar sensitivity and ensuring consis- tent detection across varying maritime scenarios. Beyond performance prediction, these models are instrumental in defining requirements for the development of new ship detection modes (SDMs). As the Department of National Defence (DND) and Canadian Armed Forces (CAF) shift focus toward detecting smaller vessels, current semiempir- ical models—based largely on ship length—fall short by neglecting key factors including standard and extended op- erating conditions[1],[2],[3],[4]. Developing more robust empirical RCS models that account for these variables is essential for improving the design of SDMs and ensuring reliable performance of future radar systems, particularly for projects, such as the defence enhanced surveillance from space (DESSP)[5]. According to theRecommended Practice for RCS Test ProceduresStd 1502-2020), developed by the Antennas and Propagation Standards Committee, RCS, de- noted byσand measured in square metres (m 2 ), quantifies a target’s ability to reflect radar energy toward the receiver. RCS is defined as 4πtimes the ratio of scattered power per unit solid angle in a specified direction to the incident power per unit area. This measure represents an area that would intercept an equivalent power, radiating isotropically, to produce the same received power at the radar receiver. Mathematically, the RCS is given by[6] σ [m 2 ]=4πlim R→∞ R 2 | E s | 2 | E i | 2 (1) where 1)Ris the distance from the target to the receiver; 2)E s is the scattered electric field; 3)E i is the incident electric field. TheR→∞assumption ensures far-field (i.e., planar wave) conditions, making RCS an intrinsic target property, independent of measurement distance. This work focuses on monostatic radar (i.e., backscatter RCS), where the in- cident and reflected scattering directions are coincident but opposite in sense. In practical applications, the RCS is often expressed in decibels relative to a square metre (dBsm) due to its wide AUTHOR PREPRINT - arXiv submission draft 1 dynamic range. That is, σ [dBsm]=10 log 10 ( σ [m 2 ] 1[m 2 ] ) .(2) ThefirstempiricalRCSmodelfornavalshipswasdevel- oped by Merrill I. Skolnik at the Naval Research Laboratory (NRL) inthe 1970s. Inhis1973NRLreport[7]andhis1974 Transactions on Aerospace and Electronic Systems paper[8], Skolnik introduced the following formula for RCS: RCS =52f 1/2 c D 3/2 (3) where 1) RCS is in (m 2 ); 2)f c is the radar center frequency in megahertz (MHz); 3)Dis the ship’s displacement in kilotons. This formula applies to ships in the microwave band, with RCS measurements based on naval vessels ranging from 2–17 kilotons. Skolnik’s model provides a median RCSvalueforshipdisplacementsatX-band(3.25-cmwave- length), S-band (10.7 cm), and L-band (23 cm). The RCS values are averaged over port and starboard bow and quarter aspects, excluding the broadside peak. As Skolnik noted in his Radar Handbook[9],[10], within classes of targets, the RCS can vary by as much as 20 to 30 dB depending on frequency, aspect angle, and target characteristics. In a separate study[11], a marine X-band radar with a 3.2-cm wavelength and horizontally polarized transmission and reception (H) was used to develop a verified and calibrated RCS dataset of 18 ships passing through the English Channel. These ships ranged from small fishing boats to large tankers, with gross tonnage ranging from 5–45 000 tons. Other RCS measurement ranges, such as those reported by the International Association of Marine Aids to Navigation and Lighthouse Authorities (IALA) in their 2022 report[12], confirm these measurements. Using Skolnik’s formula, estimates of the median RCS for var- ious ships showed errors ranging from 7–17 dBsm when compared to measured values, consistent with Skolnik’s observation that RCS errors can reach up to 30 dBsm[3]. Furtherresearchanalyzingtherelationshipbetweenship displacement and ship length using automatic identification system (AIS) data has shown that it is possible to replace the displacement parameterDin Skolnik’s formula with ship length[13]. This relationship has been leveraged in recent works to estimate the RCS based on ship dimensions[3]. Skolnik’s model has been refined over time by various researchers[14],[15], and one notable variation, introduced by Vachon et al.[1],[14],[15],[16],[17],[18],[19]and more recently by AstroCom Associates Inc[20], incorpo- rates radar incidence angle. This empirical formula takes the form RCS =ax p truthLength (b+cθ)(4) where 1)x truthLength is the ship’s true length in metres; 2)θis the radar incidence angle; 3) the parametersa,b,c,andpare inferred through fitting with the empirical data acquired at a certain center frequency, such as C-band of RADARSAT-2. In curve-fitting literature, Skolnik’s equation and its variants are referred to as power series models[21], while in machine learning, these are typically modeled using multiple linear regression[22]. In this context, “multiple” refers to the number of predictor variables, such asf c and Din Skolnik’s model [(3)], orx truthLength andθin the more recent variants [(4)]. “Linear” refers to the linear nature of the regression after appropriate transformations are ap- plied to the data. However, RCS modeling using power series and multiple linear regression suffers from several key issues, including nonconstant variance, autocorrelation, multicollinearity, and overfitting[22]. These problems can lead to unreliable predictions and inaccurate confidence intervals. From a broader machine learning perspective, multi- ple regression models are a limited special case of more advanced techniques[23],[24],[25]. These modern tech- niques not only address the abovementioned challenges more effectively but also they are not restricted to akth percentile and provide better RCS models with tighter error bounds. Therefore, the objective of this article is to utilize recent advancements in deep learning to develop a reliable and holistic RCS model. The main innovation in this article is the introduction of a novel deep sigma point process (DSPP) model[26]. This Bayesian model not only predicts RCS values but also learns the uncertainty estimates (error bounds), enhancing its practical utility. By quantifying predictive uncertainty, the DSPP model offers a more robust approach to RCS modeling, addressing the limitations inherent in traditional regression methods. The main contributions of this work are as follows. 1)Comprehensive Feature Utilization:Unlike previous works that primarily focused on a limited set of fea- tures, namely, radar center frequency, target length, and radar incidence angle, this work promotes the incorporation of a wider range of relevant features. Thisincludesnotonlyradarparametersbutalsoenvi- ronmental variables and associated AIS parameters, among others. The inclusion of these features allows the model to characterize both standard and extended operating conditions more comprehensively. 2)Framework Development for Model Training and Evaluation:A robust framework was developed for training, validation, and testing of deep learning models tailored to empirical RCS modeling of tar- gets in SAR imagery. This framework streamlines model development and benchmarking across di- verse datasets and operational scenarios. 3)Novel Bayesian DSPP Model:This work introduces an innovative Bayesian DSPP model with the dual capability of RCS prediction and error bound esti- mation, enhancing prediction reliability. The model also adheres to explainable AI (XAI) principles[27], AUTHOR PREPRINT - arXiv submission draft 2 providing insight into feature relevance by generat- ing a ranked importance of each feature. This inter- pretability is integral to validating model predictions in practical applications. 4)Adaptability to Multisensor Data:The model is de- signed to handle data from a wide range of radar sensors, whether operating at the same or different center frequencies, and featuring varying configura- tions. By incorporating center frequency, radar pa- rameters, environmental conditions, and other criti- cal features as inputs, the model achieves robust gen- eralization across diverse sensor platforms, making it highly versatile for various operational scenarios. The rest of this article is organized as follows. In SectionII, we formalize the problem definition and provide an overview of regression modeling for RCS, including our novel Bayesian approach. SectionIIIdescribes the AIS- verified RADARSAT-2 dataset used in this study, detailing both ship detection records and associated radar parameters, providing insights into the unique features of the dataset that enable rigorous model evaluation. SectionIVoutlines the baseline models, with multiple linear regression vari- ants used for comparative analysis. Our primary contribu- tion, the novel DSPP model, is introduced in SectionV, where we elaborate on its architecture, kernel selection, and uncertainty quantification capabilities. SectionVIdis- cusses the enhanced hyperparameter optimization frame- work, incorporating tree-structured Parzen estimator (TPE) and asynchronous successive halving algorithm (ASHA), which significantly improves model training efficiency. In SectionVII, we present a comparative evaluation of the baseline models and the DSPP, supported by visualizations and performance metrics. Finally, SectionVIIIconcludes this article. I. PROBLEM DEFINITION A. Regression Modeling of RCS In this article, we address the problem of predicting RCS values using regression modeling, where the relationship between a set of input features (i.e., predictor variables) X and the RCS values (i.e., target variable)yis learned. This problem can be formalized as y=f(X)(5) where 1) Xis the feature matrix of sizeN×D; 2) yrepresents the RCS in dBsm; 3)Nis the number of samples; 4)Drepresents the number of features. Each feature vector x i ∈Xfori=1,...,Dis associ- ated with either AIS parameters, radar meta-parameters, or other radar-related variables. In this study, we selected 25 features, x 1 tox 25 , which are used to predict the RCS of ships from SAR images. The RCS yis computed from spaceborne RADARSAT- 2 SAR images, specifically from AIS-verified SAR chips. The RCS for a given ship is calculated from the SAR chip using the following formula[28]: RCS = ∑ א 2 A G 2 ref (6) where 1) ∑ א 2 is the sum of the squared Digital Numbers in the SAR chip; 2) Ais the pixel area in square metres; 3)G ref is the reference gain value used for calibration. Equation(6)estimates RCS of an extended target by summing the calibrated backscatter power (i.e., squared pixel amplitudes) over all pixels in the detected ship region. This provides an integrated RCS estimate, which represents the aggregate radar reflectivity of the target and is widely used in operational SAR applications[19],[29]. This yields a measured RCS, which serves as the ground truth label for training and evaluating the regression model. The model aims to predict this observed RCS from a set of interpretable physical and environmental features, enabling RCS esti- mation under varying conditions without direct reliance on SAR imagery. The functionfis the regression model that learns the relationship between the features Xand the RCS values y. Various forms of the regression functionfare used in this study, including machine learning and deep learning approaches. As baselines, we consider four variants of multiple linear regression as follows. 1)Linear regression[30]. 2)Ridge regression[31]. 3)Lasso regression[32]. 4)ElasticNet[33]. In addition to these baseline models, we propose a novel Bayesian model based on DSPP[26], which not only predicts the RCS but also learns the error bounds for the predictions, allowing for uncertainty quantification. B. Splitting and Normalizing the Dataset To ensure that the trained model generalizes well, the dataset is divided into three splits: 1) training (70%); 2) validation (15%); and 3) testing (15%). The training split is used to learnf, while the validation split is used to tune hyperparameters and prevent overfitting. Once the model is trained, it is evaluated on the test split to assess its real-world performance. The dataset is randomly shuffled before being split. Random shuffling is performed to ensure that the data are distributed evenly across the splits, reducing the risk of bias and ensuring that the training, validation, and testing sets contain diverse samples[34]. In this study, both the input features Xand the target variable yare normalized to ensure that they are on the same scale. More precisely, each feature vector x i is normalized to the same range∈ [−1,1]as follows: x norm i =−1+2 x i −min(x i ) max(x i )−min(x i ) (7) AUTHOR PREPRINT - arXiv submission draft 3 where 1) min (x i )is the minimum value for the feature samples of the training split; 2) max (x i )is the maximum value for the feature sam- ples of the training split. This scaling helps improve convergence and perfor- mance in machine learning models by ensuring that features are comparable in magnitude, preventing any one feature from disproportionately influencing the model. Similarly, the target variable y(i.e., RCS in dBsm) is normalized[35] y norm = y −μ y σ y (8) where 1)μ y is the mean for the target variable across all samples in the training split of the dataset; 2)σ y is the standard deviation for the target variable across all samples in the training split of the dataset. C. Performance Metrics The performance of models is evaluated using two key metrics as follows. 1)Root Mean Squared Error (RMSE): This measures the average magnitude of the error between the pre- dicted (i.e., ˆy i ) and true (i.e.,y i ) RCS values[36] RMSE = √ √ √ √ 1 N N ∑ i=1 (y i −ˆy i ) 2 .(9) 1)Range of RMSE:RMSE≥ 0; 2) A value of 0 indicates perfect predictions. Higher values indicate larger deviations between predicted and true values; 3) The unit of RMSE is the same asy i ; 4)Interpretation:Lower RMSE values are better, indi- cating more accurate predictions. 2)R 2 Score (Coefficient of Determination):This mea- sures the proportion of variance in ythat is explained by X[37] R 2 =1− ∑ N i =1 (y i −ˆy i ) 2 ∑ N i =1 (y i − ̄y) 2 (10) where ̄yis the mean of the true valuesy i 1)Range of R 2 :−∞<R 2 ≤1; 2) AnR 2 =1indicates perfect prediction, whileR 2 =0 means no variance is explained; 3) Negative values indicate that the model performs worse than predicting the mean; 4)Interpretation:HigherR 2 values are better, indicat- ing that more variance in RCS is captured by the model. To calculate RMSE andR 2 for unnormalized data, the following conversions are used. Fig. 1. Geographical distribution of the 208,191 AIS-verified RADARSAT-2 ships used in this study. 1)Converting RMSE to Unnormalized Scale:Once the RMSE is calculated for the normalized data (RMSE norm ), the corresponding RMSE for unnor- malized data (RMSE unnorm ) can be obtained by mul- tiplying it by the standard deviation (σ y ) of the orig- inal target variabley RMSE unnorm =RMSE norm ·σ y .(11) 2)Converting R 2 to Unnormalized Scale:TheR 2 value remains the same for both normalized and unnor- malized data because it measures the proportion of variance explained by the model. Therefore R 2 unnorm =R 2 norm .(12) These performance metrics provide insight into how well the model predicts the RCS values, and the unnor- malized RMSE gives an accurate understanding of the prediction error in real-world units (dBsm for RCS). I. AIS-VERIFIED RADARSAT-2 DATASET This study is built on and evaluated using focused, detected SAR intensity products that have undergone stationary-world matched filtering—the same magnitude- onlyimageryroutinelyusedinoperationalmaritimesurveil- lance. These data feed processors, such as OceanSuite, the DND framework for ship detection[19],andSUMO, the open-source SAR-based ship detection module main- tained by the European Commission’s Joint Research Cen- tre (JRC)[29]. The spaceborne SAR dataset used in this study was acquired between January 2013 to May 2019 and it com- prises 208 191 AIS-verified ship detections, offering a com- prehensive overview of maritime vessel activity captured via RADARSAT-2. The geographical distribution of these detections is represented in Fig.1, which highlights areas of varying vessel densities across the globe. High-density clusters are primarily found along the east coast of North America, extending from the northeastern United States to southeastern Canada, and along major Pacific Ocean shipping routes near the west coast of Central America. AUTHOR PREPRINT - arXiv submission draft 4 Fig. 2. Count of unique MMSIs for the AIS-verified RADARSAT-2 ships used in this study. Each tick on the x-axis represents a unique MMSI (anonymized index), and they-axis shows the number of AIS-verified RADARSAT-2 detections recorded for that vessel. A notable cluster is also observed in Southeast Asia, re- flecting important regional maritime routes. Each SAR chip in the dataset is complemented by meta files that provide a wealth of information offering a com- prehensive perspective on each vessel’s detection scenario, including the following. 1) The RCS in dBsm, calculated from the SAR chip using the sum of the squared digital numbers and scaled by the area of the radar chip. 2) Radar parameters, such as polarization, beam mode, and incidence angle. 3) AIS data like vessel identification and dimensions. 4) Environmental data, such as wind speed and direc- tion, at the time of detection. A more detailed description of the dataset can be found in[38]. A histogram of unique maritime mobile service identity (MMSI) counts is shown in Fig.2, highlighting the distribution of vessel encounters in the dataset. The distribu- tion demonstrates a long-tail pattern, where a small number of vessels are frequently detected, while the majority are detected less frequently. This is crucial for understanding the variability in the RCS values across different ships. In this study, the feature matrix Xconsists of 25 features that provide detailed information about the SAR detection, radar parameters, and AIS data. The target variable yis the RCS, which represents the radar signal’s reflected power from a ship in dBsm. The feature matrix Xhas dimensions N×D, whereN =208191andD=25. It should be noted that these features are not exclusive and in future works additional features can be incorporated, for example by including the center frequency when considering datasets with different carrier frequencies. 2-D histogram plots for all the dataset samples considered in this study, pertaining to RCS values plotted against their corresponding features x 1 tox 25 are provided in Figs.3and4. These features are defined as follows. 1)ChipPolarization( x 1 ): The transmit/receive polar- ization of the SAR chip∈ H,HV,VH,V.H and V, respectively, represent horizontal and ver- tical polarization. There are 191 628 chips in the dataset that are H-polarized, while 16 563 chips are V-polarized. There are neither HV nor VH- polarized chips in the dataset. In the feature vector x 1 , the polarization labels H and V, respectively, are encoded to 1 and 2. 2)SatelliteTrackDegrees( x 2 ): The angle in degrees (from North) at which the satellite tracks along the Earth’s surface. 3)SatelliteLookDirectionDegrees( x 3 ): The direction in degrees (from North) the sensor is pointing (i.e., for default right looking configuration this is equal to SatelliteTrack+90 degrees). 4)BeamMode( x 4 ): The beam mode in the dataset∈ W, 1 FW, 2 EH, 3 F, 4 UW, 5 SCN, 6 SCW, 7 DVWF, 8 andOSVN 9 [39].Inthefeaturevectorx 4 ,thebeam mode labels are encoded from 1 to 9, respectively, and the corresponding number of chips per each beam mode∈237, 1769, 10, 34, 1, 5658, 797, 63717, 135968 5)NumberOfBeams( x 5 ): The product number of beams∈1, 2, 3, 4, 7, 8, respectively, with a corresponding number of chip samples∈2079, 5457, 201, 797, 63717, 135940. 6)ProductNumberOfPolarizations( x 6 ): The product number of polarizations∈1, 2, with a correspond- ing number of chip samples∈144380, 63811, respectively. 7)ProductNumberOfRangeLooks( x 7 ): The product number of independent range looks∈1, 2, 4, 5, with a corresponding number of chip samples∈ 2079, 5658, 63717, 136737, respectively. 8)ProductNumberOfAzimuthLooks( x 8 ): The product number of independent azimuth looks∈1, 2, 4, with a corresponding number of chip samples∈ 201489, 6455, 247, respectively. 9)ImageProjection( x 9 ): Projection of the SAR chip ∈ground-range, slant-range. In this study, all the chips are in the ground-range projection. 10)LineSpacingMetres( x 10 ): The product line spacing in metres∈1.5625, 5, 6.25, 12.5, 20, 25, 50, with a corresponding number of chip samples∈1, 28, 1803, 247, 63717, 141598, 797, respectively. 11)PixelSpacingMetres( x 11 ): The product pixel spac- ing in metres∈1.5625, 5, 6.25, 12.5, 20, 25, 35, 50, with a corresponding number of chip samples 1 Wide 2 Wide Fine 3 Extended High 4 Fine 5 Wide Ultra Fine 6 ScanSAR Narrow 7 ScanSAR Wide 8 Detection of Vessels, Wide swath, Far incidence angle 9 Ocean Surveillance, Very wide swath, Near incidence AUTHOR PREPRINT - arXiv submission draft 5 Fig. 3. 2-D histogram plots of RCS versus featuresx 1 tox 20 . AUTHOR PREPRINT - arXiv submission draft 6 Fig. 4. 2-D histogram plots of RCS versus featuresx 21 tox 25 ,andx 23 versusx 20 . ∈1, 28, 1803, 247, 63717, 5658, 135940, 797, respectively. 12)TargetCenterLatitudeDegrees( x 12 ): The latitude of the target’s center in the SAR image in degrees. 13)TargetCenterLongitudeDegrees( x 13 ): The longi- tude of the target’s center in the SAR image in degrees. 14)TargetIncidenceAngleDegrees( x 14 ): The radar in- cidence angle at target in degrees. 15)TargetGroundRangeResolutionMetres( x 15 ): The ground-range resolution of the SAR chip in metres. 16)TargetAzimuthResolutionMetres( x 16 ): The az- imuth resolution of the SAR chip in metres. 17)TargetEquivalentNumberOfLooks( x 17 ):Theequiv- alent number of looks for the SAR chip. 18)ReferenceNoiseLeveldB( x 18 ): The reference noise level in dB at the target in the SAR chip. 19)AISTargetMMSI( x 19 ): The AIS MMSI identifier of the verified ship in the SAR image. 20)AISTargetLengthMetres( x 20 ): The AIS length in metres for the verified ship in the SAR image. 21)SARRelativeWindDirectionDegrees( x 21 ): The rela- tive Canadian Meteorological Centre (CMC) wind direction in degrees for the verified ship in the SAR image. 22)SARDerivedWindSpeedMPS( x 22 ): The wind speed in m/s derived from the SAR product. 23)AISTargetWidthMetres( x 23 ): The AIS width in metres for the verified ship in the SAR image. 24)AISTargetSpeedMPS( x 24 ): The AIS speed in m/s for the verified ship in the SAR image. 25)AISTargetCourseOverGroundDegrees( x 25 ): The AIS course over ground in degrees for the verified ship in the SAR image. IV. BASELINE MODELS: MULTIPLE LINEAR REGRES- SION VARIANTS To provide a robust comparison against our proposed model, we have employed four baseline multiple linear regression models, namely,linear regression,ridge regres- sion,lasso regression,andElasticNet. Each model applies different regularization techniques to improve the gener- alization of the linear regression model, especially when dealing with high-dimensional or correlated data. A. Linear Regression Linear regression is a fundamental model that estimates the relationship between a dependent variable yand in- dependent variables Xby minimizing the residual sum of squares between observed values and the values predicted by the linear approximation[30]. The mathematical form of linear regression is as follows: ˆ y=X β+(13) where X∈R N×D represents the input features,β∈R D are the model parameters, and∈ R N denotes the error term. The goal of the model is to find ˆ βthat minimizes the following cost function: min β y−Xβ 2 2 .(14) Linear regression assumes that the relationship between the input features and the output variable is linear, and no regularization is applied to the model coefficients. B. Ridge Regression Ridge regression extends linear regression by introduc- ing anL 2 regularization term to the cost function, which penalizes large model coefficients. This helps in preventing AUTHOR PREPRINT - arXiv submission draft 7 overfitting, particularly when multicollinearity exists in the data[31]. The ridge regression optimization problem is formulated as min β y−Xβ 2 2 +λβ 2 2 (15) whereλ≥ 0is the regularization parameter that controls the strength of the penalty. Asλincreases, the magnitude of the model coefficients is shrunk, which can reduce variance but may introduce bias. C. Lasso Regression Lasso (least absolute shrinkage and selection operator) regression appliesL 1 regularization to the model coeffi- cients, encouraging sparsity by driving some coefficients to zero. This makes lasso particularly useful for feature se- lection[32]. The optimization problem for lasso regression is given by min β y−Xβ 2 2 +λβ 1 .(16) Here,β 1 represents theL 1 norm, andλ≥0controls the degree of regularization. Unlike ridge, lasso can produce sparse solutions by forcing some of the coefficients to be exactly zero, effectively selecting a subset of the features. D. ElasticNet ElasticNet combines bothL 1 andL 2 regularization tech- niques by linearly combining the penalties of ridge and lasso. This allows the model to benefit from the sparsity of lasso and the stability of ridge, making it suitable for cases where the number of predictors exceeds the number of observations, or when there are highly correlated fea- tures[33]. The optimization problem for ElasticNet is as follows: min β y−Xβ 2 2 +λ 1 β 1 +λ 2 β 2 2 (17) whereλ 1 andλ 2 are the regularization parameters for the L 1 andL 2 penalties, respectively. ElasticNet is useful when neither ridge nor lasso alone provides an adequate model for the data. V. NOVEL MODEL – DEEP SIGMA POINT PROCESS DSPP is a hierarchical Gaussian process (GP) model designed to capture complex, nonlinear relationships in data while estimating predictive uncertainty. As illustrated in Fig.12, the DSPP architecture comprises multiple layers where input features flow through several GPs, each trans- formingthedatawhilepropagatinguncertainty.Thissection discusses the model’s architecture, including the roles of sigma points, kernels, inducing points, and quadrature rules, and how dimensionality changes throughout the process. A. Summary of Key Variables 1)W: The width of each GP layer, representing the number of GPs in that layer. 2)L: The number of layers in the DSPP model, where each layer processes the outputs from the previous layer. 3)M: The number of inducing points used to approxi- mate the GP posterior for large datasets. 4)S: The number of sigma points (quadrature points) used to approximate the integral over latent func- tions. 5)ξ (s) w : Sigma points that approximate the latent vari- able distribution, transforming the integral into a finite mixture. 6)Kernel Function k (x i ,x j ):Defines the covariance between input points and determines how outputs are related in each GP. The kernel is applied separately to each GP node at each layer, with various kernels selectable based on data characteristics. B. Why DSPP? DSPP offers a robust framework for modeling nonlinear relationships while estimating mean predictions and their associated uncertainties. This capability is crucial in fields that require precise uncertainty assessments. As a Bayesian model, DSPP provides distributions over predictions rather than just point estimates, facilitating the quantification of uncertainty[26]. Key advantages of DSPP include the fol- lowing. 1)Mean and Variance Estimation:DSPP predicts both the mean and variance of outputs, providing valuable error bounds for risk assessment. 2)Uncertainty Propagation:The model captures how uncertainty evolves through multiple GP layers, al- lowing for the assessment of prediction reliability. DSPP excels at propagating uncertainty across layers of GPs, enabling it to: 1)IdentifyRisk:confidenceintervalsprovidedbyDSPP clarify prediction reliability, enhancing decision- making processes; 2)Model Uncertainty in Complex Systems:DSPP ef- fectively handles noisy, sparse data from intricate systems, yielding predictions with error bounds that reflect inherent variability. C. Why Use Sigma Points? the Link to Kalman Filtering and Other Fields Sigma points in DSPP draw inspiration from their application in unscented Kalman filtering (UKF) and other domains where uncertainty propagation is essen- tial. In Kalman filtering, particularly in control systems and robotics, the challenge lies in propagating uncertainty through nonlinear systems. UKF utilizes sigma points to approximate the distribution of state variables, capturing the mean and covariance without relying on stochastic sampling, thus enabling efficient updates of estimates[40]. In DSPP, sigma points serve similar functions as follows. AUTHOR PREPRINT - arXiv submission draft 8 1)Deterministic Quadrature Points:They approximate integrals over latent functions in GPs, effectively capturing both mean and variability. 2)Efficient Uncertainty Propagation:DSPP leverages sigma points to propagate uncertainty across mul- tiple GP layers, avoiding the computational costs associated with Monte Carlo sampling. Sigma points also find applications beyond Kalman filtering as follows. 1)Control Systems:They approximate uncertainty in evolving system states[40]. 2)Sensor Fusion:Sigma points facilitate data fu- sion from multiple sensors by propagating uncer- tainty[41]. 3)Robotics:They are used to track objects and navigate autonomous systems, where noise and uncertainty must be carefully managed[42],[43]. D. Input Layer and Dimensionality The input to the DSPP model consists of feature vectors x=[x 1 ,x 2 ,...,x D ], whereDdenotes the number of fea- tures per sample. For a dataset withNsamples, the input data is represented as a matrix X∈R N×D . As the input features progress through multiple layers of GPs, they are transformed into latent representations, with dimensionality evolving throughout the hierarchy. In the first GP layer, the input features are processed byWindependent GPs, where Wrepresents the width of the layer. Each GP computes a latent variable with a multivariate Gaussian distributionN for each input sample x i , characterized by a meanμ g w (x i ) and varianceσ 2 g w (x i ) g w (x i )∼N(μ g w (x i ),σ 2 g w (x i )),w=1,2,...,W.(18) E. Role of the Kernel The kernel functionk(x i ,x j )defines the covariance structure between input points, playing a crucial role in determining correlations among different GPs based on input data. Specifically, the kernel function: 1)Defines Covariance:it computes the covariance ma- trix K, with each entryK ij =k(x i ,x j )representing the covariance between two inputs; 2)Governs Uncertainty Propagation:it shapes how uncertainty is propagated across layers by defining correlations between latent variables. Each GP utilizes its own kernel function, with various types available (e.g., radial basis function, Matérn, periodic, etc.) depending on the data characteristics[3]. In addition, additive and multiplicative composites of these kernels can be employed to capture more complex patterns in the data. F. Inducing Points in DSPP Computing and inverting large covariance matrices in GPs can be computationally expensive for large datasets (complexityO (N 3 )). To enhance scalability, DSPP intro- duces inducing pointsM. These points approximate the GP’s latent functions, allowing the model to reduce com- putational complexity toO (NM 2 )while still capturing essential GP structures. Inducing points approximate the covariance between latent outputs and input data points at each GP layer. The kernel function computes the covariance between the inducing points and input features, facilitating efficient uncertainty propagation Cov [f(z),f(x)]=k(z,x)(19) where z∈R M×D are the inducing points, andxare the input data points. G. Sigma Points Sigma points (denoted asξ (s) w ) approximate intractable integrals over latent variables in DSPP. Serving as determin- istic quadrature points, they enable efficient approximation of nonlinear transformations of latent functions without the need for stochastic sampling, thereby propagating both the mean and uncertainty through the GP layers. For each GP, the latent outputg w (x i )is approximated using a weighted combination of the mean and variance g w (x i )=μ g w (x i )+ξ (s) w σ g w (x i ),s=1,2,...,S(20) where 1)Sis the total number of sigma points; 2)ξ (s) w are the learnable sigma points capturing variabil- ity and uncertainty. Sigma points effectively approximate the integrals needed to propagate uncertainty across GP layers. H. Quadrature Rules in DSPP Quadrature rules are techniques for approximating in- tegrals when exact computation is intractable. In DSPP, quadrature rules are implemented via sigma points, allow- ing efficient approximation of integrals over latent variable distributions. Instead of calculating exact integrals over latent functions, DSPP employs sigma points as discrete representations of latent variable distributions, forming a weighted sum that approximates the integral[26] ∫ p (f)df≈ S ∑ s=1 w s p(ξ s )(21) where w s are the weights associated with each sigma point. This transforms the integral into a manageable sum, fa- cilitating efficient computation while capturing necessary uncertainty. Utilizing quadrature rules through sigma points enables DSPP to maintain high prediction accuracy while avoiding the computational burden of traditional sampling methods, making it suitable for high-dimensional inputs and large datasets. AUTHOR PREPRINT - arXiv submission draft 9 I. Hidden GP Layers DSPP consists ofLlayers of GPs, each withWGPs (nodes). In each layer, inducing points approximate the latent function, while sigma points propagate uncertainty. For instance, in Fig.12, the DSPP model features param- etersL =2,W=4,andS=10, serving as a reference for arbitrary configurations. For each GP at layerl, the latent output is computed based on the kernel function f (l) w (g (l−1) w (x))∼N ( μ f (l) w (g (l−1) w (x)),σ 2 f (l) w (g (l−1) w (x)) ) , w=1,2,...,W. (22) The kernel function determines the covariance between inputs and inducing points at each layer, while inducing points help reduce computational complexity and sigma points approximate integrals over latent variables. J. Predictive Distribution The predictive distribution for the targety i given the input x i involves both multiplication (due to the hierarchical GP structure) and integration (to account for uncertainty over the latent variables) p (y i |x i )= ∫ L ∏ l=1 p(y i |f (L) (x i )) L ∏ l=1 p(f (l) |f (l−1) ,x i ) df (1) ...df (L) .(23) This integral is approximated using sigma points p DSPP (y i |x i )≈ S ∑ s 1 =1 · S ∑ s W =1 ω s 1 ,...,s W N(y i |μ (L) f (g (L−1) w (x i )), σ (L) f (g (L−1) w (x i ))+σ 2 obs ) (24) where 1)μ (L) f andσ (L) f are the mean and variance from the final GP layer; 2)ω s 1 ,...,s W are the mixture weights for each combina- tion of sigma points; 3)σ obs is the variance of the Normal likekihoodp(y|.). The kernel function computes the covariance between inputs and inducing points, while sigma points propagate uncertainty through the GPs. K. Loss Function The loss function for DSPP consists of a log-likelihood term that evaluates the fit of the model to the data and a regularization term that penalizes Kullback–Leibler (KL) divergence from the GP prior[26] L DSPP = N ∑ i=1 logp DSPP (y i |x i )−β reg ∑ KL(25) where: 1)Log-Likelihood:The log of the predictive distribu- tion (approximated using sigma points and kernels) is computed for each data point. 2)Regularization (KL Divergence):A regularization termβ reg that penalizes the KL divergence between the posterior distribution over inducing points and their GP prior, ensuring smoothness and preventing overfitting. L. Training the DSPP Model Training the DSPP model typically involves using op- timization algorithms, such as stochastic gradient descent (SGD) or Adam to minimize the loss function over a spec- ified number of epochs. The choice of epochs and learning rate is critical for convergence and model performance. 1)Epochs:The number of epochs required for training depends on the complexity of the data and model architecture. Monitoring the validation loss is essen- tial to avoid overfitting and ensure that the model generalizes well to unseen data. 2)Optimization Algorithm:Adam, known for its adap- tive learning rates, is often preferred due to its ef- fectiveness in handling sparse gradients and faster convergence, while SGD can be beneficial for larger datasets requiring lower memory overhead. RM- SProp is closely related to Adam[44]. By incorporating inducing points, sigma points, quadra- ture rules, and kernel functions, DSPP efficiently approxi- mates the propagation of both mean and uncertainty across layers, enablingscalable andaccurate modelingof complex, nonlinear data. The kernels define the covariance structure at each GP, inducing points reduce computational costs, and sigma points transform intractable integrals into manage- able sums, allowing DSPP to handle large datasets and high- dimensional inputs effectively. This Bayesian framework enhancespredictivecapabilitiesandempowerspractitioners with tools to assess the reliability and risk associated with their predictions. VI. ENHANCED HYPERPARAMETER OPTIMIZATATION USING TPE AND ASHA Hyperparameter optimization plays a critical role in machine learning, as it significantly impacts model per- formance. Traditional methods, such as grid and random searches are often inefficient, requiring substantial compu- tational resources and time. Recent advances in probabilis- tic and heuristic approaches have demonstrated significant potential for improving scalability and performance in hy- perparameter optimization tasks[45]. The TPE, introduced by Bergstra et al.[46], is a Bayesian optimization algorithm that models the likelihood of hyperparameter configura- tions yielding lower losses. Unlike traditional methods, which search the hyperparameter space uniformly, TPE dynamically adjusts its search based on historical perfor- mance, focusing on regions of the search space that are more likely to result in better model performance. By AUTHOR PREPRINT - arXiv submission draft 10 Algorithm 1:Enhanced Hyperparameter Optimiza- tion using TPE and ASHA. honing in on promising areas, TPE effectively balances the exploration–exploitation tradeoff—a fundamental chal- lenge in optimization[47]. The ASHA, a resource-efficient variant of Hyperband introduced by Li et al.[48], enhances hyperparameter tuning by allowing early termination of poorly performing trials. This method reallocates resources to more promising configurations without waiting for all trials to complete, significantly improving computational efficiency. ASHA is particularly well-suited for large-scale hyperparameter optimization tasks because it facilitates efficient exploration of the hyperparameter space while ensuring that promising configurations are exploited more effectively over time[48],[49]. This implementation lever- ages the combined strengths of TPE and ASHA. These well- established methodologies are integrated to manage com- putational resources more effectively and accelerate model selection. Building on the foundational work of Bergstra et al.[46]andLietal.[48], we use the TPE’s informed sam- pling method to guide the exploration of promising hyper- parameter configurations, while ASHA ensures efficient re- sourceallocationthroughitsearlystoppingmechanism.The interplay between TPE’s probabilistic model and ASHA’s bandit-based strategy significantly reduces computational overheadandshortensthetime-to-convergencecomparedto traditional methods. The steps of this approach are outlined in Algorithm1, where the hyperparameter search space is defined along with budget constraints. The experiments were conducted on NVIDIA DGX A100 system[50].Key parameters include the following. 1)CPUsandGPUs: Resources allocated to each trial. In this setup, each trial is allocated 16 CPUs and 0.5 GPUs, ensuring that parallel computations are distributed efficiently. 2)n_startup_trials: Set to 64, this parameter defines the number of initial trials that are evaluated before TPE starts using its model to make informed sugges- tions. This allows for an exploration phase to gather sufficient data for understanding the hyperparameter landscape. 3)seed: Set to 42, ensuring reproducibility by initial- izing random number generators consistently across trials. 4)n_max_trials: Limited to 512, serving as a budget constraint, ensuring that the optimization process does not exceed a predefined number of hyperpa- rameter evaluations. 5)num_epochs: Defines the maximum number of it- erations or epochs that a single trial can run. Set to 4096, ensuring that each model is trained sufficiently without consuming excessive resources. 6)grace_period:Setto16,definingtheminimumnum- ber of iterations each trial is allowed to run before ASHA evaluates its performance for potential early stopping. 7)reduction_factor: Set to 2, this parameter defines the rate at which the number of trials is halved in each successive round of evaluation, concentrating resources on the most promising configurations. This hybrid optimization strategy has been applied across various benchmark datasets, consistently demon- strating improvements in model accuracy and reductions in computational costs[45]. These results underscore the benefits of combining Bayesian optimization techniques, like TPE, with bandit-based strategies, such as ASHA, to provide a robust solution for hyperparameter optimization. By balancing exploration and exploitation, this approach ensures that the tuning process is efficient and resource- effective. AUTHOR PREPRINT - arXiv submission draft 11 Fig. 5. Parallel coordinates plot for the baseline models, with the top-performing model in terms of validation RMSE is highlighted in green. TABLE I Comparison of Model Performance VII. RESULTS AND DISCUSSION A. Training and Hyperparameter Optimization of the Baseline Models As discussed in SectionIV, the baseline models were chosen for their ability to generalize linear relationships, with each model incorporating different regularization tech- niques to handle overfitting and multicollinearity, partic- ularly in high-dimensional or correlated data. The imple- mentation of these baseline models was carried out using scikit-learn[51]andPyTorch[52], with hyperpa- rameter optimization performed via Optuna[53],awidely used framework for hyperparameter search and optimiza- tion. The search strategy for the optimization follows the method outlined in Algorithm1, where TPE was used to efficiently explore the hyperparameter space, and ASHA was employed to terminate underperforming trials early, optimizing computational resources. A total of 512 trials were conducted, with performance evaluated based on the training and validation RMSE andR 2 scores. The best- performing model for each variant was selected based on the lowest validation RMSE. The training data constituted 70% of the total dataset, while 15% was set aside for validation. The remaining 15% of the data was reserved for testing. The search space for these baseline models was defined as follows. 1)Model Type (Model):Linear regression, ridge, lasso, and ElasticNet. 2)Alpha Parameter (α):Log-uniformly sampled be- tween 0.01 and 100.0 (used for ridge, lasso, and ElasticNet). 3)Fit Intercept (Intercept):True, False. 4)L1 Ratio (Ratio):Uniformly sampled between 0.0 and 1.0 (only applicable to ElasticNet). For these models, no epoch-based training was needed, as they directly minimize a regularized objective. A parallel coordinates plot is provided in Fig.5to visualize the search space and illustrate how different hyperparameter configu- rations influenced model performance. The plot depicts that the best-performing baseline model is linear regression with validation val_RMSE norm =0.6408and val_R 2 =0.5905. The corresponding val_RMSE unnorm =3.2298[dBsm]. The performance metrics for the training, validation, and test data splits, respectively, are summarized in TableI. Among these four baseline variants, the top-performing linear re- gression (highlighted by yellow in the table) is closely followed by ridge regression. If a grid search was used instead of TPE-ASHA, the search space size would have been much larger due to the need to explore all possible combinations of hyperparameters. The search space would have included the following. 1)Model Type:Four choices (linear regression, ridge, lasso, and ElasticNet). 2)Fit Intercept:Two choices (True or False). 3)Alpha:100 possible values (log-uniform). 4)L1 Ratio:100 choices (only for ElasticNet). The total grid search space for each model would be as follows. 1) ForLinear Regression:Two options (no alpha or L1 ratio involved). 2) ForRidge and Lasso Regression:200 options each (alpha and fit intercept). AUTHOR PREPRINT - arXiv submission draft 12 Fig. 6. Scatter plot of predicted versus true RCS for the top-performing baseline (i.e., linear regression) model. Fig. 7. 2-D histogram plot of true RCS versus residuals for the top-performing baseline (i.e., linear regression) model. Residuals have an IQR =3.6 dBsm and a MAD=1.8 dBsm. 3) ForElasticNet:20 000 options (due to both alpha and L1 ratio). Fig.6provides a scatter plot for predicted versus true RCS [dBsm] pertaining to the test split of the dataset processed by the top-performing linear regression model. Further, Fig.7provides a 2-D histogram plot for the true RCS versus residual, along with 1-D histograms for the residual and true RCS. The interquartile range (IQR) and median absolute deviation (MAD) for the residual are 3.6 and 1.8 dBsm, respectively. These baseline models serve as useful benchmarks, illustrating the strengths of regularized linear approaches while allowing comparison with the more complex, uncertainty-aware DSPP model. B. Training and Hyperparameter Optimization of the DSPP Model For training the DSPP model introduced in SectionV, we utilized PyTorch[52]for model implementation and GPyTorch[54]to efficiently handle the GP layers within the DSPP structure. Hyperparameter optimization was car- ried out using Optuna[53], following the method outlined in Algorithm1. A total of 512 trials were conducted, with performance evaluated based on the training and valida- tion RMSE andR 2 scores. The best-performing model was selected based on the lowest validation loss. In DSPP models, validation loss is preferred over validation RMSE when selecting the best model because it captures both the accuracy of the predictions and the model’s confidence in those predictions. By using validation loss, we ensure a consistent evaluation with the model’s training objective, helping to identify a model that is not only accurate but also well-calibrated in its uncertainty estimates. The training, validation, and testing was based on the same 70%, 15%, and 15% data splits, respectively. The search space for the DSPP model was defined as follows. 1)Batch Size (B):16, 32, 64, 128, 256, 512, 1024, 2048. 2)Hidden Dimensions (Per Layer) (W ):2, 3, 4, 5, 6, 7, 10. 3)Number of Layers (L):2, 3, 4, 5, 6, 7. 4)Number of Inducing Points (M):300, 400, 500. 5)Number of Quadrature Sites (Sigma Points, S):6, 8, 10, 12, 15. 6)Beta Parameter (β):0.01, 0.05, 0.1, 0.2. 7)Initial Learning Rate (α 0 ):Uniformly sampled be- tween 0.001 and 0.1. 8)Optimizer:Adam, SGD, and RMSprop. 9)Base Kernel:RBFKernel, MaternKernel, Period- icKernel, LinearKernel, PiecewisePolynomialKer- nel, and RQKernel. 10)Nu (Smoothness parameter for MaternKernel) (ν): 0.5, 1.5, 2.5. 11)Smoothness Parameter (For PiecewisePolynomi- alKernel) (q):0, 1, 2, 3. 12)Composite Kernel:True, False. 13)Composite Kernel Type:Additive, Multiplicative. 14)Second Kernel for Composite Type:Same choices as for the base kernel. The total search space size for a grid search over the provided hyperparameters can be calculated as follows: Total Search Space =8×7×6×3×5×4×1000 ×3×6×3×4×2×2×6. (26) This would be approximately 1.05×10 11 possible combi- nations. This vast search space highlights the impracticality of performing grid search for this problem, reinforcing the importance of more efficient search techniques like TPE- ASHA. In addition, the training employed cosine annealing with warm restarts to adjust the learning rate dynamically over time, which has been shown to improve model con- vergence and performance in neural networks[55].During the optimization process, several performance metrics were monitored, including RMSE,R 2 , and loss scores for both training and validation phases. As depicted in Fig.8, these metrics were visualized using a parallel coordinates plot AUTHOR PREPRINT - arXiv submission draft 13 Fig. 8. Parallel coordinates plot for the 512 DSPP trials, with the top-performing model in terms of validation loss is highlighted in green. (a) Plot 1 of 3. (b) Plot 2 of 3. (c) Plot 3 of 3. acrossallthe512trials,providinginsightsintohowdifferent hyperparameters impacted the model’s performance. The best-performing trial has the lowest validation loss val_loss =0.2226at epoch no. 299. TableIlists the RMSE andR 2 performance metrics for this best-performing DSPP model in terms of the training, validation, and test data splits, respectively. Figs.9–11, respectively, show the loss, RMSE, andR 2 plots versus epochs for the best-performing trial. The model architecture for the best-performing DSPP trial is depicted in Fig.12, with the configuration parameters provided in TableII.Fig.13provides a scatter plot for predicted versus true RCS [dBsm] pertaining to the test split ofthedatasetprocessedbythetop-performingDSPPmodel, without and with the error bounds estimated by the model, respectively. Further, Fig.14provides a 2-D histogram plot for the true RCS versus residual, along with 1-D histograms for the residual and true RCS. For the DSPP’s residual, IQR =2dBsmandMAD=1 dBsm. C. Performance Comparison and Observations Performance of the top DSPP model against the corre- sponding baseline linear regression model for the same test data split is summarized as follows. 1)Error Metrics: 1) The DSPP model demonstrates superior perfor- mance compared to the linear regression baseline AUTHOR PREPRINT - arXiv submission draft 14 Fig. 9. Loss plots for the top-performing DSPP trial. Fig. 10. RMSE plots for the top-performing DSPP trial. Fig. 11. R 2 plots for the top-performing DSPP trial. (TableI), with a test RMSE of 2.57 dBsm, sig- nificantly lower than the baseline linear regression model’s RMSE of 3.24 dBsm. This improvement indicates a higher predictive accuracy in the DSPP model. 2) The IQR and MAD values also highlight the DSPP’s robustnessinerrorprediction,withanIQRof2dBsm andMADof1dBsm(Fig.7), outperforming the baseline’s IQR of 3.6 dBsm and MAD of 1.8 dBsm (Fig.14). These metrics illustrate that the DSPP model yields more consistent predictions, with fewer large deviations from the true RCS values. 2)Correlation and Model Fit: 1) The scatter plots (Figs.6and13) illustrate the rela- tionship between predicted and true RCS values for both models: 1) in thebaseline linear regression plot(Fig.6), a broaderspreadaroundtheidentitylinesuggestsmore significant errors, particularly at higher RCS values. This observation aligns with the baseline’s higher RMSE and lowerR 2 ; TABLE I Model Configuration for the Top-Performing DSPP Model 2) the DSPPplot without error bounds[Fig.13(a)] shows a tighter clustering around the identity line, reflecting the DSPP model’s better fit and predictive accuracy; 3) the DSPPplot with error bounds[Fig.13(b)] adds valuable interpretability by visualizing the model’s confidence in its predictions, with smaller residuals closer to the identity line and error bounds that highlight prediction uncertainty. This feature is an advantage of the DSPP, as it can predict RCS values along with confidence intervals, aiding in its practi- cal utility. 3)Predictive Power and Explainability: 1) The DSPP model not only achieves lower RMSE but also demonstrates higherR 2 values(0.74onthe test set) compared to the baseline’sR 2 of 0.59. This improvedR 2 indicates that the DSPP model explains a larger portion of the variance in RCS, providing stronger predictive insights. 2) The DSPP’s error bounds further enhance explain- ability, offering an estimate of the model’s uncer- tainty, which is essential for applications requiring reliable RCS predictions under varying conditions. 4)Implications for Model Selection: 1) The DSPP model’s ability to predict error bounds provides an edge over the linear regression baseline, making it a more suitable choice for applications in radar and sensor modeling where both accuracy and reliability are essential. 2) By incorporating both improved predictive perfor- mance and uncertainty estimation, the DSPP model addresses limitations observed in simpler models, such as the inability of the linear regression model to handle complex, nonlinear relationships in RCS data effectively. AUTHOR PREPRINT - arXiv submission draft 15 Fig. 12. Architecture of the top-performing novel DSPP model for RCS modeling. Fig. 13. Scatter plot of predicted versus true RCS for the top-performing DSPP model. (a) Without 95%confidence interval. (b) With 95% confidence interval. In summary, the DSPP model exhibits clear advantages over the linear regression baseline, with improved accu- racy, reliability, and interpretability, as evidenced by lower RMSEandMADvalues,higherR 2 ,andtheaddedcapability of error bound predictions. These benefits make DSPP a valuable tool for enhanced RCS modeling and reliable decision-making in radar applications. D. On Explainability of the DSPP Model Automatic relevance determination (ARD) is a tech- nique in GP models that assigns individual lengthscales to each input feature, allowing the model to determine their relative importance[56]. For the top-performing DSPP model, ARD operates within the Matérn kernel structure, with dedicated lengthscales for individual features. This setup enables interpretable feature ranking as the model iter- atively optimizes these lengthscales alongside other hyper- parameters across multiple training epochs, refining them based on each feature’s impact on the model’s predictive performance. The Matérn kernel demonstrated superior performance in this study due to its flexibility in capturing complex spatial relationships while controlling the smoothness of the function through the parameterν. For this work, ν =1.5was found to be optimal, striking a balance between modeling smooth yet localized variations in RCS values, particularly those influenced by factors, such as incidence AUTHOR PREPRINT - arXiv submission draft 16 Fig. 14. 2-D histogram plot of true RCS vs. residuals for the top-performing DSPP model. Residuals have an IQR =2dBsmanda MAD =1 dBsm. angle and environmental conditions. This aligns with prior studies highlighting the kernel’s effectiveness in geospatial and radar-related applications due to its ability to model data with varying degrees of smoothness[2],[3],[57],[58],[59]. The Matérn kernel with ARD is represented as k matern (x,x )=σ 2 matern 2 1−ν (ν) ( √ 2νx−x 2 matern i ) ν K ν ( √ 2νx−x 2 matern i ) (27) where 1)σ 2 matern is a variance parameter controlling the ampli- tude of the oscillation, learned during training; 2) matern i is the lengthscale of the Matérn kernel for a featurei, learned during training. Shorter length- scales imply high sensitivity to nearby points, focus- ing on localized variations; 3)νis the smoothness parameter; for this study, it is set toν =1.5; 4)K ν isthemodifiedBesselfunctionofthesecondkind; 5) (ν)is the gamma function. Upon extensive training, ARD assigns relevance scores to each feature, as observed in the top-performing DSPP model’s configuration (see Fig.15). Feature relevance, “Importance i %,” is calculated by normalizing the inverse of each lengthscale as follows: Importance i = 1 matern i ∑ 25 j=1 1 matern j ×100.(28) Features with shorter lengthscales receive higher relevance scores, indicating greater influence on model predictions. For instance, a feature with a lengthscale ten times smaller than another will possess significantly higher predictive importance due to its increased sensitivity in the DSPP’s trained structure. The Matérn kernel’s robustness in balancing predictive accuracy with uncertainty quantification makes it an ideal choice for this DSPP model. Its ability to differentiate features based on their specific contributions to RCS pre- dictions enhances model explainability. ARD provides an interpretable and nuanced view of feature importance that aligns with the DSPP model’s architecture and operational behavior. Each feature in the ARD-based ranking, shown in Fig.15, is analyzed with reference to its correla- tion with RCS in Figs.3–4. Features critical for RCS prediction—such as incidence angle, wind speed, and target dimensions—are accurately captured and ranked by the ARD-enhanced Matérn kernel, providing clear insights into the underlying data structure. 1)Target Incidence Angle Degrees (x 14 ) – 14.12%: This feature has the highest ARD score, as expected for RCS modeling, where incidence angle is criti- cal. It directly affects the radar return strength by in- fluencing the scattering properties of the target. The Matérn kernel, particularly withν =1.5, captures smooth variations in RCS over small angle changes while accommodating localized effects. The corre- sponding correlation plot of x 14 versus RCS shows a strong dependence of RCS on incidence angle, justifying its high ARD score. 2)SAR Derived Wind Speed MPS (x 22 ) – 9.44%: Wind speed impacts sea surface roughness, influ- encing radar backscatter, particularly for maritime targets. The gradual variation in RCS due to wind speed is effectively modeled by the Matérn kernel, which balances smoothness and flexibility. The x 22 versus RCS correlation plot shows significant vari- ation in RCS with wind speed, explaining the high ARD score. 3)AIS Target Speed MPS (x 24 ) – 7.46%: The speed of the target influences how it interacts with radar signals. The Matérn kernel, with its abil- ity to assign feature-specific length scales, models these gradual changes efficiently. The x 24 versus RCS correlation plot shows a clear dependency of RCS onspeed, validatingitsimportance inthe ARD ranking. 4)AIS Target Course Over Ground Degrees (x 25 )– 5.74%: The target course over ground benefits from the Matérn kernel’s capability to model smooth tran- sitions across varying orientations. Its flexibility, enhanced by ARD, ensures the feature’s impor- tance is well-represented. This is confirmed by the corresponding correlation plot for x 25 versus RCS, where distinct RCS patterns correspond to different headings. AUTHOR PREPRINT - arXiv submission draft 17 Fig. 15. ARD ranking of the features for the top-performing DSPP model. 5)Target Azimuth Resolution Meters (x 16 ) – 5.53%: Azimuth resolution affects the spatial fidelity of radar returns. The Matérn kernel captures grad- ual variations in resolution effectively, aided by ARD for precise length scale adjustments. The x 16 versus RCS correlation plot illustrates how higher azimuth resolution improves RCS detail, reflected in the model’s high ranking for this feature. 6)AIS Target Length Metres (x 20 ) – 5.51%: Though traditionally one of the most critical fea- tures for RCS, it ranks slightly lower here, indi- cating the added predictive value of other param- eters. The Matérn kernel, with ARD, captures the smooth variations across target sizes while tailoring the length scale to this feature. The corresponding correlation plot for x 20 versus RCS shows a clear relationship between target length and RCS, sup- porting the ARD score. 7)AIS Target Width Metres (x 23 ) – 5.38%: Width impacts the radar cross-sectional area, con- tributing to RCS prediction. The Matérn kernel smooths the influence of width while ARD ensures feature-specific modeling. The x 23 versus RCS cor- relation plot affirms its relevance in defining the physical target dimensions. 8)Target Center Latitude Degrees (x 12 ) – 5.36%: Latitude offers spatial context that indirectly im- pacts RCS through environmental factors. The Matérn kernel models gradual geographical ef- fects effectively, while ARD allows fine-grained adjustmentsforfeaturerelevance.Thisisconfirmed by the x 12 versus RCS correlation plot. 9)Target Center Longitude Degrees (x 13 ) – 5.26%: Similar to latitude, longitude provides essential spatial context. Together with latitude, it influences environmental conditions affecting radar backscat- ter. The Matérn kernel and ARD capture these nuanced interactions, evident in the x 13 versus RCS correlation plot. 10)Relative Wind Direction Degrees (x 21 ) – 4.94%: Relative wind direction, combined with wind speed, affects sea surface roughness and radar backscatter. The Matérn kernel, augmented with ARD, captures gradual environmental impacts ef- fectively. The x 21 versus RCS correlation plot high- lights this feature’s importance in maritime RCS prediction. 11)AIS Target MMSI (x 19 ) – 4.30%: The MMSI identifier serves a dual purpose in this study. First, it correlates with specific vessel types or behaviors, indirectly affecting RCS. The Matérn kernel effectively captures these implicit relation- ships, with ARD adjusting relevance to reflect their impact. Second, the MMSI provides a mechanism for tracking repeated acquisitions of the same ves- sel, enabling insights into temporal variations in radar returns. This dual functionality contributes to its moderate importance in the ARD ranking, as evidenced by the feature’s influence on the model’s predictions and the corresponding correlation plot for x 19 versus RCS. 12)Chip Polarization (x 1 ) – 3.58%: Polarization affects radar signal interaction with the target. The Matérn kernel models smooth transitions while ARD ensures feature-specific AUTHOR PREPRINT - arXiv submission draft 18 tuning. The correlation plot forx 1 versus RCS demonstrates distinct RCS variations based on polarization. 13)Product Number of Polarizations (x 6 ) – 3.18%: This feature indicates the diversity of polarizations in radar measurements, providing additional in- formation on target structure. The Matérn kernel smooths the variations in RCS, with ARD ensuring accurate relevance ranking, as shown in the x 6 versus RCS correlation plot. 14)Satellite Look Direction Degrees (x 3 ) – 3.07%: Look direction affects the aspect angle, with the Matérn kernel modeling smooth directional tran- sitions. ARD refines its length scale for feature- specific importance. The x 3 versus RCS correlation plot supports the relevance of this feature. 15)Reference Noise Level dB (x 18 ) – 3.06%: This feature calibrates signal-to-noise in RCS pre- dictions, especially in low-signal conditions. The Matérn kernel captures this effect effectively, as reflected in the x 18 versus RCS correlation plot. 16)Satellite Track Degrees (x 2 ) – 3.02%: Track angle subtly impacts radar-target orientation, captured by the Matérn kernel. ARD ensures this feature’s relevance is properly modeled. The x 2 ver- sus RCS correlation plot suggests a minor influence on RCS, consistent with the ARD score. 17)Number of Beams (x 5 ) – 2.65%: Thenumberofbeamsaffectsradarcoverage.The x 5 versus RCS correlation plot shows limited impact, explaining the lower ARD score. 18)Pixel Spacing Metres (x 11 ) – 2.48%: Pixel spacing impacts image resolution, indirectly affecting RCS. The Matérn kernel smooths this feature’s impact, as confirmed by the x 11 versus RCS correlation plot. 19)Beam Mode (x 4 ) – 1.50%: Beam mode affects radar coverage pattern but has limited influence on RCS. This low relevance is consistent with the x 4 versus RCS correlation plot. 20)Product Number of Range Looks (x 7 ) – 1.45%: Range looks add data redundancy, with minimal impact on RCS, as shown in the x 7 versus RCS correlation plot. 21)Product Number of Azimuth Looks (x 8 ) – 0.88%: Similar to range looks, azimuth looks hold limited relevance, with low impact confirmed in the x 8 versus RCS correlation plot. 22)Target Equivalent Number of Looks (x 17 ) – 0.79%: This measure of data robustness has limited impact on RCS prediction, as reflected in the x 17 versus RCS correlation plot. 23)Line Spacing Metres (x 10 ) – 0.52%: Line spacing affects SAR image quality but has minimal relevance for RCS, as confirmed in the x 10 versus RCS correlation plot. 24)Target Ground Range Resolution Metres (x 15 )– 0.46%: Ground range resolution provides spatial detail but is secondary in RCS modeling. This is reflected in the x 15 versus RCS correlation plot and ARD score. 25)Image Projection (x 9 ) – 0.33%: This feature has a constant value (ground- range) across the dataset, rendering it nonin- formative for RCS prediction. The x 9 versus RCS correlation plot confirms this minimum impact. In summary, the Matérn kernel with ARD demonstrates remarkableflexibilityandcapabilityinmodelingthediverse relationships between input features and RCS. Its ability to assign unique length scales per feature ensures both interpretability and precision in feature ranking. Validation from the ARD-based ranking (Fig.15) and RCS correlation plots (Figs.3and4) underscores its effectiveness in this application. Because the 25 input features are treated as design variables, in testing scenarios where some of these features are unavailable, the missing parameters can be held at their nominal values or the DSPP model can be fine-tuned or retrained on the remaining inputs. The ARD importance ranking in Fig.15identifies the most critical features, ensuring these simplifications incur minimal loss of accuracy. VIII. CONCLUSION In this study, we have developed a novel DSPP model for RCS prediction in spaceborne SAR imagery, using a dataset of AIS-verified RADARSAT-2 ship detections. By leveraging the DSPP model’s Bayesian framework, we achieved not only improved prediction accuracy but also quantifiable error bounds, addressing key limitations of traditional regression approaches in handling uncer- tainty and complex, nonlinear relationships within RCS data. Our results demonstrate that the DSPP model consis- tently outperformed traditional multiple linear regression variants, as evidenced by lower RMSE values and higher R 2 scores across training, validation, and test splits. The DSPP’s use of the Matérn kernel was instrumental in cap- turing intricate interactions among the 25 features. Each node in the DSPP architecture employs its own kernel, with parameters, such as feature-specific lengthscales in- dependently learned through ARD. This nodewise tun- ing allows the model to adaptively capture hierarchical relationships, with lower layers focusing on broader pat- terns and higher layers refining representations to account for localized, complex feature interactions. Features with shorter learned lengthscales emerged as more influential, highlighting their sensitivity to fine-grained variations in the data. This architecture enabled nuanced modeling of both gradual environmental changes, such as radar-derived wind speed, and abrupt variations, such as ship-specific parame- ters, resulting in a more robust and interpretable predictive framework. The model’s ARD feature ranking underscored the im- portance of incidence angle, wind speed, and target-specific AUTHOR PREPRINT - arXiv submission draft 19 AIS parameters, reaffirming these as critical variables in RCS prediction. This interpretability, facilitated by ARD within the DSPP framework, aligns well with XAI princi- ples, making the model particularly suited for operational deployment in defence and maritime applications where transparency in predictions is vital. Looking forward, several avenues exist for extending this work. First, incorporating additional datasets from sensors with similar C-band center frequencies, such as RADARSAT Constellation Mission and Sentinel-1, would enhance the model’s cross-sensor robustness and applica- bility across broader operational scenarios. To enable this cross-sensor generalization, the sensor’s name will be in- cluded as an explicit feature, allowing the model to account for potential variations in sensor characteristics and data properties. In addition, including data from sensors with other center frequencies, such as X-band or L-band, would further extend the model’s generalization capacity, with the center frequency used as an explicit feature to capture frequency-dependent RCS variations. Second, addressing data imbalances, particularly in ship type diversity, will be crucial in refining the model’s generalizability. Employing balanced sampling or data augmentation strategies could mitigate the model’s sensitivity to overrepresented classes and improve performance across diverse maritime targets. Finally, an exciting direction for future research is the potential to reverse the RCS-prediction framework to es- timate key maritime parameters, such as the ship’s radial speed, course over ground, and environmental variables (e.g., wind speed and direction). By leveraging the model’s sensitivity to these features, we aim to explore how the DSPP model could support inverse predictions for improved situational awareness and decision-making in maritime surveillance. Through these contributions, our DSPP-based approach advances the state of RCS modeling by offering a pow- erful, interpretable, and flexible model for accurate RCS predictions in complex radar data, laying the groundwork for enhanced SAR-based ship detection, tracking, and clas- sification applications. ACKNOWLEDGMENT The authors gratefully acknowledge Dr. Jessica M. Topple of the DRDC – Atlantic Research Centre for her careful reading of the manuscript and for providing helpful com- ments that improved its clarity. REFERENCES [1] P. W. Vachon and J. Wolfe, “GMES Sentinel-1 analysis of marine ap- plications potential (AMAP),” Defence R&D Canada, Ottawa, ON, Canada, DRDC Ottawa ECR 2008-218, 2008. [Online]. Available: https://apps.dtic.mil/sti/pdfs/ADA639048.pdf [2] K. El-Darymli, C. H. Gierull, and A. Damini, “Gaussian process regression for empirical radar cross section modeling based on the SCATR ISAR dataset,” inProc.Int. Geosci. Remote Sens. Symp., 2024, p. 8908–8911, doi:10.1109/IGARSS53475. 2024.10640501. [3] K. El-Darymli, C. H. Gierull, and A. Damini, “Radar cross section modelling based on the inverse synthetic aperture radar (SAR) small craft automatic target recognition dataset, Part I data analysis,” DRDC, Ottawa, ON, Canada, Sci. Rep. DRDC-RDDC-2024-R036, 2024. [Online]. Available: https://cradpdf.drdc-rddc.gc.ca/PDFS/ unc463/p817886_A1b.pdf [4] K. El-Darymli, E. W. Gill, P. Mcguire, D. Power, and C. Moloney, “Automatic target recognition in synthetic aperture radar imagery: A state-of-the-art review,”Access, vol. 4, p. 6014–6058, 2016, doi:10.1109/ACCESS.2016.2611492. [5] Directorate of Capability Integration, “Defence enhanced surveil- lance from space - project (DESSP),” Accessed: Oct. 24, 2024. [On- line]. Available: https://apps.forces.gc.ca/en/defence-capabilities- blueprint/project-details.asp?id=1791 [6]Recommended Practice for Radar Cross-Section Test Pro- cedures,Standard 1502-2020 (Revision ofStandard 1502-2007), p. 1–78, 2020, doi: [7] M. I. Skolnik and W. N. Shaddix, “Potential applications of other- radar technology to civil marine radar,” Nav. Res. Lab, Washington, DC, USA, Tech. Rep. AD0754408, Jan. 1973. [Online]. Available: https://apps.dtic.mil/sti/citations/AD0754408 [8] M. I. Skolnik, “An empirical formula for the radar cross sec- tion of ships at grazing incidence,”Trans. Aerosp. Elec- tron. Syst., vol. AES- 10, no. 2, p. 292–292, Mar. 1974, doi:10.1109/TAES.1974.307935. [9] M.I.Skolnik,Radar Handbook, 2nd ed. New York, NY, USA: McGraw–Hill, 1990. [10] M. I. Skolnik,Radar Handbook, 3rd ed. New York, NY, USA: McGraw–Hill, 2008. [11] P. Williams, H. Cramp, and K. Curtis, “Experimental study of the radar cross-section of maritime targets,”IEE J. Electron. Circuits Syst.,vol.2,no.4,p. 121–136,1978,doi:10.1049/ij-ecs.1978.0026. [12] IALA VTS comittee, “Preparation of operational and technical per- formance requirements for VTS systems, producing requirements for RADAR,” Int. Assoc. Mar. Aids Navigation Lighthouse Authorities, St Germain en Laye, France, Tech. Rep. G1111, Jan. 2022. [Online]. Available: https://w.iala-aism.org/product/g1111-3/ [13] S. Zhang, P. T. Pedersen, and R. Villavicencio,Probability and Mechanics of Ship Collision and Grounding. Oxford, U.K.: Butterworth-Heinemann, 2019. [14] P. W. Vachon, J. Campbell, C. Bjerkelund, F. Dobson, and M. Rey, “Ship detection by the RADARSAT SAR: Validation of detection model predictions,”Can. J. Remote Sens., vol. 23, no. 1, p. 48–59, 1997, doi:10.1080/07038992.1997.10874677. [15] F. Askari and B. Zerr, “An automatic approach to ship detection in spaceborne synthetic aperture radar imagery: An assessment of ship detection capability using RADARSAT,” NATO, Saclant Undersea Res. Centre, La Spezia, Italy, Tech. Rep. SR- 338, 2000. [Online]. Available: https://openlibrary.cmre.nato.int/handle/ 20.500.12489/566 [16] P. W. Vachon, R. A. English, and J. Wolfe, “Validation of RADARSAT-1 vessel signatures with AISLive data,”Can. J. Remote Sens., vol. 33, no. 1, p. 20–26, 2007, doi:10.5589/m07-004. [17] P. W. Vachon, R. A. English, and J. Wolfe, “Ship signatures in synthetic aperture radar imagery,” inProc.Int. Geosci. Re- mote Sens. Symp., 2007, p. 1393–1396, doi:10.1109/IGARSS. 2007.4423066. [18] P. W. Vachon, R. English, N. Sandirasegaram, and J. Wolfe, “Devel- opment of an x-band SAR ship detectability model,” DRDC-Ottawa, Ottawa, ON, Canada, Tech. Memo TM 2013-120, 2013. [Online]. Available: https://publications.gc.ca/site/eng/9.821307/publication. html [19] P. W. Vachon, C. Kabatoff, and R. Quinn, “Operational ship detection in Canada using RADARSAT,” inProc.Geosci. Remote Sens. Symp., 2014, p. 998–1001, doi:10.1109/IGARSS.2014.6946595. [20] AstroCom Assoc Inc., “Small ship model–small ship RCS model derivation,” Earth Observ. Serv. Continuity, Can. Space Agency, Ottawa, ON, Canada, Rep. 9F044-190081 2022. AUTHOR PREPRINT - arXiv submission draft 20 [21] MathWorks, “Fit power series models,” 2025. [Online]. Available: https://w.mathworks.com/help/curvefit/power.html [22] I. Pardoe,Applied Regression Modeling. Hoboken, NJ, USA: Wiley, 2020. [23] H. Liu, Y.-S. Ong, X. Shen, and J. Cai, “When Gaus- sian process meets Big Data: A review of scalable GPs,” Int. J. Neural Syst., vol. 31, no. 11, p. 4405–4423, 2020, doi:10.1109/TNNLS.2019.2957109. [24] F. Stulp and O. Sigaud, “Many regression algorithms, one uni- fied model: A review,”Neural Netw., vol. 69, p. 60–79, 2015, doi:10.1016/j.neunet.2015.05.005. [25] S. Aigrain and D. Foreman-Mackey, “Gaussian process regression for astronomical time series,”Annu. Rev. Astron. Astrophys., vol. 61, p. 329–371, 2023, doi:10.1146/annurev-astro-052920-103508. [26] M. Jankowiak, G. Pleiss, and J. R. Gardner, “Deep sigma point processes,” inProc. 36th Conf. Uncertainty Artif. Intell., 2020, p. 789–798. [Online]. Available: https://proceedings.mlr. press/v124/jankowiak20a/jankowiak20a.pdf [27] D. V. Carvalho, E. M. Pereira, and J. S. Cardoso, “Machine learning interpretability: A survey on methods and metrics,”Electronics, vol. 8, no. 8, 2019, Art. no. 832. [Online]. Available: https://w. mdpi.com/2079-9292/8/8/832 [28] K. El-Darymli, P. McGuire, E. W. Gill, D. Power, and C. Moloney, “Understanding the significance of radiometric calibration for syn- thetic aperture radar imagery,” inProc.27th Can. Conf. Elect. Comput. Eng., 2014, p. 1–6, doi:10.1109/CCECE.2014. 6901104. [29] H. Greidanus, M. Alvarez, C. Santamaria, F.-X. Thoorens, N. Kourti, and P. Argentieri, “The SUMO ship detector algorithm for satellite radar images,”Remote Sens., vol. 9, no. 3, 2017, Art. no. 246. [Online]. Available: https://w.mdpi.com/2072-4292/9/3/246 [30] D. C. Montgomery, E. A. Peck, and G. G. Vining,Introduction to Linear Regression Analysis. Hoboken, NJ, USA: Wiley, 2021. [31] A. E. Hoerl and R. W. Kennard, “Ridge regression: Biased esti- mation for nonorthogonal problems,”Technometrics, vol. 12, no. 1, p. 55–67, 1970, doi:10.2307/1271436. [32] R. Tibshirani, “Regression shrinkage and selection via the lasso,”J. Roy. Stat. Soc.: Ser. B (Methodol.), vol. 58, no. 1, p. 267–288, 1996. [Online]. Available: https://w.jstor.org/stable/2346178 [33] H. Zou and T. Hastie, “Regularization and variable selection via the elastic net,”J. Roy. Stat. Soc.: Ser. B (Stat. Methodol.), vol. 67, no. 2, p. 301–320, 2005, doi:10.1111/j.1467-9868.2005.00503.x. [34] I. Goodfellow, Y. Bengio, and A. Courville,Deep Learning. Cam- bridge, MA, USA: MIT Press, 2016. [35] Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller, “Efficient Back- Prop,” inNeural Networks: Tricks of the Trade. Berlin, Germany: Springer, 2012, p. 9–48, doi:10.1007/978-3-642-35289-8_3. [36] N. Beck and J. N. Katz, “What to do (and not to do) with time- series cross-section data,”Amer. Political Sci. Rev., vol. 89, no. 3, p. 634–647, 1995, doi:10.2307/2082979. [37] A. C. Cameron and F. A. G. Windmeijer, “An R-squared mea- sure of goodness of fit for some common nonlinear regression models,”J. Econometrics, vol. 77, no. 2, p. 329–342, 1997. [On- line]. Available: https://w.sciencedirect.com/science/article/pii/ S0304407696018180 [38] K. Biron, P. W. Vachon, J. Wolfe, and M. A. Salciccioli, “A RADARSAT-2 dataset for the application of machine learning in maritime domain awareness,” DRDC, Ottawa, ON, Canada, Sci. Rep. DRDC-RDDC-2020-R136, 2020. [Online]. Available: https: //cradpdf.drdc-rddc.gc.ca/PDFS/unc350/p812539_A1b.pdf [39] Canadian Space Agency, “RADARSAT satellites: Technical comparison,” Accessed: Oct. 30, 2024. [Online]. Available: https://w.asc-csa.gc.ca/eng/satellites/radarsat/technical- features/radarsat-comparison.asp [40] S. Särkkä and L. Svensson,Bayesian Filtering and Smoothing. Cambridge U.K.: Cambridge Univ. Press, 2023, vol. 17. [41] D. Koller and N. Friedman,Probabilistic Graphical Models: Prin- ciples and Techniques. Cambridge, MA, USA: MIT Press, 2009. [42] M. P. Deisenroth, D. Fox, and C. E. Rasmussen, “Gaussian processes for data-efficient learning in robotics and control,”Trans. Pattern Anal. Mach. Intell., vol. 37, no. 2, p. 408–423, Feb. 2015, doi:10.1109/TPAMI.2013.218. [43] R. Ghanem, D. Higdon, and H. Owhadi,Handbook of Uncer- tainty Quantification, vol. 6, New York, NY, USA: Springer, 2017. [44] D. P. Kingma, “Adam: A method for stochastic optimization,” in Proc. Int. Conf. Learn. Representations, San Deigo, California, 2015, p. 1–15. [Online]. Available: https://iclr.c/archive/w/ doku.php\%3Fid=iclr2015:main.html [45] L. Yang and A. Shami, “On hyperparameter optimization of ma- chine learning algorithms: Theory and practice,”Neurocomput- ing, vol. 415, p. 295–316, 2020. [Online]. Available: https://w. sciencedirect.com/science/article/pii/S0925231220311693 [46] J. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl, “Algo- rithms for hyper-parameter optimization,” inProc. Int. Conf. Neural Inf. Process. Syst., vol. 24, 2011, p. 2546–2554. [On- line]. Available: https://proceedings.neurips.c/paper_files/paper/ 2011/file/86e8f7ab32cfd12577bc2619bc635690-Paper.pdf [47] J. Snoek, H. Larochelle, and R. P. Adams, “Practical Bayesian optimization of machine learning algorithms,” inProc. 25th Int. Conf. Neural Inf. Process. Syst., vol. 2, 2012, p. 2951–2959. [Online]. Available: https://papers.nips.c/paper_files/paper/2012/ hash/05311655a15b75fab86956663e1819cd-Abstract.html [48] L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar, “Hyperband: A novel bandit-based approach to hyperparameter op- timization,”J. Mach. Learn. Res., vol. 18, no. 1, p. 6765–6816, Jan. 2017, doi:10.5555/3122009.3242042. [49] S. Bubeck and N. Cesa-Bianchi, “Regret analysis of stochas- tic and nonstochastic multi-armed bandit problems,”Founda- tions Trends Mach. Learn., vol. 5, no. 1, p. 1–122, 2012, doi:10.1561/2200000024. [50] NVIDIA, “NVIDIA DGX A100: The universal systems for AI infrastructure,” 2020. [Online]. Available: https://w.nvidia. com/content/dam/en-z/Solutions/Data-Center/nvidia-dgx-a100- datasheet.pdf [51] F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” J. Mach. Learn. Res., vol. 12, p. 2825–2830, Nov. 2011, doi:10.5555/1953048.2078195. [52] A. Paszke et al., “PyTorch: An imperative style, high- performance deep learning library,” inProc. Int. Conf. Neural Inf. Process. Syst., 2019, p. 2623–2631, doi:10.5555/3454287. 3455008. [53] T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A next-generation hyperparameter optimization frame- work,” inProc. 25th ACM SIGKDD Int. Conf. Knowl. Dis- cov. Data Mining, 2019, p. 2623–2631, doi:10.1145/3292500. 3330701. [54] J. R. Gardner, G. Pleiss, D. Bindel, K. Q. Weinberger, and A. G. Wilson, “GPyTorch: Blackbox matrix-matrix Gaussian pro- cess inference with GPU acceleration,” inProc. 32nd Int. Conf. Neural Inf. Process. Syst., 2018, p. 7587–7597, doi:10.5555/ 3327757.3327857. [55] I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” inProc. Int. Conf. Learn. Representations, Toulon, France, 2017, p. 1–16. [Online]. Available: https://arxiv.org/abs/ 1608.03983 [56] R. M. Neal,Bayesian Learning for Neural Networks . Berlin, Ger- man: Springer-Verlag, 1996. [57] M. L. Stein,Interpolation of Spatial Data: Some Theory for Kriging. New York, NY, USA: Springer, 2012. [58] C. E. Rasmussen and C. K. Williams,Gaussian Processes for Ma- chine Learning. Cambridge, MA, USA: MIT Press, 2006, vol. 17. [59] S. Banerjee, B. P. Carlin, and A. E. Gelfand,Hierarchical Modeling and Analysis for Spatial Data, 2nd ed. Boca Ra- ton, FL, USA: Chapman and Hall/CRC, 2014, doi:10.1201/ b17115. AUTHOR PREPRINT - arXiv submission draft 21 Khalid El-Darymli(Senior Member, received the B.Sc. degree in electrical engi- neering from Garyounis University, Benghazi, Libya, in 2001, the M.Sc. degree in computer and information engineering from the Interna- tional Islamic University of Malaysia, Kuala Lumpur, Malaysia, in 2006, and the Ph.D. degree (with distinction) in electrical engineering from Memorial University, St. John’s, NL, Canada, in 2015. From 2010 to 2014, he was a Doctoral Re- searcher with C-CORE, St. John’s, NL, Canada, where he developed algorithms for target recognition in SAR imagery. From 2014 to 2017, he worked as a Senior Engineer with Northern Radar Inc., St. John’s, NL, Canada, codeveloping a high-frequency software-defined radar for coastal ocean applications. Concurrently, he served as an instructor for both graduate and undergraduate courses in antennas with the Department of Electrical Engineering, Faculty of Engineering and Applied Science, Memorial University, St. John’s, NL, Canada, where he was appointed Adjunct Professor with the Department of Electrical Engineering, Faculty of Engineering and Applied Science, in 2025, and is currently a Fellow of the School of Graduate Studies. From mid-2017 to early 2021, he was a Research and Development Scientist with MDA Systems, Richmond, BC, Canada, where he contributed to several projects, including laser guide stars for optical communications, ground moving target indication, velocity estimation in SAR imagery, and target recognition in inverse SAR imagery. Since mid-2021, he has been a Defence Scientist with Defence Research and Development Canada (DRDC), Ottawa, ON, Canada, focus- ing on various aspects of SAR and ISAR imagery, along with other research and development activities. In June 2023, he was appointed Affiliate Faculty with the Fowler School of Engineering, Chapman University, Orange, CA, USA. Dr. El-Darymli is a licensed Professional Engineer (P.Eng.) with the Association of Professional Engineers and Geoscientists of Alberta. He was the recipient of the Ocean Industries Student Research Award from the Research and Development Corporation (InnovateNL), Government of Newfoundland and Labrador, for his work on SAR Detection in Cluttered Environments during his Ph.D. studies. He was also the recipient of theAnnual Newfoundland Electrical and Computer Engineering Conference 2013 Wally Read Best Student Paper Award. During his M.Sc. studies, he developed an award-winning Speech to American Sign Language Interpreter System using machine learning techniques. Christoph H. Gierull(Senior Member, received the Dr.-Ing. degree in electrical engineering from Ruhr-Universität Bochum, Bochum, Germany, in 1995. From 1991 to 1994, he was a Scientist with theResearchEstablishmentforAppliedScience, Wachtberg, Germany. In 1994, he joined the German Aerospace Center, Oberpfaffenhofen, Germany, where he led the SAR Simulation Group. As part of a DLR team, he assured the X-band interferometric SAR performance dur- ing the Space-Shuttle Radar Topography Mission with NASA Mission Control, Houston, TX, USA. From 2006 to 2009, he was with Fraunhofer FHR, Wachtberg, Germany. Since 2000, he has been a Senior Defence Scientist with Defence Research and Development Canada, Ottawa, ON, Canada. In 2010, he became the Leader of the Space-Based Radar Group. In 2011, he was appointed as an Adjunct Professor with Université Laval, Quebec City, QC, Canada, and since 2017, he has been also with Simon Fraser University, Burnaby, BC, Canada. He has authored or coauthored numerous journal publications, scientific reports, and a chapter inAppli- cations of Space-Time Adaptive Processing(IEE, 2004). He has coedited and authored a chapter in,Novel Radar Techniques and Applications(IET, 2017). He filed several Report-of-Inventions and was granted a CAN/U.S. patent on vessel detection in SAR imagery. Dr. Gierull was elected as a Fellow of IET. He was the recipient of the Best Annual Paper Award of the Association of German Electrical Engineers in 1998, and the co-recipient of the best paper awards at the International Radar Conference 2004, and the EUSAR 2006, as well as EUSAR 2016. He was also the recipient of DRDC’s highest S&T Performance Excellence Award three times, in 2013 and 2020, and again in 2022 in addition to theGeoscience and Remote Sensing Society 2016 J-STARS Paper Award. He and his team were the recipients of the SquareDanceArnoldAwardin2021.Heinitiatedandco-organizedthefirst Special Issue on Multichannel Space-Based SAR in theJ OURNAL OF SELECTEDTOPICS INAPPLIEDEARTHOBSERVATIONS ANDREMOTE SENSINGin 2015. He was an Associate Editor (AE) of EURASIP’s Signal Processing in 2004. Since 2019, has been serving as an AE for T RANSACTIONS ONGEOSCIENCE ANDREMOTESENSING. He has been a Technical Advisor to the Canadian Space Agency on several SAR missions including RADARSAT-2, RCM, and its follow-on. In 2021, the European Space Agency selected him as a Member on its Sentinel-1 Next-Generation Mission Advisory Group. Katerina Bironreceived theB.Sc.degreein me- chanical engineering from Old Dominion Uni- versity in Virginia, Norfolk, VA, USA, 2007, and the M.Sc. and Ph.D. degrees in electrical engineering and signal processing from the Uni- versity of New Brunswick, Fredericton, NB, Canada, in 2010 and 2017. She currently is a Research Scientist with the Space and Intelligence, Surveillance and Recon- naissance (ISR) Applications Section, Defence Research and Development Canada, Ottawa Re- search Centre, Ottawa, ON, Canada. Her research interests include ma- chine learning and Big Data analysis, data processing and exploitation, and geospatial data management, for remote sensing and space-based ISR for defence and civilian applications. Weimin Huang(Senior Member,re- ceived the B.S., M.S., and Ph.D. degrees in radio physics from Wuhan University, Wuhan, China, in 1995, 1997, and 2001, respectively, and the M.Eng. degree in electrical engineering from the Memorial University of Newfoundland, St. John’s, NL, Canada, in 2004. From 2008 to 2010, he was a Design Engi- neer with Rutter Technologies, St. John’s, NL, Canada.Since2010,hehasbeenwiththeFaculty of Engineering and Applied Science, Memorial University of Newfoundland, where he is currently a Professor. He has authored or coauthored more than 300 refereed research articles. He has edited the bookOcean Remote Sensing Technologies: High Frequency, Marine,andGNSS-BasedRadar.Hisresearchinterestsincludemappingof oceanic surface parameters via high-frequency ground wave radar, X-band marine radar, synthetic aperture radar, and global navigation satellite systems. Dr. Huang was a Member and the Co-Chair of the Electrical and Com- puter Engineering Evaluation Group for Natural Sciences and Engineering Research Council of Canada Discovery Grants, from 2018 to 2021. He was the recipient of a Postdoctoral Fellowship from the Memorial University of Newfoundland, the Discovery Accelerator Supplements Award from NSERC in 2017, andGeoscience and Remote Sensing Society 2019 Letters Prize Paper Award as well as some other teaching and research awards. He is anOceanic Engineering Society (OES) Distinguished Lecturer. He is also an Area Editor forC ANADIAN JOURNAL OFELECTRICAL ANDCOMPUTERENGINEERING, an Associate Editor forT RANSACTIONS ONGEOSCIENCE ANDREMOTESENSING, G EOSCIENCE ANDREMOTESENSINGLETTERS,JOURNAL OF OCEANICENGINEERING,Remote Sensing,Frontiers in Marine Science, Frontiers in Remote Sensing, and has been a Guest Editor for J OURNAL OFSELECTEDTOPICS INAPPLIEDEARTHOBSERVATIONS AND REMOTESENSINGand six other journals. He has been a Reviewer for more than 120 international journals and a reviewer for manyinternational conferences. He has been a Technical Program Committee Member. He is the General Co-Chair and the Technical Program Co-Chair for the 20th International Symposium on Antenna Technology and Applied Electromagnetics, andOES 13th Currents, Waves, and Turbulence Measurement Workshop. AUTHOR PREPRINT - arXiv submission draft 22