Paper deep dive
Physics-Informed Diffusion Model for Generating Synthetic Extreme Rare Weather Events Data
Marawan Yakout, Tannistha Maiti, Monira Majhabeen, Tarry Singh
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/13/2026, 12:20:08 AM
Summary
The paper introduces a physics-informed diffusion model based on the Context-UNet architecture to generate synthetic, multi-spectral satellite imagery of rare, extreme weather events (specifically tropical cyclones). By conditioning the model on atmospheric parameters like wind speed and ocean heat content, the authors address severe class imbalance (e.g., 202 samples of Class 4 vs 79,768 of Class 0) and improve data augmentation for operational weather detection algorithms, achieving a Log-Spectral Distance of 4.5dB.
Entities (5)
Relation Signals (3)
Context-UNet → iscoreof → Physics-informed Diffusion Model
confidence 100% · We propose a physics-informed diffusion model based on the Context-UNet architecture
Physics-informed Diffusion Model → generates → Synthetic Satellite Imagery
confidence 95% · generate synthetic, multi-spectral satellite imagery of extreme weather events
Physics-informed Diffusion Model → mitigates → Class Imbalance
confidence 90% · Results demonstrate that our model successfully learns discriminative features across ten distinct context classes, effectively mitigating the data bottleneck.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Data scarcity is a primary obstacle in developing robust Machine Learning (ML) models for detecting rapidly intensifying tropical cyclones. Traditional data augmentation techniques (rotation, flipping, brightness adjustment) fail to preserve the physical consistency and high-intensity gradients characteristic of rare Category 4-equivalent events, which constitute only 0.14\% of our dataset (202 of 140,514 samples). We propose a physics-informed diffusion model based on the Context-UNet architecture to generate synthetic, multi-spectral satellite imagery of extreme weather events. Our model is conditioned on critical atmospheric parameters such as average wind speed, type of Ocean and stage of development (early, mature, late etc) -- the known drivers of rapid intensification. Using a controlled pre-generated noise sampling strategy and mixed-precision training, we generated $16\times16$ wind-field samples that are cropped from multi-spectral satellite imagery which preserve realistic spatial autocorrelation and physical consistency. Results demonstrate that our model successfully learns discriminative features across ten distinct context classes, effectively mitigating the data bottleneck. Specifically, we address the extreme class imbalance in our dataset, where Class 4 (Ocean 2, early stage with average wind speed 50kn hurricane) contains only 202 samples compared to 79,768 samples in Class 0. This generative framework provides a scalable solution for augmenting training datasets for operational weather detection algorithms. The average Results yield an average Log-Spectral Distance (LSD) of 4.5dB, demonstrating a scalable framework for enhancing operational weather detection algorithms.
Tags
Links
- Source: https://arxiv.org/abs/2603.06782v1
- Canonical: https://arxiv.org/abs/2603.06782v1
Trouble viewing inline? Open PDF directly →
Full Text
66,550 characters extracted from source content.
Expand or collapse full text
Abstract Data scarcity is a primary obstacle in developing robust Machine Learning (ML) models for detecting rapidly intensifying tropical cyclones. Traditional data augmentation techniques (rotation, flipping, brightness adjustment) fail to preserve the physical consistency and high-intensity gradients characteristic of rare Category 4-equivalent events, which constitute only 0.14% of our dataset (202 of 140,514 samples). We propose a physics-informed diffusion model based on the Context-UNet architecture to generate synthetic, multi-spectral satellite imagery of extreme weather events. Our model is conditioned on critical atmospheric parameters such as average wind speed, type of Ocean and stage of development (early, mature, late etc)—the known drivers of rapid intensification. Using a controlled pre-generated noise sampling strategy and mixed-precision training, we generated 16×1616× 16 wind-field samples that are cropped from multi-spectral satellite imagery which preserve realistic spatial autocorrelation and physical consistency. Results demonstrate that our model successfully learns discriminative features across ten distinct context classes, effectively mitigating the data bottleneck. Specifically, we address the extreme class imbalance in our dataset, where Class 4 (Ocean 2, early stage with average wind speed 50kn hurricane) contains only 202 samples compared to 79,768 samples in Class 0. This generative framework provides a scalable solution for augmenting training datasets for operational weather detection algorithms. The average Results yield an average Log-Spectral Distance (LSD) of 4.5dB, demonstrating a scalable framework for enhancing operational weather detection algorithms. keywords: Machine Learning; Data Augmentation; Context-UNet; Physics-informed Diffusion Models; Tropical Cyclones; Synthetic Satellite Imagery; Atmospheric Physics; Extreme Weather Events Prediction; Generative Data Augmentation; Natural Science; Rapid Intensification; Synthetic Data Generation 1 1 0 ://doi.org/ -Informed Diffusion Model for Generating Synthetic Extreme Rare Weather Events Data Yakout 1 , Tannistha Maiti 2 , Monira Majhabeen 2 and Tarry Singh 2 Yakout, Tannistha Maiti, Monira Majhabeen and Tarry Singh : mmyay1@student.london.ac.uk; yakout@marawan.net (M.Y.); tannistha.maiti@deepkapha.com (T.M.) 1 Introduction Data scarcity poses a significant challenge for extreme weather events in the context of developing Machine Learning (ML) models for rare weather events. Forecasting of physical systems over both space and time is a major problem that has many real-world applications, including natural sciences, transportation, and energy systems. Numerical Weather Prediction (NWP) systems currently predict the weather using complex physical models and large supercomputers [6]. In the past decade, the data generated from spatiotemporal Earth Observations have increased in an unprecedented way, allowing data-driven forecasting models to increase and improve using deep learning techniques. Currently, ML models’ data enhancement is used for the detection of rapidly intensifying tropical cyclones. However, these models are not able to capture extreme events due to a lack of data, generating physically implausible predictions. This is mainly due to the rarity of such events, blurred forecasts, and the need for massive data to train such models. In addition to the high computational cost of training such models [5; 6]. However, the most extreme events (e.g, Category 5 hurricanes that undergoing rapid intensification just before landfall) are statistically rare. Recently, Machine Learning Weather Prediction (MLWP) models have emerged, challenging the performance of the existing NWP systems’ approaches. These models are not physics-based but mainly data-driven, mainly made as a result of deep learning algorithmic advancements and discoveries. [7] We propose leveraging diffusion models to generate synthetic data samples of rare weather events to augment training data for other AI/ML models. A critical requirement is ensuring these synthetic images exhibit physical consistency, real enough with the known physics of the atmosphere, to ensure that the synthetic data do not confuse downstream model training. Moreover, the specific atmospheric conditions that lead to the Rapid Intensification (RI) of tropical cyclones are well known. Our generative model is conditioned on parameters defining RI. (e.g., low vertical wind shear + high ocean heat content). We leverage multimodal satellite data from NASA Global Precipitation Measurement (GPM) satellite data, combined with NASA Geostationary Operational Environmental Satellite (GOES) imagery RI data to identify intense convective bursts [3]. This research designs a physics-informed diffusion model to generate synthetic, multispectral satellite images of rare tropical cyclones undergoing rapid intensification, conditioned on known atmospheric parameters, to augment robust training of operational detection algorithms. 2 Related Work Our work intersects several research areas: tropical cyclone prediction, data augmentation for imbalanced datasets, generative modeling, and physics-informed machine learning. We review relevant work in each area and position our contributions. 2.1 Tropical Cyclone Rapid Intensification Rapid intensification (RI) of tropical cyclones—defined as wind speed of at least 30 knots (15 m/s) in 24 hours [11] —remains one of the most challenging problems in operational meteorology. The atmospheric and oceanic conditions favorable for RI are well-established: low vertical wind shear, high sea surface temperatures (>26∘C), high ocean heat content, moist mid-level atmosphere and favorable upper-level outflow [11]. Despite this physical understanding, accurately predicting which storms will undergo RI remains difficult due to the complex interplay of environmental factors and internal storm dynamics. Traditional statistical approaches such as the Statistical Hurricane Intensity Prediction Scheme (SHIPS) [12] incorporate environmental predictors but show limited skill for RI forecasting. Recent deep learning approaches have shown promise: Wimmers et al. [15] use deep learning with passive microwave satellite imagery to estimate TC intensity, while Lee et al. [16] apply multi-dimensional CNNs to geostationary satellite data. Chen et al. [17] provide a comprehensive review of machine learning approaches in tropical cyclone forecast modeling. However, these methods typically focus on current intensity estimation rather than RI prediction, and their performance degrades for extreme intensification events. A fundamental limitation of data-driven approaches is the scarcity of RI events in historical records. The problem is particularly acute for Category 5-equivalent tropical cyclones, which constitute less than 0.2% of observations. In our dataset of 140,514 storm observations, the most extreme class (Class 4) contains only 202 samples compared to 79,768 in the baseline class—a nearly 400-fold imbalance. This severe class imbalance prevents supervised learning models from effectively capturing the subtle signatures preceding the most dangerous storms—precisely the problem our work addresses through synthetic data generation. 2.2 Class Imbalance and Data Augmentation Class imbalance poses fundamental challenges for supervised learning, particularly when minority classes represent the most consequential events. For image data, augmentation typically involves geometric transformations (rotation, flipping, cropping) and photometric adjustments (brightness, contrast, color jittering) [14]. While effective for general computer vision tasks, these techniques have significant limitations for atmospheric data. Arbitrary rotation of hurricane imagery violates the relationship between latitude and Coriolis-induced rotation direction (counterclockwise in the Northern Hemisphere, clockwise in the Southern Hemisphere). Photometric transformations corrupt the physical relationship between pixel intensity and meteorological quantities like wind speed or precipitation rate measured by satellite sensors [3; 4]. More fundamentally, traditional augmentation creates variations of existing samples rather than expanding coverage of the underlying data manifold. For extreme events represented by only hundreds of examples among hundreds of thousands, this limitation is particularly severe—augmented samples remain tethered to the specific storms in the training set rather than exploring the broader space of physically plausible extreme events. 2.3 Generative Models for Data Augmentation Generative Adversarial Networks (GANs) [10] learn to generate synthetic samples by training a generator network to fool a discriminator network. Conditional GANs [23] extend this framework to allow controlled generation based on class labels or other conditioning information. While GANs have achieved impressive results in high-fidelity image synthesis, they suffer from well-known training instabilities, mode collapse (failure to capture full data diversity), and difficulty generating samples for underrepresented classes—problems that are exacerbated when minority classes are already severely underrepresented. Denoising Diffusion Probabilistic Models (DDPM) [8] offer an alternative generative framework with more stable training dynamics. By gradually adding Gaussian noise to data over many timesteps and learning to reverse this process, diffusion models achieve high sample quality and diversity. Recent work has demonstrated that diffusion models can surpass GANs in image synthesis quality and sample diversity [22]. Latent diffusion models [9] extend this approach to high-resolution generation by operating in compressed latent spaces, improving computational efficiency. 2.4 Application to weather and climate data Recent work has begun exploring diffusion models for meteorological applications. GenCast [20] applies diffusion models to ensemble weather forecasting, demonstrating competitive skill with traditional numerical weather prediction. Leinonen et al. [18] use latent diffusion models for precipitation nowcasting with accurate uncertainty quantification. Ravuri et al. [19] demonstrate that deep generative models can produce skillful precipitation forecasts. These works validate the feasibility of diffusion models for generating physically plausible atmospheric fields, though they focus on forecasting rather than synthetic data augmentation for class imbalance—the gap our work addresses.The stable training properties and sample diversity of diffusion models make them particularly attractive for augmenting minority classes. However, their application to scientific data requires careful consideration of domain-specific constraints to ensure generated samples are not only statistically plausible but also physically realistic. 2.5 Physics-informed Generation Incorporating physical knowledge into machine learning models has gained increasing attention across scientific and engineering domains. For atmospheric modeling, physical consistency is particularly important. Wind fields should satisfy mass continuity, respect geostrophic balance at large scales, and maintain realistic spatial autocorrelation structures. Our work makes several novel contributions relative to existing literature: • Diffusion for extreme weather enhancement: Building on recent applications of diffusion models to weather data [20; 18; 19], we present the first work that specifically targets synthetic data enhancement for extreme weather class imbalance rather than forecasting. While prior work validates that diffusion models can generate physically plausible atmospheric fields, we extend this to addressing the severe data scarcity problem in extreme event prediction. • Addressing severe class imbalance: We demonstrate effective generation for a minority class with only 202 samples within a dataset of 140,514 observations—a nearly 400-fold imbalance more severe than typically addressed in the machine learning literature. Our pre-generated noise strategy ensures fair representation of rare classes during training, addressing a subtle but important challenge when applying diffusion models to severely imbalanced datasets. • Comparison to GAN-based approaches. As discussed in Section 2.3, compared to standard GAN-based augmentation [10], our physics-informed diffusion approach avoids mode collapse and yields greater sample diversity—properties that are crucial for training downstream detection models that must generalize across varying storm morphologies. 3 Proposed Workflow Overview As shown in the figure 1 below, our approach consists of three main stages. In the forward process (left), clean wind field data x0x_0 —(16×1616× 16 spatial resolution) progressively corrupted by adding Gaussian noise following the schedule q(xt|x0)q(x_t|x_0), utilizing a pre-generated noise strategy where ϵ∼(0,I)ε (0,I) is stored and consistently reused across training epochs to ensure fair representation of rare Class 4 samples. The training phase (center) employs a Context-UNet architecture that takes three inputs: the noisy sample xtx_t, a sinusoidal timestep embedding t, and a physics-aware context vector c encoding atmospheric parameters (wind shear, ocean heat content, SST anomalies) as one-hot classes. The U-Net (with 64 base features) predicts the noise component ϵθ(xt,t,c) _θ(x_t,t,c), which is optimized via MSE loss to match the actual noise. The reverse process (right) uses the trained model to iteratively denoise pure Gaussian noise xTx_T over 500 steps following pθ(xt−1|xt)p_θ(x_t-1|x_t), conditioned on the desired physics parameters, ultimately generating novel synthetic extreme weather events x^0 x_0 that maintain physical consistency while addressing the severe class imbalance in the original dataset. Figure 1: Figure illustrates the three-phase pipeline for generating synthetic extreme events: Forward Process: x0x_0 (original wind field) is transformed into Gaussian noise xtx_t using a pre-generated noise strategy. Context-UNet Training: A UNet (n_feat=64n\_feat=64) predicts noise by conditioning the noisy input (xtx_t) on sinusoidal time steps (t) and physics-based context (c), such as one-hot encoded wind shear classes. Reverse Process: The trained model performs 500 iterative denoising steps starting from xTx_T to produce a synthetic extreme event x0^x _0. 4 Model Architecture The implementation utilizes the Context-UNet architecture, which is trained on single-channel 16×1616× 16 spatial wind data with pre-generated noise samples. This framework is scalable to higher spatial resolutions, though at increased computational and memory cost. In the following, we detail all hyperparameters, architectural choices, optimization strategies, and the mathematical framework underlying the forward and reverse diffusion processes. 4.1 Context U-Net Configuration The fundamental model used is Context U-Net [1], adapted for diffusion modeling with the following specifications: • Input channels: cin=1c_in=1 (single-channel grayscale wind data) • Base feature channels: nfeat=64n_feat=64 • Context feature channels: ncfeat=10n_cfeat=10 (conditional generation capacity) • Spatial resolution: h=w=16h=w=16 (16×1616× 16 pixel grid) The U-Net architecture comprises an encoder (contracting path) and a decoder (expanding path), with skip connections that preserve spatial information across hierarchical feature representations. The model predicts the noise vector ϵθ(t,t) ε_θ(x_t,t) conditioned on the noisy input tx_t and the normalized timestep t/Tt/T. 4.2 Timestep Conditioning The discrete timestep t∈1,2,…,Tt∈\1,2,…,T\ is normalized to [0,1][0,1] and embedded in the network through sinusoidal positional encodings inspired by Transformer [2] architectures. For a timestep t and embedding dimension index i, the positional encoding can be presented as the following: PE(t,2i)=sin(t100002i/d),PE(t,2i+1)=cos(t100002i/d)PE(t,2i)= ( t10000^2i/d ), (t,2i+1)= ( t10000^2i/d ) (1) where d is the embedding dimensionality. The normalized timestep t^=t/T∈(0,1] t=t/T∈(0,1] is converted into a high-dimensional sinusoidal embedding PE(t^)∈ℝdPE( t) ^d where each dimension alternates between sine and cosine functions with exponentially increasing frequencies. Specifically, the embedding at dimension index i is computed as sin(t^/100002i/d) ( t/10000^2i/d) for even dimensions and cos(t^/100002i/d) ( t/10000^2i/d) for odd numbered dimensions. • t t → Normalized timestep (0,1](0,1]. • i → Dimension index of the embedding, with range 0 to d/2−1d/2-1. • 2i2i → Even dimension index. • 2i+12i+1 → Odd dimension index. • d → Total embedding dimensionality (e.g., 256, 512) • 100002i/d10000^2i/d → Frequency scaling term. The resulting embedding is processed through a learned multi-layer perceptron and integrated into the U-Net’s convolutional blocks enabling the model to condition its denoising operations on the current noise level. 5 Diffusion Process Formulation 5.1 Pre-generated Noise Strategy Unlike standard Denoising Diffusion Probabilistic Models (DDPM) [8]that sample noise dynamically during each training iteration, this implementation utilizes a pre-generated noise strategy. Noise samples ϵ∼(0,)ε (0,I) are generated offline and stored in a single large-scale binary file to ensure reproducibility across training runs and maintain pairing integrity with the sparse dataset. As illustrated in Figure 5, the extreme scarcity of certain event classes—specifically Class 4 with only 202 samples compared to 79,768 in Class 0. This creates a nearly 400-fold imbalance and requires a highly controlled training environment. By pre-assigning a dedicated noise sequence to each image, we ensure that the model’s exposure to rare intensities is consistent across epochs. The noise tensor is structured with the shape (Nimages,T,C,H,W)(N_images,T,C,H,W), where Nimages=140,514N_images=140,514 is the total dataset size, T=500T=500 represents the diffusion timesteps, C=1C=1 for grayscale atmospheric data, and the spatial dimensions are 16×1616× 16. To verify the integrity of the generation, eight random samples were extracted and analyzed. As shown in Figure 2, each 16×1616× 16 patch exhibits the expected "white noise" characteristics of a standard normal distribution. 5.2 Forward Diffusion Process The forward process gradually corrupts clean data 0∼q(0)x_0 q(x_0) by adding Gaussian noise over T discrete timesteps following a Markov chain: q(t|t−1)=(t;1−βtt−1,βt)q(x_t|x_t-1)=N(x_t; 1- _tx_t-1, _tI) (2) where βt _t is the variance schedule controlling noise addition at each step. However, the code implementation employs an optimization: instead of iterating sequentially as 0→1→2→⋯→tx_0 _1 _2→·s _t, it computes tx_t directly in a single step via the reparameterization trick, significantly reducing computational overhead. Figure 2: Visual and statistical verification of eight random noise samples. Each patch represents a 16×1616× 16 realization of ϵ∼(0,1)ε (0,1). The consistent Gaussian distribution across samples ensures that rare event detection is not biased by noise artifacts. where ϵ∼(0,I)ε (0,I) is Gaussian noise. 5.3 Reverse Diffusion The reverse diffusion process is then used, the model learns to denoise samples by iteratively removing noise, reconstructing 0x_0 from pure Gaussian noise T∼(,)x_T (0,I). The reverse transition is modeled as a Gaussian distribution with learned parameters: The reverse process learns to denoise samples by predicting the noise component at each timestep: pθ(xt−1|xt)=(xt−1;μθ(xt,t),Σθ(xt,t))p_θ(x_t-1|x_t)=N(x_t-1; _θ(x_t,t), _θ(x_t,t)) (3) The mean is computed as: μθ(xt,t)=1αt(xt−1−αt1−α¯tϵθ(xt,t,c)) _θ(x_t,t)= 1 _t (x_t- 1- _t 1- α_t _θ(x_t,t,c) ) (4) where ϵθ _θ is the neural network that predicts the noise, and c represents optional context conditioning. 5.3.1 Linear Variance Schedule We employ a linear variance schedule with the following parameters: • Timesteps: T=500T=500 • Initial variance: β1=1×10−3=0.001 _1=1× 10^-3=0.001 • Final variance: βT=2×10−2=0.02 _T=2× 10^-2=0.02 The schedule is then constructed via linear interpolation: βt=β1+βT−β1T⋅(t−1),t∈1,2,…,T _t= _1+ _T- _1T·(t-1), t∈\1,2,…,T\ (5) This creates a sequence βtt=1T\ _t\_t=1^T that linearly increases from β1 _1 to βT _T. 5.3.2 Derived Coefficients The complementary variance (signal retention coefficient) is defined as: αt=1−βt _t=1- _t (6) The cumulative product of αt _t values, denoted α¯t α_t, is computed as: α¯t=∏s=1tαs=exp(∑s=1tlogαs) α_t= _s=1^t _s= ( _s=1^t _s ) (7) with the initial condition α¯0=1 α_0=1. This coefficient governs the signal-to-noise ratio decay throughout the diffusion trajectory. The log-sum-exp formulation is paramount for numerical stability, preventing float32 underflow when computing products over large numbers of timesteps. 5.3.3 Efficient Forward Perturbation An essential advantage of DDPM is the ability to sample tx_t at any arbitrary timestep t directly from 0x_0 without iterating through intermediate steps. The closed-form reparameterization can be presented as follows: t=α¯t0+1−α¯tϵ,ϵ∼(,)x_t= α_tx_0+ 1- α_t ε, ε (0,I) (8) This is implemented in the perturb_input function: perturb_input(,t,ϵ)=α¯t+1−α¯tϵ perturb\_input(x,t, ε)= α_tx+ 1- α_t ε (9) The noise vector ϵ ε is sourced from pregenerated noise samples. A critical design choice in this implementation is the use of pregenerated noise stored in the ./pregenerated_noise directory. This approach differs from standard DDPM implementations where noise is sampled randomly during each training iteration. In contrast, standard implementations generate noise dynamically: • Fresh noise samples drawn from (,)N(0,I) each iteration • Implemented as torch.randn_like(x) • No storage overhead, fully stochastic training 5.4 Reverse Diffusion Process The reverse diffusion process is then used, the model learns to denoise samples by iteratively removing noise, reconstructing 0x_0 from pure Gaussian noise T∼(,)x_T (0,I). The reverse transition is modeled as: pθ(t−1|t)=(t−1;θ(t,t),β~t)p_θ(x_t-1|x_t)=N(x_t-1; μ_θ(x_t,t), β_tI) (10) 5.4.1 Posterior Mean Estimation The posterior mean θ μ_θ is derived from Tweedie’s formula and the predicted noise ϵθ(t,t) ε_θ(x_t,t): θ(t,t)=1αt(t−βt1−α¯tϵθ(t,t)) μ_θ(x_t,t)= 1 _t (x_t- _t 1- α_t ε_θ(x_t,t) ) (11) This equation represents the expected value of t−1x_t-1 given tx_t and the model’s noise prediction. 5.4.2 Denoising with Stochasticity The sampling step combines the posterior mean with injected noise (except at the final timestep): t−1=θ(t,t)+βt,∼(,)x_t-1= μ_θ(x_t,t)+ _tz, (0,I) (12) For t=1t=1 (final denoising step), we set =z=0 to obtain a deterministic reconstruction. The implementation in denoise_add_noise follows: denoise_add_noise(t,t,ϵθ,)=1αt(t−(1−αt)1−α¯tϵθ)+βt denoise\_add\_noise(x_t,t, ε_θ,z)= 1 _t (x_t- (1- _t) 1- α_t ε_θ )+ _tz (13) 6 Training Configuration 6.1 Optimization Hyperparameters Table 1: Core Training Hyperparameters Parameter Value Batch size 64 Training epochs 120 Initial learning rate (ηmax _max) 1×10−41× 10^-4 Minimum learning rate (ηmin _min) 1×10−61× 10^-6 Optimizer Adam EMA decay coefficient (βEMA _EMA) 0.995 6.2 Adam Optimizer The Adam (Adaptive Moment Estimation) optimizer is employed with default PyTorch parameters: • Learning rate: α=1×10−4α=1× 10^-4 • Beta coefficients: β1=0.9 _1=0.9 (first moment), β2=0.999 _2=0.999 (second moment) • Epsilon: ϵ=1×10−8ε=1× 10^-8 (numerical stability) Adam computes adaptive learning rates for each parameter by maintaining exponential moving averages of gradients (mtm_t) and squared gradients (vtv_t): mt m_t =β1mt−1+(1−β1)gt = _1m_t-1+(1- _1)g_t (14) vt v_t =β2vt−1+(1−β2)gt2 = _2v_t-1+(1- _2)g_t^2 (15) θt+1 _t+1 =θt−αm^tv^t+ϵ = _t-α m_t v_t+ε (16) where m^t m_t and v^t v_t are bias-corrected estimates. 6.3 Cosine Annealing Learning Rate Schedule To improve convergence and generalization, we employ Cosine Annealing with warm restarts (specifically, CosineAnnealingLR in PyTorch): ηt=ηmin+12(ηmax−ηmin)(1+cos(πtTmax)) _t= _min+ 12( _max- _min) (1+ ( π tT_max ) ) (17) where: • ηt _t is the learning rate at epoch t • Tmax=120T_max=120 (total epochs) • ηmax=1×10−4 _max=1× 10^-4 (initial learning rate) • ηmin=1×10−6 _min=1× 10^-6 (minimum learning rate) This schedule provides a smooth, continuous decay from ηmax _max to ηmin _min following a cosine curve, avoiding abrupt learning rate changes and promoting stable convergence. The cosine schedule has been shown to delay difficult denoising tasks until after the midpoint of training, leading to enhanced sample quality and faster convergence compared to linear or step decay schedules. 6.4 Mixed Precision Training To accelerate training and reduce memory footprint, Automatic Mixed Precision (AMP) is employed using PyTorch’s torch.cuda.amp module: • Autocast: Automatically selects FP16 (half precision) or FP32 (single precision) for each operation based on numerical stability requirements • GradScaler: Scales loss values to prevent gradient underflow in FP16 arithmetic, then unscales gradients before optimizer updates The training loop integrates AMP as follows: Algorithm 1 Training loop approach Initialize GradScaler S for each batch (,ϵ)(x, ε) do Zero gradients with autocast(): Forward pass: ϵθ←model(t,t/T) ε_θ (x_t,t/T) Compute loss: ℒ=‖ϵθ−ϵ‖22L=\| ε_θ- ε\|_2^2 Scaled backward: .scale(ℒ).backward()S.scale(L).backward() Optimizer step: .step(optimizer)S.step(optimizer) Update scaler: .update()S.update() Update EMA model end for Mixed precision training typically yields 1.51.5–3×3× speedup on modern GPUs (e.g., NVIDIA V100, A100) with Tensor Cores, while maintaining numerical accuracy. This part is covered in more detail in Section 13.2. 7 Loss Function The training objective is the noise prediction mean squared error (MSE): ℒsimple(θ)=t,0,ϵ,[‖ϵ−ϵθ(t,t,)‖22]L_simple(θ)=E_t,x_0, ε,c [\| ε- ε_θ(x_t,t,c)\|_2^2 ] (18) where: • t∼Uniform1,2,…,Tt \1,2,…,T\ (random timestep) • 0∼q(0)x_0 q(x_0) (clean data sample) • c (conditioning vector, e.g., wind speed or class label) • ϵ∼(,) ε (0,I) (Gaussian noise) • t=α¯t0+1−α¯tϵx_t= α_tx_0+ 1- α_t ε (noisy sample) This simplified objective (omitting weighting factors) is empirically shown to produce high-quality samples and stable training dynamics. The model learns to denoise by predicting the noise component ϵ ε that was added to the clean data. 7.1 Context Conditioning For the context-conditioned variant, we implement classifier-free guidance training by randomly masking the context during training: cmasked=c⋅m,m∼Bernoulli(0.9)c_masked=c· m, m (0.9) (19) where c is the one-hot encoded context vector and m is a binary mask. Each training sample has a 10% probability of having its context completely zeroed out, allowing the model to learn both conditional and unconditional generation. This enables classifier-free guidance during inference by interpolating between conditional and unconditional predictions. 8 Data Configuration and Noise Generation 8.1 Dataset Characteristics and Preprocessing The training dataset consists of single-channel wind field data representing intense convective bursts within rare tropical cyclones. Each data sample x0x_0 is a 16×1616× 16 spatial grid, a resolution selected to balance physical detail with computational efficiency for the Context-UNet architecture. The raw imagery is sourced from NASA Global Precipitation Measurement (GPM) and Geostationary Operational Environmental Satellite (GOES) [4] datasets as shown in Figure 3. The training data consists of: • Source: Wind field data stored in ./data/wind_1D16X16.npy • Data type: NumPy array, single-channel (1D physical field) • Spatial resolution: 16×1616× 16 grid points • Normalization: Data values are clipped to [0,1][0,1] range during sample generation 8.2 Train-Validation Split The dataset is partitioned into training and validation subsets: • Training set: 90% of total data • Validation set: 10% of total data • Splitting performed via torch.utils.data.random_split • Ensures statistical representativeness with sufficient validation samples latex Figure 3: Cyclone event capture from emergence to dissipation. Note: Images were originally 1024×10241024× 1024 pixels and have been cropped to 16×1616× 16 pixels for training. To ensure numerical stability during the diffusion process, all input values are normalized using a StandardScaler to remove mean and unit variance bias before being clipped to a [0,1][0,1] range for consistency. As shown in Figure 4, the resulting intensity fields exhibit high-gradient spatial features characteristic of rapid intensification (RI). Figure 4: Storm Wind Fields representing physical intensity on a 16×1616× 16 grid. Samples are conditioned on specific atmospheric parameters (Param 0 and 1) representing Rapid Intensification (RI) conditions. The colormap highlights the high-intensity gradients preserved by the physics-informed model. 8.3 Label Distribution and Data Scarcity A significant challenge addressed in this work is the scarcity of data for extreme rare weather events. To categorize the training data, we employ a multi-class labeling system based on physical parameters defining RI, such as low vertical wind shear and high ocean heat content. Figure 5: Distribution of unique physical parameter labels within the dataset. The extreme scarcity of samples in Class 4 (202 samples) demonstrates the data bottleneck for rare weather events that this research aims to mitigate through synthetic augmentation. In the context of training diffusion models with grayscale images, the choice of scaling is critical because diffusion processes—specifically the forward and reverse Gaussian noise additions—are mathematically designed to operate on a specific range. While basic image processing often uses Min-Max to scale pixels to [0,1][0,1], diffusion models almost exclusively use Min-Max to scale pixels to the [−1,1][-1,1] range. For almost all modern diffusion architectures (like DDPM or Stable Diffusion) [9], Min-Max Scaling is the standard approach, but with a specific target range. DDPM (Denoising Diffusion Probabilistic Models) paper by Ho et al. (2020) Figure 5 above illustrates the distribution of these unique physical parameter labels across the training set. The distribution reveals a stark imbalance, with extreme classes (e.g., Category 4 and 5 equivalents) significantly underrepresented compared to baseline storm conditions. Specifically, Class 0 contains approximately 79,768 samples, while the most extreme rare event class (Class 4) contains only 202 samples. This 400-fold difference in class frequency underscores the critical need for the proposed generative model to augment training data for downstream operational detection algorithms. 9 Model Architecture 9.1 Network Configuration The model employs a Context U-Net architecture designed for conditional diffusion modeling. The configuration parameters are presented in Table 2. Parameter Value Architecture Context U-Net Input channels 1 (grayscale) Feature dimension (nfeatn_feat) 64 Context features (ncfeatn_cfeat) 5 (standard), 10 (context) Image resolution 16 × 16 pixels Activation function SiLU (Swish) Normalization Group Normalization Table 2: Model architecture configuration parameters. 9.2 Exponential Moving Average To improve sample quality and training stability, we maintain an exponential moving average (EMA) of the model weights: θEMA(t)=βEMA⋅θEMA(t−1)+(1−βEMA)⋅θ(t) _EMA^(t)= _EMA· _EMA^(t-1)+(1- _EMA)·θ^(t) (20) where βEMA=0.995 _EMA=0.995 is the decay factor, θ(t)θ^(t) are the model parameters at step t, and θEMA(t) _EMA^(t) are the EMA parameters. The EMA model is updated after each gradient step and is used for inference and sample generation, while the standard model is used for training. 10 Training Methodology 10.1 Training Hyperparameters The complete set of training hyperparameters is detailed in Table 3. Parameter Value Batch size 64 Training epochs 120 Initial learning rate 10−410^-4 Optimizer Adam Adam β1 _1 0.9 Adam β2 _2 0.999 Weight decay 0 Learning rate scheduler Cosine Annealing Minimum learning rate 10−610^-6 EMA decay rate 0.995 Data split ratio 90% train, 10% validation Context masking probability 0.1 (context model) Table 3: Training hyperparameters and optimization settings. 10.2 Learning Rate Schedule A cosine annealing learning rate schedule is employed to gradually reduce the learning rate over training: ηt=ηmin+12(ηmax−ηmin)(1+cos(tTmaxπ)) _t= _ + 12( _ - _ ) (1+ ( tT_ π ) ) (21) where ηmax=10−4 _ =10^-4 is the initial learning rate, ηmin=10−6 _ =10^-6 is the minimum learning rate, Tmax=120T_ =120 represents the total number of epochs, and t is the current epoch. This schedule provides smooth decay from the initial learning rate to the minimum, promoting stable convergence in later training stages. 11 Training Procedure 11.1 Training Loop (Per Epoch) Algorithm 2 Training Loop for Diffusion Model 1: Initialize running loss ℒtrain=0L_train=0 2: for each batch (,ϵ)(x, ε) in training DataLoader do 3: Zero optimizer gradients 4: Move data to GPU: ←.to(device)x .to(device), ϵ←ϵ.to(device) ε← ε.to(device) 5: Sample random timesteps: t∼Uniform1,…,500t \1,…,500\ for each batch element 6: Perturb input: t←perturb_input(,t,ϵ)x_t← perturb\_input(x,t, ε) 7: with autocast(): 8: Predict noise: ϵθ←model(t,t/500) ε_θ (x_t,t/500) 9: Compute loss: ℒ←MSE(ϵθ,ϵ)L ( ε_θ, ε) 10: Scaled backward pass: scaler.scale(ℒ).backward() scaler.scale(L).backward() 11: Optimizer step: scaler.step(optimizer) scaler.step(optimizer) 12: Update scaler: scaler.update() scaler.update() 13: Update EMA model: θEMA←0.995⋅θEMA+0.005⋅θ _EMA← 0.995· _EMA+0.005·θ 14: Accumulate loss: ℒtrain←ℒtrain+ℒL_train _train+L 15: end for 16: Compute average training loss: ℒ¯train=ℒtrain/Nbatches L_train=L_train/N_batches 17: Step learning rate scheduler: scheduler.step() scheduler.step() 12 Sampling and Inference 12.1 Generation Procedure Sample generation employs the EMA model and follows the reverse diffusion process: 12.2 Validation Loop (Per Epoch) Algorithm 3 Validation Step for Diffusion Model 1: Set model to evaluation mode: model.eval() 2: Initialize validation loss ℒval=0L_val=0 3: with torch.no_grad(): 4: for each batch x in validation DataLoader do 5: Move data to GPU and cast to float32 6: Extract single channel: ←[:,0:1,:,:]x [:,0:1,:,:] 7: Sample random timesteps: t∼Uniform1,…,500t \1,…,500\ 8: Sample noise: ϵ∼(,) ε (0,I) 9: Perturb input: t←perturb_input(,t,ϵ)x_t← perturb\_input(x,t, ε) 10: Predict noise: ϵθ←model(t,t/500) ε_θ← model(x_t,t/500) 11: Compute loss: ℒ←MSE(ϵθ,ϵ)L ( ε_θ, ε) 12: Accumulate loss: ℒval←ℒval+ℒL_val _val+L 13: end for 14: Compute average validation loss: ℒ¯val=ℒval/Nbatches L_val=L_val/N_batches 15: Set model back to training mode: model.train() Note: Validation evaluates the standard model θ, not the EMA model θEMA _EMA, to monitor the primary training trajectory. Algorithm 4 Sample Generation for Diffusion Model 1: Set EMA model to evaluation mode: ema_model.eval()ema\_model.eval() 2: Initialize from pure noise: T∼(,)x_T (0,I), shape (N,1,16,16)(N,1,16,16) 3: for i=T,T−1,…,1i=T,T-1,…,1 do 4: Normalize timestep: t←i/Tt← i/T 5: if i>1i>1 then 6: Sample noise: ∼(,)z (0,I) 7: else 8: Set deterministic: ←z 0 9: end if 10: Predict noise: ϵθ←ema_model(i,t) ε_θ \_model(x_i,t) 11: Denoise: i−1←denoise_add_noise(i,i,ϵθ,)x_i-1← denoise\_add\_noise(x_i,i, ε_θ,z) 12: end for 13: Clamp to valid range: 0←clamp(0,0,1)x_0 (x_0,0,1) 14: Save as image grid: save_image(0,path,nrow=4) save\_image(x_0,path,nrow=4) This procedure generates N=16N=16 samples by default, arranged in a 4×44× 4 grid. Samples are generated every 4 epochs during training for qualitative assessment. Table 4: Complete Training Configuration Summary Component Specification Model Architecture Model Context U-Net Input channels 1 (grayscale) Base features 64 Context features 5 Spatial resolution 16×1616× 16 Training Configuration Batch size 64 Epochs 120 Optimizer Adam (β1=0.9 _1=0.9, β2=0.999 _2=0.999) Learning rate (initial) 1×10−41× 10^-4 Learning rate (minimum) 1×10−61× 10^-6 LR scheduler Cosine Annealing EMA decay 0.995 Loss function MSE (noise prediction) Mixed precision Enabled (autocast + GradScaler) Checkpointing Frequency Every 4 epochs + final Saved models Standard + EMA Generated samples 16 (4×44× 4 grid) 13 Results To evaluate the model’s ability to generate class-specific wind patterns, we conducted conditional generation experiments using one-hot encoded context vectors. The model was trained with ncfeat = 10 context features, corresponding to 10 distinct wind pattern classes (labeled 0–9). The model architecture consists of a U-Net backbone with contextual embedding layers, trained for 500 diffusion timesteps using a linear beta schedule ranging from β1=10−4 _1=10^-4 to β2=0.02 _2=0.02. All experiments were conducted on 16×1616× 16 grayscale wind field images, with pixel values normalized to the range [−1,1][-1,1] using a min-max scaler fitted on the training dataset. The dataset consists of different storm characteristics organized into two ocean groups. 13.1 Context-Conditioned Generation Ocean Group A encompasses storms primarily forming in Ocean 1, following a clear lifecycle progression from early development through peak intensity to eventual dissipation. The journey begins with nascent systems (Contexts 7 & 9) averaging 33 hours old with winds around 47 knots, then strengthens through Context 0 at approximately 50 hours and 51 knots. These storms reach their peak maturity in Context 1 around the 5-day mark (120 hours), achieving maximum average winds of 56 knots for this ocean group. As systems age beyond this point, they enter prolonged weakening phases (Contexts 6 & 5) spanning 6–10 days with diminishing winds of 41–43 knots. Notably, Context 3 represents an exceptional subset of extraordinarily persistent storms that survive beyond 17 days (416 hours), though these rare events constitute only a small fraction of Ocean 1’s storm population. Ocean Group B, dominated by Ocean 2 systems, exhibits a more concentrated development pattern with notably higher intensities at maturity. The majority of these storms (Contexts 2 & 4) are captured during their formative stages around 41 hours old, displaying wind speeds between 47–53 knots—comparable to early Ocean 1 development but progressing through a different evolutionary trajectory. The defining characteristic of Ocean Group B emerges in Context 8, which represents large, exceptionally powerful mature cyclones that maintain their strength far longer than their Ocean 1 counterparts. These Context 8 systems reach an impressive average of 65 knots—the highest wind speed across all contexts—while persisting for approximately one week (168 hours), demonstrating Ocean 2’s capacity to sustain significantly more intense storms at extended durations compared to Ocean Group A’s typical lifecycle. Table 5: Storm lifecycle characteristics across different contexts and ocean basins Context Ocean Stage Avg. Wind Speed Description 7, 9 1 Early 47 kn Initial formation/Early development 0 1 Mid-Early 51 kn Intensifying tropical system 1 1 Mature 56 kn Peak intensity phase 6, 5 1 Late 42 kn Long-duration, weakening storms 3 1 Extreme 43 kn Rare, ultra-long-lived systems 2, 4 2 Early 50 kn Standard development in Ocean 2 8 2 Late-Peak 65 kn Major mature storms (Typhoons/Cyclones) To evaluate the model’s ability to generate class-specific wind patterns, we conducted conditional generation experiments using one-hot encoded context vectors. The model was trained with ncfeat=10n_cfeat=10 context features, corresponding to 10 distinct wind pattern classes (labeled 0–9). Figure 6 demonstrates the model’s capacity for controlled generation across all 10 context classes. Each sample corresponds to a specific context label, generated by passing the respective one-hot encoded vector through the context conditioning mechanism. Figure 6: Context-conditioned generation for labels 0-9. Each image represents a wind field pattern generated conditioned on its corresponding class label using one-hot encoding. The model successfully produces distinct patterns for different context inputs. The results in Figure 6 reveal that the model learns discriminative features for each context class. Classes 2,4 and 7 produce more structured patterns with pronounced intensity variations. This demonstrates the effectiveness of the context embedding architecture in capturing class-specific characteristics during the denoising process. To further investigate the model’s learned representations, we analyzed the generation quality for individual context classes. Figures 7(a) and 7(b) show multiple samples generated for context labels 1 and 5, respectively. ((a)) ((b)) Figure 7: Multiple samples generated for specific context labels. (a) Context label 1 produces predominantly low-intensity, smooth gradient fields. (b) Context label 5 generates high-contrast patterns with distinct spatial structures and intensity clusters. Qualitative assessment of the generated outputs reveals a significant correlation between the conditioning level and the structural complexity of the wind fields. At a lower conditioning level (Context 1), the model produces relatively uniform, high-frequency patterns characterized by soft morphology and low contrast, suggesting a steady-state or calm atmospheric representation. In contrast, Context 5 triggers the generation of highly structured features, including distinct localized "cells" and vortex-like structures with sharp gradients. These results indicate that the contextual embedding layers effectively modulate the U-Net’s feature maps, enabling the model to transition from generating basic Gaussian-like noise to complex, turbulent-like meteorological states as the context value increases. A comparison of the generated wind fields across various conditioning states reveals a clear progression in structural resolution and morphological detail. Figures 8(a) and 8(b) compare samples for Context 8 at epochs 116 and 4, respectively. At lower training checkpoints (epoch 4) or lower-order conditioning (Context 8), the model exhibits limited capacity for high-fidelity reconstruction, resulting in "blurry" or "hazy" outputs characterized by high-frequency noise and poorly defined boundaries. These early-stage results represent a global mean of the training data, where the U-Net backbone has yet to learn the localized, sharp gradients necessary for realistic wind modeling. As conditioning increases through training (epoch 116), the model successfully transitions from stochastic noise to coherent structural features. This is evidenced by the emergence of distinct localized "eyes" and cellular vortex structures, suggesting that the contextual embedding layers effectively guide the denoising process to capture the non-linear fluid dynamics inherent in the storm dataset. Despite the coarse 16×1616× 16 pixel resolution, the higher-context samples demonstrate a marked improvement in contrast and spatial organization, effectively utilizing the limited pixel space to represent complex atmospheric phenomena. ((a)) ((b)) Figure 8: As the conditioning increases to epoch 116, the U-Net backbone successfully synthesizes well-defined localized gradients and cellular structures. At lower conditioning levels, the model produces high-frequency, low-contrast outputs indicative of early-stage training or weak prior guidance. Although visual inspection demonstrates qualitative success, we note several key observations regarding model performance. • Spatial coherence: Generated samples maintain realistic spatial autocorrelation patterns typical of wind field data, avoiding checkerboard artifacts or high-frequency noise. • Diversity: The model produces varied samples within each context class, suggesting adequate mode coverage and avoiding mode collapse. • Context fidelity: The clear visual distinctions between context classes indicate successful conditioning, with the model responding appropriately to different one-hot encoded inputs. • Scaling consistency: The inverse StandardScaler transformation successfully maps generated samples back to the physically meaningful [0, 255] intensity range, with proper clipping to handle edge cases. The results demonstrate that the proposed context-conditioned DDPM architecture effectively learns to generate realistic wind field patterns while maintaining controllability through context conditioning. 13.2 Implementation Details All experiments were implemented in PyTorch 2.0+ and executed on NVIDIA GPUs with CUDA support. The denoising model architecture consists of a ContextUnet with 64 base features (nfeat=64n_feat=64), processing single-channel grayscale images of size 16×1616× 16 pixels. The diffusion process employs 500 timesteps with a linear variance schedule, and generation is performed using the DDPM sampling algorithm without classifier-free guidance. The model achieved a Log-Spectral Distance (LSD) of 4.5 dB, indicating that the generated samples successfully captured the fundamental global structure [21]. Although this score reflects a strong statistical alignment, the remaining distance suggests minor discrepancies in high-frequency texture or fine-grained detail compared to the ground truth. Data pre-processing involves flattening each 16×1616× 16 image to a 256-dimensional vector, followed by standardization using scikit-learn’s StandardScaler. The scaler parameters (mean and standard deviation) are fitted on the training set and saved for consistent inverse transformation during inference. Context conditioning is achieved through one-hot encoding of class labels, which are embedded and concatenated with intermediate U-Net features at multiple resolution levels. The sampling procedure initializes with Gaussian noise T∼(,)x_T (0,I) and iteratively denoises according to: t−1=1αt(t−1−αt1−α¯tϵθ(t,t,))+βt,∼(,)x_t-1= 1 _t (x_t- 1- _t 1- α_t ε_θ(x_t,t,c) )+ _tz, (0,I) (22) where ϵθ ε_θ is the learned noise predictor, c is the context vector, αt=1−βt _t=1- _t, and α¯t=∏s=1tαs α_t= _s=1^t _s. The code is made available at [https://github.com/MarawanYakout/SERWED.git] for reproducibility. 14 Discussion Our context-conditioned diffusion model effectively addresses the extreme class imbalance in meteorological datasets. The nearly 400-fold disparity between Class 0 (79,768 samples) and Class 4 (202 samples) presents a fundamental challenge for supervised learning. Traditional augmentation techniques (rotation, flipping, brightness adjustment) merely create variations of existing samples and can violate physical constraints inherent in atmospheric data. By learning the underlying generative process, our model generates genuinely novel samples that maintain physical and statistical characteristics of each storm class. This approach provides particular value for rare extreme events where limited examples make traditional augmentation insufficient. Compared to standard GAN-based augmentation, our physics-informed diffusion approach avoids mode collapse and provides higher sample diversity, which is crucial for training downstream detection models that must generalize across varying storm morphologies. A critical requirement for synthetic meteorological data is adherence to physical constraints. The context-conditioning mechanism embeds key atmospheric parameters, including wind shear, ocean heat content, and storm age, directly into the generation process, ensuring that synthesized samples reflect realistic relationships between environmental conditions and storm characteristics. Qualitative analysis reveals clear differentiation across context classes. Low-intensity contexts (0, 1, 7, 9) generate smooth gradient fields characteristic of developing or weakening systems, while high-intensity contexts (4, 8) produce localized high-gradient regions and pronounced vortex structures. The progression from Context 0 (51 knots) through Context 1 (56 knots) to Context 8 (65 knots) demonstrates that the model captures realistic storm intensification patterns rather than overfitting to specific training examples. Additionally, the coherent spatial organization of generated samples, combined with the absence of checkerboard artifacts or nonphysical high-frequency noise, suggests that the model has learned spatially correlated wind field structures rather than arbitrary pixel-level patterns. This is particularly important for applications requiring physically interpretable synthetic data. The progression from epoch 4 to epoch 116 illustrates the model’s evolving capacity to capture fine-grained atmospheric structures. Early in training (epoch 4), generated samples exhibit high-frequency noise and lack coherent spatial organization, characteristic of under-fitted generative models. By epoch 116, the model generates well-defined vortex structures and sharp intensity gradients, indicating successful learning of hierarchical atmospheric features. This evolution demonstrates that the Context-UNet architecture effectively learns multi-scale representations. Skip connections preserve spatial coherence while the encoder-decoder pathway captures increasingly abstract structural patterns. These observations align with theoretical expectations for diffusion models, where iterative denoising enables progressive refinement of spatial structure. Several methodological design choices were influenced by the need to balance computational feasibility with model performance: • Pre-generated Noise Strategy- Ensures fair representation of rare classes by providing consistent training conditions across all 120 epochs. Each image faces identical denoising challenges, eliminating variance that could disproportionately affect learning for the 202 Class 4 examples. The trade-off is substantial storage requirements and reduced stochasticity in training. • Spatial resolution- The 16×1616× 16 resolution balances computational feasibility with physical detail. While this sacrifices fine-scale features from the original 1024×10241024× 1024 imagery, it enabled efficient training and rapid iteration. Higher resolutions would quadratically increase memory requirements. • Mixed-precision training- Achieved 1.51.5–3×3× speedup on modern GPUs with Tensor Cores while maintaining numerical stability, partially addressing the high computational costs inherent in diffusion model training. 14.1 Computational Limitations Training diffusion models presents significant computational challenges, as noted in our introduction. The high computational cost stems from several factors. The model requires 120 epochs over 140,514 samples, with each sample processed at multiple random timesteps (1-500), necessitating repeated forward and backward passes through the U-Net for computing noise schedules and gradients. The pregenerated noise strategy, while ensuring reproducibility and fair representation of rare classes, requires storing a large tensor (140,514 images × 500 timesteps × 1 channel × 16 × 16 pixels), consuming substantial disk space. The 16×1616× 16 resolution was chosen specifically to make training computationally feasible; scaling to higher resolutions would dramatically increase requirements—moving to 64×6464× 64 would increase memory by 16×16× and training time by approximately 10−15×10-15×. While training is computationally intensive, generation is more efficient once the model is trained. However, each generated sample still requires 500 forward passes through the U-Net (one per denoising step), precluding real-time generation for operational deployment. These computational constraints influenced our architectural choices and represent practical limitations for scaling this approach to operational forecasting systems that require higher-resolution images or real-time generation capabilities. Beyond computational constraints, three additional limitations require acknowledgment. The spatial resolution of 16×1616× 16 sacrifices fine-scale atmospheric features such as mesoscale convection and detailed eye structure. Moreover, single-timestep generation cannot capture temporal storm evolution. Finally, while conditioning guides generation toward physical plausibility, physical relationships are learned implicitly from data rather than explicitly enforced. To remain within the T4’s 16GB VRAM envelope, we purposefully balanced the Context-UNet’s depth (nfeat=64n_feat=64) with a compressed 16×1616× 16 spatial manifold. This allowed for a 500-step diffusion trajectory that prioritizes thermodynamic consistency over raw pixel density. 14.2 Future Directions Several directions emerge as important avenues for future investigation. First, scaling to higher resolutions (64×6464× 64 or 128×128128× 128) would capture the mesoscale features currently lost at 16×16 resolution, potentially using progressive training strategies to manage the associated computational costs. Second, extending the model to time-series generation would enable modeling of storm evolution and intensification dynamics over multiple timesteps rather than single snapshots. Third, incorporating physics-informed loss functions that explicitly enforce physical constraints such as mass continuity and momentum balance would provide stronger consistency guarantees beyond the implicit learning achieved through data-driven conditioning. This methodology applies to other rare atmospheric phenomena (tornadogenesis, flash flooding, severe convection) and, more broadly, to any domain where extreme class imbalance limits ML development for applications governed by physical laws. By combining generative modeling with domain-specific conditioning, we can expand limited datasets while maintaining physical fidelity. 15 Conclusions We have developed and validated a physics-informed, context-conditioned diffusion model for generating synthetic wind field data associated with rare tropical cyclone events. By integrating key atmospheric parameters directly into the Context-UNet conditioning mechanism, we ensure that generated samples reflect realistic relationships between environmental conditions and storm characteristics. Our methodology, which includes a pre-generated noise strategy, mixed-precision training, and cosine annealing learning rate schedule, achieved stable convergence and produced spatially coherent outputs. A central contribution of this work is to address the severe class imbalance inherent in rapid intensification datasets, where rare extreme events are vastly underrepresented. By learning the underlying generative process rather than relying on conventional augmentation, the model is able to synthesize novel, physically consistent examples of rare events, thereby mitigating a critical bottleneck in extreme weather machine learning applications. The results further indicate that the model successfully learns discriminative features across ten distinct context classes, with clear differentiation between low-intensity developing storms and high-intensity mature cyclones. The training progression from epoch 4 to epoch 116 shows the model’s evolution from producing high-frequency noise to generating well-defined vortex structures with sharp intensity gradients characteristic of rapidly intensifying systems. Future work should focus on scaling to higher spatial resolutions to capture mesoscale atmospheric features, extending to temporal sequences to model storm evolution over time, and incorporating physics-informed constraints. The methodology presented here is broadly applicable to other rare atmospheric phenomena and, more generally, to any domain where extreme class imbalance limits machine learning model development for applications governed by known physical laws. References References [1] Ronneberger, O., Fischer, P., & Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (p. 234–241). Springer, Cham. [2] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., … & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. [3] Huffman, G. J., Stocker, E. F., Bolvin, D. T., Nelkin, E. J., & Tan, J. (2019). GPM IMERG Final Precipitation L3 Half Hourly 0.1 degree x 0.1 degree V06. Greenbelt, MD, Goddard Earth Sciences Data and Information Services Center (GES DISC). [cite_start]doi:10.5067/GPM/IMERG/3B-H/07[cite: 78, 474]. [4] NOAA Office of Satellite and Product Operations. (1994). NOAA Geostationary Operational Environmental Satellite (GOES) I-M and N-P Series Imager Data. NOAA National Centers for Environmental Information. [cite_start]doi:10.25921/Z9JQ-K976[cite: 78, 474]. [5] Gao, Z., Shi, X., Han, B., Wang, H., Jin, X., Maddix, D., Zhu, Y., Li, M. & Wang, Y. Prediff: Precipitation nowcasting with latent diffusion models. Advances In Neural Information Processing Systems. 36 p. 78621-78656 (2023) [6] Bauer, P., Thorpe, A. & Brunet, G. The quiet revolution of numerical weather prediction. Nature. 525, 47-55 (2015,9,1), https://doi.org/10.1038/nature14956 [7] Andrae, M., Landelius, T., Oskarsson, J. & Lindsten, F. Continuous Ensemble Weather Forecasting with Diffusion models. (2024,10,4), https://openreview.net/forum?id=ePEZvQNFDW [8] Ho, J., Jain, A. & Abbeel, P. Denoising Diffusion Probabilistic Models. Advances In Neural Information Processing Systems. 33 p. 6840-6851 (2020), https://proceedings.neurips.c/paper_files/paper/2020/file/4c5bcfec8584af0d967f1ab10179ca4b-Paper.pdf [9] Rombach, R., Blattmann, A., Lorenz, D., Esser, P. & Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. (arXiv,2022,4,13), http://arxiv.org/abs/2112.10752 [10] Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. & Bengio, Y. Generative Adversarial Networks. (arXiv,2014,6,10), http://arxiv.org/abs/1406.2661 [11] Kaplan, J., & DeMaria, M. (2003). Large-scale characteristics of rapidly intensifying tropical cyclones in the North Atlantic basin. Weather and Forecasting, 18(6), 1093–1108. [12] DeMaria, M., Mainelli, M., Shay, L. K., Knaff, J. A., & Kaplan, J. (2005). Further improvements to the Statistical Hurricane Intensity Prediction Scheme (SHIPS). Weather and Forecasting, 20(4), 531–543. [13] Wimmers, A., Velden, C. & Cossuth, J. Using Deep Learning to Estimate Tropical Cyclone Intensity from Satellite Passive Microwave Imagery. Monthly Weather Review. 147, 2261-2282 (2019,6,1), https://journals.ametsoc.org/view/journals/mwre/147/6/mwr-d-18-0391.1.xml [14] Shorten, C., & Khoshgoftaar, T. M. (2019). A survey on image data augmentation for deep learning. Journal of Big Data, 6(1), 1–48. [15] Wimmers, A. J., Velden, C. S., & Cossuth, J. H. (2019). Using deep learning to estimate tropical cyclone intensity from satellite passive microwave imagery. Monthly Weather Review, 147(6), 2261–2282. [16] Lee, J., Im, J., Cha, D., Park, H. & Sim, S. Tropical Cyclone Intensity Estimation Using Multi-Dimensional Convolutional Neural Networks from Geostationary Satellite Data. Remote Sensing. 12 (2019,12,28), https://w.mdpi.com/2072-4292/12/1/108 [17] Chen, R., Zhang, W., & Wang, X. (2019). Machine learning in tropical cyclone forecast modeling: A review. Atmosphere, 10(7), 676. [18] Leinonen, J., Hamann, U., Sideris, I. V., & Germann, U. (2023). Latent diffusion models for generative precipitation nowcasting with accurate uncertainty quantification. arXiv preprint arXiv:2304.12891. [19] Ravuri, S., Lenc, K., Willson, M., Carver, A., Chung, A., Creswell, A., … Mohamed, S. (2021). Skilful precipitation nowcasting using deep generative models of radar. Nature, 597(7878), 672–677. [20] Price, I., Sanchez-Gonzalez, A., Alet, F., Andersson, T., El-Kadi, A., Masters, D., Ewalds, T., Stott, J., Mohamed, S., Battaglia, P., Lam, R. & Willson, M. GenCast: Diffusion-based ensemble forecasting for medium-range weather. (arXiv,2024,5,1), http://arxiv.org/abs/2312.15796 [21] Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A. & Kavukcuoglu, K. WaveNet: A Generative Model for Raw Audio. (arXiv,2016,9,12), http://arxiv.org/abs/1609.03499 [22] Dhariwal, P. & Nichol, A. Diffusion Models Beat GANs on Image Synthesis. (arXiv,2021,6,1), http://arxiv.org/abs/2105.05233 [23] Mirza, M. & Osindero, S. Conditional Generative Adversarial Nets. (arXiv,2014,11,6), http://arxiv.org/abs/1411.1784