Paper deep dive
SALIENT: Frequency-Aware Paired Diffusion for Controllable Long-Tail CT Detection
Yifan Li, Mehrdad Salimitari, Taiyu Zhang, Guang Li, David Dreizin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/20/2026, 8:43:20 AM
Summary
The paper introduces SALIENT, a mask-conditioned wavelet-domain diffusion framework for controllable synthetic data augmentation in long-tail CT detection. SALIENT performs structured diffusion over discrete wavelet coefficients to separate low-frequency brightness from high-frequency structural detail, enabling interpretable attribute-level regulation. It utilizes a 3D VAE for volumetric lesion mask generation and a semi-supervised teacher for pseudo-labeling. The method improves generative realism (MS-SSIM 0.63 to 0.83, FID 118.4 to 46.5) and downstream detection performance (AUPRC gains) under extreme class imbalance, specifically for mediastinal hematoma detection.
Entities (10)
Relation Signals (8)
SALIENT â uses â Wavelet-Domain Diffusion
confidence 95% · SALIENT performs structured diffusion over discrete wavelet coefficients, explicitly separating low-frequency brightness from high-frequency structural detail.
SALIENT â generates â Synthetic CT-Mask Pairs
confidence 92% · SALIENT synthesizes paired lesion-masking volumes for controllable CT augmentation under long-tail regimes.
3D VAE â generates â Volumetric Lesion Masks
confidence 90% · A 3D VAE generates diverse volumetric lesion masks... for downstream mask-guided detection.
SALIENT â improves â Long-Tail Detection Performance
confidence 90% · SALIENT-augmented training improves long-tail detection performance, yielding disproportionate AUPRC gains across low prevalences and target-to-volume ratios.
SALIENT â improves â Generative Realism
confidence 90% · SALIENT improves generative realism, as reflected by higher MS-SSIM (0.63 to 0.83) and lower FID (118.4 to 46.5).
ResNet-50 â isusedfor â Slice-Level Classification
confidence 88% · The resulting synthetic CTâmask pairs augment training for a slice-level mask-guided ResNet-50 classifier.
EViT â aggregates â Slice-Level Predictions
confidence 85% · Slice-level predictions are further aggregated into subject-level decisions using an Embedded Vision Transformer (EViT).
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Detection of rare lesions in whole-body CT is fundamentally limited by extreme class imbalance and low target-to-volume ratios, producing precision collapse despite high AUROC. Synthetic augmentation with diffusion models offers promise, yet pixel-space diffusion is computationally expensive, and existing mask-conditioned approaches lack controllable attribute-level regulation and paired supervision for accountable training. We introduce SALIENT, a mask-conditioned wavelet-domain diffusion framework that synthesizes paired lesion-masking volumes for controllable CT augmentation under long-tail regimes. Instead of denoising in pixel space, SALIENT performs structured diffusion over discrete wavelet coefficients, explicitly separating low-frequency brightness from high-frequency structural detail. Learnable frequency-aware objectives disentangle target and background attributes (structure, contrast, edge fidelity), enabling interpretable and stable optimization. A 3D VAE generates diverse volumetric lesion masks, and a semi-supervised teacher produces paired slice-level pseudo-labels for downstream mask-guided detection. SALIENT improves generative realism, as reflected by higher MS-SSIM (0.63 to 0.83) and lower FID (118.4 to 46.5). In a separate downstream evaluation, SALIENT-augmented training improves long-tail detection performance, yielding disproportionate AUPRC gains across low prevalences and target-to-volume ratios. Optimal synthetic ratios shift from 2x to 4x as labeled seed size decreases, indicating a seed-dependent augmentation regime under low-label conditions. SALIENT demonstrates that frequency-aware diffusion enables controllable, computationally efficient precision rescue in long-tail CT detection.
Tags
Links
- Source: https://arxiv.org/abs/2602.23447v1
- Canonical: https://arxiv.org/abs/2602.23447v1
Trouble viewing inline? Open PDF directly â
Full Text
42,192 characters extracted from source content.
Expand or collapse full text
SALIENT: Frequency-Aware Paired Diffusion for Controllable Long-Tail CT Detection â Yifan Li 1,2 , Mehrdad Salimitari 1,2 , Taiyu Zhang 1,2 , Guang Li 2 , and David Dreizin 1,2 1 1 Trauma Radiology AI Laboratory (TRAIL), Department of Radiology and Nuclear Medicine, University of Maryland School of Medicine 2 2 Department of Radiology and Nuclear Medicine, University of Maryland School of Medicine Abstract. Detection of rare lesions in whole-body CT is fundamentally limited by extreme class imbalance and low target-to-volume ratios, pro- ducing precision collapse despite high AUROC. Synthetic augmentation with diffusion models offers promise, yet pixel-space diffusion is compu- tationally expensive, and existing mask-conditioned approaches lack con- trollable attribute-level regulation and paired supervision for accountable training. We introduce SALIENT, a mask-conditioned wavelet-domain diffusion framework that synthesizes paired lesionâmask volumes for con- trollable CT augmentation under long-tail regimes. Instead of denois- ing in pixel space, SALIENT performs structured diffusion over dis- crete wavelet coefficients, explicitly separating low-frequency brightness from high-frequency structural detail. Learnable frequency-aware ob- jectives disentangle target and background attributes (structure, con- trast, edge fidelity), enabling interpretable and stable optimization. A 3D VAE generates diverse volumetric lesion masks, and a semi-supervised teacher produces paired slice-level pseudo-labels for downstream mask- guided detection. SALIENT improves generative realism, as reflected by higher MS-SSIM (0.63â0.83) and lower FID (118.4â46.5).In a sep- arate downstream evaluation, SALIENT-augmented training improves long-tail detection performance, yielding disproportionate AUPRC gains across low prevalences and target-to-volume ratios. Optimal synthetic ra- tios shift from 2Ă to 4Ă as labeled seed size decreases, indicating a seed- dependent augmentation regime under low-label conditions. SALIENT demonstrates that frequency-aware diffusion enables controllable, com- putationally efficient precision rescue in long-tail CT detection. Keywords: Diffusion Models· Wavelet-Domain Generation· Long-Tail Detection· Synthetic Data Augmentation· Medical Imaging â Code and trained models are planned for public release under terms permitting non- commercial academic research use. Final licensing details will be specified at the time of release. arXiv:2602.23447v1 [eess.IV] 26 Feb 2026 2Y. Li et al. 1 Introduction Whole-body CT (WBCT) is widely used for cancer staging, inflammatory dis- ease evaluation, and polytrauma assessment [9, 10, 13, 28, 32]. Despite advances in deep learning, detection of uncommon or small lesions in WBCT remains fun- damentally difficult. Two compounding failure modes drive this challenge. First, within-patient signal dilution arises from low target-to-volume ratios (TVRs) in large torso fields of view [8, 21]. Second, cross-dataset prevalence dilution pro- duces extreme class imbalance in long-tail detection settings [6, 7]. Together, these effects create a precision ceiling that cannot be resolved solely through architectural modification [4,24,30]. Segmentation networks such as nnU-Net [14] and hybrid CNNâTransformer architectures [29,33] have achieved strong voxel-level performance. However, le- sion detection under severe imbalance remains sensitive to background domi- nance. Even when AUROC appears high, models frequently suffer from poor precision, low AUPRC, and unstable F1 scores [11,30]. In deployment, low pre- cision leads to spurious saliency and increased false alarms, limiting clinical trust and usability [2, 17]. Attention mechanisms and mask-guided training can im- prove feature localization [1, 22, 34, 35], yet they do not address the underlying scarcity of informative positive samples. Synthetic data augmentation has long been proposed to mitigate data spar- sity [20]. Diffusion probabilistic models (DDPMs) have recently surpassed GANs in image quality metrics for high-dimensional medical imaging [18,23]. However, existing mask-conditioned diffusion approaches often rely on fixed input geome- tries or limited sources of structural diversity [5, 12, 36]. Pixel-space DDPMs are computationally prohibitive in 3D and often require aggressive downsam- pling, degrading fine-grained in-plane detail critical for small lesion modeling. Frequency-domain diffusion methods offer computational speedups [25], but rely on manually tuned band weights and do not disentangle interpretable image at- tributes such as brightness, structure, and detail. Moreover, augmentation is typically assumed to provide monotonic bene- fit [3,19,27], yet augmentation dose-response behavior remains uncharacterized. We are not aware of prior work that defines optimal ("therapeutic dose") syn- thetic augmentation levels or identifies performance degradation under excessive synthetic sampling (âtoxic dosesâ) in long-tail detection [16]. Consequently, syn- thetic augmentation remains heuristic rather than prescriptive. In this work, we introduce SALIENT (Structured Attention-Leveraged In- ference for Edge-aware Neural Training), a mask-conditioned wavelet-domain diffusion framework for controllable CT augmentation. SALIENT replaces pixel- space denoising with structured diffusion over discrete wavelet coefficients, ex- plicitly separating global brightness from high-frequency detail. We introduce learnable frequency-domain weighting that disentangles target and background attributes into interpretable optimization âdialsâ governing structure, detail, con- trast, brightness, and spatial context. This design yields substantial computa- tional speedup while preserving high-resolution boundaries essential for small lesions. Title Suppressed Due to Excessive Length3 Fig. 1: Overview of the proposed SALIENT synthetic data and classification pipeline. Real CT volumes and lesion masks are first processed by a 3D VAE to generate diverse volumetric masks, which are projected into 2D slice space and used as conditioning signals for the wavelet-domain diffusion model. SALIENT operates on discrete wavelet coefficients to synthesize mask-guided CT slices, which are subsequently pseudo-labeled by a semi-supervised segmentation teacher (UCMT). The resulting synthetic CTâmask pairs augment training for a slice-level mask-guided ResNet-50 classifier. Slice-level pre- dictions are further aggregated into subject-level decisions using an Embedded Vision Transformer (EViT) [15]. Crucially, SALIENT generates anatomically coherent lesionâmask pairs sam- pled from a learned latent pathology manifold, enabling direct training of mask- guided detectors. We further characterize the augmentation dose-response un- der varying prevalence and seed sizes, revealing a stable therapeutic regime and a rightward dose shift under low-label conditions. These findings suggest that frequency-aware diffusion enables controllable precision rescue in long-tail detec- tion settings. Our contributions are: â A mask-conditioned wavelet-domain diffusion framework with learnable fre- quency weighting for attribute-specific control. â Paired synthetic lesionâmask generation enabling accountable mask-guided detection training. â Empirical characterization of augmentation dose-response behavior under varying prevalence and labeled seed sizes. 4Y. Li et al. â A practical foundation for archetype-guided augmentation scaling in long- tail CT detection. 2 Related Work 2.1 Long-Tail Detection in Medical Imaging Detection of rare or small targets in cross-sectional imaging is challenged by extreme class imbalance and low target-to-volume ratios [8,21,30]. While voxel- level segmentation networks such as nnU-Net [14] and modern CNNâTransformer hybrids [29, 33] achieve strong segmentation performance, detection precision under imbalance remains limited. Attention-based and mask-guided paradigms improve feature localization and accountability [1,22,34,35], but do not funda- mentally address the scarcity of informative positive samples. 2.2 Synthetic Augmentation with Diffusion Models Generative augmentation has been widely explored to address data sparsity [20]. GAN-based approaches were historically dominant but suffered from training instability and mode collapse. Diffusion probabilistic models (DDPMs) have demonstrated superior perceptual realism and stability in medical imaging [18, 23]. Mask-conditioned diffusion variants allow for structural control [5, 12, 36], yet most approaches bind lesion geometry to conditioning masks or preserve the surrounding background, which limits morphological diversity. Pixel-space diffusion in 3D is computationally expensive and often requires reduced spatial resolution. Latent-space diffusion reduces cost but may obscure fine-grained structural detail. Recent frequency-domain diffusion methods intro- duce spectral acceleration [25], yet typically rely on manually tuned band weights without interpretable control over image attributes. In contrast, SALIENT integrates mask conditioning with wavelet-domain diffusion and learnable attribute-specific frequency weighting. This enables ex- plicit disentanglement of target and background structure, detail, contrast, and brightness, providing interpretable and task-aligned generative control. 2.3 Augmentation Dose-Response Most augmentation studies assume monotonic performance gains with increasing synthetic data [3,19,27]. However, oversampling is known to risk overfitting and degraded generalization in imbalanced learning [16]. Systematic characterization of augmentation dose-response curves, therapeutic regimes, or toxicity effects remains underexplored. No existing framework provides prescriptive guidance for optimal augmentation under varying prevalence or labeled seed sizes. We address this gap by empirically quantifying augmentation dose-response behavior under controlled prevalence and seed conditions, demonstrating a re- producible therapeutic regime and a seed-dependent dose shift. Title Suppressed Due to Excessive Length5 3 Proposed Method We propose an end-to-end pipeline for controllable CT augmentation and long- tail detection (Fig. 1). SALIENT generates paired synthetic CTâmask samples via wavelet-domain diffusion conditioned on VAE [26]-sampled lesion masks, and performance is evaluated using mask-guided slice-level classification with subject-level aggregation. 3.1 Dataset The dataset comprises contrast-enhanced whole-body CT examinations from an adult trauma cohort. Mediastinal hematoma occurs in approximately 3% of stud- ies, yielding a naturally long-tail detection problem characterized by low target- to-volume ratios and irregular, multiscale morphology. Mediastinal hematomas are potentially clinically significant, making this a practical stress-test for preci- sion under signal dilution. We evaluate on an internal contrast-enhanced torso CT cohort comprising 5,205 subjects (200 mediastinal hematomaâpositive, 5,005 negative controls) following IRB approval. Positive subjects were confirmed by board-certified radiologists using clinical reports and dedicated image review. All scans underwent unified preprocessing, including reorientation to a consis- tent coordinate convention, resampling to common voxel spacing, soft-tissue HU windowing, intensity normalization, and anatomical cropping to the mediastinal region of interest. All experiments are conducted on the mediastinal hematoma task under controlled prevalence shifts to simulate long-tail detection regimes. 3.2 SALIENT Overview. As illustrated in Fig. 2, SALIENT is a wavelet-domain conditional diffusion model designed to generate anatomically coherent and mask-consistent CT slices. Rather than operating directly in pixel space, the model performs diffusion over discrete wavelet coefficients, enabling explicit control over low- frequency structure and high-frequency detail. Mask and 2.5D anatomical con- text are incorporated both architecturally and through structured guidance dur- ing sampling. Wavelet-Domain Formulation. Let X â R ZĂHĂW denote a preprocessed CT volume and m z â 0, 1 HĂW the lesion mask at slice z. For each central slice x z â R HĂW , we define a 2.5D neighborhood N (z) =x z+â : ââS, where S contains small axial offsets capturing through-plane continuity. We apply a single-level Haar discrete wavelet transform: w z =W(x z ) = [L z ,LH z ,HL z ,H z ]â R 4Ă H 2 Ă W 2 ,(1) 6Y. Li et al. Fig. 2: Architecture of SALIENT in the wavelet domain. A central CT slice and its ax- ial neighbors are transformed into wavelet coefficients (L, LH, HL, H). A mask-gated frequency scaling (FSA) module modulates the noisy coefficients before concatenation with a 2.5D conditioning stack. A time-conditioned UNet predicts clean wavelet co- efficients at each diffusion step, which are reconstructed into synthetic CT slices via inverse DWT. where L encodes low-frequency structure and LH, HL, H capture oriented high-frequency components. We learn a conditional diffusion model p Ξ (w z | c z ),(2) where c z encodes the down-sampled central mask and selected neighbor wavelet bands. The forward process follows q(w t | w 0 ) =N w t ; â Ìα t w 0 , (1â Ìα t )I ,(3) and the reverse model predicts clean coefficients: Ëw 0 = f Ξ (w t ,t,c z ).(4) Synthetic CT slices are reconstructed via x syn z =W â1 ( Ëw 0 ).(5) Mask-Guided Wavelet UNet. At timestep t, noisy coefficients w t â R 4Ă H 2 Ă W 2 are first modulated by a mask-gated frequency scaling (FSA) module to produce Ìw t . The conditioning tensor c z â R C cond Ă H 2 Ă W 2 Title Suppressed Due to Excessive Length7 includes the down-sampled mask and selected neighbor wavelet bands. The UNet input is u t = [ Ìw t â„c z ].(6) The backbone is a four-level encoderâdecoder UNet with residual blocks, symmetric skip connections, and self-attention at the deepest resolutions. Dif- fusion timesteps are encoded via sinusoidal embeddings and injected through FiLM-style scaleâshift conditioning. Objective and Modeling Principle SALIENT is a mask-conditioned wavelet- domain diffusion model designed to generate anatomically coherent CT slices and paired lesion masks under extreme class imbalance. Let x z â R HĂW denote a central CT slice and m z â 0, 1 HĂW its lesion mask. Instead of modeling pixel intensities directly, we operate in the discrete wavelet domain: w z =W(x z ) = [L z ,LH z ,HL z ,H z ]â R 4Ă H 2 Ă W 2 ,(7) where L captures global structure and brightness, and LH/HL/H encode oriented high-frequency detail. Our goal is to learn a conditional generative model: p Ξ (w z | c z ),(8) where c z includes the down-sampled lesion mask and selected neighboring wavelet features (2.5D context). The forward diffusion process follows: q(w t | w 0 ) =N ( â Ìα t w 0 , (1â Ìα t )I),(9) and the reverse model predicts clean coefficients: Ëw 0 = f Ξ (w t ,t,c z ).(10) Synthetic slices are reconstructed via inverse wavelet transform: x syn z =W â1 ( Ëw 0 ).(11) Operating in wavelet space provides two advantages: (i) explicit separation of global brightness (L) from high-frequency boundaries, (i) structured control over targetâbackground detail without full 3D pixel-space diffusion. Wavelet-Aware Training Objective Unlike standard diffusion trained with uniform â 2 loss, SALIENT optimizes a frequency-structured objective that dis- entangles global structure from boundary detail. We define a band-weighted reconstruction: L wavelet = E [â„W â ( Ëw 0 â w 0 )â„ 1 ],(12) 8Y. Li et al. where W applies higher weights near lesion boundaries and moderates diagonal H amplification. To stabilize brightness and prevent drift in low-prevalence regimes, we intro- duce low-frequency moment regularization: L L = λ ÎŒ â„ÎŒ pred L â ÎŒ tgt L â„ 2 2 + λ Ï â„ logÏ pred L â logÏ tgt L â„ 2 2 .(13) High-frequency variance control ensures texture fidelity without noise ampli- fication: L HF = X bâLH,HL,H λ b â„ logÏ pred b â logÏ tgt b â„ 2 2 .(14) Finally, mild pixel-space auxiliary constraints encourage edge alignment and prevent intensity saturation in lesion and peri-lesional regions. The overall objective is: L total =L wavelet +L L +L HF +L aux .(15) This formulation enables interpretable control over structure, detail, and brightness while maintaining computational efficiency. Structured Classifier-Free Guidance To disentangle lesion conditioning from anatomical context, we adopt structured classifier-free guidance. We com- pute three forward passes: (i) unconditional, (i) mask-only, and (i) mask+neighbor conditioned predictions. The final estimate is: Ëw SALIENT 0 = f Ξ (x t ,t, â )(16) + s mask [f Ξ (x t ,t,c mask )â f Ξ (x t ,t, â )](17) + s nei (t) [f Ξ (x t ,t,c mask+nei )â f Ξ (x t ,t,c mask )].(18) The neighbor guidance scale s nei (t) decays over time, encouraging global anatomical coherence early in diffusion and lesion-focused refinement in later steps. This structured guidance allows SALIENT to sample morphologically diverse lesions while preserving anatomical plausibility. 3.3 3D VAE for Volumetric Lesion Mask Generation To introduce mask-guided generative diversity beyond the limited observed pos- itive set, we train a 3D variational autoencoder (MaskVAE3D) on volumetric lesion masks. Contiguous axial mask slices are stacked and resampled to a fixed size, forming spatially coherent binary volumes. The encoder maps each volume to a low-dimensional latent representation, and the decoder reconstructs anatomically plausible lesion masks. Reconstruc- tion combines Dice and boundary-weighted losses to preserve topology and sur- face fidelity. A KL regularization term with free-bits stabilization prevents latent collapse. Title Suppressed Due to Excessive Length9 At inference, latent codes are sampled from a Gaussian prior and decoded into volumetric masks, which are sliced to provide diverse conditioning inputs for SALIENT. 3.4 Semi-Supervised Segmentation for Paired Masks To obtain slice-aligned lesion masks for synthetic CT images, we apply a semi- supervised segmentation model based on Uncertainty-aware Cross-Model Train- ing (UCMT) [31]. UCMT uses two student networks with an exponential moving average (EMA) teacher, combining supervised Dice-style loss on labeled real slices with cross- pseudo supervision and uncertainty-guided mixing on unlabeled data. After training, the frozen EMA teacher is applied to synthetic slices to pro- duce binary lesion masks. These pseudo-labels provide geometrically consistent CTâmask pairs for downstream mask-guided classification. 3.5 Mask-Guided Detection Evaluation We evaluate whether SALIENT-generated paired CTâmask samples improve downstream detection under severe class imbalance. We train a mask-guided slice-level classifier and aggregate slice evidence for subject-level prediction. We adopt a ResNet-50 backbone augmented with two lightweight mask- guided attention (MGA) blocks inserted at intermediate and deep feature stages, following prior mask-guided attention paradigms [22, 35]. Given an input slice (or 2.5D triplet), each MGA block produces a spatial attention map encouraged to align with the lesion mask during training. Classification is optimized using focal loss to address class imbalance. In ad- dition, we apply an attention-alignment loss (mean-squared error) between nor- malized attention maps and resized lesion masks to enforce spatial accountability. The overall objective combines focal and attention-alignment losses, weighted by a balancing coefficient. For patient-level prediction, we extract slice-level embeddings from the trained backbone and aggregate them using a lightweight Transformer encoder with po- sitional embeddings. Following the Embedded Vision Transformer (EViT) [15] paradigm, token aggregation emphasizes informative slices before producing a subject-level probability. The same aggregation scheme is used across all abla- tions to isolate the effect of synthetic paired augmentation. 4 Experiments 4.1 Experimental Setting In this study, SALIENT was trained using AdamW with cosine learning-rate de- cay and exponential moving average (EMA) stabilization. Diffusion employed a 10Y. Li et al. cosine noise schedule and wavelet-domain band weighting to balance global struc- ture and high-frequency detail. Low-frequency stabilization, high-frequency vari- ance control, and mild pixel-space auxiliary constraints were applied as described in Sec. 3. Full hyperparameter configurations are provided in the Supplement. MaskVAE3D was trained with AdamW and EMA stabilization using boundary- aware reconstruction losses and KL regularization with free-bits stabilization. Training details and loss weights are provided in the Supplement. UCMT em- ployed DeepLabv3+ backbones with supervised Dice-style loss and cross-pseudo supervision under uncertainty-aware mixing. Models were trained using AdamW with cosine decay. Full training parameters are included in the Supplement. The mask-guided ResNet50 backbone was trained with focal loss and attention- alignment supervision. To address imbalance, all positive slices were included per epoch with balanced negative sampling. Optimization used AdamW with cosine decay. Detailed settings are provided in the Supplement. For subject-level predic- tion, we extracted per-slice 2048-D features from the frozen slice-level backbone. Sequences were ordered by slice index and padded/truncated to length 500. A lightweight Transformer encoder (4 layers, 8 heads, embedding dimension 512) processed slice tokens with learned positional encodings. Training used AdamW (learning rate 10 â4 ), cosine decay for 200 epochs, and balanced subject-level sampling. 4.2 Generation Results We compare SALIENT against a pixel-space MedDDPM [5] baseline trained with the same VAE mask generation and semi-supervised segmentation pipeline. The baseline operates directly in pixel space using 2D or 2.5D denoising, while SALIENT performs diffusion over discrete wavelet coefficients. Qualitative Comparison. Figure 3 shows representative synthetic CT slices generated by MedDDPM and SALIENT alongside corresponding masks. Med- DDPM outputs exhibit visible high-frequency noise, localized brightness am- plification near the mediastinum, and occasional structural oversmoothing. In contrast, SALIENT produces sharper vascular boundaries, improved soft-tissue contrast, and more anatomically coherent lungâmediastinum interfaces. Lesion regions remain consistent with conditioning masks without introducing periph- eral artifacts. These qualitative differences are consistent across subjects and slice positions, particularly in challenging superior mediastinal regions where small hematomas occupy a limited fraction of the field of view. A board-certified radiologist per- formed blinded grading of synthetic CT slices containing mediastinal hematoma from 20 randomly selected cases using a 5-point Likert scale (1=poor, 5=near- indistinguishable from real). Graded criteria included global structural realism, lesion plausibility, lesionâbackground integration, high-frequency artifact pres- ence (reverse scored), brightness and contrast realism, and mask fidelity (def- initions in Appendix). SALIENT demonstrated higher brightness and contrast Title Suppressed Due to Excessive Length11 Fig. 3: Comparison of synthetic CT slices generated by pixel-space MedDDPM and wavelet-domain SALIENT. SALIENT produces sharper anatomical boundaries, re- duced high-frequency noise, and improved contrast stability in mediastinal regions. realism, improved lesionâbackground integration, fewer high-frequency artifacts, and markedly superior mask fidelity compared to pixel-space MedDDPM, with modest differences in global structural realism. These qualitative trends were consistent with quantitative segmentation fidelity (Dice: 0.72±0.24 vs 0.27±0.16) and aligned with the effects of SALIENT augmentation on precision in down- stream detection dose-response experiments. Objective Realism Metrics. We quantify synthesis quality using multi-scale structural similarity (MS-SSIM) and FrĂ©chet Inception Distance (FID). SALIENT improves MS-SSIM from 0.63 to 0.83 and reduces FID from 118.4 to 46.5: MethodMS-SSIM â FID â MedDDPM (pixel-space)0.63118.4 SALIENT (wavelet-domain) 0.8346.5 The increase in MS-SSIM indicates improved structural fidelity across scales, while the substantial FID reduction reflects better alignment with real CT fea- ture distributions. Wavelet-Band Analysis. We analyzed energy distributions across the four wavelet bands (L, LH, HL, H). Pixel-space MedDDPM exhibits unstable low- frequency brightness shifts and excessive high-frequency variance. In contrast, SALIENTâs band-wise diffusion with explicit L/HF regularization produces en- ergy profiles that more closely match real mediastinal hematoma data. As shown in Fig. 4, the L-band standard deviation indicates that MedDDPM underesti- mates global contrast and displays greater variability, consistent with brightness drift. SALIENT restores L variability toward the real CT distribution, improv- ing low-frequency stability. For high-frequency bands (LH, HL, H), MedDDPM shows directional imbalance and variance distortion, indicating noisy or overly smoothed textures. SALIENT aligns more closely with real variance across detail bands, preserving anisotropic edge structure without amplifying H artifacts. 12Y. Li et al. Fig. 4: Quantitative frequency and intensity comparison between Real CT, pixel-space MedDDPM, and SALIENT. (Left) Per-slice L standard deviation. (Middle) High- frequency variance per slice (LH/HL/H). (Right) ROI intensity histograms using method-specific masks. ROI intensity histograms further demonstrate that MedDDPM shifts lesion in- tensities toward higher values, with broader tails, whereas SALIENT maintains contrast distributions closer to those of real CT while avoiding saturation. Computational Efficiency. Wavelet-domain modeling also improves efficiency. SALIENT achieves approximately 4Ă faster training relative to 2.5D MedDDPM, and 28Ă speedup compared to full 3D MedDDPM (single NVIDIA H100), while preserving 512Ă 512 in-plane resolution. This enables practical high-resolution synthesis without the computational burden of volumetric diffusion. Overall, these results indicate that wavelet-domain diffusion with structured mask guidance yields sharper, more stable, and more computationally efficient synthesis than pixel-space MedDDPM. 4.3 Detection Results We evaluate whether SALIENT-generated paired CTâmask samples provide functional benefit in downstream detection, beyond perceptual realism. Using the mask-guided ResNet50 backbone, we report subject-level performance un- der severe class imbalance by controlling test prevalence (1â5%) and varying the synthetic-to-real augmentation ratio in training. Performance is summarized using AUPRC (primary under imbalance) and AUROC. Dose-response under synthetic augmentation. Table 1 reports the dose- response of SALIENT augmentation (ratios 0Ă to 10Ă). With a labeled seed of n=50 positive studies, SALIENT exhibits a consistent therapeutic regime at 2Ă synthetic augmentation across all prevalences, improving AUPRC byâ0.05â 0.06 absolute. When the labeled seed is reduced to n=25, the therapeutic dose shifts rightward to 4Ă, with substantially larger AUPRC gains (up to â0.12 at 1% prevalence). Notably, AUROC remains high throughout, indicating that the primary effect of SALIENT augmentation is precision rescue (AUPRC) rather than trivially inflating separability. Title Suppressed Due to Excessive Length13 SeedPrev.AUPRC (0Ă)Best AUPRCOpt. doseâAUPRC n=50 1%0.88090.94142Ă+0.0605 2%0.90620.96772Ă+0.0615 3%0.91100.96842Ă+0.0574 4%0.92150.97452Ă+0.0530 5%0.92230.97452Ă+0.0522 n=25 1%0.81690.94084Ă+0.1239 2%0.87760.96984Ă+0.0922 3%0.88030.97364Ă+0.0933 4%0.88310.98264Ă+0.0995 5%0.88430.98264Ă+0.0983 Table 1: Subject-level dose-response of SALIENT augmentation. âOpt. doseâ denotes the synthetic-to-real ratio achieving the best AUPRC for each prevalence. SALIENT yields a stable therapeutic dose of 2Ă at n=50 and a right-shift to 4Ă at n=25, with larger gains in the low-label regime. To isolate the contribution of paired mask-conditioning, we compare against training without mask guidance (same real/synthetic sampling protocol). At 1% prevalence (n = 50), the baseline model without augmentation achieves an AUPRC of 0.8809. Synthetic augmentation without mask guidance fails to improve performance (best â0.8408 across ratios). In contrast, mask-guided SALIENT augmentation increases AUPRC to 0.9414 at 2Ă. These results sug- gest that performance improvements stem from paired imageâmask supervision rather than synthetic image quantity alone. We further stratified performance by target-to-volume ratio (TVR) at lo- cal (4%) prevalence to assess robustness under within-patient signal dilution. SALIENT augmentation produced the largest AUPRC gains in the small-TVR regime (+0.1103), followed by large-TVR (+0.0832) and middle-TVR (+0.0764). In contrast, AUROC gains were modest (â0.0005, +0.0028, +0.0129 for small, middle, and large TVR, respectively), indicating that SALIENT primarily im- proves precision rather than ranking separability. These findings suggest that paired wavelet-domain augmentation is particularly effective when the lesion signal is diluted by low target-to-volume ratios. Qualitative evidence via saliency alignment. Figure 5 visualizes saliency maps under three training conditions: (i) without mask guidance, (i) without synthetic data, and (i) with SALIENT paired augmentation. Without masks or synthetic pairing, saliency frequently concentrates on irrelevant structures (e.g., body wall), suggesting shortcut learning. In contrast, SALIENT paired augmen- tation increases overlap between saliency and the lesion region, supporting the claim that wavelet-domain diffusion enables clinically meaningful augmentation that improves where the detector attends, not only its aggregate metrics. Across prevalences (1â5%) and labeled seed sizes (n=25, 50), SALIENT ex- hibits a reproducible augmentation dose-response with a clear therapeutic regime. The systematic right-shift in optimal augmentation under reduced label availabil- 14Y. Li et al. Fig. 5: Saliency alignment improves with paired augmentation. Top: training without mask guidance; middle: without synthetic data; bottom: with SALIENT paired CTâmask augmentation. SALIENT encourages lesion-focused evidence and reduces spurious activations on irrelevant anatomy. ity suggests that higher synthetic ratios may be beneficial in low-label regimes, provided mask guidance preserves spatial accountability. 5 Conclusion We presented SALIENT, a mask-conditioned wavelet-domain diffusion frame- work for controllable CT augmentation under extreme class imbalance. By re- placing pixel-space denoising with band-wise diffusion over discrete wavelet coef- ficients (L/LH/HL/H), SALIENT explicitly separates global brightness struc- ture from high-frequency edge and texture components. This structured fre- quency decomposition enables stable optimization with targetâbackground reg- ularization, improved artifact control, and substantially reduced computational cost compared to volumetric diffusion. Across mediastinal hematoma detection, SALIENT demonstrated consistent improvements in both perceptual realism and functional utility. Wavelet-domain modeling increased MS-SSIM and reduced FID while preserving high in-plane resolution with 4Ă faster training relative to 2.5D pixel-space diffusion. More im- portantly, paired CTâmask synthesis translated into reproducible downstream precision gains under severe class imbalance. We observed a stable augmenta- tion dose-response: with sufficient labeled data (n = 50), the therapeutic regime occurs at 2Ă synthetic augmentation, while in low-label settings (n = 25), the optimal dose shifts to 4Ă. Mask guidance proved essential for accountable gains, Title Suppressed Due to Excessive Length15 improving both saliency alignment and AUPRC relative to unpaired synthetic augmentation. Notably, gains were strongest in low prevalence and target-to- volume ratio regimes, highlighting SALIENTâs ability to counteract prevalence- and within-patient signal dilution in precision-sensitive settings through com- bined lesion-mask augmentation. These findings suggest that frequency-aware diffusion provides a practical mechanism for controllable precision rescue in long-tail detection problems. By enabling explicit regulation of brightness and high-frequency structure, SALIENT transforms synthetic data from a heuristic augmentation strategy into a tunable component of the training pipeline. Future work will explore extensions to addi- tional long-tail CT detection tasks involving irregular, spatially heterogeneous, and multiscale lesion morphology. References 1. Arrieta, A.B., DĂaz-RodrĂguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., GarcĂa, S., Gil-LĂłpez, S., Molina, D., Benjamins, R., et al.: Explainable artifi- cial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information fusion 58, 82â115 (2020) 2. Bernstein, M.H., Atalay, M.K., Dibble, E.H., Maxwell, A.W., Karam, A.R., Agar- wal, S., Ward, R.C., Healey, T.T., Baird, G.L.: Can incorrect artificial intelligence (ai) results impact radiologists, and if so, what can we do about it? a multi-reader pilot study of lung cancer detection with chest radiography. European radiology 33(11), 8263â8269 (2023) 3. Cha, K.H., Petrick, N., Pezeshk, A., Graff, C.G., Sharma, D., Badal, A., Sahiner, B.: Evaluation of data augmentation via synthetic images for improved breast mass detection on mammograms using deep learning. Journal of Medical Imaging 7(1), 012703â012703 (2020) 4. Daneshjou, R., Vodrahalli, K., Novoa, R.A., Jenkins, M., Liang, W., Rotemberg, V., Ko, J., Swetter, S.M., Bailey, E.E., Gevaert, O., et al.: Disparities in derma- tology ai performance on a diverse, curated clinical image set. Science advances 8(31), eabq6147 (2022) 5. Dorjsembe, Z., Pao, H.K., Odonchimed, S., Xiao, F.: Conditional diffusion mod- els for semantic 3d brain mri synthesis. IEEE Journal of Biomedical and Health Informatics 28(7), 4084â4093 (2024) 6. Dreizin, D., Munera, F.: Blunt polytrauma: evaluation with 64-section whole-body ct angiography. Radiographics 32(3), 609â631 (2012) 7. Dreizin, D., Munera, F.: Multidetector ct for penetrating torso trauma: state of the art. Radiology 277(2), 338â355 (2015) 8. Dreizin, D., Staziaki, P.V., Khatri, G.D., Beckmann, N.M., Feng, Z., Liang, Y., Delproposto, Z.S., Klug, M., Spann, J.S., Sarkar, N., et al.: Artificial intelligence cad tools in trauma imaging: a scoping review from the american society of emer- gency radiology (aser) ai/ml expert panel. Emergency radiology 30(3), 251â265 (2023) 9. Dreizin, D., Zhou, Y., Chen, T., Li, G., Yuille, A.L., McLenithan, A., Morrison, J.J.: Deep learning-based quantitative visualization and measurement of extraperitoneal hematoma volumes in patients with pelvic fractures: potential role in personalized forecasting and decision support. Journal of Trauma and Acute Care Surgery 88(3), 425â433 (2020) 16Y. Li et al. 10. Dreizin, D., Zhou, Y., Fu, S., Wang, Y., Li, G., Champ, K., Siegel, E., Wang, Z., Chen, T., Yuille, A.L.: A multiscale deep learning method for quantitative visual- ization of traumatic hemoperitoneum at ct: assessment of feasibility and compari- son with subjective categorical estimation. Radiology: Artificial Intelligence 2(6), e190220 (2020) 11. Hasani, N., Farhadi, F., Morris, M.A., Nikpanah, M., Rhamim, A., Xu, Y., Pariser, A., Collins, M.T., Summers, R.M., Jones, E., et al.: Artificial intelligence in medical imaging and its impact on the rare disease community: threats, challenges and opportunities. PET clinics 17(1), 13 (2022) 12. Heo, C., Jung, J.: Controllable mask diffusion model for medical annotation syn- thesis with semantic information extraction. Computers in Biology and Medicine 196, 110807 (2025) 13. Huang, W., Liu, W., Zhang, X., Yin, X., Han, X., Li, C., Gao, Y., Shi, Y., Lu, L., Zhang, L., et al.: Lidia: Precise liver tumor diagnosis on multi-phase contrast- enhanced ct via iterative fusion and asymmetric contrastive learning. In: Interna- tional Conference on Medical Image Computing and Computer-Assisted Interven- tion. p. 394â404. Springer (2024) 14. Isensee, F., Jaeger, P.F., Kohl, S.A., Petersen, J., Maier-Hein, K.H.: nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods 18(2), 203â211 (2021) 15. Islam, N.U., Zhou, Z., Gehlot, S., Gotway, M.B., Liang, J.: Seeking an optimal approach for computer-aided diagnosis of pulmonary embolism. Medical image analysis 91, 102988 (2024) 16. Johnson, J.M., Khoshgoftaar, T.M.: Survey on deep learning with class imbalance. Journal of big data 6(1), 27 (2019) 17. Katal, S., York, B., Gholamrezanezhad, A.: Ai in radiology: From promise to practice- a guide to effective integration. European Journal of Radiology 181, 111798 (2024) 18. Khader, F., MĂŒller-Franzes, G., Tayebi Arasteh, S., Han, T., Haarburger, C., Schulze-Hagen, M., Schad, P., Engelhardt, S., BaeĂler, B., Foersch, S., et al.: De- noising diffusion probabilistic models for 3d medical image generation. Scientific reports 13(1), 7303 (2023) 19. Khosravi, B., Li, F., Dapamede, T., Rouzrokh, P., Gamble, C.U., Trivedi, H.M., Wyles, C.C., Sellergren, A.B., Purkayastha, S., Erickson, B.J., et al.: Syntheti- cally enhanced: unveiling synthetic dataâs potential in medical imaging research. EBioMedicine 104 (2024) 20. Langlotz, C.P., Allen, B., Erickson, B.J., Kalpathy-Cramer, J., Bigelow, K., Cook, T.S., Flanders, A.E., Lungren, M.P., Mendelson, D.S., Rudie, J.D., et al.: A roadmap for foundational research on artificial intelligence in medical imaging: from the 2018 nih/rsna/acr/the academy workshop. Radiology 291(3), 781â791 (2019) 21. Lee, S., Summers, R.M.: Clinical artificial intelligence applications in radiology: chest and abdomen. Radiologic Clinics 59(6), 987â1002 (2021) 22. Li, K., Wu, Z., Peng, K.C., Ernst, J., Fu, Y.: Tell me where to look: Guided attention inference network. In: Proceedings of the IEEE conference on computer vision and pattern recognition. p. 9215â9223 (2018) 23. MĂŒller-Franzes, G., Niehues, J.M., Khader, F., Arasteh, S.T., Haarburger, C., Kuhl, C., Wang, T., Han, T., Nolte, T., Nebelung, S., et al.: A multimodal compar- ison of latent denoising diffusion probabilistic models and generative adversarial networks for medical image synthesis. Scientific reports 13(1), 12098 (2023) Title Suppressed Due to Excessive Length17 24. Oakden-Rayner, L., Dunnmon, J., Carneiro, G., RĂ©, C.: Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. In: Proceedings of the ACM conference on health, inference, and learning. p. 151â159 (2020) 25. Phung, H., Dao, Q., Tran, A.: Wavelet diffusion models are fast and scalable image generators. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 10199â10208 (2023) 26. Pinheiro Cinelli, L., AraĂșjo Marins, M., Barros da Silva, E.A., Lima Netto, S.: Variational autoencoder. In: Variational Methods for Machine Learning with Applications to Deep Networks, p. 111â149. Springer, Cham (2021). https: //doi.org/10.1007/978-3-030-73531-9_4 27. Prakash, E., Valanarasu, J.M.J., Chen, Z., Reis, E.P., Johnston, A., Pareek, A., Bluethgen, C., Gatidis, S., Olsen, C., Chaudhari, A.S., et al.: Evaluating and im- proving the effectiveness of synthetic chest x-rays for medical image analysis. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 4413â4421 (2025) 28. Roth, H.R., Xu, Z., Tor-DĂez, C., Jacob, R.S., Zember, J., Molto, J., Li, W., Xu, S., Turkbey, B., Turkbey, E., et al.: Rapid artificial intelligence solutions in a pandemicâthe covid-19-20 lung ct lesion segmentation challenge. Medical image analysis 82, 102605 (2022) 29. Roy, S., Koehler, G., Ulrich, C., Baumgartner, M., Petersen, J., Isensee, F., Jaeger, P.F., Maier-Hein, K.H.: Mednext: transformer-driven scaling of convnets for med- ical image segmentation. In: International conference on medical image computing and computer-assisted intervention. p. 405â415. Springer (2023) 30. Salmi, M., Atif, D., Oliva, D., Abraham, A., Ventura, S.: Handling imbalanced med- ical datasets: review of a decade of research. Artificial intelligence review 57(10), 273 (2024) 31. Shen, Z., Cao, P., Yang, H., Liu, X., Yang, J., Zaiane, O.R.: Co-training with high- confidence pseudo labels for semi-supervised medical image segmentation. arXiv preprint arXiv:2301.04465 (2023), https://arxiv.org/abs/2301.04465 32. Vorontsov, E., Cerny, M., RĂ©gnier, P., Di Jorio, L., Pal, C.J., Lapointe, R., Vandenbroucke-Menu, F., Turcotte, S., Kadoury, S., Tang, A.: Deep learning for automated segmentation of liver lesions at ct in patients with colorectal cancer liver metastases. Radiology: Artificial Intelligence 1(2), 180014 (2019) 33. Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I.S., Xie, S.: Convnext v2: Co-designing and scaling convnets with masked autoencoders. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 16133â 16142 (2023) 34. Wu, S., Liu, S., Zhong, M., de Loos, E.R., Hartert, M., Fuentes-MartĂn, Ă., Lenzini, A., Wang, D., Qian, Q.: Development and validation of a self-attention network- based algorithm to detect mediastinal lesions on computed tomography images. Journal of Thoracic Disease 16(5), 3306â3316 (2024) 35. Yan, J., Zeng, Y., Lin, J., Pei, Z., Fan, J., Fang, C., Cai, Y.: Enhanced object detection in pediatric bronchoscopy images using yolo-based algorithms with cbam attention mechanism. Heliyon 10(12) (2024) 36. Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models. In: Proceedings of the IEEE/CVF international conference on computer vision. p. 3836â3847 (2023)