Paper deep dive
Cross-Scenario Deraining Adaptation with Unpaired Data: Superpixel Structural Priors and Multi-Stage Pseudo-Rain Synthesis
Kangbo Zhao, Miaoxin Guan, Xiang Chen, Yukai Shi, Jinshan Pan
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/26/2026, 2:32:02 AM
Summary
The paper introduces a cross-scenario deraining adaptation framework that enables deep learning models to generalize to unseen Out-of-Distribution (OOD) domains without requiring paired rainy/clean data in the target domain. The method utilizes a Superpixel Generation (Sup-Gen) module to extract structural priors, a Resolution-adaptive Fusion strategy to align these with target backgrounds, and a multi-stage pseudo-label re-synthesis mechanism to simulate realistic rain, achieving significant PSNR gains.
Entities (7)
Relation Signals (3)
Superpixel Generation (Sup-Gen) → uses → Simple Linear Iterative Clustering (SLIC)
confidence 99% · This module employs the Simple Linear Iterative Clustering (SLIC) algorithm.
Cross-scenario deraining adaptation framework → improves → Out-of-Distribution (OOD) generalization
confidence 95% · Extensive evaluations on state-of-the-art models demonstrate that our approach accelerates convergence and significantly improves out-of-distribution generalization
Cross-scenario deraining adaptation framework → integrates → Superpixel Generation (Sup-Gen)
confidence 95% · This framework functions as a versatile plug-and-play module capable of seamless integration into arbitrary deraining architectures.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Image deraining plays a pivotal role in low-level computer vision, serving as a prerequisite for robust outdoor surveillance and autonomous driving systems. While deep learning paradigms have achieved remarkable success in firmly aligned settings, they often suffer from severe performance degradation when generalized to unseen Out-of-Distribution (OOD) scenarios. This failure stems primarily from the significant domain discrepancy between synthetic training datasets and the complex physical dynamics of real-world rain. To address these challenges, this paper proposes a pioneering cross-scenario deraining adaptation framework. Diverging from conventional approaches, our method obviates the requirements for paired rainy observations in the target domain, leveraging exclusively rain-free background images. We design a Superpixel Generation (Sup-Gen) module to extract stable structural priors from the source domain using Simple Linear Iterative Clustering. Subsequently, a Resolution-adaptive Fusion strategy is introduced to align these source structures with target backgrounds through texture similarity, ensuring the synthesis of diverse and realistic pseudo-data. Finally, we implement a pseudo-label re-Synthesize mechanism that employs multi-stage noise generation to simulate realistic rain streaks. This framework functions as a versatile plug-and-play module capable of seamless integration into arbitrary deraining architectures. Extensive experiments on state-of-the-art models demonstrate that our approach yields remarkable PSNR gains of up to 32% to 59% in OOD domains while significantly accelerating training convergence.
Tags
Links
- Source: https://arxiv.org/abs/2603.21661v1
- Canonical: https://arxiv.org/abs/2603.21661v1
Trouble viewing inline? Open PDF directly →
Full Text
55,274 characters extracted from source content.
Expand or collapse full text
1 Cross-Scenario Deraining Adaptation with Unpaired Data: Superpixel Structural Priors and Multi-Stage Pseudo-Rain Synthesis Kangbo Zhao, Miaoxin Guan, Xiang Chen, Yukai Shi, Jinshan Pan Abstract—Image deraining plays a pivotal role in low-level computer vision, serving as a prerequisite for robust outdoor surveillance and autonomous driving systems. While deep learn- ing paradigms have achieved remarkable success in firmly aligned settings, they often suffer from severe performance degradation when generalized to unseen Out-of-Distribution (OOD) scenar- ios. This failure stems primarily from the significant domain discrepancy between synthetic training datasets and the complex physical dynamics of real-world rain. To address these challenges, this paper proposes a pioneering cross-scenario deraining adap- tation framework. Diverging from conventional approaches, our method obviates the requirements for paired rainy observations in the target domain, leveraging exclusively rain-free background images. We design a Superpixel Generation (Sup-Gen) module to extract stable structural priors from the source domain using Simple Linear Iterative Clustering. Subsequently, a Resolution- adaptive Fusion strategy is introduced to align these source structures with target backgrounds through texture similarity, ensuring the synthesis of diverse and realistic pseudo-data. Finally, we implement a pseudo-label re-Synthesis mechanism that employs multi-stage noise generation to simulate realistic rain streaks. This framework functions as a versatile plug- and-play module capable of seamless integration into arbitrary deraining architectures. Extensive experiments on state-of-the-art models demonstrate that our approach yields remarkable PSNR gains of up to 32% to 59% in OOD domains while significantly accelerating training convergence. Index Terms—Image Deraining, Image Restoration, Cross- scenario, Domain Adaptation, Pseudo-label Learning. I. INTRODUCTION I Mage deraining, which aims to remove rain streaks from rainy observations to recover the latent clean images, serves as a fundamental problem in low-level computer vi- sion [1]–[5]. Propelled by the rapid advancements in deep learning, deraining methodologies have witnessed remarkable improvements in both accuracy and robustness [3], [6]–[8]. Conventionally, deep learning-based deraining is formulated as a supervised learning problem, necessitating extensive datasets of paired rainy inputs and ground-truth clean images for model training [9]–[13]. However, acquiring high-quality, large-scale paired datasets in real-world rainy scenarios poses a significant challenge, constrained by unpredictable weather K. Zhao, M. Guan and Y. Shi are with School of Information Engi- neering, Guangdong University of Technology, Guangzhou, 510006, China (email: zhaokangbo329@gmail.com; guanmiaoxin1@mails.gdut.edu.cn; yk- shi@gdut.edu.cn) . X. Chen and J. Pan are with the School of Computer Science and Engineering, Nanjing University of Science and Technology (email: chenxi- ang@njust.edu.cn; sdluran@gmail.com). conditions, hardware limitations, and prohibitive annotation costs. Consequently, most existing models resort to training on synthetic data, typically generated by superimposing simulated rain streaks onto clean backgrounds. While this synthetic paradigm demonstrates efficacy in controlled settings and enables models to excel within the source domain, it often fails to faithfully replicate the complex physical dynamics of natural rain, including varying illumi- nation, perspective occlusions, and haze interference [14]. Consequently, these models suffer from substantial perfor- mance degradation when generalized to unseen target scenes, also known as Out-of-Distribution domains [15]–[17]. We regard datasets exhibiting distinct stylistic characteristics [18] as diverse scenarios and evaluate the performance of models pre-trained on Rain200L [9], including NeRD-Rain [19], DRS- former [20], and FADformer [21]. While the models achieve impressive PSNR scores in the source domain, they suffer a precipitous performance decline ranging from 30% to 60% when applied to target domains such as Rain200H [9], DID- Data [22], and DDN-Data [12]. These results underscore the severe domain shifts inherent in cross-domain scenarios. Em- pirically, the generated outputs frequently exhibit noticeable artifacts, including residual rain streaks, textural distortions, and even a loss of background fidelity. Consequently, generalizing image deraining to unseen sce- narios has emerged as a pivotal research frontier demanding urgent attention in recent years [23], [24]. To thoroughly investigate the implications and challenges associated with this issue, we conceptualize ”unseen scenes” as a series of datasets characterized by distinct stylistic divergences. Subsequently, we conduct a systematic evaluation of the generalization performance of deraining models trained on a single source dataset across multiple target datasets. As shown in Fig. 1, experimental results demonstrate that the majority of con- temporary deraining models suffer from a marked decline in performance during cross-dataset adaptation, frequently manifesting as residual rain streaks or over-deraining artifacts. These findings corroborate the ubiquity and severity of ”cross- scenario domain discrepancies. ” Thus, domain discrepancies are pervasive in real-world applications, including intelligent security surveillance [25] and autonomous driving systems [26], [27]. Given the high heterogeneity of imaging devices and environmental condi- tions in these contexts, coupled with the scarcity of training data, there is an imperative need for a cross-scenario deraining methodology capable of autonomously adapting to stylistic arXiv:2603.21661v1 [cs.CV] 23 Mar 2026 2 MethodNeRD-RainDRSformerFADformer Training Set: Rain200L(IID) Metrics:PSNR↑SSIM↑PSNR↑SSIM↑PSNR↑SSIM↑ Testing Set: Rain200L (IID) 41.710.990341.230.989441.590.9901 DDN-Data (OOD) 27.32(34%↓)0.8716(11%↓)27.43(33%↓)0.8530(13%↓)26.33(36%↓)0.8254(16%↓) DID-Data (OOD) 24.13(42%↓)0.7983(19%↓)24.09(41%↓)0.7500(24%↓)23.75(42%↓)0.7357(25%↓) Rain200H (OOD) 16.95(59%↓)0.5358(45%↓)18.70(54%↓)0.5848(40%↓)16.74(59%↓)0.5254(46%↓) Fig. 1. The challenge of removing raindrop across various Out-of-distribution (OOD) scenarios. We first use the Rain200L dataset to train the NeRD-Rain, DRSformer, and FADformer methods respectively. In order to implement the evaluation on Out-of-distribution(OOD) scenarios, we use Rain200L as the source domain data, and Rain200H, DID, and DDN as the OOD domains for testing. Empirical results demonstrate that while the model achieves exemplary deraining performance on the source domain (Rain200L), the performance of state-of-the-art derain models deteriorates significantly when applied to unseen OOD domains. Thus, we call for a solid derain pipeline is required to better handle unknown scene conditions. variations. To this end, this study presents a pioneering exploration of a cross-scenario deraining adaptation framework that operates independently of paired rainy and clean observations within the target domain. Requiring exclusively clean background images from the target dataset, the proposed method leverages a rain streak synthesis mechanism to progressively guide the model’s adaptation from the source to the target domain, thereby enabling it to capture the latent statistical charac- teristics intrinsic to the target scene. With the help of that, our model avoids the reliance on real paired data of the target domain, being more feasible and scalable in practical deployment. In contrast to conventional Synthetic→ Real or Real→ Real adaptation paradigms, our proposed pseudo-label scene adaptation methodology aligns more closely with the intrinsic constraints and data characteristics of real-world applications. By effectively mitigating the domain discrepancy between training datasets and testing environments, this approach holds substantial potential for practical deployment toward non- aligned real-world scenarios. The primary contributions of this paper are summarized as follows: • We propose the first cross-scenario deraining adaptation framework that eliminates the need for rainy target ob- servations. By training solely on rain-free target images, our method significantly enhances the feasibility of real- world deployment. • We introduce key information extraction and unsuper- vised augmentation mechanisms combining superpixel structural priors with texture matching. Utilizing Simple Linear Iterative Clustering (SLIC) for structural con- sistency and an MSE-based sliding window for align- ment, we seamlessly blend source structures onto target backgrounds. This approach ensures natural fusion and improves the diversity of pseudo-data. • Our method presents a plug-and-play flexibility toward existing deraining architectures without modification. Ex- tensive evaluations on state-of-the-art models demon- strate that our approach accelerates convergence and significantly improves out-of-distribution generalization, achieving PSNR gains of 32% to 59%. I. RELATED WORK A. Image Deraining Image deraining aims to restore clean background scenes from rain-degraded observations, thereby enhancing the recog- nition accuracy of downstream computer vision tasks [9], [28]–[31]. Advancements in the field have driven a paradigm shift in deraining methodologies, evolving from conventional model-based approaches to data-driven deep learning tech- niques, and then to fusing attention mechanisms and Trans- formers [2], [3], [6]. Early iterations of image deraining methodologies were primarily predicated on physical models and hand-crafted priors [2], [3], [10], [32]. Notably, Zhang et al. introduced the Image Deraining Conditional Generative Adversarial Network (ID-CGAN) [33], which enhances both the visual fidelity and discriminative performance of the recovered images by incorporating generative adversarial mechanisms. 3 Initialize Cluster Center I target psdo Adapt Progress Sup-Gen ⊙ α 1-α 퐼 input fuse . Lab Transform . A(l,a, b) . A(풙 풂 , 풚 풂 ) 푥 푎 푦 푎 Location Matrix . A(l , a , b , 풙 풂 , 풚 풂 ) Five-dimensional Matrix Re-Cluster Recalculate Cluster Center End Cluster Re-Cluster 퐼 target superpixel (A) Superpixel Generation Judge 퐷= d 푐 m 2 + d 푠 S 2 d 푠 =푥 푗 −푥 푖 2 +푦 푗 −푦 푖 2 , d 푐 =푙 푗 −푙 푖 2 +푎 푗 −푎 푖 2 +푏 푗 −푏 푖 2 , 푚 푖 = 1 퐶 푖 σ 푧 푘 ∈퐶 푖 푧 푘 , 푖=1,2,3⋯,푘 A(r ,g , b) Element-wise Add * Convolution ⊙ Element-wise Multiply Judge Block Psedo-Syn (C) Pseudo-label Re-Synthesize YUV Transform RGB Transform 퐼 푌 Saltpepper Gaussian Motionblur X X s X g X m x 1 x 2 x 3 x 4 x 5 x 6 x 7 x 8 x 9 x 10 x 11 x 12 x 13 x 14 x 15 x 16 X x 1 255 x 3 x 4 x 5 255 x 7 x 8 x 9 x 10 255 x 12 255 x 14 x 15 255 Random Change X s * Kernel WeightFunction I ′ uv= x=−k k y=−k k Iu−xv−y⋅Gxy I input psdo I target psdo I GM uv= x=−L L y=−L L I G u−xv−yM θ xy M θ xy= ቐ 1 L ,ifxy∈D θ ,xy≤ 퐿 2 0,else 퐼 푌 푝푠푑표 β 1-β P p ∗ M random I target psdo 퐼 target fuse 퐼 target superpixel Adapt Derain Network Judge 퐼 source 퐼 source Texture-Aware Match Psedo - Syn 퐼 input fuse 퐼 target fuse (B) Resolution-adaptive Fusion Fig. 2. An illustration of our proposed method: (A) Superpixel Generation: This module takes the source domain rain-free image I source as input. Utilizing a superpixel segmentation algorithm, it parses the image into a collection of superpixel blocks rich in structural and textural information, denoted as I superpixel target . (B) Resolution-adaptive Fusion: This module performs local matching between the extracted superpixel blocks I superpixel target and the target domain’s pseudo rain-free image I psdo target . It locates optimal integration regionsP ∗ p based on semantic consistency and texture similarity(MSE). By employing a random mask matrix M random for proportional extraction and applying an α-fusion strategy with a fusion coefficient of β, it generates an information-enhanced target image, I fuse target . (C) Pseudo-label Re-Synthesis: To synthesize high-fidelity pseudo rain streak layers, the module implements a three-stage degradation process comprising Salt-and-Pepper noise (X s ), Gaussian blurring (X g ), and motion blurring (X m ). This streak layer is superimposed onto the luminance channel of I fuse target using α fusion to create the corresponding rainy image, I fuse input . With the synergistic operation of these three modules, a final set of pseudo-paired samples(I fuse input ,I fuse target ) is obtained, significantly enhancing the model’s generalization toward out-of-distribution (OOD) domains. In recent years, deep learning methods have become the mainstream [3], [9], [11], [12]. Convolutional Neural Networks (CNN) are widely applied to image deraining tasks [34]–[39]. For example, the DID-MDN method [22] significantly improves the deraining effect by jointly estimat- ing rain density and deraining through a multi-stream densely connected convolutional neural network. Propelled by the remarkable success of Transformers in computer vision [40], [41], image deraining methodologies have increasingly incorporated Transformer-based architec- tures [42]. A notable instance is Restormert [43], which mit- igates computational complexity by calculating self-attention along the channel dimension. Additionally, TransWeather uti- lizes Transformer to handle multiple adverse weather condi- tions. Furthermore, the introduction of attention mechanisms has further improved the deraining effect [7], [12], [29], [44]. For example, networks based on residual channel attention mechanisms [28] recover clear images by accurately modeling the structure and details of the images [45]. To recapitulate, image deraining technology has undergone a significant evolution, progressing from conventional model- based approaches to data-driven deep learning paradigms, and subsequently to diverse architectures integrating Transformers and attention mechanisms [1]–[3], [46], and more recently, Bayesian-based restoration frameworks [47]. Driven by the ongoing refinement of models and the proliferation of compre- hensive datasets [1], [17], image deraining technology holds great promise for achieving superior performance in real-world applications. B. Pseudo-Label Learning Pseudo-labeling represents a semi-supervised learning strat- egy widely adopted in deep learning [16], [48]–[51], designed to leverage unlabeled data to enhance the performance of the model. Sun et al. [52] use predictions of the model on unlabeled data as pseudo-labels, and then incorporate them to expand the training scale. In the field of image deraining, the application of pseudo- label data augmentation is relatively rare, but the potential of it is gradually being focused on by researchers [27], [53]. Traditional image deraining methods mostly rely on manually annotated rainy images and clean image pairs [1], [17], which presents problems such as high annotation costs and limited data volume in practical applications. To solve this problem, researchers attempt to utilize pseudo-label data augmentation strategies. For instance, Huang et al. introduced the Multi-pseudo Reg- ularized Label (MpRL) [48] technique. By assigning multiple 4 Initialize Cluster Center . Lab Transform . A(l,a, b) . A(풙 풂 , 풚 풂 ) 푥 푎 푦 푎 Location Matrix . A(l , a , b , 풙 풂 , 풚 풂 ) Re-Cluster Recalculate Cluster Center End Cluster Re-Cluster 퐼 target superpixel Judge 퐷= d 푐 m 2 + d 푠 S 2 d 푠 =푥 푗 −푥 푖 2 +푦 푗 −푦 푖 2 , d 푐 =푙 푗 −푙 푖 2 +푎 푗 −푎 푖 2 +푏 푗 −푏 푖 2 , 푚 푖 = 1 퐶 푖 σ 푧 푘 ∈퐶 푖 푧 푘 , 푖=1,2,3⋯,푘 A(r ,g , b) 퐼 source YUV Transform RGB Transform 퐼 푌 Saltpepper Gaussian Motionblur X X s X g X m x 1 x 2 x 3 x 4 x 5 x 6 x 7 x 8 x 9 x 10 x 11 x 12 x 13 x 14 x 15 x 16 X x 1 255 x 3 x 4 x 5 255 x 7 x 8 x 9 x 10 255 x 12 255 x 14 x 15 255 Random Change X s * Kernel WeightFunction I ′ uv= x=−k k y=−k k Iu−xv−y⋅Gxy I input psdo I target psdo I GM uv= x=−L L y=−L L I G u−xv−yM θ xy M θ xy= ቐ 1 L ,ifxy∈D θ ,xy≤ 퐿 2 0,else 퐼 푌 푝푠푑표 β 1-β Fig. 3. The framework of Superpixel Generation, the process initiates by incorporating spatial positional information(x,y) into the image data(l,a,b). Subsequently, cluster centroids(l i ,a i ,b i ,x i ,y i ) are initialized based on predefined parameters. Clustering is then executed utilizing the SLIC [60] algorithm, involving the re-calculation of updated cluster centroids. pseudo-labels to generated imagery, this approach bolsters the capacity of the model to learn from unlabeled data, thereby yielding substantial performance gains in person re- identification tasks. Analogous strategies have also been investigated within the context of image deraining [15], [54]. By employing tech- niques such as Generative Adversarial Networks (GANs) [15], [33], [55], [56], researchers generate pseudo-labeled data, which is subsequently co-trained with ground-truth labeled data to enhance the generalization capability of the model. In addition, combining self-supervised learning and pseudo-label strategies further improves the performance of the model on unlabeled data [53], [57]–[59] . Nevertheless, the deployment of pseudo-label data augmen- tation strategies within image deraining remains in a nascent stage, confronted with critical hurdles including ensuring the quality of pseudo-labels and the mitigation of label noise [6]. In conclusion, pseudo-labeling strategies hold significant promise for image deraining applications. Future research can conduct in-depth exploration in aspects such as quality control of pseudo-labels and strategy optimization, to improve the performance and generalization ability of the model [1]. I. METHODOLOGY A. Problem As Fig. 1 shows, we first apply the Rain200L dataset to train a few de-raining methods (e.g., NeRD-Rain, DRSformer, and FADformer). To test the results of state-of-the-art models on the source domain and the target domain respectively, we use Rain200L as the source domain data. And Rain200H, DID, and DDN are adopted as the Out-of-Distribution (OOD) domains for testing. Empirical results demonstrate that while the model achieves exemplary deraining performance on the source domain. Otherwise, the performance of the model de- teriorates significantly when applied to unseen environments. To address the challenge of suboptimal deraining perfor- mance within unseen target domains, we propose across- domain adaptation module in Fig. 2. It is mainly composed of three sub-modules: • Superpixel Generation (Sup-Gen) takes the source do- main rain-free image I source as input, utilizing a super- pixel segmentation algorithm to parse it into a set of superpixel patches I superpixel target rich in structure and texture informationI superpixel target . • Resolution-adaptive Fusion performs local matching be- tween I superpixel target and the rain-free pseudo image of the target domain I psdo target . It locates the best integration region through correspondence relationships based on seman- tic consistency and texture similarity, utilizes a random mask matrix to extract superpixel blocks proportionally, and adopts an α-blending strategy to synthesize the information-augmented target image I fuse target . • Pseudo-label Re-Synthesis (Psedo-Syn) generats high- fidelity pseudo-rain streak layers and superimposing them onto the two images respectively in the luminance chan- nel in an α-blending strategy the corresponding rainy images I fuse input are synthesized. The synergistic integration of these three modules cul- minates in the generation of a set of pseudo-paired samples(I fuse input ,I fuse target ). This outcome significantly bolsters the generalization and adaptive capabilities of the model within Out-of-Distribution (OOD) domains. B. Superpixel Generation (Sup-Gen) To acquire high-quality local image patches that facilitate subsequent cross-domain fusion, this module employs the Sim- ple Linear Iterative Clustering (SLIC) algorithm. This process partitions the rain-free source image I source into a collection of superpixel patches I superpixel target , which are characterized by structural compactness and feature consistency. As illustrated in Fig. 3, the input image is initially trans- formed from the RGB to the CIELAB color space. Specif- ically, for each pixel p i , we construct a 5D feature vector V i = [l i ,a i ,b i ,x i ,y i ] T , where (l,a i ,b i ) represents the color coordinates in the CIELAB space and (x i ,y i ) denotes the spatial position. To effectively associate these heterogeneous parameters while resolving the scale disparity between color and spatial domains, we employ a normalized distance metric: d c = q (l j − l i ) 2 + (a j − a i ) 2 + (b j − b i ) 2 , d s = q (x j − x i ) 2 + (y j − y i ) 2 , D ′ = s d c m 2 + d s S 2 . (1) where C k and P i denote the feature vectors of the cluster center and the candidate pixel respectively. d c and d s are the Euclidean distances in the color and spatial domains, respectively. S = p N/k is the initial grid interval (where N is the pixel count and k is the number of superpixels), and m is the compactness parameter that balances boundary adherence and shape regularity. And we set k to 50 here. Regarding computational complexity, the integration of these 5D parameters does not impose a prohibitive burden. By constraining the search space for each cluster center to a 2S × 2S local neighborhood, each pixel is compared only 5 Initialize Cluster Center I target psdo Adapt Progress Sup-Gen ⊙ α 1-α 퐼 input fuse . Lab Transform . A(l,a, b) . A(풙 풂 , 풚 풂 ) 푥 푎 푦 푎 Location Matrix . A(l , a , b , 풙 풂 , 풚 풂 ) Five-dimensional Matrix Re-Cluster Recalculate Cluster Center End Cluster Re-Cluster 퐼 target superpixel Part A: Sup-Gen Judge 퐷= d 푐 m 2 + d 푠 S 2 d 푠 =푥 푗 −푥 푖 2 +푦 푗 −푦 푖 2 , d 푐 =푙 푗 −푙 푖 2 +푎 푗 −푎 푖 2 +푏 푗 −푏 푖 2 , 푚 푖 = 1 퐶 푖 σ 푧 푘 ∈퐶 푖 푧 푘 , 푖=1,2,3⋯,푘 A(r ,g , b) Element-wise Add * Convolution ⊙ Element-wise Multiply Judge Block Psedo-Syn Part B: Psedo-Syn YUV Transform RGB Transform 퐼 푌 Saltpepper Gaussian Motionblur X X s X g X m x 1 x 2 x 3 x 4 x 5 x 6 x 7 x 8 x 9 x 10 x 11 x 12 x 13 x 14 x 15 x 16 X x 1 255 x 3 x 4 x 5 255 x 7 x 8 x 9 x 10 255 x 12 255 x 14 x 15 255 Random Change X s * Kernel WeightFunction I ′ uv= x=−k k y=−k k Iu−xv−y⋅Gxy I input psdo I target psdo I GM uv= x=−L L y=−L L I G u−xv−yM θ xy M θ xy= ቐ 1 L ,ifxy∈D θ ,xy≤ 퐿 2 0,else 퐼 푌 푝푠푑표 β 1-β P p ∗ M random I target psdo 퐼 target fuse 퐼 target superpixel Adapt Derain Network Judge 퐼 source 퐼 source Texture-Aware Match Psedo - Syn 퐼 input fuse 퐼 target fuse Fig. 4. The framework of Resolution-Adaptive Fusion. We propose a novel fusion paradigm driven by bidirectional region matching and adaptive fusion mechanisms. This approach simultaneously incorporates structural priors from the source domain while preserving the intrinsic background distribution of the target domain, thereby facilitating subsequent domain adaptation tasks. with a limited number of adjacent centers . Consequently, the complexity is reduced from the standard K-means O(kNI) to a strictly linear complexity of O(N). This ensures that our Sup-Gen module maintains exceptional efficiency even for high-resolution images across diverse scenarios. The algorithm initializes cluster centers and iteratively per- forms pixel assignment and cluster center updates within a 2S×2S region around each center until the residual converges. The following is the formula for recalculating cluster centers each time: m i = 1 |C i | X z k ∈C i z k ,i = 1, 2, 3· ,k (2) where m i denotes the updated cluster centroid, z k represents the feature vectors of constituent pixels within the cluster, and C i corresponds to the cardinality of the cluster region. This process ultimately yields the set of superpixel patches I superpixel target , serving as foundational building blocks rich in structural information for the subsequent fusion stage. C. Resolution-adaptive Fusion To incorporate structural information from the source do- main, we propose a Resolution-Adaptive Fusion anchored in texture similarity and Mean Squared Error (MSE) matching. The inputs to this module comprise the set of source domain superpixel patches, denoted as I superpixel target ∈R H s ×W s ×3 , and the rain-free pseudo-image of the target domain. Here, H s ,W s and H p ,W p the corresponding dimensions of the target image represent the spatial resolution of the superpixel patches and the target image, respectively. This module aims to localise specific regions within the tar- get image to facilitate Resolution-Adaptive Fusion with source domain superpixels. The procedural framework involves iden- tifying the region that exhibits maximal texture similarity to the target image within the source domain superpixel patches. This is executed via a sliding window matching mechanism, where the alignment is optimised by minimizing the Mean Squared Error (MSE): (y ∗ ,x ∗ ) = arg min y,x 1 3H p W p I superpixel target [y : y + H p ,x : x + W p ]− I psdo target 2 2 (3) where I superpixel target [y : y + H s ,x : x + W s ] represents the candidate region extracted from the target image. After obtaining the optimal matching regionP ∗ s , random sampling is performed on it, followed by adaptive fusion with I psdo target : I fuse target = I psdo target · α +P ∗ s · M random · (1− α)(4) where M random is a random mask matrix controlling the sam- pling position of superpixel blocks, and α is a hyperparameter for balancing the fusion weight. I fuse target is the pseudo data after fusion. If we discover that I superpixel target does not meet the fusion conditions, we directly match a region of I psdo target to perform replacement transfer: P fuse p =P ∗ p · α + I superpixel target · M random · (1− α)(5) I fuse target = I psdo target −P ∗ p +P fuse p (6) whereP ∗ p is the optimal matching region obtained in I psdo target . Through the application of bidirectional region matching and adaptive fusion mechanisms, this module incorporates structural priors from the source domain while simultaneously preserving the background distribution intrinsic to the target domain. As evidenced in Fig. 4, the imagery processed by this module effectively integrates structural information from the source domain while maintaining the integrity of the target scene content. Consequently, the resultant fused pseudo-data exhibits enhanced diversity and stochasticity, supplying high- quality training samples for the subsequent domain adaptation of the deraining model. D. Pseudo-label Re-Synthesis (Psedo-Syn) As shown in Fig. 5, specifically, given the rain-free input imagery from the preceding module (i.e., comprising I psdo target and I fuse target ) the process initiates by transforming these inputs into 6 Initialize Cluster Center . Lab Transform . A(l,a, b) . A(풙 풂 , 풚 풂 ) 푥 푎 푦 푎 Location Matrix . A(l , a , b , 풙 풂 , 풚 풂 ) Re-Cluster Recalculate Cluster Center End Cluster Re-Cluster 퐼 target superpixel Judge 퐷= d 푐 m 2 + d 푠 S 2 d 푠 =푥 푗 −푥 푖 2 +푦 푗 −푦 푖 2 , d 푐 =푙 푗 −푙 푖 2 +푎 푗 −푎 푖 2 +푏 푗 −푏 푖 2 , 푚 푖 = 1 퐶 푖 σ 푧 푘 ∈퐶 푖 푧 푘 , 푖=1,2,3⋯,푘 A(r ,g , b) 퐼 source YUV Transform RGB Transform 퐼 푌 Saltpepper Gaussian Motionblur X X s X g X m x 1 x 2 x 3 x 4 x 5 x 6 x 7 x 8 x 9 x 10 x 11 x 12 x 13 x 14 x 15 x 16 X x 1 255 x 3 x 4 x 5 255 x 7 x 8 x 9 x 10 255 x 12 255 x 14 x 15 255 Random Change X s * Kernel WeightFunction I ′ uv= x=−k k y=−k k Iu−xv−y⋅Gxy I input psdo I target psdo I GM uv= x=−L L y=−L L I G u−xv−yM θ xy M θ xy= ቐ 1 L ,ifxy∈D θ ,xy≤ 퐿 2 0,else 퐼 푌 푝푠푑표 β 1-β Fig. 5. Pseudo-label Re-Synthesis introduces a pseudo-rain streak genera- tion methodology grounded in multi-stage noise synthesis. By sequentially executing noise initialization, morphological transformation, and luminance channel fusion within the YUV color space, the proposed method achieves the realistic simulation of rain effects on rain-free images within the target domain. the YUV color space and extracting the luminance component I Y . Subsequently, we generate rain streaks through three stages: First, create a rain streak mask X and inject salt-and- pepper noise with a density of p to obtain the point-like noise base X s . Subsequently, apply a k× k Gaussian kernel (stan- dard deviation σ g ) to perform blurring processing, forming the patch-like noise X g with optical gradient characteristics. Finally, convert the patch-like noise into linear rain streaks through motion blur, generating the final rain streak mask X m : M θ (x,y) = ( 1 L r , if (x,y)∈ D θ , |(x,y)|≤ L r 2 0,otherwise (7) I GM (u, v) = L r X x=−L r L r X y=−L r I G (u ′ − xv− y)M θ (xy)(8) where the rain streak length L r ∈ [L min ,L max ], in- clination angle θ r ∈ [θ min ,θ max ], and rain streak width W r ∈ [W min ,W max ] are all adjustable parameters. In the superposition stage, a weighted fusion strategy is adopted to superimpose rain streaks onto the luminance channel: I rainy Y = (1− β)· X m + β· I Y (9) where β is the fusion coefficient. Finally, merge the pro- cessed luminance channel with the original UV channels, and convert back to RGB space, obtaining the rainy image I fuse input . This module finally outputs a set of high-correlation pseudo- paired samples (I fuse input ,I fuse target ), providing key training data for the target domain transfer of the cross-domain deraining model. E. Losses The proposed methodology is characterized by a plug-and- play property. In our experimental evaluations, we applied this approach to a selection of representative deraining archi- tectures. During the training phase, the network is governed by a joint optimization strategy incorporating Charbonnier Reconstruction Loss, Frequency Domain Consistency Loss and Edge Preservation Loss. Charbonnier Reconstruction Loss: This loss term serves to quantify pixel-wise discrepancies between the reconstructed derained image and the ground-truth rain-free image. Formu- lated as a smoothed L 1 penalty, it enhances the stability of convergence: L char = q ( ˆ I − I gt ) 2 + ε 2 (10) where ε represents a constant typically set within the range of 10 −3 to 10 −6 , employed to mitigate the issue of gradient discontinuity. Frequency Domain Consistency Loss : This loss term im- poses constraints on spectral consistency between the re- constructed image and the ground-truth image within the frequency domain, placing particular emphasis on preserving textural details and structural distributions. To implement this, the Two-Dimensional Fast Fourier Transform (FFT) is first applied to the image: F( ˆ I) = FFT( ˆ I), F(I gt ) = FFT(I gt )](11) The frequency domain loss is defined as: L fft = 1 N X u,v ∥F( ˆ I) u,v |−|F(I gt ) u,v ||(12) Edge Preservation Loss: This loss term enhances the con- sistency of structural edges by imposing constraints on image gradients. It specifically utilizes Sobel or Laplacian operators to compute the edge maps: E( ˆ I) =∇( ˆ I), E(I gt ) =∇(I gt )(13) The loss function is defined as follows: L edge =|E( ˆ I)− E(I gt )| 1 (14) where ∇ denotes the gradient operator (specifically, the Sobel convolution kernel). Finally,the aggregate objective function is formulated as follows: Ltotal = λ 1 L char + λ 2 L fft + λ 3 L edge (15) where L char denotes the Charbonnier Reconstruction Loss (facilitating stable convergence within the pixel domain),L fft represents the Frequency Domain Consistency Loss (con- straining high-frequency details), andL edge signifies the Edge Preservation Loss (enhancing structural restoration). IV. EXPERIMENTAL VALIDATION A. Experimental Setup Implementation Details: The proposed framework is imple- mented using PyTorch, with all experiments conducted on an NVIDIA GeForce RTX 3090 GPU. During the training phase, the patch size is configured to 128×128 and the batch size to 12. The learning rate is initialized at 5×10 −5 and subsequently decayed to 5×10 −7 utilizing a cosine annealing strategy over 200 epochs. Out-of-Distribution Derain Benchmark: To simulate Out- of-Distribution (OOD) scenarios, we utilize HQ-RAIN [1] as the source dataset and Rain200L [9] as the target dataset. Specifically, the source dataset comprises paired samples with and without rain (w/ rain and w/o rain), whereas the target dataset contains exclusively clean, rain-free images (w/o rain) 7 NeRD-Rain(w/ ours) DRSformer(w/ ours) FADformer(w/ ours) DFSSM(w/ ours) NeRD-Rain(w/o ours)DRSformer(w/o ours)FADformer(w/o ours)DFSSM(w/o ours) Fig. 6. Derained results on the Rain200L dataset. Compared with the derained results without our method, our method recovers a high-quality image with clearer details. Zooming in the figures offers a better view at the deraining capability. TABLE I QUANTITATIVE COMPARISON ON OOD DATASET. Baseline Testing Set: Rain200L (OOD) PSNR↑SSIM↑LPIPS↓FSIM↑NIQE↓PI↓ NeRD-Rain [19] w/o ours25.440.83720.24830.88724.16452.7664 NeRD-Rain [19] w/ ours33.67 (32%↑)0.9495 (13%↑)0.1000(60%↑)0.9559(8%↑)3.2822(21%↑)2.3327(16%↑) DRSformer [20] w/o ours28.350.87870.19690.90683.66582.5371 DRSformer [20] w/ ours32.68 (15%↑)0.9478 (7%↑)0.0849(57%↑)0.9547(5%↑)3.2469(11%↑)2.3505(11%↑) FADformer [21] w/o ours27.180.85350.23750.89154.13802.7486 FADformer [21] w/ ours33.85 (24%↑)0.9550 (11%↑)0.0747(69%↑)0.9609(8%↑)3.2139(22%↑)2.3227(15%↑) DFSSM [61] w/o ours27.090.84990.24020.88794.15272.8573 DFSSM [61] w/ ours33.30 (22%↑)0.9490 (11%↑)0.0804(67%↑)0.9589(8%↑)3.2245(22%↑)2.3367(18%↑) TABLE I COMPARISON OF TRAINING EFFICIENCY AND PERFORMANCE BaselineVariantBest Ep.PSNR↑SSIM↑ NeRD-Rain random8427.700.8779 Ours15 (5.6×)33.670.9495 DRSformer random2624.520.8515 Ours6 (4.3×)32.680.9478 FADformer random15028.700.8894 Ours101 (1.5×)33.850.9550 B. Experimental Results On the Out-of-Distribution Derain Benchmark, we inte- grated our approach into four representative methods: NeRD- Rain [19], FADformer [21], DRSformer [20], and DF- SSM [61]. For a holistic assessment of deraining efficacy across both pixel-level fidelity and perceptual quality, we expand our quantitative framework to encompass a bifurcated suite of metrics: full-reference (PSNR, SSIM, LPIPS, and FSIM) and no-reference (NIQE and PI) indicators. Table. I presents the experimental results for Out-of- Distribution (OOD) scenarios. It is observed that the three model frameworks, when trained originally on the source domain (independent identically distributed, IID) dataset, yield suboptimal results when evaluated on the target domain (OOD) test set. To address this, we utilize the source-domain model as a pre-trained baseline and conduct transfer training leveraging the newly synthesized pseudo-data. Experimental results demonstrate that our approach yields a substantial gain in PSNR (18% to 35%) and SSIM (8% to 15%) across various benchmarks. The marked optimization in perceptual metrics quantitatively substantiates the efficacy of the Super- pixel Generation (Sup-Gen) module; specifically, the dramatic 8 Loss Nerd-Rain FADformer DRSformer Fig. 7. The ablation study of our algorithm on different baseline methods. Our algorithm has greatly enhanced the convergence speed and generalization capability on cross-scenario environments. TABLE I ABLATION STUDY ON RESOLUTION-ADAPTIVE FUSION SCALE (α) α0.80.60.40.2 PSNR↑32.8432.8332.5933.08 SSIM↑0.94270.94200.94120.9460 reduction in LPIPS alongside a peak FSIM of 0.9625 under- scores the module’s capability in preserving intrinsic structural priors and intricate textures within the target domain. Addi- tionally, the no-reference indicators NIQE and PI as proxies for image naturalness exhibit significant declines of 2% and 20% respectively to validate the superior fidelity of the restored images. Comparisons in Fig. 6 illustrate the significant advan- tages of our method in image deraining tasks. Specifically, in the regions highlighted by green and red boxes, the results demonstrate superior preservation of background content and textural details. Existing benchmark models frequently suffer from persistent rain streak artifacts or induce undesirable blur- ring of the underlying background details. Consequently, the proposed method exhibits remarkable adaptability to the target domain. Our approach reconstructs images characterized by aesthetic appeal, high naturalness, and sharp edge definitions, which is in seamless alignment with the superior FSIM and NIQE metrics achieved. C. Comparison with Random Rain Streaks As shown in Table. I, ”w/ random pseudo rain” denoted as simple random rain streaks are directly superimposed onto the target domain images. We conducted comparative experiments benchmarking this simple rain generation baseline against our proposed method under identical environmental conditions. Although the random rain streak benchmark achieves limited OOD adaptation, it fails to capture the physical dynamics and optical gradients of real-world rain, leading to poor perceptual quality. In contrast, our multi-stage Pseudo-label Re-Synthesis (Pseudo-Syn) module provides physics-inspired high-fidelity pseudo-paired samples, significantly accelerating the conver- gence of various architectures. Experimental results confirm that the proposed framework not only substantially improves training efficiency but also ensures robust and perceptually natural image restoration under severe domain shifts. D. Training Efficiency Comparison As shown in Table. I, in the three sets of efficiency comparison experiments, we recorded the trajectories of the loss curves for each group, as illustrated in Fig. 7. In terms of intra-group comparison, our method demonstrates faster convergence to the optimal state. Regarding random group comparison, our method maintains a lower average loss compared to the random group baselines, indicating superior robustness across different experimental settings. E. Ablations Ablation on Superpixel Generation: To evaluate the ef- fectiveness of the Sup-Gen module, specific ablation studies were conducted. Initially, we established a baseline simu- lating the absence of our method (denoted as ”w/o ours method”) by the direct superposition of random image patches. Subsequently, we investigated an alternative configuration by substituting the SLIC [60] algorithm with the NC05 [62] algorithm. Quantitatively, the NC05 configuration yields a PSNR of 32.56 dB, compared to 32.96 dB for the ’w/o ours’ variant. Benchmarking these variations against our proposed approach which yields a PSNR of 33.35db revealed that our method yields substantial improvements compared to the ”w/o ours method” baseline. Conversely, the utilization of the NC05 [62] algorithm resulted in observable performance degradation within the deraining task. Ablation on Resolution-adaptive Fusion Rate α: As Table. I shows, this table investigates the influence of the fusion coefficient α within the Resolution-Adaptive Fusion Equation. (5). We synthesized four distinct sets of pseudo- data corresponding to α values of 0.8, 0.6, 0.4, and 0.2, while maintaining all other variables constant, and subsequently performed transfer training. Empirical observations indicate that the optimal transfer performance is achieved when α is set to 0.2. V. CONCLUSION This paper presents a cross-scenario image deraining adap- tation methodology anchored in superpixel partitioning and pseudo-rain streak synthesis. By leveraging superpixel patch generation and fusion modules, the framework effectively incorporates structural priors from the source domain. Fur- thermore, by integrating a multi-stage rain streak synthesis 9 strategy, it facilitates the generation of high-fidelity pseudo- labeled data. Exhibiting robust generalization and portability, this approach offers an effective solution for image deraining in real-world applications. REFERENCES [1] X. Chen, J. Pan, J. Dong, and J. Tang, “Towards unified deep image deraining: A survey and a new benchmark,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. [2] H. Wang, Y. Wu, M. Li, Q. Zhao, and D. Meng, “A survey on rain removal from video and single image. arxiv 2019,” arXiv preprint arXiv:1909.08326, 2019. [3] W. Yang, R. T. Tan, S. Wang, Y. Fang, and J. Liu, “Single image derain- ing: From model-based to data-driven and beyond,” IEEE Transactions on pattern analysis and machine intelligence, vol. 43, no. 11, p. 4059– 4077, 2020. [4] Y. Yang and H. Lu, “Single image deraining via recurrent hierarchy enhancement network,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, no. 5, p. 1328–1339, 2020. [5] Z. Weng, X. Liu, C. Liu, X. Guo, Y. Shi, and L. Lin, “Dronesr: Rethinking few-shot thermal image super-resolution from drone-based perspective,” IEEE Sensors Journal, 2025. [6] Z. Zhang, Y. Wei, H. Zhang, Y. Yang, S. Yan, and M. Wang, “Data- driven single image deraining: A comprehensive review and new per- spectives,” Pattern Recognition, vol. 143, p. 109740, 2023. [7] X. Chen, J. Pan, K. Jiang, Y. Li, Y. Huang, C. Kong, L. Dai, and Z. Fan, “Detail-recovery image deraining via context aggregation networks,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 2, p. 537–550, 2021. [8] M. Wang, Y. Song, P. Wei, X. Xian, Y. Shi, and L. Lin, “Idf-cr: Iterative diffusion process for divide-and-conquer cloud removal in remote-sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, p. 1–14, 2024. [9] W. Yang, R. T. Tan, J. Feng, J. Liu, Z. Guo, and S. Yan, “Deep joint rain detection and removal from a single image,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, p. 1357–1366. [10] L.-W. Kang, C.-W. Lin, and Y.-H. Fu, “Automatic single-image-based rain streaks removal via image decomposition,” IEEE transactions on image processing, vol. 21, no. 4, p. 1742–1755, 2011. [11] X. Fu, J. Huang, X. Ding, Y. Liao, and J. Paisley, “Clearing the skies: A deep network architecture for single-image rain removal,” IEEE Transactions on Image Processing, vol. 26, no. 6, p. 2944–2956, 2017. [12] X. Fu, J. Huang, D. Zeng, Y. Huang, X. Ding, and J. Paisley, “Removing rain from single images via a deep detail network,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, p. 3855–3863. [13] Y. Zhang, Y. Zhang, and J. Sun, “Dynamic kernel-based progressive network for single image deraining,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 1, p. 207–219, 2022. [14] Y. Wang, D. Gong, J. Yang, Q. Shi, A. van den Hengel, D. Xie, and B. Zeng, “Deep single image deraining via modeling haze-like effect,” IEEE Transactions on Multimedia, vol. 23, p. 2481–2492, 2021. [15] R. Yasarla, V. A. Sindagi, and V. M. Patel, “Syn2real transfer learning for image deraining using gaussian processes,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, p. 2726–2736. [16] W. Wei, D. Meng, Q. Zhao, Z. Xu, and Y. Wu, “Semi-supervised transfer learning for image rain removal,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, p. 3877– 3886. [17] W. Li, Q. Zhang, J. Zhang, Z. Huang, X. Tian, and D. Tao, “Toward real-world single image deraining: A new benchmark and beyond,” arXiv preprint arXiv:2206.05514, 2022. [18] L. Yu, B. Wang, J. He, G.-S. Xia, and W. Yang, “Single image deraining with continuous rain density estimation,” IEEE Transactions on Multimedia, vol. 25, p. 443–456, 2023. [19] X. Chen, J. Pan, and J. Dong, “Bidirectional multi-scale implicit neural representations for image deraining,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, p. 25 627–25 636. [20] X. Chen, H. Li, M. Li, and J. Pan, “Learning a sparse transformer network for effective image deraining,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, p. 5896– 5905. [21] N. Gao, X. Jiang, X. Zhang, and Y. Deng, “Efficient frequency- domain image deraining with contrastive regularization,” in European Conference on Computer Vision. Springer, 2024, p. 240–257. [22] H. Zhang and V. M. Patel, “Density-aware single image de-raining using a multi-stream dense network,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, p. 695–704. [23] X. Li, Y. Feng, F. Zhou, Y. Liang, Z. Su, and K. Larson, “Revitalizing real image deraining via a generic paradigm towards multiple rainy patterns.” in IJCAI, 2024, p. 1029–1037. [24] X. Fu, J. Xiao, Y. Zhu, A. Liu, F. Wu, and Z.-J. Zha, “Continual image deraining with hypergraph convolutional networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 8, p. 9534– 9551, 2023. [25] W. Sultani, C. Chen, and M. Shah, “Real-world anomaly detection in surveillance videos,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, p. 6479–6488. [26] B. R. Kiran, I. Sobh, V. Talpaert, P. Mannion, A. A. Al Sallab, S. Yo- gamani, and P. P ́ erez, “Deep reinforcement learning for autonomous driving: A survey,” IEEE transactions on intelligent transportation systems, vol. 23, no. 6, p. 4909–4926, 2021. [27] W. Yang, S. Ji, R. T. Tan, and J. Liu, “Semi-supervised video deraining with dynamical rain generator,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 11, p. 7457–7470, 2022. [28] Q. Yi, J. Li, Q. Dai, F. Fang, G. Zhang, and T. Zeng, “Structure- preserving deraining with residue channel prior guidance,” in Proceed- ings of the IEEE/CVF international conference on computer vision, 2021, p. 4238–4247. [29] T. Wang, X. Yang, K. Xu, S. Chen, Q. Zhang, and R. W. Lau, “Spatial attentive single-image deraining with a high quality real rain dataset,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, p. 12 270–12 279. [30] Y. Ba, H. Zhang, E. Yang, A. Suzuki, A. Pfahnl, C. C. Chandrappa, C. M. De Melo, S. You, S. Soatto, A. Wong et al., “Not just streaks: Towards ground truth for single image deraining,” in European Conference on Computer Vision. Springer, 2022, p. 723–740. [31] C. Shi, L. Fang, H. Wu, X. Xian, Y. Shi, and L. Lin, “Nitedr: Nighttime image de-raining with cross-view sensor cooperative learning for dynamic driving scenes,” IEEE Transactions on Multimedia, vol. 26, p. 9203–9215, 2024. [32] Y. Li, R. T. Tan, X. Guo, J. Lu, and M. S. Brown, “Rain streak removal using layer priors,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, p. 2736–2744. [33] H. Zhang, V. Sindagi, and V. M. Patel, “Image de-raining using a con- ditional generative adversarial network,” IEEE transactions on circuits and systems for video technology, vol. 30, no. 11, p. 3943–3956, 2019. [34] K. He, Z. Li, and W. Zhang, “A dual-path interaction network for single image deraining,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 11, p. 4263–4276, 2021. [35] K. Li, J. Wu, L.-J. Deng, and T.-Z. Lu, “Single image deraining via scale- aware multi-stage recurrent network,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 12, p. 4803–4817, 2021. [36] Z. Huo, H. Li, and N. Wang, “Uncertainty-guided multi-scale residual learning-using a cycle spinning cnn for single image deraining,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 12, p. 4685–4698, 2021. [37] S. Suh, H. Cha, W. Lee, H.-E. Yoo, T. Kim, and H. G. Kang, “Hierarchy regulated dual-path network for single image deraining,” IEEE Transac- tions on Circuits and Systems for Video Technology, vol. 32, no. 9, p. 6071–6084, 2022. [38] X. Lin, L. Ma, B. Sheng, Z.-J. Wang, and W. Chen, “Utilizing two-phase processing with fbls for single image deraining,” IEEE Transactions on Multimedia, vol. 23, p. 664–676, 2021. [39] C. Wang, Y. Xing, Y. Wu, Z. Su, and J. Wu, “Dcsfn: Deep cross-scale fusion network for single image rain removal,” IEEE Transactions on Multimedia, vol. 22, no. 11, p. 2892–2907, 2020. [40] H. Chen, Y. Wang, T. Guo, C. Xu, Y. Deng, Z. Liu, S. Ma, C. Xu, C. Xu, and W. Gao, “Pre-trained image processing transformer,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, p. 12 299–12 310. [41] K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu et al., “A survey on vision transformer,” IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 1, p. 87–110, 2022. [42] J. Xiao, X. Fu, A. Liu, F. Wu, and Z.-J. Zha, “Image de-raining trans- former,” IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 11, p. 12 978–12 995, 2022. 10 [43] S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang, “Restormer: Efficient transformer for high-resolution image restoration,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, p. 5728–5739. [44] C. Yu, H. Zhu, X. Li, and H. Zhao, “Msaff-net: Multiscale attention feature fusion networks for single image dehazing and beyond,” IEEE Transactions on Multimedia, vol. 24, p. 2620–2633, 2022. [45] H. Chen, X. Chen, J. Lu, and Y. Li, “Rethinking multi-scale repre- sentations in deep deraining transformer,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 2, 2024, p. 1046– 1053. [46] X. Fu, Q. Qi, Z.-J. Zha, X. Ding, F. Wu, and J. Paisley, “Successive graph convolutional network for image de-raining,” International Journal of Computer Vision, vol. 129, no. 5, p. 1691–1711, 2021. [47] J. Xiao, X. Fu, Y. Zhu, and Z.-J. Zha, “Bayesian window transformer for image restoration,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. [48] Y. Huang, J. Xu, Q. Wu, Z. Zheng, Z. Zhang, and J. Zhang, “Multi- pseudo regularized label for generated data in person re-identification,” IEEE Transactions on Image Processing, vol. 28, no. 3, p. 1391–1403, 2018. [49] Y. Lu, Y. Lin, H. Wu, X. Xian, Y. Shi, and L. Lin, “Sirst-5k: Exploring massive negatives synthesis with self-supervised learning for robust infrared small target detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, p. 1–11, 2024. [50] Y. Li, Y. Lu, H. Wu, S. Zhang, L. Lin, and Y. Shi, “Ivan-istd: Rethinking cross-domain heteroscedastic noise perturbations in infrared small target detection,” IEEE Transactions on Geoscience and Remote Sensing, 2026. [51] Y. Lu, Y. Li, X. Guo, S. Yuan, Y. Shi, and L. Lin, “Rethinking gen- eralizable infrared small target detection: A real-scene benchmark and cross-view representation learning,” IEEE Transactions on Geoscience and Remote Sensing, 2025. [52] S. Sun, W. Ren, J. Zhou, S. Wang, J. Gan, and X. Cao, “Semi- supervised state-space model with dynamic stacking filter for real-world video deraining,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, p. 26 114–26 124. [53] Y. Liu, Z. Yan, T. Li, and X.-P. Sun, “Embedding consistency logic into models for unsupervised single image deraining,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 5, p. 2087– 2100, 2023. [54] Y. Wei, Z. Zhang, Y. Wang, H. Zhang, M. Zhao, M. Xu, and M. Wang, “Semi-deraingan: A new semi-supervised single image deraining net- work,” arXiv preprint arXiv:2001.08388, 2020. [55] H. Zhu, X. Peng, J. T. Zhou, S. Yang, V. Chanderasekh, L. Li, and J.-H. Lim, “Singe image rain removal with unpaired information: A differentiable programming perspective,” in Proceedings of the AAAI conference on artificial intelligence, vol. 33, no. 01, 2019, p. 9332– 9339. [56] Y. Wei, Z. Zhang, Y. Wang, M. Xu, Y. Yang, S. Yan, and M. Wang, “Deraincyclegan: Rain attentive cyclegan for single image deraining and rainmaking,” IEEE Transactions on Image Processing, vol. 30, p. 4788– 4801, 2021. [57] X. Jin, Z. Chen, J. Lin, Z. Chen, and W. Zhou, “Unsupervised single image deraining with self-supervised constraints,” in 2019 IEEE Inter- national Conference on Image Processing (ICIP).IEEE, 2019, p. 2761–2765. [58] Y. Liu, Z. Yue, J. Pan, and Z. Su, “Unpaired learning for deep image deraining with rain direction regularizer,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, p. 4753– 4761. [59] C. Yu, Y. Chang, Y. Li, X. Zhao, and L. Yan, “Unsupervised image deraining: Optimization model driven deep cnn,” in Proceedings of the 29th ACM International Conference on Multimedia, 2021, p. 2634– 2642. [60] R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua, and S. S ̈ usstrunk, “Slic superpixels compared to state-of-the-art superpixel methods,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 34, no. 11, p. 2274–2282, 2012. [61] S. Yamashita and M. Ikehara, “Image deraining with frequency- enhanced state space model,” in Proceedings of the Asian Conference on Computer Vision, 2024, p. 3655–3671. [62] T. Cour, F. Benezit, and J. Shi, “Spectral segmentation with multiscale graph decomposition,” in 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), vol. 2.IEEE, 2005, p. 1124–1131. Kangbo Zhao received the B.S. degree in 2024, from the School of Information Engineering, Guang- dong University of Technology, Guangzhou, China, where he is currently working towards a M.S. de- gree. His research interests include computer vision and machine learning. Miaoxin Guanreceived the B.S. degree from the School of Electronic Engineering, South China Agricultural University, Guangzhou, China, in 2025. She is currently working towards the M.S. degree at the School of Information Engineering, Guangdong University of Technology, Guangzhou, China. Her research interests include computer vision and ma- chine learning. Xiang Chen received the MS degree from the College of Electronic and Information Engineering, Shenyang Aerospace University, China, in 2022. He is currently working toward the PhD degree with the School of Computer Science and Engineer- ing, Nanjing University of Science and Technology, China. His research interests include computer vi- sion and deep learning, with special emphasis on image restoration and enhancement under adverse weather conditions. Yukai Shi received the Ph.D. degree from the School of Data and Computer Science, Sun Yat- sen University, Guangzhou China, in 2019. He is currently an associate professor at the School of Information Engineering, Guangdong University of Technology, China. His research interests include computer vision and machine learning. Jinshan Pan (Senior Member, IEEE) received the PhD degree in computational mathematics from the Dalian University of Technology, China, in 2017. He was a joint-training PhD student with the School of Mathematical Sciences, Dalian University of Tech- nology and also with Electrical Engineering and Computer Science, University of California, Merced, CA. He is currently a professor with the School of Computer Science and Engineering, Nanjing Univer- sity of Science and Technology. His research inter- ests include image deblurring, image/video analysis and enhancement, and related vision problems. He is a senior member of IEEE.