Paper deep dive
Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image Classification
Fangyan Zhang, Fan Zhang, Shiqi Zhou, Jun Ni, Carlos Lรณpez-Martรญnez, Qiang Yin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/23/2026, 2:38:51 AM
Summary
The paper introduces CV-SSMNet, a physics-aware complex-valued state-space network for PolSAR image classification. It addresses limitations in existing methods by modeling long-range spatial dependencies in the complex domain and using seven polarimetric scattering priors (H, A, alpha, Ps, Pd, Pv, Span) as FiLM-style modulation signals to guide feature evolution. The architecture integrates multi-scale complex convolutions, branch-wise CV-SSM encoding, and prior-guided recalibration, demonstrating improved accuracy and boundary preservation on L-band and P-band datasets.
Entities (13)
Relation Signals (6)
CV-SSMNet โ processes โ PolSAR
confidence 95% ยท CV-SSMNet for PolSAR image classification
CV-SSMNet โ uses โ CV-SSM
confidence 95% ยท The proposed method builds a complex-valued state-space model (CV-SSM) in the original complex domain
CV-SSMNet โ uses โ FiLM
confidence 92% ยท encoded as FiLM-style modulation signals to adaptively recalibrate complex-valued representations
CV-SSMNet โ evaluatedon โ L-band
confidence 90% ยท Experiments on three L-band benchmark datasets
CV-SSMNet โ evaluatedon โ P-band
confidence 90% ยท an additional P-band BIOMASS evaluation
Physical Priors โ modulates โ CV-SSMNet
confidence 90% ยท seven physically meaningful scattering priors... are encoded as FiLM-style modulation signals
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Polarimetric synthetic aperture radar (PolSAR) image classification is a representative task for physics-aware GeoAI, where land-cover semantics are closely coupled with electromagnetic scattering mechanisms. Many existing complex-valued networks can preserve amplitude-phase information, but they are often limited in long-range spatial dependency modeling and usually incorporate polarimetric priors only as input-level or shallow auxiliary features. As a result, physical knowledge is insufficiently used to guide deep feature evolution. To address this issue, this paper proposes CV-SSMNet, a physics-aware complex-valued state-space network with scattering-aware feature modulation for PolSAR image classification. The proposed method builds a complex-valued state-space model (CV-SSM) in the original complex domain to capture long-range spatial dependencies while preserving polarimetric amplitude-phase coupling. Meanwhile, seven physically meaningful scattering priors, are encoded as FiLM-style modulation signals to adaptively recalibrate complex-valued representations during feature evolution. CV-SSMNet further integrates multi-scale complex convolutions, branch-wise CV-SSM encoding, prior-guided recalibration, and lightweight global context aggregation, enabling physically guided representation learning from local scattering structures to global spatial context. Experiments on three L-band benchmark datasets and an additional P-band BIOMASS evaluation demonstrate that CV-SSMNet achieves competitive accuracy, improved regional consistency, and better boundary preservation, supporting the effectiveness of embedding polarimetric scattering mechanisms into complex-valued long-range GeoAI representation learning.
Tags
Links
- Source: https://arxiv.org/abs/2607.19787v1
- Canonical: https://arxiv.org/abs/2607.19787v1
Trouble viewing inline? Open PDF directly โ
Full Text
91,010 characters extracted from source content.
Expand or collapse full text
Highlights Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image Classification Fangyan Zhang, Fan Zhang, Shiqi Zhou, Jun Ni, Carlos Lรณpez-Martรญnez, Qiang Yin โข A physics-aware complex-valued state-space learning framework is developed for GeoAI-oriented PolSAR image classification. โข Polarimetric scattering priors, including H/A/alpha, Freeman-Durden powers, and Span, are encoded as conditional constraints rather than simple auxiliary inputs. โข Leakage-free spatial block experiments on L-band and P-band PolSAR datasets verify the effectiveness, robustness, and physical interpretability of the proposed model. arXiv:2607.19787v1 [cs.CV] 22 Jul 2026 Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image Classification Fangyan Zhang a,d , Fan Zhang a , Shiqi Zhou a , Jun Ni b , Carlos Lรณpez-Martรญnez c and Qiang Yin a,โ a The College of Information Science and Technology, Beijing University of Chemical Technology, Beijing, China b The School of Information Science and Engineering, Yunnan University, Kunming, China c The Department of Signal Theory and Communications, Polytechnic University of Catalonia, Barcelona, Spain d The School of Information and Cyberspace Security, Ningxia University, Yinchuan, China A R T I C L E I N F O Keywords: Physics-aware machine learning GeoAI Polarimetric scattering priors Complex-valued state space model Scattering-aware feature modulation A B S T R A C T Polarimetric synthetic aperture radar (PolSAR) image classification is a representative task for physics-aware GeoAI, where land-cover semantics are closely coupled with electromagnetic scattering mechanisms. Many existing complex-valued networks can preserve amplitudeโphase information, but they are often limited in long-range spatial dependency modeling and usually incorporate polarimetric priors only as input-level or shallow auxiliary features. As a result, physical knowledge is insufficiently used to guide deep feature evolution. To address this issue, this paper proposes CV-SSMNet, a physics- aware complex-valued state-space network with scattering-aware feature modulation for PolSAR image classification. The proposed method builds a complex-valued state-space model (CV-SSM) in the original complex domain to capture long-range spatial dependencies while preserving polarimetric amplitudeโphase coupling. Meanwhile, seven physically meaningful scattering priors, including ํป, ํด, ํผ, ํ ํ , ํ ํ , ํ ํฃ , and Span, are encoded as FiLM-style modulation signals to adaptively recalibrate complex-valued representations during feature evolution. CV-SSMNet further integrates multi-scale complex convolutions, branch-wise CV-SSM encoding, prior-guided recalibration, and lightweight global context aggregation, enabling physically guided representation learning from local scattering structures to global spatial context. Experiments on three L-band benchmark datasets and an additional P-band BIOMASS evaluation demonstrate that CV-SSMNet achieves competitive accuracy, improved regional consistency, and better boundary preservation, supporting the effectiveness of embedding polarimetric scattering mechanisms into complex-valued long-range GeoAI representation learning. 1. Introduction Polarimetric Synthetic Aperture Radar (PolSAR) is an active microwave remote sensing modality that can acquire scattering information of ground targets regardless of cloud cover and illumination conditions. From the perspective of physics-aware GeoAI, PolSAR image classification requires not only data-driven representation learning but also the ex- plicit incorporation of electromagnetic scattering knowledge into the feature evolution process. Owing to its sensitivity to surface geometry, dielectric properties, and scattering mech- anisms, PolSAR has become an important data source for Earth observation, resource survey, and environmental mon- itoring. With the continuous improvement of sensor spatial resolution and acquisition capabilities, PolSAR data have rapidly increased in scale, dimensionality, and scene com- plexity, creating new opportunities for fine-grained land- cover mapping and scene understanding. The recent release โ Corresponding author: Qiang Yin. This work was supported in part by the National Natural Science Founda- tion of China under Grant No. 62331026 and in part by the Natural Science Foundation of Shandong Province under Grant No. ZR2024ZD19. lucia@nxu.edu.cn (F. Zhang); zhangf@mail.buct.edu.cn (F. Zhang); 1285565899@q.com (S. Zhou); jun.ni@ynu.edu.cn (J. Ni); carlos.lopezmartinez@upc.edu (C. Lรณpez-Martรญnez); yinq@mail.buct.edu.cn (Q. Yin) ORCID(s): 0000-0003-0739-0307 (F. Zhang); 0000-0002-2058-2373 (F. Zhang); 0009-0004-0975-8771 (S. Zhou); 0000-0002-7105-8475 (J. Ni); 0000-0002-1366-9446 (C. Lรณpez-Martรญnez); 0000-0002-8413-4756 (Q. Yin) of large-scale complex-scene datasets such as AIR-PolSAR- Seg-2.0 further highlights the growing practical demand for robust PolSAR terrain classification under realistic condi- tions Wang et al. (2025b). As a fundamental task in Earth observation, PolSAR image classification plays a vital role in land-cover mapping, agricultural monitoring, change de- tection, and target recognition Wang et al. (2025a). PolSAR image classification is a natural scenario for physics-aware machine learning in GeoAI because land-cover semantics are strongly coupled with electromagnetic scattering mecha- nisms. Therefore, an effective GeoAI model should not only learn discriminative features from data, but also preserve complex-valued scattering information and use physically meaningful polarimetric descriptors to guide representation learning. Traditional PolSAR image classification methods mainly rely on scattering-mechanism modeling, polarimetric tar- get decomposition, and statistical distribution modeling An and Lin (2021); Zhuang et al. (2024); Wu et al. (2019). Representative physical and statistical descriptors, such as ํปโํดโํผ, FreemanโDurden scattering powers, Span, and Wishart-based models, provide interpretable cues for land- cover recognition Zhang et al. (2024b); Freeman and Durden (1998); Nie et al. (2015); Lee et al. (1999). However, these methods usually depend on handcrafted features and prede- fined assumptions, which limits their ability to capture high- level semantic relationships, multi-scale spatial structures, and long-range contextual dependencies in complex scenes. Zhang et al.: Preprint submitted to ElsevierPage 1 of 20 CV-SSMNet Deep learning has been widely introduced to improve PolSAR representation learning Xiao and Liu (2019); Hu et al. (2019). CNN-based, attention-based, self-supervised, and Transformer-style models have shown strong capabil- ity in spatial feature extraction and end-to-end classifica- tion Zhang et al. (2017); Kuang et al. (2024); Li et al. (2025); Jamali et al. (2023). For time-series or multi-temporal Pol- SAR data, recurrent learning, tensor-GCN, and attention- ViT models have also been explored to capture spatial- temporal dependencies Ni et al. (2021); Cheng et al. (2022); Yin et al. (2023). Nevertheless, many deep models operate in the real domain and require splitting or mapping complex- valued PolSAR observations into real-valued representa- tions, which may weaken the inherent amplitudeโphase cou- pling. To preserve complex-valued information, complex-valued deep neural networks have attracted increasing attention Alkhatib et al. (2025). Recent studies have explored complex-valued diffusion models, hybrid complex-valued networks, and complex-valued contourlet neural networks for PolSAR rep- resentation learning Kuang et al. (2025b); Alkhatib (2024); Liu et al. (2024a). These methods improve complex feature modeling, but they are still largely built on convolutional or transformer-style operators, making it challenging to effi- ciently capture long-range spatial dependencies in complex PolSAR scenes. Another important direction is to incorporate polarimet- ric scattering priors into deep networks. Physical descriptors such as ํปโํดโํผ, Span, and scattering powers have been used as auxiliary inputs, parallel branches, or fusion features to enhance the sensitivity of networks to scattering mecha- nisms Imani (2022); Ghazvinizadeh et al. (2023); Shi et al. (2023); Zhang et al. (2024a). However, most existing meth- ods treat these priors as additional features rather than using them as conditional information to regulate deep feature evolution. Therefore, the interaction between physical priors and complex-valued representations remains insufficiently explored. Recently, state-space models (SSMs) have emerged as an efficient alternative for long-range sequence modeling with linear complexity. SSM-based PolSAR classification has also been preliminarily investigated to enhance contextual dependency modeling Kuang et al. (2025a). However, exist- ing designs are mostly real-valued, and dedicated complex- valued SSM architectures for jointly modeling amplitudeโ phase coupling, regional scattering consistency, and long- range context remain underdeveloped. In summary, existing PolSAR image classification meth- ods still face three major limitations, as illustrated in Fig. 1 (a): 1) physical priors are usually used as input-level features instead of feature-evolution constraints; 2) amplitude-phase coupling and long-range spatial dependency are not jointly modeled in the complex domain; and 3) existing models lack a unified physics-aware framework that combines local scattering structures, global context, and interpretable po- larimetric priors. Local Modeling (Ignores distant context) Global Modeling Modeling-Long-Range Dependency Short-to-medium-range dependencies (Transformer) Captures long-range dependencies. Global context integration Physics-Aware Prior Modulation Auxiliary Information Constraints Unified Hierarchical Architecture Isolated Architecture Unified Hierarchical CV-SSMNet Raw Data Raw Data Discrimination Improved Discrimination (a) Existing Methods(b) Proposed Method Global Aggregation CV-CNN / Transformer Module A Module B Module C Module D Module E ๏ก , , A H Span v d s P P P , , H d P Multi-Branch Branch-Wise Pri or-Guided CV- Head representation learning ... Weakly constrained scattering-aware learning Prior modulation constraints Figure 1: Conceptual comparison between existing PolSAR classification methods (a) and the proposed CV-SSMNet (b) along three dimensions. To address these challenges, this paper proposes a prior- guided complex-valued state-space network (CV-SSMNet) for PolSAR image classification. The proposed method uses a state-space model (SSM) as the core operator for effi- cient long-range dependency modeling, and extends the state update process to the complex domain to directly process complex-valued PolSAR scattering features. Furthermore, a scattering-aware prior modulation mechanism is introduced to transform physically meaningful polarimetric priors into conditional signals aligned with the learned representation space, thereby enabling explicit guidance of physical knowl- edge during complex-valued state-space feature learning. As shown in Fig. 1 (b), the proposed network first ex- tracts local scattering structures using multi-scale complex- valued convolutions, then performs spatial-to-sequence re- arrangement and horizontalโvertical bidirectional scanning for long-range context modeling, and finally injects prior modulation and feature enhancement during representation evolution to produce pixel-wise classification outputs. The main contributions of our method are summarized as follows. โข Physics-aware scattering-prior-conditioned state- space learning framework: A polarimetric-prior- conditioned complex state-space learning framework is proposed, where seven polarimetric scattering pri- ors, namely ํป, ํด, ํผ, ํ ํ , ํ ํ , ํ ํฃ , and Span, are trans- formed from input-side auxiliary information into conditional constraints for the state-space represen- tation learning process. Specifically, conditional rep- resentations are generated by a Prior Encoder and in- jected into complex-valued feature evolution through Zhang et al.: Preprint submitted to ElsevierPage 2 of 20 CV-SSMNet feature-wise linear modulation (FiLM), scattering- prior-conditioned affine modulation, and prior-guided channel recalibration. In this way, the learned long- range dependencies are driven not only by data statis- tics but also by physically meaningful scattering mechanisms. In this work, physics-aware learning means that deterministic polarimetric scattering de- scriptors are transformed into conditional modulation signals that constrain deep complex-valued feature evolution. โข Complex-valued state-space modeling paradigm for PolSAR classification: A complex-valued state- space modeling paradigm is proposed to leverage the linear-complexity long-range dependency modeling capability of SSMs. By means of complex-valued state updates, horizontalโvertical bidirectional scan- ning, and stabilized state transition design, amplitudeโ phase coupling, regional scattering consistency, and cross-region contextual relationships are jointly mod- eled in the complex domain, thereby mitigating the locality limitation of conventional convolution-based methods. โข Hierarchical CV-SSMNet collaborative architec- ture: A hierarchical architecture is developed to unify local features, multi-scale context, global dependen- cies, and physical priors. By integrating three-branch multi-scale complex-valued convolutions, branch-wise CV-SSM modules, prior-guided fusion, and a lightweight global CV-SSM module, local scattering structures, multi-scale context, long-range global dependencies, and physics-based prior knowledge are jointly mod- eled. Furthermore, given the inherent strong spatial autocor- relation of PolSAR imagery, pixel-wise random partitioning may assign neighboring pixels to different subsets, which can introduce spatial information leakage and lead to overly optimistic classification results. Although such random sam- pling protocols have been widely used in PolSAR image classification, they may be less suitable for evaluating cross- region performance in highly autocorrelated scenes. There- fore, all experiments in this study adopt a spatial block-based partitioning strategy to reduce spatial information leakage between training and test regions. 2. Related Work 2.1. State Space Models State space models (SSMs) Gu et al. (2021) provide an efficient framework for long-range sequence modeling through recurrent state updates and convolutional inference. Mamba, introduced by Gu and Dao (2023), further intro- duces data-dependent selective scanning, enabling linear- time modeling of long sequences with improved computa- tional efficiency. Compared with self-attention, SSMs offer a favorable balance between global dependency modeling and scalability. Recent visual SSMs, such as Vision Mamba Zhu et al. (2024) and VMamba Liu et al. (2024b), extend selective scanning to two-dimensional visual inputs through bidi- rectional or multi-path scan strategies. In remote sensing, MSFMamba Gao et al. (2025) further explores Mamba- based modeling for multi-source image classification by designing multi-scale spatial, spectral, and fusion Mamba blocks, demonstrating the potential of SSMs for heteroge- neous remote sensing feature representation and fusion. For PolSAR classification, however, the complex-valued nature of the data introduces additional challenges, including amplitudeโphase coupling, scattering-mechanism preserva- tion, and stable complex-valued state evolution. Existing visual and remote-sensing SSMs are mainly designed for real-valued visual or multi-source representations, which motivates the development of a complex-valued SSM with scattering-prior-conditioned feature modulation for PolSAR representation learning. 2.2. Physics-Aware and Prior-Guided Learning in GeoAI Recent studies have increasingly explored the incorpo- ration of physical priors into remote sensing representation learning to enhance both discriminability and interpretabil- ity. In PolSAR analysis, polarimetric decomposition and scattering-characteristic modeling provide explicit physical cues related to surface scattering, double-bounce scattering, volume scattering, and randomness-dominated scattering mechanisms Han et al. (2023); Hu et al. (2023); Duan et al. (2024). Such priors are especially beneficial for distinguish- ing categories that exhibit similar visual appearances but dif- fer in their microwave scattering responses. Existing prior- guided PolSAR methods typically exploit physical descrip- tors through direct concatenation, parallel feature branches, or late fusion. Although these strategies can improve classi- fication performance, they generally lack an explicit mech- anism to regulate the evolution of intermediate representa- tions. In contrast, the proposed method regards polarimetric decomposition features as physics-aware conditioning vari- ables and injects them into the complex-valued state-space learning process via bounded FiLM modulation and prior- guided recalibration, thereby enabling physically informed feature refinement throughout representation learning. Learning-based methods further exploit such physical information through feature fusion, attention interaction, knowledge-guided learning, and conditional modeling Guo et al. (2024); Imani (2025); Hua et al. (2026); Geng et al. (2025). Instead of relying only on data-driven spatial fea- tures, these methods show that polarimetric and scattering priors can guide networks toward mechanism-consistent rep- resentations. However, direct concatenation or late fusion may treat physical priors as ordinary auxiliary channels, lim- iting their ability to adaptively regulate intermediate feature evolution. Conditional feature modulation provides a more flexible way to inject prior information by generating feature-wise scale and shift parameters from conditioning variables Dai Zhang et al.: Preprint submitted to ElsevierPage 3 of 20 CV-SSMNet et al. (2025). Motivated by this idea, the proposed SFM formulates the seven polarimetric priors, including ํป, ํด, ํผ, ํ ํ , ํ ํ , ํ ํฃ , and Span, as scattering-aware conditioning vari- ables for complex-valued feature modulation. Different from prompt-learning methods that introduce learnable prompt tokens or rely on pre-trained foundation models, SFM is implemented as a FiLM-style scattering-prior-conditioned modulation module tailored for complex-valued PolSAR representation learning. 3. Methodology 3.1. PolSAR Data Processing For a monostatic PolSAR system, the scattering matrix is denoted as ํบ = [ ํ ํป ํ ํปํ ํ ํ ํป ํ ํ ํ ] ,(1) where ํป and ํ represent horizontal and vertical polar- izations, respectively. Under the reciprocity assumption, ํ ํปํ = ํ ํ ํป . The scattering matrix is then represented in the Pauli basis as ํ = 1 โ 2 [ ํ ํป + ํ ํ ํ , ํ ํป โ ํ ํ ํ , 2ํ ํปํ ] ํ . (2) Based on the Pauli scattering vector, the multilook co- herency matrix is computed as ํป = 1 ํฟ ํฟ โ ํ=1 ํ ํ ํ ํป ํ ,(3) where ํฟ is the number of looks and the superscript ํป denotes the conjugate transpose. Since ํป is Hermitian, its diagonal elements are real- valued and its off-diagonal elements are complex-valued. Therefore, the six upper-triangular elements ํ 11 ,ํ 12 ,ํ 13 , ํ 22 ,ํ 23 ,ํ 33 are used as the complex-valued input chan- nels. For each center pixel, a local neighborhood with a size of 13 ร 13 is extracted, resulting in an input tensor of size 13 ร 13 ร 6. The patch representation preserves both polarimetric scattering information and local spatial context. To avoid information leakage during preprocessing, the channel-wise mean and standard deviation are computed only from the training set and then applied to the test set. In addition to the complex-valued coherency-matrix in- put, seven physically interpretable polarimetric descriptors are used as prior information, including the CloudeโPottier parameters ํป, ํด, and ํผ, the FreemanโDurden scattering powers ํ ํ , ํ ํ , and ํ ํฃ , and the total scattering power Span. The physical meanings of these seven scattering priors are summarized in Fig. 2. These descriptors form the physical prior vector ํ = [ํป,ํด,ํผ,ํ ํ ,ํ ํ ,ํ ํฃ , Span].(4) Low Entropy High Entropy simple surface scattering depolarizedscattering from forest (water, roads, bare soil) Single reflection Surface Scattering (low) Volume Scattering (mid) Double-Bounce Scattering (high) Dihedral scattering Vegetation canopy,forest, crops Total backscattered energy overall scattering intensity s P d P v P ๏ผ ๏ผ = Span Scattering Entropy Anisotropy Relative strength of secondary mechanisms 01 01 0 101 ๏ฏ 0 ๏ฏ 45 ๏ฏ 90 3 2 1 I I I ๏ฝ ๏พ 3 2 1 I I I ๏พ ๏ฝ 1 I 2 I 3 I H A Mean Scattering Angle ๏ก Surface Scattering Power s P Double-bounce Power d P Volume Scattering Power v P Total Scattering Span Power Span Figure 2: Illustration of the seven polarimetric scattering priors used in CV-SSMNet, including entropy ํป, anisotropy ํด, mean scattering angle ํผ, surface scattering power ํ ํ , double-bounce scattering power ํ ํ , volume scattering power ํ ํฃ , and total scattering power Span. 3.2. Overall framework We first describe the input representation and then intro- duce the overall pipeline of the proposed method. Our pro- posed CV-SSMNet consists of two key components: a multi- scale CV-SSM and scattering-aware feature modulation. As illustrated in Fig. 3, CV-SSMNet is organized into four func- tional parts: input PolSAR data and physical prior encod- ing, conceptual-to-discrete CV-SSM modeling, horizontalโ vertical complex state scanning, and prior-modulated feature enhancement for classification. The input of CV-SSMNet consists of a 13 ร 13 ร 6 complex-valued PolSAR patch and a seven-dimensional physical prior vector ํ = [ํป,ํด,ํผ,ํ ํ ,ํ ํ ,ํ ํฃ , Span] โค . The six complex channels are formed by the upper-triangular elements of the coherency matrix ํป , which preserves po- larimetric amplitude, phase, and amplitudeโphase coupling information. Therefore, CV-SSMNet directly processes the complex-valued PolSAR patch using complex-valued 3D convolutions, rather than converting it into a purely real- valued representation. 3.2.1. Stage 1: Multi-Branch Multi-Scale Complex-Valued Convolution A shallowโmediumโdeep three-branch ComplexConv3D structure with receptive fields of 3 ร 3, 5 ร 5, and 7 ร 7 extracts multi-scale amplitudeโphase coupled features. This transforms local spatial structures such as textures, edges, and point targets into higher-dimensional complex channel representations. The resulting 2D feature maps are spatially rearranged into 1D sequences via row-wise and column-wise scans. Each branch is followed by a CV-SSM module, which efficiently aggregates long-range contextual information. The row and column outputs are averaged and reshaped back to the spatial domain. Finally, the three branches are concatenated along the channel dimension to obtain a 13 ร 13 ร 6 ร 48 complex-valued feature map. 3.2.2. Stage 2: Scattering-aware Feature Modulation The physical prior ํ enters a conditional branch to inject physically interpretable preferences into the backbone Zhang et al.: Preprint submitted to ElsevierPage 4 of 20 CV-SSMNet Physical Prior Vector (7 Dim) Prior Encoder (MLP) H A ๏ก s P d P v P Span FiLM Generator Gate Fusion Complex feature Complex Tensor Prior - Guided SE Light Weight CV - SSM Prior - Guided SE 3 Complex - Valued Classification Head Classification Output Stage 2 -- Physical-aware Feature Modulation Real part Imag part Modulated Complex Feature Dense + ReLU+BN Stage1 -- Multi-scale Feature Extraction Prior Modulation(FiLM-Guided) Physical Modulated Vector Coherency Matrix Patch T Concat Prior Encoder (MLP) H A ๏ก s P d P v P Span H A ๏ก s P d P v P Span B. CV-SSM: Continuous to Discrete Complex Model Discrete &Complex m ht+1=A*ht+B Conceptual continuous dynamics A. Input PolSAR Data & Physical Priors B. CV-SSM: From Conceptual Dynamics to Discrete Complex Recurrence Complex PolSAR Physical scattering priors Complex Matrices Complex-Valued State Space Model Multi - scale pyramid features Complex sequence Complex recurrent state linear-complex memory with stable long-range dependency State update Output projection Skip connection Context sequence CV - SSM Blocks CV - SSM Blocks Class head PolSAR Land Cover Classification D. Complex Feature Enhancement & Classification C. Horizontal-Vertical Complex State Scan ... ... CV - SSM Horizontal Vertical Bidirectional scan H ๏ก d P Complex Features Bidirectional Complex States (Physics-Aware Feed-Forward Network) Prior Modulation Re Im ComplexMatrix Complex Discrete CV- SSM recurrence Complex Input tskiptoutt xDhCy๏ซ๏ฝ t in t t x B h A h ๏ซ ๏ฝ ๏ซ sn 1 ) ( ) ( ) ( t x B t h A t h c c ๏ซ ๏ฝ ๏ข ) ( ) ( ) ( t x D t h C t y c c ๏ซ ๏ฝ ... Figure 3: Overall architecture of CV-SSMNet. Complex PolSAR inputs and physical priors ํป,ํด,ํผ,ํ ํ ,ํ ํ ,ํ ํฃ , Span guide scattering-aware feature modulation. The CV-SSM illustrates conceptual continuous dynamics and implements a discrete complex recurrence with ํด sn , ํต in , ํถ out , and ํท skip , while the subscript ํ in ํด ํ , ํต ํ , ํถ ํ , and ํท ํ denotes conceptual continuous-domain parameters. features. These priors are not raw measurements but low- dimensional summaries of scattering mechanisms (e.g., wa- ter typically exhibits high ํ ํ and low ํป, urban areas show highํ ํ , and forests exhibit highํ ํฃ andํป). The prior vector is encoded via a Multilayer Perceptron(MLP) to obtain a conditional embedding ํ = MLP(ํ). FiLM performs feature-wise conditional affine modu- lation by generating scaling and shifting parameters (ํ,ํฝ) from ํ: ํ = ํ ํ (ํง) ํฝ = ํ ํฝ (ํง) (5) where ํ ํ (โ ) and ํ ํฝ (โ ) denote parameterized mapping functions. Subsequently, the calibrated complex features are directly fed into a lightweight CV-SSM for long-range dependency modeling, while prior-guided SE blocks oper- ate on the magnitude responses of complex channels and broadcast the resulting recalibration weights symmetrically to both real and imaginary parts, so that the complex-valued algebraic structure and phase information are preserved. Finally, adaptive fusion is achieved through Gate Fusion: ํ = ํฟ โ SSM(ํ) + (1 โ ํฟ) โ๎น(ํ,ํ)(6) where โ denotes element-wise multiplication, ํฟ denotes learnable gating coefficients, ํฟ โ โ ํถ and๎น(ํ,ํ) denotes the scattering-prior-conditioned modulation output. SSM(โ ) typically denotes the output features resulting from applying State Space Modeling to the input features, serving to cap- ture long-range dependencies. Here,ํ denotes the complex- valued feature representation extracted from the coherency matrix ํป via the backbone network. This mechanism dy- namically balances data-driven representation learning and physically informed guidance. 3.2.3. Stage 3: Classification head The fused complex-valued features are flattened and passed through two complex-valued fully connected layers (128 and 64 units) with Dropout regularization. The final complex logits are converted to real-valued probabilities by applying a modulus operator followed by Softmax, produc- ing the final class probabilities. Overall, this design preserves the phase-related informa- tion throughout the network while incorporating physically interpretable prior knowledge, which helps improve robust- ness and cross-scene classification performance. 3.3. Physics-aware conditional modulation The proposed SFM is the core physics-aware component of CV-SSMNet. It converts polarimetric scattering priors into bounded scaling and shifting parameters, so that the evolution of complex-valued features is guided by physi- cally interpretable scattering mechanisms. SFM explicitly injects polarimetric domain knowledge into the network by Zhang et al.: Preprint submitted to ElsevierPage 5 of 20 CV-SSMNet using physical scattering priors as conditioning variables to modulate complex-valued backbone features. As illustrated in Fig. 4, this module modulates complex-valued features through adaptive scaling and shifting, enabling effective integration of physically meaningful attributes. The mechanism consists of two components: 3.3.1. Scattering-aware Feature Modulation Design A seven-dimensional physical prior vector ํ โ โ ํตร7 is first encoded by a Prior Encoder implemented as a MLP, producing a conditional embedding: ํ = MLP(ํ)(7) Based on ํ, two modulation parameters are generated: a scaling vector ํ and a shifting vector ํฝ, both constrained to bounded ranges for stable conditioning. 3.3.2. Complex Feature Modulation Given an input complex-valued feature map ํน, the mod- ulation is applied separately to its real and imaginary com- ponents: ํ ํ ( ํน โฒ ) = ํ ํ โ ํฅ ํ + ํฝ ํ ํผํ ( ํน โฒ ) = ํ ํ โ ํฅ ํ + ํฝ ํ (8) whereํฅ โ is a complex-valued feature tensor with dimen- sions ํต ร ํป ร ํ ร ํน, and ํ is a channel-level modulation coefficient with dimensions ํต ร ํน. The resulting ํน โฒ is referred to as the modulated complex feature representation. This adaptive modulation mechanism enables the backbone network to leverage physical priors in a controllable manner, improving interpretability and robustness in complex-valued PolSAR feature learning. 3.3.3. Scattering-aware Feature Modulation Injection Mechanisms The scattering-aware feature modulation explicitly in- jects land-cover scattering characteristics into the network through conditional feature modulation, thereby guiding the learning of complex-valued feature representations. Specif- ically, a seven-dimensional physical prior vector [ํป,ํด,ํผ, ํ ํ ,ํ ํ ,ํ ํฃ , Span], derived from polarimetric decomposition, is employed as input. Each component corresponds to a physically interpretable property, including scattering en- tropy, anisotropy, mean scattering angle, and the power distribution among different scattering mechanisms. Within the network architecture, the physical priors are first processed by a Prior Encoder, which performs non- linear mapping to project the low-dimensional prior vector into a high-dimensional conditional space. This conditional representation is aligned with the channel dimension of the backbone features, enabling effective conditioning. Based on this representation, channel-wise scaling parameters ํ and bias parameters ํฝ are generated using the FiLM mechanism. The generated modulation parameters are applied sepa- rately to the real and imaginary components of the complex- valued features, preserving the complete complex-valued H A s P d P v P ๏ก ... ... ... ... ... ... Prior Encoder Relu Dense(64) Tanh Dense(128)Dense(48) Scale Gen ๏ณ Shift Gen ๏ณ Physical Priors Input (Bx7) Condition Vector(z) Complex Features Real Part Imaginary part Modulated Complex Features Scale ๏ข Shift Scale Shift ๏ข Complex Feature Modulation Flow Condition Generation Flow SFM:Scattering-aware Feature Modulation Span element-wise multiplication element-wise addition ๏ฌ ๏ข ๏ฌ ๏ซ ๏ฝ in out Feature Feature ๏ฌ Figure 4: Structure of the scattering-aware feature modulation (SFM) module. The physical prior vector is encoded to generate channel-wise scaling ํ and shifting ํฝ, which modulate the real and imaginary components of complex features. structure required for phase recovery. Simultaneously, shared scaling modulation helps maintain the consistency of this complex-valued structure, thus preventing the loss of phase- related information. In selected network layers, a gated fusion strategy is further introduced to dynamically balance the contributions of data-driven features and scattering- aware feature modulation. Through this design, the scattering-aware feature modu- lation integrates polarimetric scattering priors into the fea- ture evolution process while preserving the structure of complex-valued representations. This provides a physically motivated conditioning mechanism that can improve fea- ture interpretability and robustness across diverse land-cover scenes. 3.4. Complex-Valued State Space Model CV-SSM is a type of sequential learning framework designed for modeling complex-valued signals, capable of simultaneously characterizing both amplitude and phase in- formation in the complex domain. Compared to traditional real-valued models, CV-SSM is more suitable for processing complex-valued data such as polarimetric SAR data, effec- tively preserving scattering mechanisms and physical char- acteristics. Its core idea is to model long-range dependencies through a recursive state-space mechanism. By stabilizing information propagation and global context modeling in the complex feature space, CV-SSM enhances both the expres- sive power and the robustness of downstream classification. 3.4.1. Key implementation details of complex-valued SSM Fig. 5 illustrates the implemented discrete CV-SSM structure. The complex feature map is first serialized through horizontal and vertical flattening to form complex sequences, where magnitude and phase information are preserved. The sequences are then processed by complex state updates, output projection, and context propagation, enabling stable long-range dependency modeling in the complex domain. Zhang et al.: Preprint submitted to ElsevierPage 6 of 20 CV-SSMNet The input feature map of a CV-SSM block is denoted as ํฟ โ โ ํตรํปรํ รํท ํ รํถ ,(9) where ํต is the batch size, ํป and ํ are the spatial dimen- sions, ํท ํ is the polarimetric depth dimension, and ํถ is the feature-channel dimension. Each element of ํฟ is a complex value with real and imaginary components. In CV-SSMNet, two sequence construction strategies are used according to the position of the CV-SSM block. Branch-wise CV-SSM. For each multi-scale branch, the feature map is ํฟ ํ โ โ ํตรํปรํ รํท ํ รํถ ํ ,(10) where ํถ ํ = 16 in our implementation. In the branch-wise CV-SSM, the polarimetric depth dimension is preserved and folded into the batch dimension during spatial scanning: ํฟ ํ โ ฬ ํฟ ํ โ โ (ํตํท ํ )ร(ํปํ )รํถ ํ .(11) Therefore, the branch-wise CV-SSM uses ํฟ ํ = ํป ร ํ , ํ model = ํถ ํ .(12) For the actual input size ํป = ํ = 13, ํท ํ = 6, and ํถ ํ = 16, this gives ํฟ ํ = 13 ร 13 = 169, ํ model = 16.(13) To capture spatial dependencies along both directions, row-wise and column-wise serializations are adopted. For row-wise serialization, the sequence is defined as ํบ (ํ) ํ,ํ [ํก, โถ] = ํฟ ํ [ํ,โ,ํค,ํ, โถ] โ โ ํถ ํ , ํก = โโ ํ + ํค, (14) where โ โ [0,ํป โ 1], ํค โ [0,ํ โ 1], and ํ โ [0,ํท ํ โ 1]. Similarly, column-wise serialization is defined as ํบ (ํ) ํ,ํ [ํก, โถ] = ํฟ ํ [ํ,โ,ํค,ํ, โถ] โ โ ํถ ํ , ํก = ํคโ ํป + โ. (15) The row-wise and column-wise outputs are averaged and then reshaped back to โ ํตรํปรํ รํท ํ รํถ ํ . Lightweight CV-SSM. After the three branch-wise out- puts are concatenated along the channel dimension, the fused feature map becomes ํฟ ํ โ โ ํตรํปรํ รํท ํ รํถ ํ ,(16) where ํถ ํ = 48. In the lightweight CV-SSM, the spatial dimensions and the polarimetric depth dimension are jointly flattened into the sequence dimension: ํฟ ํ โ ฬ ํฟ ํ โ โ ํตร(ํปํ ํท ํ )รํถ ํ .(17) Thus, the lightweight CV-SSM uses ํฟ ํ = ํป ร ํ ร ํท ํ , ํ model = ํถ ํ .(18) For ํป = ํ = 13, ํท ํ = 6, and ํถ ํ = 48, this gives ํฟ ํ = 13 ร 13 ร 6 = 1521, ํ model = 48.(19) A single-direction scan with a diagonal state transition ma- trix is adopted in this lightweight block to reduce computa- tional cost. For a general complex-valued sequence ํบ โ โ ํต โฒ รํฟรํ model , the discrete-time complex SSM is formulated as ํ ํก+1 = ํจํ ํก + ํฉ in ํ ํก , ํ ํก = ํช out ํ ํก + ํซ skip ํ ํก , (20) where ํ ํก โ โ ํ model is the input token at position ํก, ํ ํก โ โ ํ is the hidden state, and ํ ํก โ โ ํ model is the output token. The learnable parameters are complex-valued: ํจ โ โ ํรํ , ํฉ in โ โ ํรํ model , ํช out โ โ ํ model รํ , ํซ skip โ โ ํ model รํ model . Here, ํจ controls state evolution, ํฉ in injects the input token into the hidden state, ํช out projects the hidden state back to the feature space, and ํซ skip provides a direct input-output path. To improve parameter efficiency, the state transition matrix is parameterized as ํ = ํ 0 + ํํ ํป , where ํ 0 is a structured base matrix and ํํ ํป is a low-rank complex adaptation term. In the lightweight CV-SSM block, a diago- nal ํ 0 is adopted to reduce computational cost. The stability of the SSM depends on the spectral radius ํ(ํจ). Since ํ(ํจ)โค ํ max (ํจ), spectral normalization is used to constrain the state transition: ํจ sn = ํจ max ( 1,ํ max (ํจ)โํ ) ,(21) whereํ max (ํจ) is the largest singular value andํ controls the upper bound of the spectral norm. In our implementation, ํ is set to 0.95โ0.98 for branch-wise CV-SSM and 0.9 for lightweight CV-SSM. The normalized matrix ํจ sn is then used in the recurrence: ํ ํก+1 = ํจ sn ํ ํก + ํฉ in ํ ํก .(22) Overall, the branch-wise CV-SSM focuses on spatial dependency modeling within each polarimetric depth slice, while the lightweight CV-SSM further captures global de- pendencies across both spatial positions and polarimetric depth. This clarification explains why the branch-wise CV- SSM usesํฟ = 169 andํ model = 16, whereas the lightweight CV-SSM uses ํฟ = 1521 and ํ model = 48. 3.5. Hierarchical CV-SSMNet Collaborative Architecture To address the insufficient use of physical priors, lim- ited long-range modeling in the complex domain, and frag- mented representation learning, CV-SSMNet integrates lo- cal scattering feature extraction, multi-scale context model- ing, complex-valued long-range dependency learning, and physical-prior guidance into a unified end-to-end architec- ture, as summarized in Table 1. Zhang et al.: Preprint submitted to ElsevierPage 7 of 20 CV-SSMNet Complex TensorsInputshtht+1=Aht + BXtht+1Inputs Outputsht =Aht+BXt Figure A: Complex-Valued State-Space Model (CV- SSM) Inputs Inputs Complex Input Feature Map Magnitude Phase Output s Complex Tensors Out puts Complex sequence linear - complex memory with stable long - range dependency State update Output projection Context sequence Complex recurrent state Col Flattening Row Flattening Figure 5: Structure of the proposed discrete CV-SSM. Row- wise and column-wise serialization convert complex feature maps into sequences, which are processed by complex state updates and output projection to model long-range dependen- cies. 3.5.1. Multi-Branch Multi-Scale Complex-Valued Convolution PolSAR scenes contain targets with different scales, spatial layouts, and scattering heterogeneity. Therefore, three parallel complex-valued convolution branches with recep- tive fields of 3 ร 3, 5 ร 5, and 7 ร 7 are used to cap- ture fine details, mesoscale patterns, and broader contextual structures, respectively. All branches operate directly in the complex domain, preserving amplitudeโphase coupling and polarimetric correlations while avoiding information loss caused by premature real-valued decomposition. Given an input complex-valued PolSAR patch ํ โ โ ํปรํ รํทรํถ ,(23) where ํป and ํ denote spatial dimensions, ํท is the polarimetric/depth-related dimension, and ํถ is the channel number, the three branches generate ํน (1) , ํน (2) , ํน (3) โ โ ํปรํ รํทรํถ .(24) These multi-scale features preserve local scattering struc- tures and provide complementary representations for subse- quent CV-SSM encoding. 3.5.2. Branch-Wise CV-SSM Encoding Although complex-valued convolutions effectively ex- tract local features, their receptive fields are still limited. To capture long-range scattering consistency and contextual dependencies, an independent CV-SSM encoder is assigned to each branch. For each branch, the spatial feature map is rearranged into a sequence and processed by complex- valued state updates with horizontalโvertical bidirectional scanning. This design connects local multi-scale extrac- tion with nonlocal contextual modeling while preserving amplitudeโphase coupling. 3.5.3. Prior-Guided Recalibration and Cross-Branch Fusion Physical priors are not simply concatenated with im- age features. Instead, the seven polarimetric priors ํป, ํด, ํผ, ํ ํ , ํ ํ , ํ ํฃ , and Span are encoded by a Prior En- coder into conditional embeddings aligned with the feature space. These embeddings guide feature-wise linear mod- ulation, scattering-prior-conditioned CV-SSM modulation, and prior-guided channel recalibration, thereby enhancing scattering-consistent responses and suppressing redundant activations. After branch-wise local-context modeling, the three branches are fused into a unified multi-scale representation and further refined by a lightweight global CV-SSM mod- ule. Since scale-specific dependencies have already been modeled in each branch, the global module mainly performs holistic integration, balancing representation capability and computational efficiency. The final representation jointly en- codes local scattering structures, multi-scale context, long- range dependencies, and physics-based semantic guidance. 3.5.4. Complex-Valued Classification Head To retain complex-valued representations until classi- fication, a complex-valued fully connected layer is used. Given ํง = ํฅ ํ + ํํฅ ํ , ํ = ํ ํ + ํํ ํ , and ํ = ํ ํ + ํํ ํ , the output is ํฆ = ํงํ + ํ,(25) which can be expanded as ํฆ ํ = ํฅ ํ ํ ํ โ ํฅ ํ ํ ํ + ํ ํ , ํฆ ํ = ํฅ ํ ํ ํ + ํฅ ํ ํ ํ + ํ ํ . (26) Thus, the real and imaginary components jointly contribute to classification. Complex-valued batch normalization, syn- chronized dropout, and Cartesian complex ReLU are used for stable optimization: ํ(ํง) = ReLU(ํฅ ํ ) + ํ ReLU(ํฅ ํ ).(27) In the prior modulation module, the scaling and shifting parameters are generated from the physical prior vector ํฉ: ํ = 1 + 0.5 tanh(ํ ํ (ํฉ)), ํฝ = 0.5 tanh(ํ ํฝ (ํฉ)). (28) The modulation is applied separately to the real and imagi- nary parts: ํฅ โฒ ํ = ํ ํ โ ํฅ ํ + ํฝ ํ , ํฅ โฒ ํ = ํ ํ โ ํฅ ํ + ํฝ ํ . (29) Finally, complex logits are converted to magnitudes be- fore Softmax. For the ํ-th class logit ํฆ ํ = ํฆ ํ,ํ + ํํฆ ํ,ํ , |ํฆ ํ | = โ ํฆ 2 ํ,ํ + ํฆ 2 ํ,ํ ,(30) and the class probability is ํ ํ = exp(|ํฆ ํ |) โ ํพ ํ=1 exp(|ํฆ ํ |) ,(31) whereํพ is the number of classes,ํ ํ โฅ 0, and โ ํพ ํ=1 ํ ํ = 1. Zhang et al.: Preprint submitted to ElsevierPage 8 of 20 CV-SSMNet Table 1 Layer-wise configuration of CV-SSMNet. StageLayerInputOutputDescription Input PolSAR_inputโ(ํต, 13, 13, 6, 1)Complex-valued patch Input prior_inputโ(ํต, 7)Physical prior vector Multi-scale branches shallow/mid/deep_conv(ํต, 13, 13, 6, 1)3 ร (ํต, 13, 13, 6, 16)1/2/3 CV-Conv3D layers Branch-wise modeling branch_cvssm(ํต, 13, 13, 6, 16)(ํต, 13, 13, 6, 16)CV-SSM per branch Feature fusion features_concat3 ร (ํต, 13, 13, 6, 16)(ํต, 13, 13, 6, 48)Channel concatenation Prior modulation concat_prior_cond(ํต, 13, 13, 6, 48)(ํต, 13, 13, 6, 48)FiLM conditioning Refinement se_block_1โ3(ํต, 13, 13, 6, 48)(ํต, 13, 13, 6, 48)Prior-guided SE Global modeling lightweight_cvssm(ํต, 13, 13, 6, 48)(ํต, 13, 13, 6, 48)Lightweight CV-SSM Classification head flatten+dense(ํต, 13, 13, 6, 48)(ํต,ํพ)Complex dense + magnitude softmax Note: Branch-wise CV-SSM uses (ํตํท ํ ,ํปํ ,ํถ ํ ) with ํฟ = 169 and ํ model = 16; lightweight CV-SSM uses (ํต,ํปํ ํท ํ ,ํถ ํ ) with ํฟ = 1521 and ํ model = 48. 4. Experiments 4.1. Dataset Description The proposed method is comprehensively evaluated on four real-world PolSAR datasets covering different sen- sors, frequency bands, spatial resolutions, and land-cover scenarios. Three widely used L-band benchmarks, namely Flevoland, San Francisco, and Oberpfaffenhofen, are adopted to evaluate classification performance in agricultural, urban, and mixed land-cover scenes. In addition, the P-band ESA BIOMASS dataset is used as a large-scale evaluation sce- nario to examine the applicability of CV-SSMNet under P- band observations with different scattering characteristics. The Pauli-RGB images and ground-truth maps of the four datasets are shown in Fig. 6. 1) Flevoland Dataset Vissers and van der Sanden (1992): The Flevoland dataset was acquired in 1989 over the Flevoland region of the Netherlands by the NASA/JPL AIRSAR sen- sor. It consists of L-band fully polarimetric SAR data with a spatial resolution of approximately 6 m and contains 15 crop categories, making it a classic benchmark for agricultural PolSAR classification. 2) San Francisco Dataset Liu et al. (2022): The San Francisco dataset was collected over the San Francisco Bay area by the AIRSAR airborne sensor. It provides L-band fully polarimetric SAR data with a spatial resolution of about 10 m and includes typical urban and suburban land-cover types, such as buildings, vegetation, and water. 3) Oberpfaffenhofen Dataset Hochstuhl et al. (2023): The Oberpfaffenhofen dataset was acquired in 2003 over Germany by the DLR E-SAR sensor. It contains L-band fully polarimetric SAR data with a spatial resolution of approximately 3 m, covering built-up areas, roads, open areas, and vegetation. 4) BIOMASS Dataset: The BIOMASS dataset is de- rived from the ESA BIOMASS mission, the first space- borne fully polarimetric P-band SAR mission for global for- est above-ground biomass estimation Quegan et al. (2019). Compared with L-band benchmarks, P-band observations provide stronger canopy penetration and different polarimet- ric scattering statistics, leading to greater spatial variability. Therefore, BIOMASS offers a challenging large-scale P- band test case. Since the model is trained and evaluated within this dataset, this experiment should be regarded as a P-band evaluation rather than a strict cross-band transfer setting. 4.2. Implementation Details 4.2.1. Mitigating Spatial Information Leakage via Block-Based Data Splitting To ensure a reproducible and leakage-free evaluation, a global spatial block partitioning protocol is adopted for all datasets. As shown in Fig. 7, the labeled map is first divided into non-overlapping spatial blocks with a size of 35 ร 35. The split is performed globally over the whole scene rather than independently for each class, so that neighboring pixels from different classes are assigned consistently to the same spatial subset. Following the common low-shot PolSAR setting, only 1% labeled pixels of each class are selected as training sample centers from the training candidate blocks, while the retained pixels in the spatially disjoint test regions are used for evaluation. Since the input patch size is 13ร13, a 6-pixel guard band, corresponding to the patch radius, is removed along train/test boundaries to prevent contextual overlap between patches from different subsets. Pixels in the guard band and unused pixels in training candidate blocks are excluded from both training and testing. The BIOMASS reference map used for classification has a size of 529 ร 378, containing 199,962 labeled pixels. The resulting sample statistics are summarized in Table 2. 4.2.2. Experimental environment and details All experiments were implemented with TensorFlow 2.6, the cvnn library, and Python 3.9, and conducted on an NVIDIA RTX A5500 GPU with 24 GB memory. For fair comparison with CV-ASDF2Net Alkhatib et al. (2025), the same 1% class-wise low-shot training setting is adopted. Test samples are selected using the proposed spa- tially disjoint block partition rather than pixel-wise random Zhang et al.: Preprint submitted to ElsevierPage 9 of 20 CV-SSMNet (a) WaterBare LandVegetation ForestBuild-up Build-up Woodland Open Areas (b) (c) (f) Bare SoilMountain Water VegetationUrban (d) Figure 6: Pauli RGB images and ground-truth maps of the four datasets used in this study. (a) Flevoland dataset (b) San Francisco dataset (c) Oberpfaffenhofen dataset (d) BIOMASS dataset. Table 2 Statistics of the leakage-free spatial split protocol. DatasetBlock size Total labeled Train Train (%)TestTest (%) Guard Guard (%) Unused Unused (%) Flevoland35 ร 35207,8322,0861.00172,16582.8424,64811.868,9334.30 San Francisco35 ร 35802,3028,0261.00666,70083.1095,66011.9231,9163.98 Oberpfaffenhofen 35 ร 351,311,618 13,1171.001,085,74682.78157,72312.0355,0324.20 Biomass35 ร 35199,9622,0021.00167,78883.9123,02011.517,1523.58 sampling, which helps reduce information leakage caused by spatial autocorrelation. Specifically, each PolSAR image is divided into non-overlapping spatial regions; training sam- ples are drawn from training regions, while labeled pixels in testing regions are used only for evaluation. The class- wise training/testing numbers are reported in Tables 3โ5. All compared methods and ablation variants use identical spatial partitions, sample numbers, random seeds, and evaluation protocols. Implementation Details. CV-SSMNet adopts a dual- input setting, including a complex-valued PolSAR patch of size 13 ร 13 ร 6 with shape (ํต, 13, 13, 6, 1) and a 7-D (a)(b)(c) Global Spatial block partitioning (d) Figure 7: Visualization of the leakage-free global spatial block partitioning protocol. (a) Flevoland, (b) San Francisco, (c) Oberpfaffenhofen, and (d) Biomass. Blue pixels denote the selected 1% training sample centers, red pixels denote the spatially disjoint test samples, gray regions denote the guard bands discarded to prevent 13 ร 13 patch overlap, and black regions denote background or unused labeled pixels. Zhang et al.: Preprint submitted to ElsevierPage 10 of 20 CV-SSMNet physical prior vector ํฉ = [ํป,ํด,ํผ,ํ ํ ,ํ ํ ,ํ ํฃ , Span]. The feature extractor contains three multi-scale ComplexConv3D branches with 1/2/3 convolutional layers for shallow, middle, and deep features. Each convolution uses a 3 ร 3 ร 3 kernel, 16 filters, same padding, and cart_relu. The branch outputs are concatenated into (ํต, 13, 13, 6, 48). Complex-valued SSM. CV-SSMNet includes three branch- wise CV-SSM blocks and one lightweight CV-SSM block. For each branch-wise CV-SSM, features of shape (ํต, 13, 13, 6, 16) are reshaped to (ํตโ 6, 169, 16), where ํฟ = 169 and ํ model = 16. For the lightweight CV-SSM, the fused feature (ํต, 13, 13, 6, 48) is reshaped to (ํต, 1521, 48), where ํฟ = 1521 and ํ model = 48. Scattering-aware Feature Modulation and Stabilization. Physical priors are injected through a FiLM conditioner. The prior encoder maps the 7-D prior vector to a (ํต, 128) condition using an MLP 7โ 64โ 128โ 128 with Tanh activation. The FiLM generator produces channel- wise modulation parameters ํ,ํฝ โ โ ํถ with ํถ = 48. Prior-guided SE blocks and gated fusion enhance scattering- aware feature selection. Spectral normalization constrains the spectral radius of ํด to 0.95โ0.98 for branch-wise CV- SSM blocks and 0.9 for the lightweight block. The classifier uses ComplexDense layers with 128/64 units, ComplexDropout with a rate of 0.25, and a magnitude-based softmax output. Training Details. The model is optimized by Adam with weight decay 1ร10 โ4 and an initial learning rate of 1ร10 โ3 . A cosine learning-rate schedule with a 5-epoch warmup and global gradient clipping with a max norm of 1.0 are used. The model is trained for 100 epochs with a batch size of 64. Unless otherwise specified, results are reported as mean ยฑ standard deviation over five independent runs, using OA, A, kappa coefficient, and per-class accuracy as evaluation metrics. 4.2.3. Comparison Methods To evaluate the effectiveness of the proposed CV-SSMNet, comparisons are conducted with seven representative Pol- SAR and remote sensing classification methods. These baselines are selected from four perspectives: label-efficient PolSAR representation learning, complex-valued spatial- scattering feature extraction, attention-based contextual mod- eling, and state-space/Mamba-based long-range dependency modeling. For fair comparison, all methods are evaluated under the same leakage-free spatial protocol, including identical train- ing sample centers, test regions, guard bands, patch size, data augmentation, batch size, training epochs, and evaluation metrics. Methods with official implementations are retrained using their released code whenever available, while methods without released code are reimplemented according to the original papers. To prevent test data leakage, the spatially disjoint test regions are used only for final evaluation and are never used for training or hyperparameter tuning. Hyperpa- rameters are fixed before final testing based on preliminary experiments conducted only within the training regions. The learning rate, weight decay, and dropout are selected from 5 ร 10 โ4 , 10 โ3 , 2 ร 10 โ3 , 10 โ5 , 10 โ4 , 5 ร 10 โ4 , and 0, 0.25, 0.5, respectively. โข Self-supervised and contrastive PolSAR represen- tation learning: PCLNet Zhang et al. (2022), PiCL Kuang et al. (2024), and SSL-MBC Li et al. (2025) are selected to assess label-efficient PolSAR feature learning. โข Complex-valued spatial-scattering feature extrac- tion: CV-3DCNN Tan et al. (2019) and CV-ASDF2Net Alkhatib et al. (2025) are adopted to evaluate complex- valued feature modeling and hierarchical scattering representation. โข Transformer-based contextual modeling: PolSAR- Former Jamali et al. (2023) is included to compare CV-SSMNet with attention-based PolSAR classifica- tion. โข State-space/Mamba-based long-range modeling: MSFMamba Gao et al. (2025) is introduced as a recent Mamba-based remote sensing baseline for evaluating efficient long-range dependency modeling. 4.3. Comparison and visualization analysis with SOTA methods This section comprehensively verifies the effectiveness of the proposed CV-SSMNet in PolSAR image classifi- cation tasks through systematic comparative experiments with seven representative state-of-the-art methods. Quanti- tative metrics and visualization results analysis evaluate the modelโs advantages in complex-valued long-range depen- dency modeling and physical prior guidance. Fig. 8 shows the visualization results on the Flevoland dataset, distinct methods exhibit significant differences re- garding the regional purity, boundary refinement, and spa- tial consistency of their classification maps. Comparative methodsโsuch as PCLNet, CV-3DCNN, and MSFMamba achieve reasonable performance in identifying major land- cover categories; however, they still manifest noticeable class mixing and speckle noise within complex bound- aries and small-scale target regions. While PiCL, PolSAR- Former, CV-ASDF2Net, and SSL-MBC demonstrate im- provements in terms of spatial continuity and regional integrity, they nonetheless remain deficient in the recovery of local details. In contrast, the proposed CV-SSMNet yields classification results that align more closely with the ground truth annotations; it not only effectively suppresses intra- regional noise but also preserves superior structural integrity and boundary clarity at parcel boundaries and within fine- grained target areas, thereby indicating the effectiveness of the proposed method in improving overall regional consis- tency and boundary preservation. As shown in Table 3, CV-SSMNet achieves the best overall performance on the Flevoland dataset, obtaining the highest OA of 97.56% and Kappaร100 of 97.05. It reaches 100% accuracy on several classes, including Water, Forest, Grass, Bare Soil, three Wheat-related classes, Barley, and Zhang et al.: Preprint submitted to ElsevierPage 11 of 20 CV-SSMNet (a) (b)(c) (d) (e)(f) (g) (h)(i) Figure 8: Visualization of classification results for the Flevoland dataset. (a) Ground Truth. (b) PCLNet. (c) CV-3DCNN. (d) MSFMamba. (e) PiCL. (f) PolSAR-Former. (g) CV-ASDF2Net. (h) SSL-MBC. (i) CV-SSMNet. Buildings, demonstrating strong global consistency and re- liable discrimination for categories with distinctive scatter- ing responses. Nevertheless, its advantages are not uniform across all classes. For visually similar crop categories, such as Lucerne, Beet, Potatoes, and Peas, the accuracies drop to 86.42%, 72.97%, 83.23%, and 89.30%, respectively, which are lower than several competing SOTA methods. This also explains why its A of 95.13% is slightly lower than SSL- MBC. These results indicate that CV-SSMNet improves overall classification consistency, while fine-grained crop discrimination remains challenging. As shown in Fig. 9, the San Francisco results show clear differences among the compared methods in complex urban scenes. PCLNet, CV-3DCNN, and MSFMamba can iden- tify major categories but still suffer from misclassifications and noise near building edges and transition areas. PiCL, PolSAR-Former, CV-ASDF2Net, and SSL-MBC improve spatial continuity to varying degrees, whereas CV-SSMNet produces maps closest to the ground truth, with sharper structural boundaries and fewer errors around buildings, roads, and small objects. The quantitative results in Table 4 further confirm this advantage. CV-SSMNet achieves the highest OA of 97.02% and Kappaร100 of 95.06, outperforming CV-ASDF2Net by 1.22% in OA and 1.04 in Kappaร100. It also obtains the best accuracies for Urban and Vegetation. Although its A is slightly lower than SSL-MBC due to weaker performance on Water and Mountain, CV-SSMNet shows stronger overall consistency and reliability. As shown in Fig. 10, the Oberpfaffenhofen visualization results reveal clear differences in regional consistency and boundary preservation. PCLNet, CV-3DCNN, and MSF- Mamba produce generally accurate maps but still contain speckle noise and local misclassifications, while PiCL, PolSAR-Former, CV-ASDF2Net, and SSL-MBC improve regional integrity to different extents. In contrast, CV- SSMNet aligns more closely with the ground truth, produc- ing cleaner regions, sharper boundaries, and better structural continuity. Zhang et al.: Preprint submitted to ElsevierPage 12 of 20 CV-SSMNet Table 3 Classification comparison results for the Flevoland dataset. The training set and test set represent the number of training samples and test samples for each class, respectively, and the samples were selected through a spatially disjoint partition. The best and second-best results are highlighted in bold and underlined, respectively. ClassTrainTestPCLNetCV-3DCNN MSFMambaPiCLPolSAR-Former CV-ASDF2Net SSL-MBC CV-SSMNet Water293 2449996.6798.6996.6999.1599.0198.8299.17100.00 Forest159 1294298.1198.87 95.7897.4896.2998.3398.61100.00 Lucerne112 994091.7697.6780.5597.1995.3495.6498.2086.42 Grass103 674191.4888.9082.2496.0688.8794.8495.18100.00 Rapeseed219 1935388.0598.1880.3791.6897.04 95.1395.1395.06 Beet148 1225994.5889.1485.9096.3994.8393.9797.2172.97 Potatoes214 1789284.7586.6586.8895.14 91.5891.9598.2583.23 Peas104 799999.9896.9791.0899.1296.7592.7298.6189.30 Stem beans85 671892.0398.02 87.4988.9694.4397.2390.09100.00 Bare Soil64 549688.9093.9189.6492.9994.68 91.5693.65100.00 Wheat177 1475997.7388.0776.5695.8595.1599.5197.81100.00 Wheat 2107 726892.1995.2992.1595.97 95.3091.9295.69100.00 Wheat 3221 1928790.6498.9796.7898.1396.5395.3997.82100.00 Barley74 659698.6297.0693.2789.9498.36 92.2397.65100.00 Buildings641699.0979.4187.6598.7981.8397.2199.12100.00 OA (%)92.09 ยฑ 0.22 94.72 ยฑ 1.32 90.12 ยฑ 1.55 95.10 ยฑ 0.1895.87 ยฑ 1.3995.33 ยฑ 0.7296.72 ยฑ 0.04ํํ.ํํ ยฑ ํ.ํํ A (%)93.63 ยฑ 0.15 93.72 ยฑ 1.10 88.20 ยฑ 1.42 95.52 ยฑ 0.1994.40 ยฑ 1.4395.10 ยฑ 0.21ํํ.ํํ ยฑ ํ.ํํ 95.13 ยฑ 0.23 Kappaร10091.37 ยฑ 0.26 92.16 ยฑ 1.13 89.56 ยฑ 1.77 94.65 ยฑ 0.2295.49 ยฑ 1.4593.68 ยฑ 0.8096.42 ยฑ 0.05ํํ.ํํ ยฑ ํ.ํํ (a) (b) (c) (d)(e) (f) (g)(h) (i) Figure 9: Visualization of classification results for the San Francisco dataset. (a) Ground Truth. (b) PCLNet. (c) CV-3DCNN. (d) MSFMamba. (e) PiCL. (f) PolSAR-Former. (g) CV-ASDF2Net. (h) SSL-MBC. (i) CV-SSMNet. Zhang et al.: Preprint submitted to ElsevierPage 13 of 20 CV-SSMNet Table 4 Classification comparison results for the San Francisco dataset. The training set and test set represent the number of training samples and test samples for each class, respectively, and the samples were selected through a spatially disjoint partition. The best and second-best results are highlighted in bold and underlined, respectively. ClassTrainTestPCLNetCV-3DCNN MSFMambaPiCLPolSAR-Former CV-ASDF2Net SSL-MBC CV-SSMNet Water3296 28308787.5086.0264.5395.6182.8185.8596.7186.57 Urban3428 27824292.4996.8493.5393.4391.7292.3494.0598.07 Vegetation536 4296193.1097.6298.9795.9499.24 93.4395.1999.40 Bare Soil138952889.6598.2497.1192.3195.9397.6395.3897.02 Mountain628 5288280.7087.4576.4085.9169.5776.0687.4283.10 OA (%)90.69 ยฑ 0.17 93.23 ยฑ 1.76 95.66 ยฑ 1.67 93.52 ยฑ 0.0994.98 ยฑ 1.4595.80 ยฑ 0.1294.09 ยฑ 0.10 ํํ.ํํ ยฑ ํ.ํํ A (%)88.69 ยฑ 4.30 93.23 ยฑ 1.6486.11 ยฑ 2.03 92.64 ยฑ 0.2187.85 ยฑ 1.4889.06 ยฑ 0.03ํํ.ํํ ยฑ ํ.ํํ 92.83 ยฑ 1.40 Kappaร10086.03 ยฑ 0.32 93.56 ยฑ 1.23 93.17 ยฑ 1.48 90.08 ยฑ 0.1992.12 ยฑ 1.4994.02 ยฑ 1.3291.83 ยฑ 0.25 ํํ.ํํ ยฑ ํ.ํํ Table 5 Classification comparison results for the Oberpfaffenhofen dataset. The training set and test set represent the number of training samples and test samples for each class, respectively, and the samples were selected through a spatially disjoint partition. The best and second-best results are highlighted in bold and underlined, respectively. ClassTrainTest PCLNet CV-3DCNN MSFMambaPiCLPolSAR-Former CV-ASDF2Net SSL-MBC CV-SSMNet Build-Up Areas 3281 26565882.5394.3966.2177.3691.8592.5090.4595.20 Wood Land2467 20055786.6595.4699.6389.0896.0092.3789.1995.17 Open Areas7369 61953189.4197.2398.8593.7992.3189.5083.8392.72 OA (%)87.17 ยฑ 0.11 95.69 ยฑ 1.47 91.26 ยฑ 1.68 88.80 ยฑ 0.3795.76 ยฑ 1.4792.50 ยฑ 0.1290.45 ยฑ 0.10 ํํ.ํํ ยฑ ํ.ํํ A (%)86.20 ยฑ 0.47 94.13 ยฑ 1.59 88.23 ยฑ 1.71 86.75 ยฑ 0.5093.39 ยฑ 1.5991.46 ยฑ 0.1587.82 ยฑ 0.30 ํํ.ํํ ยฑ ํ.ํํ Kappaร10078.47 ยฑ 0.35 93.54 ยฑ 1.54 85.14 ยฑ 1.87 86.96 ยฑ 1.0093.76 ยฑ 1.5489.50 ยฑ 0.1883.83 ยฑ 0.29 ํํ.ํํ ยฑ ํ.ํ The quantitative results in Table 5 further support this observation. CV-SSMNet achieves the best OA of 96.07% and Kappaร100 of 94.72, with the highest accuracy for Build-Up Areas and competitive performance on Wood Land and Open Areas. Although its A is slightly lower than CV-3DCNN, CV-SSMNet provides the most balanced performance across OA, A, and kappa, indicating effective modeling of structural consistency and regional dependen- cies. 4.4. Effect of Input Representation: Complex-Valued PolSAR Input vs. 107-D Real-Valued Features To verify whether preserving the native complex-valued structure is beneficial, CV-SSMNet is compared with a real-valued baseline (RV-SSMNet) using 107-dimensional handcrafted polarimetric features Ni et al. (2022) as input on the Flevoland dataset. As reported in Table 6, employing 107-dimensional handcrafted features, On the Flevoland dataset, CV-SSMNet outperforms the RV-SSMNet baseline, suggesting that preserving the native complex-valued co- variance representation is beneficial in this setting. In con- trast, direct complex-domain modeling allows CV-SSMNet to more realistically capture long-range dependencies and physically meaningful feature interactions. CV-SSMNet out- performs RV-SSMNet by 2.86, 2.90, and 2.84 percentage points in OA, A, and kappa coefficients, respectively, indicating that preserving the native complex-valued repre- sentation of PolSAR data is more effective than converting it to high-dimensional real-valued descriptors. 4.5. Ablation Experiments To dissect the contribution of each design choice in CV- SSMNet, a unified ablation study is conducted on three benchmark datasets. All variants share the same training protocol and spatial block-based partitioning to ensure a fair comparison. The aggregated results, including (A) compo- nent ablation, (B) prior ablation, and (C) scanning-direction ablation, are summarized in Table 7. 4.5.1. Effect of Different Modules CV-ASDF2Net is adopted as the baseline (Base), and the proposed CV-SSM and SFM modules are progressively incorporated. As shown in Table 7-A, the two modules provide complementary improvements. CV-SSM mainly en- hances long-range contextual modeling, leading to clear gains on Oberpfaffenhofen, where OA and A increase from 92.50% and 92.37% to 94.61% and 93.61%, respectively. SFM brings more consistent class-balanced improvements, especially on Flevoland, where A increases from 89.90% to 94.20%, indicating that scattering priors are effective for fine-grained crop discrimination. When CV-SSM and SFM are combined, CV-SSMNet achieves the best OA, A, and kappaร100 among all component variants on the three datasets. Compared with the baseline, the OA improvements are 2.23%, 1.22%, and 3.57% on Flevoland, San Francisco, and Oberpfaffenhofen, respectively, confirming the com- plementarity between complex-valued long-range modeling and scattering-aware feature modulation. 4.5.2. Contribution of Each Prior Parameter Table 7-B evaluates the influence of different physical prior groups. Using all priors achieves the highest OA on Zhang et al.: Preprint submitted to ElsevierPage 14 of 20 CV-SSMNet (a) (b) (c) (d) (e) (f) (g) (h) (i) Figure 10: Visualization of the classification results for the Oberpfaffenhofen dataset. (a) Ground Truth. (b) PCLNet. (c) CV- 3DCNN. (d) MSFMamba. (e) PiCL. (f) PolSAR-Former. (g) CV-ASDF2Net. (h) SSL-MBC. (i) CV-SSMNet. Table 6 Comparison between complex-valued input and 107-D real-valued feature input on the Flevoland dataset. MethodInput typeRepresentationOA (%)A (%)Kappaร100 RV-SSMNetReal-valued107-D handcrafted polarimetric features94.70 ยฑ 1.2592.23 ยฑ 2.0394.21 ยฑ 1.76 CV-SSMNetComplex-valuedComplex-valued covariance representationํํ.ํํ ยฑ ํ.ํํํํ.ํํ ยฑ ํ.ํํํํ.ํํ ยฑ ํ.ํํ all three datasets, with 97.56%, 97.02%, and 96.07% on Flevoland, San Francisco, and Oberpfaffenhofen, respec- tively. It also obtains the best A on Flevoland and San Fran- cisco, demonstrating that the joint use ofํปโํดโํผ, Freemanโ Durden powers, and Span provides complementary scatter- ing information. The FreemanโDurden powers (ํ ํ ,ํ ํ ,ํ ํฃ ) show strong performance on San Francisco and Oberpfaf- fenhofen, particularly in kappaร100 and A, suggesting that surface, double-bounce, and volume scattering components are highly discriminative for structurally complex scenes. In contrast, using Span alone gives relatively lower perfor- mance, indicating that total backscattering power is insuffi- cient without complementary polarimetric descriptors. 4.5.3. Effect of Scanning Directions Table 7-C compares horizontal, vertical, and bidirec- tional scanning strategies. Single-direction scanning ex- hibits dataset-dependent behavior: vertical scanning per- forms better on Flevoland and San Francisco, while hor- izontal scanning gives a higher A on Oberpfaffenhofen. This indicates that unidirectional scanning is sensitive to the dominant spatial orientation of land-cover structures. By combining row-wise and column-wise dependencies, the bidirectional setting achieves the best OA and kappaร100 on all three datasets, and the best A on Flevoland and San Francisco. Although horizontal scanning obtains a slightly higher A on Oberpfaffenhofen, the bidirectional strategy Zhang et al.: Preprint submitted to ElsevierPage 15 of 20 CV-SSMNet Table 7 Ablation results of different components, prior settings, and scanning strategies on the three datasets. The best and second-best results within each ablation group are highlighted in bold and underlined, respectively. Method FlevolandSan FranciscoOberpfaffenhofen OA (%)A (%)Kappaร100OA (%)A (%)Kappaร100OA (%)A (%)Kappaร100 A) Component ablation Base95.3389.9093.6895.8091.4694.0292.5092.3789.50 Base + CV-SSM95.3191.4693.6695.8290.0594.82 94.6193.6190.35 Base + SFM95.6694.2095.2596.2691.8594.2394.2393.2289.69 Base + CV-SSM + SFM97.5695.1397.0597.0292.8395.0696.0794.3694.72 B) Prior ablation Only ํปโํดโํผ96.5493.2696.31 96.7891.1794.6494.6894.7693.58 Only ํ ํ โํ ํ โํ ํฃ 97.11 94.7295.8496.9291.5995.6095.7795.4295.06 Only Span95.7193.4595.3296.2890.7694.1594.3294.2693.91 All priors97.5695.1397.0597.0292.83 95.0696.0794.3694.72 C) Scanning direction ablation Horizontal scanning94.2392.0793.9495.4590.0593.6895.6295.4293.85 Vertical scanning95.14 92.3494.1596.0290.3394.2694.3794.2193.89 All directions97.5695.1397.0597.0292.8395.0696.0794.3694.72 Table 8 Physical interpretability comparison of different prior-injection strategies. MethodISCโ ISDโ SRโ OA (%) No Prior0.23 1.45 0.67 95.33 Shallow Concat0.41 2.18 0.74 95.66 FiLM Modulation 0.68 3.52 0.89 97.56 Note: SR denotes speckle robustness. provides the most stable overall performance, demonstrat- ing its ability to capture complementary spatial dependen- cies and reduce orientation bias in heterogeneous PolSAR scenes. 4.5.4. Prior-injection strategy comparison. We further compare three prior-injection strategies: no prior, shallow concatenation, and FiLM-based modulation. As shown in Table 8, FiLM modulation achieves the highest intra-class scattering consistency (ISC) and inter-class dis- criminability (ISD). Compared with shallow concatenation, FiLM modulation improves ISC by approximately 2.9ร and ISD by approximately 2.4ร. These results suggest that con- ditional modulation is more effective than input-level fusion for preserving physically meaningful scattering information during feature evolution. Robustness under speckle noise. Finally, we evaluate the stability of different prior-injection strategies under mul- tiplicative speckle perturbation. The prediction consistency of FiLM modulation remains 0.89, compared with 0.74 for shallow concatenation and 0.67 for the no-prior base- line. This indicates that scattering-aware FiLM modulation improves robustness to noise-sensitive statistical variations and helps maintain more stable predictions under degraded PolSAR observations. Table 9 Computational complexity and run-level performance stability of CV-SSMNet. A) Computational complexity comparison MethodParams FLOPs Mem. Infer. Training time (M)(G)(MB) (ms)(h) Base0.077 0.0100 5120.408.5 Base + CV-SSM0.0810.01026400.4810.2 Base + SFM0.0860.10507040.5210.8 CV-SSMNet0.0910.11207120.5511.5 B) Statistical significance test: Base vs. CV-SSMNet Datasetํ 2 ํ-valueSignificant Flevoland16.370.0005Yes San Francisco34.860.00001Yes Oberpfaffenhofen18.480.00002Yes 4.6. Computational Complexity and Performance Stability To further evaluate the efficiency and reliability of the proposed method, the computational cost of different model variants is reported, and a statistical significance test against the baseline is performed. As shown in Table 9-A, reports the computational complexity of different model variants, while Table 9-B presents the statistical significance test between the baseline and CV-SSMNet. CV-SSMNet introduces a small increase in parameters, while the FLOPs increase due to the additional CV-SSM and modulation modules. Nevertheless, the absolute computational cost remains low, with an inference time below 0.6 ms per sample. Although FLOPs increase from 0.0100G to 0.1120G, the absolute computational cost remains low, with an inference time below 0.6 ms per sample.the overall computational overhead remains acceptable considering the consistent improvements in classification accuracy. These results indicate that the proposed long-range modeling and scattering-aware mod- ulation modules improve representation capability without introducing prohibitive computational cost. Zhang et al.: Preprint submitted to ElsevierPage 16 of 20 CV-SSMNet In addition, Table 9-B reports the statistical significance test between the baseline and CV-SSMNet on the three benchmark datasets. The resulting ํ-values are all below 0.001, indicating that the performance gains achieved by CV-SSMNet are stable across five independent runs. 4.7. P-band BIOMASS Application Beyond the three L-band benchmarks, CV-SSMNet is further evaluated on the BIOMASS Level-1b dataset as a large-scale P-band experiment. Compared with the L-band scenes, BIOMASS covers a much larger area and contains five land-cover categories, leading to stronger spatial hetero- geneity and intra-class scattering variability. Since training and testing are performed within each dataset, this experi- ment is not a strict L-to-P cross-frequency transfer evalu- ation. Instead, it assesses the applicability of CV-SSMNet under different frequency-band and scene-scale conditions, especially in the presence of frequency-dependent scattering variations. As shown in Table 10, CV-SSMNet achieves OA values of 97.56%, 97.02%, and 96.07% on Flevoland, San Fran- cisco, and Oberpfaffenhofen, respectively. On the P-band BIOMASS dataset, it obtains an OA of 88.87%, an A of 82.32%, and a kappaร100 of 84.57. Compared with the Base model, CV-SSMNet improves OA, A, and kappaร100 by 2.57, 2.86, and 0.99 points, respectively, indicating that CV-SSM and scattering-aware feature modulation remain effective in the large-scale P-band scenario. The relatively lower accuracy on BIOMASS can be mainly attributed to three factors: 1. P-band SAR has greater penetration depth than L- band SAR, resulting in different scattering responses and polarimetric statistics. 2. The BIOMASS scene covers a larger geographical area, causing stronger spatial non-stationarity and intra-class variation. 3. The spatial block-based partition reduces spatial leak- age but increases the difficulty of cross-region classi- fication. Despite these challenges, CV-SSMNet adapts to P-band observations through shared complex-valued representation learning and prior-conditioned modulation. The coherency matrix ํ 3 provides a unified input form, dataset-specific channel normalization reduces band-dependent magnitude shifts, and SFM recalibrates intermediate features using the scattering priors ํป, ํด, ํผ, ํ ํ , ํ ํ , ํ ํฃ , and Span. The ablation results in Fig. 11 further demonstrate the effectiveness of CV-SSM and SFM on the large-scale BIOMASS scene. 4.8. Feature Visualization Analysis 4.8.1. Feature visualization To evaluate the discriminative power of the learned features, t-SNE is employed to visualize the learned feature distributions. Specifically, in Fig. 12 (a)-(c) , the features extracted by the Base model are scattered across the three datasets and exhibit substantial inter-class overlap. The inter- class mixing is particularly prominent in the areas marked by red ellipses. This indicates that the features extracted by the Base model have limited discriminative power, making it difficult to form clear and stable intra-class clustering and inter-class separation in the high-dimensional space, especially for land cover categories with similar scattering characteristics or complex structures, which easily leads to confusion. This insufficient feature representation also ex- plains its limited performance in cross-regional or complex scenarios. In contrast, the CV-SSMNet feature visualiza- tion results shown in Fig. 12 (d)-(f) demonstrate significant improvement. On all three datasets, samples of the same class exhibit a more compact clustering structure in the low- dimensional embedded space, and the separation between different classes is significantly increased. The areas that were severely overlapping in the Base model are effec- tively separated. This phenomenon indicates that the features learned by CV-SSMNet have stronger discriminative power and consistency. 4.8.2. Physics-Aware Activation and Layer-wise Feature Analysis To analyze how scattering priors and long-range model- ing influence the learned representation, Fig. 13 visualizes the activation responses of the seven physical prior channels and their fused output. Different priors show spatial patterns broadly consistent with PolSAR scattering mechanisms. The ํป channel responds strongly to regions with high scattering randomness, while ํ ํ , ํ ํ , and ํ ํฃ highlight surface-like, double-bounce, and volume-scattering areas, respectively. In contrast, ํด, ํผ, and Span provide smoother complementary contextual responses. The fused activation map integrates these scattering-dependent cues, indicating that SFM uses physical priors as conditional modulation signals rather than simple concatenated features. Fig. 14 further compares intermediate feature maps of the baseline and CV-SSMNet across representative stages, including shallow convolution, middle convolution, deep convolution with CV-SSM, and final logits. Both mod- els capture basic local structures at shallow layers, but CV-SSMNet produces stronger boundary responses and more uniform activations over heterogeneous regions. At deeper layers, the baseline tends to focus on central areas with weaker peripheral responses, whereas CV-SSMNet shows broader spatial coverage, confirming that complex- valued state-space modeling captures long-range depen- dencies more effectively. The output maps also exhibit sharper boundaries and more separable class responses, demonstrating improved regional consistency and boundary discriminability over the CV-ASDF2Net baseline. 5. Discussion and Conclusion This paper presents CV-SSMNet, a physics-aware complex- valued state-space network for PolSAR image classification. Unlike methods that use polarimetric descriptors only as shallow auxiliary inputs, CV-SSMNet encodes ํป,ํด,ํผ,ํ ํ , ํ ํ ,ํ ํฃ , Span as FiLM-style modulation signals to guide the evolution of complex-valued features. By integrating Zhang et al.: Preprint submitted to ElsevierPage 17 of 20 CV-SSMNet Table 10 Multi-dataset evaluation on L-band benchmarks and the P-band BIOMASS dataset. DatasetModel Data descriptionPerformance BandClassesImage sizeOA (%)A (%)Kappaร100 FlevolandCV-SSMNetL15750 ร 102497.5695.1397.05 San FranciscoCV-SSMNetL5900 ร 102497.0292.8395.06 OberpfaffenhofenCV-SSMNetL31300 ร 120096.0794.3694.72 BIOMASSCV-SSMNetP521172 ร 150988.8782.3284.57 BIOMASSBaseP521172 ร 150986.30 79.4683.58 Base+CV-SSMBase+SFMBase + CV-SSM + SFMBase Figure 11: Visualization of ablation classification results on the BIOMASS dataset. (a) Base. (b) Base + CV-SSM. (c) Base + SFM. (d) CV-SSMNet. The progressive incorporation of CV-SSM and SFM improves regional consistency, boundary preservation, and misclassification suppression. (a) (b) (c) (d) (e)(f) Figure 12: t-SNE feature visualization on three benchmark datasets. The first, second, and third columns correspond to Flevoland, Oberpfaffenhofen, and San Francisco, respectively. The top row shows the Base model, and the bottom row shows CV-SSMNet. CV-SSMNet yields more compact and separable feature distributions across different scenes. complex-valued convolution, CV-SSM-based long-range dependency modeling, and scattering-aware feature mod- ulation, the proposed framework learns physically guided representations from local scattering structures to global spatial context. The proposed physics-aware mechanism follows a soft- constraint strategy rather than a hard physics-constrained formulation. The seven polarimetric priors are not directly imposed through first-principles electromagnetic equations, such as Maxwell equations or radiative transfer models. Instead, they serve as physically meaningful conditioning variables for data-driven feature evolution. This design pre- serves the flexibility of deep learning while introducing Statistical Features (Entropy) (Anisotropy) Alpha (Surface) (Double) (Volume) Combined Scattering Mechanisms Activation Intensity H A s P d P v P Span Figure 13: Prior-conditioned activation responses of the seven polarimetric scattering descriptors. The responses are broadly consistent with typical PolSAR scattering mechanisms, indicat- ing that SFM learns physically meaningful modulation patterns rather than purely data-driven attention weights. scattering-mechanism awareness into intermediate feature modulation, making the framework adaptable to heteroge- neous land-cover scenes and extensible to other polarimetric descriptors. Experiments on three L-band benchmark datasets and an additional P-band BIOMASS dataset demonstrate that CV- SSMNet achieves competitive OA, A, and kappa values. Zhang et al.: Preprint submitted to ElsevierPage 18 of 20 CV-SSMNet Input PolSAR PatchShallow ConvMid ConvDeep Conv+CV-SSMOutput Base Ours Figure 14: Feature maps at different layers of the baseline (top) and the proposed CV-SSMNet (bottom). Columns: input patch, shallow Conv, mid Conv, deep Conv + CV-SSM, and output. Ablation results show that CV-SSM and SFM provide com- plementary gains: CV-SSM improves long-range contextual modeling, while SFM enhances scattering-aware feature re- calibration. Visualization results further indicate that the learned prior-conditioned responses are broadly consistent with typical PolSAR scattering mechanisms, supporting the physical interpretability of the proposed modulation strat- egy. Several limitations remain. The lower performance on BIOMASS suggests that frequency-dependent scattering variations are not fully modeled, especially wavelength- dependent penetration depth, volume scattering, and di- electric response differences between L-band and P-band observations. In addition, the fixed 13 ร 13 patch size may limit scale adaptability, and the CV-SSM and SFM modules introduce extra computational cost, although the absolute inference time remains low. Decomposition-based priors may also be affected by speckle noise, low-SNR observations, and decomposition errors. Future work will explore band-aware prior encoding, adaptive multi-scale patch sampling, efficient CV-SSM de- sign, and uncertainty-aware prior gating. Another direction is to incorporate explicit electromagnetic scattering con- straints or differentiable scattering models, moving from soft physics-guided modulation toward stronger physics- constrained GeoAI representation learning. The code is available at https://github.com/Lucia-210/CV-SSMNet. Acknowledgments The authors gratefully acknowledge the providers of the datasets and the developers of the PolSARpro software. This study was conducted within the framework of the Dragon- 6 Cooperation Programme between the European Space Agency (ESA) and the Ministry of Science and Technology (MOST) of China. References Alkhatib, M.Q., 2024. Polsar image classification using a hybrid complex- valued network (hybridcvnet). IEEE Geosci. Remote Sens. Lett. 21, 1โ5. Alkhatib, M.Q., Zitouni, M.S., Al-Saad, M., et al., 2025. Polsar image clas- sification using shallow to deep feature fusion network with complex- valued attention. Scientific Reports 15, 24315. An, W., Lin, M., 2021. Generalized polarimetric entropy: Polarimetric in- formation quantitative analyses of model-based incoherent polarimetric decomposition. IEEE Trans. Geosci. Remote Sens. 59, 2041โ2057. Cheng, J., Xiang, D., Yin, Q., Zhang, F., 2022. A novel crop classification method based on the tensor-gcn for time-series polsar data. IEEE Trans. Geosci. Remote Sens. 60, 1โ14. Art. no. 5237614. Dai, L.Y., Li, M.D., Chen, S.W., 2025. Polsar image super-resolution with polarimetric degradation modulation network. IEEE Trans. Aerosp. Electron. Syst. 61, 17938โ17954. Duan, D., Wang, Y., Zhang, Y., 2024. The critical role of cross-polarized backscatter in understanding l-band polsar data in forested and urban environments. Remote Sens. Environ. 311. Art. no. 114265. Freeman, A., Durden, S.L., 1998. A three-component scattering model for polarimetric sar data. IEEE Trans. Geosci. Remote Sens. 36, 963โ973. Gao, F., Jin, X., Zhou, X., Dong, J., Du, Q., 2025. Msfmamba: Multi-scale feature fusion state space model for multi-source remote sensing image classification. IEEE Trans. Geosci. Remote Sens. Early Access. Geng, J., Dong, L., Zhang, Y., Jiang, W., 2025. Masked auto-encoding and scatter-decoupling transformer for polarimetric sar image classification. Pattern Recognit. 166. Art. no. 111660. Ghazvinizadeh, A.H., Imani, M., Ghassemian, H., 2023. Residual network based on entropy-anisotropy-alpha target decomposition for polarimetric sar image classification. Earth Sci. Informat. 16, 357โ366. doi:10.1007/ s12145-023-00944-6. Gu, A., Dao, T., 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752 . Gu, A., Goel, K., Rรฉ, C., 2021. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396 . Guo, Z., Zhang, H., Ge, J., Shi, Z., Xu, L., Tang, Y., Wu, F., Wang, Y., Wang, C., 2024. Built-up area extraction in polsar imagery using real-complex polarimetric features and feature fusion classification network. Int. J. Appl. Earth Obs. Geoinf. 134. Art. no. 104144. Han, W., Fu, H., Zhu, J., Zhang, S., Xie, Q., Hu, J., 2023. A polarimetric projection-based scattering characteristics extraction tool and its appli- cation to polsar image classification. ISPRS J. Photogramm. Remote Sens. 202, 314โ333. Hochstuhl, S., Pfeffer, N., Thiele, A., Hinz, S., Scheiber, R., Reigber, A., 2023. Pol-insar-island: A benchmark dataset for multi-frequency pol- insar data land cover classification. ISPRS Open J. Photogramm. Remote Sens. 3. Art. no. 100047. Hu, C., Wang, Y., Sun, X., Quan, S., Xiang, D., 2023. Model-based polarimetric target decomposition with power redistribution for urban areas. IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. 16, 8795โ 8808. Hu, T., Li, W., Qin, X., et al., 2019. Terrain classification of polarimetric synthetic aperture radar images based on deep learning and conditional random field model. Journal of Radars 8, 471โ478. doi:10.12000/ JR18065. Hua, W., Wang, Y., Yang, Z., 2026. Knowledge and data co-driven deep learning model for polsar image classification. Results Eng. 29. Art. no. 108947. Imani, M., 2022. Low frequency and radarโs physical based features for improvement of convolutional neural networks for polsar image classification. Egypt. J. Remote Sens. Space Sci. 25, 55โ62. doi:10. 1016/j.ejrs.2021.12.007. Imani, M., 2025. Attention based network for fusion of polarimetric and contextual features for polarimetric synthetic aperture radar image classification. Eng. Appl. Artif. Intell. 139. Art. no. 109665. Jamali, A., Roy, S.K., Bhattacharya, A., Ghamisi, P., 2023. Local window attention transformer for polarimetric sar image classification. IEEE Geosci. Remote Sens. Lett. 20, 1โ5. Kuang, Z., Bi, H., Li, F., Xu, C., 2025a. Ecp-mamba: An efficient multiscale self-supervised contrastive learning method with state space model for polsar image classification. IEEE Trans. Geosci. Remote Sens. 63, 1โ18. Kuang, Z., Bi, H., Li, F., et al., 2024. Polarimetry-inspired contrastive learning for class-imbalanced polsar image classification. IEEE Trans. Geosci. Remote Sens. 62, 1โ19. Zhang et al.: Preprint submitted to ElsevierPage 19 of 20 CV-SSMNet Kuang, Z., Liu, K., Bi, H., Li, F., 2025b. Polsar image classification with complex-valued diffusion model as representation learners. IEEE Trans. Aerosp. Electron. Syst. 61, 1โ21. doi:10.1109/TAES.2025.3572877. Lee, J.S., et al., 1999. Unsupervised classification using polarimetric decomposition and the complex wishart classifier. IEEE Trans. Geosci. Remote Sens. 37, 2249โ2258. Li, W., Xia, H., Xi, B., Wang, Y., He, Y., Han, Y., 2025. Ssl-mbc: Self- supervised learning with multibranch consistency for few-shot polsar image classification. IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. 18, 4696โ4710. doi:10.1109/JSTARS.2025.3528529. Liu, M., Jiao, L., Liu, X., Li, L., Liu, F., Yang, S., Guo, Y., Chen, P., 2024a. C2n2: Complex-valued contourlet neural network. IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. 17, 4478โ4491. Liu, X., Jiao, L., Liu, F., Zhang, D., Tang, X., 2022. Polsf: Polsar image datasets on san francisco, in: Proc. Int. Conf. Intell. Sci., Springer, Cham, Switzerland. p. 214โ219. Liu, Y., Tian, Y., Zhao, Y., et al., 2024b. Vmamba: Visual state space model, in: Adv. Neural Inf. Process. Syst., p. 103031โ103063. Ni, J., Xiang, D., Lin, Z., et al., 2022. Dnn-based polsar image classification on noisy labels. IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. 15, 3697โ3713. Ni, J., Zhang, F., Yin, Q., Zhou, Y., Li, H.C., Hong, W., 2021. Random neighbor pixel-block-based deep recurrent learning for polarimetric sar image classification. IEEE Trans. Geosci. Remote Sens. 59, 7557โ7569. Nie, X., Qiao, H., Zhang, B., 2015. A variational model for polsar data speckle reduction based on the wishart distribution. IEEE Trans. Image Process. 24, 1209โ1222. Quegan, S., Le Toan, T., Chave, J., Dall, J., Exbrayat, J.F., Minh, D.H.T., Lomas, M., DโAlessandro, M.M., Paillou, P., Papathanassiou, K., Rocca, F., Saatchi, S., Scipal, K., Shugart, H., Smallman, T.L., Soja, M.W.J., Tebaldini, S., Ulander, L., Villard, L., Williams, M., 2019. The european space agency biomass mission: Measuring forest above-ground biomass from space. Remote Sens. Environ. 227, 44โ60. Shi, J., Nie, M., Ji, S., Shi, C., Liu, H., Jin, H., 2023. Polarimetric synthetic aperture radar image classification based on double-channel convolution network and edge-preserving markov random field. Remote Sens. 15. doi:10.3390/rs15235458. art. no. 5458. Tan, X., Li, M., Zhang, P., Wu, Y., Song, W., 2019. Complex-valued 3- d convolutional neural network for polsar image classification. IEEE Geosci. Remote Sens. Lett. 17, 1022โ1026. Vissers, M.A.M., van der Sanden, J.J., 1992. Groundtruth collection for the JPL-SAR and ERS-1 campaign in Flevoland and the Veluwe (NL) 1991. Technical Report 31. Landbouwuniversiteit Wageningen. Wageningen, The Netherlands. 68 p. Wang, S., Sun, Z., Bian, T., Guo, Y., Dai, L., Guo, Y., Jiao, L., 2025a. Cdfnet: Cross-domain feature fusion network for polsar terrain classi- fication. IEEE Trans. Geosci. Remote Sens. 63, 1โ15. Wang, Z., Zhao, L., Wang, Y., et al., 2025b. Air-polsar-seg-2.0: Polarimetric sar ground terrain classification dataset for large-scale complex scenes. Journal of Radars 14, 353โ365. doi:10.12000/JR24237. Wu, Q., Hou, B., Wen, Z., Jiao, L., 2019. Variational learning of mixture wishart model for polsar image classification. IEEE Trans. Geosci. Remote Sens. 57, 141โ154. Xiao, D., Liu, C., 2019. Polsar terrain classification based on fine-tuned dilated group-cross convolution neural network. Journal of Radars 8, 479โ489. doi:10.12000/JR19039. Yin, Q., Lin, Z., Hu, W., Lรณpez-Martรญnez, C., Ni, J., Zhang, F., 2023. Crop classification of multitemporal polsar based on 3-d attention module with vit. IEEE Geosci. Remote Sens. Lett. 20, 1โ5. Zhang, L., Zhang, S., Zou, B., Dong, H., 2022. Unsupervised deep representation learning and few-shot classification of polsar images. IEEE Trans. Geosci. Remote Sens. 60. Art. no. 5100316. Zhang, S., Cui, L., Dong, Z., An, W., 2024a. A deep learning classification scheme for polsar image based on polarimetric features. Remote Sens. 16. doi:10.3390/rs16101676. art. no. 1676. Zhang, Y., Wang, W., Guo, Z., Li, N., 2024b. Enhanced pga for dual- polarized isar imaging by exploiting cloudeโpottier decomposition. IEEE Geosci. Remote Sens. Lett. 21, 1โ5. Zhang, Z., Wang, H., Xu, F., Jin, Y.Q., 2017. Complex-valued convolutional neural network and its application in polarimetric sar image classifica- tion. IEEE Trans. Geosci. Remote Sens. 55, 7177โ7188. Zhu, L., Liao, B., Zhang, Q., et al., 2024. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv preprint arXiv:2401.09417 . Zhuang, D., Zhang, L., Zou, B., 2024. Model-based polarimetric sar target decomposition: A scheme to introduce repeat-pass polinsar coherence. IEEE Trans. Geosci. Remote Sens. 62, 1โ16. Zhang et al.: Preprint submitted to ElsevierPage 20 of 20