Paper deep dive
CDGC-Net: 3D Medical Image Segmentation with Cooperative Dual-Scale Self-Attention and Grouped Channel Modeling
Zheyang Jing, Qin Lu, Jianwang Li, Yujie Yang, Chen Yi, Shaofeng Jiang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/12/2026, 2:01:50 AM
Summary
The paper introduces CDGC-Net, a 3D medical image segmentation network designed to address semantic mismatches and redundant representations in existing methods. It combines Cooperative Dual-Scale Self-Attention (CDSA) for capturing both local and global spatial contexts and Grouped Hierarchical Channel Attention (GHCA) for modeling channel dependencies. The model achieves state-of-the-art performance on Synapse, ACDC, BraTS, and LA datasets with reduced computational complexity compared to UNETR++.
Entities (10)
Relation Signals (10)
CDGC-Net → containsmodule → CDSA
confidence 95% · Within each CDGC block, Cooperative Dual-Scale Self-Attention (CDSA) assigns attention heads...
CDGC-Net → containsmodule → GHCA
confidence 95% · Their outputs are concatenated... and directly passed to Grouped Hierarchical Channel Attention (GHCA).
CDGC-Net → evaluatedon → Synapse
confidence 95% · On the Synapse, ACDC, BraTS, and LA datasets, CDGC-Net achieved mean DSC values...
CDGC-Net → evaluatedon → ACDC
confidence 95% · On the Synapse, ACDC, BraTS, and LA datasets, CDGC-Net achieved mean DSC values...
CDGC-Net → evaluatedon → BraTS
confidence 95% · On the Synapse, ACDC, BraTS, and LA datasets, CDGC-Net achieved mean DSC values...
CDGC-Net → evaluatedon → LA
confidence 95% · On the Synapse, ACDC, BraTS, and LA datasets, CDGC-Net achieved mean DSC values...
CDSA → captures → global-context
confidence 90% · The two branches capture fine spatial details and long-range anatomical context...
CDSA → captures → local-spatial-details
confidence 90% · The two branches capture fine spatial details and long-range anatomical context...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing methods often model global and local features in separate modules or feature levels and perform channel recalibration independently. This may cause semantic mismatch between global context and local boundaries, insufficient channel relationship modeling, weak spatial-channel interaction, and redundant representations. We propose CDGC-Net, a 3D medical image segmentation network that combines cooperative dual-scale spatial attention with grouped hierarchical channel modeling. With-in each CDGC block, Cooperative Dual-Scale Self-Attention (CDSA) assigns attention heads to parallel local-window and global-sparse branches. The two branches capture fine spatial details and long-range anatomical context at the same feature level. Their outputs are concatenated into an $N\times C$ spatial representation and directly passed to Grouped Hierarchical Channel Attention (GHCA). GHCA organizes the channels into $r$ groups and models both within-group and cross-group dependencies. CDSA and GHCA reuse a shared key projection to maintain a consistent feature reference. Residual feature alignment subsequently integrates the refined features with the original representation. On the Synapse, ACDC, BraTS, and LA datasets, CDGC-Net achieved mean DSC values of 86.96\%, 92.91\%, 82.56\%, and 93.52\%, respectively, exceeding the next-highest reported values by 0.39, 0.47, 0.17, and 0.32 percentage points. CDGC-Net contains 25.83M parameters and 28.62G FLOPs for an input size of $64\times128\times128$, reducing these quantities by 39.87\% and 40.30\%, respectively, relative to UNETR++. These results indicate a favorable trade-off between segmentation accuracy and computational complexity.
Tags
Links
- Source: https://arxiv.org/abs/2608.08575v1
- Canonical: https://arxiv.org/abs/2608.08575v1
Trouble viewing inline? Open PDF directly →
Full Text
34,861 characters extracted from source content.
Expand or collapse full text
11institutetext: Nanchang Hangkong University, Nanchang, China 11email: 1255813468@q.com, luq0715@163.com, 2037405746@q.com, 1796104091@q.com, 470835297@q.com, jsphone@163.com CDGC-Net: 3D Medical Image Segmentation with Cooperative Dual-Scale Self-Attention and Grouped Channel Modeling Zheyang Jing Qin Lu Jianwang Li Yujie Yang Chen Yi Shaofeng Jiang Abstract Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing methods often model global and local features in separate modules or feature levels and perform channel recalibration independently. This may cause semantic mismatch between global context and local boundaries, insufficient channel relationship modeling, weak spatial-channel interaction, and redundant representations. We propose CDGC-Net, a 3D medical image segmentation network that combines cooperative dual-scale spatial attention with grouped hierarchical channel modeling. With-in each CDGC block, Cooperative Dual-Scale Self-Attention (CDSA) assigns attention heads to parallel local-window and global-sparse branches. The two branches capture fine spatial details and long-range anatomical context at the same feature level. Their outputs are concatenated into an N×CN× C spatial representation and directly passed to Grouped Hierarchical Channel Attention (GHCA). GHCA organizes the channels into r groups and models both within-group and cross-group dependencies. CDSA and GHCA reuse a shared key projection to maintain a consistent feature reference. Residual feature alignment subsequently integrates the refined features with the original representation. On the Synapse, ACDC, BraTS, and LA datasets, CDGC-Net achieved mean DSC values of 86.96%, 92.91%, 82.56%, and 93.52%, respectively, exceeding the next-highest reported values by 0.39, 0.47, 0.17, and 0.32 percentage points. CDGC-Net contains 25.83M parameters and 28.62G FLOPs for an input size of 64×128×12864× 128× 128, reducing these quantities by 39.87% and 40.30%, respectively, relative to UNETR++. These results indicate a favorable trade-off between segmentation accuracy and computational complexity. 1 Introduction Volumetric medical image segmentation is a core task in medical image processing. It has wide applications in tumor assessment, quantitative organ analysis, surgical navigation, and treatment response tracking [7, 9]. These tasks require models to generate accurate 3D segmentation masks while preserving both global anatomical context and fine boundary details. Therefore, they place high demands on spatial representation and feature discrimination. Following the success of CNN-based methods, Transformers have been widely introduced into 3D medical image segmentation. Through self-attention, they can effectively model long-range dependencies and improve segmentation performance [7, 23]. However, standard self-attention has a quadratic complexity of O(n2)O(n^2). When processing 3D volumetric data, it brings high memory usage and training cost. To reduce the computational burden, existing methods usually adopt window attention, local self-attention, deformable attention, or CNN-Transformer hybrid designs [4, 15, 2, 18]. For example, window-based or local attention methods reduce complexity by limiting the attention range. CoTr [18] uses deformable attention to model key spatial positions. TransUNet [3] and UNETR++ [13] combine local feature extraction with global context modeling to enhance feature representation. Although these methods improve the balance between long-range dependency modeling and computational efficiency, they often extract global context and local details from different modules, branches, or feature levels. These features differ in spatial resolution, receptive field, and semantic abstraction. Without sufficient cross-scale interaction and alignment, feature fusion may produce a mismatch between global semantics and local boundary details. In addition, when multiple attention branches lack explicit interaction or constraints, their responses may overlap. This can reduce feature complementarity and increase the risk of redundant representations. Besides scale differences in spatial features, semantic representation in the channel dimension also affects segmentation performance. Existing channel attention methods, such as SE [8], CBAM [17], and ECA [14], mainly enhance feature representation by recalibrating channel responses. However, these methods usually generate channel weights through global statistical descriptions or local cross-channel interactions. They often do not explicitly model hierarchical relationships among different channel subspaces. For medical images, responses from background, normal tissues, and lesion regions often overlap in the channel dimension. Therefore, a single channel recalibration strategy may fail to emphasize both high-response salient regions and low-response fine-grained regions. As a result, some useful features may be weakened. Furthermore, many segmentation networks introduce both spatial attention and channel attention to enhance feature representation [17, 22, 13, 21]. For example, CBAM-UNet [21] adopts a serial weighting strategy with channel attention and spatial attention. TransFuse [22] uses the BiFusion module to fuse attention features from different branches. The EPA module in UNETR++ [13] also improves feature representation through paired spatial and channel attention. However, existing methods usually still rely on single-scale attention for spatial modeling. They also rarely model grouped hierarchical relationships in channel modeling. Meanwhile, spatial and channel features are often fused by addition, concatenation, or convolution. Their correlation modeling and feature alignment still have room for improvement. Therefore, how to cooperatively model dual-scale spatial information and structured channel relationships in a unified module, and how to build effective alignment between spatial and channel features, remain important problems in 3D medical image segmentation. To address these problems, we propose CDGC-Net, a cooperative dual-scale self-attention and grouped channel modeling network for 3D medical image segmentation. CDGC-Net uses the U-shaped UNETR++ [13] backbone and replaces each computation block with the proposed CDGC block. This block integrates cooperative dual-scale spatial modeling, grouped hierarchical channel modeling, and residual feature alignment to enhance the interaction between spatial and channel features. The main contributions of this work are summarized as follows: (1) We design a Cooperative Dual-Scale Self-Attention (CDSA) mechanism. It combines global sparse attention and local window attention to capture coarse-scale global context and fine-scale local structures within a unified spatial attention module. (2) We introduce a Grouped Hierarchical Channel Attention (GHCA) module. It models intra-group and inter-group channel relationships to improve channel discrimination and feature complementarity. (3) We connect CDSA and GHCA through shared key projection, spatial-guided channel query, and residual feature alignment. This forms a continuous feature refinement process and promotes effective interaction between dual-scale spatial features and grouped channel features. Experiments on Synapse, ACDC, BraTS, and LA show that CDGC-Net achieved the highest mean DSC among the compared methods under matched evaluation protocols. 2 Methods 2.1 Overall Architecture Figure 1: Overall architecture of the proposed CDGC-Net. As shown in Fig. 1, CDGC-Net adopts the U-shaped encoder-decoder architecture of UNETR++ [13]. Given a 3D medical image, the network first maps the input volume into embedded features through a patch embedding layer. Meanwhile, a shallow convolution branch processes the original input and preserves high-resolution spatial information for the final prediction. During the encoding stage, the network extracts multi-scale volumetric features through four hierarchical stages. Each encoder stage stacks three CDGC blocks as the main feature extraction units. As the network goes deeper, the spatial resolution gradually decreases, while the channel dimension gradually increases. After the encoder obtains hierarchical features, the decoder gradually restores the spatial resolution through upsampling. At each decoding stage, the upsampled feature is fused with the corresponding encoder feature through a skip connection. The stacked CDGC blocks then further refine the fused feature. Within each CDGC block, positional embedding and layer normalization produce the normalized feature X¯ X. Linear projections generate the spatial query QsQ_s, shared key K, and spatial value VsV_s. CDSA produces the spatial representation X^s X_s, which conditions the channel query QcQ_c in GHCA. GHCA reuses K with the channel value VcV_c, after which residual feature alignment and convolutional refinement produce the block output. After the final decoding stage, the network fuses the decoder output with the high-resolution feature from the shallow convolution branch. Finally, a prediction head with 3×3×33× 3× 3 and 1×1×11× 1× 1 convolutions maps the fused feature to the target classes and uses a softmax operation to generate the final segmentation probability map. 2.2 Cooperative Dual-Scale Self-Attention (CDSA) CDSA captures local and global spatial dependencies at the same feature level. Given X∈ℝH×W×D×CX ^H× W× D× C, CDSA first flattens the spatial dimensions to obtain Xf∈ℝN×CX_f ^N× C, where N=HWDN=HWD. After positional embedding and layer normalization, the resulting feature X¯ X is linearly projected as Qs=X¯WsQ,K=X¯WK,Vs=X¯WsV.Q_s= XW_s^Q, K= XW^K, V_s= XW_s^V. (1) The key K is computed once and subsequently reused by GHCA. For CDSA, QsQ_s, K, and VsV_s are partitioned into HlH_l local heads and HgH_g global heads. Here, Hl+Hg=HaH_l+H_g=H_a, where HaH_a is the total number of attention heads and dh=C/Had_h=C/H_a is the dimension of each head. The two head sets are routed to window attention and sparse global attention, respectively. For each local head, the query, key, and value are rearranged as Ql,Kl,Vl∈ℝWin×(N/Win)×dhQ_l,K_l,V_l ^Win×(N/Win)× d_h, corresponding to WinWin non-overlapping windows with N/WinN/Win tokens each. Local attention is then computed as X^l=Softmax(QlKlTdh+B)Vl. X_l=Softmax ( Q_lK_l^T d_h+B )V_l. (2) Here, dhd_h is the dimension of each attention head. For a 3D window of size hw×w×dwh_w× w_w× d_w, let M=hwwdwM=h_ww_wd_w denote its number of tokens. The learnable relative position bias B∈ℝM×MB ^M× M encodes the relative spatial displacement between every pair of tokens within the window and is added independently to each local head. For each global head, the key and value are compressed along the token dimension using learnable projection matrices EK,EV∈ℝP×NE_K,E_V ^P× N: K^g=EKKg,V^g=EVVg, K_g=E_KK_g, V_g=E_VV_g, (3) where Kg,Vg∈ℝN×dhK_g,V_g ^N× d_h and K^g,V^g∈ℝP×dh K_g, V_g ^P× d_h, with P≪NP N. Sparse global attention is then computed as: X^g=Softmax(QgK^gTdh)V^g. X_g=Softmax ( Q_g K_g^T d_h ) V_g. (4) The reduced key and value preserve long-range interactions while lowering the spatial attention cost. After merging the heads, the local and global branch outputs have dimensions X^l∈ℝN×Hldh X_l ^N× H_ld_h and X^g∈ℝN×Hgdh X_g ^N× H_gd_h, respectively. They are then concatenated along the channel dimension: X^s=Concat(X^l,X^g)∈ℝN×C. X_s=Concat ( X_l, X_g ) ^N× C. (5) Thus, CDSA integrates local-window details and sparse global context without requiring features from different network levels. 2.3 Grouped Hierarchical Channel Attention (GHCA) GHCA receives X^s∈ℝN×C X_s ^N× C and models channel correlations through within-group and cross-group attention. The spatially conditioned query and channel value are defined as Qc=(X¯+X^s)WcQQ_c=( X+ X_s)W_c^Q and Vc=X¯WcVV_c= XW_c^V, respectively. The key K=X¯WKK= XW^K is reused from CDSA without an additional key projection. Before channel grouping, the representations from all HaH_a attention heads are concatenated along the channel dimension, restoring Qc,K,Vc∈ℝN×CQ_c,K,V_c ^N× C. Each tensor is then divided into r channel groups, where Qci,Ki,Vci∈ℝN×dgQ_c^i,K^i,V_c^i ^N× d_g, dg=C/rd_g=C/r, and i=1,…,ri=1,…,r. Within-group channel attention is computed as Gi=VciSoftmax((Qci)TKidg),i=1,…,r.G_i=V_c^iSoftmax ( (Q_c^i)^TK^i d_g ), i=1,…,r. (6) Here, (Qci)TKi∈ℝdg×dg(Q_c^i)^TK^i ^d_g× d_g represents the channel affinity within the i-th group, and Gi∈ℝN×dgG_i ^N× d_g is the corresponding group-enhanced feature. The group outputs are concatenated as G=Concat(G1,…,Gr)∈ℝN×CG=Concat(G_1,…,G_r) ^N× C. For cross-group modeling, G is reshaped along the channel dimension into r groups. Without additional linear projections, the query, key, and value are constructed as QG=KG=Reshape(NormN(G))Q_G=K_G=Reshape(Norm_N(G)) and VG=Reshape(G)V_G=Reshape(G), where QG,KG,VG∈ℝN×r×dgQ_G,K_G,V_G ^N× r× d_g. Here, NormN(⋅)Norm_N(·) denotes ℓ2 _2 normalization along the token dimension. For each spatial token n, cross-group attention is computed as AG(n)=Softmax(QG(n)(KG(n))Tdg),G~(n)=AG(n)VG(n).A_G^(n)=Softmax ( Q_G^(n) (K_G^(n) )^T d_g ), G^(n)=A_G^(n)V_G^(n). (7) where AG(n)∈ℝr×rA_G^(n) ^r× r and G~(n)∈ℝr×dg G^(n) ^r× d_g. The softmax operation is applied along the last group dimension. The outputs for all spatial tokens are reshaped from N×r×dgN× r× d_g back to N×CN× C, producing the channel-enhanced feature X^c X_c. The off-diagonal entries of AG(n)A_G^(n) propagate information across different channel groups. 2.4 Residual Feature Alignment (RFA) After GHCA produces X^c X_c, residual feature alignment applies a channel-wise scaling vector initialized as γ(0)=10−6Cγ^(0)=10^-61_C. The aligned feature is defined as Xa=X¯+γ⊙X^c,X_a= X+γ X_c, (8) where ⊙ denotes channel-wise multiplication with broadcasting over the spatial tokens. This residual pathway retains the normalized input representation while controlling the magnitude of the channel correction. The aligned feature is further refined through residual convolution: Xout=Xa+Conv1×1×1(Conv3×3×3(Xa)).X_out=X_a+Conv_1× 1× 1 (Conv_3× 3× 3(X_a) ). (9) 2.5 Loss Function Following UNETR++ [13], UNETR [7], and nnFormer [23], we use a combination of soft Dice loss and cross-entropy loss for supervision. The soft Dice loss measures the overlap between the prediction and the ground truth, while the cross-entropy loss encourages voxel-wise classification accuracy. The total loss is defined as: ℒ(Y,P)=1−1I∑i=1I2∑v=1VYv,iPv,i+ϵ∑v=1VYv,i2+∑v=1VPv,i2+ϵ−1V∑v=1V∑i=1IYv,ilog(Pv,i),L(Y,P)=1- 1I _i=1^I 2 _v=1^VY_v,iP_v,i+ε _v=1^VY_v,i^2+ _v=1^VP_v,i^2+ε- 1V _v=1^V _i=1^IY_v,i (P_v,i), (10) where I denotes the number of classes and V denotes the number of voxels. Yv,iY_v,i and Pv,iP_v,i denote the ground-truth label and predicted probability for class i at voxel v, respectively. ϵε is a small constant used to avoid numerical instability. 3 Experiments 3.1 Experimental Configuration Datasets. We conduct experiments on four public 3D medical image segmentation datasets, including Synapse Multi-organ CT Segmentation [10], ACDC [1], BraTS [11], and Left Atrium (LA) Segmentation [19]. 1) The Synapse dataset contains 30 abdominal CT scans with annotations of eight organs, including the spleen, right kidney, left kidney, gallbladder, liver, stomach, aorta, and pancreas. Following previous methods [3], we use 18 cases for training and 12 cases for testing. 2) The ACDC dataset consists of cardiac MRI scans from 100 patients. It provides annotations for three cardiac structures: the right ventricle (RV), myocardium (MYO), and left ventricle (LV). Following nnFormer [23], we split the dataset into 70 training cases, 10 validation cases, and 20 testing cases. 3) The BraTS dataset contains 484 brain tumor MRI cases. Each case includes four modalities: FLAIR, T1w, T1gd, and T2w. The segmentation targets include whole tumor (WT), enhancing tumor (ET), and tumor core (TC). We divide the dataset into training, validation, and testing sets with a ratio of 80:5:15. 4) The LA dataset contains 100 cardiac MRI scans with left atrium annotations. We use 70 cases for training, 10 cases for validation, and 20 cases for testing. Evaluation Metrics. We use the Dice Similarity Coefficient (DSC) and the 95% Hausdorff Distance (HD95) as evaluation metrics. DSC measures the overlap between the predicted segmentation and the ground truth, while HD95 evaluates boundary errors between them. A higher DSC and a lower HD95 indicate better segmentation performance. Implementation Details. We implemented CDGC-Net using the PyTorch-based MONAI framework. All experiments were conducted on a single NVIDIA GeForce RTX 2080 Ti GPU with 11 GB of memory. The model was trained for 1,000 epochs using the SGD optimizer with an initial learning rate of 0.01, a momentum of 0.99, a weight decay of 3×10−53× 10^-5, and a batch size of 2. For all datasets, we used the same input sizes and preprocessing strategies as the compared methods, without using additional training data. Sliding window inference with 50% overlap was adopted during testing, and all reported results were obtained from a single model without ensemble strategies. For Synapse, the input size was set to 128×128×64128× 128× 64. For ACDC and LA, the input size was set to 160×160×16160× 160× 16. For BraTS, the input size was set to 128×128×128128× 128× 128. Other training hyperparameters and data augmentation settings followed nnFormer [23]. Table 1: Comparison of segmentation performance on the Synapse dataset in terms of DSC (%) and HD95 (m). Spl: spleen; RKid: right kidney; LKid: left kidney; Gal: gallbladder; Liv: liver; Sto: stomach; Aor: aorta; Pan: pancreas. Method Spl RKid LKid Gal Liv Sto Aor Pan Mean HD95 U-Net [5] 86.67 68.60 77.77 69.72 93.43 75.58 89.07 53.98 76.85 39.70 TransUNet [3] 85.08 77.02 81.87 63.16 94.08 75.62 87.23 55.86 77.49 31.69 Swin-Unet [2] 90.66 79.61 83.23 66.53 94.29 76.60 85.47 56.58 79.13 21.55 Swin UNETR [6] 94.59 85.88 86.51 66.72 95.33 78.20 90.75 70.07 83.51 14.78 UNETR [7] 87.81 84.80 85.66 60.56 94.46 73.99 89.99 59.25 79.56 22.97 CoTr [18] 88.58 83.62 85.45 68.93 93.89 76.23 85.42 63.77 80.78 19.15 nnFormer [23] 90.51 86.25 86.57 70.17 96.84 86.83 92.04 83.35 86.57 10.63 UNETR++ [13] 89.67 87.23 86.04 71.94 96.56 84.48 92.58 82.30 86.35 10.43 nnU-Net [9] 91.16 86.21 86.92 69.77 96.49 85.92 91.78 83.23 86.44 10.91 CDGC-Net 91.90 87.58 87.56 71.07 96.82 85.57 92.83 82.32 86.96 8.33 3.2 Comparison with State-of-the-Art Methods Table 1 reports the segmentation results on the Synapse dataset. CDGC-Net achieved the highest mean DSC of 86.96% and the lowest HD95 of 8.33 m among the compared methods. Compared with nnFormer, which achieved the next-highest mean DSC, CDGC-Net improved the mean DSC by 0.39 percentage points and reduced HD95 from 10.63 to 8.33 m. Relative to UNETR++, CDGC-Net improved the mean DSC by 0.61 percentage points and reduced HD95 from 10.43 to 8.33 m. CDGC-Net also achieved the highest DSC values for the right kidney, left kidney, and aorta, exceeding the corresponding next-highest values by 0.35, 0.64, and 0.25 percentage points, respectively. In Fig. 2, the Synapse 3D visualization further shows that CDGC-Net produces more complete organ shapes and clearer boundaries, while the compared methods show local misclassification and boundary errors. Figure 2: Qualitative comparison of 3D segmentation results on the Synapse (top) and ACDC (bottom) datasets. Orange dashed boxes highlight local segmentation errors produced by the comparison methods. On the ACDC dataset, CDGC-Net achieved the highest DSC among the listed methods for all three cardiac structures, with a mean DSC of 92.91% (Table 3). Its mean DSC exceeded those of UNETR++, nnU-Net, and nnFormer by 0.47, 0.50, and 0.85 percentage points, respectively. The improvement over UNETR++ was statistically significant based on paired per-case mean DSC values (two-sided Wilcoxon signed-rank test, p<0.03p<0.03). At the structure level, the largest margin over the next-best result was 0.90 percentage points for the RV. The representative example in Fig. 2 illustrates that CDGC-Net better preserved the myocardial ring and reduced local boundary irregularities. For the BraTS dataset, CDGC-Net obtained the highest mean DSC of 82.56% and the lowest HD95 of 5.73 m (Table 3). Its mean DSC exceeded those of Swin UNETR and UNETR++ by 0.17 and 0.53 percentage points, respectively, while reducing HD95 from 5.92 to 5.73 m relative to UNETR++. CDGC-Net also achieved the highest DSC for WT and ET, with improvements of 0.14 and 0.20 percentage points, but its TC result was 0.17 percentage points below that of Swin UNETR. In the displayed example in Fig. 3, its predicted tumor extent and subregions more closely matched the ground truth. On the LA dataset, CDGC-Net achieved the highest DSC of 93.52%, exceeding UNETR++ by 0.32 percentage points (Table 4). Its HD95 of 3.92 m was slightly lower than that of UNETR++. The LA example in Fig. 3 shows that CDGC-Net produced a more continuous atrial region and reduced the fragmented predictions highlighted for UNETR++. Table 2: Segmentation performance on the ACDC dataset in terms of DSC (%). Method RV MYO LV Mean TransUNet [3] 88.86 84.54 95.73 89.71 Swin-Unet [2] 88.55 85.62 95.83 90.00 UNETR [7] 85.29 86.52 94.02 88.61 CoTr [18] 89.13 88.40 95.16 91.04 nnFormer [23] 90.94 89.58 95.65 92.06 UNETR++ [13] 91.08 90.28 95.97 92.44 nnU-Net [9] 90.96 90.34 95.92 92.41 CDGC-Net 91.98 90.66 96.09 92.91 Table 3: Segmentation performance on the BraTS dataset in terms of DSC (%) and HD95 (m). Method WT ET TC Mean HD95 TransUNet [3] 70.60 54.20 68.40 64.40 12.98 Swin UNETR [6] 91.12 77.65 78.41 82.39 6.43 UNETR [7] 90.35 76.30 77.02 81.22 8.82 CoTr [18] 91.01 77.52 77.43 81.99 9.70 nnFormer [23] 91.23 77.84 77.91 82.32 6.12 UNETR++ [13] 91.00 77.73 77.35 82.03 5.92 TransBTS [16] 90.91 77.86 76.10 81.62 9.65 CDGC-Net 91.37 78.06 78.24 82.56 5.73 Figure 3: Qualitative comparison on the LA (top) and BraTS (bottom) datasets. Table 4: Segmentation performance on the LA dataset in terms of DSC (%) and HD95 (m). Method DSC HD95 V-Net [12] 91.14 5.75 DPBNet [20] 92.57 2.74 UNETR++ [13] 93.20 3.95 nnFormer [23] 90.60 7.10 Swin UNETR [6] 91.40 6.20 CDGC-Net 93.52 3.92 3.3 Ablation Experiments We performed ablation experiments on ACDC to quantify the contribution of each CDGC component. As reported in Table 6, the complete CDGC-Net achieved a mean DSC of 92.91%, improving upon the UNETR++ baseline of 92.44% by 0.47 percentage points. Among the evaluated components, excluding the global branch caused the largest decrease, lowering the mean DSC by 0.31 percentage points. Removing residual feature alignment produced a decrease of 0.25 percentage points. The window branch and GHCA contributed improvements of 0.18 and 0.19 percentage points, respectively. Replacing the shared key with separate key projections decreased the mean DSC from 92.91% to 92.77%, indicating a contribution of 0.14 percentage points from key sharing. Table 6 summarizes the model complexity for an input size of 64×128×12864× 128× 128. CDGC-Net contains 25.83M parameters and requires 28.62G FLOPs. Relative to UNETR++, these values correspond to reductions of 39.87% and 40.30%, respectively. Table 5: Ablation results of different components in the CDGC block on the ACDC dataset. Ablation setting RV MYO LV Mean Baseline 91.08 90.28 95.97 92.44 w/o Global branch 91.34 90.48 95.97 92.60 w/o Window branch 91.65 90.64 95.92 92.73 w/o GHCA 91.55 90.59 96.02 92.72 w/o Key sharing 91.69 90.61 96.01 92.77 w/o RFA 91.37 90.55 96.05 92.66 CDGC-Net 91.98 90.66 96.09 92.91 Table 6: Model complexity comparison under the same input size. Method Params (M) FLOPs (G) UNETR [7] 115.38 586.12 TransUNet [3] 96.07 97.36 Swin UNETR [6] 62.19 383.57 nnFormer [23] 150.05 213.37 UNETR++ [13] 42.96 47.94 CDGC-Net 25.83 28.62 4 Conclusion In this paper, we propose CDGC-Net, an efficient 3D medical image segmentation network based on cooperative spatial-channel feature modeling. The core CDGC block integrates CDSA and GHCA within the same module. CDSA captures complementary global and local spatial dependencies, while GHCA models grouped channel relationships to improve channel discrimination. A shared key K connects the two modules in a unified feature space, allowing spatial information to guide channel recalibration and reducing redundant projections. Residual feature alignment further stabilizes the feature refinement process and preserves the original representation. Experiments on Synapse, ACDC, BraTS, and LA show that CDGC-Net achieves competitive or superior Dice and HD95 performance across multi-organ, cardiac, brain tumor, and left atrium segmentation tasks. It also uses fewer parameters and lower FLOPs, demonstrating a better balance between segmentation accuracy and computational efficiency. In future work, we will further evaluate CDGC-Net under cross-domain and multi-center settings and extend it to more heterogeneous medical imaging scenarios. credits 4.0.1 Acknowledgements This work was supported by the National Natural Science Foundation of China under Grant 62261039 and the Postgraduate Innovation Special Fund of Nanchang Hangkong University under Grant YC2025-060. 4.0.2 The authors have no competing interests to declare that are relevant to the content of this article. References [1] O. Bernard, A. Lalande, C. Zotti, F. Cervenansky, X. Yang, P. Heng, I. Cetin, K. Lekadir, O. Camara, M. A. Gonzalez Ballester, G. Sanroma, S. Napel, S. Petersen, G. Tziritas, E. Grinias, M. Khened, V. A. Kollerathu, G. Krishnamurthi, M. Rohé, X. Pennec, M. Sermesant, F. Isensee, P. Jäger, K. H. Maier-Hein, P. M. Full, I. Wolf, S. Engelhardt, C. F. Baumgartner, L. M. Koch, J. M. Wolterink, I. Išgum, Y. Jang, Y. Hong, J. Patravali, S. Jain, O. Humbert, and P. Jodoin (2018) Deep learning techniques for automatic MRI cardiac multi-structures segmentation and diagnosis: is the problem solved?. 37 (11), p. 2514–2525. External Links: Document Cited by: §3.1. [2] H. Cao, Y. Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang (2022) Swin-Unet: unet-like pure transformer for medical image segmentation. In European conference on computer vision, p. 205–218. Cited by: §1, Table 1, Table 3. [3] J. Chen, J. Mei, X. Li, Y. Lu, Q. Yu, Q. Wei, X. Luo, Y. Xie, E. Adeli, Y. Wang, M. P. Lungren, S. Zhang, L. Xing, L. Lu, A. Yuille, and Y. Zhou (2024) TransUNet: rethinking the u-net architecture design for medical image segmentation through the lens of transformers. Medical Image AnalysisIEEE Transactions on Image ProcessingIEEE Transactions on Medical ImagingarXiv preprint arXiv:1904.10509arXiv preprint arXiv:2006.04768Journal of Applied Remote SensingIEEE Transactions on Medical ImagingIEEE Transactions on Medical ImagingMedical Image AnalysisIEEE Transactions on Medical Imaging 97, p. 103280. External Links: ISSN 1361-8415, Document, Link Cited by: §1, §3.1, Table 1, Table 3, Table 3, Table 6. [4] R. Child, S. Gray, A. Radford, and I. Sutskever (2019) Generating long sequences with sparse transformers. Cited by: §1. [5] Ö. Çiçek, A. Abdulkadir, S. S. Lienkamp, T. Brox, and O. Ronneberger (2016) 3D U-Net: learning dense volumetric segmentation from sparse annotation. In International conference on medical image computing and computer-assisted intervention, p. 424–432. Cited by: Table 1. [6] A. Hatamizadeh, V. Nath, Y. Tang, D. Yang, H. R. Roth, and D. Xu (2021) Swin UNETR: swin transformers for semantic segmentation of brain tumors in mri images. In International MICCAI brainlesion workshop, p. 272–284. Cited by: Table 1, Table 3, Table 4, Table 6. [7] A. Hatamizadeh, Y. Tang, V. Nath, D. Yang, A. Myronenko, B. Landman, H. R. Roth, and D. Xu (2022) UNETR: transformers for 3d medical image segmentation. In 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Vol. , p. 1748–1758. External Links: Document Cited by: §1, §2.5, Table 1, Table 3, Table 3, Table 6. [8] J. Hu, L. Shen, and G. Sun (2018) Squeeze-and-excitation networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vol. , p. 7132–7141. External Links: Document Cited by: §1. [9] F. Isensee, P. F. Jaeger, S. A. A. Kohl, J. Petersen, and K. H. Maier-Hein (2021) nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature Methods 18, p. 203–211. External Links: Document Cited by: §1, Table 1, Table 3. [10] B. Landman, Z. Xu, J. E. Iglesias, M. Styner, T. R. Langerak, and A. Klein (2015) MICCAI multi-atlas labeling beyond the cranial vault – workshop and challenge. In Proceedings of the MICCAI Multi-Atlas Labeling Beyond Cranial Vault – Workshop and Challenge, Vol. 5, p. 12. Cited by: §3.1. [11] B. H. Menze, A. Jakab, S. Bauer, J. Kalpathy-Cramer, K. Farahani, J. Kirby, Y. Burren, N. Porz, J. Slotboom, R. Wiest, L. Lanczi, E. Gerstner, M. Weber, T. Arbel, B. B. Avants, N. Ayache, P. Buendia, D. L. Collins, N. Cordier, J. J. Corso, A. Criminisi, T. Das, H. Delingette, Ç. Demiralp, C. R. Durst, M. Dojat, S. Doyle, J. Festa, F. Forbes, E. Geremia, B. Glocker, P. Golland, X. Guo, A. Hamamci, K. M. Iftekharuddin, R. Jena, N. M. John, E. Konukoglu, D. Lashkari, J. A. Mariz, R. Meier, S. Pereira, D. Precup, S. J. Price, T. R. Raviv, S. M. S. Reza, M. Ryan, D. Sarikaya, L. Schwartz, H. Shin, J. Shotton, C. A. Silva, N. Sousa, N. K. Subbanna, G. Szekely, T. J. Taylor, O. M. Thomas, N. J. Tustison, G. Unal, F. Vasseur, M. Wintermark, D. H. Ye, L. Zhao, B. Zhao, D. Zikic, M. Prastawa, M. Reyes, and K. Van Leemput (2015) The multimodal brain tumor image segmentation benchmark (BRATS). 34 (10), p. 1993–2024. External Links: Document Cited by: §3.1. [12] F. Milletari, N. Navab, and S. Ahmadi (2016) V-net: fully convolutional neural networks for volumetric medical image segmentation. In 2016 Fourth International Conference on 3D Vision (3DV), Vol. , p. 565–571. External Links: Document Cited by: Table 4. [13] A. Shaker, M. Maaz, H. Rasheed, S. Khan, M. Yang, and F. Shahbaz Khan (2024) UNETR++: delving into efficient and accurate 3d medical image segmentation. 43 (9), p. 3377–3390. External Links: Document Cited by: §1, §1, §1, §2.1, §2.5, Table 1, Table 3, Table 3, Table 4, Table 6. [14] Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, and Q. Hu (2020) ECA-Net: efficient channel attention for deep convolutional neural networks. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , p. 11531–11539. External Links: Document Cited by: §1. [15] S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma (2020) Linformer: self-attention with linear complexity. Cited by: §1. [16] W. Wang, C. Chen, M. Ding, H. Yu, S. Zha, and J. Li (2021) TransBTS: multimodal brain tumor segmentation using transformer. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2021, M. de Bruijne, P. C. Cattin, S. Cotin, N. Padoy, S. Speidel, Y. Zheng, and C. Essert (Eds.), Cham, p. 109–119. External Links: ISBN 978-3-030-87193-2, Document Cited by: Table 3. [17] S. Woo, J. Park, J. Lee, and I. S. Kweon (2018) CBAM: convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV), p. 3–19. Cited by: §1. [18] Y. Xie, J. Zhang, C. Shen, and Y. Xia (2021) CoTr: efficiently bridging cnn and transformer for 3d medical image segmentation. In International conference on medical image computing and computer-assisted intervention, p. 171–180. Cited by: §1, Table 1, Table 3, Table 3. [19] Z. Xiong, Q. Xia, Z. Hu, N. Huang, C. Bian, Y. Zheng, S. Vesal, N. Ravikumar, A. Maier, X. Yang, P. Heng, D. Ni, C. Li, Q. Tong, W. Si, E. Puybareau, Y. Khoudli, T. Géraud, C. Chen, W. Bai, D. Rueckert, L. Xu, X. Zhuang, X. Luo, S. Jia, M. Sermesant, Y. Liu, K. Wang, D. Borra, A. Masci, C. Corsi, C. de Vente, M. Veta, R. Karim, C. J. Preetha, S. Engelhardt, M. Qiao, Y. Wang, Q. Tao, M. Núñez-Garcia, O. Camara, N. Savioli, P. Lamata, and J. Zhao (2021) A global benchmark of algorithms for segmenting the left atrium from late gadolinium-enhanced cardiac magnetic resonance imaging. 67, p. 101832. External Links: ISSN 1361-8415 Cited by: §3.1. [20] F. Xu, W. Tu, F. Feng, M. Gunawardhana, J. Yang, Y. Gu, and J. Zhao (2024) Dynamic position transformation and boundary refinement network for left atrial segmentation. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2024, p. 209–219. Cited by: Table 4. [21] Y. Zhang, J. Kong, and Z. F. He (2022) Convolutional block attention module u-net: a method to improve attention mechanism and u-net for remote sensing images. 16 (2), p. 026516–1–026516–15. Cited by: §1. [22] Y. Zhang, H. Liu, and Q. Hu (2021) TransFuse: fusing transformers and CNNs for medical image segmentation. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2021, Cham, p. 14–24. Cited by: §1. [23] H. Zhou, J. Guo, Y. Zhang, X. Han, L. Yu, L. Wang, and Y. Yu (2023) NnFormer: volumetric medical image segmentation via a 3d transformer. 32 (), p. 4036–4045. External Links: Document Cited by: §1, §2.5, §3.1, §3.1, Table 1, Table 3, Table 3, Table 4, Table 6.