Paper deep dive
SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification
Harshit Mittal, Arash Rabbani
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/1/2026, 2:24:18 AM
Summary
The paper introduces SAFViT, a Vision Transformer-based model for nucleus segmentation and classification that replaces standard skip connections with a novel Spatial Attention Fusion (SAF) Gating module. This module generates a per-pixel 'heatmap of trust' to dynamically weight encoder and decoder features, significantly improving the detection of minority classes like 'Dead' cells and achieving state-of-the-art multi-class panoptic quality (mPQ) on the PanNuke dataset.
Entities (10)
Relation Signals (7)
SAFViT â evaluatedon â PanNuke
confidence 99% · SAF Gating is compared against six gating alternatives... on PanNuke and MoNuSeg datasets.
SAFViT â uses â Spatial Attention Fusion Gating
confidence 98% · This study proposes replacing conventional skip connections in a CellViT-based model with a novel Spatial Attention Fusion (SAF) Gating module.
SAFViT â evaluatedon â MoNuSeg
confidence 95% · SAF Gating is compared against six gating alternatives... on PanNuke and MoNuSeg datasets.
SAFViT â outperforms â CellViT
confidence 95% · SAF Gating achieves the highest mPQ (0.471), a gain driven primarily by a 14.5-point improvement in Dead-class F1 score compared to ungated CellViT baseline.
SAFViT â usesbackbone â Swin-Tiny
confidence 93% · The encoder employs a Swin-Tiny transformer backbone.
Spatial Attention Fusion Gating â improvesdetectionof â Dead
confidence 92% · The resulting fused features improve the model's ability to detect the minority 'Dead' class.
SAFViT â achievesmetric â mPQ
confidence 90% · SAF Gating achieves the highest mPQ (0.471).
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Accurate cell segmentation and classification are foundational to digital pathology, enabling quantitative tissue analysis for diagnosis and treatment planning. Encoder-decoder architectures that fuse multi-scale features through skip connections have become the dominant paradigm for this task, yet standard direct skip connections treat every spatial location equally, which leads to redundant and potentially conflicting information reaching the decoder. To overcome this problem, various gating mechanisms have been introduced, but most of them operate solely on filtering encoder information, neglecting the benefit of global contextual information from the decoder. This study proposes replacing conventional skip connections in a CellViT-based model with a novel Spatial Attention Fusion (SAF) Gating module. Each SAF gate concatenates the encoder skip and upsampled decoder features, compresses them through two pointwise convolutions with an intermediate ReLU, and applies a channel-wise softmax to produce a per-pixel "heatmap of trust" that sums to unity at every spatial location, allowing the network to learn where each source is most trustworthy. The resulting fused features improve the model's ability to detect the minority "Dead" class, which in turn enhances the multi-class panoptic quality (mPQ) on the PanNuke dataset. SAF Gating is compared against six gating alternatives including no gating, attention gates, squeeze-and-excitation, CBAM, cross-attention, and attentional feature fusion on PanNuke and MoNuSeg datasets. SAF Gating achieves the highest mPQ (0.471), a gain driven primarily by a 14.5-point improvement in Dead-class F1 score compared to ungated CellViT baseline.
Tags
Links
- Source: https://arxiv.org/abs/2607.27835v1
- Canonical: https://arxiv.org/abs/2607.27835v1
Trouble viewing inline? Open PDF directly â
Full Text
42,287 characters extracted from source content.
Expand or collapse full text
SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification Harshit Mittal Arash Rabbani a.rabbani@leeds.ac.uk Abstract Accurate cell segmentation and classification are foundational to digital pathology, enabling quantitative tissue analysis for diagnosis and treatment planning. Encoderâdecoder architectures that fuse multi-scale features through skip connections have become the dominant paradigm for this task, yet standard direct skip connections treat every spatial location equally, which leads to redundant and potentially conflicting information reaching the decoder. To overcome this problem, various gating mechanisms have been introduced, but most of them operate solely on filtering encoder information, neglecting the benefit of global contextual information from the decoder. This study proposes replacing conventional skip connections in a CellViT-based model with a novel Spatial Attention Fusion (SAF) Gating module. Each SAF gate concatenates the encoder skip and upsampled decoder features, compresses them through two pointwise convolutions with an intermediate ReLU, and applies a channel-wise softmax to produce a per-pixel âheatmap of trustâ that sums to unity at every spatial location, allowing the network to learn where each source is most trustworthy. The resulting fused features improve the modelâs ability to detect the minority âDeadâ class, which in turn enhances the multi-class panoptic quality (mPQ) on the PanNuke dataset. SAF Gating is compared against six gating alternatives including no gating, attention gates, squeeze-and-excitation, CBAM, cross-attention, and attentional feature fusion on PanNuke and MoNuSeg datasets. SAF Gating achieves the highest mPQ (0.471), a gain driven primarily by a 14.5-point improvement in Dead-class F1F_1 score compared to ungated CellViT baseline. keywords: Cell Segmentation , Encoderâdecoder architectures , Spatial Attention Fusion (SAF) Gating , Heatmap of Trust , multi-class panoptic quality (mPQ) â journal: Computerized Medical Imaging and Graphics [label1]organization=School of Computer Science, University of Leeds, addressline=Woodhouse, city=Leeds, postcode=LS2 9JT, country=United Kingdom graphicalabstract highlights Proposes Spatial Attention Fusion Gating to build a CellViT-based model SAFViT. Generates âheatmap of trustâ to learn most trustworthy source for each pixel. Achieves highest F1F_1 score for minority class âDeadâ among all gating methods. Achieves highest mPQ of 0.471 among all compared models including the baseline. 1 Introduction Cancer is the second-leading cause of mortality after cardiovascular diseases, with millions of new cases registered each year (Bray et al., 2024). The disease manifests differently across organs, producing distinct patterns of cellular deformation that must be identified for accurate diagnosis. Even with new and effective non-invasive radiological imaging techniques, using tissue samples for examining them under a microscope is still a common practice for efficient diagnosis. By identifying types of anomalies in the tissue, a pathologist can decide the stage for possible treatment strategies and also utilise them for additional research. However, manual histopathology is highly time-consuming and may cause variability in observations across different viewers. Advancement in computational histopathology, driven by high-level research in deep learning (Voulodimos et al., 2018), has enabled automated analysis of whole-slide images (WSIs) at a scale and consistency that led to digitised cell segmentation and classification (Meijering, 2012; Vicar et al., 2019). Digital cell segmentation and classification accurately outlines individual cell nuclei and identifies the class they belong to, helping in tumour grading, prognosis, and treatment planning. The morphological and spatial distribution of nuclei such as neoplastic, inflammatory, connective, dead, and epithelial cells provides critical diagnostic information that can augment and accelerate clinical decision-making if made reliably (Gamper et al., 2019). Early nuclei segmentation methods relied on classical morphological processing, with Watershed-based approaches treating image intensity as a topographic surface to delineate boundaries (Levner and Zhang, 2007). While computationally efficient, these methods proved sensitive to staining variation and noise, often leading to over-segmentation. The introduction of U-Net (Ronneberger et al., 2015; Schmidt et al., 2018; Graham et al., 2019) marked a turning point, establishing the encoderâdecoder architecture with skip connections as the dominant paradigm for biomedical image segmentation (Doan et al., 2022; Baumann et al., 2024). The introduction of the Vision Transformer (ViT) showed that pure self-attention mechanisms could match or exceed convolutional neural networks (CNNs) on image recognition tasks (Dosovitskiy et al., 2021; Vaswani et al., 2017). By treating an image as a sequence of patches, ViT captures long-range dependencies across the entire input, rather than the local receptive fields typical of convolutional layers (Chen et al., 2021; Cao et al., 2023). CellViT extended this idea to nuclei segmentation by leveraging large-scale pretrained ViT encoders, achieving strong performance on the PanNuke dataset (Hörst et al., 2024, 2026). Despite their architectural differences, all these encoderâdecoder models depend on skip connections to fuse multi-scale features. Standard skip connections employ direct concatenation or addition, treating all spatial regions and channels as equally informative, which often propagates redundant noise from early encoder layers. Several gating mechanisms have been proposed to address this (Oktay et al., 2018; Hu et al., 2018; Khanh et al., 2020). While these methods improve feature selection, most operate only on the encoder stream. They treat the decoder features merely as a fixed gating signal, rather than as a source that could itself benefit from selective modulation. This one-sided filtering neglects the potential of jointly learning where local boundary detail from the encoder and global contextual information from the decoder are each most reliable. This paper addresses this limitation by proposing SAFViT (Spatial Attention Fusion Vision Transformer), a CellViT-based architecture that replaces conventional skip connections with a novel Spatial Attention Fusion (SAF) Gating module. Unlike existing approaches that gate only the encoder stream, each SAF gate treats both the encoder skip and the upsampled decoder feature as equal candidates for fusion. The two streams are concatenated and compressed through two pointwise convolutions with an intermediate ReLU, followed by a channel-wise softmax that produces a two-channel, per-pixel âheatmap of trustâ summing to unity at every spatial location. This heatmap adaptively weights local detail from the encoder against global context from the decoder before a refinement convolution, allowing the network to learn where each source is most trustworthy. Contributions: We propose Spatial Attention Fusion (SAF) Gating, a novel per-pixel dual-stream gating mechanism that applies a softmax-based heatmap of trust to jointly modulate encoder and decoder features, rather than filtering only one stream. We implement SAFViT by integrating SAF Gating into a CellViT-based architecture with a Swin-Tiny encoder and specialised segmentation, horizontalâvertical (HoVer) distance, and cell-type classification heads (Graham et al., 2019). We conduct a controlled comparison study of SAF Gating against six alternative gating mechanisms, namely standard skip connections, attention gates, squeeze-and-excitation, the Convolutional Block Attention Module, cross-attention and attentional feature fusion, under identical training conditions on the PanNuke pan-cancer dataset using 3-fold cross-validation. We demonstrate that SAF Gating achieves the highest mPQ (0.4710.471) among all compared methods, with the improvement attributable specifically to minority Dead-class detection (F1=0.518F_1=0.518), and show that on class-agnostic metrics (Dice, Aggregated Jaccard Index [AJI], and binary Panoptic Quality [bPQ]) all gating variants perform comparably, confirming that the choice of fusion strategy does not affect majority-class segmentation but has decisive impact on underrepresented classes. We also validate out-of-domain generalisation on the MoNuSeg test dataset (Kumar et al., 2020). 2 Proposed Network Architecture The proposed SAFViT architecture builds upon the CellViT framework, retaining its core encoderâdecoder topology while introducing a fundamentally different feature fusion strategy at the skip connections. The encoder employs a Swin-Tiny transformer backbone (Liu et al., 2021) pretrained on ImageNet (Deng et al., 2009), which processes a 224Ă224Ă3224Ă 224Ă 3 haematoxylin and eosin (H&E) input patch through four hierarchical stages, producing multi-scale feature maps E1,E2,E3,E4\E_1,E_2,E_3,E_4\ with channel dimensions 96,192,384,768\96,192,384,768\ at spatial resolutions 562,282,142,72\56^2,28^2,14^2,7^2\, respectively. The decoder follows a bottom-up pathway: the deepest encoder output E4E_4 is progressively upsampled through three transposed convolution stages, each doubling the spatial resolution. At each decoder level iâ1,2,3iâ\1,2,3\, the upsampled decoder feature and the corresponding encoder skip connection EiE_i are fused through the proposed SAF Gating module, described in Section 2.1, rather than through direct concatenation or addition as in the original CellViT. After three stages of gated fusion, the decoder produces a feature tensor 1ââ96Ă56Ă56x_1 ^96Ă 56Ă 56 at the highest resolution. Figure 1: The Step Diagram for the Architecture of SAFViT Model from initial H&E patches to various stages which includes 4 encoder stages downsampling the images in blue boxes, SAF gating of each skip connection in pink, Decoder upsampling in violet, further leading to different heads in green, post procesing and finally instance maps as outputs in orange. This shared representation is then fed into three independent, task-specific prediction heads. The Nuclei Presence (NP) head outputs a single-channel binary map Y^nâpâ[0,1]HĂW Y_npâ[0,1]^HĂ W indicating whether each pixel belongs to any nucleus. The HorizontalâVertical (HV) head produces a two-channel map Y^hâvââ2ĂHĂW Y_hv ^2Ă HĂ W encoding the horizontal and vertical distances of each nuclear pixel to its instance centroid, following the formulation introduced in HoVer-Net (Graham et al., 2019). The Nucleus Type (NT) head outputs a (K+1)(K+1) channel classification map Y^nâtââ(K+1)ĂHĂW Y_nt ^(K+1)Ă HĂ W, where K is the number of cell classes and the additional channel represents the background. Each head comprises two 3Ă33Ă 3 convolutional layers with batch normalisation and ReLU activation, followed by a final 1Ă11Ă 1 convolution that projects to the respective output dimensionality. At inference time, the NP map provides a binary nuclear mask, the HV maps are used to compute gradient-based markers for watershed-based instance separation, and the NT map assigns a cell-type label to each segmented instance via majority voting within each predicted region. The architecture is visually defined in Figure 1. 2.1 Spatial Attention Fusion (SAF) Gating Figure 2: Spatial Attention Fusion (SAF) Gating module. Encoder skip localF_local and upsampled decoder globalF_global are concatenated, compressed through two 1Ă11Ă1 convolutions, and normalised by channel-wise softmax to produce a two-channel per-pixel heatmap of trust [local;global][w_local;w_global]. The central contribution of SAFViT is the SAF Gating module, which replaces the conventional direct skip connection at each decoder level. Unlike standard attention gates that produce a single scalar weight per spatial location to filter only the encoder stream, SAF Gating treats the encoder and decoder features as dual information streams and learns a per-pixel reliability map, a âheatmap of trustâ that jointly modulates both before fusion. The gate pursues a single goal: at every spatial location (h,w)(h,w), decide how much to trust local boundary detail from the encoder versus global contextual information from the decoder, and blend the two streams in proportion to that learned decision. The complete mechanism is illustrated in Figure 2 and formalised below. Let localââCĂHĂWF_local ^CĂ HĂ W denote the encoder skip-connection features carrying local boundary detail, and globalââCĂHĂWF_global ^CĂ HĂ W denote the upsampled decoder features carrying global contextual information, where both have been aligned to the same spatial resolution and channel dimensionality C. The two streams are first concatenated along the channel axis to form a joint representation: concat=concat[local:global]ââ2âCĂHĂWF_concat=concat[F_local:F_global] ^2CĂ HĂ W (1) This concatenated tensor is then passed through a lightweight gate network comprising two successive 1Ă11Ă 1 convolutional layers with a bottleneck structure. The first convolution compresses the channel dimension by a factor of four, followed by a ReLU non-linearity to introduce a representational bottleneck: =ReLUâ(1âconcat+1)ââ(C/2)ĂHĂWG=ReLU\! (W_1*F_concat+b_1 ) ^(C/2)Ă HĂ W (2) where 1ââ(C/2)Ă2âCĂ1Ă1W_1 ^(C/2)Ă 2CĂ 1Ă 1 and 1ââC/2b_1 ^C/2 are learnable parameters. The second convolution projects the bottleneck representation to exactly two channels: =2â+2ââ2ĂHĂWA=W_2*G+b_2 ^2Ă HĂ W (3) where 2ââ2Ă(C/2)Ă1Ă1W_2 ^2Ă(C/2)Ă 1Ă 1 and 2ââ2b_2 ^2. A softmax is then applied along the channel dimension to obtain normalised trust weights: [local;global]=Softmaxâ(,dim=channel)ââ2ĂHĂW[w_local;w_global]=Softmax(A,dim=channel) ^2Ă HĂ W (4) where for every spatial position (h,w)(h,w): wlocal(h,w)+wglobal(h,w)=1,wlocal(h,w),wglobal(h,w)â(0,1)w_local^(h,w)+w_global^(h,w)=1, w_local^(h,w),\ w_global^(h,w)â(0,1) (5) The resulting weight maps localw_local and globalw_global constitute the heatmap of trust: at spatial locations where the gate network determines that local encoder detail is more informative (e.g. at nuclear boundaries), wlocalw_local approaches unity. Conversely, in homogeneous tissue regions where global decoder context is more reliable, wglobalw_global dominates. These weights are broadcast across all C channels and applied through element-wise multiplication to their respective feature streams, followed by summation: blend=localâlocal+globalâglobalF_blend=w_local _local+w_global _global (6) where â denotes element-wise multiplication with channel-wise broadcasting. Finally, a refinement layer consisting of a 3Ă33Ă 3 convolution followed by batch normalisation and ReLU activation is applied to smooth the blended features and allow local spatial mixing: fused=ReLUâ(BNâ(râblend+r))ââCĂHĂWF_fused=ReLU\! (BN\! (W_r*F_blend+b_r ) ) ^CĂ HĂ W (7) This fused tensor fusedF_fused replaces the conventional skip-connected output at the corresponding decoder level and proceeds to the next upsampling stage or, at the shallowest level, to the three prediction heads. The softmax normalisation in Eq. (4) guarantees that the two trust weights are complementary at every pixel, enforcing a zero-sum trade-off between local and global reliance. Attention gates (Oktay et al., 2018) instead produce an independent scalar attention coefficient for the encoder alone, and SE blocks (Hu et al., 2018) recalibrate channels globally without spatial specificity. SAF Gating does not discard information: it redistributes it between the two sources based on learned spatial reliability, so the gating is both selective and information-preserving. 3 Experimental Setup 3.1 PanNuke Dataset PanNuke (Gamper et al., 2019) is a large-scale, semi-automatically generated pan-cancer nuclei dataset comprising 7,901 image patches of size 256Ă256256Ă 256 pixels extracted from 19 distinct tissue types. Each patch contains pixel-level instance segmentation masks annotated across five clinically relevant cell classes: neoplastic, inflammatory, connective, dead, and epithelial. The dataset exhibits substantial class imbalance, with neoplastic nuclei dominating and the dead class being severely underrepresented, making it a challenging benchmark for multi-class segmentation. All input patches are centre-cropped to 224Ă224224Ă 224 pixels to match the Swin-Tiny encoderâs expected input resolution. During training, standard data augmentation is applied, including random horizontal and vertical flips, random rotation, colour jitter, Gaussian blur, and elastic deformation. All models, SAFViT and the six gating baselines, are trained under identical conditions: the same fold splits, augmentation pipeline, learning rate schedule (cosine annealing with warm restarts (Loshchilov and Hutter, 2017)), optimiser (AdamW (Loshchilov and Hutter, 2019)), and number of epochs, so that observed performance differences are attributable solely to the gating mechanism. 3.2 MoNuSeg Dataset The MoNuSeg dataset (Kumar et al., 2020) is used to evaluate out-of-domain generalisation. It consists of H&E-stained tissue images from seven organs, with the official test set containing 14 images of size 1000Ă10001000Ă 1000 pixels accompanied by XML-format instance annotations. Unlike PanNuke, MoNuSeg provides only binary instance masks without cell-type labels, making it a class-agnostic segmentation benchmark. Models trained on PanNuke are applied directly to MoNuSeg without any fine-tuning, testing whether the learned feature representations and gating behaviour transfer to unseen tissue types and staining conditions. Table 1: Segmentation performance on PanNuke (3-fold cross-validation) and MoNuSeg datasets. PanNuke values are reported as mean ± std across five folds. Best results per metric are shown in bold. Inference time (I. Time) is reported in milliseconds per image. Model Dice AJI bPQ mPQ PanNuke Dataset CellViT 0.832±.003 0.672±.004 0.625±.003 0.453±.030 SAFViT 0.834±.003 0.674±.006 0.627±.007 0.471±.032 AG 0.834±.002 0.670±.005 0.620±.005 0.430±.007 SE 0.833±.002 0.674±.003 0.628±.001 0.457±.029 CBAM 0.834±.004 0.674±.005 0.627±.003 0.453±.026 CrossAttn 0.827±.002 0.663±.005 0.608±.005 0.449±.026 AFF 0.833±.003 0.670±.003 0.622±.000 0.429±.001 MoNuSeg Dataset Model Dice AJI I. Time(ms)(Colab T4) CellViT 0.738±.033 0.539±.035 5.8 SAFViT 0.737±.033 0.540±.035 5.7 AG 0.745±.028 0.544±.031 5.6 SE 0.740±.033 0.542±.035 5.6 CBAM 0.739±.032 0.542±.034 5.6 CrossAttn 0.741±.030 0.541±.034 5.6 AFF 0.740±.031 0.542±.034 5.7 3.3 Evaluation Metrics We evaluate all models using four complementary metrics. The Dice Score measures pixel-level overlap between the predicted binary nuclear mask and the ground truth. The Aggregated Jaccard Index (AJI) extends the standard Jaccard index to the instance level by computing a one-to-one matching between predicted and ground-truth instances and aggregating the intersection-over-union across all matched pairs, penalising both missed and spurious detections. The Binary Panoptic Quality (bPQ) evaluates class-agnostic instance segmentation by combining detection (F1F_1 at IoUâ„0.5IoUâ„ 0.5) with segmentation quality (mean IoU of matched pairs). The Multi-class Panoptic Quality (mPQ) extends bPQ to the multi-class setting by computing panoptic quality independently for each cell class and averaging across classes. Since mPQ weights all classes equally regardless of prevalence, it is particularly sensitive to performance on rare classes such as the Dead class, making it the most discriminative metric for evaluating the impact of gating mechanisms on underrepresented cell types. For MoNuSeg, which lacks cell-type annotations, only Dice and AJI are reported. Inference time per image is also recorded to confirm negligible computational overhead. 4 Results and Discussion 4.1 Model Nomenclature All seven compared models share the same Swin-Tiny encoder, decoder topology, three prediction heads, composite loss function, and training protocol. The only architectural difference is the skip-connection fusion strategy. CellViT is the ungated baseline (Hörst et al., 2024), SAFViT (proposed) employs SAF Gating, and the five remaining gating variants are attention gates (AG) (Oktay et al., 2018), squeeze-and-excitation (SE) (Hu et al., 2018), the Convolutional Block Attention Module (CBAM) (Woo et al., 2018), gated cross-attention (CrossAttn) (Jia et al., 2024), and attentional feature fusion (AFF) (Dai et al., 2021). Figure 3: SAF Gate âHeatmap of Trustâ for a representative PanNuke patch. From left: H&E patch, ground-truth mask coloured by class, and wlocalw_local overlaid at Gate 3 (deepest, 14Ă1414Ă14), Gate 2 (28Ă2828Ă28), and Gate 1 (shallowest, 56Ă5656Ă56). Red indicates high encoder trust, blue indicates high decoder trust. At nuclear boundaries especially Dead cells (blue in GT) wlocalw_local approaches 1.01.0, confirming that the gate amplifies fine-scale encoder detail precisely where rare-class discrimination is needed. Figure 4: Visual comparison of selected models with ground truth including proposed SAFViT. Where Red: Neoplastic, Green: Inflammatory, Yellow: Connective, Blue: Dead and Orange: Epithelial. This image proves the superiority of SAFViT in recognizing minority cells such as âDead Cellsâ much more efficiently than compared to other benchmark models. Table 2: Class-wise F1F_1-scores on the PanNuke dataset (out-of-fold). Best per class in bold. Model Neop. Infl. Conn. Dead Epit. CellViT 0.902 0.801 0.781 0.373 0.898 SAFViT 0.903 0.790 0.782 0.518 0.898 AG 0.898 0.790 0.778 0.000 0.895 SE 0.903 0.800 0.781 0.295 0.903 CBAM 0.898 0.792 0.777 0.364 0.900 CrossAttn 0.892 0.773 0.756 0.462 0.885 AFF 0.893 0.785 0.778 0.000 0.881 Table 3: Wilcoxon signed-rank test results comparing the proposed SAFViT against alternative models on per-image AJI and Dice scores (N=7,901N=7,901). Metric Comparison Model Mean Diff. p-value AJI CellViT +0.0013+0.0013 2.16Ă10â22.16Ă 10^-2 AG +0.0033+0.0033 5.47Ă10â185.47Ă 10^-18 SE +0.0002+0.0002 8.80Ă10â18.80Ă 10^-1 CBAM â0.0005-0.0005 2.19Ă10â12.19Ă 10^-1 CrossAttn +0.0105+0.0105 2.08Ă10â942.08Ă 10^-94 AFF +0.0033+0.0033 3.39Ă10â53.39Ă 10^-5 DICE CellViT +0.0013+0.0013 5.53Ă10â35.53Ă 10^-3 AG â0.0002-0.0002 4.91Ă10â24.91Ă 10^-2 SE +0.0008+0.0008 8.91Ă10â18.91Ă 10^-1 CBAM â0.0002-0.0002 2.06Ă10â12.06Ă 10^-1 CrossAttn +0.0065+0.0065 1.94Ă10â1321.94Ă 10^-132 AFF +0.0010+0.0010 7.51Ă10â27.51Ă 10^-2 4.2 PanNuke Results Table 1 presents the quantitative results of all seven models on the PanNuke dataset, evaluated using 3-fold cross-validation with metrics reported as mean ± standard deviation. Table 2 provides the corresponding per-class F1F_1-scores computed from out-of-fold predictions, where each image is scored exactly once by the checkpoint that did not train on it. SAFViT achieves the highest multi-class panoptic quality (mPQ=0.471mPQ=0.471), outperforming the ungated CellViT baseline (0.4530.453) by 1.8 percentage points. This gain is notable because mPQ weights all five cell classes equally regardless of prevalence, making it the metric most sensitive to performance on rare categories. Among the gating alternatives, SE (0.4570.457) and CBAM (0.4530.453) follow as the next strongest on mPQ, while AG (0.4300.430) and AFF (0.4290.429) fall below even the ungated baseline, indicating that purely encoder-side gating can be counterproductive when it lacks the capacity to preserve informative encoder activations for underrepresented classes. On binary metrics, all models cluster tightly: Dice 0.8270.827â0.8340.834, AJI 0.6630.663â0.6740.674 (SAFViT, SE, and CBAM tie for the highest AJI at 0.6740.674), bPQ 0.6080.608â0.6280.628 (SE leads marginally at 0.6280.628, SAFViT at 0.6270.627). These narrow ranges confirm that the comparison is controlled and that gating strategy does not affect majority-class foreground segmentation. The per-class F1F_1-scores in Table 2 reveal the mechanism behind SAFViTâs mPQ advantage. SAFViT achieves the highest Dead-class F1F_1-score (0.5180.518), representing a 14.5-point lead over the ungated CellViT baseline (0.3730.373) and a 5.6-point lead over the next best gating variant, CrossAttn (0.4620.462). Two gating methods, AG and AFF, score exactly 0.0000.000 on the Dead class, meaning they completely fail to detect this minority category. AG gates only the encoder stream through a single sigmoid coefficient, which tends to suppress weak encoder activations associated with rare cell types. AFF, although it does fuse encoder and decoder streams via a learned channel-attention weight, derives that weight from a squeeze-and-excitation-style global context vector rather than a per-pixel softmax; this coarser, channel-wise (rather than spatial) weighting appears insufficient to preserve the sparse, spatially localised activations that signal Dead-class nuclei. On the four majority classes (Neoplastic, Inflammatory, Connective, Epithelial), SAFViT remains competitive with the best-performing variants: it ties SE for the highest Neoplastic F1F_1-score (0.9030.903), leads on Connective (0.7820.782), and falls within 1.1 points of CellViT on Inflammatory (0.7900.790 vs. 0.8010.801). The slight Inflammatory trade-off is more than offset by the substantial Dead-class gain when averaged into mPQ. To assess whether the observed metric differences reflect systematic variation rather than noise, we applied a Wilcoxon signed-rank test to the per-image Dice and AJI scores across the full out-of-fold evaluation set (N=7,901N=7,901). Results are reported in Table 3. SAFViT significantly outperforms AG, AFF, and CrossAttn on AJI, and shows a significant but modest advantage over CellViT. SAFViT is statistically indistinguishable from SE and CBAM on both metrics, consistent with the narrow numeric gaps in Table 1. As mPQ is a fold-level quantity (N=3N=3 folds), formal significance testing is not applicable. The numerical mPQ lead over SE and CellViT should be interpreted alongside the fold-level standard deviations in Table 1 and Figures 5, 6(a), 6(b), 6(c), 6(d). Figure 3 visualises the per-pixel trust weight wlocalw_local at each of the three SAF gates for a representative patch containing Dead and Neoplastic nuclei. At nuclear boundaries particularly the thin, irregular membranes of Dead cells wlocalw_local approaches 1.01.0, indicating that the gate relies predominantly on local encoder detail in these regions. In homogeneous stromal regions, wglobalw_global dominates, reflecting that global decoder context is more reliable where fine-grained boundary information is absent. This spatial pattern provides a mechanistic explanation for the Dead-class F1F_1 gain: SAF Gating preferentially amplifies encoder skip features at exactly the fine-scale loci where Dead cells are most distinguishable from background which is proved in Figure 4. Figure 5: The difference between different models with respect to their f1-score for each class of cells including Neoplastic, Inflammatory, Connective, Dead and Epithelial. The CellViT is pink color, SAFViT is yellow, AG is green, AFF is pink, CrossAttention is purple, SE is blue and CBAM is violet. 4.3 MoNuSeg Results To assess out-of-domain generalisation, all models trained on PanNuke are evaluated on the MoNuSeg test set without any fine-tuning. Since MoNuSeg provides only binary instance annotations without cell-type labels, only Dice and AJI are reported. As shown in the MoNuSeg columns of Table 1, all seven models cluster tightly (Dice 0.7370.737â0.7450.745, AJI 0.5390.539â0.5440.544), with AG leading marginally (Dice 0.7450.745, AJI 0.5440.544). SAFViT (Dice 0.7370.737, AJI 0.5400.540) performs on par with the ungated CellViT baseline and the remaining gating variants. This convergence is expected: MoNuSegâs binary evaluation cannot capture multi-class discrimination, which is precisely where SAFViTâs advantage lies. AGâs slight MoNuSeg lead contrasts sharply with its poor PanNuke mPQ (0.4300.430) and zero Dead-class F1F_1-score, illustrating that strong class-agnostic boundary detection does not imply strong cell-type classification. The consistent MoNuSeg performance across all variants confirms that SAF Gating does not overfit to PanNuke-specific class distributions and that the learned spatial weighting strategy generalises to unseen tissue types and staining conditions. (a) Dice Score comparison. (b) AJI comparison. (c) mPQ comparison. (d) bPQ comparison. Figure 6: Quantitative metric performance comparisons on the PanNuke validation split (33-fold cross-validation, mean ± std) where (a) evaluates the pixel-level overlap accuracy via Dice Score, (b) measures multi-nucleus boundary matches using the Aggregated Jaccard Index (AJI), (c) captures instance detection performance through mean Panoptic Quality (mPQ), and (d) assesses tissue-wide panoptic segmentation boundaries via binary Panoptic Quality (bPQ). Red accent borders identify our proposed SAFViT network. 4.4 Computational Overhead Inference times in Table 1 confirm that all seven models operate within 5.65.6â5.85.8 ms per image, with SAFViT (5.75.7 ms) introducing no measurable latency increase relative to CellViT (5.85.8 ms). The two pointwise convolutions and softmax operation in each SAF gate add a negligible number of parameters relative to the Swin-Tiny encoder and prediction heads. SAFViTâs mPQ improvement therefore comes at effectively zero computational cost, which is important for clinical digital pathology workflows. 5 Conclusion This study introduced SAFViT, a CellViT-based architecture that replaces conventional skip connections with Spatial Attention Fusion (SAF) Gating for nucleus instance segmentation and classification in H&E-stained histopathology images. The core contribution is a per-pixel, softmax-normalised âheatmap of trustâ that jointly modulates encoder and decoder feature streams before fusion, rather than filtering only one stream as in existing gating mechanisms. Through a controlled comparison against six alternative gating strategies under identical training conditions on the PanNuke dataset, we demonstrated that SAF Gating achieves the highest mPQ (0.4710.471) while remaining competitive across Dice, AJI, and bPQ. The per-class analysis revealed that this improvement is driven primarily by SAFViTâs superior detection of the minority Dead class (F1=0.518F_1=0.518), where two competing gating mechanisms (AG and AFF) failed entirely (F1=0.000F_1=0.000). By learning where local boundary detail and global contextual information are each most trustworthy at every spatial position, dual-stream gating preserves subtle morphological cues that single-stream approaches systematically suppress. Evaluation on the MoNuSeg dataset without fine-tuning confirmed that the learned gating behaviour generalises to unseen tissue types, and inference time analysis showed that SAF Gating introduces negligible computational overhead (5.75.7 ms per patch), which supports its use for whole-slide image processing in clinical settings. A limitation of the Dead-class result warrants explicit acknowledgement: the Dead class comprises fewer than 2% of all annotated instances in PanNuke, meaning that a small number of correctly classified instances can produce large relative changes in F1F_1 when the class is this rare. The 14.5-point F1F_1 improvement is consistent across all five cross-validation folds, but evaluation on a dataset with higher Dead-class prevalence would further strengthen this finding. Several directions remain for future investigation. First, the current SAF Gating module uses a fixed bottleneck ratio (2âCâC/2â22Câ C/2â 2) at all decoder levels, an adaptive or level-specific compression strategy could allow the gate network to allocate more capacity at decoder stages where the encoderâdecoder feature gap is largest. Second, while SAFViT employs a Swin-Tiny backbone, the SAF Gating module is architecture-agnostic and could be integrated into larger pretrained encoders, potentially yielding further gains from richer encoder representations. CRediT authorship contribution statement Harshit Mittal: Conceptualization, Methodology, Software, Validation, Investigation, Writing - original draft, Visualization. Arash Rabbani: Supervision, Writing - review & editing, Resources, Project administration. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Data and Code availability The Pannuke Data and MonuSeg Data are publicly available on their respective institutesâ website and the authors of the same have been cited. The complete source code for SAFViT are publicly available in our GitHub repository at https://github.com/itsmittalharshit/SAFViT. References E. Baumann, B. Dislich, J. L. Rumberger, I. D. Nagtegaal, M. R. MartĂnez, and I. Zlobec (2024) HoVer-next: a fast nuclei segmentation and classification pipeline for next generation histopathology. In Proceedings of Machine Learning Research, Vol. 250, p. 61â86. Cited by: §1. F. Bray, M. Laversanne, H. Sung, J. Ferlay, R. L. Siegel, I. Soerjomataram, and A. Jemal (2024) Global cancer statistics 2022: globocan estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: A Cancer Journal for Clinicians 74 (3), p. 229â263. External Links: Document Cited by: §1. H. Cao, Y. Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang (2023) Swin-unet: unet-like pure transformer for medical image segmentation. In Lecture Notes in Computer Science, Vol. 13803 LNCS, Germany, p. 205â218. External Links: Document Cited by: §1. J. Chen, Y. Lu, Q. Yu, X. Luo, E. Adeli, Y. Wang, L. Lu, A. L. Yuille, and Y. Zhou (2021) TransUNet: transformers make strong encoders for medical image segmentation. External Links: Link Cited by: §1. Y. Dai, F. Gieseke, S. Oehmcke, Y. Wu, and K. Barnard (2021) Attentional feature fusion. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), p. 3560â3569. External Links: Document Cited by: §4.1. J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei (2009) ImageNet: a large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 248â255. External Links: Document Cited by: §2. T. N. N. Doan, B. Song, T. T. L. Vuong, K. Kim, and J. T. Kwak (2022) SONNET: a self-guided ordinal regression neural network for segmentation and classification of nuclei in large-scale multi-tissue histology images. IEEE Journal of Biomedical and Health Informatics 26, p. 3218â3228. External Links: Document Cited by: §1. A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby (2021) An image is worth 16x16 words: transformers for image recognition at scale. Cited by: §1. J. Gamper, N. A. Koohbanani, K. Benet, A. Khuram, and N. Rajpoot (2019) PanNuke: an open pan-cancer histology dataset for nuclei instance segmentation and classification. In Lecture Notes in Computer Science (including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Vol. 11435 LNCS, United Kingdom, p. 11â19. External Links: Document Cited by: §1, §3.1. S. Graham, Q. D. Vu, S. E. A. Raza, A. Azam, Y. W. Tsang, J. T. Kwak, and N. Rajpoot (2019) HoVer-net: simultaneous segmentation and classification of nuclei in multi-tissue histology images. Medical Image Analysis 58, p. 101563. External Links: Document Cited by: §1, §1, §2. F. Hörst, M. Rempe, H. Becker, L. Heine, J. Keyl, and J. Kleesiek (2026) CellViT++: energy-efficient and adaptive cell segmentation and classification using foundation models. Computer Methods and Programs in Biomedicine 277, p. 109206. External Links: Document Cited by: §1. F. Hörst, M. Rempe, L. Heine, C. Seibold, J. Keyl, G. Baldini, S. Ugurel, J. Siveke, B. GrĂŒnwald, J. Egger, and J. Kleesiek (2024) CellViT: vision transformers for precise cell segmentation and classification. Medical Image Analysis 94, p. 103143. External Links: Document Cited by: §1, §4.1. J. Hu, L. Shen, and G. Sun (2018) Squeeze-and-excitation networks. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, USA, p. 7132â7141. External Links: Document Cited by: §1, §2.1, §4.1. X. Jia, S. Jian, Y. Tan, Y. Che, W. Chen, and Z. Liang (2024) Gated cross-attention network for depth completion. External Links: Link Cited by: §4.1. T. L. B. Khanh, D. P. Dao, N. H. Ho, H. J. Yang, E. T. Baek, G. Lee, S. H. Kim, and S. B. Yoo (2020) Enhancing u-net with spatial-channel attention gate for abnormal tissue segmentation in medical imaging. Applied Sciences (Switzerland) 10. External Links: Document Cited by: §1. N. Kumar, R. Verma, D. Anand, Y. Zhou, O. F. Onder, E. Tsougenis, H. Chen, P. A. Heng, J. Li, Z. Hu, Y. Wang, N. A. Koohbanani, M. Jahanifar, N. Z. Tajeddin, A. Gooya, N. Rajpoot, X. Ren, S. Zhou, Q. Wang, D. Shen, C. K. Yang, C. H. Weng, W. H. Yu, C. Y. Yeh, S. Yang, S. Xu, P. H. Yeung, P. Sun, A. Mahbod, G. Schaefer, I. Ellinger, R. Ecker, O. Smedby, C. Wang, B. Chidester, T. V. Ton, M. T. Tran, J. Ma, M. N. Do, S. Graham, Q. D. Vu, J. T. Kwak, A. Gunda, R. Chunduri, C. Hu, X. Zhou, D. Lotfi, R. Safdari, A. Kascenas, A. OâNeil, D. Eschweiler, J. Stegmaier, Y. Cui, B. Yin, K. Chen, X. Tian, P. Gruening, E. Barth, E. Arbel, I. Remer, A. Ben-Dor, E. Sirazitdinova, M. Kohl, S. Braunewell, Y. Li, X. Xie, L. Shen, J. Ma, K. D. Baksi, M. A. Khan, J. Choo, A. Colomer, V. Naranjo, L. Pei, K. M. Iftekharuddin, K. Roy, D. Bhattacharjee, A. Pedraza, M. G. Bueno, S. Devanathan, S. Radhakrishnan, P. Koduganty, Z. Wu, G. Cai, X. Liu, Y. Wang, and A. Sethi (2020) A multi-organ nucleus segmentation challenge. IEEE Transactions on Medical Imaging 39, p. 1380â1391. External Links: Document Cited by: §1, §3.2. I. Levner and H. Zhang (2007) Classification-driven watershed segmentation. IEEE Transactions on Image Processing 16, p. 1437â1445. External Links: Document Cited by: §1. Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo (2021) Swin transformer: hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 9992â10002. External Links: Document Cited by: §2. I. Loshchilov and F. Hutter (2017) SGDR: stochastic gradient descent with warm restarts. External Links: Link Cited by: §3.1. I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. External Links: Link Cited by: §3.1. E. Meijering (2012) Cell segmentation: 50 years down the road [life sciences]. IEEE Signal Processing Magazine 29, p. 140â145. External Links: Document Cited by: §1. O. Oktay, J. Schlemper, L. L. Folgoc, M. Lee, M. Heinrich, K. Misawa, K. Mori, S. McDonagh, N. Y. Hammerla, B. Kainz, B. Glocker, and D. Rueckert (2018) Attention u-net: learning where to look for the pancreas. Cited by: §1, §2.1, §4.1. O. Ronneberger, P. Fischer, and T. Brox (2015) U-net: convolutional networks for biomedical image segmentation. In Lecture Notes in Computer Science (including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Vol. 9351, Germany, p. 234â241. External Links: Document Cited by: §1. U. Schmidt, M. Weigert, C. Broaddus, and G. Myers (2018) Cell detection with star-convex polygons. In Lecture Notes in Computer Science (including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Vol. 11071 LNCS, Spain, p. 265â273. External Links: Document Cited by: §1. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin (2017) Attention is all you need. Cited by: §1. T. Vicar, J. Balvan, J. Jaros, F. Jug, R. Kolar, M. Masarik, and J. Gumulec (2019) Cell segmentation methods for label-free contrast microscopy: review and comprehensive comparison. BMC Bioinformatics 20. External Links: Document Cited by: §1. A. Voulodimos, N. Doulamis, A. Doulamis, and E. Protopapadakis (2018) Deep learning for computer vision: a brief review. Computational Intelligence and Neuroscience. External Links: Document Cited by: §1. S. Woo, J. Park, J. Lee, and I. S. Kweon (2018) CBAM: convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, p. 3â19. External Links: Document Cited by: §4.1.