Paper deep dive
MMA-Former: Multi-Window Mixture-of-Head Attention Transformer for Adaptive PNI Prediction in 3D MRI
Youngung Han, Induk Um, Kyeonghun Kim, Junga Kim, Hyunsu Go, Jaewon Jung, Woo Kyoung Jeong, Won Jae Lee, Pa Hong, Ken Ying-Kai Liao, Hyuk-Jae Lee, Nam-Joon Kim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/18/2026, 1:46:39 PM
Summary
The paper introduces MMA-Former, a novel 3D Transformer architecture for non-invasive prediction of Perineural Invasion (PNI) in cholangiocarcinoma using T1-weighted MRI. The model utilizes a Coarse-Fine Transformer (CFT) structure for multi-scale feature extraction and a Window-Specific Mixture-of-Head (WS-MoH) attention mechanism that dynamically routes entire 3D windows to specialized or common attention heads. Evaluated on a dataset of 168 patients, MMA-Former achieved an AUC of 0.752, outperforming standard CNNs and Transformers.
Entities (9)
Relation Signals (8)
MMA-Former → achievesmetric → 0.752
confidence 95% · MMA-Former achieved an AUC of 0.752
MMA-Former → proposessolutionfor → Perineural Invasion
confidence 95% · We propose the Multi-window Mixture-of-Head Attention Transformer (MMA-Former)... for Adaptive PNI Prediction
MMA-Former → usescomponent → Coarse-Fine Transformer
confidence 92% · featuring a Coarse-Fine Transformer (CFT) structure
MMA-Former → usescomponent → WS-MoH
confidence 92% · integrating a novel Window-Specific Mixture-of-Head attention (WS-MoH) mechanism
MMA-Former → outperforms → ResNet
confidence 90% · outperforming other 3D architectures, including the best CNN (AUC of 0.708)
MMA-Former → outperforms → Swin Transformer
confidence 90% · outperforming other 3D architectures... and Transformer baselines (AUC of 0.681)
WS-MoH → replaces → Multi-head Self-Attention
confidence 90% · Unlike standard Multi-Head Self Attention (MSA), WS-MoH generates a representation for each 3D window
MMA-Former → evaluatedondataset → Samsung Medical Center
confidence 80% · Evaluated on a retrospective dataset... acquired from... Samsung Medical Center
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. Non-invasive prediction from 3D MRI is challenging, demanding models that efficiently capture both fine-grained details and global context. We propose the Multi-window Mixture-of-Head Attention Transformer (MMA-Former), a novel end-to-end 3D architecture featuring a Coarse-Fine Transformer (CFT) structure for parallel multi-scale feature extraction. We advance this structure by integrating a novel Window-Specific Mixture-of-Head attention (WS-MoH) mechanism. Unlike standard Multi-Head Self Attention (MSA), WS-MoH generates a representation for each 3D window and dynamically routes the entire window to specialized or common attention heads. This enables spatially adaptive feature extraction tailored to the local context of each window, enhancing specialization and reducing redundancy without increasing parameters. Evaluated on a retrospective dataset of 168 T1-weighted MRI scans, MMA-Former achieved an AUC of 0.752, outperforming other 3D architectures, including the best CNN (AUC of 0.708) and Transformer baselines (AUC of 0.681).
Tags
Links
- Source: https://arxiv.org/abs/2607.10988v1
- Canonical: https://arxiv.org/abs/2607.10988v1
Trouble viewing inline? Open PDF directly →
Full Text
23,316 characters extracted from source content.
Expand or collapse full text
MMA-FORMER: MULTI-WINDOW MIXTURE-OF-HEAD ATTENTION TRANSFORMER FOR ADAPTIVE PNI PREDICTION IN 3D MRI Youngung Han 1,2 , Induk Um 3 , Kyeonghun Kim 2 , Junga Kim 1 , Hyunsu Go 1 , Jaewon Jung 1 , Woo Kyoung Jeong 4 , Won Jae Lee 5 , Pa Hong 5 , Ken Ying-Kai Liao 6 , Hyuk-Jae Lee 1 , Nam-Joon Kim 1,† 1 Seoul National University, Seoul, Republic of Korea 2 OUTTA, Seoul, Republic of Korea 3 Chung-Ang University, Seoul, Republic of Korea 4 Samsung Medical Center, Sungkyunkwan University School of Medicine, Seoul, Republic of Korea 5 Samsung Changwon Hospital, Changwon, Republic of Korea 6 NVIDIA AI Technology Center, Taipei, Taiwan † Corresponding author: knj01@snu.ac.kr ABSTRACT Perineural invasion (PNI) is a critical prognostic factor in cholangio- carcinoma. Non-invasive prediction from 3D MRI is challenging, demanding models that efficiently capture both fine-grained details and global context. We propose the Multi-window Mixture-of-Head Attention Transformer (MMA-Former), a novel end-to-end 3D ar- chitecture featuring a Coarse-Fine Transformer (CFT) structure for parallel multi-scale feature extraction. We advance this structure by integrating a novel Window-Specific Mixture-of-Head attention (WS-MoH) mechanism. Unlike standard Multi-Head Self Attention (MSA), WS-MoH generates a representation for each 3D window and dynamically routes the entire window to specialized or com- mon attention heads. This enables spatially adaptive feature extrac- tion tailored to the local context of each window, enhancing spe- cialization and reducing redundancy without increasing parameters. Evaluated on a retrospective dataset of 168 T1-weighted MRI scans, MMA-Former achieved an AUC of 0.752, outperforming other 3D architectures, including the best CNN (AUC of 0.708) and Trans- former baselines (AUC of 0.681). Index Terms— Vision Transformer, Mixture-of-Head attention, Adaptive Feature Extraction, Window-level Routing 1. INTRODUCTION Perineural invasion (PNI), the insidious infiltration of cancer cells along nerve sheaths, is a critical route of metastasis in cholangiocar- cinoma. Its presence significantly escalates the risk of recurrence, correlates with poor survival, and dictates surgical planning [1, 2, 3]. Accurate preoperative identification of PNI is therefore paramount for personalized treatment. Despite its clinical urgency, non-invasive PNI prediction re- mains challenging due to subtle MRI features [4, 5]. Furthermore, most PNI studies are hampered by small cohorts, often numbering in the low hundreds [2, 6], which frequently leads to class imbalance. Methodologically, existing approaches often rely on radiomic features [7], which may fail to capture complex 3D spatial patterns. While end-to-end deep learning offers potential, standard architec- tures struggle. CNNs [8, 9] are limited in modeling long-range de- pendencies, while Vision Transformers (ViT) [10, 11] and their 3D adaptations [12] incur prohibitive computational costs in 3D MRI. AxialCoronalSagittal (a) (b) Fig. 1. Grad-CAM visualizations, highlighting critical regions at the tumor interface. (a) PNI-positive and (b) PNI-negative cases local- ized by MMA-Former. Hierarchical approaches like Swin Transformer [13] improve efficiency by computing attention within local windows. Build- ing on this, architectures incorporating parallel multi-scale window processing, such as MViT [14] or Focal Transformers [15], have emerged to explicitly capture features at different scales simultane- ously. However, the uniform processing of standard Multi-Head Self Attention (MSA) in these methods limits adaptive feature selection and causes attention head redundancy.[16, 17, 18]. The Mixture-of-Head attention (MoH) mechanism [16], inspired by Mixture-of-Experts (MoE) principles [19, 20], addresses this re- dundancy by treating attention heads as experts. MoH employs a router to dynamically select a subset of specialized heads for each input. This enhances specialization and efficiency. However, MoH typically operates at the token level, which remains computationally intensive for 3D data. We propose the Multi-window Mixture-of-Head Attention Transformer (MMA-Former), an adaptive 3D architecture that syn- ergizes multi-scale processing with a novel adaptation of MoH. We introduce the Coarse-Fine Transformer (CFT) structure to capture parallel multi-scale information. Crucially, we integrate Window- Specific MoH (WS-MoH). Instead of routing every token, WS-MoH dynamically routes the entire window to specialized heads. This arXiv:2607.10988v1 [cs.CV] 13 Jul 2026 (a)(b)(c) WS-MoHWS-MoHSWS-MoHSWS-MoH Multi-Head Cross- Attention Downsampling Concat Multi-Head Cross- Attention Spatial Attention Spatial Attention Patch Merging MMA Block Input Image Patch Partioning Linear Embedding Encoding phaseK × Stage Xpace Block MLP MMA Block LNLNLNLN LNLN MLPMLP Sub-Block 퓁 f P fc s f Q K V K Q Sub-Block 퓁+1 MMA Block Xpace Block s+1 P fc Fig. 2. Overview of the MMA-Former architecture. (a) The framework utilizes a hierarchical structure incorporating Multi-window Mixture- of-Head Attention (MMA) blocks and Cross-Spatial Attention (Xpace) blocks. (b) The MMA block implements the CFT structure across consecutive blocks (block l and l + 1). It utilizes parallel Coarse (P c ) and Fine (P f ) pathways with different window sizes, employing Window-Specific MoH (WS-MoH) and Shifted Window-Specific MoH (SWS-MoH). (c) The Xpace block for hierachical feature fusion through Multi-Head Cross-Attention and Spatial Attention. enables spatially adaptive feature extraction where different regions leverage different attention heads efficiently. Our main contributions are these: • MMA-Former, a novel 3D end-to-end architecture utilizing a CFT structure for parallel multi-scale feature extraction for PNI prediction. • WS-MoH, a novel approach applying MoH routing at the win- dow level, enabling adaptive feature extraction based on the spa- tial context of the window. • Xpace Block, an encoder utilizing hierarchical feature fusion be- tween different stages. 2. METHODOLOGY 2.1. Datasets and Preprocessing We utilized a retrospective dataset of anonymized T1-weighted, contrast-enhanced MR images (hepatobiliary phase) acquired from multiple MRI scanners at Samsung Medical Center over a decade. Images were provided in NIfTI format with ground truth annotations for the liver and tumor. After quality control, the final analysis co- hort comprised 168 cholangiocarcinoma patients. The presence of PNI was confirmed by post-surgical histopathological examination, comprising 67 PNI-positive cases and 101 PNI-negative cases. We employed a localization strategy, extracting cropped vol- umes of size 96× 96× 48 centered on the tumor. The necessity of this localization is validated in Table 2 (A). 2.2. MMA-Former Architecture The MMA-Former, illustrated in Fig. 2 (a), employs a hierarchical design. Input 3D volumes undergo patch partitioning and linear em- bedding. The architecture consists of an encoding phase followed by K stages. Each stage comprises patch merging layer and an MMA block for feature extraction. An Xpace block, shown in Fig. 2 (c), fuses hierarchical features across stages by integrating downsampled representations from the previous stage with features from the cur- rent stage. 2.3. The MMA Block: CFT with WS-MoH 2.3.1. CFT Structure The Coarse-Fine Transformer (CFT) structure captures multi-scale information by processing features through parallel pathways with different window sizes. Given an input feature map z l−1 at sub-block l, we apply layer normalization (LN). The normalized features z ′ are processed in par- allel through two pathways. The Fine path, denoted as P f , utilizes a small window size W F (e.g., 3× 3× 3). The Coarse path, denoted as P c , employs a larger window size W C (e.g., 6× 6× 6). Both pathways use WS-MoH as the attention mechanism. A F = WS-MoH W F (z ′ ), A C = WS-MoH W C (z ′ )(1) The outputs are fused via concatenation and linear projection (Fu- sion), then combined with the input via residual connections, fol- lowed by an MLP: A fused = Fusion(A F ,A C ) + z l−1 (2) z l = MLP(LN(A fused )) + A fused (3) Subsequent blocks (Block l+1) employ shifted window partitioning (SWS-MoH) [13] for cross-window communication (Fig. 2 (b)). 2.3.2. WS-MoH We replace standard MSA with the adaptive Window-Specific MoH (WS-MoH), detailed in Fig. 3. MSA (Eq. 4) activates all h heads uniformly, regardless of the input context. MSA(X) = h X i=1 H i (X)W i O (4) MoH [16] introduces a router to dynamically weight heads, de- noted as g i . We propose WS-MoH, which adapts this routing to the window level for efficiency in 3D data. First, standard Q, K, V projections are computed for the input features within a window X W . Q = X W W Q , K = X W W K , V = X W W V (5) h h h h h h Attention Router Attention Top-K Router Attention Top-K WS-MoHW-MSASWS-MoH Shared Heads Routed Heads s s+1 h s h s s h h h h s s h h h 1 1 h s+1 h h 1 1 h s+1 h s+1 h s+2 h s+2 h s+3 h s+3 h s+1 h s+1 h s+2 h s+2 h s+3 h s+3 h h 1 1 (a)(b) (c) Fig. 3. Comparison of attention mechanisms. (a) W-MSA: All heads are activated uniformly for the window. (b) WS-MoH (Window-Specific MoH): A router utilizes window-representative features to select a Top-K subset of Routed Heads, which are combined with always-active Shared Heads. This allows different windows to utilize different head combinations. (c) SWS-MoH (Shifted Window-Specific MoH): WS- MoH applied to a shifted window configuration. As the window is shifted, the router dynamically activates different head combinations for each respective window. Crucially, instead of routing each token, we generate a window representative feature X rep by average pooling the features X W . The router uses this single representative feature X rep to deter- mine the routing scores g i for the entire window. The final WS-MoH output is the weighted sum of the outputs from the selected heads: WS-MoH(X W ) = h X i=1 g i (X rep )H i (X W )W i O (6) This enables spatially adaptive feature extraction, as different windows activate different combinations of heads based on their lo- cal context. We employ the two-stage routing strategy [16]. Heads are di- vided into Shared Heads, denoted as H S , which capture common knowledge and are always active, and Routed Heads, denoted asH R , which handle specialized patterns. A router network processesX rep to produce probabilities for the routed heads, and the Top-K heads are selected. P R = S(W r X rep )(7) The routing score g i (Eq. 8) is determined by balancing the con- tributions of H S and the Top-K selected H R via coefficients α 1 ,α 2 , where [α 1 ,α 2 ] = S(W h X rep ) and S(·) denotes the Softmax func- tion. g i (X rep ) = α 1 S(W s X rep ) i , if i∈ H S α 2 (P R ) i ,if i∈ Top-K(H R ) 0,otherwise. (8) 2.3.3. Load Balance Loss In mixture models utilizing routing, training often leads to an im- balance where a few heads process the majority of inputs, leaving others undertrained [19]. To mitigate this in WS-MoH and ensure effective utilization of all Routed Heads for robust PNI prediction, Table 1. AUC comparison of PNI classifiers on cropped 3D images (5-fold CV). CategoryModel (3D)Mean AUC CNNResNet [8]0.708 DenseNet [9]0.688 EfficientNet[21]0.673 TransformerSwin Transformer [13]0.681 MMA-Former (CFT + WS-MoH)0.752 we incorporate a load balance loss L LB . This loss encourages an even distribution of windows across the Routed Heads. L LB = X i∈H R P i · f i (9) where P i is the average routing probability for head i, and f i is the fraction of windows that selected head i. 2.4. Xpace Block The Cross-Spatial Attention (Xpace) block fuses features across en- coder stages. First, it downsamples previous-stage features. The fusion mechanism operates through bidirectional cross-attention be- tween downsampled and current-stage features. Spatial attention is applied to highlight important regions in both representations. The resulting features are then concatenated along the channel dimension and forwarded to the next stage. 3. EXPERIMENTS AND RESULTS 3.1. Experimental Setup All experiments were performed on NVIDIA A100 GPUs. We em- ployed stratified 5-fold cross-validation across the 168 cases. Mod- Table 2. Comprehensive ablation study of MMA-Former components and configurations on the PNI dataset. CategoryConfigurationDescriptionAUC BaselineMMA-Former (Full Model)Cropped Input, CFT (2 paths), WS-MoH (S=2, Top-K 75%), Xpace0.752 (A) Input StrategyUncropped InputMMA-Former trained on whole MRI volumes0.705 Cropped Input (Default)MMA-Former trained on localized tumor volumes (96× 96× 48)0.752 (B) Component Ablationw/o WS-MoH (CFT+MSA)WS-MoH replaced by standard MSA (All heads active)0.731 w/o CFT StructureStandard Transformer blocks + WS-MoH (Single window path)0.690 w/o Xpace BlockHierarchical features not explicitly fused via Xpace0.709 (C) CFT: Window PathsFine Path OnlySingle path utilizing only the small window size0.704 Coarse Path OnlySingle path utilizing only the large window size0.718 Fine + Coarse (Default)Parallel Fine + Coarse paths0.752 (D) WS-MoH: Shared Heads (S)S=0No Shared Heads; Only Routed heads active (Top-K 75%)0.748 (Top-K ratio fixed at 75%)S=11 Shared head + Routed heads (Top-K 75%)0.737 S=2 (Default)2 Shared heads + Routed heads (Top-K 75%)0.752 S=AllAll Shared; No routing (Equivalent to MSA)0.731 (E) WS-MoH: Top-K Ratio50%S=2, Top-K 50% of Routed heads active0.743 (S fixed at 2)75% (Default)S=2, Top-K 75% of Routed heads active0.752 90%S=2, Top-K 90% of Routed heads active0.733 els were trained using the AdamW optimizer [22] with a learning rate of 8× 10 −5 and a batch size of 4. The total loss function com- bines the Weighted BCE Loss, denoted asL W CE , utilizing a posi- tive weight of 1.5 to address class imbalance, and the load balance lossL LB , weighted by a factor β = 0.01. L T otal =L W CE + βL LB (10) 3.2. PNI Prediction Performance Table 1 summarizes the performance comparison. MMA-Former achieved the highest mean AUC of 0.752. It significantly outper- formed the best-performing CNN, ResNet (which achieved an AUC of 0.708), and the standard 3D Swin Transformer (which achieved an AUC of 0.681). 3.3. Ablation Study We conducted ablation studies to validate the design choices of MMA-Former (Table 2). The default configuration uses 2 Shared heads (S=2) and a Top-K ratio of 75%. Impact of Input Strategy (A): Training MMA-Former on un- cropped images decreased the AUC by 0.047. This confirms that localization to the tumor region is crucial for detecting the subtle features of PNI, validating our preprocessing strategy. Impact of Core Components (B): Replacing WS-MoH with standard MSA decreased the AUC by 0.021. This highlights that the adaptive, window-specific feature extraction provided by WS- MoH contributes to performance. Removing the CFT structure also reduced performance, resulting in an AUC of 0.690. Additionally, removing the Xpace Block resulted in degraded performance, low- ering the AUC to 0.709. Impact of CFT Configuration (C): The parallel configuration achieved the best performance, outperforming both the Fine path only (AUC of 0.704) and the Coarse path only (AUC of 0.718). This validates the CFT structure, confirming that synergizing local and global contexts outperforms single-scale approaches. Impact of WS-MoH Shared Heads (D): The configuration with 2 Shared heads provided the optimal balance. Relying solely on Routed heads failed to capture common features effectively, while using only Shared heads lacked adaptive specialization. Impact of WS-MoH Top-K Ratio (E): Activating 75% of the routed heads yielded the best result. Using fewer heads reduced ca- pacity, while using more likely reintroduced redundancy. 3.4. Qualitative Results and Visualization We utilized 3D Grad-CAM [23] to visualize the regions most influ- ential to the model’s predictions (Fig. 1). The visualizations confirm that MMA-Former focuses precisely on the tumor and the immediate peritumoral interface, the areas most critical for identifying PNI. 4. DISCUSSION AND CONCLUSION The MMA-Former introduces a novel approach to 3D medical image analysis by enabling spatially adaptive feature extraction within a hi- erarchical transformer framework. The effectiveness of the approach is strongly supported by the necessity of input localization. The core innovation is the WS-MoH within the CFT framework. The CFT structure provides multi-scale context. WS-MoH signifi- cantly outperforms CFT+MSA via window-level adaptation. Dy- namically routing window representations to specialized heads cap- tures nuanced local context. This contrasts with MSA’s uniform pro- cessing and avoids the computational burden of token-level MoH. The primary limitation is the single-institution nature of the dataset, comprising 168 cases. Although the cohort size is com- parable to related PNI studies, external validation on multi-center datasets is required to ensure generalizability. In conclusion, we proposed the MMA-Former, an adaptive 3D Transformer for PNI prediction. By synergizing the CFT structure with a novel WS-MoH, MMA-Former effectively captures multi- scale context while enabling efficient, spatially adaptive feature ex- traction. Our experiments demonstrate the superiority of MMA- Former (AUC of 0.752) over existing 3D backbones (best AUC of 0.708), highlighting the benefit of window-level adaptive attention for complex 3D medical imaging tasks. 5. ACKNOWLEDGMENTS This work was supported by the Institute of Information & Commu- nications Technology Planning & Evaluation (IITP), funded by the Korea government (MSIT), under the Artificial Intelligence Semi- conductor Support Program to nurture the best talents (IITP-2023- RS-2023-00256081) and the grant for the Development of an AI Deep Learning Processor and Module for a 2,000 TFLOPS Server (No. 2020-0-01305) 6. REFERENCES [1] M. Zou, J. Sheng, M. Ruan, W. Zhou, F. Ye, G. Yang, Y. Qian, J. Wang, R. Wang, S. Liu et al., “Perineural invasion con- fers poorer clinical outcomes in patients with t1/t2 intrahep- atic cholangiocarcinoma: a single center, retrospective cohort study,” Journal of Gastrointestinal Oncology, vol. 14, no. 6, p. 2500, 2023. [2] Z. Zhang, Y. Zhou, K. Hu, D. Wang, Z. Wang, and Y. Huang, “Perineural invasion as a prognostic factor for intrahepatic cholangiocarcinoma after curative resection and a potential in- dication for postoperative chemotherapy: a retrospective co- hort study,” Bmc Cancer, vol. 20, no. 1, p. 270, 2020. [3] Z. Liu, C. Luo, X. Chen, Y. Feng, J. Feng, R. Zhang, F. Ouyang, X. Li, Z. Tan, L. Deng et al., “Noninvasive predic- tion of perineural invasion in intrahepatic cholangiocarcinoma by clinicoradiological features and computed tomography ra- diomics based on interpretable machine learning: a multicenter cohort study,” International Journal of Surgery, vol. 110, no. 2, p. 1039–1051, 2024. [4] S. Doran, R. Whiriskey, N. Sheehy, C. Johnston, and D. Byrne, “Perineural tumour spread in head and neck cancer: a picto- rial review,” Clinical Radiology, vol. 79, no. 10, p. 749–756, 2024. [5] Z. Qi, H. Yuan, Q. Li, P. Chen, D. Li, K. Chen, B. Meng, P. Ning, H. Yu, and D. Li, “An mri-based fusion model for preoperative prediction of perineural invasion status in patients with intrahepatic cholangiocarcinoma,” World Journal of Sur- gical Oncology, vol. 23, no. 1, p. 164, 2025. [6] P.-C. Zhan, P.-j. Lyu, Z. Li, X. Liu, H.-X. Wang, N.-N. Liu, Y. Zhang, W. Huang, Y. Chen, and J.-b. Gao, “Ct-based ra- diomics analysis for noninvasive prediction of perineural inva- sion of perihilar cholangiocarcinoma,” Frontiers in Oncology, vol. 12, p. 900478, 2022. [7] X. Huang, J. Shu, Y. Yan, X. Chen, C. Yang, T. Zhou, and M. Li, “Feasibility of magnetic resonance imaging-based ra- diomics features for preoperative prediction of extrahepatic cholangiocarcinoma stage,” European Journal of Cancer, vol. 155, p. 227–235, 2021. [8] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, p. 770– 778. [9] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proceedings of the IEEE conference on computer vision and pattern recog- nition, 2017, p. 4700–4708. [10] A. Dosovitskiy, “An image is worth 16x16 words: Trans- formers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020. [11] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Sys- tems, vol. 30, 2017, p. 6000–6010. [12] A. Hatamizadeh, Y. Tang, V. Nath, D. Yang, A. Myronenko, B. Landman, H. R. Roth, and D. Xu, “Unetr: Transformers for 3d medical image segmentation,” in Proceedings of the IEEE/CVF winter conference on applications of computer vi- sion, 2022, p. 574–584. [13] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF in- ternational conference on computer vision, 2021, p. 10 012– 10 022. [14] H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” in Pro- ceedings of the IEEE/CVF international conference on com- puter vision, 2021, p. 6824–6835. [15] J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao, “Focal self-attention for local-global interactions in vi- sion transformers,” arXiv preprint arXiv:2107.00641, 2021. [16] P. Jin, B. Zhu, L. Yuan, and S. Yan, “Moh:Multi- head attention as mixture-of-head attention,” arXiv preprint arXiv:2410.11842, 2024. [17] E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov, “Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned,” in Proceedings of the 57th Annual Meeting of the Association for Computational Lin- guistics, 2019, p. 5797–5808. [18] P. Michel, O. Levy, and G. Neubig, “Are sixteen heads really better than one?” in Advances in Neural Information Process- ing Systems, vol. 32, 2019, p. 14 014–14 024. [19] N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in Proceedings of the International Conference on Learning Representations, 2017. [20] W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research, vol. 23, no. 120, p. 1–39, 2022. [21] M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in International conference on machine learning. PMLR, 2019, p. 6105–6114. [22] I. Loshchilov and F. Hutter, “Decoupled weight decay regular- ization,” arXiv preprint arXiv:1711.05101, 2017. [23] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep net- works via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision, 2017, p. 618–626.