Paper deep dive
MAESIL: Masked Autoencoder for Enhanced Self-supervised Medical Image Learning
Kyeonghun Kim, Hyeonseok Jung, Youngung Han, Junsu Lim, YeonJu Jean, Seongbin Park, Eunseob Choi, Hyunsu Go, SeoYoung Ju, Seohyoung Park, Gyeongmin Kim, MinJu Kwon, KyungSeok Yuh, Soo Yong Kim, Ken Ying-Kai Liao, Nam-Joon Kim, Hyuk-Jae Lee
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/2/2026, 3:28:33 AM
Summary
MAESIL is a novel self-supervised learning framework for 3D medical imaging (specifically CT scans) that addresses the limitations of 2D-slice-based methods. By introducing a 'superpatch' input unit and a dual-masking strategy, MAESIL effectively captures 3D structural context and axial coherence, outperforming traditional baselines like AE, VAE, and VQ-VAE in reconstruction metrics such as PSNR and SSIM.
Entities (6)
Relation Signals (3)
MAESIL → uses → Superpatch
confidence 98% · The core innovation is the 'superpatch', a 3D chunk-based input unit
MAESIL → evaluatedon → BTCV
confidence 95% · We validated our approach on three diverse large-scale public CT datasets... BTCV
MAESIL → outperforms → AE
confidence 95% · MAESIL demonstrates significant improvements over existing methods such as AE, VAE and VQ-VAE
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Training deep learning models for three-dimensional (3D) medical imaging, such as Computed Tomography (CT), is fundamentally challenged by the scarcity of labeled data. While pre-training on natural images is common, it results in a significant domain shift, limiting performance. Self-Supervised Learning (SSL) on unlabeled medical data has emerged as a powerful solution, but prominent frameworks often fail to exploit the inherent 3D nature of CT scans. These methods typically process 3D scans as a collection of independent 2D slices, an approach that fundamentally discards critical axial coherence and the 3D structural context. To address this limitation, we propose the autoencoder for enhanced self-supervised medical image learning(MAESIL), a novel self-supervised learning framework designed to capture 3D structural information efficiently. The core innovation is the 'superpatch', a 3D chunk-based input unit that balances 3D context preservation with computational efficiency. Our framework partitions the volume into superpatches and employs a 3D masked autoencoder strategy with a dual-masking strategy to learn comprehensive spatial representations. We validated our approach on three diverse large-scale public CT datasets. Our experimental results show that MAESIL demonstrates significant improvements over existing methods such as AE, VAE and VQ-VAE in key reconstruction metrics such as PSNR and SSIM. This establishes MAESIL as a robust and practical pre-training solution for 3D medical imaging tasks.
Tags
Links
- Source: https://arxiv.org/abs/2604.00514v1
- Canonical: https://arxiv.org/abs/2604.00514v1
Trouble viewing inline? Open PDF directly →
Full Text
20,771 characters extracted from source content.
Expand or collapse full text
MAESIL: Masked Autoencoder for Enhanced Self-supervised Medical Image Learning Kyeonghun Kim OUTTA kyeonghun.kim@outta.ai Hyeonseok Jung Chung-Ang University mjmk0820@cau.ac.kr Youngung Han Seoul National University yuhan@snu.ac.kr Junsu Lim Sangmyung University 202115042@sangmyung.ac.kr YeonJu Jean Ewha Womans University ahxlzjt@ewhain.net Seongbin Park Seoul National University tjdqls@snu.ac.kr Eunseob Choi GIST eunseobchoi@gm.gist.ac.kr Hyunsu Go Seoul National University hsmail02@snu.ac.kr Seoyoung Ju Sangmyung University 202115055@sangmyung.ac.kr Seohyoung Park Ewha Womans University 03nobel@ewhain.net Gyeongmin Kim Chung-Ang University rnrn6442@cau.ac.kr MinJu Kwon Chung-Ang University kmjcap4@cau.ac.kr Kyungseok Yuh Dankook University 32192743@dankook.ac.kr Soo Yong Kim AI Matics ksyint@aimatics.ai Ken Ying-Kai Liao NVIDIA ATC, Taiwan kenyingkail@nvidia.com Nam-Joon Kim † Seoul National University knj01@snu.ac.kr Hyuk-Jae Lee Seoul National University hjlee@capp.snu.ac.kr Abstract—Training deep learning models for three-dimensional (3D) medical imaging, such as Computed Tomography (CT), is fundamentally challenged by the scarcity of labeled data. While pre-training on natural images is common, it results in a significant domain shift, limiting performance. Self-Supervised Learning (SSL) on unlabeled medical data has emerged as a powerful solution, but prominent frameworks often fail to exploit the inherent 3D nature of CT scans. These methods typically process 3D scans as a collection of independent 2D slices, an approach that fundamentally discards critical axial coherence and the 3D structural context. To address this limitation, we pro- pose the autoencoder for enhanced self-supervised medical image learning(MAESIL), a novel self-supervised learning framework designed to capture 3D structural information efficiently. The core innovation is the ‘superpatch,’ a 3D chunk-based input unit that balances 3D context preservation with computational effi- ciency. Our framework partitions the volume into superpatches and employs a 3D masked autoencoder strategy with a dual- masking strategy to learn comprehensive spatial representations. We validated our approach on three diverse large-scale public CT datasets. Our experimental results show that MAESIL demonstrates significant improvements over existing methods such as AE, VAE and VQ-VAE in key reconstruction metrics such as PSNR and SSIM. This establishes MAESIL as a robust and practical pre-training solution for 3D medical imaging tasks. Index Terms—Self-Supervised Learning, Medical Imaging, Masked Autoencoder, 3D CT, Representation Learning I. INTRODUCTION Deep learning techniques in medical imaging are rapidly becoming essential for clinical diagnostics and treatment planning [1]. Among various imaging modalities, Computed † Corresponding author Tomography (CT) is widely utilized as it provides three- dimensional (3D) scans of complex anatomical structures within the human body, making it indispensable for compre- hensive disease detection. However, fully leveraging the potential of these CT scans with deep learning faces several fundamental challenges. First, the scarcity of labeled data remains the most significant barrier to training deep learning models [2]. The expert an- notation process required for creating large-scale datasets is exceptionally time-consuming and costly. Second, to circum- vent this issue, models pre-trained on natural images (e.g., ImageNet [3]) are often used. However, this approach leads to a domain shift problem, as the visual representations learned from natural images are fundamentally different from the unique characteristics of CT imaging, resulting in limited performance [4]. To address these core problems, Self-Supervised Learning (SSL) has emerged as a powerful alternative that utilizes large- scale unlabeled data [5]. In particular, pre-training on domain- specific datasets (i.e., medical images) rather than natural images has proven effective in mitigating the domain shift problem, demonstrating the strong potential of this direction. However, these preceding approaches do not fully exploit the potential of 3D CT volumes, which is the inherent nature of CT scans. Many existing methods, including prominent SSL frameworks, treat the 3D CT scans as a mere collection of independent 2D slices for pre-training. While this 2D- based approach may be computationally convenient, it has an intrinsic limitation: it fails to capture the axial coherence and 3D structural context between slices, which are critical for accurate CT analysis [6]. arXiv:2604.00514v1 [cs.CV] 1 Apr 2026 Superpatch Reconstructed Superpatch Positional Embedding Encoder Decoder Fig. 1. Overview of the proposed MAESIL framework, which masks a high ratio of an input 3D superpatch, feeds only the visible portions to an encoder, and uses a decoder to reconstruct the original superpatch from the encoded information and learnable mask tokens. The goal of this paper is to overcome this critical limita- tion of 2D-based SSL approaches. We propose a novel self- supervised learning framework that introduces a new 3D input unit capable of capturing 3D structural information efficiently. This approach avoids the information loss seen in 2D slice methods while mitigating the prohibitive computational bur- den associated with processing entire 3D scans, enabling a practical and effective pre-training solution for 3D CT scans. I. RELATED WORK Masked Autoencoders (MAE) have recently emerged as one of the most scalable and effective self-supervised learn- ing (SSL) frameworks for visual representation learning [7]. Inspired by masked language modeling, MAE learns rich representations by reconstructing randomly masked patches from a small subset of visible patches [8]. This powerful paradigm has naturally been extended to the medical imaging domain to address its unique challenges. A prominent example is MedMAE, which directly tackles the critical domain shift problem [9]. The authors of MedMAE correctly identified that models pre-trained on natural images (e.g., ImageNet) perform poorly on medical tasks.To solve this, they demonstrated that pre-training on a large-scale, diverse, unlabeled dataset (MID) allows the model to learn domain-specific visual representations. Their results convinc- ingly show that this domain-specific backbone significantly outperforms ImageNet-trained models across various down- stream tasks, including classification and segmentation. However, while MedMAE successfully addresses the do- main shift problem, its methodology for handling 3D data like CT reveals a critical limitation. Although their extensive MID dataset includes 3D modalities such as CT and MRI, the framework itself processes these complex volumes as a collection of independent 2D slices. This 2D-based approach, while computationally convenient, fundamentally discards the 3D structural context and axial coherence between slices. This inter-slice information is essential for a comprehensive understanding of anatomical structures and pathologies in CT analysis. As the work on MedMAE illustrates, the value of domain- specific pre-training is well-established. Yet, a methodology that effectively and efficiently captures the intrinsic 3D con- textual information of CT scans, without resorting to 2D-slice simplification, remains an open challenge in the field. I. METHODS The overall architecture of our proposed MAESIL frame- work is illustrated in Fig. 1. The central innovation of MAESIL is its input processing pipeline, which introduces a 3D chunk- based unit termed the ‘superpatch’. This design was conceived to balance two competing objectives: (1) mitigating the ‘3D context loss’ from 2D slice-based methods and (2) managing the ‘high computational cost’ of processing full 3D volumes. A. Superpatching and Masking Standardized Patches ... Superpatches CT Scan ... Fig. 2.The 3D input processing pipeline. A full 3D CT scans is first partitioned into ‘superpatches’. Each superpatch is then densely tokenized into smaller ‘standardized patches’, which serve as the input tokens for the Transformer. As illustrated in Fig. 2, an input CT scans, assumed to be 512×512×512, is first processed by our model. To balance 3D context against computational load, we partition this volume into 64 non-overlapping superpatches, each with dimensions of 128×128×128. This superpatch unit is then further tokenized. It is densely divided into 8×8×8 local patches, yielding 4,096 tokens per superpatch. A 3D convolution layer embeds these local patches into a latent vector space (e.g., 768-dimensional) suitable for the Transformer. To preserve spatial relationships, positional encodings are added to each token. For the pre-training objective, we adapt the MAE strategy for 3D data. A dual-masking strategy is applied to the super- patch tokens. First, ‘plane-wise’ masking randomly discards 75% of the patches within each 2D plane. Second, ‘axis-wise’ masking removes a contiguous block of 50% of the slices along the Superior-Inferior (S-I) axis. B. Encoder-Decoder Architecture The encoder-decoder architecture is responsible for recon- structing the original superpatch from the masked input. The decoder first receives the compressed embeddings from the encoder. Concurrently, learnable mask tokens are inserted into the original positions of the masked patches to complete the full sequence. These mask tokens serve as signals, inform- ing the decoder which positions need to be reconstructed. This full sequence is then reordered to its original order and passes through a stack of Transformer blocks within the decoder. The Transformer blocks learn the contextual relationships within the sequence via self-attention. Through this process, predictions for each patch are made. Finally, all predicted patches are reassembled into a single superpatch, restoring the original volume. IV. EXPERIMENTS AND RESULTS A. Dataset Our model’s performance and generalization capabilities are validated across three distinct, large-scale, publicly available medical CT datasets: BTCV, LIDC-IDRI, and TotalSegmen- tator. These datasets were chosen to cover a wide spectrum of anatomical regions, from abdominal organs and thoracic scans to a full-body. A summary of these datasets is provided in Table I. TABLE I OVERVIEW OF THE DATASETS USED FOR PRE-TRAINING. Dataset# ScansRegion BTCV [10]30Abdominal LIDC-IDRI [11]1,018Thoracic TotalSegmentator [12]1,204Full-body BTCV [10] The “Multi-Atlas Labeling Beyond the Cranial Vault” (BTCV) dataset is a standard benchmark for abdominal organ segmentation. It contains 30 subjects with contrast- enhanced abdominal CT scans and provides high-quality man- ual segmentations for 13 abdominal structures. LIDC-IDRI [11] The “Lung Image Database Consortium and Image Database Resource Initiative” (LIDC-IDRI) is a large-scale public database of thoracic CT scans. It consists of 1,018 cases, with detailed XML annotations for lung nodules (lesions) independently delineated by four experienced radiologists. TotalSegmentatorV2 [12] This is a recent, large-scale benchmark for robust, segmentation, containing 1,204 CT examinations [13]. Its key feature is the comprehensive an- notation of 104 different anatomical structures (including 27 organs, 59bones, 10 muscles, and 8 vessels). The dataset is intentionally diverse, representing a wide variety of patient ages, scanners, and pathologies. While this diversity poses a significant challenge, it is precisely what makes the dataset ideal for evaluating model robustness to varied scan protocols and patient conditions. For our pre-training, these three distinct collections were combined into a single, unified dataset to force the model to learn generalized representations across all domains. B. Comparison Results To evaluate the reconstruction performance of our proposed MAESIL framework, we conduct a comprehensive quantitative comparison against several standard generative baseline mod- els. These baselines include a standard Autoencoder (AE), a Variational Autoencoder (VAE), and a Vector Quantized Varia- tional Autoencoder (VQ-VAE) [14]–[16]. While our MAESIL is trained on a masked reconstruction (inpainting) task, the baselines are trained on the standard full reconstruction task, i.e., reconstructing the original uncorrupted input. To rigor- ously test generalization, all models were pre-trained under fair conditions on a single, unified dataset that combines all three previously mentioned collections (BTCV, LIDC- IDRI, and TotalSegmentator). This forced the models to learn robust representations from a highly diverse mix of anatomical structures, body parts, and scan protocols. The performance is measured using three widely accepted metrics for image reconstruction quality: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS). For PSNR and SSIM, higher values indicate better reconstruction fidelity, while for the perceptually-based LPIPS metric, lower values signify that the reconstructed image is closer to the original in human perception. The consolidated results of this comparison are presented in Table I. TABLE I COMPARISON WITH OTHER MODELS. ModelPSNR (↑)SSIM (↑)LPIPS (↓) AE [14]25.230.970.39 VAE [15]11.910.250.91 VQ-VAE [16]20.040.880.49 MAESIL (Ours)30.280.980.26 The results clearly demonstrate the superiority of our MAE- SIL framework in all evaluation metrics. MAESIL achieves the highest PSNR score of 30.28, significantly surpassing all baselines and indicating a much higher pixel-level accuracy in our reconstructions. Similarly, in terms of structural fi- delity, our model obtains a near-perfect SSIM score of 0.99. The VAE, in particular, shows a notable failure in capturing structural information, scoring only 0.25 on this metric. Most critically, MAESIL achieves the lowest LPIPS score of 0.26, compared to 0.39 (AE), 0.49 (VQ-VAE), and 0.91 (VAE). This low perceptual distance score confirms that our model’s re- constructions are not only pixel-accurate but also perceptually more realistic and less blurrier than those from the baseline methods. This strong performance validates the effectiveness of our 3D superpatch and a dual-masking strategy in learning comprehensive 3D representations. C. Discussion In this section, we analyze the quantitative and qualitative results of our framework. TABLE I PSNR PERFORMANCE OF OUR MODEL (MAESIL) ACROSS DIFFERENT DATASETS. DatasetPSNR (↑) BTCV [10]32.09 LIDC-IDRI [11]22.58 TotalSegmentatorV2 [12]18.72 Table I presents the quantitative results, showing the PSNR performance of our model when evaluated on each of the three datasets individually. The performance varies, achieving the highest PSNR of 32.09 on BTCV and the lowest of 18.72 on TotalSegmentator. This variation is expected; the TotalSegmentatorV2 dataset is by far the most diverse, mixing 104 anatomical structures, various body parts, and different scan protocols, which makes achieving a high average PSNR highly challenging. Critically, these results were obtained using a total masking ratio of 75% and an embedding di- mension of 768. This high ratio was implemented via our dual-masking scheme, which includes a difficult 50% axial- wise masking. This aggressive masking strategy makes the inpainting task extremely challenging and prevents the model from trivially copying visible patches. The fact that MAESIL still achieves robust PSNR scores under this difficult condition strongly suggests that the model is forced to, and successfully does, learn meaningful 3D contextual representations from the limited visible data. Reconstruction Original Fig. 3. Qualitative reconstruction results of MAESIL across different anatom- ical views. The original input and reconstruction output are compared. From left to right: Axial, Coronal and Sagittal views. Figure 3 provides a qualitative comparison of our re- construction results across three different anatomical views (Axial, Coronal, and Sagittal). As shown, MAESIL success- fully reconstructs complex anatomical structures—such as the vertebrae and various soft tissues—with high fidelity to the original input across all three axes. This visually confirms that our model effectively learns and preserves 3D contextual information rather than treating slices independently. However, fixed-pattern artifacts are observable in the reconstruction outputs. This is a common phenomenon in patch-based recon- struction methods related to the decoder’s upsampling process, and addressing these artifacts is a clear direction for future work [7]. V. CONCLUSION In this paper, we proposed MAESIL, a novel self-supervised learning framework based on a superpatch input unit and a dual-masking strategy to address the 3D structural context loss in existing 2D-based methods. Experiments on a unified dataset of three public benchmarks showed that MAESIL significantly outperforms standard reconstruction baselines like AE and VQ-VAE across all key metrics. This validates MAESIL as an effective pre-training strategy for 3D medical data. Future work will focus on refining the decoder architec- ture to reduce reconstruction artifacts and applying the pre- trained backbone to downstream tasks such as segmentation and classification. VI. ACKNOWLEDGMENT This work was supported by the Next Generation Semicon- ductor Convergence and Open Sharing System, and by the Institute of Information & Communications Technology Plan- ning & Evaluation (IITP) through the Artificial Intelligence Semiconductor Support Program to Nurture the Best Talents (IITP-2023-RS-2023-00256081), funded by the Ministry of Science and ICT of Korea (MSIT). REFERENCES [1] G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. A. Van Der Laak, B. Van Ginneken, and C. I. S ́ anchez, “A survey on deep learning in medical image analysis,” Medical Image Analysis, vol. 42, p. 60–88, 2017. [2] S.-C. Huang, A. Pareek, M. Jensen, M. P. Lungren, S. Yeung, and A. S. Chaudhari, “Self-supervised learning for medical image classification: a systematic review and implementation guidelines,” npj Digital Medicine, vol. 6, no. 1, p. 74, 2023. [3] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “Imagenet large scale visual recognition challenge,” International journal of computer vision, vol. 115, no. 3, p. 211–252, 2015. [4] H. Guan and M.-H. Liu, “Domain adaptation for medical image analysis: A survey,” IEEE Transactions on Biomedical Engineering, vol. 69, no. 7, p. 2245–2259, 2022. [5] E. L. Gao, W. K. Mark, A. I. McKinney, J. Xiao, D. A. Goldman, J. M. Holcomb, and R. J. Young, “Comparing 3d, 2.5 d, and 2d approaches to brain image auto-segmentation,” Tomography, vol. 9, no. 2, p. 478–489, 2023. [6] S. Singh, P. Kencha, and G. Prakash, “Leveraging 2d deep learning imagenet-trained models for native 3d medical image analysis,” Scientific Reports, vol. 13, no. 1, p. 18803, 2023. [7] K. He, X. Chen, S. Xie, Y. Li, P. Doll ́ ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” arXiv preprint arXiv:2111.06377, 2021. [8] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, 2018. [9] A. Gupta, I. Osman, M. S. Shehata, and J. W. Braun, “Medmae: A self-supervised backbone for medical imaging tasks,” arXiv preprint arXiv:2407.14784, 2024. [10] B. A. Landman, Z. Xu et al., “Btcv: Multi-atlas labeling beyond the cranial vault (synapse),” Synapse Repository, 2015. [11] S. G. Armato I et al., “Lidc/idri collection (tcia),” The Cancer Imaging Archive, 2015. [12] J. Wasserthal, H.-C. Breit, M. T. Meyer, M. Pradella, D. Hinck, A. W. Sauter, T. Heye, D. T. Boll, J. Cyriac, S. Yang et al., “Totalsegmentator: robust segmentation of 104 anatomic structures in ct images,” Radiology: Artificial Intelligence, vol. 5, no. 5, p. e230024, 2023. [13] P. Guo, C. Zhao, D. Yang, Z. Xu, V. Nath, Y. Tang, B. Simon, M. Belue, S. Harmon, B. Turkbey et al., “Maisi: Medical ai for synthetic imaging,” in 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2025, p. 4430–4441. [14] G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” science, vol. 313, no. 5786, p. 504–507, 2006. [15] D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, 2013. [16] A. Van Den Oord, O. Vinyals et al., “Neural discrete representation learning,” Advances in neural information processing systems, vol. 30, 2017.