Paper deep dive
MATHENA: Mamba-based Architectural Tooth Hierarchical Estimator and Holistic Evaluation Network for Anatomy
Kyeonghun Kim, Jaehyung Park, Youngung Han, Anna Jung, Seongbin Park, Sumin Lee, Jiwon Yang, Jiyoon Han, Subeen Lee, Junsu Lim, Hyunsu Go, Eunseob Choi, Hyeonseok Jung, Soo Yong Kim, Woo Kyoung Jeong, Won Jae Lee, Pa Hong, Hyuk-Jae Lee, Ken Ying-Kai Liao, Nam-Joon Kim
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/2/2026, 3:28:50 AM
Summary
MATHENA is a unified, Mamba-based framework for dental diagnosis from Orthopantomograms (OPGs), integrating tooth detection (MATHE) and multi-task analysis (HENA) for caries segmentation, anomaly detection, and dental developmental staging. It utilizes a novel Global Context State Token (GCST) and a coarse-to-fine architecture to achieve high performance with linear O(N) complexity, supported by the newly curated PARTHENON benchmark.
Entities (6)
Relation Signals (4)
MATHENA → comprises → MATHE
confidence 100% · MATHENA integrates MATHE, a multi-resolution SSM-driven detector
MATHENA → comprises → HENA
confidence 100% · These crops are processed by HENA, a lightweight Mamba-UNet
MATHENA → evaluatedon → PARTHENON
confidence 100% · We also curate PARTHENON, a benchmark... MATHENA achieves 93.78% mAP@50 in tooth detection
MATHE → uses → Mamba
confidence 100% · MATHENA incorporates Mamba... into MATHE for tooth detection
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Dental diagnosis from Orthopantomograms (OPGs) requires coordination of tooth detection, caries segmentation (CarSeg), anomaly detection (AD), and dental developmental staging (DDS). We propose Mamba-based Architectural Tooth Hierarchical Estimator and Holistic Evaluation Network for Anatomy (MATHENA), a unified framework leveraging Mamba's linear-complexity State Space Models (SSM) to address all four tasks. MATHENA integrates MATHE, a multi-resolution SSM-driven detector with four-directional Vision State Space (VSS) blocks for O(N) global context modeling, generating per-tooth crops. These crops are processed by HENA, a lightweight Mamba-UNet with a triple-head architecture and Global Context State Token (GCST). In the triple-head architecture, CarSeg is first trained as an upstream task to establish shared representations, which are then frozen and reused for downstream AD fine-tuning and DDS classification via linear probing, enabling stable, efficient learning. We also curate PARTHENON, a benchmark comprising 15,062 annotated instances from ten datasets. MATHENA achieves 93.78% mAP@50 in tooth detection, 90.11% Dice for CarSeg, 88.35% for AD, and 72.40% ACC for DDS.
Tags
Links
- Source: https://arxiv.org/abs/2604.00537v1
- Canonical: https://arxiv.org/abs/2604.00537v1
Trouble viewing inline? Open PDF directly →
Full Text
25,248 characters extracted from source content.
Expand or collapse full text
11institutetext: OUTTA, Seoul, South Korea 11email: kyeonghun.kim@outta.ai 22institutetext: Seoul National University, Seoul, South Korea 22email: knj01@snu.ac.kr 33institutetext: Sangmyung University, Seoul, South Korea 44institutetext: Gwangju Institute of Science and Technology, Gwangju, South Korea 55institutetext: Samsung Medical Center, Sungkyunkwan University, Seoul, South Korea 66institutetext: Samsung Changwon Hospital, Sungkyunkwan University, Changwon, South Korea 77institutetext: NVIDIA AI Technology Center, Taipei, Taiwan MATHENA: Mamba-based Architectural Tooth Hierarchical Estimator and Holistic Evaluation Network for Anatomy Kyeonghun Kim Jaehyung Park Youngung Han Anna Jung Seongbin Park Sumin Lee Jiwon Yang Jiyoon Han Subeen Lee Junsu Lim Hyunsu Go Eunseob Choi Hyeonseok Jung Soo Yong Kim Woo Kyoung Jeong Won Jae Lee Pa Hong Hyuk-Jae Lee Ken Ying-Kai Liao Nam-Joon Kim(🖂) Abstract Dental diagnosis from Orthopantomograms (OPGs) requires coordination of tooth detection, caries segmentation (CarSeg), anomaly detection (AD), and dental developmental staging (DDS). We propose Mamba-based Architectural Tooth Hierarchical Estimator and Holistic Evaluation Network for Anatomy (MATHENA), a unified framework leveraging Mamba’s linear-complexity State Space Models (SSM) to address all four tasks. MATHENA integrates MATHE, a multi-resolution SSM-driven detector with four-directional Vision State Space (VSS) blocks for O(N)O(N) global context modeling, generating per-tooth crops. These crops are processed by HENA, a lightweight Mamba-UNet with a triple-head architecture and Global Context State Token (GCST). In the triple-head architecture, CarSeg is first trained as an upstream task to establish shared representations, which are then frozen and reused for downstream AD fine-tuning and DDS classification via linear probing, enabling stable, efficient learning. We also curate PARTHENON, a benchmark comprising 15,062 annotated instances from ten datasets. MATHENA achieves 93.78% mAP50mAP_50 in tooth detection, 90.11% Dice for CarSeg, 88.35% for AD, and 72.40% ACC for DDS. 1 Introduction Orthopantomogram (OPG) is the widely used dental radiographs that provide a comprehensive overview of dental arches, maxillary and mandibular bones, and surrounding anatomical structures [7]. Clinical diagnosis requires integrated evaluation of tooth detection, caries segmentation (CarSeg), anomaly detection (AD), and dental developmental staging (DDS) [27, 2]. While clinical dentistry distinguishes congenital anomalies from acquired pathoses [27], computer vision defines anomaly detection more broadly as identifying morphological deviation from normal distributions [16]. Adopting this computational perspective, we formulate AD as a generalized pixel-wise segmentation task. By unifying dental lesions into a single anomalous class, our approach integrates AD with other diagnostic tasks and emphasizes pathological localization. Current deep learning models often address these tasks in isolation: CNNs lack global context [19], while Transformers incur quadratic computational overhead [10, 32]. To overcome these limitations, we propose MATHENA, a unified framework reflecting the clinical coarse-to-fine workflow. MATHENA incorporates Mamba [4, 28, 34]–a selective State Space Model with linear O(N)O(N) complexity and global receptive fields–into MATHE for tooth detection and HENA for per-tooth multi-task analysis. To address fragmentation across existing datasets, we curate PARTHENON, a large-scale benchmark unifying ten datasets into 15,062 annotated instances. As shown in Fig. 1, MATHENA consistently outperforms existing baselines in tooth detection, CarSeg, and AD across the individual datasets comprising PARTHENON. Our main contributions can be summarized as follows: • We propose MATHENA, a unified framework for tooth detection, CarSeg, AD, and DDS. • We integrate directional Vision State Space (VSS) blocks to achieve O(N)O(N) global context modeling without Transformer overhead. • We enable per-tooth multi-task prediction via a novel Global Context State Token (GCST) mechanism and triple-head design. Figure 1: Quantitative results on PARTHENON: (a) MATHE variants outperform baseline models in tooth detection (mAP50); (b, c) MATHENA variants show superior Dice scores in CarSeg and AD. († : baseline enhanced with P2, BiFPN, and WIoU). 2 Dataset 2.1 PARTHENON Dataset As shown in Table 1, PARTHENON aggregates ten dental datasets (8 panoramic, 2 periapical; 15,062 instances) with annotations spanning 14 original diagnostic categories, which are merged into task-specific binary labels (Sec. 2.2). Annotations consist of tooth-level bounding boxes, either manually annotated or generated from existing segmentation masks. For subsets with developmental metadata, dental maturity is categorized using the Demirjian method (A-H) [2]. Each dataset supports one or more tasks summarized in Table 1. Table 1: PARTHENON composition: ①-④ denote tooth detection, Caries Segmentation (CarSeg), Anomaly Detection (AD), and Dental Developmental Staging (DDS). No. Dataset Image Annotation CarSeg Mask AD Mask Task D1 DC1000 [25] 597 597 591 591 ② ③ D2 PRAD [33] 5,000 5,000 — 669 ① ③ D3 PRDA [3] 532 532 — 516 ③ D4 Dentex [5] 2,326 2,326 694 724 ② ③ D5 ADCD [14] 1,808 1,808 1,068 1,559 ① ② ③ D6 DDSNet [24] 380 380 — — ① ④ D7 DVCTNet [11] 2,000 2,000 1,771 1,771 ② ③ D8 TSD [26] 994 994 611 611 ① ② ③ D9 Tufts [15] 1,000 1,000 — — ① D10 UFBA-425 [1] 425 425 — — ① Total 15,062 15,062 4,735 6,411 2.2 Data Preprocessing 2.2.1 Semi-Supervised Pseudo-Label Generation. Ground-truth bounding boxes for tooth detection are available for PARTHENON subsets D2, D5, D6, D8, D9, D10. RT-DETR-L, trained on these subsets, achieves 93.7 mAP50mAP_50 and serves as the teacher. We apply it to the other datasets in a semi-supervised framework [9] to generate pseudo-ground-truth bounding boxes [29, 32]. 2.2.2 Pseudo-Label Quality Filtering. Teacher-generated bounding boxes are filtered using a confidence threshold and NMS. Anatomically implausible predictions are rejected by Mahalanobis distance-based outlier detection [13]. Each box is mapped to normalized spatial features v=[cx/W,cy/H,log(w/W),log(h/H)]⊤v=[c_x/W,\;c_y/H,\; (w/W),\; (h/H)] , and predictions whose squared distance exceeds χ42χ^2_4 threshold at p<0.001p<0.001 are discarded [21, 20]. The resulting bounding boxes are used to train MATHE. 2.2.3 Label Merging. For CarSeg, the multi-stage annotations in D1, D4, D5, D7, and D8 are collapsed into a binary mask (caries vs. background) by mapping all non-zero pixel classes to 1. For AD, per-tooth labels in D1, D2, D3, D4, D5, and D7 are unified into a binary normal/anomalous label. 2.2.4 Cropping and Augmentation. Cropped image-mask pairs are extracted from the OPG and corresponding mask for HENA training. Offline augmentations–random rotations and horizontal flip–are applied to each instance. 3 Methodology: The MATHENA Framework We propose MATHENA, a unified framework motivated by the coarse-to-fine clinical review process, as illustrated in Fig. 2. Our framework is twofold: MATHE for tooth detection (Sec. 3.1) and HENA for multi-task analysis, including CarSeg, AD, and DDS (Sec. 3.2). Figure 2: MATHENA architecture: Left shows MATHE backbone, BiFPN, and detection head; Right depicts HENA encoder-decoder with GCST skip fusion. SRA (Spatial Re-Alignment) remaps the output mask to the original OPG for visualization. 3.1 MATHE: Mamba-based Architectural Tooth Hierarchical Estimator Given an OPG image, MATHE extracts multi-scale features, fuses them across resolutions, and outputs per-tooth bounding boxes. 3.1.1 Hybrid CNN-SSM Backbone. Early stages (P2P_2, P3P_3) use standard convolutions to capture local features. Deeper stages (P4P_4, P5P_5) replace bottlenecks with C2fSSM blocks containing VSS units that perform four-directional selective scanning. This allocates O(N)O(N) global context modeling at semantically rich stages while retaining efficient convolutions at high resolution. 3.1.2 Bidirectional Feature Pyramid and Head. Multi-scale features from four backbone stages are fused through a Bidirectional Feature Pyramid Network (BiFPN) [22] with learnable per-node fusion weights. Unlike standard FPN, BiFPN adds a bottom-up pathway and weighted feature combination at each fusion node: Piout=∑jwj⋅Pjin∑jwj+ϵ,P_i^out= _jw_j· P_j^in _jw_j+ε, where wjw_j are learnable scalar weights and ϵε ensures numerical stability. We include P2P_2 (stride 4) to improve detection of small periapical structures lost at coarser scales [31]. The head uses decoupled convolutional towers for box regression at each pyramid level and is optimized with Wise-IoU (WIoU) [23], which dynamically adjusts gradients based on box quality. 3.2 HENA: Holistic Evaluation Network for Anatomy Each detected tooth region with its paired mask is cropped, resized to 224×224224× 224, and passed through HENA’s pipeline before being routed to task-specific heads. 3.2.1 Lightweight Encoder. We employ a lightweight U-shaped encoder-decoder inspired by MobileUNETR [17], where we replace all Transformer blocks with Mamba VSS blocks [28] to achieve O(N)O(N) complexity for intra-tooth dependency modeling. The encoder consists of depthwise-separable convolution (DWSep) blocks [6] to progressively downsample spatial resolution while expanding channel depth. 3.2.2 Mamba Bottleneck with GCST. At the 28×2828×28 bottleneck, the feature map Fbot∈ℝB×256×28×28F_bot ^B×256×28×28 is flattened to a sequence X∈ℝB×784×256X ^B×784×256. We prepend a learnable global context token Tg∈ℝ1×256T_g ^1×256 (initialized to zeros) to form X′=[Tg;X]∈ℝB×785×256X =[T_g;X] ^B×785×256. A VSS block processes X′X with four-directional selective scanning, producing H=VSS(X′)∈ℝB×785×256H=VSS(X ) ^B×785×256. We extract the global token state hg=H0∈ℝCh_g=H_0 ^C and the spatial hidden states Z=H1:L∈ℝL×CZ=H_1:L ^L×C, then broadcast-add the global context to all spatial positions: Yout=Z+Lhg⊤,Y_out=Z+1_Lh_g , where hgh_g acts as a global context aggregator. It accumulates holistic tooth-level semantics through Mamba’s linear recurrence, providing dense spatial modulation at an efficient O(N)O(N) cost. The modulated features are then reshaped to ℝB×256×28×28R^B×256×28×28 for the decoder. 3.2.3 Decoder with GCST Skip Fusion. The decoder progressively restores spatial resolution through transposed convolutions with skip connections. At each decoder level s∈1,2,3s∈\1,2,3\, the skip features Ss∈ℝB×Cs×Hs×WsS_s ^B×C_s×H_s×W_s are modulated by GCST before fusion. The mechanism flattens SsS_s to Xs∈ℝB×Ls×CsX_s ^B×L_s×C_s (Ls=Hs×WsL_s=H_s×W_s), prepends a learnable scale token TsT_s, and applies a lightweight Mamba block to [Ts;Xs][T_s;X_s]. The output token state hsh_s is projected to Feature-wise Linear Modulation (FiLM) [18] parameters: (γs,βs)=ψ(hs),S^s=γs⊙Ss+βs,( _s, _s)=ψ(h_s), S_s= _s S_s+ _s, where γs,βs∈ℝCs _s, _s ^C_s are broadcast spatially. This replaces the Transformer-based cross-attention in MobileUNETR with an O(Ls)O(L_s) alternative that captures cross-scale dependencies through Mamba’s selective recurrence. Table 2: Tooth detection on PARTHENON: ⋆baseline YOLOv8; †enhanced with P2, BiFPN, WIoU. Method Backbone mAP50(%) mAP75(%) mAP50:95(%) RetinaNet ResNet-50 71.01 61.81 51.83 YOLOv8n⋆ CSPDarkNet(C2f) 71.30 66.30 61.90 FCOS ResNet-50 71.86 62.20 53.20 SSD300 VGG-16 72.77 65.60 58.80 YOLOv8s⋆ CSPDarkNet(C2f) 73.66 70.05 65.02 Faster R-CNN MobileNetV3 80.08 66.56 49.30 Faster R-CNN ResNet-50 80.32 68.97 50.72 YOLOv8m⋆ CSPDarkNet(C2f) 80.32 75.71 67.80 YOLOv8n† CSPDarkNet(C2f) 90.11 85.30 77.90 YOLOv8s† CSPDarkNet(C2f) 92.24 89.89 79.32 YOLOv8m† CSPDarkNet(C2f) 92.78 88.77 79.91 MATHE Mamba-SSM 93.78 91.89 81.32 MATHE + TTA Mamba-SSM 94.89 92.94 83.45 3.2.4 Triple-Head Multi-Task Learning and Training Strategy. HENA employs a shared encoder-decoder backbone with three task-specific heads, optimized via sequential transfer learning. First, the entire network is trained on the upstream CarSeg task to establish robust dental representations. Next, treating AD and DDS as downstream tasks, we freeze the shared backbone to efficiently transfer these learned features. The AD head is attached to the decoder, while the DDS head is applied to the encoder’s bottleneck (GAP(Fbot)GAP(F_bot)) and fine-tuned via linear probing. At inference, the predicted stage is s^=∑j=18[σ(y^j)>0.5] s= _j=1^81[σ( y_j)>0.5], mapped to A-H. Freezing the common backbone for downstream tasks minimizes computational overhead. Compared to a fully learnable setup (90.03% CarSeg Dice), our frozen sequential approach maintained 90.11% Dice while reducing training and inference times by 3.5× and 1.4×, respectively. 3.3 Loss Functions The MATHE detector and HENA analyzer are trained with: ℒMATHE=λwiouℒWIoU(B,B^)+λl1ℒL1(B,B^)+λdflℒDFL(B,B^)L_MATHE= _wiouL_WIoU(B, B)+ _l1L_L1(B, B)+ _dflL_DFL(B, B) (1) ℒHENA=ℒDice(S,S^)+ℒDice(A,A^)+ℒOrd(Ystg,Y^stg)L_HENA=L_Dice(S, S)+L_Dice(A, A)+L_Ord(Y_stg, Y_stg) (2) where ℒWIoUL_WIoU [23] is the Wise-IoU loss, ℒL1L_L1 is the bounding box regression loss, ℒDFLL_DFL is the Distribution Focal Loss addressing the bounding box coordinate distribution, ℒDiceL_Dice [12] supervises CarSeg and AD, and ℒOrdL_Ord [8] is the cumulative ordinal loss with Kstg=8K_stg=8 stages and Kstg−1K_stg-1 binary thresholds. 4 Experiments 4.1 Implementation Details MATHENA was implemented in PyTorch on an NVIDIA A100 GPU. MATHE builds on YOLOv8m [30], replacing C2f blocks at P4P_4 and P5P_5 with C2fSSM, FPN with BiFPN, and IoU with WIoU. Training was performed for 100 epochs at 1024×5121024×512 using AdamW (lr=1×10−4lr=1×10^-4, weight decay 5×10−25×10^-2) with linear warmup and cosine annealing. HENA was trained for 150 epochs on 224×224224×224 crops with batch size 16. TTA applies horizontal flip with NMS (threshold 0.5), improving MATHE from 93.78% to 94.89% mAP50. 4.2 Comparative Analysis We compared MATHE and MATHENA with standard object detectors and leading segmentation models on the PARTHENON test set. Quantitative performance is summarized in Table 2 and 3, with a visual comparison in Fig. 3. Table 3: Multi-task performance on PARTHENON: CarSeg and AD averaged across subsets; DDS on DDSNet. CarSeg AD DDS Method Dice (%) IoU (%) Dice (%) IoU (%) ACC (%) F1 (%) FPN 76.41 63.93 74.74 61.29 – – MAnet 77.51 64.48 75.29 62.03 – – PSPNet 79.21 66.12 76.55 63.37 – – UNet 80.81 67.24 78.16 64.65 – – Linknet 82.21 68.75 79.87 66.14 – – UNet++ 83.11 69.57 81.30 67.83 – – SEResUNet 83.28 69.89 81.65 68.06 – – UNet3+ 84.12 70.64 82.38 68.92 – – TransUNet 84.64 71.13 83.03 69.59 – – MobileUNETR 84.82 71.47 83.40 69.90 – – nnU-Net 85.23 72.12 83.91 69.87 – – DeepLabv3+ 85.91 72.83 84.48 70.74 – – MATHENA 90.11 76.94 88.35 74.77 72.40 70.10 MATHENA + TTA 91.31 78.05 89.59 75.92 74.10 72.30 Figure 3: Visual comparison: Left shows Tooth Detection; Right depicts CarSeg. Tooth Detection. MATHE achieves 93.78% mAP50, improving over baseline YOLOv8 configurations. TTA, merging multi-view predictions via NMS, increases mAP50 to 94.89%. Multi-task Performance. MATHENA reaches 90.11% Dice for CarSeg and 88.35% for AD, outperforming baselines including DeepLabv3+ (85.91%, 84.48%). As shown in Fig. 3, MATHENA provides precise tooth detection and segmentation across tasks. DDS. The DDS head achieves 72.40% ACC and 70.10% F1 on DDSNet. 4.3 Ablation Study Table 4 validates the architectural design of MATHENA across all four tasks. Removing GCST from MATHE drops detection performance by 4.36% mAP50, indicating that global context is critical for resolving crowded dentition. Replacing BiFPN with a standard FPN reduces detection by 3.14% mAP50, confirming the necessity of bidirectional multi-scale fusion for accurate tooth detection. For per-tooth analysis, removing the GCST skip fusion degrades CarSeg by 3.15% Dice, highlighting its essential role in cross-scale spatial modulation. Replacing the Mamba bottleneck with standard convolutions degrades CarSeg by 2.36% Dice and DDS by 2.45% ACC, proving the efficacy of linear recurrence for holistic feature extraction. Finally, replacing Mamba blocks with Vision Transformers universally underperforms, demonstrating Mamba’s superior architectural stability and efficiency in capturing intra-tooth dependencies without quadratic computational overhead. Table 4: Ablation study on PARTHENON: impact of each component across all tasks. Configuration Detection (%) CarSeg (%) AD (%) DDS (%) Full MATHENA 93.78 90.11 88.35 72.40 w/o GCST in MATHE 89.42 90.05 88.29 72.36 w/o Mamba in HENA BN 93.74 87.75 85.99 69.95 w/o GCST skip fusion 93.75 86.96 85.20 72.32 w/ standard FPN(no BiFPN) 90.64 90.03 88.27 72.25 w/ Transformer(vs. Mamba) 92.13 88.66 86.90 70.94 5 Conclusion We present MATHENA, a unified framework integrating MATHE and HENA for tooth detection and multi-task analysis. We also introduce PARTHENON, a benchmark unifying ten datasets under a common schema. The clinically motivated coarse-to-fine pipeline addresses tooth detection, CarSeg, AD, and DDS. Experiments demonstrate that MATHENA provides a structured baseline for multi-task panoramic analysis. References [1] D. Budagam, A. Z. Imanbayev, I. R. Akhmetov, A. Sinitca, S. Antonov, and D. Kaplun (2025) OralBBNet: spatially guided dental segmentation of panoramic x-rays with bounding box priors. arXiv preprint arXiv:2406.03747. Cited by: Table 1. [2] A. Demirjian, H. Goldstein, and J. M. Tanner (1973) A new system of dental age assessment. Human biology, p. 211–227. Cited by: §1, §2.1. [3] A. Fatima, I. Shafi, H. Afzal, K. Mahmood, I. d. l. T. Díez, V. Lipari, J. B. Ballester, and I. Ashraf (2023) Deep learning-based multiclass instance segmentation for dental lesion detection. In Healthcare, Vol. 11, p. 347. Cited by: Table 1. [4] A. Gu and T. Dao (2024) Mamba: linear-time sequence modeling with selective state spaces. In First conference on language modeling, Cited by: §1. [5] I. E. Hamamci, S. Er, E. Simsar, A. E. Yuksel, S. Gultekin, S. D. Ozdemir, K. Yang, H. B. Li, S. Pati, B. Stadlinger, et al. (2023) DENTEX: an abnormal tooth detection with dental enumeration and diagnosis benchmark for panoramic x-rays. arXiv preprint arXiv:2305.19112. Cited by: Table 1. [6] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam (2017) Mobilenets: efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861. Cited by: §3.2.1. [7] R. Izzetti, M. Nisi, G. Aringhieri, L. Crocetti, F. Graziani, and C. Nardi (2021) Basic knowledge and new advances in panoramic radiography imaging techniques: a narrative review on what dentists and radiologists should know. Applied Sciences 11 (17), p. 7858. Cited by: §1. [8] H. Li and Z. Lin (2006) Learning to rank with nonsmooth cost functions. NeurIPS. Cited by: §3.3. [9] Y. Liu, C. Ma, Z. He, C. Kuo, K. Chen, P. Zhang, B. Wu, Z. Kira, and P. Vajda (2021) Unbiased teacher for semi-supervised object detection. arXiv preprint arXiv:2102.09480. Cited by: §2.2.1. [10] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo (2021) Swin transformer: hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, p. 10012–10022. Cited by: §1. [11] T. Luo, H. Wu, T. Yang, D. Shen, and Z. Cui (2025-09) Adapting Foundation Model for Dental Caries Detection with Dual-View Co-Training . In proceedings of Medical Image Computing and Computer Assisted Intervention – MICCAI 2025, Vol. LNCS 15975. Cited by: Table 1. [12] F. Milletari, N. Navab, and S. Ahmadi (2016) V-net: fully convolutional neural networks for volumetric medical image segmentation. Cited by: §3.3. [13] M. Mueller and M. Hein (2025) Mahalanobis++: improving ood detection via feature normalization. arXiv preprint arXiv:2505.18032. Cited by: §2.2.2. [14] Cited by: Table 1. [15] K. Panetta, R. Rajendran, A. Ramesh, S. P. Rao, and S. Agaian (2021) Tufts dental database: a multimodal panoramic x-ray dataset for benchmarking diagnostic systems. IEEE journal of biomedical and health informatics 26 (4), p. 1650–1659. Cited by: Table 1. [16] G. Pang, C. Shen, L. Cao, and A. V. D. Hengel (2021) Deep learning for anomaly detection: a review. ACM computing surveys (CSUR) 54 (2), p. 1–38. Cited by: §1. [17] S. Perera, Y. Erzurumlu, D. Gulati, and A. Yilmaz (2024) MobileUNETR: a lightweight end-to-end hybrid vision transformer for efficient medical image segmentation. In European Conference on Computer Vision, p. 281–299. Cited by: §3.2.1. [18] E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville (2018) Film: visual reasoning with a general conditioning layer. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32. Cited by: §3.2.3. [19] S. Ren, K. He, R. Girshick, and J. Sun (2016) Faster r-cnn: towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence 39 (6), p. 1137–1149. Cited by: §1. [20] P. J. Rousseeuw and B. C. Van Zomeren (1990) Unmasking multivariate outliers and leverage points. Journal of the American Statistical association 85 (411), p. 633–639. Cited by: §2.2.2. [21] B. G. Tabachnick, L. S. Fidell, and J. B. Ullman (2007) Using multivariate statistics. Vol. 5, pearson Boston, MA. Cited by: §2.2.2. [22] M. Tan, R. Pang, and Q. V. Le (2020) Efficientdet: scalable and efficient object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 10781–10790. Cited by: §3.1.2. [23] Z. Tong, Y. Chen, Z. Xu, and R. Yu (2023) Wise-iou: bounding box regression loss with dynamic focusing mechanism. arXiv preprint arXiv:2301.10051. Cited by: §3.1.2, §3.3. [24] P. Wang, A. He, A. Wang, Z. Zhou, X. Guan, and T. Li (2025) Towards automated pediatric dental development staging: a dataset and model. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), Cited by: Table 1. [25] X. Wang, S. Gao, K. Jiang, H. Zhang, L. Wang, F. Chen, J. Yu, and F. Yang (2023) Multi-level uncertainty aware learning for semi-supervised dental panoramic caries segmentation. Neurocomputing 540, p. 126208. Cited by: Table 1. [26] X. Wang, L. Wang, Z. Yang, J. Zhou, Y. Zheng, F. Chen, R. Hong, J. Yu, and F. Yang (2024) DSIS-dpr: structured instance segmentation and diffusion prior refinement for dental anatomy learning. IEEE Transactions on Multimedia. Cited by: Table 1. [27] S. C. White and M. J. Pharoah (2014) Oral radiology: principles and interpretation. 7th edition, Elsevier Health Sciences, St. Louis, Missouri. Cited by: §1. [28] Z. Xing, T. Ye, Y. Yang, G. Liu, and L. Zhu (2024-10) SegMamba: Long-range Sequential Modeling Mamba For 3D Medical Image Segmentation . In proceedings of Medical Image Computing and Computer Assisted Intervention – MICCAI 2024, Vol. LNCS 15008. Cited by: §1, §3.2.1. [29] X. Yang, Z. Song, I. King, and Z. Xu (2022) A survey on deep semi-supervised learning. IEEE transactions on knowledge and data engineering 35 (9), p. 8934–8954. Cited by: §2.2.1. [30] M. Yaseen (2025) What is yolov8: an in-depth exploration of the internal features of the next-generation object detector (2024). Accessed: Sep 10. Cited by: §4.1. [31] J. Yu, H. Zheng, L. Xie, L. Zhang, M. Yu, and J. Han (2023) Enhanced yolov7 integrated with small target enhancement for rapid detection of objects on water surfaces. Frontiers in Neurorobotics 17, p. 1315251. Cited by: §3.1.2. [32] Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y. Liu, and J. Chen (2023) DETRs beat yolos on real-time object detection. arXiv preprint arXiv:2304.08069. Cited by: §1, §2.2.1. [33] Z. Zhou, Y. Zhang, R. Xu, X. Zhao, and T. Li (2025) PRAD: periapical radiograph analysis dataset and benchmark model development. External Links: 2504.07760, Link Cited by: Table 1. [34] L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang (2024) Vision mamba: efficient visual representation learning with bidirectional state space model. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. Cited by: §1.