Paper deep dive
Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization
Dinh Tan Nguyen, Quang-Hien Kha, Le-Hoang Nguyen, Minh-Toan Dinh, Xuan-Huy Nguyen, Dac Phu Ho, Cao Truong Tran, Sai Ho Ling, Lan T Ho-Pham, Liem Pham, Nguyen Quoc Khanh Le
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a multi-task DETR framework, where shared representations support both image-level malignancy prediction and lesion localization, and evaluate its performance on OPTIMAM and a biopsy-confirmed SGM1k cohort. Across both datasets, modern backbones consistently outperformed older ResNet-style features, with ConvNeXtV2 and DINOv3 giving the strongest overall results, whereas MambaVision was less competitive. On OPTIMAM, ConvNeXtV2 achieved the best overall performance, reaching 97.96% AUC, 99.89% sensitivity, 25.08% mAP@.5, and 74.38% recall@.25. On SGM1k, DINOv3 gave the strongest overall results, with 90.97% AUC, 86.28% sensitivity, 82.00% specificity, 27.04% mAP@.5, and 77.32% recall@.25. These findings suggest that backbone quality is a critical factor in effective multi-task mammography, with ConvNeXtV2 emerging as a particularly strong and well-matched CNN backbone for mammography in this framework.
Tags
Links
- Source: https://arxiv.org/abs/2608.09801v1
- Canonical: https://arxiv.org/abs/2608.09801v1
Trouble viewing inline? Open PDF directly →
Full Text
18,518 characters extracted from source content.
Expand or collapse full text
Medical Imaging with Deep Learning 2026Short Paper Track Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization Dinh Tan Nguyen 1,2,† dinhtan.nguyen@uts.edu.au Quang-Hien Kha 2,3,4,† d142111015@tmu.edu.tw Le-Hoang Nguyen 3 lehoangnguyen510@gmail.com Minh-Toan Dinh 7 toandinh6501@outlook.com Xuan-Huy Nguyen 2 huylop99@gmail.com Dac Phu Ho 1 c3514490@uon.edu.au Cao Truong Tran 6 truongct@lqdtu.edu.vn Sai Ho Ling 1 steve.ling@uts.edu.au Lan T Ho-Pham 3 lan.hopham@saigonmec.org Liem Pham 2 liem.pham@saigonmec.org Nguyen Quoc Khanh Le 3,4∗ khanhlee@tmu.edu.tw 1 University of Technology Sydney, Australia. 2 Saigon Precision Medicine Research Center, Viet- nam. 3 College of Medicine, Taipei Medical University, Taiwan. 4 AIBioMed Research Group, Taipei Medical University, Taiwan. 6 Le Qui Don Technical University, Vietnam. 7 International Graduate Program in Artificial Intelligence, National Central University, Taiwan. † These authors contributed equally to this work. Abstract Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a multi-task DETR framework, where shared representations support both image-level malignancy prediction and lesion localization, and evaluate its performance on OPTIMAM and a biopsy-confirmed SGM1k cohort. Across both datasets, modern backbones consistently outperformed older ResNet- style features, with ConvNeXtV2 and DINOv3 giving the strongest overall results, whereas MambaVision was less competitive. On OPTIMAM, ConvNeXtV2 achieved the best over- all performance, reaching 97.96% AUC, 99.89% sensitivity, 25.08% mAP@.5, and 74.38% recall@.25. On SGM1k, DINOv3 gave the strongest overall results, with 90.97% AUC, 86.28% sensitivity, 82.00% specificity, 27.04% mAP@.5, and 77.32% recall@.25. These find- ings suggest that backbone quality is a critical factor in effective multi-task mammography, with ConvNeXtV2 emerging as a particularly strong and well-matched CNN backbone for mammography in this framework. 1 Keywords: Mammography, multi-task learning, object detection, classification, DETR, lesion localization 1. Introduction Screening mammography requires both exam-level risk assessment and spatial evidence that can direct reader attention. This remains challenging because suspicious findings may be 1. Code is publicly available at https://github.com/saigonmec/mammo2detr. © 2026 C-BY 4.0, D.T. Nguyen et al. arXiv:2608.09801v1 [cs.CV] 10 Aug 2026 Nguyen Kha Nguyen Dinh Nguyen Ho Tran Ling Ho-Pham Pham Le Figure 1: Overview of the proposed multi-task DETR architecture subtle, small, and partially masked by dense tissue; in a large screening study, mammo- graphic sensitivity dropped substantially in the densest breasts (Kolb et al., 2002). Screen- ing decisions must also balance benefit and harm: a systematic review of breast cancer screening reported a non-trivial cumulative risk of false-positive biopsy findings, especially with more frequent screening (Myers et al., 2015). For AI systems intended as decision sup- port, it is therefore valuable to assess not only classification performance but also whether the model can return plausible candidate regions. Multi-task learning is attractive in this setting because a shared representation can support both image-level malignancy prediction and lesion localization within a single framework (Kha et al., 2024). DETR-style detectors provide an appealing basis for this approach because they predict a set of object instances end to end without hand-crafted anchors (Carion et al., 2020). Here, we study a shared multi-task DETR architecture for joint mammography classification and lesion localization, with a particular focus on backbone suitability. Because the quality of the shared repre- sentation is central to multi-task performance, not all backbone families may be equally effective in this setting. We therefore compare ResNet50 (He et al., 2016), ConvNeXtV2- Tiny (Woo et al., 2023), MambaVision-Tiny (Hatamizadeh and Kautz, 2025), and DINOv3 ViT-B/16 (Sim ́eoni et al., 2025) across OPTIMAM (OMI-DB) (Halling-Brown et al., 2020) and a biopsy-confirmed SGM1k cohort (Kha et al., 2024). 2. Method We use a fixed multi-task DETR framework for joint mammography classification and lesion localization. As illustrated in Figure 1, the proposed architecture combines an interchange- able visual backbone with a shared feature projection layer, an image-level classification branch, and a query-based localization branch based on a Deformable DETR-style decoder (Zhu et al., 2020). This design enables a controlled comparison of backbone effects while keeping the downstream prediction heads unchanged across experiments. We evaluated the framework on two mammography datasets, including OPTIMAM (Halling-Brown et al., 2020) and the biopsy-confirmed SGM1k cohort, comprising 24,643 and 3,525 images after preprocessing, respectively. Data preprocessing, training, and evaluation were standardized 2 Multi-task DETR’s backbone for Mammography Table 1: Classification and localization performance OPTIMAM (OMI-DB)Oncology Hospital (SGM1k) MetricResNet50 ConvNeXtV2 Mamba DINOv3 ResNet50 ConvNeXtV2 Mamba DINOv3 Classification (%) Acc89.3789.9991.7191.9476.3783.7177.8484.25 AUC97.3697.9696.9297.3586.6290.4484.7490.97 Sens99.7899.8990.6295.3590.7081.4073.2686.28 Spec83.0084.0088.0090.0057.0087.0084.0082.00 F189.5490.1591.8392.0275.4583.7977.9684.25 Detection (%) IoU20.3033.1919.4327.6721.4631.6519.6632.65 mAP@.518.4125.0812.9018.0511.2623.7912.2027.04 mAP@.2548.7255.3832.2738.0840.8041.6527.5846.00 R@.2563.5274.3852.1567.15 64.4672.4055.5877.32 across experiments following the pipeline described in Appendices B and C, while additional architectural details are provided in Appendix A. 3. Results and Discussion The results show that the proposed multi-task DETR framework performs best when paired with a suitable modern backbone. Across both datasets, ConvNeXtV2 and DINOv3 achieved the strongest overall performance, whereas ResNet50 was consistently less com- petitive and MambaVision showed weaker overall results. On OPTIMAM, ConvNeXtV2 gave the best overall metrics, suggesting that an improved CNN backbone is particularly well suited to mammography images. On SGM1k, DINOv3 performed best overall and achieved the strongest localization results, while ConvNeXtV2 again preserved the highest specificity. From a clinical perspective, the detection metrics in Table 1 should be interpreted as candidate-region support rather than precise lesion delineation. The Grad-CAM visualiza- tions in Appendix D show that stronger backbones more consistently focused on clinically relevant suspicious regions. Such approximate localization may still be useful for directing reader attention in mammography, where small abnormalities and tissue overlap make ex- act boundaries difficult, particularly in dense breasts (Kolb et al., 2002). Overall, backbone quality is a key determinant of effective multi-task mammography, with modern CNN and ViT representations providing the most suitable shared features for joint classification and localization. 4. Conclusion Our findings suggest that the success of multi-task DETR in mammography depends strongly on representation quality. Modern backbones provided a better balance between exam-level prediction and candidate-region localization, highlighting backbone suitability 3 Nguyen Kha Nguyen Dinh Nguyen Ho Tran Ling Ho-Pham Pham Le as a key design factor in multi-task breast imaging. In particular, ConvNeXtV2 appears especially well matched to the fine-grained visual patterns of mammography. References Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020. Mark D Halling-Brown, Lucy M Warren, Dominic Ward, Emma Lewis, Alistair Mackenzie, Matthew G Wallis, Louise S Wilkinson, Rosalind M Given-Wilson, Rita McAvinchey, and Kenneth C Young. Optimam mammography image database: A large-scale resource of mammography images and clinical data. Radiology: Artificial Intelligence, 3(1):e200103, 2020. doi: 10.1148/ryai.2020200103. Ali Hatamizadeh and Jan Kautz. Mambavision: A hybrid mamba-transformer vision back- bone. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025. Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. Hien Q. Kha, Dinh-Tan Nguyen, Thinh B. Lam, Thanh-Huy Nguyen, Cao T. Tran, Manh D. Vu, Lan T. Ho-Pham, Liem Pham, and Nguyen Quoc Khanh Le. M2net: Two-stage multi-label breast cancer detection networks. In 2024 IEEE International Symposium on Biomedical Imaging (ISBI), pages 1–4, 2024. doi: 10.1109/ISBI56570.2024.10635406. Thomas M Kolb, Jacob Lichy, and Jeffrey H Newhouse. Comparison of the performance of screening mammography, physical examination, and breast us and evaluation of factors that influence them: An analysis of 27,825 patient evaluations. Radiology, 225(1):165–175, 2002. doi: 10.1148/radiol.2251011667. Evan R Myers, Patricia Moorman, Jennifer M Gierisch, Laura J Havrilesky, Lars J Grimm, Sujata Ghate, Brittany Davidson, Ranee Chatterjee Montgomery, Matthew J Crowley, Douglas C McCrory, Amy Kendrick, and Gillian D Sanders. Benefits and harms of breast cancer screening: A systematic review. JAMA, 314(15):1615–1634, 2015. doi: 10.1001/jama.2015.13183. Oriane Sim ́eoni, Huy V Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Micha ̈el Ramamonjisoa, et al. Dinov3. arXiv preprint arXiv:2508.10104, 2025. Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie. Convnext v2: Co-designing and scaling convnets with masked autoen- coders. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16133–16142, 2023. 4 Multi-task DETR’s backbone for Mammography Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai.De- formable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020. 5 Nguyen Kha Nguyen Dinh Nguyen Ho Tran Ling Ho-Pham Pham Le Appendix A. Architecture Details A.1. Overall multi-task design Given an input mammogram I ∈R 3×H×W , the model jointly predicts an image-level ma- lignancy label and a set of query-based lesion proposals: f θ (I) = ˆy,( ˆ b k , ˆs k ) K k=1 , where ˆy ∈R C denotes the image-level logits for C classes, ˆ b k ∈ [0, 1] 4 is the k-th predicted bounding box in normalized coordinates, and ˆs k ∈ [0, 1] is its corresponding objectness score. The architecture consists of an interchangeable visual backbone, a classification branch for image-level prediction, and a localization branch for lesion proposal generation. A.2. Backbone interface To enable a controlled comparison across heterogeneous backbone families, the backbone output is projected into a common feature representation using a 1×1 convolution: F = φ(F backbone )∈R B×256×H ′ ×W ′ . This projection standardizes the channel dimension while keeping the downstream detection and classification heads unchanged across experiments. The compared backbones were ResNet50, ConvNeXtV2-Tiny, MambaVision-Tiny, and DINOv3 ViT-B/16. A.3. Classification branch The classification branch applies global average pooling to the projected feature map F, followed by a learnable linear classifier: ˆy = W · GAP(F) + b where W and b denote the weights and bias of the final linear layer. This branch produces the image-level logits for malignancy classification. A.4. Localization branch The localization branch applies a lightweight multi-scale spatial module to enhance local lesion cues before decoding. Specifically, three parallel 3×3 convolutions with different dilation rates are used to capture multiple receptive fields: (d,p)∈(1, 1), (2, 2), (4, 4). The resulting features are aggregated and passed to a Deformable DETR-style decoder, which uses a fixed set of learned object queries to predict lesion proposals. Each query outputs a bounding box ˆ b k and an objectness score ˆs k . This design is intended to provide candidate regions for review rather than pixel-accurate delineation. Deformable DETR is a suitable choice here because it preserves the end-to-end set-prediction formulation of DETR while improving convergence and small-object handling. 6 Multi-task DETR’s backbone for Mammography A.5. Loss formulation The model is trained with a joint objective: L =L cls + λL det , where L cls denotes the image-level classification loss and L det denotes the detection loss. The classification loss is standard cross-entropy: L cls =− C X c=1 y c log ˆp c . The detection loss combines bipartite matching with box regression and objectness super- vision: L det =L box +L obj . Here, L box includes an L 1 term and a generalized IoU term, while L obj is a binary loss on the objectness score. Appendix B. Datasets We evaluated the model on two mammography datasets: OPTIMAM (OMI-DB) (Halling- Brown et al., 2020) and SGM1k, a biopsy-confirmed cohort from HCM Oncology Hospital. For both datasets, all images were preprocessed to retain only the breast region by cropping away the background and non-breast dark areas before model training and evaluation. Bounding-box annotations were then converted into the format required by the DETR- based framework, and all splits were performed at the patient level to avoid data leakage. Table 2 summarizes the final dataset composition after preprocessing. OPTIMAM yielded 24,643 unique images from 7,851 patients, with 19,780 training images and 4,863 test images. SGM1k yielded 3,525 unique images from 1,002 patients, with 2,776 training images and 749 test images. The OPTIMAM cohort showed a lower proportion of images with multiple boxes but a higher maximum number of boxes per image, whereas SGM1k had fewer total images but a higher proportion of malignant cases. Appendix C. Experiments All backbone variants were compared under a controlled experimental setting designed to isolate the effect of representation choice within the shared multi-task DETR architecture. Experiments were conducted on two mammography datasets, OPTIMAM (OMI-DB) and SGM1k; for both datasets, images were preprocessed by cropping to the breast region and removing background dark areas outside the breast, bounding-box annotations were converted to the DETR format, and all splits were performed at the patient level to prevent leakage. Input images were resized to 512×512, and each model was trained with a batch size of 32 for up to 400 epochs using a learning rate of 1×10 −4 and early stopping with a patience of 150 epochs. The decoder used 3 object queries and allowed at most 3 target objects per image during training. Image-level classification was optimized with focal loss, whereas the detection objective combined bounding-box regression, generalized IoU, and objectness 7 Nguyen Kha Nguyen Dinh Nguyen Ho Tran Ling Ho-Pham Pham Le Table 2: Dataset statistics after preprocessing and patient-level splitting. All images were cropped to retain only the breast region. OPTIMAM (OMI-DB) Oncology Hospital (SGM1k) StatisticTrain TestTotalTrain TestTotal Images (unique)19780 48632464327767493525 Patients (unique)6332 151978518022001002 Images with > 1 box1291202149342488512 Maximum boxes/image16916444 Benign12106 30561516211353191454 Malignant7674 1807948116414302071 terms with weights λ bbox =5.0, λ GIoU =2.0, and λ obj =1.0. Performance was evaluated using AUC, sensitivity, and specificity for classification, and mAP@.5 together with recall@.25 for lesion localization. Aside from backbone initialization, all training and evaluation settings were identical across experiments. Experiments were performed on a Linux server with 202.4 GB RAM and two NVIDIA L40 GPUs. Each L40 is an Ada Lovelace data-center GPU with 48 GB GDDR6 ECC memory and a 300 W maximum power rating. Appendix D. Grad-CAM Visualization To provide qualitative insight into model behavior, we examined Grad-CAM visualizations for representative mammography cases across the evaluated backbones. Figure 2 shows that the stronger-performing backbones, particularly ConvNeXtV2 and DINOv3, more consis- tently concentrated attention on clinically relevant suspicious regions, whereas ResNet50 and MambaVision tended to produce less focused or less well-aligned responses in more dif- ficult cases. These qualitative patterns are broadly consistent with the quantitative results in Table 1, where ConvNeXtV2 and DINOv3 achieved stronger overall classification and localization performance. From an interpretability perspective, these maps should be viewed as supportive visual cues rather than definitive lesion localization. In mammography, exact lesion boundaries can be difficult to define because abnormalities are often small, subtle, and partially obscured by overlapping dense tissue. In this setting, approximate attention to suspicious regions may still be useful for highlighting candidate areas for review, even when the highlighted region does not precisely match the annotated box. Appendix E. Ethical Approval Use of the OPTIMAM data in this study was conducted under the data access agreement dated 17/05/2023 between Cancer Research Horizons, Saigon Precision Medicine Research Center (SaigonMEC), and Royal Surrey NHS Foundation Trust. Use of the SGM1k cohort 8 Multi-task DETR’s backbone for Mammography Figure 2: Qualitative comparison of Grad-CAM maps across backbones was approved by Ho Chi Minh Oncology Hospital. All data use and analysis were performed in accordance with the relevant institutional approvals. Acknowledgments The authors sincerely thank the doctors and staff at Pham Ngoc Thach University of Medicine for their valuable clinical and academic support. We also gratefully acknowl- edge the doctors and staff at Ho Chi Minh City Oncology Hospital for their assistance with clinical coordination and data-related activities. We further thank the student volunteers for their dedicated support in data collection and preparation, and the UTS eResearch team for their technical support and research infrastructure assistance. 9