Paper deep dive
Empirical investigation of 3D CT Foundation Models and Unsupervised Adaptation for Head and Neck Cancer Recurrence Prediction
Bilel Guetarni, Feryal Windal, David Pasquier, Halim Benhabiles
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/4/2026, 3:36:21 AM
Summary
This study benchmarks 3D CT foundation models (VFMs) for predicting recurrence-free survival in head and neck cancer, comparing them against traditional radiomics. Using datasets RADCURE and Head-Neck-PET-CT, the authors evaluate various VFM architectures (CT-CLIP, CT-FM, VISTA3D, SUPREM) and adaptation strategies (ProtoNet, Cox-like ProtoNet) with modality fusion techniques (concatenation, gating, attention). Results indicate that while VISTA3D performs best on internal validation, significant performance drops occur on external cohorts, highlighting generalization challenges. Integrating imaging features with clinical data remains the most accurate approach, though universal generalization across diverse clinical contexts remains a substantial challenge.
Entities (12)
Relation Signals (8)
Integration of imaging features with clinical data → ismostaccurateapproachfor → prognostic_prediction
confidence 95% · integration of imaging features with clinical data remains the most accurate approach for prognostic prediction
Head-Neck-PET-CT → usedfor → external_validation
confidence 95% · HN-PETCT is exclusively used as an external validation dataset
RADCURE → usedfor → training_and_internal_validation
confidence 95% · For RADCURE, we use the train/test split... Training patients are used for hyperparameters search and test patients are held out for evaluation
3D CT Foundation Models → arealternativeto → Radiomics
confidence 90% · offering a compelling alternative to traditional radiomics
Radiomics → suffersfrom → reproducibility_issues
confidence 90% · radiomics which is known to suffer from reproducibility issues
VISTA3D → outperforms → CT-CLIP
confidence 85% · On RADCURE, VISTA3D features clearly outperform the other VFMs
VISTA3D → achievesbestperformancewith → imaging_features_alone
confidence 80% · VISTA3D reaches its best performance when imaging features are used alone
CT-CLIP → exhibitsbettergeneralizationthan → VISTA3D
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid emergence of 3D CT foundation models has opened new avenues for predictive modeling from CT imaging, offering a compelling alternative to traditional radiomics which is known to suffer from reproducibility issues and sensitivity to acquisition protocol variations. Yet, as these models grow in availability, a critical need arises to evaluate how well their learned representations generalize across diverse clinical settings and whether adaptation to specific downstream tasks is necessary to unlock their full potential. To address these questions, we benchmarked several 3D CT foundation models for predicting recurrence-free survival in head and neck cancer across two public datasets totaling 3,644 patients, evaluating various adaptation strategies and modality fusion mechanisms. Our findings reveal persistent difficulty in identifying features that generalize consistently across different imaging distributions, as evidenced by significant performance drops on external validation cohorts. Ultimately, the integration of imaging features with clinical data remains the most accurate approach for prognostic prediction, though achieving universal generalization across varied clinical contexts continues to represent a substantial challenge for the current generation of models.
Tags
Links
- Source: https://arxiv.org/abs/2608.00071v1
- Canonical: https://arxiv.org/abs/2608.00071v1
Trouble viewing inline? Open PDF directly →
Full Text
36,577 characters extracted from source content.
Expand or collapse full text
Empirical investigation of 3D CT Foundation Models and Unsupervised Adaptation for Head and Neck Cancer Recurrence Prediction Bilel Guetarni ⋆1 , Feryal Windal 3,4,5,6,7 , David Pasquier 1,2 , and Halim Benhabiles 8 1 University of Lille, CRISTAL UMR CNRS 9189, France 2 Academic Department of Radiation Oncology, Centre Oscar Lambret, Lille, France 3 Junia, Lille, France 4 UMR 8520, CNRS, France 5 Centrale Lille, France 6 Univerity of Polytechnique Hauts-de-France, Lille, France 7 University of Lille, France 8 University of Lille, Centre for Digital Systems, IMT Nord Europe, Institut Mines-Télécom, Lille, France Abstract. The rapid emergence of 3D CT foundation models has opened new avenues for predictive modeling from CT imaging, offering a com- pelling alternative to traditional radiomics which is known to suffer from reproducibility issues and sensitivity to acquisition protocol variations. Yet, as these models grow in availability, a critical need arises to evaluate how well their learned representations generalize across diverse clinical settings and whether adaptation to specific downstream tasks is neces- sary to unlock their full potential. To address these questions, we bench- marked several 3D CT foundation models for predicting recurrence-free survival in head and neck cancer across two public datasets totaling 3,644 patients, evaluating various adaptation strategies and modality fusion mechanisms. Our findings reveal persistent difficulty in identifying fea- tures that generalize consistently across different imaging distributions, as evidenced by significant performance drops on external validation co- horts. Ultimately, the integration of imaging features with clinical data remains the most accurate approach for prognostic prediction, though achieving universal generalization across varied clinical contexts contin- ues to represent a substantial challenge for the current generation of models. Keywords: Medical Imaging· Foundation models· Benchmarking· Head and Neck cancer· Modality fusion 1 Introduction Head and neck cancer is a major public health concern, accounting for more than 1.7 million new cases and over 500,000 deaths annually [19]. Despite advances ⋆ Corresponding author (bilel.guetarni@univ-lille.fr). arXiv:2608.00071v1 [cs.CV] 29 Jul 2026 2B. Guetarni et al. in multimodal therapies, disease recurrence remains frequent, with over 50% of patients experiencing relapse within two years after initial treatment [19]. This high rate of recurrence substantially compromises overall survival and under- scores the urgent need for more effective risk stratification, as well as adaptive treatment and follow-up strategies. In this context, developing reliable recurrence prognostic tools is essential to support clinicians in identifying high-risk patients and enabling more personal- ized treatment. In this sense, patients may benefit from treatment escalation [16] or de-escalation which may reduce risks of radiotherapy-induced toxicities [1]. As 3D Computed Tomography imaging (CT scans) is routinely used to visualize head and neck tumors and for post-treatment follow-up, it may provide a rele- vant modality for recurrence prediction. Moreover, its integration with patient clinical information could further enhance predictive performance. A prominent line of research for characterizing these CT scans involves ra- diomics, which utilizes mathematically defined descriptors to characterize the shape, intensity, and texture of regions-of-interest. While radiomics provides a structured framework for analyzing medical images, it is hindered by signifi- cant reproducibility issues; indeed radiomics are known to be protocol-specific and not generalize across diverse imaging protocols and reconstruction kernels [25,11]. In particular, a recent study showed a difficulty in reproducing associ- ations between radiomics and clincial endpoints identified in a prior work [3]. In parallel, an alternative approach has emerged through the development of 3D vision foundation models (VFMs) that are pre-trained on large quantities of CT scans. Unlike traditional methods, these models leverage unsupervised pre- text tasks to learn general invariant features directly from the data distribution [15]. By capturing the latent structure of volumetric data without the need of annotations, these models generate high-level representations offering a general characterization of complex morphological patterns. Despite the availability of these 3D CT foundation models, their utility in prognostic tasks depends heavily on how they are integrated into downstream pipelines. When these models are employed as frozen feature extractors, the primary challenge shifts from the pre-training itself to the design and adapta- tion of the subsequent predictor. There is currently a lack of systematic studies addressing how to map these high-dimensional latent embeddings to specific clinical outcomes, particularly regarding how to fuse them with complementary modalities or optimize the classification head. Consequently, there is a need for comprehensive benchmarks permitting to compare different 3D CT foundation models and evaluate various approaches for adapting downstream predictors to these frozen backbones. Such a framework is essential for moving toward a more structured understanding of how to leverage large-scale pre-trained models for reliable prognostic tasks in oncology. In this work, we address the aforementioned questions by benchmarking sev- eral state-of-the-art VFMs on the task of recurrence prediction at two years in Empirical investigation of 3D CT Foundation Models3 patients treated with radiotherapy in a curative intent for head and neck can- cer. We intentionally exclude traditional radiomics from this benchmark, as our primary objective is an empirical investigation of CT foundation models. We have systematically investigated different strategies for training the predictor, as well as multiple approaches for multimodal data fusion combining CT imag- ing and clinical variables (including dosimetric information). The performance evaluation of the predictors has been conducted using a comprehensive set of metrics to ensure robust and reliable comparisons. All experiments have been performed on a large-scale public dataset, namely RADCURE [24,23], compris- ing 3,346 patients, and all classifiers have been further validated on an external public dataset, Head-Neck-PET-CT [21], comprising 298 patients, to evaluate their potential for generalizability. For reproducibility purposes, the code and extracted VFM features will be made publicly available upon publication. 2 Related Works Most studies on recurrence predictive models employ radiomics as imaging fea- tures due to their genericity and easy computation simplified by standardized libraries (e.g., PyRadiomics) [7,20,14,17]. Applied to CT volumes with delin- eated regions-of-interest, they can be used to build risk stratification models using for example Cox’s proportional hazards [14] or random forests [7]. Specific time endpoints models are also widely used (e.g., recurrence at 2 or 5 years) [17,20] as they have the advantage of fixing a particular point in time for the event in question. More recently, for the reproducibility reasons stated earlier [25,11], several studies have leveraged deep learning-based imaging characterization driven by the availability of public datasets such as RADCURE [23]. These studies employ diverse types of end-to-end deep learning architectures such as Graph Neural Networks [2] or Convolutional Neural Networks [12]. Although these methods achieve strong performances, they often require computationally heavy hyper- parameters tuning steps and can fail to generalize to external cohorts with differ- ent imaging features distribution due to misalignment in acquisition protocols. VFMs have been proposed as a solution to these challenges as they do not require any hyperparameter tuning once trained, and are pre-trained on large heteroge- neous datasets covering many centers in a task agnostic manner [15,8,9,13]. Pai et al. [15] pre-trained a SegResNet on a self-supervised contrastive task to generate similar latent embeddings for two augmented versions of the same patch. They propose a variant of the SimCLR task [4] to guarantee that positive and negative pairs are different. The model showed strong results on several downstream task scenarios as organ segmentation, tumor segmentation and head abnormality clas- sification. The authors also found the model to be robust under perturbations. Similarly, Hamamci et al. [8] performed contrastive self-supervised learning to train a combination of language-image model composed of a text and an image encoders. These are jointly optimized to generate similar latent embedding for 4B. Guetarni et al. pairs of text and image extracted from a curated dataset collected beforehand. When evaluated, the image encoder outperforms fully supervised approaches on multi-abnormality detection and case retrieval. Li et al. [13] pre-trained diverse deep learning architectures for segmentation of 25 anatomical structures on a de- tailed voxel-level annotated dataset. They show that their trained models have stronger transfer abilities to new 3D segmentation tasks than existing models, especially on classes with limited training samples. He et al. [9] proposed a foun- dation model for 3D CT segmentation that supports 127 different structures including organs, bones and arteries. Their interactive training framework al- lows to generate, during the training, new annotations on unlabelled images by allowing human annotators to be solicited. Their model was shown to be com- petitive with specialized models on the full range of supported classes, as well as, outperforming state-of-the-art methods in zero-shot transfer tasks. Fine-tuning these large-scale models comes with a substantial computational cost. Utilizing them as feature extractors provides a more practical alternative, similar to radiomics. In this setup, only the subsequent, smaller predictor is trained, often from scratch, to map the VFM features to the target outcome. A significant challenge remains, however, as the limited availability of labeled data may be insufficient for this downstream classifier to effectively adapt to the fea- ture space of the foundation model. Indeed, such constraints make it challenging to adapt to never seen before categories especially in the context of deep neural networks, prone to overfitting. To fix this issue, few-shot learning is a paradigm in which a classifier is accommodated to new classes with only a few samples at hand. In this context, Snell et al. [18] proposed prototypical networks (Pro- toNet), in which a model is optimized to generate a metric space inside which new samples can be classified by computing the distance towards prototypes (i.e., class centroids). To this end the model is trained, from a few number of sam- ples, to minimize the distance between the samples of the same class and their centroid, and maximize the distance with respect to all other classes centroids. Comparative studies on natural images datasets demonstrated the advantage of ProtoNet over other competitive methods of the literature. Hu et al. [10] pro- posed an extension of this approach by incorporating it into a three stages pro- cess including a pre-training stage with self-supervised contrastive tasks. Their experiments suggest a positive margin of downstream performance compared to previous methods. The promising results demonstrated by these diverse foundation models sug- gest they could provide a more robust alternative to traditional radiomics for clinical applications. However, their specific utility for predicting oncological out- comes and the most effective way to adapt their features to such tasks remains to be fully established. This highlight the need to systematically benchmark these models to evaluate their predictive potential and generalization capabilities, in- cluding in the context of head and neck cancer recurrence. Empirical investigation of 3D CT Foundation Models5 3 Foundation model-based prognosis prediction PATIENT DATA Clinical data FUSION BLOCK minimize distance maximize distance 풛 푰 풛 푪 •Age •Sex •Treatment •GTV dose •... Merged features 풛 CLASSIFICATION HEAD Prediction of recurrence probability 3D CT volume (centerd around GTV) 3D Foundation Model UNSUPERVISED PRE-TRAINING STAGE Fig. 1: Proposed empirical investigation of combining Vision Foundation Model and unsupervised adaptation for recurrence prediction. In the bottom grey box we illustrate the difference between ProtoNet (left) and our proposed variant (right). As illustrated in Figure 1, we study three key aspects of VFM-based ap- proaches that build prognostic predictive models. We study the impact of the foundation models conception on the features relevance for the prediction task. This includes the model architecture, pre-training objective, as well as the data quantity and diversity. The second element concerns how the imaging features and clinical data are fused together. As various strategies exist and selecting the most appropriate one is not trivial, we compare different approaches. The third element is the most efficient way for training the predictor from the modalities features. 3.1 Impact of foundation model conception on features transfer We selected four 3D CT VFMs from the literature, chosen to cover different pre-training strategies, model size, as well as different training data quantity and source. CT-FM [15] is a 77M parameters SegResNet encoder pre-trained on 148K CT scans with contrastive self-supervised learning. CT-CLIP [8] is an image-language model pre-trained using contrastive image-language pre-training on 25K pairs of scans and textual reports. We only use the 25M parameters im- age encoder consisting of a 3D vision transformer. VISTA3D [9] is a SegResNet 6B. Guetarni et al. encoder trained on 127 anatomical classes for 3D CT segmentation. The authors curated a dataset of 11K 3D CT scans to train the model. SUPREM [13] is a suite of pre-trained 3D segmentation models covering 25 anatomical structures trained on 9K CT scans; we employ the 4.7M parameters SegResNet version published by the authors in our experiments. In this work, VFMs are used as feature extractors of the CT volume, centered around the gross tumor volume (GTV). More precisely, expert annotation of the GTV is used to center crop the CT, while respecting the input size requirements of each VFM. For segmentation models, the output of the encoder’s last layer is converted into an embedding vector through a global average pooling. For the other models, the same operation is applied on their last layer. 3.2 Modality fusion functions Once imaging features are extrated by the VFM, they are combined with clin- ical data to produce a multimodal patient representation. The choice of this fusion function directly affects how much each modality contributes to the final prediction. We compare three fusion mechanisms. Concatenation Imaging and clinical features are projected to a common em- bedding dimension by a linear projection. The projected features are then con- catenated followed by a 2-layer MLP with GELU activation and dropout. Gating Inspired by the gating mechanism used in recurrent neural networks [5], we introduce weighting factors that modulate the contribution of each modality. Similar to concatenation, imaging and clinical features are first projected to a common embedding followed by concatenation and batch normalization. The concatenated vector is then used to compute gating scores through a linear layer, which are then used to weight the modalities features vectors in the fused vector. Finally we apply a 2-layer MLP with GELU activation and dropout. Attention Finally, we consider the attention mechanism from the Transformer architecture [22]. The modalities features vectors are projected to a common latent embedding dimension, and we compute attention scores using a learn- able query vector (initialized randomly before training). The fused vector is the weighted sum of the modalities vectors using these scores. We also apply a 2-layer MLP with GELU activation and dropout to the fused vector. 3.3 Predictor training strategies The overall predictor, consisting of the fusion module and the classification head, is trained in two different ways. First, we train the predictor from scratch in a classical supervised setting. Second, we propose an unsupervised adaptation of Empirical investigation of 3D CT Foundation Models7 the predictor to the VFM features. This is implemented as a pre-training ob- jective that optimizes the fusion module to cluster patients according to their prognostic class. This is due to the fact that the available patient cohorts may be insufficient for a downstream predictor to naturally adapt to frozen founda- tion model embeddings. We therefore employ an unsupervised adaptation step to align the fusion module with the latent imaging representations. More specifically, for the second training strategy we base our work on the few-shot learning paradigm [18]. In a few-shot learning setting, a classifier is trained to learn new classes with a few samples of each class. This can be chal- lenging and lead to overfitting, especially with high-dimensional features. For example, Snell et al. [18] proposed ProtoNet to train a model by optimizing its latent embedding such that samples belonging to the same class are close to one another (i.e., close to the class centroid). In ProtoNet, one is given a small sup- port of samples belonging to two classes 9 , S =(x i ,y i )|i = 1...N, with x∈R D and y ∈0, 1. S + is the subset of positives samples (y = 1), and similarly, S − the set of negative ones (y = 0). For a sample belonging to the positive class, the training objective consists of minimizing the distance with the class centroid defined as: c + = 1 |S + | X (x i ,y i )∈S + f θ (x i )(1) The negative class prototype c − is defined equivalently. The model f θ :R D → R M is trained to maximize the likelihood of a sample belonging to its class. In the case of a positive sample (y = 1), this can be formulated as: p θ (y = 1|x) = exp(−d(f θ (x),c + )) exp(−d(f θ (x),c + )) + exp(−d(f θ (x),c − )) (2) where d is a distance function. In the standard ProtoNet, a sample is compared to the centroid of the op- posing class (see denominator of Eq. 2). We investigate whether it is better to compare it to each of its individual samples. We propose a variant, inspired by the Cox survival model that maximizes the partial likelihood, i.e., the proba- bility of observing an event for a given sample relative to all others. Following this, we modify the class-level distance computation with an instance-level one for the opposite class: p θ (y = 1|x) = exp(−d(f θ (x),c + )) exp(−d(f θ (x),c + )) + P (x ′ i ,y i )∈S − exp(−d(f θ (x),f θ (x ′ i ))) (3) In both cases, the loss function is the negative log-likelihood which is defined as: − logp θ (y|x). 9 this is specific to our downstream task, the approach is generic to any number of classes and we refer the reader to the original paper. 8B. Guetarni et al. We qualify this method as unsupervised since the patients labels are not directly used in the loss formulation. Instead, the class is used to group patients together for computing the prototypes and samples distances. 4 Experiments Datasets. To perform our study we employed two publicly available datasets: RADCURE [24,23] and Head-Neck-PET-CT (HN-PETCT) [21]. Each patient has a planning CT, from which we extract VFM features, and its associated clinical data which includes: age, sex, treatment type (radiotherapy alone or combined with chemotherapy) and the TNM stage. We also include the GTV total dose computed as the median voxel value in the dose maps using the GTV contours provided with the CT. Age and dose are normalized and categor- ical variables one-hot encoded. To avoid under-represented clinical categories, we apply the following inclusion criteria: no surgery prior to radiotherapy, non metastatic (M0) oropharyngeal cancer. We studied 2-year recurrence as the clin- ical endpoint of prediction. This endpoint was selected because most recurrences in head and neck cancer occur during the first two years following treatment [19] Survival data are binarized accordingly: patients with an event before the end- point are labeled positive, others negative. Patients with last follow-up before the endpoint and no observed event are censored and therefore excluded of the study. For RADCURE, we use the train/test split provided in the TCIA repos- itory [24] after applying the previously mentioned inclusion criteria. Training patients are used for hyperparameters search and test patients are held out for evaluation. HN-PETCT is exclusively used as an external validation dataset. In total, we have 1,554 training, 499 validation and 532 testing samples. Evaluation metrics. We report three classification metrics: AUC, F1-score and balanced accuracy (BA, the average of specificity and sensitivity). AUC and F1-score are the most common metrics in healthcare literature [6]. BA is included because survival datasets can become highly imbalanced once binarized: 25% of RADCURE and 20% of HN-PETCT patients are positive. Therefore, selecting appropriate metrics that are not impacted by class imbalance is necessary to perform clinically relevant comparisons. BA is robust to this imbalance and offers a good trade-off between sensitivity and specificity. Reported metrics are averaged over 10 bootstraps. Training setup. As stated earlier, the VFMs are used as feature extractors and are therefore, not trained, unlike the fusion module and classification head. Unless stated otherwise, all models were trained for 200 epochs with a batch size of 16, an Adam optimizer with learning rate of 5×10 −5 and dropout (p = 0.5). To counter class imbalance, we apply random undersampling on the majority class to match the size of the minority one. For unsupervised pre-training, models are trained for 2000 steps with a batch size of 128 using Adam with a cosine Empirical investigation of 3D CT Foundation Models9 Table 1: Comparison of different pre-training strategies and Vision Foundation Models. RADCUREHN-PETCT VFMpre-trainingmultimodalAUC BA F1AUC BA F1 CT-CLIP none✓0.625 0.591 0.3900.641 0.600 0.391 ProtoNet✓0.589 0.502 0.2870.573 0.503 0.275 Cox-like ProtoNet✓0.637 0.596 0.4370.658 0.582 0.406 SUPREM none✓0.596 0.528 0.2220.572 0.507 0.087 ProtoNet✓0.619 0.508 0.3080.584 0.500 0.229 Cox-like ProtoNet✓0.607 0.574 0.4100.485 0.500 0.013 CT-FM none✓0.643 0.611 0.4420.539 0.515 0.349 ProtoNet✓0.655 0.617 0.4510.500 0.491 0.338 Cox-like ProtoNet✓0.647 0.613 0.4400.541 0.526 0.348 VISTA3D none0.668 0.638 0.4760.545 0.518 0.305 ProtoNet0.661 0.627 0.4630.545 0.528 0.315 Cox-like ProtoNet0.662 0.627 0.4630.560 0.530 0.317 learning rate schedule: initialized at 10 −6 , warmed-up over 50 steps to 5× 10 −5 then decayed back to 10 −6 . L2 weights decay (λ = 0.1) is applied in both stages and we use binary cross-entropy as the loss function. 5 Results 5.1 Foundation model embedding Tables 1 reports the performance of the four VFMs under different pre-training strategies. On RADCURE, VISTA3D features clearly outperform the other VFMs, particularly when the predictor is trained without pre-training (AUC 0.668, BA 0.638 and F1 0.476). Notably, VISTA3D reaches its best performance when imag- ing features are used alone: adding clinical data decreases performance by 0.029 AUC, 0.026 BA and 0.029 F1. However, when evaluated on the external cohort, we observe a significant drop of performance across every training strategies (i.e., with or without pre- training). In contrast, CT-CLIP displays better generalization capability than other VFM: combined with our Cox-like ProtoNet pre-training, it reaches an AUC of 0.658, 0.582 BA and 0.406 F1. This indicates that CT-CLIP features are more robust and can better generalize to other acquisition protocols. In- deed, the difference of AUC between the internal and external cohorts is 0.021 for CT-CLIP, 0.035 for SUPREM, 0.123 for VISTA3D and 0.155 for CT-FM. Robustness to data distribution shifts is, indeed, a characteristic expected from foundation model embeddings. We only found that CT-CLIP exhibits such ro- bustness across all considered metrics. 10B. Guetarni et al. Table 2: Comparison of fusion functions and unsupervised adaptation. We use CT-CLIP for imaging features. RADCUREHN-PETCT modalityfusion functionAUCBAF1AUCBAF1 w/o pre-training clinical0.5530.5270.2930.5520.5300.291 image0.5190.5020.1320.4840.5090.143 multiconcat0.5940.5010.2190.5930.5080.244 multigating0.6250.5910.3900.641 0.600 0.391 multiattention0.5600.5100.1520.5790.5110.154 Cox-like ProtoNet clinical0.5870.5200.2750.5750.5070.218 image0.5060.5030.2740.5050.5020.232 multiconcat0.5900.5000.0820.5800.5000.078 multigating0.637 0.596 0.4370.658 0.582 0.406 multiattention0.5270.5200.3100.5630.5390.354 Nonetheless, the reported classification performances are not sufficient to argue in favor of VFMs as ready-to-use embedding models for recurrence pre- diction, specifically with limited training data. We observe that most AUCs on the external cohorts are below 0.6 and every model shows an F-score lower than 0.5. 5.2 Contribution of multimodality for recurrence prediction Since CT-CLIP showed the highest generalization capability on the external cohort (see section 5.1), we use it with our Cox-like ProtoNet pre-training as the base configuration to compare the three fusion functions described in section 3.2. Results are reported in Table 2 and Figure 2. We can observe that gating is clearly the most effective fusion function on both cohorts, with an AUC of 0.637 on RADCURE and 0.658 on HN-PETCT. It outperforms both concatenation (0.590 and 0.580) and attention (0.527 and 0.563). Moreover, we can note that imaging features only do not compete with clinical data that consistently perform better across the all metrics and datasets. 5.3 Pre-training to improve generalization Table 1 shows that unsupervised pre-training improves external generalization on both RADCURE and HN-PETCT for most VFMs. On RADCURE, pre- training also improves performance for the majority of them, with the exception of VISTA3D whose features are already well-suited to the internal cohort with- out adaptation. Our Cox-like ProtoNet variant is the best pre-training strat- egy in three out of four cases. On HN-PETCT, for CT-CLIP, VISTA3D, and SUPREM the best AUC is achieved with pre-training, with respective gains of Empirical investigation of 3D CT Foundation Models11 +0.017, +0.015, and +0.012 compared to no pre-training. CT-FM shows benefit of pre-training in terms of AUC (+0.017) and F-score (+0.015). 0.0 0.2 0.4 0.6 0.8 1.0 AUC RADCURE (AUC)HN-PETCT (AUC) 0.0 0.2 0.4 0.6 0.8 1.0 BA RADCURE (BA)HN-PETCT (BA) noneprotonetcox+protonet pre-training strategy 0.0 0.2 0.4 0.6 0.8 1.0 F-score RADCURE (F-score) noneprotonetcox+protonet pre-training strategy HN-PETCT (F-score) modality clinical image both (concat) both (gated) both (attention) Fig. 2: AUC, BA and F-score reported on both RADCURE and HN-PETCT using CT-CLIP for imaging features. 12B. Guetarni et al. Table 2 also provides results indicating benefits from unsupervised pre-training. In RADCURE most models AUC improve after applying pre-training. This find- ing does not translate however on the external cohort; which may be explained by the fact that pre-training is performed on the training cohort data. 6 Conclusion In this work, we benchmarked four 3D vision foundation models (VFMs) for 2- year recurrence prediction in head and neck cancers from computed tomography images. We included three modality fusion functions, to merge clinical and imag- ing features, and two training strategies. Our results show that no single VFM dominates across all settings: VISTA3D achieves the best internal performances, while CT-CLIP features demonstrated better generalization on the external co- hort. As an attempt to explain the superiority of CT-CLIP features generaliza- tion on the external cohort, one may notice that it is the only imaging foundation model that was pre-trained via an image-text contrastive task. Its superior per- formance could suggests that text-derived signals during self-supervised training, help the image encoder learns imaging features that are better at generalizing be- yond its pre-training data. We leave this as an open question for further research. Additionally, we investigate the choice of the fusion function. Gating mechanism, inspired by recurrent neural network, showed improved prediction compared to other fusion functions (AUC +0.078 and +0.095 over concatenation and atten- tion respectively on HN-PETCT). Additionally, we observed that unsupervised adaptation, in the form of few-shot learning pre-training, helps generalization. We propose a variant of an existing few-shot learning paradigm, that yields the best pre-training strategy results in three out of four VFMs. Despite its wide adoption, we found the ROC AUC metric insufficient for comparing and clearly identifying the best predictive models when dealing with imbalanced datasets. Instead, the addition of balanced accuracy (i.e., average of specificity and sensi- tivity) and F1-score was crucial to assess the capacity of each model to correctly identify patients at risk of recurrence while minimizing the rate of false posi- tives. Finally, this benchmark therefore identifies the main barriers to leveraging large-scale pre-trained models for individualized risk profiling in head and neck cancer. Acknowledgments. This study was funded by the AAP SEQ-RTH22 from the French National Cancer Institute. Disclosure of Interests. The authors have no competing interests to declare that are relevant to the content of this article. References 1. Adelstein, D.J., Ismaila, N., Ku, J.A., Burtness, B., Swiecicki, P.L., Mell, L., Beitler, J.J., Gross, N., Jones, C.U., Kaufman, M., et al.: Role of treatment dein- tensification in the management of p16+ oropharyngeal cancer: Asco provisional clinical opinion. Journal of Clinical Oncology 37(18), 1578–1589 (2019) Empirical investigation of 3D CT Foundation Models13 2. Bae, J., Kapse, S., Zhou, L., Mani, K., Prasanna, P.: Hog-net: Hierarchical multi- organ graph network for head and neck cancer recurrence prediction from ct images. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. p. 317–327. Springer (2024) 3. Berger, T., Noble, D.J., Yang, Z., Shelley, L.E., McMullan, T., Bates, A., Thomas, S., Carruthers, L.J., Beckett, G., Duffton, A., et al.: Assessing the generalisability of radiomics features previously identified as predictive of radiation-induced sticky saliva and xerostomia. Physics and Imaging in Radiation Oncology 25, 100404 (2023) 4. Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for con- trastive learning of visual representations. In: International conference on machine learning. p. 1597–1607. PmLR (2020) 5. Cho, K., Van Merriënboer, B., Gulçehre, Ç., Bahdanau, D., Bougares, F., Schwenk, H., Bengio, Y.: Learning phrase representations using rnn encoder–decoder for statistical machine translation. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). p. 1724–1734 (2014) 6. Collins, G.S., Dhiman, P., Ma, J., Schlussel, M.M., Archer, L., Van Calster, B., Harrell, F.E., Martin, G.P., Moons, K.G., Van Smeden, M., et al.: Evaluation of clinical prediction models (part 1): from development to external validation. bmj 384 (2024) 7. Haider, S.P., Sharaf, K., Zeevi, T., Baumeister, P., Reichel, C., Forghani, R., Kann, B.H., Petukhova, A., Judson, B.L., Prasad, M.L., et al.: Prediction of post- radiotherapy locoregional progression in hpv-associated oropharyngeal squamous cell carcinoma using machine-learning analysis of baseline pet/ct radiomics. Trans- lational oncology 14(1), 100906 (2021) 8. Hamamci, I.E., Er, S., Wang, C., Almas, F., Simsek, A.G., Esirgun, S.N., Dogan, I., Durugol, O.F., Hou, B., Shit, S., et al.: Generalist foundation models from a multimodal dataset for 3d computed tomography. Nature Biomedical Engineering p. 1–19 (2026) 9. He, Y., Guo, P., Tang, Y., Myronenko, A., Nath, V., Xu, Z., Yang, D., Zhao, C., Si- mon, B., Belue, M., et al.: Vista3d: A unified segmentation foundation model for 3d medical imaging. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 20863–20873 (2025) 10. Hu, S.X., Li, D., Stühmer, J., Kim, M., Hospedales, T.M.: Pushing the limits of simple pipelines for few-shot learning: External data and fine-tuning make a difference. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 9068–9077 (2022) 11. Lambin, P., Leijenaar, R.T., Deist, T.M., Peerlings, J., De Jong, E.E., Van Tim- meren, J., Sanduleanu, S., Larue, R.T., Even, A.J., Jochems, A., et al.: Radiomics: the bridge between medical imaging and personalized medicine. Nature reviews Clinical oncology 14(12), 749–762 (2017) 12. Le, W.T., Vorontsov, E., Romero, F.P., Seddik, L., Elsharief, M.M., Nguyen-Tan, P.F., Roberge, D., Bahig, H., Kadoury, S.: Cross-institutional outcome prediction for head and neck cancer patients using self-attention neural networks. Scientific reports 12(1), 3183 (2022) 13. Li, W., Yuille, A., Zhou, Z.: How well do supervised 3d models transfer to medical imaging tasks? arXiv preprint arXiv:2501.11253 (2025) 14. Lv, W., Xu, H., Han, X., Zhang, H., Ma, J., Rahmim, A., Lu, L.: Context-aware saliency guided radiomics: Application to prediction of outcome and hpv-status from multi-center pet/ct images of head and neck cancer. Cancers 14(7), 1674 (2022) 14B. Guetarni et al. 15. Pai, S., Hadzic, I., Bontempi, D., Bressem, K., Kann, B.H., Fedorov, A., Mak, R.H., Aerts, H.J.: Vision foundation models for computed tomography. arXiv preprint arXiv:2501.09001 (2025) 16. Rühle, A., Sprave, T., Kalckreuth, T., Stoian, R., Haehl, E., Zamboglou, C., Laszig, R., Knopf, A., Grosu, A.L., Nicolay, N.H.: The value of moderate dose escalation for re-irradiation of recurrent or second primary head-and-neck cancer. Radiation Oncology 15(1), 81 (2020) 17. Shannon, N.B., Lyer, N.G., Chua, M.L.K.: Leveraging artificial intelligence and ra- diomics for improved nasopharyngeal carcinoma prognostication. Cancer Medicine 14(6), e70706 (2025) 18. Snell, J., Swersky, K., Zemel, R.: Prototypical networks for few-shot learning. Ad- vances in neural information processing systems 30 (2017) 19. Sun, S., Lu, M., Wei, S., Liang, Y., Zhang, Z., Wang, H., Si, L.: Global burden and cross-country inequalities in head and neck cancer from 1992 to 2021: results from the global burden of disease study. Health Economics Review 15(1), 84 (2025) 20. Teng, X., Zhang, J., Ma, Z., Zhang, Y., Lam, S., Li, W., Xiao, H., Li, T., Li, B., Zhou, T., et al.: Improving radiomic model reliability using robust features from perturbations for head-and-neck carcinoma. Frontiers in oncology 12, 974467 (2022) 21. Vallières, M., Kay-Rivest, E., Perrin, L.J., Liem, X., Furstoss, C., Aerts, H.J., Khaouam, N., Nguyen-Tan, P.F., Wang, C.S., Sultanem, K., et al.: Radiomics strategies for risk assessment of tumour failure in head-and-neck cancer. Scientific reports 7(1), 10117 (2017) 22. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Advances in neural information pro- cessing systems 30 (2017) 23. Welch, M.L., Kim, S., Hope, A.J., Huang, S.H., Lu, Z., Marsilla, J., Kazmierski, M., Rey-McIntyre, K., Patel, T., O’Sullivan, B., et al.: Radcure: An open-source head and neck cancer ct dataset for clinical radiation therapy insights. Medical Physics 51(4), 3101–3109 (2024) 24. Welch, M., Kim, S., Hope, A., Huang, S., Lu, Z., Marsilla, J., Kazmierski, M., Rey-McIntyre, K., Patel, T., O’Sullivan, B., et al.: Computed tomography images from large head and neck cohort (radcure). The Cancer Imaging Archive 4 (2023). https://doi.org/https://doi.org/10.7937/J47W-NM11 25. Zhao, B., Tan, Y., Tsai, W.Y., Qi, J., Xie, C., Lu, L., Schwartz, L.H.: Reproducibil- ity of radiomics for deciphering tumor phenotype with imaging. Scientific reports 6(1), 23428 (2016)