Paper deep dive
Topology-Driven Transferability Estimation of Medical Foundation Models for Segmentation
Jiaqi Tang, Shaoyang Zhang, Xiaoqi Wang, Jiaying Zhou, Yang Liu, Qingchao Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/20/2026, 7:20:46 AM
Summary
The paper proposes a Topology-Driven Transferability Estimation (TTE) framework to select optimal medical foundation models for segmentation tasks without fine-tuning. It introduces Global Representation Topology Divergence (GRTD) and Local Boundary-Aware Topological Consistency (LBTC) to evaluate feature-label structural isomorphism and boundary separability, respectively. These metrics are fused via a Task-Adaptive mechanism based on semantic cardinality. Validated on the OpenMind benchmark, the method significantly outperforms existing baselines in predicting fine-tuning performance.
Entities (9)
Relation Signals (8)
Topology-Driven Transferability Estimation → includescomponent → Local Boundary-Aware Topological Consistency
confidence 95% · ...(2) Local Boundary-Aware Topological Consistency (LBTC), which assesses manifold separability...
Topology-Driven Transferability Estimation → includescomponent → Global Representation Topology Divergence
confidence 95% · Our approach introduces three components: (1) Global Representation Topology Divergence (GRTD)...
Topology-Driven Transferability Estimation → validateson → OpenMind benchmark
confidence 92% · Validated on the large-scale OpenMind benchmark across diverse anatomical targets
Local Boundary-Aware Topological Consistency → assesses → manifold separability
confidence 90% · assesses manifold separability specifically at critical anatomical boundaries
Global Representation Topology Divergence → usesalgorithm → Minimum Spanning Tree
confidence 90% · Global Representation Topology Divergence (GRTD), utilizing Minimum Spanning Trees to quantify feature-label structural isomorphism
Topology-Driven Transferability Estimation → usesmetricforevaluation → Weighted Kendall's tau
confidence 88% · significantly outperforms state-of-the-art baselines by around 31% relative improvement in the weighted Kendall metric
Topology-Driven Transferability Estimation → outperforms → LogME
confidence 85% · traditional classification-based metrics (LogME, LEEP) exhibit severe negative correlations
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The advent of large-scale self-supervised learning (SSL) has produced a vast zoo of medical foundation models. However, selecting optimal medical foundation models for specific segmentation tasks remains a computational bottleneck. Existing Transferability Estimation (TE) metrics, primarily designed for classification, rely on global statistical assumptions and fail to capture the topological complexity essential for dense prediction. We propose a novel Topology-Driven Transferability Estimation framework that evaluates manifold tractability rather than statistical overlap. Our approach introduces three components: (1) Global Representation Topology Divergence (GRTD), utilizing Minimum Spanning Trees to quantify feature-label structural isomorphism; (2) Local Boundary-Aware Topological Consistency (LBTC), which assesses manifold separability specifically at critical anatomical boundaries; and (3) Task-Adaptive Fusion, which dynamically integrates global and local metrics based on the semantic cardinality of the target task. Validated on the large-scale OpenMind benchmark across diverse anatomical targets and SSL foundation models, our approach significantly outperforms state-of-the-art baselines by around 31% relative improvement in the weighted Kendall metric, providing a robust, training-free proxy for efficient model selection without the cost of fine-tuning. The code will be made publicly available upon acceptance.
Tags
Links
- Source: https://arxiv.org/abs/2602.23916v2
- Canonical: https://arxiv.org/abs/2602.23916v2
Trouble viewing inline? Open PDF directly →
Full Text
27,177 characters extracted from source content.
Expand or collapse full text
Topology-Driven Transferability Estimation of Medical Foundation Models for Segmentation Jiaqi Tang 1,4,5,6⋆ , Shaoyang Zhang 2⋆ , Xiaoqi Wang 3 , Jiaying Zhou 1 , Yang Liu 1 , and Qingchao Chen 1,4,5,6 1 Peking University, Beijing, China 2 Hohai University, Nanjing, China 3 Beijing Normal University-Hong Kong Baptist University United International College, Zhuhai, China 4 National Institute of Health Data Science, Peking University, Beijing, China 5 Institute of Medical Technology, Peking University, Beijing, China 6 State Key Laboratory of General Artificial Intelligence, Peking University jiaqi_tang, qingchao.chen@pku.edu.cn, zsyhhu04@163.com Qingchao Chen is the corresponding author Abstract. The advent of large-scale self-supervised learning (SSL) has produced a vast zoo of medical foundation models. However, selecting optimal medical foundation models for specific segmentation tasks re- mains a computational bottleneck. Existing Transferability Estimation (TE) metrics, primarily designed for classification, rely on global statis- tical assumptions and fail to capture the topological complexity essen- tial for dense prediction. We propose a novel Topology-Driven Transfer- ability Estimation framework that evaluates manifold tractability rather than statistical overlap. Our approach introduces three components: (1) Global Representation Topology Divergence (GRTD), utilizing Minimum Spanning Trees to quantify feature-label structural isomorphism; (2) Lo- cal Boundary-Aware Topological Consistency (LBTC), which assesses manifold separability specifically at critical anatomical boundaries; and (3) Task-Adaptive Fusion, which dynamically integrates global and local metrics based on the semantic cardinality of the target task. Validated on the large-scale OpenMind benchmark across diverse anatomical tar- gets and SSL foundation models, our approach significantly outperforms state-of-the-art baselines by around 31% relative improvement in the weighted Kendall’sτ, providing a robust, training-free proxy for efficient model selection without the cost of fine-tuning. The code will be made publicly available upon acceptance. Keywords: transferability estimation· transfer learning· foundation model selection ⋆ These authors contributed equally to this work. Accepted at MICCAI 2026. arXiv:2602.23916v2 [cs.CV] 23 May 2026 2J. Tang, S. Zhang et al. Previous Method(CCFV) VoCo Features SwinUNETR Features Topology-driven VoCo Topology SwinUNETR Topology Ground Truth: Finetuning On VoCo Encoder On SwinUNETR Encoder Dice Score Our Metric CCFV Metric Fig. 1: Left: Fine-tuning performance of diverse foundation models varies across downstream datasets, indicating that the optimal pre-trained encoder is task- dependent. Right: Visualizing transferability: Statistics vs. Topology. We com- pare the ground truth fine-tuning performance (Left) against rankings from the statistical metric (Middle) and our topology-driven metric (Right). 1 Introduction The rise of large-scale self-supervised learning (SSL) has created a diverse zoo of medical foundation models trained on massive unlabeled data to learn general- purpose representations [2, 3, 6, 16, 21, 25]. Yet, the best pre-trained encoder is highly task-dependent: as illustrated in Fig.1 (left), different downstream seg- mentation datasets favor different source models, making exhaustive fine-tuning a costly combinatorial search[20]. This motivates a training-free Transferability Estimation (TE) framework that predicts post-fine-tuning segmentation perfor- mance directly from pre-training features, enabling fast and resource-efficient model selection. Existing TE metrics, largely developed for image classification [4,13–15,23, 24], can be misaligned with the “SSL encoder + segmentation decoder” setting. Methods such as LEEP [13] and LogME [24] implicitly favor linear separability, while embedding-based scores such as GBC [14] and CCFV [23] often assume simple parametric statistics (e.g., Gaussianity). However, segmentation quality depends less on coarse global class separation and more on whether features preserve local geometric structure near high-frequency boundaries. As shown in Fig. 1 (right), purely statistical similarity may yield rankings inconsistent with fine-tuning outcomes, whereas a topology-aware view of the feature manifold better reflects boundary separability and thus transferability. Our method bridges these gaps through a topology-driven approach. Rather than forcing complex feature spaces into predefined statistical molds, we utilize non-parametric, graph-theoretic structures to quantify man- ifold alignment. By explicitly modeling local boundary regions and dynami- cally balancing global and local priors, our framework could match the specific structural complexity of diverse downstream medical targets. Specifically, our Topology-Driven Transferability Estimation3 method introduces three novel components tailored for medical segmentation, as shown in Figure 2: (1) Representation Topology Divergence (GRTD) quantifies structural alignment by measuring the discrepancy between Minimum Spanning Trees (MST) constructed in the feature and label spaces, capturing the over- all manifold isomorphism. (2) Local Boundary-Aware Topological Consistency (LBTC) explicitly assesses manifold separability at critical boundary patches, utilizing local MST graphs to ensure distinct decision boundaries in the pres- ence of background heterogeneity, where segmentation typically fails. (3) Task- Adaptive Topological Fusion dynamically calibrates the global and local metrics based on the target task’s complexity, optimally balancing the need for broad anatomical context versus fine-grained boundary details. We validate our framework on the large-scale OpenMind benchmark [20], covering 6 diverse anatomical segmentation tasks and a model zoo consisting of 7 mainstream SSL methods pre-trained on 114,000 3D volumes. Extensive exper- iments demonstrate that our topology-driven approach significantly outperforms existing baselines by around 31% relative improvement in weighted Kendall’s τ, offering a robust, source-free solution for efficient model selection in the era of medical foundation models. 2 Methodology 2.1 Problem Formulation and Framework Overview We address the problem of selecting the optimal pre-trained encoder from a SSL foundation model zoo M = φ k K k=1 for a target segmentation task D = (x i ,y i ) N i=1 without incurring the computational cost of fine-tuning. Let P(φ,D) denote the ground-truth performance (e.g., Dice score) after fine-tuning. Our ob- jective is to derive a training-free transferability score T(φ,D) that serves as a reliable proxy for P. Ideally, the score should preserve the performance rank- ing across the model zoo, satisfying the monotonicity condition for any pair of models φ i ,φ j ∈M: T(φ i ,D) > T(φ j ,D) ⇐⇒ P(φ i ,D) > P(φ j ,D)(1) While existing metrics typically rely on distributional statistics and neglect the complex geometry of SSL manifolds required for dense prediction, we argue that for segmentation tasks, topological tractability, i.e., the preservation of struc- tural connectivity or boundary separation, is a rather reliable predictor of trans- ferability than mere statistical overlap. We propose a topology-driven framework that assesses feature-label alignment at three levels: (1) Global Representa- tion Topology Divergence (GRTD), which quantifies the structural dis- crepancy between feature-induced and label-induced Minimum Spanning Trees (MST); (2) Local Boundary-Aware Topological Consistency (LBTC), which specifically examines the feature space distinctness at anatomical bound- aries, ensuring that the pre-trained features remain separable in these critical transition zones where segmentation failures most frequently occur; and (3) 4J. Tang, S. Zhang et al. #2 Local Local Boundary-Aware Topological Consistency(LBTC) #1 Global Global Representation Topology Divergence (GRTD) Decoder (initialized) Boundary Detection MST Weights Discrepancy 푻 푮푹푻푫 Transfer 푻 푳푩푻푪 → 1 Transfer Topology Comparison K Pretrained Foundation SSL Models Feature MST Graph Semantic Label MST Graph Minimum Spanning Tree (MST) Extracted Patch MST Encoder (frozen) Input Label Downstream Dataset Decoder (initialized) Stratified Sampling Features 푻 푮푹푻푫 =− ퟏ 푫 & 풅&ퟏ &흎 풊풋 풇풆풂풕 −&흎 풊풋 풔풆풎 푻 푳푩푻푪 =ퟏ− ퟏ 푵 & 풌&ퟏ 푵 흆 풌 (흉 풌 ) N : sampled boundary patch number 흎 풊풋 풇풆풂풕 흎 풊풋 풇풆풂풕 흎 풊풋 풔풆풎 =ퟎ 흎 풊풋 풔풆풎 ≠ퟎ Leakage Rate: 흆 풌 (흉 풌 )= 푵풖풎풃풆풓 풆풅품풆 풄풓풐풔 푵풖풎풃풆풓 풆풅품풆 풂풍 푫 : number of decoder stages : cross boundary edge connecting different classes : normal edge class A feature : boundary class B feature Task Complexity: 흒=풍풐품(∁) Rank models by 푺 흋 Gating Factor: 휶=흈(휸/흒 +휷) 1 2 3 4 5 6 7 푵(·) : Min–Max normalization over candidate models (per dataset) Few Leakage Obvious Leakage · #3Task-Adaptive Topological Fusion 푺 흋 =휶%푵푻 푮푹푻푫 + (ퟏ−휶)%푵(푻 푳푩푻푪 ) Fig. 2: The overall framework. First, we extract multi-scale features via stratified sampling and construct topological graphs via MST. We then compute GRTD to quantify overall manifold alignment and LBTC to evaluate separability at crit- ical anatomical boundaries. These metrics are integrated via a task-complexity gating factor, producing a final score S φ that predicts fine-tuning performance without training. Task-Adaptive Aggregation strategy that dynamically weights global and local metrics based on the inherent topological complexity of the target anatomy. 2.2 Global Representation Topology Divergence (GRTD) To capture the complex, non-linear geometry without imposing parametric as- sumptions, we model the data distribution using the Minimum Spanning Tree (MST), which serves as a robust descriptor of the manifold’s 1-skeleton, enabling us to evaluate transferability through the lens of topological isomorphism: a highly transferable model should yield a feature topology that naturally aligns with the semantic hierarchy. Let X = (v i ,y i ) N i=1 be the sampled feature-label pairs. We construct two graphs to evaluate topological alignment, as shown in Figure 2. The first is the Native Feature Graph G feat , where edge weights represent Euclidean distances in the embedding space: w feat ij =∥v i −v j ∥ 2 . The second is the Semantic Label- Induced Graph G sem , which represents an ideal label-driven topology. It explicitly injects semantic constraints by forcing samples of the same class to perfectly cluster while preserving the original feature distances for inter-class pairs up to a maximum penalty: w sem ij = ( 0if y i = y j , min(∥v i −v j ∥ 2 ,λ) if y i ̸= y j . (2) Topology-Driven Transferability Estimation5 where λ penalizes inter-class transitions. The MSTs derived from these graphs, denoted T feat and T sem , represent the natural clustering tendency and the ideal semantic connectivity, respectively. We define the GRTD as the discrepancy between the total weights of these two topological structures. Aggregating across D decoder stages, the final score is formulated as: T GRTD =− 1 D D X d=1 X (i,j)∈T d feat w feat ij − X (i,j)∈T d sem w sem ij .(3) A higher T GRTD (closer to 0) indicates that the encoder’s native geometry nat- urally respects semantic boundaries. 2.3 Local Boundary-Aware Topological Consistency (LBTC) Although GRTD captures the global manifold topology, it is susceptible to the class imbalance typical of medical images, where low-frequency background re- gions dominate the feature space statistics. However, the success of segmentation often hinges on the model’s ability to preserve high-frequency details at critical anatomical boundaries. Hence, we propose the Local Boundary-Aware Topolog- ical Consistency (LBTC), shifting the focus from global isomorphism to local separability, explicitly scrutinizing whether the pre-trained features maintain distinct decision boundaries within ambiguous boundary patches. We first define the set of boundary anchors ∂Y via the morphological gradient of the ground truth masks. For each anchor c k ∈ ∂Y, we extract a local patch P k and construct a local graph G k . The Minimum Spanning Tree of this local neighborhood, T k , serves as a probe for local cluster purity. Ideally, T k should traverse all intra-class nodes before bridging the semantic gap. We quantify the Topological Leakage Rate ρ k , which measures the proportion of edges in the local MST that erroneously connect distinct semantic classes: ρ k (T k ) = 1 |V k |− 1 X (u,v)∈E(T k ) I(y u ̸= y v ),(4) where E(T k ) denotes the edge set of the local MST and I(·) is the indicator function. The final LBTC score is derived by aggregating these local inconsis- tencies over N sampled boundary regions, formulated as the complement of the expected leakage: T LBTC = 1− 1 N N X k=1 ρ k (T k ).(5) A score approaching 1 implies that the encoder preserves strict topological sep- aration even within the ambiguous transition zones of the manifold. 6J. Tang, S. Zhang et al. Table 1: Weighted Kendall’s τ for transferability estimation on the OpenMind benchmark. Method ID (Same Region)OOD (Different Region) Avg MSF ISL HNT TPCACDKITAll LogME [24] -0.524 -0.223 -0.315 -0.5450.575-0.649-0.280 LEEP [13] -0.220 -0.427 -0.492 -0.051-0.6210.225-0.264 GBC [14] -0.707 -0.601 -0.537 -0.256-0.325-0.621-0.508 CCFV [23] 0.277 0.503 0.869 0.6640.8170.1800.552 Ours0.705 0.814 0.905 0.6640.6710.5780.723 2.4 Task-Adaptive Topological Fusion Medical segmentation targets exhibit distinct topological regimes: multi-organ tasks demand global structural preservation (high structural complexity), while small lesion extraction relies on local boundary contrast (high boundary com- plexity). A static combination of global and local metrics fails to generalize across these heterogeneous distributions. To resolve this, we propose a dynamic fusion mechanism driven by the semantic cardinality of the target task. We define a task complexity prior κ = log(|C|), where |C| is the number of semantic classes. This prior modulates a gating factor α ∈ (0, 1) via a sigmoid function σ(·), shifting the focus between macro-structure and micro-details. Let N(·) denote the Min-Max normalization operator over the model zoo. The final transferability score S φ is computed as a convex combination of the global and local topological metrics: α = σ(γ· κ + β), S φ = α·N(T GRTD ) + (1− α)·N(T LBTC ),(6) where γ and β are scaling constants. This formulation ensures that for complex anatomical tasks (α → 1), the ranking prioritizes global layout isomorphism, whereas for focal pathologies (α → 0), it emphasizes the sharpness of local decision boundaries. 3 Experiment and Results 3.1 The Foundation Model Zoo and Downstream Tasks We utilize the OpenMind benchmark [20] model zoo, featuring ResEnc-L [11] backbones pre-trained on 114,000 unlabeled 3D brain MRIs. To rigorously evalu- ate transferability without prior label exposure, we select models trained via re- construction (MAE[6], SimMIM[3], ModelsGenesis[25], S3D [19]) and contrastive learning (VoCo[21], SimCLR[2], SwinUNETR[16]). We evaluate these models across diverse downstream tasks spanning various anatomies and modalities. Topology-Driven Transferability Estimation7 Fig. 3: Correlation between the fine-tuning performance and transferability met- rics using MSF as an example. The vertical axis represents the average Dice, while the horizontal axis represents the standardized transferability metric. We want to observe a positive relationship between higher performance and higher transferability estimations. In-Distribution (ID) Tasks: We select four datasets closely aligned with the pre-training domain (head-and-neck) to test intrinsic transferability: ISLES (ISL) [9] (stroke lesions on DWI/ADC), HNTS-MRG (HNT) [18] (head-and- neck tumors on MR), MS FLAIR (MSF) [12] (multiple-sclerosis lesions), and ToP-CoW (TPC) [22] (Circle of Willis vessels). Out-of-Distribution (OOD) Tasks: To assess generalization under distribution shifts, we include two OOD tasks: ACDC (ACD) [1] (cardiac structures on cine-MRI), and KiTS19 (KIT) [8] (kidney and tumors on CT). Notably, KIT represents a challenging cross-anatomy and cross-modality (MR→CT) transfer scenario with clear tumor boundaries [11]. 3.2 Implementation Details Considering the SSL foundation model just have well-trained encoders, we attach a randomly initialized nnU-Net [10] decoder to each candidate model, extracting multi-scale features via sliding-window inference. To ensure topological fidelity despite severe class imbalance, we employ a stratified sampling strategy that balances class-conditional foreground voxels against background and global dis- tributions. We differentiate the extraction depth to align with topological roles: GRTD is computed on the final k decoder layers to capture semantic layout, while Local LBTC is evaluated on the last k encoder layers to assess intrinsic boundary separability. The task-adaptive gating weights are estimated using a computationally efficient pilot set of K=10 cases, with final rankings quantified via the weighted Kendall correlation (τ ∗,w ) [17]. 3.3 Results and Analysis Overall Performance We report the weighted Kendall’s τ between the esti- mated transferability scores and the actual fine-tuning performance. As shown in Table 1 and Figure 3, traditional classification-based metrics (LogME, LEEP) 8J. Tang, S. Zhang et al. Table 2: Left: weighted Kendall’s τ ∗ w averaged over five random seeds for each decoder initialization scheme. Right: transferability time and fine-tuning time are averaged over five repeated timing trials; both are reported as the total wall- clock time for all 7 models. All times are in minutes. Decoder initialization robustness Dataset Kaiming Xavier Gaussian TPC0.664 0.6460.652 ACD0.671 0.6240.652 MSF0.705 0.7060.718 Avg.0.680 0.6590.674 Time comparison Dataset CCFV Ours Fine-tune TPC1050.5 10.07 3000+ ACD756.9 8.02 3000+ MSF22.3 2.89 3000+ Avg.609.9 6.99 3000+ exhibit severe negative correlations, confirming their inadequacy for dense pre- diction tasks. While the embedding-based CCFV shows positive correlations, it struggles on challenging out-of-distribution (OOD) transfers like KIT (τ = 0.180). In contrast, our method demonstrates robust and consistent performance across both in-distribution and OOD scenarios, achieving the highest average correlation of 0.723, validating that topological tractability is a superior proxy for segmentation transferability. Robustness to Decoder Initialization. Although the decoder is randomly initialized, our metric should primarily reflect the intrinsic quality of the pre- trained representation rather than stochasticity from initialization. As summa- rized in Table 2 (left), the weighted Kendall’s τ ∗ w remains consistently stable across Kaiming [7]/Xavier [5]/Gaussian initializations with small variance, sug- gesting that the topology-driven signal dominates over initialization noise. Computational Time Comparison. Table 2 (right) contrasts the practi- cal cost of model selection via training-free scoring versus exhaustive down- stream fine-tuning. Compared with CCFV, our approach substantially reduces the metric-computation overhead, enabling rapid screening over a model zoo. More importantly, it avoids the prohibitive end-to-end cost of repeating fine- tuning across multiple candidates, making transferability estimation a realistic substitute for brute-force selection in resource-constrained settings. Impact of Target Topology Table 3 reveals the complementary nature of our topological priors. While Global (GRTD) and Local (LBTC) metrics suffer significant performance drops on fragmented and structured targets respectively, our adaptive fusion effectively bridges this gap and achieves a robust average of 0.714, a substantial improvement over the single-stream baselines. 4 Conclusion In this paper, we present a novel topology-driven transferability estimation framework of medical foundation model for segmentation. By shifting the eval- uation paradigm from statistical overlap to manifold topology, our approach Topology-Driven Transferability Estimation9 Table 3: Ablation study on topological priors of different datasets. We stratify tasks into Fragmented (ISL, MSF) and Structured (ACD, TPC) targets. Our Adaptive fusion achieves robust performance across all regimes. MethodISL MSFACD TPCAvg. Global (GRTD)0.122 0.3330.624 0.6640.436 Local (LBTC) 0.814 0.7050.319 0.0620.475 Adaptive (Ours)0.814 0.7050.671 0.6640.714 uniquely captures both global structural isomorphism and local boundary sep- arability. Extensive validation on the OpenMind benchmark demonstrates that our task-adaptive method significantly outperforms existing baselines across di- verse in-distribution and out-of-distribution tasks. By providing a robust, training- free proxy for model selection, our framework eliminates the prohibitive com- putational costs of exhaustive fine-tuning, paving the way for the efficient and scalable clinical deployment of medical foundation models. References 1. Bernard, O., Lalande, A., Zotti, C., Cervénanský, J., et al.: Deep learning tech- niques for automatic mri cardiac multi-structures segmentation and diagnosis: Is the problem solved? IEEE Transactions on Medical Imaging (2018) 2. Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for con- trastive learning of visual representations. In: International conference on machine learning. p. 1597–1607. PmLR (2020) 3. Chen, Z., Agarwal, D., Aggarwal, K., Safta, W., Balan, M.M., Brown, K.: Masked image modeling advances 3d medical image analysis. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. p. 1970– 1980 (2023) 4. Ding, N., Chen, X., Levinboim, T., Changpinyo, S., Soricut, R.: Pactran: Pac- bayesian metrics for estimating the transferability of pretrained models to classifi- cation tasks. In: European Conference on Computer Vision. p. 252–268. Springer (2022) 5. Glorot, X., Bengio, Y.: Understanding the difficulty of training deep feedforward neural networks. In: Proceedings of the thirteenth international conference on ar- tificial intelligence and statistics. p. 249–256. JMLR Workshop and Conference Proceedings (2010) 6. He, K., Chen, X., Xie, S., Li, Y., Dollár, P., Girshick, R.: Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 16000–16009 (June 2022) 7. He, K., Zhang, X., Ren, S., Sun, J.: Delving deep into rectifiers: Surpassing human- level performance on imagenet classification. In: Proceedings of the IEEE interna- tional conference on computer vision. p. 1026–1034 (2015) 8. Heller, N., Isensee, F., Maier-Hein, K.H., Hou, X., Xie, C., Li, F., Nan, Y., Mu, G., Lin, Z., Han, M., et al.: The state of the art in kidney and kidney tumor 10J. Tang, S. Zhang et al. segmentation in contrast-enhanced ct imaging: Results of the kits19 challenge. Medical image analysis 67, 101821 (2021) 9. Hernandez Petzsche, M.R., De La Rosa, E., Hanning, U., Wiest, R., Valenzuela, W., Reyes, M., Meyer, M., Liew, S.L., Kofler, F., Ezhov, I., et al.: Isles 2022: A multi-center magnetic resonance imaging stroke lesion segmentation dataset. Scientific data 9(1), 762 (2022) 10. Isensee, F., Jaeger, P.F., Kohl, S.A., Petersen, J., Maier-Hein, K.H.: nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods 18(2), 203–211 (2021) 11. Isensee, F., Wald, T., Ulrich, C., Baumgartner, M., Roy, S., Maier-Hein, K., Jaeger, P.F.: nnu-net revisited: A call for rigorous validation in 3d medical image segmen- tation. In: International Conference on Medical Image Computing and Computer- Assisted Intervention. p. 488–498. Springer (2024) 12. Muslim, A.M., Mashohor, S., Al Gawwam, G., Mahmud, R., Hanafi, M.B., Al- nuaimi, O., Josephine, R., Almutairi, A.D.: Brain mri dataset of multiple sclerosis with consensus manual lesion segmentation and patient meta information. Data in Brief 42, 108139 (2022). https://doi.org/10.1016/j.dib.2022.108139, https: //w.sciencedirect.com/science/article/pii/S235234092200347X 13. Nguyen, C., Hassner, T., Seeger, M., Archambeau, C.: Leep: A new measure to evaluate transferability of learned representations. In: International Conference on Machine Learning. p. 7294–7305. PMLR (2020) 14. Pándy, M., Agostinelli, A., Uijlings, J., Ferrari, V., Mensink, T.: Transferability es- timation using bhattacharyya class separability. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 9172–9182 (2022) 15. Tan, Y., Li, Y., Huang, S.L.: Otce: A transferability metric for cross-domain cross- task representations. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 15779–15788 (2021) 16. Tang, Y., Yang, D., Li, W., Roth, H.R., Landman, B., Xu, D., Nath, V., Hatamizadeh, A.: Self-supervised pre-training of swin transformers for 3d medi- cal image analysis. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 20730–20740 (June 2022) 17. Vigna, S.: A weighted correlation index for rankings with ties. In: Proceedings of the 24th international conference on World Wide Web. p. 1166–1176 (2015) 18. Wahid, K., Dede, C., Naser, M., Fuller, C.: Training dataset for hntsmrg 2024 challenge. https://zenodo.org/doi/10.5281/zenodo.11199559 (2024) 19. Wald, T., Ulrich, C., Lukyanenko, S., Goncharov, A., Paderno, A., Miller, M., Maerkisch, L., Jaeger, P., Maier-Hein, K.: Revisiting mae pre-training for 3d med- ical image segmentation. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 5186–5196 (2025) 20. Wald, T., Ulrich, C., Suprijadi, J., Ziegler, S., Nohel, M., Peretzke, R., Kohler, G., Maier-Hein, K.: An openmind for 3d medical vision self-supervised learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 23839–23879 (2025) 21. Wu, L., Zhuang, J., Chen, H.: Voco: A simple-yet-effective volume contrastive learn- ing framework for 3d medical image analysis. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 22873– 22882 (June 2024) 22. Yang, K., Musio, F., Ma, Y., Juchler, N.: Topcow: Benchmarking topology-aware anatomical segmentation of the circle of willis (cow) for cta and mra. Journal of Medical Imaging Technology 34, 123–135 (2024) Topology-Driven Transferability Estimation11 23. Yang, Y., Wei, M., He, J., Yang, J., Ye, J., Gu, Y.: Pick the best pre-trained model: Towards transferability estimation for medical image segmentation. In: In- ternational Conference on Medical Image Computing and Computer-Assisted In- tervention. p. 674–683. Springer (2023) 24. You, K., Liu, Y., Wang, J., Long, M.: Logme: Practical assessment of pre-trained models for transfer learning. In: International Conference on Machine Learning. p. 12133–12143. PMLR (2021) 25. Zhou, Z., Sodha, V., Pang, J., Gotway, M.B., Liang, J.: Models genesis. Medical image analysis 67, 101840 (2021)