Paper deep dive
Robustifying and Selecting Cohort-Appropriate Prognostic Models under Distributional Shifts
Dimitris Bertsimas, Carol Gao, Angelos G. Koulouras, Georgios Antonios Margonis
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/27/2026, 12:21:50 AM
Summary
The paper addresses the challenge of prognostic model transportability under distributional shifts (covariate and concept shifts). It proposes two strategies: 1) A model-developer perspective using meta-analysis-informed importance weighting to create 'best-on-average' models, and 2) An end-user perspective using cohort outcome similarity (Euclidean distance of means) to select the 'best-for-their-cohort' model. Testing on six surgical cohorts for colorectal liver metastases (CRLM), the authors demonstrate that higher KL divergence between cohorts correlates with higher Integrated Calibration Index (ICI) and that meta-analysis-informed weighting improves calibration and clinical utility (DCA).
Entities (10)
Relation Signals (4)
Integrated Calibration Index → assesses → Calibration
confidence 100% · calibration was assessed by the Integrated Calibration Index (ICI)
Meta-analysis-informed weighting → improves → Calibration
confidence 95% · Meta-analysis-informed weighting improved calibration in most settings
Cox Proportional Hazards → isevaluatedby → Harrell's Concordance Index
confidence 90% · Discrimination was evaluated using Harrell's concordance index.
Kullback-Leibler Divergence → quantifies → Cohort Similarity
confidence 90% · Cohort similarity was quantified using Kullback–Leibler (KL) divergence
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:External validation is widely regarded as the gold standard for prognostic model evaluation. In this study, we challenge the assumption that successful external calibration guarantees model generalizability and propose two complementary strategies to improve transportability of prognostic models across cohorts. Using six real-world surgical cohorts from tertiary academic centers, we tested whether successful external calibration depends largely on similarity in covariates and outcomes between training and validation cohorts, quantified using Kullback-Leibler (KL) divergence, with calibration assessed by the Integrated Calibration Index (ICI). From the model-developer's perspective, we trained the "best-on-average" prognostic model by tuning toward a meta-analysis-derived covariate and outcome distribution as an approximation of the broader target population. From the end-user perspective, we proposed a simple measure for cohort outcome similarity to identify, among published models, the one most suitable for a given target cohort in terms of both calibration and clinical utility. External calibration worsened as distributional mismatch increased. Higher KL divergence was associated with higher ICI in both surgery-alone (Spearman $\rho=0.614$, $p=0.004$) and surgery + adjuvant chemotherapy cohorts (Spearman $\rho=0.738$, $p<0.001$). Meta-analysis-informed weighting improved calibration in most settings without materially affecting discrimination, with the clearest benefit when evaluated on the aggregated external population ($p=0.037$). Models developed in more similar cohorts achieved lower ICI in surgery-alone (Spearman $\rho=0.803$, $p<0.001$) and surgery + adjuvant chemotherapy cohorts (Spearman $\rho=0.737$, $p<0.001$), and provided greater clinical utility on DCA.
Tags
Links
- Source: https://arxiv.org/abs/2604.16537v1
- Canonical: https://arxiv.org/abs/2604.16537v1
Trouble viewing inline? Open PDF directly →
Full Text
62,403 characters extracted from source content.
Expand or collapse full text
Robustifying and Selecting Cohort-Appropriate Prognostic Models under Distributional Shifts Dimitris Bertsimas 1,2 , Carol Gao 1 , Angelos Koulouras 1 , Georgios Antonios Margonis 3,4* 1 Operations Research Center, Massachusetts Institute of Technology, Cambridge, MA, 02139, USA. 2 Sloan School of Management, Massachusetts Institute of Technology, Cambridge, MA, 02139, USA. 3 Department of Surgery, Memorial Sloan Kettering Cancer Center, New York, NY, 10065, USA. 4 Charit ́e – Universit ̈atsmedizin Berlin, Berlin, 10117, Germany. *Corresponding author(s). E-mail(s): margonig@mskcc.org; Abstract Introduction: External validation is widely regarded as the gold standard for prognostic model evaluation. In this study, we challenge the assumption that successful external validation of calibration guarantees model generalizabil- ity and propose two complementary strategies to improve the transportability of prognostic models across cohorts. Methods: Using six real-world surgical cohorts from tertiary academic centers, we first tested the hypothesis that successful external calibration does not neces- sarily guarantee model generalizability, but instead depends largely on similarity in covariates and outcomes between the development and validation cohorts. Cohort similarity was quantified using Kullback–Leibler (KL) divergence, and calibra- tion was assessed using the Integrated Calibration Index (ICI). Second, from the model-developer’s perspective, we aimed to train the “best-on-average” prognos- tic model by tuning models toward a meta-analysis-derived covariate and outcome distribution, treating this distribution as an approximation of the broader target population, and then evaluating whether this improved calibration. Third, from the end-user perspective, we proposed a simple measure of cohort outcome similarity to identify, among available published models, the prognostic model most suitable 1 arXiv:2604.16537v1 [stat.ME] 16 Apr 2026 for a given target cohort with respect to both calibration and clinical utility, as assessed by Decision Curve Analysis (DCA). Results: External calibration worsened as distributional mismatch increased. Higher KL divergence was associated with higher ICI in both surgery-alone (Spearman ρ = 0.614,p = 0.004) and surgery + adjuvant chemotherapy cohorts (Spearman ρ = 0.738,p < 0.001). Meta-analysis-informed weighting improved calibration in most settings without materially affecting discrimination, with the clearest benefit observed when performance was evaluated against the aggre- gated external population (p=0.037), consistent with the goal of approximating a broader target population. From the end-user perspective, models developed in cohorts more similar to the target cohort achieved lower ICI in both surgery alone (Spearman ρ = 0.803,p < 0.001) and surgery + adjuvant chemotherapy cohorts (Spearman ρ = 0.737,p < 0.001), and also provided greater clinical utility on DCA. Conclusion: Successful external calibration depends strongly on the similar- ity between development and validation cohorts and therefore does not, by itself, guarantee model generalizability to additional settings. We propose one strategy to train the “best-on-average” prognostic model and another to help users select the “best-for-their-cohort” prognostic model. 1 Introduction Accurate prognostication is one of the pillars of clinical practice. Prognostic model out- puts are commonly used to inform patients about what to expect after treatment, to guide physicians in deciding whether treatment escalation is warranted (for example, surgery plus chemotherapy vs surgery alone), and to support shared decision-making between physicians and patients. Accordingly, evaluating the performance of prognos- tic models is a central task in clinical outcome research. The current standard for prognostic model evaluation includes two complementary dimensions. The first is discrimination, that is, the ability of the model to correctly rank patients according to their risk. In survival analysis, this is typically assessed using metrics such as the concordance index (C-index) or time-dependent area under the curve (AUC). The second, and arguably even more clinically relevant, dimension is calibration: whether the absolute risk estimates assigned to individual patients, such as the predicted probability of 5-year survival, are accurate.[1] The most informative assessment of both discrimination and calibration occurs in an independent external cohort from another center, where good performance is generally interpreted as evi- dence that the model can generalize beyond the cohort used to develop it. Nonetheless, calibration is not purely an intrinsic property of the model. Rather, it depends in part on the cohort in which the model is evaluated. In clinical outcome research, external validation is often limited to one or, at most, two outside cohorts. However, successful validation in such a small number of datasets may be mislead- ing: performance may appear favorable simply because the external cohort happens to share a similar underlying risk distribution with the training cohort, rather than because the model is truly generalizable. In practical terms, this challenges the credi- bility of many published prognostic tools, particularly those supported by only one or two external validations. 2 This problem is especially important because, in most areas of medicine, the fac- tors that determine patient outcomes are only incompletely understood. As a result, the risk distributions of different cohorts are not known a priori. Consequently, when a model performs well in an outside cohort, it often remains unclear whether this reflects genuine robustness or merely favorable alignment between the derivation and valida- tion populations. This raises a fundamental question: is it truly feasible to develop a single prognostic model that generalizes well across all settings, or should model selec- tion be tailored to the specific characteristics and needs of the target cohort? In this study, we address the problem both from the perspective of the model developer, who aims to build a robust and generalizable model across various cohorts, and from the perspective of the end user, who wishes to identify the most appropriate available model for their specific real-world cohort, even when that model is not the best on average. Our contributions are as follows: 1. We demonstrate empirically on 6 real-world cohorts that successful external validation does not necessarily imply generalization and provide a theoretical justification. 2. We develop a prognostic weighting approach that is more robust to distribution shifts. Specifically, we show across a range of external cohorts that the mod- els weighted towards a meta-analysis distribution, on average, perform better in calibration than competing models. 3. We provide a structured framework for selecting, among the models available in the literature, the one that best fits the prognostic needs of a given external cohort. 2 Methods The Methods are organized into five parts. First, we describe the study cohorts used for empirical evaluation. Second, we describe and formulate the problem of distribution shifts we attempt to solve. Third, from the perspective of the model developer, we present a framework to train prognostic models that are more robust to distributional shifts across cohorts. Fourth, from the perspective of the end user, we describe a framework for selecting the prognostic model that best fits a specific target cohort. Lastly, we summarize the statistical analyses used to evaluate model performance and clinical utility. 2.1 Study cohorts As noted above, the aim of this study was not to address a specific clinical problem, such as prognostication in a single disease setting, but rather to propose: (1) a struc- tured approach for training the best average prognostic model from the perspective of the model developer, and (2) a framework for selecting, from the perspective of the end user, the prognostic model from the literature that best fits the needs of a given target cohort. 3 To demonstrate the proposed methodology, we performed empirical analyses using high-quality data from tertiary institutions across distant geographic regions, includ- ing centers in the United States, Europe, and Japan. All cohorts consisted of patients undergoing surgery for colorectal liver metastases (CRLM), thereby providing a clin- ically coherent setting in which to evaluate model transportability across different populations. Specifically, we included patients who underwent surgical treatment for colorectal liver metastases at six academic institutions: Johns Hopkins Hospital (JHH, USA), Cleveland Clinic Foundation (CCF, USA), Erasmus University of Rotterdam (UOR, the Netherlands), Yokohama City University (YCU, Japan), CHU Clermont-Ferrand (CHU, France), and the University of Graz (Austria). Eligible patients were aged 18 years or older and had undergone complete (R0 or R1) resection of both the primary colorectal tumor and all colorectal liver metastases. 2.2 Problem formulation: distribution shifts in real-world medical data Consider a dataset of CRLM patients, where each patient i is characterized by a set of prognostic covariates x i , a treatment status t i ∈ 0, 1, where t i = 1 denotes receipt of adjuvant chemotherapy and a t i = 0 denotes no adjuvant chemotherapy and a recurrence outcome y i ∈0, 1. A central challenge in prognostic modeling is that a model trained in one real-world cohort may later be applied to external cohorts whose underlying data distributions differ. In this study, we consider two types of distribution shift that may affect model transportability across cohorts. The first is covariate shift, which refers to differences in the distribution of observed prognostic factors across cohorts. These differences may arise from variation in patient selection criteria and differences in the “pool” of candidate patients across countries and centers. The second is concept shift, where the conditional outcome distribution differs across cohorts: p train (y | x)̸= p test (y | x). Concept shifts may occur due to differences in standards of care between centers or because of varying prevalence of unobserved prognostic factors such as genomic alterations, radiomics features, or other latent clinical variables not captured in the dataset. Formally, given n train patients each with covariates x i and outcome y i in a training cohort, model training involves minimizing the loss function L train = Z ℓ(f (x),y)p train (x,y)dxdy.(1) When the joint distribution in the test cohort is similar to that of the training cohort, that is, p train (x,y)≈ p test (x,y), the test loss is expected to be similar: L train = Z ℓ(f (x),y)p train (x,y)dxdy ≈ Z ℓ(f (x),y)p test (x,y)dxdy =L test .(2) 4 Accordingly, the methodological problem addressed in this study is how to improve prognostic modeling when the training and target cohorts differ because of covariate shift, concept shift, or both. In the following sections, we address this problem from two complementary perspectives: first, from the perspective of the model developer, who seeks to train a model that is more robust across heterogeneous external cohorts, and second, from the perspective of the end user, who seeks to identify the model most appropriate for a specific target cohort. 2.3 Part I. Proposed solution from the model-developer perspective: training a more generalizable prognostic model Having formalized the problem of distributional shift across cohorts, we next address it from the perspective of the model developer. Specifically, our goal is to train a prog- nostic model that is less dependent on the idiosyncrasies of a single training cohort and instead better aligned with the broader population in which the model may even- tually be applied. To this end, we propose an importance-weighting framework based on internal prognostic risk distribution that shifts the effective training distribution toward a meta-analysis-derived target distribution. Ideally, if the target test distribution p test (x,y) is known, we can introduce sample weights w(x,y) = p test (x,y) p train (x,y) to recover the true test loss from training data, namely L w train = Z w(x,y)ℓ(f (x),y)p train (x,y)dxdy(3) = Z p test (x,y) p train (x,y) ℓ(f (x),y)p train (x,y)dxdy(4) = Z ℓ(f (x),y)p test (x,y)dxdy(5) =L test .(6) In practice, however, the target test distribution is unknown when the prognostic model is trained. Moreover, there is no single universal target distribution, because each cohort in which the model may eventually be used may have a different distri- bution. Therefore, from the perspective of model development, it is desirable to train a prognostic model that is more robust to distributional shifts. We aim to improve generalizability by adjusting the training process toward the distribution of the broader population of CRLM patients. To this end, we use summary statistics from a meta-analysis as a proxy for the general population distribution, under the rationale that a meta-analysis samples multiple studies across regions and is therefore more likely to reflect the broader distribution encountered in practice. 5 Empirically, in a training set, the training process is given by minimizing the empirical loss ˆ L train = 1 n train n train X i=1 w i ℓ(f (x i ),y i ),(7) where w i is the sample weight on observation i We first consider the scenario where the data only suffer from concept shift. We propose a method that computes the concept weight. w conc i = ˆ p meta (y i | x i ) ˆ p train (y i | x i ) ,(8) where ˆ p train (y i | x i ) can be directly estimated by fitting a survival model S i in the training cohort and making inference S i (x i ) for patient i. ˆ p meta (y i | x i ) is obtained from a simulated distribution with aggregated statistics from the meta-analysis, namely the mean and standard deviation of recurrence. This approach fully addresses distribution shift when the shift is purely due to differences in the conditional outcome distribution and there is no covariate shift. However, concept-based weighting alone is not generally sufficient to correct pure covariate shift, because in that setting the discrepancy arises through p(x i ) rather than p(y i | x i ). Still, because differences in covariate distributions often manifest through differences in predicted-risk strata, the proposed approach that relies on an internal risk prognosis may partially reduce some forms of covariate mismatch in practice. The model used to compute weights must be trained on the training cohort. Since the weights are applied during training of the final prognostic model in that same cohort, in order to recover the general distribution ˆ p meta (y i | x i ), it is necessary that ˆ p train (y i | x i ) used in the weights is an accurate estimate of the ˆ p train (y i | x i ) in the final model, as shown in equation (4). Let n train denote the number of observations in the training set and m the number of risk strata, such that s(i) = kif ℓ k < ˆ p train (y i | x i )≤ u k ∀k = 1, 2,...,m, where ℓ k ,u k are the lower and upper thresholds of stratum k respectively. Then, the importance of stratum k in the training cohort is given by ˆ q k train = n k n train = number of patients in bucket k total number of patients . Specifically, ˆ q k = P( ˆ p train (y | x)∈ bucket k). We then generate a simulated cohort with synthetic recurrence risk as a nor- mally distributed random variable R∼N (μ y ,σ y ) where μ y ,σ y are the reported mean and standard deviation of recurrence in the meta-analysis. We then compute the 6 importance of each stratum k = 1, 2,...,m as ˆ q k meta = P(R≤ u k )− P(R≤ ℓ k ) k = 1, 2, 3,...,m(9) The weights are then computed for each stratum k as w conc k = ˆ q k meta ˆ q k train ,(10) or equivalently for each patient i in the training set as w conc i = ˆ q s(i) meta ˆ q s(i) train = importance of its bucket in general population importance of its bucket in training data . Next, observe that by conditional probability, ˆ p meta (x i ,y i ) ˆ p train (x i ,y i ) = ˆ p meta (y i | x i ) ˆ p train (y i | x i ) |z concept weight · ˆ p meta (x i ) ˆ p train (x i ) |z covariate weight To adjust when there exist both concept shifts and covariate shifts, we introduce the joint weight given by w joint i = w conc i ˆ p meta (x i ) ˆ p train (x i ) ,(11) where ˆp meta (x i ) ˆp train (x i ) accounts for the covariate shift. Assumption 1 The covariance between each pair of covariates is constant across cohorts. Let Σ train ∈ R d×d denote the covariance matrix of covariates X train in the train- ing data. By Assumption 1, we obtain the covariance matrix of the meta-analysis population Σ meta by the following: Σ meta j,j = (σ meta j ) 2 , Σ meta j,l = Σ train j,l for j ̸= l, where (σ meta j ) 2 is the reported variance of feature j in the meta-analysis. We can then simulate a random vector X meta ∼ N (μ meta , Σ meta ) whereμ meta are the reported means of the features in the meta-analysis. After obtaining the simulated general pop- ulation feature space, we then follow the literature and use Kernel Density Estimation (KDE) with a Gaussian kernel to obtain the probability density function of X meta 7 and X train [2]. Specifically, ˆ p train (x i ) = 1 n train n train X n=1 1 b d K( x i − x n b ), ˆ p meta (x i ) = 1 n meta n meta X n=1 1 b d K( x i − x n b ), where b is a chosen bandwidth and K is the Multivariate Gaussian kernel. Algorithm 1 Weighted Prognostic Model Adjusting for Concept Shift Input: Training dataset (X train , t train , y train ), reported outcome mean and vari- ance in a meta-analysis μ y ,σ 2 y . Output: Predicted 5-year recurrence risks on testing cohort patients ˆy test Step 1: Risk prediction and stratification on training cohort Obtain ˆ p train (y i |x i ) for each training patient i. Stratify training patients into m strata where s(i) = k if ℓ k ≤ ˆ p train (y i | x i ) ≤ u k for k = 1, 2,...,m. Obtain stratum density ˆ q k train = n k n train . Step 2: Risk distribution approximation through meta-analysis Simulate recurrence risk in the general patient population R ∼ N (μ y ,σ y ) where μ y ,σ y are the reported recurrence mean and standard deviation in meta-analysis. Obtain stratum density ˆ q k meta = P(r ≤ u k )− P(r ≤ ℓ k ). Step 3: Concept weight computation Compute sample weight for each training patient i to be w i = ˆq s(i) meta ˆq s(i) train where s(i) is the corresponding stratum of patient i. Step 4: Train weighted prognostic model 2.4 Part I. Proposed solution from the end-user perspective: selecting the best model for a target cohort Even if a model is optimized to perform well on average across heterogeneous popu- lations, it will not necessarily be the best-performing model for every specific cohort. From the perspective of the end user, the relevant question is therefore different: rather than asking which model performs best on average, the goal is to identify which avail- able model is most likely to be well calibrated and clinically useful in the user’s own patient population. We therefore propose a cohort-specific model-selection strategy based on similarity between the target cohort and the original development cohorts of candidate models. Specifically, we hypothesized that models developed in cohorts whose outcome pro- file more closely resembles that of the target cohort are more likely to show better calibration when externally applied to that cohort. Let μ test y and μ train y , denote the mean Kaplan Meier estimates at the time point of interest in the target cohort and a model-development cohort, respectively. We 8 quantified the distance between the two cohorts using the Euclidean norm of the means, given by dist train,test =∥μ train y ,−μ test y ∥ 2 . Smaller values indicate greater similarity between the development and target cohorts. 2.5 Statistical analysis To quantify the training and testing distribution similarity, we compute the Kull- back–Leibler (KL) divergence between the two cohorts. To capture both concept shift and covariate shift, we define z i = [x i ,y i ] as the concatenated vector of the covariates and outcome. The probability density function is obtained similarly using KDE with a Gaussian kernel. Namely, for a pair of training and testing sets, ˆ q train (z i ) = 1 n train n train X n=1 1 b d K( z i − z n h ), ˆ q test (z i ) = 1 n test n test X n=1 1 b d K( z i − z n h ). Then, the KL divergence between the two cohorts is given by KL( ˆ q train | ˆ q test ) = E z∼ˆq train log ˆ p train (z) ˆ p test (z) . A smaller KL divergence indicates that the training and testing cohort are more similar in distribution. Three prognostic model classes were evaluated: Cox Proportional Hazards, Ran- dom Survival Forests (RSF), representing the ensemble-tree class, and Optimal Survival Trees (OST), representing the single-tree class [3, 4]. Cox models were imple- mented in R using the survival package. RSF models were implemented in Python using the scikit-survival package, and OST models were implemented in R using the Interpretable AI modules (version 1.10.2). Model performance was assessed along two complementary dimensions: discrimina- tion and calibration. Discrimination was evaluated using Harrell’s concordance index. Calibration was assessed using the Integrated Calibration Index (ICI), which sum- marizes the average absolute difference between predicted and observed recurrence probabilities across the risk spectrum and provides a single interpretable measure that enables direct comparison across models. Formally, ICI on a test set is given by ICI train,test = P n test i=1 | ˆ p train (y i | x i )− P(y i = 1| ˆ p train (y i | x i ))| n test ∈ [0, 1].(12) 9 To formally evaluate the correlation between distribution similarity and calibra- tion, we compute the Spearman correlation coefficient, given by ρ = 1− 6 X train X test̸=train (R[ICI train,test ]− R[KL( ˆ q train | ˆ q test )]) 2 n(n 2 − 1) ∈ [−1, 1],(13) where R[ICI train,test ] is the ranking of ICI train,test and R[KL( ˆ q train | ˆ q test )] is the ranking of KL( ˆ q train | ˆ q test ) among all pairs of cohorts considered. A large positive ρ would indicate that calibration and cohort distribution similarity are highly positively correlated. Next, we want to formally evaluate whether weighting on average improves calibra- tion. To this end, we compute the Wilcoxon signed-rank test that tests the hypothesis that whether the median ICI among the weighted models is statistically different from the median ICI among the unweighted models [5]. Lastly, to assess potential clinical utility in guiding treatment decisions, we also performed decision curve analysis (DCA). In the empirical evaluation of our frame- work, the treatment of interest was adjuvant chemotherapy. Therefore, DCA was restricted to patients treated with surgery alone, such that observed recurrence served as a pragmatic approximation of baseline risk. Net benefit was calculated across a range of threshold probabilities and compared with the default “treat all” and “treat none” strategies. For figure clarity, we emphasized the threshold range from 0.10 to 0.70, which broadly reflects clinically relevant risk-benefit tradeoffs. For each model, we recorded the maximum net benefit and the range of thresholds over which the model outperformed both default strategies. All analyses were performed using Python and R. 3 Results The baseline characteristics, sample sizes, and long-term outcomes of the six real-world CRLM cohorts are summarized in Table 1. The 5-year RFS reported correspond to the mean 5- year Kaplan-Meier estimates. Marked heterogeneity was observed across cohorts. Among surgery-alone cohorts, the mean 5-year mean recurrence free survival ranged from 22% to 62%; among surgery + adjuvant chemotherapy cohorts, it ranged from 10% to 60%. Covariates used to train the prognostic models were selected on the basis of relevant surgical literature and included metastatic lymph node status, number of liver metastases, maximum metastasis size, laterality of liver metastases, carcinoembryonic antigen (CEA), extrahepatic disease, and resection margin status. 3.1 Distributional similarity and external calibration To first establish the practical importance of distributional shift, we examined whether similarity between training and target cohorts was associated with external calibration. For each surgery-alone training cohort, we trained an unweighted prognostic model and evaluated it across each of the other available surgery alone cohorts. We then compared pairwise cohort similarity, quantified by KL divergence — which reflects combined 10 Cohort # of Patients 5-Year RFS CEA (ng/ml) Extrahepatic Disease Positive Margin Maximum Tumor Size Number of Metastases Metastatic Lymph Nodes Right-sided tumor Left-sided tumor JHH1570.21640.6110 [6.37%]20 [12.74%]3.552.85100 [63.69%]46 [29.30%]111 [70.70%] CCF2250.54463.1731 [13.78%]53 [23.56%]3.982.06152 [67.56%]70 [31.11%]101 [44.89%] YCU1170.24464.3625 [21.37%]41 [35.04%]3.555.0779 [67.52%]31 [26.50%]48 [41.03%] Graz1650.61632.7223 [13.94%]0 [0.00%]4.961.09126 [76.36%]73 [44.24%]92 [55.76%] UOR10580.24683.0587 [8.24%]166 [15.70%]3.622.87640 [60.51%]196 [18.53%]840 [79.40%] JHH3880.29551.7535 [9.02%]58 [14.95%]3.0252.814272 [70.10%]127 [32.73%]261 [67.27%] CCF3080.604115.0839 [12.66%]87 [28.25%]3.6182.405209 [67.86%]97 [31.49%]111 [36.04%] YCU700.09932.3312 [17.14%]15 [21.43%]3.2835.54347 [67.14%]16 [22.86%]31 [44.29%] Graz1090.5331.3620 [18.35%]0 [0.00%]3.9861.08758 [53.21%]21 [19.27%]88 [80.73%] CHU6480.49277.846 [7.10%]31 [4.78%]4.092.264388 [59.88%]160 [24.69%]395 [60.96%] Surgery alone Surgery + chemotherapy 5-year RFS is the mean 5-year Kaplan-Meier estimate. CEA, maximum tumor size, and number of metastasers are reported as means. Remaining features are reported as counts with percentages in square brackets. Maximum tumor size is reported in centimeters (cm). Table 1: Characteristics of the CRLM cohorts mismatch arising from both covariate shift and concept shift — with calibration in the target cohort, quantified by ICI. The same analysis was repeated for surgery + chemotherapy cohorts. Across the analyses of patients who underwent surgery alone, models consistently achieved their best external calibration in target cohorts that were most similar in dis- tribution to the corresponding training cohort. In other words, the smaller the distance between cohorts, the greater the distributional similarity and the lower the ICI, indi- cating better calibration (Figure 1). The same qualitative pattern was observed in the analyses on patients who received adjuvant chemotherapy (Figure 2). Overall, these results suggest that similarity in the joint distribution of covariates and outcomes is associated with better external calibration. The magnitude of this effect was substantial. For a given model trained on a single cohort, ICI ranged from values below 0.10 in distributionally similar target cohorts to values approaching 0.30 or higher in dissimilar target cohorts. For example, the model trained on surgery-alone patients from JHH, a cohort with high recurrence risk, achieved ICI values below 0.10 when applied to the UOR and YCU cohorts, which had similarly high recurrence risk, but its ICI increased to nearly 0.35 when applied to Graz, a substantially lower-risk cohort. Had external validation been performed only in the more similar cohorts, the model might have appeared broadly general- izable despite poor performance in distributionally distinct settings. Moreover, the surgery-alone cohorts and surgery + chemotherapy cohorts yield statistically signif- icant positive correlations of 0.614 (with p-value = 0.004) and 0.738 (with p-value < 0.001) respectively between KL divergence and ICI. Appendix Figure 14 plots these correlations in a scatterplot. These results further formally test the hypothesis that cohorts that are similar in distributions tend to have better calibration. These findings, which underscore the limited generalizability of prognostic models in the presence of covariate and/or concept shift, motivate the first methodological question addressed in this study: whether concept weighting or joint weighting can be used to train a prognostic model that generalizes better across heterogeneous external cohorts. 11 A. Trained on JHHB. Trained on CCFC. Trained on YCU D. Trained on GrazE. Trained on UOR Fig. 1: Pairwise KL divergence between training and test cohorts and the correspond- ing ICI for models trained in each surgery-alone cohort. ICIs are averaged across three models — Cox, Optimal Survival Trees (OST), and Random Survival Forest (RSF) — and the error bars represent the minimum and maximum ICIs across the three mod- els. KL divergence is computed deterministically with fixed bandwidth. Each panel represents one training cohort, and the bars within that panel show results for all four test cohorts. Panels A–E correspond to models trained in JHH, CCF, YCU, Graz, and UOR, respectively. 3.2 Part I. Model-developer perspective: meta-analysis-informed weighting improves average external calibration We next evaluated whether the importance-weighting strategies described in Section 2.3 improved average external calibration across cohorts. We applied the weighting framework using statistics derived from the meta-analysis by Bosma and colleagues [6]. 1 For each training cohort, we estimated predicted 5-year recurrence risk and strat- ified patients into eight risk strata. We then trained unweighted models, models weighted by concept weights w conc , and models weighted by joint weights w joint , and evaluated each model separately in the remaining external cohorts. For each train- ing cohort, the reported ICI represents the average across the corresponding external testing cohorts. Among surgery-alone cohorts, either concept weighting or joint weighting improved average external calibration for 3 of the 5 training cohorts (Figure 3, Panel A). Among the surgery-plus-chemotherapy cohorts, concept weighting improved average 1 The individual studies included in the meta-analysis are [7–10]. When 5-year recurrence was not explicitly reported in [6], it was extracted from the individual studies. 12 A. Trained on JHHB. Trained on CCFC. Trained on YCU D. Trained on GrazE. Trained on CHU Fig. 2: Pairwise KL divergence between training and test cohorts and the correspond- ing test-set ICI for models trained in each surgery + adjuvant chemotherapy cohort. ICIs are averaged across three models — Cox, Optimal Survival Trees (OST), and Random Survival Forest (RSF) — and the error bars represent the minimum and maximum ICIs across the three models. KL divergence is computed deterministically with fixed bandwidth. Each panel represents one training cohort, and the bars within that panel show results for all four test cohorts. Panels A–E correspond to models trained in JHH, CCF, YCU, Graz, and CHU, respectively. external calibration for 4 of the 5 training cohorts (Figure 3, Panel B). The magni- tude of improvement varied across settings, ranging from approximately 5% in the Graz surgery-plus-chemotherapy cohort to more than 50% in the YCU surgery-plus- chemotherapy cohort. Concept weight tends to outperform joint weight, likely due to data limitations and covariate noise. Meta-analytic studies vary in their reporting practices — not all provide means and standard deviations for every model feature, and some report continuous means while others report binary counts relative to a threshold. This inconsistency produces noisy estimates of the proxy meta-analysis distribution. Our results suggest that correcting for concept shift alone is generally sufficient to improve model robustness, even when the covariate distribution of the general population is not directly estimable. We further compare calibration on an aggregated cohort excluding the training cohort between weighted and unweighted models. This yields a Wilcoxon stat of 7.0 with p-value of 0.037. This suggests that weighted models validated on the remaining cohorts aggregated have significantly lower median ICI than the unweighted models and therefore better calibration. Notably, the gains from weighting were largest when the baseline external cal- ibration of the unweighted model was poorest. This pattern is consistent with the 13 underlying rationale of the approach: when a training cohort is distributionally dis- tinct from the external cohorts in which the model is likely to be used, shifting the effective training distribution toward a broader meta-analysis-derived target can yield more substantial improvements in average transportability. By contrast, weighting had little effect on discrimination. As shown in Figure 4, Harrell’s C-index remained largely unchanged after applying either concept weights or joint weights. Taken together, these findings suggest that the proposed weight- ing framework primarily improves calibration, rather than rank ordering, across heterogeneous external cohorts. A. Surgery-aloneB. Surgery + chemotherapy Fig. 3: Average test-set ICI by weighting strategy. For each training cohort, the reported ICI represents the average across the corresponding external test cohorts and across the models (OST and RSF). Error bars represent the minimum and maximum model-specific average ICIs. Panel A: surgery-alone cohorts. Panel B: surgery-plus- chemotherapy cohorts. 3.3 Part I. End-user perspective: cohort-specific model selection improves calibration and clinical utility We next turned to the second perspective of the framework: the perspective of the end user, for whom the most broadly transportable model may not necessarily be the most appropriate model for a specific target cohort. We therefore examined whether simple cohort-level similarity metrics could guide selection of the model most likely to be well calibrated in a given external cohort. For each target cohort, we computed the distance dist train,test between that cohort and each candidate training cohort, using the cohort-level mean 5-year recurrence probability obtained from Kaplan-Meier estimates described in the Methods. We then compared this distance with the observed calibration of each externally applied model. Among surgery-alone cohorts, in all of the target cohorts, the model trained on the most similar cohort also yielded either the lowest or second lowest ICI in the tar- get cohort (Figure 5). In the surgery-alone UOR and JHH cohorts where the most similar cohort yields the second lowest ICI, the difference in both ICI and cohort 14 A. Surgery aloneB. Surgery + chemotherapy Fig. 4: Average test-set Harrell’s C-index by weighting strategy. For each training cohort, the reported C-index represents the average across the corresponding external test cohorts and across the models (OST and RSF). Error bars represent the minimum and maximum model-specific average ICIs. Panel A: surgery-alone cohorts. Panel B: surgery-plus-chemotherapy cohorts. similarity between the lowest and second lowest cohorts is marginal. Similar pat- terns were observed in the surgery-plus-chemotherapy cohorts (Figure 6). In addition, a formal examination of the correlation yields a correlation index of 0.803 between cohort similarity and ICI among the surgery-alone cohorts and 0.737 among the surgery-plus-chemotherapy cohorts, both with p-values < 0.001. This correlation is shown in Appendix Figure 15. These findings suggest that simple cohort-level outcome summaries may provide a useful and pragmatic guide for selecting among existing prognostic models when direct recalibration or local model development is not feasible. We then examined whether this improved cohort matching also translated into greater clinical utility. Decision curve analysis showed that models trained on cohorts with outcome profiles similar to the target cohort generally produced higher net benefit across clinically relevant threshold ranges than models trained on more dissimilar cohorts. For example, when patients from JHH who underwent surgery alone served as the target cohort, models trained on YCU and UOR—two cohorts with similarly high recurrence risk—outperformed both the “treat all” and “treat none” strategies over a meaningful range of thresholds, whereas models trained on the lower-risk CCF and Graz cohorts offered little or no clinical value beyond trivial strategies (Figure 7). A similar pattern was observed when patients from YCU who underwent surgery alone were used as the target cohort (Figure 8). We next examined whether the same pattern held for the two lower-risk surgery- alone cohorts, CCF and Graz. As shown in Figure 9, when CCF was used as the target cohort, the model trained on Graz patients who underwent surgery alone out- performed the other externally trained models, consistent with the fact that Graz was the most similar cohort to CCF in terms of risk of recurrence. However, despite this relative advantage, the corresponding DCA curve largely overlapped with the “treat all” strategy, indicating limited added clinical value. To better understand this finding, we examined the distribution of predicted risks (Figure 10). In contrast to the better-performing models in the JHH target 15 cohort—specifically, the models trained on YCU and UOR, which generated a broad range of predicted risks from approximately 0.3 to 0.9—the model trained on Graz pro- duced a much narrower range of predictions, approximately 0.2 to 0.45, when applied to the CCF cohort. This suggests that the model assigned relatively low recurrence risk to most patients in the cohort. Because CCF itself is a predominantly low-risk cohort, such predictions can still yield favorable calibration and low ICI, as shown in Panel B of Figure 5. However, a model that predicts a narrow low-risk range for nearly all patients, even if reasonably calibrated on average, offers limited ability to distinguish patients with meaningfully different prognoses and therefore limited clinical utility. This interpretation is further supported by the calibration plots in Figure 11, where we compare random survival forest models trained on the surgery-alone UOR and Graz cohorts. When validated on JHH patients who underwent surgery alone, the model trained on UOR patients—whose risk profile more closely resembles that of JHH—produced a calibration curve that closely followed the 45-degree line, indi- cating good agreement between predicted and observed risks. By contrast, the model trained on Graz patients underestimated risk in JHH, consistent with the lower-risk composition of the Graz cohort. Conversely, when validated on CCF patients who underwent surgery alone, the Graz-trained model produced a calibration curve close to the 45-degree line, whereas the UOR-trained model systematically overestimated risk because UOR consists of substantially higher-risk patients. Importantly, however, even in the setting where the Graz model was well calibrated in CCF, its predicted risks remained concentrated within a relatively narrow range. Thus, similarity between development and target cohorts may improve calibration, but clinical usefulness also depends on whether the model produces a sufficiently informative spread of predicted risks to support treatment decisions. Similar patterns were observed in the surgery-plus-chemotherapy cohorts. For example, when patients from CCF who received adjuvant chemotherapy were used as the target cohort, models trained on the French and Graz surgery-plus-chemotherapy cohorts achieved greater net benefit than those trained on JHH and YCU, consis- tent with their closer similarity to the target population. Additional DCA plots and predicted-risk histograms are provided in the Supplementary Appendix. 4 Discussion In this methodological study, we addressed a central challenge in prognostic modeling: that calibration is not an intrinsic property of a model, but depends in part on the population in which the model is applied. Across geographically distinct cohorts of patients with colorectal liver metastases, we found that external calibration varied substantially according to the degree of alignment between the training and target cohorts. More specifically, calibration deteriorated when cohorts differed either in the distribution of observed prognostic factors or in the relationship between those factors and outcomes. These findings provide empirical support, in a multicenter real-world setting, for the principle articulated by Andrew Vickers that calibration is a joint property of the model and the cohort to which it is applied.[11] 16 A. Tested on JHHB. Tested on CCFC. Tested on YCU D. Tested on GrazE. Tested on UOR Fig. 5: Pairwise cohort distance between training and test cohorts and the correspond- ing test-set ICI for models trained in each surgery-alone cohort. ICIs are averaged across three models — Cox, Optimal Survival Trees (OST), and Random Survival For- est (RSF) — and the error bars represent the minimum and maximum ICIs across the three models. Cohort distance is computed deterministically as the Euclidean distance between the cohorts’ 5-year Kaplan Meier estimates. Each panel represents one train- ing cohort, and the bars within that panel show results for all four test cohorts. Panels A–E correspond to models tested in JHH, CCF, YCU, Graz, and UOR, respectively. This observation is particularly important in light of expert guidance on prognos- tic model evaluation, which has identified calibration as the most important property of a prognostic model.[1] Nonetheless, in the current literature, discrimination metrics such as the C-index are reported routinely, whereas calibration is often underreported despite its more direct relevance to many clinical decisions. In fact, a recent review of prognostic models for colorectal liver metastases found that fewer than half of the mod- els assessed calibration, while nearly all reported discrimination.[12] This imbalance is problematic because discrimination and calibration answer fundamentally different questions. The C-index evaluates whether a model ranks patients correctly, whereas calibration evaluates whether predicted probabilities correspond to actual outcome fre- quencies. In many clinical settings, particularly when the goal is to estimate absolute risk and guide management, this latter question may be more important. Importantly, a model may provide clinically meaningful estimates of absolute risk even if its abil- ity to rank individual patients is only modest. For example, consider three CRLM patients for whom a model predicts 5-year recurrence risks of 40%, 42%, and 43%. These estimates are close to the cohort’s empirical recurrence risk and are therefore well calibrated. Yet if the actual recurrence order of these patients were reversed, the model’s C-index would be poor This is especially relevant in complex diseases such as 17 A. Tested on JHHB. Tested on CCFC. Tested on YCU D. Tested on GrazE. Tested on CHU Fig. 6: Pairwise KL divergence between training and test cohorts and the correspond- ing test-set ICI for models trained in each surgery + adjuvant chemotherapy cohort. ICIs are averaged across three models — Cox, Optimal Survival Trees (OST), and Random Survival Forest (RSF) — and the error bars represent the minimum and maximum ICIs across the three models. Cohort distance is computed deterministically as the Euclidean distance between the cohorts’ 5-year Kaplan Meier estimates. Each panel represents one training cohort, and the bars within that panel show results for all four test cohorts. Panels A–E correspond to models tested in JHH, CCF, YCU, Graz, and CHU, respectively. CRLM, where high discrimination may be difficult to achieve because important prog- nostic determinants remain unmeasured. By contrast, well-calibrated risk estimates can still support clinical decision-making even in the presence of incomplete biological knowledge. The first methodological contribution of this study addressed this problem from the perspective of the model developer. We proposed a meta-analysis-informed importance-weighting strategy that shifts the effective training distribution toward a broader target population using only summary-level outcome and covariate informa- tion that s typically available in meta-analyses. Empirically, this approach improved average external calibration across cohorts without materially affecting discrimination. The gain was greatest when the original training cohort was most dissimilar to the external cohorts in which the model was evaluated, which is consistent with the under- lying rationale of the method: the more idiosyncratic the derivation cohort, the more benefit there is in moving training toward a broader reference population. Concep- tually, this improvement appears to arise primarily through addressing concept shift. In our setting, concept shift may reflect both differences in standards of care across centers and differences in unobserved prognostic factors across cohorts. In CRLM, 18 0.20.30.40.50.60.70.8 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:41:23:44:32:14:1 Threshold Probability Cost:Benefit Ratio A. Trained on CCF 0.40.50.60.70.80.91.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 2:310:95:33:110:1100:1 Threshold Probability Cost:Benefit Ratio B. Trained on YCU ] 0.10.20.30.40.5 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:51:31:27:101:1 Threshold Probability Cost:Benefit Ratio C. Trained on Graz 0.20.40.60.81.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:41:210:92:15:1100:1 Threshold Probability Cost:Benefit Ratio D. Trained on UOR Fig. 7: Decision curve analysis showing the clinical utility of models trained in four different cohorts and externally validated in the JHH surgery-alone cohort. for example, prognosis may depend not only on routinely recorded clinicopathologic variables, but also on tumor genomics, immune microenvironment, micrometastatic burden, and other latent biological features that are incompletely captured in struc- tured datasets. By aligning the training process to outcome distributions observed in a broader external reference population, the proposed framework may partially account for these otherwise unmeasured differences and thereby yield better calibrated predictions across heterogeneous cohorts. This methodological contribution should also be understood in the context of the broader literature on distribution shift. A large body of statistical work has studied importance weighting as a means of addressing covariate shift, typically by reweight- ing the training data so that its covariate distribution better matches that of the target population. Early work by Shimodaira et al introduced this idea in the con- text of maximum likelihood estimation under covariate shift,[13] while Bickel and colleagues proposed to formulate classification tasks as an integrated optimization problem to account for train-test distribution mismatch without directly estimating each distribution.[14] Other methods, such as kernel mean matching, were developed to align distributions in feature space by matching moments of the training and test sets.[15, 16] Related work has also shown that standard cross-validation can become biased under covariate shift and proposed importance-weighted alternatives.[17] For a broader overview of this area, see the review by Nair and colleagues.[2] In most of these 19 0.50.60.70.80.91.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:13:25:24:110:1100:1 Threshold Probability Cost:Benefit Ratio A. Trained on JHH 0.20.30.40.50.60.70.8 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:41:23:44:32:14:1 Threshold Probability Cost:Benefit Ratio B. Trained on CCF 0.20.40.60.81.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:41:210:92:15:1100:1 Threshold Probability Cost:Benefit Ratio C. Trained on UOR 0.10.20.30.40.5 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:51:31:27:101:1 Threshold Probability Cost:Benefit Ratio D. Trained on Graz Fig. 8: Decision curve analysis showing the clinical utility of models trained in four different cohorts and externally validated in the YCU surgery-alone cohort. approaches, however, the target-domain distribution is assumed to be directly observ- able during model development, either through patient-level target data or through explicit access to the target covariate distribution. That assumption is often unrealis- tic in clinical prognostic modeling, where models are commonly developed in a single institutional cohort and later applied to external centers for which individual-level data are unavailable at the time of training. Our approach differs in that it does not require direct access to the eventual target cohort. Instead, it uses summary-level information from the published literature to approximate a broader reference population and then shifts the effective training distribution toward that population. Our framework is also related to the smaller literature on concept shift, which con- siders settings in which the relationship between predictors and outcomes differs across training and test populations. Nguyen et al., for example, showed that even ridge regression with infinite data may fail to generalize under strong concept shift[18] while Tian et al. proposed using Geometric Sensitivity Decomposition to improve the detec- tion of concept shift.[19] Prior work in this area has focused mainly on characterizing or detecting such shifts. In medicine, however, concept shift may arise from differences in standards of care, case mix, follow-up intensity, or unmeasured biological factors across centers. Our goal was therefore more pragmatic: to improve calibration across heterogeneous real-world cohorts by using external summary information to partially account for both covariate and concept mismatch between training and validation sets. 20 0.50.60.70.80.91.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:13:25:24:110:1100:1 Threshold Probability Cost:Benefit Ratio A. Trained on JHH 0.40.50.60.70.80.91.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 2:310:95:33:110:1100:1 Threshold Probability Cost:Benefit Ratio B. Trained on YCU 0.10.20.30.40.5 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:51:31:27:101:1 Threshold Probability Cost:Benefit Ratio C. Trained on Graz 0.20.40.60.81.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:41:210:92:15:1100:1 Threshold Probability Cost:Benefit Ratio D. Trained on UOR Fig. 9: Decision curve analysis showing the clinical utility of models trained in four different cohorts and externally validated in the CCF surgery-alone cohort. 0.200.250.300.350.400.45 Predicted Risk 0 5 10 15 20 Frequency Histogram of OST 0.2750.3000.3250.3500.3750.4000.4250.450 Predicted Risk Histogram of RSF A. Trained on Graz and Validated on CCF 0.30.40.50.60.70.80.91.0 Predicted Risk 0 1 2 3 4 5 Frequency Histogram of OST 0.550.600.650.700.750.800.850.90 Predicted Risk Histogram of RSF B. Trained on UOR and Validated on JHH 0.40.50.60.70.80.91.0 Predicted Risk 0 2 4 6 8 10 Frequency Histogram of OST 0.50.60.70.80.9 Predicted Risk Histogram of RSF C. Trained on YCU and Validated on JHH Fig. 10: Histogram of Predicted Risks for Surgery-alone Cohorts In this sense, our framework extends prior weighting approaches approaches to a set- ting in which neither the future deployment cohort nor its full patient-level distribution is known in advance. The second methodological contribution of this study addressed the same problem from the perspective of the end user. Even if a model is optimized to perform well on average across external cohorts, the most generalizable model will not necessarily 21 A. Trained on UOR and Validated on JHHB. Trained on Graz and Validated on JHH C. Trained on Graz and Validated on CCFD. Trained on UOR and Validated on CCF Fig. 11: Calibration plots for recurrence-free survival, showing external validation in the surgery-alone JHH and CCF cohorts of models developed in the surgery-alone UOR and Graz cohorts. Top left: UOR model validated in JHH. Top right: Graz model validated in JHH. Bottom left: Graz model validated in CCF. Bottom right: UOR model validated in CCF. be the best model for a specific center. Clinicians typically face a different practical question: among the many prognostic models already available in the literature, which one is most likely to be well calibrated in their own patient population? We therefore proposed a simple model-selection strategy based on proximity in outcome space, using summary-level statistics such as KM-based recurrence-free survival to quantify similarity between the target cohort and candidate development cohorts. Across most 22 cohort comparisons, the model trained on the most similar cohort also showed the best calibration in the target population. This suggests that even simple cohort-level outcome summaries may provide a useful and pragmatic guide for selecting among published prognostic models. This strategy differs from conventional approaches that prioritize model reputa- tion, prior external validation, or discrimination alone. It is also distinct from post-hoc recalibration approaches such as Platt scaling or isotonic regression. Those methods recalibrate predictions after model development and usually require data from the same population in which performance is later assessed, which may blur the distinc- tion between model selection and model evaluation. By contrast, our approach uses external summary-level information to guide model choice before deployment and without recalibrating on the target data itself. In this sense, it is better aligned with the real-world problem faced by clinicians deciding which published model to adopt. A particularly important finding of our study was that calibration and clinical util- ity, although related, are not identical. In many settings, models trained on cohorts similar to the target population achieved both lower ICI and greater net benefit on decision curve analysis. However, our results also revealed an important nuance: a model may be reasonably calibrated on average yet still provide limited clinical value if it predicts risks within a narrow range and therefore fails to separate patients into meaningfully different prognostic groups. This was illustrated by the low-risk Graz- to-CCF comparison, where the model trained on the Graz cohort was relatively well calibrated in the CCF cohort but generated only a narrow band of low predicted risks, leading to limited added value over trivial treatment strategies. By contrast, when models trained on YCU or UOR cohorts were applied to the JHH cohort, they generated a broader range of predicted risks and yielded greater clinical util- ity. These findings suggest that clinically useful prognostic models require not only good calibration, but also sufficient spread in predicted risk to support actionable stratification. This study has limitations. First, the empirical analyses were conducted in colorec- tal liver metastases, and although this setting is clinically relevant, generalizability to other diseases remains to be established. Second, the meta-analysis-informed target distributions used in the weighting framework are themselves approximations and may not fully represent the true deployment population. Third, the outcome-space model- selection strategy intentionally relies on simple summary statistics for practicality, but richer summaries may further improve performance. In conclusion, this study shows that calibration is strongly shaped by alignment between the development and target populations, and that poor calibration can be addressed from two directions: by training more generalizable models through meta- analysis-informed weighting, and by helping clinicians identify better-fitting models using simple outcome-based measures of cohort similarity. More broadly, our findings argue for a shift in how model generalizability and cohort-specific fit are assessed. This framework is highly practical for clinicians, both when training models intended to perform well on average across diverse settings and when selecting the model best suited to their own cohort for treatment-guiding predictions. 23 References [1] Alba, A.C., Agoritsas, T., Walsh, M., Hanna, S., Iorio, A., Devereaux, P., McGinn, T., Guyatt, G.: Discrimination and calibration of clinical prediction models: users’ guides to the medical literature. Jama 318(14), 1377–1384 (2017) [2] Nair, N.G., Satpathy, P., Christopher, J., et al.: Covariate shift: A review and analysis on classifiers. In: 2019 Global Conference for Advancement in Technology (GCAT), p. 1–6 (2019). IEEE [3] Ishwaran, H., Kogalur, U.B., Blackstone, E.H., Lauer, M.S.: Random survival forests (2008) [4] Bertsimas, D., Dunn, J., Gibson, E., Orfanoudaki, A.: Optimal survival trees. Machine learning 111(8), 2951–3023 (2022) [5] Woolson, R.F.: Wilcoxon signed-rank test. Wiley encyclopedia of clinical trials, 1–3 (2007) [6] Bosma, N.A., Keehn, A.R., Lee-Ying, R., Karim, S., MacLean, A.R., Brenner, D.R.: Efficacy of perioperative chemotherapy in resected colorectal liver metas- tasis: A systematic review and meta-analysis. European Journal of Surgical Oncology 47(12), 3113–3122 (2021) [7] Portier, G., Elias, D., Bouche, O., Rougier, P., Bosset, J.-F., Saric, J., Belghiti, J., Piedbois, P., Guimbaud, R., Nordlinger, B., et al.: Multicenter randomized trial of adjuvant fluorouracil and folinic acid compared with surgery alone after resection of colorectal liver metastases: Ffcd achbth aurc 9002 trial. Journal of Clinical Oncology 24(31), 4976–4982 (2006) [8] Nordlinger, B., Sorbye, H., Glimelius, B., Poston, G.J., Schlag, P.M., Rougier, P., Bechstein, W.O., Primrose, J.N., Walpole, E.T., Finch-Jones, M., et al.: Peri- operative folfox4 chemotherapy and surgery versus surgery alone for resectable liver metastases from colorectal cancer (eortc 40983): long-term results of a randomised, controlled, phase 3 trial. The lancet oncology 14(12), 1208–1215 (2013) [9] Hasegawa, K., Saiura, A., Takayama, T., Miyagawa, S., Yamamoto, J., Ijichi, M., Teruya, M., Yoshimi, F., Kawasaki, S., Koyama, H., et al.: Adjuvant oral uracil-tegafur with leucovorin for colorectal cancer liver metastases: a randomized controlled trial. PloS one 11(9), 0162400 (2016) [10] Kanemitsu, Y., Shimizu, Y., Mizusawa, J., Inaba, Y., Hamaguchi, T., Shida, D., Ohue, M., Komori, K., Shiomi, A., Shiozawa, M., et al.: Hepatectomy followed by mfolfox6 versus hepatectomy alone for liver-only metastatic colorectal can- cer (jcog0603): a phase i or i randomized controlled trial. Journal of Clinical Oncology 39(34), 3789–3799 (2021) 24 [11] Vickers, A.J., Cronin, A.M.: Everything you always wanted to know about eval- uating prediction models (but were too afraid to ask). Urology 76(6), 1298–1301 (2010) [12] Kokkinakis, S., Ziogas, I.A., Llaque Salazar, J.D., Moris, D.P., Tsoulfas, G.: Clinical prediction models for prognosis of colorectal liver metastases: A compre- hensive review of regression-based and machine learning models. Cancers 16(9), 1645 (2024) [13] Shimodaira, H.: Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of statistical planning and inference 90(2), 227–244 (2000) [14] Bickel, S., Br ̈uckner, M., Scheffer, T.: Discriminative learning under covariate shift. Journal of Machine Learning Research 10(9) (2009) [15] Huang, J., Gretton, A., Borgwardt, K., Sch ̈olkopf, B., Smola, A.: Correcting sam- ple selection bias by unlabeled data. Advances in neural information processing systems 19 (2006) [16] Gretton, A., Smola, A., Huang, J., Schmittfull, M., Borgwardt, K., Sch ̈olkopf, B., et al.: Covariate shift by kernel mean matching. Dataset shift in machine learning 3(4), 5 (2009) [17] Sugiyama, M., Krauledat, M., M ̈uller, K.-R.: Covariate shift adaptation by impor- tance weighted cross validation. Journal of Machine Learning Research 8(5) (2007) [18] Nguyen, A., Schwab, D.J., Ngampruetikorn, V.: Generalization vs. specialization under concept shift. arXiv preprint arXiv:2409.15582 (2024) [19] Tian, J., Hsu, Y.-C., Shen, Y., Jin, H., Kira, Z.: Exploring covariate and concept shift for out-of-distribution detection. In: NeurIPS 2021 Workshop on Distribution Shifts: Connecting Methods and Applications (2021) 25 Supplementary Appendix 0.50.60.70.80.91.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:13:25:24:110:1100:1 Threshold Probability Cost:Benefit Ratio A. Trained on JHH 0.20.30.40.50.60.70.8 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:41:23:44:32:14:1 Threshold Probability Cost:Benefit Ratio B. Trained on CCF 0.40.50.60.70.80.91.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 2:310:95:33:110:1100:1 Threshold Probability Cost:Benefit Ratio C. Trained on YCU 0.20.40.60.81.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:41:210:92:15:1100:1 Threshold Probability Cost:Benefit Ratio D. Trained on UOR Fig. 12: Decision curve analysis showing the clinical utility of models trained in four different cohorts and externally validated in the Graz surgery-alone cohort. 26 0.50.60.70.80.91.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:13:25:24:110:1100:1 Threshold Probability Cost:Benefit Ratio A. Trained on JHH 0.20.30.40.50.60.70.8 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:41:23:44:32:14:1 Threshold Probability Cost:Benefit Ratio B. Trained on CCF 0.40.50.60.70.80.91.0 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 2:310:95:33:110:1100:1 Threshold Probability Cost:Benefit Ratio C. Trained on YCU 0.10.20.30.40.5 0.0 0.2 0.4 0.6 0.8 1.0 Net Benefit OST RSF All None 1:51:31:27:101:1 Threshold Probability Cost:Benefit Ratio D. Trained on Graz Fig. 13: Decision curve analysis showing the clinical utility of models trained in four different cohorts and externally validated in the UOR surgery-alone cohort. 27 Fig. 14: Correlation between ICI and KL Divergence Fig. 15: Correlation between ICI and Mean 5-Year Survival Distance 28