Paper deep dive
Distilling CT Foundation Models into Editable Concept Bottlenecks for Lung Nodule Malignancy Prediction
Fakrul Islam Tushar, Stephen Adamo, Geoffrey D. Rubin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/11/2026, 4:13:50 AM
Summary
The paper presents a method for distilling frozen CT foundation models (CT-FM and FMCIB) into editable concept bottleneck models (CBMs) to predict lung nodule malignancy. By mapping foundation model embeddings to eight radiologist-defined attributes (e.g., spiculation, lobulation) and combining them with nodule size, the authors create transparent, interpretable models. The study evaluates concept fidelity, internal/external discrimination (AUROC), and editability on the LUNA25 and DLCS cohorts, finding that while concept fidelity is modest, the CBMs provide transparent explanations with performance comparable to nodule size alone.
Entities (10)
Relation Signals (8)
Concept Bottleneck Model → predicts → Malignancy Prediction
confidence 95% · predict malignancy from the estimated concepts and nodule size
CT-FM → usedin → Concept Bottleneck Model
confidence 95% · The models included CT-FM... We developed concept bottleneck models that map two frozen CT foundation-model representations
FMCIB → usedin → Concept Bottleneck Model
confidence 95% · and FMCIB, a nodule-focused contrastive encoder... We developed concept bottleneck models that map two frozen CT foundation-model representations
Concept Bottleneck Model → achievesauroc → AUROC
confidence 90% · Internally, the CT-FM and FMCIB concept+size models achieved AUROCs of 0.86
DLCS → usedforevaluation → Concept Bottleneck Model
confidence 90% · evaluated on a held-out internal test set and the external DLCS cohort
LUNA25 → usedfortraining → Concept Bottleneck Model
confidence 90% · Malignancy models were trained on LUNA25
CT-FM → hashigherfidelityfor → Spiculation
confidence 85% · Concept fidelity was modest but higher for FMCIB than CT-FM for ... spiculation (0.17 vs. 0.08)
FMCIB → hashigherfidelityfor → Lobulation
confidence 85% · Concept fidelity was modest but higher for FMCIB than CT-FM for ... lobulation (0.15 vs. 0.05)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Foundation models provide transferable CT representations, but predictions based directly on these embeddings are difficult to interpret. We developed concept bottleneck models that map two frozen CT foundation-model representations to eight radiologist-defined pulmonary-nodule attributes and predict malignancy from the estimated concepts and nodule size. The models included CT-FM, a whole-CT self-supervised encoder using a 96^3-voxel nodule-centered patch, and FMCIB, a nodule-focused contrastive encoder using a 50-mm crop. Eight ridge-regression concept heads were trained on 2,610 LIDC-IDRI nodules. Malignancy models were trained on LUNA25 and evaluated on a held-out internal test set and the external DLCS cohort. Concept fidelity was assessed using five-fold cross-validated R^2, and malignancy discrimination was assessed using AUROC with 95% confidence intervals estimated by patient-grouped bootstrap resampling. Concept fidelity was modest but higher for FMCIB than CT-FM for subtlety (R2, 0.24 vs. 0.11), spiculation (0.17 vs. 0.08), texture (0.17 vs. 0.07), and lobulation (0.15 vs. 0.05). Internally, the CT-FM and FMCIB concept+size models achieved AUROCs of 0.86 (95% CI, 0.80-0.92) and 0.86 (0.79-0.92), respectively. Externally, AUROCs were 0.72 (0.68-0.75) and 0.73 (0.70-0.76), compared with 0.73 for nodule size alone and 0.60 and 0.67 for the corresponding embedding only probes. Additive predictions could be decomposed into feature-level contributions and modified through controlled concept interventions. Concept bottlenecks provided transparent malignancy predictions with discrimination similar to nodule size alone, while differences in concept fidelity suggest that concept recovery depends on the underlying foundation-model representation.
Tags
Links
- Source: https://arxiv.org/abs/2608.07857v1
- Canonical: https://arxiv.org/abs/2608.07857v1
Trouble viewing inline? Open PDF directly →
Full Text
16,571 characters extracted from source content.
Expand or collapse full text
Further author information: (Send correspondence to F.I.T.) F.I.T.: E-mail: fitushar@arizona.edu Distilling CT Foundation Models into Editable Concept Bottlenecks for Lung Nodule Malignancy Prediction Fakrul I. Tushar Stephen Adamo Geoffrey D. Rubin Abstract Foundation models provide transferable CT representations, but predictions based directly on these embeddings are difficult to interpret. We developed concept bottleneck models that map two frozen CT foundation-model representations to eight radiologist-defined pulmonary-nodule attributes and predict malignancy from the estimated concepts and nodule size. The models included CT-FM, a whole-CT self-supervised encoder using a 96396^3-voxel nodule-centered patch, and FMCIB, a nodule-focused contrastive encoder using a 50-m crop. Eight ridge-regression concept heads were trained on 2,610 LIDC-IDRI nodules. Malignancy models were trained on LUNA25 and evaluated on a held-out internal test set and the external DLCS cohort. Concept fidelity was assessed using five-fold cross-validated R2R^2, and malignancy discrimination was assessed using AUROC with 95% confidence intervals estimated by patient-grouped bootstrap resampling. Concept fidelity was modest but higher for FMCIB than CT-FM for subtlety (R2R^2, 0.24 vs. 0.11), spiculation (0.17 vs. 0.08), texture (0.17 vs. 0.07), and lobulation (0.15 vs. 0.05). Internally, the CT-FM and FMCIB concept+size models achieved AUROCs of 0.86 (95% CI, 0.80–0.92) and 0.86 (0.79–0.92), respectively. Externally, AUROCs were 0.72 (0.68–0.75) and 0.73 (0.70–0.76), compared with 0.73 for nodule size alone and 0.60 and 0.67 for the corresponding embedding-only probes. Additive predictions could be decomposed into feature-level contributions and modified through controlled concept interventions. Concept bottlenecks provided transparent malignancy predictions with discrimination similar to nodule size alone, while differences in concept fidelity suggest that concept recovery depends on the underlying foundation-model representation. keywords: lung nodule, malignancy prediction, concept bottleneck model, foundation model, interpretable AI. Figure 1: Overview of the proposed concept-bottleneck framework. Frozen CT-FM[8] and FMCIB[7] embeddings are mapped to eight LIDC-IDRI[1] radiologist concepts, which are combined with nodule size in an additive model for malignancy prediction. LUNA25[9] is used for training and internal testing, DLCS[11, 10] for external testing, and embedding-only linear probes serve as black-box comparators. 1 INTRODUCTION Estimating pulmonary-nodule malignancy risk is central to management decisions in CT lung-cancer screening. Existing approaches include clinical risk models such as Brock/PanCan,[5] radiomic signatures, and deep-learning classifiers,[7, 10] yet nodule size remains a strong standalone predictor. More recently, CT foundation models have provided transferable representations for downstream imaging tasks,[7, 8] but classifiers operating directly on these embeddings remain difficult to interpret. Interpretability methods include post hoc feature-attribution approaches such as SHAP,[4] additive glass-box models,[6] and concept bottleneck models (CBMs),[3] with concept-embedding variants extending the framework.[2] Unlike post hoc explanations, CBMs route predictions through human-defined intermediate attributes, enabling feature-level explanations and controlled concept interventions. However, their clinical value depends on whether the predicted concepts faithfully represent the intended radiologic characteristics. This fidelity may vary across foundation models because their pretraining objectives, spatial context, and representation dimensionality differ. We therefore developed editable CBMs from two frozen CT foundation models: a whole-CT self-supervised encoder and a nodule-focused contrastive encoder.[8, 7] Each embedding was mapped to eight radiologist-defined LIDC-IDRI attributes,[1] and an additive classifier predicted malignancy from the estimated concepts and nodule size. We evaluated cross-validated concept fidelity, internal and external malignancy discrimination,[9, 11] and concept-level editability, while benchmarking against nodule size and embedding-based linear probes. An overview of the study framework is shown in Fig. 1. Table 1: Cohort characteristics and study roles. Continuous variables are median [IQR]; splits are reported as nodule counts. LIDC-IDRI is the concept-annotation reference cohort (age and sex not applicable). Characteristic LUNA25[9] DLCS[11] LIDC-IDRI[1] Source/region NLST, 33 US centers Duke, single US center 7 US centers (public) Patients 2,120 1,613 1,010 Nodules 6,163 2,487 2,610 Nodule size, m [IQR] 6.0 [5.0–8.1] 5.4 [4.4–7.6] 5.7 [4.5–8.2] Malignant, n (%) 555 (9.0) 264 (10.6) Radiologist ratings Age, y [IQR] 63 [59–67] 67 [62–72] — Female, n (%) 909 (42.9) 803 (49.8) — Use (nodules) Train + validation: 5,302; internal test: 581 (280 excluded) External test: 2,487 Concept-head training (8 attributes) 2 Methods 2.1 Datasets and concept labels We used two publicly derived low-dose CT lung-cancer screening cohorts and one radiologist-annotated reference set for the concept labels (Table 1). Screening cohorts. LUNA25,[9] derived from the National Lung Screening Trial, comprises 6,163 nodules from 2,120 participants, of which 555 (9.0%) are malignant. The Duke Lung Cancer Screening cohort (DLCS),[11] from a single US academic center, comprises 2,487 nodules from 1,613 patients, including 264 malignant nodules (10.6%). The malignancy reference standard was the cohort-provided nodule label. Concept-label set. The eight radiologist concepts were taken from LIDC-IDRI,[1] a public collection from seven US academic centres in which Up to four thoracic radiologists assessed each nodule ≥3≥ 3 m using eight semantic characteristics. Median reader ratings were used for concept-head training. The eight concepts are subtlety (conspicuity against surrounding lung), internal structure (soft tissue, fluid, fat, or air), calcification (pattern or absence), sphericity (3-D roundness), margin (sharpness of the border), lobulation (lobulated contour), spiculation (spiculated margin), and texture (solid, part-solid, or ground-glass). Spiculation, lobulation, subtlety, and texture are the morphologic hallmarks clinicians use to judge malignancy. Partitions. Each screening cohort carries a fixed, patient-grouped train/validation/test split (Table 1). The concept heads were trained on all 2,610 LIDC nodules. The glass-box malignancy model and the black-box probe were trained on the union of the LUNA25 training and validation partitions (5,302 nodules; 280 unassigned nodules excluded) and evaluated (i) internally on the held-out LUNA25 test partition (581 nodules) and (i) externally on the full DLCS cohort (2,487 nodules), which contributes no training data. Splits were grouped by patient, and bootstrap resampling used patients as the sampling unit. 2.2 Foundation-Model Features and Glass-Box Concept Bottleneck Two contrasting foundation models. Each nodule was represented by frozen embeddings from two encoders selected to differ in pretraining and field of view: CT-FM,[8] a whole-CT self-supervised encoder evaluated on a 96396^3 nodule patch (512 dimensions), and FMCIB,[7] a nodule-contrastive encoder evaluated on a tight 50×50×5050× 50× 50 m crop at 1-m isotropic resolution (4,096 dimensions). Neither encoder was fine-tuned; only linear readouts were trained, isolating the information contained in each frozen representation. Concept bottleneck. For each FM, we trained eight ridge heads, with regularization tuned by cross-validation, to map its embedding to the eight LIDC-IDRI attributes, forming the interpretable bottleneck =g(embedding)c=g(embedding). A glass-box additive classifier[6] then predicted malignancy from the eight predicted concepts and nodule size, y^=f(,size) y=f(c,size) (Fig. 1). Because the classifier is additive, each prediction can be decomposed into feature-level contributions and evaluated under controlled concept interventions. Baselines. For each FM, we evaluated a raw-embedding logistic probe (the black-box comparator, without concepts) and a shared nodule-size-only reference model. 2.3 Experiments For each FM, we evaluated (1) Concept fidelity was assessed using five-fold cross-validated R2R^2 of each concept head on LIDC-IDRI (in-sample R2R^2 overstates fidelity, particularly for the 4,096-dimensional FMCIB embedding); (2) malignancy discrimination of the CBM relative to the black-box and nodule-size baselines, both internally on LUNA25 and externally on DLCS, with patient-grouped bootstrap 95% confidence intervals based on 2,000 resamples; and (3) editability, measured as the mean change in predicted malignancy risk when a concept was moved from its low value (10th percentile) to its high value (90th percentile), while all other concepts were held fixed. Figure 2: Malignancy discrimination on (a) the LUNA25 internal test set and (b) the external DLCS cohort. AUROC and ROC curves compare nodule size, the eight-concept CBMs, and embedding-only linear probes for CT-FM and FMCIB. Error bars indicate patient-grouped bootstrap 95% confidence intervals. Externally, both CBMs performed similarly to size and outperformed their corresponding embedding-only probes. 3 Results Concept fidelity was modest and differed between foundation models. Five-fold cross-validated R2R^2 was consistently higher for FMCIB than for CT-FM for subtlety (0.24 vs. 0.11), spiculation (0.17 vs. 0.08), texture (0.17 vs. 0.07), and lobulation (0.15 vs. 0.05). Fidelity was near zero for calcification, sphericity, and internal structure for both models. In-sample R2R^2 substantially overestimated fidelity, reaching approximately 0.3–0.4 for CT-FM and 0.98 for FMCIB. These findings suggest that the nodule-focused FMCIB representation captured radiologist-defined morphology more effectively, although differences in architecture, pretraining, and input context prevent attributing this effect to field of view alone. Concept-plus-size models performed similarly to nodule size alone. The CT-FM CBM achieved AUROCs of 0.86 (95% CI, 0.80–0.92) on the LUNA25 internal test set and 0.72 (0.68–0.75) on DLCS. Corresponding FMCIB results were 0.86 (0.79–0.92) and 0.73 (0.70–0.76), respectively (Fig. 2). Both CBMs had higher AUROC point estimates than their embedding-only probes, particularly externally, where the CT-FM and FMCIB probes achieved 0.60 and 0.67, respectively. However, CBM performance was similar to nodule size alone, which achieved AUROCs of 0.86 internally and 0.73 externally. This pattern indicates that most malignancy discrimination was carried by nodule size, while the concept bottleneck primarily added interpretability. The additive models supported concept-level editing and prediction decomposition. Moving each concept from its 10th to 90th percentile while holding the remaining inputs fixed produced the largest mean risk changes for lobulation (+0.04+0.04 for CT-FM and +0.07+0.07 for FMCIB), subtlety (+0.02+0.02 and +0.05+0.05), and spiculation (+0.03+0.03 for both models) (Fig. 3a). Sphericity, margin, and calcification produced minimal changes. For the representative malignant and benign nodules, predictions were decomposed into contributions from nodule size and each predicted concept (Fig. 3b,c). Size contributed most strongly, while lobulation and spiculation provided smaller, directionally consistent contributions. Thus, the models enabled transparent feature-level explanations and controlled model-level interventions, although the concepts contributed less predictive information than size. Figure 3: Concept editability and local explanations of the eight-concept CBMs. (a) Mean change in malignancy risk after shifting each concept from its 10th to 90th percentile on the LUNA25 test set. (b,c) Per-feature contributions to malignancy log-odds for representative malignant and benign nodules. CT-FM is shown in blue and FMCIB in gray. 4 Discussion Distilling two frozen CT foundation models into radiologist-defined concepts produced interpretable and editable malignancy models that maintained performance on an external cohort. Both concept-plus-size models had higher AUROC point estimates than their embedding-only probes but performed similarly to nodule size alone, indicating that most discrimination was size-driven. Thus, the concept bottleneck primarily added transparent feature-level explanations rather than improved accuracy. Concept fidelity was modest but higher for FMCIB for several morphologic attributes, suggesting that concept recovery depends on the underlying representation. However, differences in pretraining, architecture, dimensionality, and field of view prevent attributing this result to field of view alone. In-sample R2R^2 substantially overstated fidelity, supporting cross-validated evaluation. Limitations include concept supervision from one cohort, frozen linear heads, and evaluation of only two foundation models. Future work will assess additional encoders, matched embedding-plus-size baselines, and fine-tuned concept models. Acknowledgements.This work was supported by startup funding from the Department of Radiology and Imaging Sciences at the University of Arizona. References [1] S. G. Armato I, G. McLennan, L. Bidaut, M. F. McNitt-Gray, C. R. Meyer, A. P. Reeves, B. Zhao, D. R. Aberle, C. I. Henschke, E. A. Hoffman, et al. (2011) The lung image database consortium (lidc) and image database resource initiative (idri): a completed reference database of lung nodules on ct scans. Medical physics 38 (2), p. 915–931. Cited by: Figure 1, Table 1, §1, §2.1. [2] M. Espinosa Zarlenga, P. Barbiero, G. Ciravegna, G. Marra, F. Giannini, M. Diligenti, Z. Shams, F. Precioso, S. Melacci, A. Weller, et al. (2022) Concept embedding models: beyond the accuracy-explainability trade-off. Advances in neural information processing systems 35, p. 21400–21413. Cited by: §1. [3] P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang (2020) Concept bottleneck models. In International conference on machine learning, p. 5338–5348. Cited by: §1. [4] S. M. Lundberg and S. Lee (2017) A unified approach to interpreting model predictions. Advances in neural information processing systems 30. Cited by: §1. [5] A. McWilliams, M. C. Tammemagi, J. R. Mayo, H. Roberts, G. Liu, K. Soghrati, K. Yasufuku, S. Martel, F. Laberge, M. Gingras, et al. (2013) Probability of cancer in pulmonary nodules detected on first screening ct. New England journal of medicine 369 (10), p. 910–919. Cited by: §1. [6] H. Nori, S. Jenkins, P. Koch, and R. Caruana (2019) Interpretml: a unified framework for machine learning interpretability. arXiv preprint arXiv:1909.09223. Cited by: §1, §2.2. [7] S. Pai, D. Bontempi, I. Hadzic, V. Prudente, M. Sokač, T. L. Chaunzwa, S. Bernatz, A. Hosny, R. H. Mak, N. J. Birkbak, et al. (2024) Foundation model for cancer imaging biomarkers. Nature machine intelligence 6 (3), p. 354–367. Cited by: Figure 1, §1, §1, §2.2. [8] S. Pai, I. Hadzic, D. Bontempi, K. Bressem, B. H. Kann, A. Fedorov, R. H. Mak, and H. J. Aerts (2025) Vision foundation models for computed tomography. arXiv preprint arXiv:2501.09001. Cited by: Figure 1, §1, §1, §2.2. [9] D. Peeters, B. Obreja, N. Antonissen, Z. Saghir, U. Pastorino, M. Silva, G. H. de Bock, H. Gietema, F. Gleeson, M. A. Heuvelmans, et al. (2026) Benchmarking of ai and radiologists for indeterminate lung nodule malignancy risk estimation on screening ct: the luna25 challenge. Note: Radiology: Artificial Intelligence, advance online publication, article e260179Published online June 24, 2026 External Links: Document Cited by: Figure 1, Table 1, §1, §2.1. [10] F. I. Tushar, A. Wang, L. Dahal, E. Samei, M. R. Harowicz, J. Kalpathy-Cramer, K. J. Lafata, T. D. Tailor, C. Rudin, and J. Y. Lo (2024) Reproducible benchmarking for lung nodule detection and malignancy classification across multiple low-dose ct datasets. arXiv preprint arXiv:2405.04605. Cited by: Figure 1, §1. [11] A. J. Wang, F. I. Tushar, M. R. Harowicz, B. C. Tong, K. J. Lafata, T. D. Tailor, and J. Y. Lo (2025) The duke lung cancer screening (dlcs) dataset: a reference dataset of annotated low-dose screening thoracic ct. Radiology: Artificial Intelligence 7 (4), p. e240248. Cited by: Figure 1, Table 1, §1, §2.1.