Paper deep dive
FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making
Pramit Dutta, Jenita Manokaran, Richa Mittal, Ryan Appleby, Eranga Ukwatta
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/18/2026, 5:08:33 AM
Summary
The paper introduces FZ-VLM, a two-stage Vision-Language Model framework for pulmonary nodule characterization in lung CT scans. Stage 1 uses a fine-tuned Florence-2 model to extract radiological attributes (anatomical location, diameter, margin characteristics, attenuation type) from 2D axial CT slices. Stage 2 uses a Zephyr-7B language model to generate structured nodule descriptions, follow-up recommendations, and longitudinal analyses based on the extracted attributes. The framework was evaluated on a dataset curated from the National Lung Screening Trial (NLST), demonstrating superior performance over GPT-4 baselines and human baselines in attribute extraction, and high clinical relevance in expert radiologist evaluations.
Entities (8)
Relation Signals (7)
FZ-VLM → uses → Zephyr-7B
confidence 98% · a Zephyr-7B model uses these attributes to generate nodule descriptions
FZ-VLM → uses → Florence-2
confidence 98% · The framework uses a fine-tuned Florence-2 model to extract radiological attributes
FZ-VLM → evaluatedon → National Lung Screening Trial
confidence 95% · dataset ... was curated from 745 patient cases derived from the National Lung Screening Trial.
Florence-2 → extracts → radiological attributes
confidence 95% · extract radiological attributes from expert-annotated 2D axial CT slices
Zephyr-7B → generates → clinical recommendations
confidence 95% · generate nodule descriptions, follow-up recommendations, and longitudinal analyses
Computed Tomography → usedfor → Lung Cancer Screening
confidence 95% · Computed Tomography (CT) is a primary imaging tool for screening and followup assessment.
FZ-VLM → outperforms → GPT-4
confidence 90% · outperforming evaluated GPT-4-based baselines as well as the human baseline.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment. After pulmonary nodule detection, radiologists manually assess anatomical location, diameter, margin characteristics, and attenuation type to support risk assessment and clinical decision-making. However, this post-detection workflow is time-consuming and can be affected by inter-observer variability. Existing Artificial Intelligence methods often focus on isolated tasks, limiting their use as a unified, clinically grounded interpretation framework. This study presents FZ-VLM, a two-stage Florence-Zephyr Vision Language Model framework for unified structured pulmonary nodule characterization in lung CT. The framework uses a fine-tuned Florence-2 model to extract radiological attributes from expert-annotated 2D axial CT slices, while a Zephyr-7B model uses these attributes to generate nodule descriptions, follow-up recommendations, and longitudinal analyses. Results showed that the Stage 1 model achieved 77.18\% accuracy for anatomical location, 67.96\% accuracy for margin characteristics, and 79.13\% accuracy for attenuation type, with a Mean Absolute Error of 2.58 mm for diameter estimation, outperforming evaluated GPT-4-based baselines as well as the human baseline. Expert radiologist evaluation of Stage 2 showed 93.9\% accuracy, 98.6\% completeness score, 76.1\% clinical relevance, and an overall score of 89.5\%. Safety analysis showed that most outputs were clinically safe, although some follow-up recommendations still required expert review. To the best of our knowledge, this study presents the first two-stage Vision-Language Model framework for structured nodule characterization and clinical decision-making.
Tags
Links
- Source: https://arxiv.org/abs/2608.15004v1
- Canonical: https://arxiv.org/abs/2608.15004v1
Trouble viewing inline? Open PDF directly →
Full Text
55,489 characters extracted from source content.
Expand or collapse full text
FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making Pramit Dutta Affiliation: College of Engineering, University of Guelph, Guelph, ON N1G 2W1, Canada Affiliation: Corresponding author: eukwatta@uoguelph.ca Jenita Manokaran Affiliation: Holland Bloorview Kids Rehabilitation Hospital, Toronto, ON M4G 1R8, Canada Richa Mittal Affiliation: Guelph General Hospital, Guelph, ON N1E 4J4, Canada Ryan Appleby Affiliation: Department of Clinical Studies, Ontario Veterinary College, University of Guelph, Guelph, ON N1G 2W1, Canada Eranga Ukwatta Affiliation: College of Engineering, University of Guelph, Guelph, ON N1G 2W1, Canada Abstract Lung cancer remains one of the leading causes of cancer related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment. After pulmonary nodule detection, radiologists manually assess anatomical location, diameter, margin characteristics, and attenuation type to support risk assessment and clinical decision making. However, this post detection workflow is time consuming and can be affected by inter observer variability. Existing Artificial Intelligence methods often focus on isolated tasks, which limits their use as a unified and clinically grounded interpretation framework. This study presents FZ-VLM, a two stage Florence-Zephyr Vision Language Model framework for unified structured pulmonary nodule characterization in lung CT. The framework uses a fine-tuned Florence-2 model to extract radiological attributes from expert annotated 2D axial CT slices, while a Zephyr-7B model uses the extracted attributes to generate nodule descriptions, follow-up recommendations, and longitudinal analysis. A structured Visual Question Answering dataset containing 8,540 image-prompt-answer triplets was curated from 745 patient cases derived from the National Lung Screening Trial. Results showed that the Stage 1 model achieved 77.18% accuracy for anatomical location, 67.96% accuracy for margin characteristics, and 79.13% accuracy for attenuation type, with a Mean Absolute Error of 2.58 m for diameter estimation, outperforming evaluated GPT-4-based baselines as well as human baseline. Expert radiologist evaluation of Stage 2 showed 93.9% accuracy, 98.6% completeness score, 76.1% clinical relevance, and an overall score of 89.5%. Safety analysis showed that most outputs were clinically safe, although some follow up recommendations still required expert review. To the best of our knowledge, this study presents the first two stage Vision Language Model framework for structured nodule characterization and clinical decision making system. keywordsVision Language Model, Explainability, National Lung Screening Trial, Pulmonary Nodule, Lung CT, Expert Radiologist Grading Introduction Lung cancer is the leading cause of cancer-related mortality worldwide, causing nearly 1.8 million deaths each year [28]. Early diagnosis can increase the five-year survival rate to approximately 65.5%, compared with less than 29.5% overall survival [21]. Low dose computed tomography (LDCT) is the primary imaging modality for lung cancer screening and follow-up assessment [29]. The clinical workflow considered in this study consists of three stages. The process begins with high-risk lung cancer screening and LDCT image acquisition, followed by CT volume review for nodule identification. After a pulmonary nodule is detected, the radiologist selects the representative CT slice where the longest nodule diameter is most visible. In the post-detection stage, key radiological attributes, including anatomical location, diameter, margin characteristics, and attenuation type, are extracted to support risk assessment, follow-up recommendations, and treatment planning [9, 6]. The proposed framework focuses on this post-detection interpretation stage, where structured nodule characterization is required for clinical decision making [1]. In this study, four parameters were considered: anatomical location, diameter, margin characteristics, and attenuation type. Anatomical location identifies the affected lung lobe [18], while in the curated NLST annotations nodule diameter is calculated as the average of the long-axis and perpendicular short-axis measurements, rounded to the nearest whole millimeter, and used to support size-based management [1]. Margin characteristics describe the nodule boundary [7, 33], whereas attenuation represents its internal density pattern [8, 36]. Accurate extraction of these attributes is therefore important for risk stratification and follow-up decisions. However, manual attribute extraction is time consuming and affected by interobserver variability, particularly for qualitative features such as margin and attenuation. Unanimous correct classification among expert radiologists has been reported in only 58% of pulmonary nodule cases [11, 25]. To address this issue several deep learning methods have been developed for detection, segmentation, location assessment, diameter estimation, margin characterization, and attenuation classification [19, 26, 7, 36]. However, most methods remain task specific and do not provide a unified framework for structured nodule characterization [6]. Vision Language Models (VLMs) offer a possible solution because they connect visual information with language prompts and structured questions. General-purpose VLMs support image-text alignment, visual question answering, grounding, and instruction following [23, 15, 32]. However, their direct use in radiology remains limited because CT findings involve subtle patterns, domain-specific labels, and clinically sensitive decisions [5, 16]. Medical adaptation through prompt learning, fine-tuning, structured supervision, and clinically grounded image evidence is therefore important [5, 31, 4]. Structured visual question answering is particularly suitable for pulmonary nodule analysis because it converts broad image interpretation into focused questions about location, size, margin, and attenuation [1]. Recent studies suggest that structured VQA and guideline-based prompting can improve clinically meaningful interpretation [3, 12, 27, 35]. However, medical VLMs may still produce fluent responses that are weakly grounded in the image, creating risks of hallucination and unsafe recommendations [6, 5]. Reliable systems should therefore use clinically constrained outputs and task-specific evaluation. Large language models can provide this reasoning layer that organizes extracted findings and generates structured clinical explanations. Their use should remain assistive because overconfident or poorly calibrated outputs may cause harm [34, 20, 10]. Explainability is also important for clinical use. Visual evidence can indicate whether a model focuses on the relevant lesion, while textual reasoning can show how extracted attributes support the generated interpretation. Prior studies have used saliency maps, attention visualization, quantitative explanation metrics, radiomics, and guideline-based reasoning to improve transparency [24, 2, 13, 14, 27]. A clinically useful framework should therefore provide both image-level and language-level explanations. Large screening datasets remain underused for structured VLM development. The National Lung Screening Trial (NLST) is a major multicenter and multivendor lung cancer screening dataset, but it was not designed for VQA [29]. Recent lung nodule VLM studies more commonly use LIDC-IDRI because it provides accessible semantic annotations and segmentation labels [12, 27, 35]. NLST therefore requires additional curation to create image-prompt-answer triplets for attribute-based learning. These limitations indicate the need for a unified framework that can extract multiple pulmonary nodule attributes from CT images and use them as controlled inputs for clinical reasoning. Therefore, this study proposes FZ-VLM, a two-stage Florence-Zephyr framework for structured pulmonary nodule characterization and clinical decision support. The first stage uses a fine-tuned Florence-2 model to extract radiological attributes from expert-selected 2D axial CT slices [32]. The second stage provides these image-derived attributes to Zephyr-7B to generate nodule descriptions, follow-up recommendations, and longitudinal analyses [30]. This attribute-to-reasoning design limits unrestricted language generation and helps ground the generated outputs in explicit radiological findings [17]. The study also transforms NLST imaging data and associated annotations into structured image-prompt-answer triplets for model development and evaluation [29]. The framework was evaluated through quantitative analysis, image-level explainability for qualitative analysis, slice-sensitivity analysis, baseline comparison, external generalization assessment, and expert radiologist review. The main contribution of this work is a unified pipeline that connects visually grounded nodule characterization with structured language-based clinical interpretation. The study further contributes an NLST-derived VQA dataset and two complementary forms of interpretability: Florence-2 attention heatmaps for image-level evidence and Zephyr-7B reasoning for language-level explanation. To the best of our knowledge, this is the first two-stage VLM framework developed specifically for structured pulmonary nodule characterization and subsequent clinical decision support. Methods This study developed a two stage Florence–Zephyr Vision Language Model (FZ-VLM) for structured interpretation of pulmonary nodules in low dose CT images. The methodological workflow comprised dataset curation, Florence-2 fine-tuning, hierarchical integration of the extracted attributes, zero-shot clinical interpretation using Zephyr-7B and overall evaluation strategy. Data Curation and Dataset Construction Low dose chest CT examinations were obtained from the National Lung Screening Trial dataset. Cases were selected from screen detected and non-screen detected lung cancer cohorts for which confirmed disease status, nodule location, representative slice index, and expert annotated nodule attributes were available. Cases with missing, indeterminate, or incomplete annotations were excluded. For each nodule, the axial slice containing the maximum nodule diameter was selected using the reference slice index. CT intensities were converted from Hounsfield units to 8-bit image intensities using a lung window with a width of 1700 HU and a level of −700-700 HU. Values below −1550-1550 HU and above 150 HU were clipped before linear intensity normalization. The selected window was applied uniformly to all images. Then, each CT image was paired with standardized questions targeting four radiological attributes, and the attribute as answer. The resulting dataset contained 8,540 image-question-answer triplets, divided into training (7,332), validation (384), and held-out test (824) sets using an approximately 85:5:10 ratio. Splitting was performed at the patient level so that all images and attribute questions associated with an individual patient remained in the same partition, thereby preventing information leakage between training and evaluation data. Florence–Zephyr Vision Language framework We developed the Florence–Zephyr Vision-Language Model, or FZ-VLM, as a two-stage framework for structured pulmonary nodule interpretation, as illustrated in Figure 1. The first stage uses a domain-adapted Florence-2 model to extract visually grounded radiological attributes from CT images. The second stage uses Zephyr-7B to transform the extracted attributes into structured nodule descriptions, follow-up recommendations, and longitudinal interpretations. Figure 1: Architecture of the FZ-VLM framework. The pipeline executes attribute extraction by a finetuned Florence-2 with explainable attention heatmaps. These parameters are processed through a hierarchical integration layer to serve as contextual input for a Zephyr-7B model that performs zero-shot inference to generate structured clinical outputs. The two-stage design separates image perception from clinical language generation. Consequently, the reasoning model receives an explicit set of radiological attributes rather than unrestricted visual embeddings. This design allows the generated clinical interpretation to be traced to specific model-extracted findings and permits the visual and reasoning components to be evaluated independently. Stage 1: Florence-2 Feature Extractor In Stage 1 of the FZ-VLM framework the large variant of Florence-2 was used as the visual attribute extraction model. The model contains approximately 0.77 billion parameters and includes a DaViT vision encoder, a text embedding block, and a multimodal encoder-decoder architecture. Each input consisted of a windowed axial CT image paired with one of four standardized questions targeting anatomical location, diameter, margin characteristics, or attenuation type. The model generated the corresponding radiological attribute as a text response. The pretrained Florence-2 model was fine-tuned using the training subset of the curated VQA dataset and evaluated during training using the validation subset. This task-specific adaptation enabled the model to learn the relationship between visual features in low-dose CT slices and the corresponding expert-annotated radiological attributes. Fine-tuning was conducted on the Rorqual computing cluster provided by the Digital Research Alliance of Canada using two NVIDIA H100 GPUs. The model was trained using the AdamW optimizer with a base learning rate of 2×10−62× 10^-6. A cosine decay learning rate scheduler with two warm up steps was applied over 10 training epochs. The batch size was two image-question-answer triplets per device. The trained model was subsequently evaluated on the held out unseen test set. Hierarchical integration of extracted attributes The four outputs generated by Florence-2 were combined through a hierarchical integration layer, as illustrated in Figure 2. At the nodule level, the extracted radiological attributes were assembled into one structured nodule representation for nodule description. At the study-year level, all nodules identified within a screening examination were grouped to represent the complete imaging findings and support follow-up recommendations. At the patient level, the study-year representations were ordered chronologically to create a longitudinal record for temporal analysis. Figure 2: Hierarchical integration of the radiological attributes extracted by Florence-2. Individual nodule features are combined into various profiles for different assessment. Stage 2: Clinical Decision Making using Zephyr-7B Zephyr-7B was used as the clinical decision making component. The model is an instruction aligned language model derived from the Mistral-7B architecture and optimized through direct preference optimization. In the proposed framework, Zephyr-7B was used without additional fine-tuning. Structured attributes generated by Stage 1 were inserted into standardized prompt templates. Three clinical tasks were evaluated: the nodule description task generated a concise radiological description from the attributes of an individual nodule. The follow-up recommendation task produced an overall recommendation based on all nodules identified during a single screening examination. The longitudinal analysis task interpreted interval changes across multiple screening examinations, including changes in nodule number, size, margin, and attenuation. Task-specific system instructions defined the model as a medical assistant and directed it to base its response exclusively on the provided structured findings. The prompts did not include demographic, pathological, or outcome information that was unavailable from Stage 1. This constraint was used to reduce unsupported inference and maintain traceability between the generated output and the extracted radiological evidence. Overview of Experimental Evaluation The proposed FZ-VLM framework was evaluated using unseen test data that was not used during model training or validation. The evaluation was designed to assess both stages of the framework separately. This two level evaluation strategy was used to examine whether the framework could first extract reliable radiological attributes from CT images and then use those attributes to generate clinically meaningful interpretation. In Stage 1, the fine-tuned Florence-2 model was assessed through five evaluation phases: standard quantitative performance analysis, image level explainability assessment, slice selection sensitivity analysis, baseline comparison, and external validation. Standard evaluation parameters were used to measure attribute extraction performance, including accuracy, precision, recall, F1-score, confusion matrix analysis, and mean absolute error for diameter estimation. In Stage 2, the Zephyr-7B model was evaluated through expert radiologist review. The evaluation focused on generated clinical outputs, and each output was reviewed using clinically relevant parameters. Accuracy assessed whether the generated output correctly represented the clinical finding or recommendation. Completeness evaluated whether the output included the necessary clinical information without missing important details. Clinical relevance assessed whether the interpretation was meaningful for pulmonary nodule assessment. Safety evaluated whether the output avoided unsupported, misleading, or potentially harmful clinical recommendations. Results This section presents the experimental evaluation of the proposed FZ-VLM framework. The evaluation is divided into two main stages, where Stage 1 is assessed for radiological attribute extraction and Stage 2 is evaluated for generated clinical interpretation via expert radiologist review. Additional analyses included explainability assessment, slice sensitivity testing, baseline comparison, and external generalization. Quantitative Performance Analysis of Florence-2 For radiological attribute extraction, the fine-tuned Florence-2 model achieved an overall accuracy of 74.42% with 77.18% for anatomical location, 67.96% for margin characteristics, and 79.13% for attenuation type on the held-out test set. Table 1 summarizes the macro average precision, recall, and F1-score for the three categorical radiological attributes. Anatomical location extraction showed the strongest macro F1-score among the categorical tasks, while margin characteristics showed moderate performance. Attenuation type achieved the highest task-level accuracy, although its lower macro average scores indicate reduced performance for minority classes. Table 1: Performance of the fine-tuned Florence-2 model on the test set. Macro average precision, recall, F1-score, and Accuracy are reported for categorical attributes. For diameter prediction, accuracy was measured using a tolerance from 0 to 5 m. Parameter Precision Recall F1-score Accuracy Anatomical location 0.75 0.72 0.70 77.18% Margin characteristics 0.65 0.64 0.64 67.96% Attenuation type 0.58 0.52 0.53 79.13% Nodule diameter MAE = 2.58 m; SD = 3.73 m; Accuracy = 21.4–87.9% For diameter prediction, the model achieved an MAE of 2.58 m with an SD of 3.73 m. The result also shows that accuracy increased from 21.4% at exact match to 87.9% within ± 5 m tolerance. Although the model showed relatively low exact match accuracy, it performed reliably within small measurement tolerances. This suggests that the predicted diameters were generally close to the reference measurements, as also reflected by the median absolute error of 1 m. Explainability Analysis Using Attention Heatmaps To analyze the decision-making behavior of the fine-tuned Florence-2 model, attention heatmaps were generated for representative test cases. These heatmaps provide a visual indication of the image regions that received higher model attention during radiological attribute extraction. In this analysis, selected cases were examined to understand how the model focused on nodule-related regions, how correct predictions were formed, and how different types of prediction errors appeared across anatomical location, diameter, margin characteristics, and attenuation type. (a) Case 1: LUL, 5 m, poorly defined, ground glass. Model: LUL, 7 m, poorly defined, ground glass. (b) Case 2: RUL, 6 m, poorly defined, soft tissue. Model fully matched the reference attributes. (c) Case 3: LUL, 15 m, poorly defined, ground glass. Model: LUL, 13 m, poorly defined, ground glass. (d) Case 4: LUL, 5 m, smooth, soft tissue. Model fully matched the reference attributes. (e) Case 5: LUL and margin were correct, but diameter and attenuation were incorrect. (f) Case 6: Severe error case with incorrect location, diameter, margin, and attenuation. Figure 3: Qualitative explainability and error analysis using attention heatmaps. Representative cases show correct predictions and error patterns in Stage 1 radiological attribute extraction. LUL: left upper lobe; RUL: right upper lobe. Figure 3 presents representative examples of correct, partially correct, and incorrect predictions. Cases 1–4 show mostly reliable categorical attribute extraction. In Cases 1 and 3, the model correctly predicted anatomical location, margin, and attenuation, while the predicted diameter differed slightly from the reference measurement. Case 2 and Case 4 showed complete agreement with the reference attributes. These examples suggest that the model can often extract key categorical nodule features when the relevant nodule region is visually clear in the heatmap. Case 5 shows a partial error pattern. The model correctly identified the anatomical location and margin, but it overestimated the diameter and misclassified the attenuation type. This suggests that the model may focus on a broader visual region or surrounding opacity when estimating size and density. Case 6 shows a more severe failure pattern, where the model output did not match the reference labels for anatomical location, margin, or attenuation, and the diameter estimate was also substantially different. This case indicates that model errors can increase when the heatmap attention is not well aligned with the true nodule region or when the visual appearance of the nodule is ambiguous. Comparison with Baseline Models To contextualize the performance of the proposed fine-tuned Florence-2 model, a comparative analysis was conducted against multiple reference baselines, including a pre-trained GPT-4 model, a domain-specific fine-tuned GPT-4 model [22], and available human reference values from prior radiological studies. The GPT-4 baselines were evaluated using the same test set and evaluation protocol to assess the effect of task-specific fine-tuning and to compare model behaviour across different vision-language architectures. Prior work on pulmonary nodule measurement has shown that manual diameter assessment is subject to interreader variability, and that reported overall variability in manual diameter measurement used here as the human reference for diameter estimation [11]. Similarly, prior observer studies have demonstrated that attenuation-related classification, such as differentiating solid from subsolid nodules, shows variability even among experienced thoracic radiologists; therefore, the reported range of radiologist performance was used to contextualize attenuation classification [25]. This comparative framework allows the proposed model to be evaluated not only against computational baselines but also against reported human variability in clinically relevant CT nodule assessment tasks. The overall comparison is summarized in Table 2. As shown in Table 2, the proposed fine-tuned Florence-2 model achieved the best overall performance across all evaluated radiological attributes. For location classification, the proposed model reached 77.18% accuracy, which was substantially higher than both the pre-trained GPT-4 baseline and the fine-tuned GPT-4 model. This indicates that the proposed model was more effective in capturing spatial information from the CT slices and map it to the correct anatomical location. Table 2: Performance comparison of the proposed fine-tuned Florence-2 model with GPT-4 baselines and available human reference values available in the literature for pulmonary nodule characterization tasks. For diameter, the reported value represents inter-reader variability in manual diameter measurement, whereas for attenuation it represents the reported range of radiologist performance for solid versus subsolid nodule classification. Attribute Pre-trained GPT-4 Fine-tuned GPT-4 Human Reference Proposed Model Location (Accuracy) 36.31% 43.68% – 77.18% Diameter (MAE) 4.74 m 3.83 m 3.2 m 2.58 m Margin (Accuracy) 38.42% 39.47% – 67.96% Attenuation (Accuracy) 37.36% 63.68% 58–77% 79.13% For diameter estimation, the proposed model achieved the lowest MAE of 2.58 m, compared with 4.74 m for pre-trained GPT-4 and 3.83 m for fine-tuned GPT-4. The proposed model also performed below the reported human interreader variability of 3.2 m. Although the human baseline represents measurement variability rather than model prediction error, this comparison shows that the proposed model produced diameter estimates within a clinically relevant range of manual measurement variation. For margin characterization, the proposed model achieved 67.96% accuracy, while the pre-trained and fine-tuned GPT-4 models achieved 38.42% and 39.47%, respectively. The limited improvement after GPT-4 fine-tuning suggests that margin classification remains difficult for GPT-4-based models. In contrast, the stronger performance of the proposed model indicates improved recognition of boundary-related visual features. For attenuation classification, the proposed model achieved 79.13% accuracy, exceeding both GPT-4 baselines and the reported human baseline range of 58 – 77%. This result is notable because attenuation classification is visually challenging, particularly when solid and subsolid appearances overlap in a single 2D CT slice. Overall, the results show that the proposed fine-tuned Florence-2 model provides stronger and more consistent performance than the GPT-4 baselines, while also reaching a level comparable to available human-reference values for diameter and attenuation assessment. Slice Selection Sensitivity Analysis Slice-selection sensitivity was evaluated by shifting the input up to three slices in both directions from the expert-annotated reference slice and measuring the resulting changes in model performance. Figure 4 shows that all four prediction tasks achieved their best performance when the expert-annotated slice was used as input. As the slice moved away from this reference position, model performance generally declined, with the poorest results occurring at an offset of three slices. Anatomical location showed the greatest sensitivity, with accuracy decreasing from 78% to 47% at both −3-3 and +3+3 slices. Margin classification accuracy decreased from 68% to 48% at −3-3 slices, while attenuation was comparatively more robust, declining from 79% to 61% at the same offset. Diameter estimation also became less accurate, with the MAE increasing from 2.58 m on the reference slice to 4.60 m at −3-3 slices. These findings indicate that the model performs most reliably when the input contains the most representative nodule appearance, while larger slice offsets reduce the visual information required for accurate characterization. Figure 4: Performance of the finetuned Florence-2 model across different input slice offsets relative to the expert annotated reference slice. External Generalization Assessment External generalization was examined using an independent veterinary CT dataset collected from Ontario Veterinary College. The dataset contained 16 CT slices, each with one pulmonary nodule and annotations for anatomical location, diameter, margin characteristics, and attenuation type. The same prompt structure and output categories used for the NLST dataset were applied, producing 64 external image-question-answer triplets. These images were not used during training or validation, and no additional finetuning was performed. Therefore, the experiment served as a preliminary domain-shift assessment of the Stage 1 Florence-2 model rather than a full external clinical validation. The model achieved 43.75% accuracy for anatomical location, 87.5% accuracy for both margin characteristics and attenuation type, and an MAE of 3.75 m for diameter estimation. The lower location accuracy likely reflects species-related differences in lung anatomy and the use of human lobar output categories for veterinary images. Margin and attenuation appeared more transferable across the two imaging domains, although their high accuracies may have been influenced by the larger proportion of smooth-margin and soft-tissue cases in the small external dataset. Overall, the results provide preliminary evidence that some learned nodule features can transfer to an independent CT source, while also showing that broader generalization claims require larger and more diverse external datasets. Stage 2 Clinical Interpretation Performance Stage 2 was evaluated by expert human assessment across 88 generated outputs, including nodule description, follow-up recommendation, and longitudinal temporal analysis. The evaluation assessed accuracy, clinical relevance, completeness, and clinical safety. Overall Human Evaluation The Stage 2 model showed strong overall performance in the expert evaluation. As shown in Figure 5, the model achieved 93.8% accuracy, 98.6% completeness, 76.2% clinical relevance, and an overall human score of 89.5%. These results indicate that the model was generally able to generate clinically meaningful outputs that were consistent with the structured input attributes. The highest score was observed for completeness, suggesting that the model included most of the required clinical information in its responses. Accuracy was also high, showing that most outputs correctly reflected the given nodule characteristics. Figure 5: Human evaluation results for Stage 2 outputs across nodule description, follow-up recommendation, longitudinal temporal analysis, and combined average performance. Clinical relevance received a comparatively lower score than accuracy and completeness. This suggests that, although the generated outputs were usually correct and complete, some responses required stronger clinical prioritization or more precise recommendation language. The safety pass rate was 77.3%, which indicates that most outputs were clinically acceptable, but a subset still required expert review before use in a clinical setting. Therefore, Stage 2 should be interpreted as an assistive clinical reasoning component rather than an autonomous decision-making system. Performance Across Clinical Interpretation Tasks Task-wise evaluation showed that model performance varied across the three clinical interpretation tasks. As shown in Figure 5, nodule description achieved 100.0% accuracy, 100.0% completeness, 71.1% relevance, and a 90.4% overall score. Follow-up recommendation achieved the highest overall task score, with 99.3% accuracy, 100.0% completeness, 78.7% relevance, and a 92.7% overall score. Longitudinal temporal analysis was more challenging, with 83.2% accuracy, 96.1% completeness, 78.1% relevance, and an 85.8% overall score. These results show that nodule description was the most stable task in terms of accuracy and completeness. This is expected because nodule description mainly requires direct conversion of structured attributes, such as location, diameter, margin, and attenuation, into descriptive clinical text. Follow-up recommendation also performed well, which suggests that the model could use the structured nodule attributes to generate reasonable management-oriented outputs. However, this task still requires caution because follow-up recommendation has direct clinical implications and depends on guideline interpretation. Longitudinal temporal analysis showed the lowest accuracy and overall score among the three tasks. This task is more complex because the model must compare findings across multiple screening rounds, identify interval changes, and summarize temporal progression. Therefore, errors in this task may arise from difficulty in tracking multiple nodules, interpreting changes over time, or assigning appropriate clinical significance to evolving nodule features. Overall, the task-wise results suggest that Stage 2 performs well when the input structure is simple and direct, but performance becomes more variable when the task requires higher-level temporal reasoning. Summary of Stage-2 Performance Evaluation The Stage 2 evaluation demonstrates that the Zephyr-based clinical reasoning component can generate accurate and complete outputs when it is grounded in structured radiological attributes. The model performed especially well for nodule description and follow-up recommendation, while longitudinal temporal analysis remained more difficult. The lower clinical relevance and safety pass rate also show that expert review remains necessary, especially for outputs involving follow-up planning or temporal disease assessment. These findings support the use of Stage 2 as a structured clinical interpretation aid, but not as a replacement for radiologist judgment. Discussion This study developed and evaluated FZ-VLM, a two-stage Vision Language Model framework for structured pulmonary nodule characterization and clinical decision support. The main finding is that separating visual attribute extraction from language-based reasoning provides a traceable connection between CT image evidence and clinical interpretation. Florence-2 extracted multiple attributes within a unified model, while Zephyr-7B used these attributes for downstream clinical interpretation. This design differs from task specific approaches that require separate models for individual radiological attributes and from direct image-to-report systems in which the source of generated clinical statements may be difficult to verify. Stage 1 demonstrated that a general-purpose VLM could be adapted for specialized lung CT interpretation. The model achieved 77.18% accuracy for anatomical location and for diameter estimation achieved a mean absolute error of 2.58 m, with a median absolute error of 1 m and 72.3% of predictions falling within ± 2 m of the reference measurement. These findings indicate that most estimates remained close to the annotated values, although small errors may still be clinically important when the measured diameter is near a management threshold. Margin and attenuation assessment depended on subtle visual differences. As illustrated in Figure 6, poorly defined and spiculated margins can share irregular boundary features, while mixed attenuation contains both ground-glass and soft-tissue components. These overlapping appearances reduce the distinction between classes and may also contribute to disagreement during expert assessment. The observed errors therefore likely reflect both model limitations and the inherent ambiguity of these radiological features. Figure 6: Representative pulmonary nodule margin characteristics and attenuation types on axial CT images. The upper row shows different margin characteristics, while the lower row shows various attenuation types. The examples illustrate the visual overlap between margin characteristics and attenuation types. The attention heatmaps provide an important interpretability strength of the proposed framework. In addition to generating radiological attribute predictions, the model highlights the image regions that contributed most strongly to its decisions. Correct predictions were generally associated with attention around the nodule or the surrounding lung region, indicating that the model used clinically relevant visual information. This feature reduces the black box nature of the model and provides experts with visual evidence that can be reviewed alongside the predicted attributes. Therefore, the heatmaps improve the transparency of Stage 1 and support the use of the framework as an assistive tool in which the expert can examine both the prediction and decision making process of the model. The slice-sensitivity analysis identified an important practical limitation of the two-dimensional approach. Performance was highest on the expert-selected reference slice and generally decreased as the input moved farther from this position. Changes in slice position can alter the visible nodule size, boundary, attenuation pattern, and anatomical context. The results therefore support the current use of representative key slices, while also showing that automated slice selection or volumetric analysis will be needed for a more independent clinical workflow. The baseline comparison provides strong evidence for the value of task specific adaptation. The finetuned Florence-2 model outperformed both the pretrained and finetuned GPT-4 baselines across all evaluated radiological attributes. More importantly, the proposed model also exceeded the available human reference values for attenuation classification and achieved a lower diameter error than the reported inter-reader variability. In contrast, both of the GPT-4 baselines remained below the reported human range for attenuation and showed higher diameter errors. These findings indicate that the proposed model not only improved over general purpose VLM baselines, but also reached or exceeded available human reference performance for clinically important nodule characteristics which was not achieved by the evaluated GPT baselines. The veterinary CT experiment provided a preliminary assessment under substantial domain shift. Margin and attenuation performance remained high, whereas anatomical location accuracy decreased considerably, likely because location depends on species-specific lung anatomy and lobar organization. The findings suggest that some visual characteristics may transfer beyond the NLST data. However, the small sample size, different anatomy, and uneven class distribution prevent broader generalization and do not represent external human clinical validation. Stage 2 results showed that structured attributes can provide an effective input for clinical language generation. Expert evaluation produced high accuracy, completeness, and overall scores, indicating that Zephyr-7B generally preserved the supplied nodule information when generating descriptions and recommendations. The high completeness score is particularly relevant because the hierarchical integration layer was designed to organize information at the nodule, study-year, and patient levels before clinical interpretation. This structured organization appears to have supported consistent generation across the three tasks. Performance differed across the three tasks because they required different levels of reasoning. Nodule description was the most stable task because it mainly required conversion of structured attributes into clinical text. Follow-up recommendation also performed well, indicating that the model could combine multiple findings into management-oriented outputs. Longitudinal analysis was more difficult because it required tracking nodules across screening years and interpreting interval changes. Structured inputs therefore improved grounding but did not fully resolve complex temporal reasoning. Clinical relevance and safety were lower than accuracy and completeness for different reasons. Nodule descriptions were generally accurate and complete, but because the task was mainly descriptive, some outputs added limited clinical value beyond restating the supplied attributes. Safety concerns were more evident in follow-up recommendations, where inappropriate timing could affect patient management. These findings show that expert review remains necessary, particularly for follow-up and longitudinal outputs, and that future work should improve guideline integration, risk calibration, and the clinical usefulness of generated responses. Although the proposed FZ-VLM framework demonstrated promising performance in structured pulmonary nodule characterization and clinical interpretation, several limitations remain. First, the framework relies on expert-selected two-dimensional slices rather than complete CT volumes, which limits spatial information and creates dependence on accurate key-slice selection. Second, it operates only in a post-detection setting and assumes that the pulmonary nodule has already been identified. Third, the external assessment used a small veterinary dataset rather than an independent human cohort. Fourth, the safety pass rate of 77.3% indicates that some Stage 2 outputs, particularly follow-up recommendations, may still be clinically inappropriate and require expert review. Finally, Zephyr-7B was used through zero-shot inference and may not consistently follow institution-specific reporting practices or management guidelines. Overall, this study shows that FZ-VLM can link structured pulmonary nodule characterization with clinically meaningful language generation within a unified two-stage framework. The strong performance of Florence-2 and the expert-evaluated outputs from Zephyr-7B support the potential of this approach as an assistive tool for post-detection interpretation. However, the remaining limitations in safety, temporal reasoning, and external validation confirm that radiologist oversight is still essential. Future work should therefore focus on automated slice selection, three-dimensional CT analysis, validation using independent human datasets, and closer integration with established clinical guidelines. References [1] American College of Radiology Committee on Lung-RADS (2022) Lung-rads assessment categories 2022. Note: Accessed: 2023-01-01 External Links: Link Cited by: Introduction, Introduction. [2] Y. Brima and M. Atemkeng (2024) Saliency-driven explainable deep learning in medical imaging: bridging visual explainability and statistical quantitative analysis. BioData Mining 17, p. 18. External Links: Document Cited by: Introduction. [3] T. Chang, Q. Gou, L. Zhao, T. Zhou, H. Chen, D. Yang, H. Ju, K. E. Smith, C. Sun, J. Pan, Y. Huang, X. He, X. Zhang, D. Xu, J. Xu, J. Bian, and A. Chen (2025) From image to report: automating lung cancer screening interpretation and reporting with vision-language models. Journal of Biomedical Informatics 171, p. 104931. External Links: ISSN 1532-0464, Document, Link Cited by: Introduction. [4] H. Chen, R. Shukla, R. Wu, S. Yang, D. Duong-Tran, D. M. H. Nguyen, M. Niepert, C. Beeche, J. Gee, J. Duda, R. Sharma, C. Davatzikos, W. Witschey, B. Hou, and L. Shen (2025) Adapting vision–language models for 3d ct/mri understanding on pmbb via slice selection and explanation analysis. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, p. 2273–2282. Cited by: Introduction. [5] N. Dutta, K. Bose, E. Syailendra, L. Chu, and P. Gupta (2026) Vision-language models in diagnostic imaging: review of technical advances, clinical validation, and practical deployment. International Journal of Medical Informatics 208, p. 106227. External Links: ISSN 1386-5056, Document, Link Cited by: Introduction, Introduction. [6] P. Dutta, J. Manokaran, R. Mittal, and E. Ukwatta (2026) Structured clinical interpretation of lung computed tomography: a post-detection application of vision language models. In Medical Imaging 2026: Imaging Informatics, Proceedings of SPIE, Vol. 13930, p. 1393009. External Links: Document Cited by: Introduction, Introduction, Introduction. [7] J. R. Ferreira, M. C. Oliveira, and P. M. de Azevedo-Marques (2018) Characterization of pulmonary nodules based on features of margin sharpness and texture. Journal of Digital Imaging 31 (4), p. 451–463. External Links: ISSN 1618-727X, Document, Link Cited by: Introduction, Introduction. [8] M. C. B. Godoy and D. P. Naidich (2009) Subsolid pulmonary nodules and the spectrum of peripheral adenocarcinomas of the lung: recommended interim guidelines for assessment and management. Radiology 253 (3), p. 606–622. External Links: Document Cited by: Introduction. [9] M. K. Gould, J. Donington, W. R. Lynch, P. J. Mazzone, D. E. Midthun, D. P. Naidich, and R. S. Wiener (2013) Evaluation of individuals with pulmonary nodules: when is it lung cancer? diagnosis and management of lung cancer, 3rd ed: american college of chest physicians evidence-based clinical practice guidelines. Chest 143 (5 Suppl), p. e93S–e120S. External Links: Document Cited by: Introduction. [10] P. Hager, F. Jungmann, R. Holland, K. Bhagat, I. Hubrecht, M. Knauer, J. Vielhauer, M. Makowski, R. Braren, G. Kaissis, and D. Rueckert (2024) Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature Medicine 30 (9), p. 2613–2622. External Links: Document, Link Cited by: Introduction. [11] D. Han, M. A. Heuvelmans, R. Vliegenthart, M. Rook, M. D. Dorrius, G. J. de Jonge, J. E. Walter, P. M. A. van Ooijen, H. J. de Koning, and M. Oudkerk (2018) Influence of lung nodule margin on volume- and diameter-based reader variability in CT lung cancer screening. British Journal of Radiology 91 (1090), p. 20170405. External Links: Document Cited by: Introduction, Comparison with Baseline Models. [12] S. Khademi, M. Shabanpour, R. Taleei, A. Oikonomou, and A. Mohammadi (2025) AutoRad-Lung: a radiomic-guided prompting autoregressive vision-language model for lung nodule malignancy prediction. Note: arXiv preprint arXiv:2503.20662 External Links: 2503.20662, Document Cited by: Introduction, Introduction. [13] P. Komorowski, H. Baniecki, and P. Biecek (2023) Towards evaluating explanations of vision transformers for medical imaging. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vol. , p. 3726–3732. External Links: Document Cited by: Introduction. [14] C. Y. Lin, S. M. Guo, J. J. J. Lien, et al. (2024) Combined model integrating deep learning, radiomics, and clinical data to classify lung nodules at chest ct. La Radiologia Medica 129, p. 56–69. External Links: Document, Link Cited by: Introduction. [15] H. Liu, C. Li, Q. Wu, and Y. J. Lee (2023) Visual instruction tuning. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, p. 34892–34916. External Links: Link Cited by: Introduction. [16] L. Liu, N. Wang, D. Liu, X. Yang, X. Gao, and T. Liu (2024) Towards specific domain prompt learning via improved text label optimization. IEEE Transactions on Multimedia 26 (), p. 10805–10815. External Links: Document Cited by: Introduction. [17] S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, and et al. (2025) Grounding dino: marrying dino with grounded pre-training for open-set object detection. In Computer Vision – ECCV 2024, p. 38–55. External Links: Document Cited by: Introduction. [18] Y. Liu, Y. Balagurunathan, T. Atwater, S. Antic, Q. Li, R. C. Walker, G. T. Smith, P. P. Massion, M. B. Schabath, and R. J. Gillies (2017) Radiological image traits predictive of cancer status in pulmonary nodules. Clinical Cancer Research 23 (6), p. 1442–1449. External Links: ISSN 1078-0432, Document Cited by: Introduction. [19] J. Manokaran, R. Mittal, and E. Ukwatta (2024) Pulmonary nodule detection in low dose computed tomography using a medical-to-medical transfer learning approach. Journal of Medical Imaging 11 (4), p. 044502. External Links: Document Cited by: Introduction. [20] D. McDuff, M. Schaekermann, T. Tu, et al. (2025) Towards accurate differential diagnosis with large language models. Nature 642, p. 451–457. External Links: Document, Link Cited by: Introduction. [21] National Cancer Institute (2026) Cancer stat facts: lung and bronchus cancer. Note: https://seer.cancer.gov/statfacts/html/lungb.htmlSurveillance, Epidemiology, and End Results Program. Accessed: 2026-07-30 Cited by: Introduction. [22] OpenAI (2024) Fine-tuning now available for gpt‑4o. Note: https://openai.com/index/gpt-4o-fine-tuning/Accessed: 30 July 2026 Cited by: Comparison with Baseline Models. [23] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever (2021) Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 139. Cited by: Introduction. [24] A. Rao and O. Aalami (2023) Towards improving the visual explainability of artificial intelligence in the clinical setting. BMC Digital Health 1, p. 23. External Links: Document Cited by: Introduction. [25] C. A. Ridge, A. Yildirim, P. M. Boiselle, T. Franquet, C. M. Schaefer-Prokop, D. Tack, P. A. Gevenois, and A. A. Bankier (2016) Differentiating between subsolid and solid pulmonary nodules at ct: inter- and intraobserver agreement between experienced thoracic radiologists. Radiology 278 (3), p. 888–896. Note: PMID: 26458208 External Links: Document, Link, https://doi.org/10.1148/radiol.2015150714 Cited by: Introduction, Comparison with Baseline Models. [26] A. A. A. Setio, A. Traverso, T. de Bel, M. Berens, C. van den Bogaard, P. Cerello, H. Chen, Q. Dou, M. Fantacci, B. Geurts, and et al. (2017) Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: the luna16 challenge. Medical Image Analysis 42, p. 1–13. Cited by: Introduction. [27] N. Shi, Z. Liu, Z. Wan, G. Yang, Y. Shi, P. Wang, and X. Liu (2026) A vision-language model-based approach for lung cancer diagnosis using lossless 3D CT images: evaluation of GPT-4.1 and GPT-4o for patient-level malignancy assessment. Health Information Science and Systems 14 (1), p. 16. External Links: Document Cited by: Introduction, Introduction, Introduction. [28] H. Sung, J. Ferlay, R. L. Siegel, M. Laversanne, I. Soerjomataram, A. Jemal, and F. Bray (2021) Global cancer statistics 2020: globocan estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: A Cancer Journal for Clinicians 71 (3), p. 209–249. Note: doi: https://doi.org/10.3322/caac.21660 Cited by: Introduction. [29] N. L. S. T. R. Team (2011) The national lung screening trial: overview and study design. Radiology 258 (1), p. 243–253. Note: PMID: 21045183 External Links: Document, Link, https://doi.org/10.1148/radiol.10091808 Cited by: Introduction, Introduction, Introduction. [30] L. Tunstall, E. Beeching, N. Lambert, N. Rajani, K. Rasul, Y. Belkada, S. Huang, L. von Werra, C. Fourrier, N. Habib, N. Sarrazin, O. Sanseviero, A. M. Rush, and T. Wolf (2023) Zephyr: direct distillation of lm alignment. arXiv preprint arXiv:2310.16944. External Links: 2310.16944, Link Cited by: Introduction. [31] B. P. Veasey and A. A. Amini (2024) Parameter-efficient fine-tuning of dinov2 vision transformers for lung nodule classification. In 2024 IEEE International Symposium on Biomedical Imaging (ISBI), Vol. , p. 1–5. External Links: Document Cited by: Introduction. [32] B. Xiao, H. Wu, W. Xu, X. D. an Houdong Hu, Y. Lu, M. Zeng, C. Liu, and L. Yuan (2023) Florence-2: advancing a unified representation for a variety of vision tasks. External Links: 2311.06242, Link Cited by: Introduction, Introduction. [33] Y. Xie, Y. Xia, J. Zhang, Y. Song, D. Feng, M. Fulham, and W. Cai (2019) Knowledge-based collaborative deep learning for benign–malignant lung nodule classification on chest CT. IEEE Transactions on Medical Imaging 38 (4), p. 991–1004. External Links: Document Cited by: Introduction. [34] R. Yang, T. F. Tan, W. Lu, A. J. Thirunavukarasu, D. S. W. Ting, and N. Liu (2023) Large language models in health care: development, applications, and challenges. Health Care Science 2 (4), p. 255–263. External Links: Document, Link, https://onlinelibrary.wiley.com/doi/pdf/10.1002/hcs2.61 Cited by: Introduction. [35] D. Zhao, J. Xi, X. Guo, J. Chai, Z. Xu, L. Li, Y. Xue, Q. Sun, Y. Zheng, and S. Liu (2026) Graphicalized vision-language modeling for comprehensive lung nodule analysis and risk stratification. npj Digital Medicine 9 (1), p. 442. External Links: Document Cited by: Introduction, Introduction. [36] C. Zhou, H. Chan, A. Chughtai, L. M. Hadjiiski, E. A. Kazerooni, and J. Wei (2020) Pathologic categorization of lung nodules: radiomic descriptors of ct attenuation distribution patterns of solid and subsolid nodules in low-dose ct. European Journal of Radiology 129, p. 109106. External Links: ISSN 0720-048X, Document, Link Cited by: Introduction, Introduction. Acknowledgements This work was supported by research funding to the AI-Driven Medical Imaging & Diagnostics Lab at the University of Guelph, including support from an Natural Sciences and Engineering Research Council of Canada (NSERC) Discovery Grant. Data used in this study were obtained from the National Lung Screening Trial (NLST), a project supported by the U.S. National Cancer Institute. This research also used high-performance computing resources provided by the Digital Research Alliance of Canada. AI-assisted writing tools were used solely and strictly for language and grammar refinement during manuscript preparation. All scientific content, experimental design, data analysis, interpretation, and conclusions were developed and verified by the authors, who take full responsibility for the accuracy and integrity of the work. Author contributions statement P.D. conducted the study design, dataset preparation, image preprocessing, model development, implementation, experiments, evaluation, analysis, and manuscript preparation. J.M. provided research guidance, supported dataset access and interpretation, and contributed to the review of the methodology and results. R.M. provided clinical and radiological guidance, evaluated the model generated clinical outputs, and reviewed representative CT cases. R.A. provided and managed the veterinary CT dataset and corresponding annotations. E.U. supervised the study, contributed to the research design and interpretation of the findings, and reviewed and revised the manuscript. All authors reviewed and approved the final manuscript. Additional information Data availability statement The NLST data used in this study were obtained from the National Lung Screening Trial (NLST) through the National Cancer Institute Cancer Data Access System under project number NLST-871. The data were provided under a formal Data Transfer Agreement and cannot be publicly redistributed. Researchers may apply independently for access through the National Cancer Institute Cancer Data Access System, subject to the applicable data access requirements and institutional approvals. The source code used to develop and evaluate the FZ-VLM framework is publicly available at https://github.com/PramitDutta1999/Florence-Zephyr-Vision-Language-Model-FZ-VLM-Framework. The repository does not include restricted NLST images, patient-level information, or trained model weights. Competing interests The authors declare no competing interests.