Paper deep dive
Dual-Modal Lung Cancer AI: Interpretable Radiology and Microscopy with Clinical Risk Integration
Baramee Sukumal, Aueaphum Aueawatthanaphisut
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/27/2026, 7:19:54 PM
Summary
This paper proposes a dual-modal artificial intelligence framework for lung cancer diagnosis that integrates CT radiology and H&E histopathology. The system uses two separate convolutional neural networks (CT-CNN and H&E-CNN) and a clinical metadata encoder (incorporating age, sex, and smoking history) to perform weighted decision-level fusion. The model classifies lung cancer into five categories: adenocarcinoma, squamous cell carcinoma, large cell carcinoma, small cell lung cancer (SCLC), and normal tissue. The framework emphasizes interpretability by utilizing XAI techniques like Grad-CAM++, which demonstrated high faithfulness and localization accuracy, bridging the gap between non-invasive screening and definitive pathological diagnosis.
Entities (14)
Relation Signals (5)
Baramee Sukumal → affiliatedwith → Hatyaiwittayalai School
confidence 100% · Baramee Sukumal 1 ... 1 Hatyaiwittayalai School
CT Radiology → usedfordiagnosisof → Lung Cancer
confidence 100% · integrates CT radiology with hematoxylin and eosin (H&E) histopathology for lung cancer diagnosis
H&E Histopathology → usedfordiagnosisof → Lung Cancer
confidence 100% · integrates CT radiology with hematoxylin and eosin (H&E) histopathology for lung cancer diagnosis
Grad-CAM++ → providesinterpretabilityfor → Dual-Modal AI Framework
confidence 95% · Grad-CAM++ achieved the highest faithfulness and localization accuracy, demonstrating strong correspondence with expert-annotated tumor regions.
LIDC-IDRI → providesdatafor → CT Radiology
confidence 90% · LIDC–IDRI have facilitated the development of computer-aided diagnostic systems based on CT imaging
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Lung cancer remains one of the leading causes of cancer-related mortality worldwide. Conventional computed tomography (CT) imaging, while essential for detection and staging, has limitations in distinguishing benign from malignant lesions and providing interpretable diagnostic insights. To address this challenge, this study proposes a dual-modal artificial intelligence framework that integrates CT radiology with hematoxylin and eosin (H&E) histopathology for lung cancer diagnosis and subtype classification. The system employs convolutional neural networks to extract radiologic and histopathologic features and incorporates clinical metadata to improve robustness. Predictions from both modalities are fused using a weighted decision-level integration mechanism to classify adenocarcinoma, squamous cell carcinoma, large cell carcinoma, small cell lung cancer, and normal tissue. Explainable AI techniques including Grad-CAM, Grad-CAM++, Integrated Gradients, Occlusion, Saliency Maps, and SmoothGrad are applied to provide visual interpretability. Experimental results show strong performance with accuracy up to 0.87, AUROC above 0.97, and macro F1-score of 0.88. Grad-CAM++ achieved the highest faithfulness and localization accuracy, demonstrating strong correspondence with expert-annotated tumor regions. These results indicate that multimodal fusion of radiology and histopathology can improve diagnostic performance while maintaining model transparency, suggesting potential for future clinical decision support systems in precision oncology.
Tags
Links
- Source: https://arxiv.org/abs/2604.16104v1
- Canonical: https://arxiv.org/abs/2604.16104v1
Trouble viewing inline? Open PDF directly →
Full Text
33,205 characters extracted from source content.
Expand or collapse full text
Dual-Modal Lung Cancer AI: Interpretable Radiology and Microscopy with Clinical Risk Integration Baramee Sukumal 1 and Aueaphum Aueawatthanaphisut 2 1 Hatyaiwittayalai School, Hat Yai, Songkhla, Thailand, 90110 Tarohatsune28@gmail.com 2 School of Information, Computer, and Communication Technology, Sirindhorn International Institute of Technology, Thammasat University, Pathum Thani, Thailand aueawatth.aue@gmail.com Abstract. Lung cancer remains one of the leading causes of cancer- related morbidity and mortality worldwide. Conventional computed to- mography (CT) imaging, while essential for detection and staging, faces limitations in distinguishing benign from malignant lesions and in pro- viding explainable diagnostic insights. To address these challenges, this study developed and validated a dual-modal artificial intelligence (AI) framework that integrates CT radiology with hematoxylin and eosin (H&E) microscopy for lung cancer diagnosis and subtype classification. The model employed a convolutional neural network (CNN) architecture to extract and combine radiologic and histopathologic features, incor- porating clinical metadata to enhance diagnostic robustness. Both CT- based and microscopy-based CNNs were trained independently and sub- sequently fused through weighted decision-level integration to achieve unified predictions across categories including adenocarcinoma, squa- mous cell carcinoma, large cell carcinoma, small cell lung cancer (SCLC), and normal tissue. Explainable AI (XAI) techniques—Grad-CAM, Grad- CAM++, Integrated Gradients, Occlusion, Saliency Maps, and Smooth- Grad—were implemented to provide visual interpretability and assess alignment with clinically meaningful regions. Quantitative evaluation demonstrated strong model performance with accuracy ranging from 0.84 to 0.87, precision between 0.87 and 0.89, recall between 0.84 and 0.87, F1-scores ranging from 0.84 to 0.88, and AUROC values exceeding 0.94. Grad-CAM++ achieved the highest faithfulness (insertion AUC≈ 0.83 for H&E, 0.81 for CT) and localization accuracy (IOU≈ 0.65 for H&E, 0.81 for CT), confirming strong correspondence with expert-annotated tumor regions. In conclusion, the proposed dual-modal AI framework demonstrates high diagnostic accuracy, interpretability, and clinical rele- vance. By fusing radiology and microscopy modalities within an explain- able deep learning system, this approach enhances both precision and transparency, showing strong potential for integration into real-world precision oncology and clinical decision support systems. arXiv:2604.16104v1 [eess.IV] 17 Apr 2026 2Baramee Sukumal and Aueaphum Aueawatthanaphisut Keywords: Lung cancer diagnosis· Dual-modal AI fusion· CT radiology imaging· H&E histopathology· Explainable AI (Grad-CAM++). 1 Introduction Lung cancer remains one of the leading causes of cancer-related mortality world- wide, accounting for a substantial proportion of global cancer deaths each year. Early detection and accurate subtype classification are critical for improving pa- tient outcomes and guiding personalized treatment strategies. In clinical prac- tice, thoracic computed tomography (CT) imaging is widely used as the primary modality for lung cancer screening and staging due to its ability to visualize pul- monary nodules and structural abnormalities. Large-scale public datasets such as the Lung Image Database Consortium and Image Database Resource Ini- tiative (LIDC–IDRI) have facilitated the development of computer-aided diag- nostic systems based on CT imaging [1]. In addition, repositories such as The Cancer Imaging Archive (TCIA) provide curated imaging datasets that support radiomics and machine learning research in oncology [3]. Radiomic analysis has demonstrated that quantitative imaging features extracted from CT scans can capture tumor phenotypes and potentially predict clinical outcomes [2]. Fig. 1: Example chest CT images demonstrating a suspected lung tumor region across axial, coronal, and sagittal planes. CT imaging is commonly used for lung cancer screening and detection. Despite these advances, CT imaging alone presents several limitations. Ra- diological features may be insufficient to reliably differentiate benign nodules from malignant lesions, particularly in early-stage disease. Moreover, CT imaging lacks the ability to directly capture cellular-level morphology, which is essential for accurate pathological classification. Consequently, histopathological exami- nation of biopsy tissue stained with hematoxylin and eosin (H&E) remains the gold standard for definitive diagnosis and subtype identification of lung cancer. Explainable Dual-Modal AI for Lung Cancer Diagnosis3 Fig. 2: Representative histopathology image of lung tumor tissue stained with hematoxylin and eosin (H&E). Histopathological analysis enables visualization of cellular structures and tumor morphology. Histopathological analysis provides detailed insight into cellular morphology and tissue architecture, allowing pathologists to distinguish between major lung cancer subtypes such as adenocarcinoma, squamous cell carcinoma, large cell carcinoma, and small cell lung cancer (SCLC). Public datasets such as the Lung and Colon Cancer Histopathological Image Dataset (LC25000) have enabled the development of deep learning approaches for automated histopathological classification [8]. However, manual microscopic analysis remains time-consuming and subject to inter-observer variability. Recent advances in deep learning have enabled the development of automated diagnostic systems capable of analyzing both radiological and histopathological images. Convolutional neural networks (CNNs), including architectures such as EfficientNet, have demonstrated strong performance in medical image classifi- cation tasks [9]. Additionally, self-configuring frameworks such as nnU-Net have shown the effectiveness of deep learning pipelines for biomedical image segmen- tation and analysis [10]. However, many existing approaches rely on a single imaging modality, which may limit diagnostic robustness and generalization. Another critical challenge in medical AI systems is the lack of interpretabil- ity. Deep neural networks are often regarded as “black-box” models, making it difficult for clinicians to understand the reasoning behind predictions. Explain- able artificial intelligence (XAI) techniques have been proposed to address this issue by highlighting image regions that contribute to model decisions. Meth- ods such as Gradient-weighted Class Activation Mapping (Grad-CAM) [11] and 4Baramee Sukumal and Aueaphum Aueawatthanaphisut Fig. 3: Major pathological subtypes of lung cancer, including adenocarcinoma, squamous cell carcinoma, large cell carcinoma, and small cell lung cancer (SCLC). its improved variant Grad-CAM++ [12] have been widely applied to visualize discriminative regions in medical imaging models. To address these limitations, this study proposes a dual-modal explainable artificial intelligence framework that integrates radiological CT imaging and histopathological H&E microscopy for lung cancer diagnosis. The proposed ap- proach combines modality-specific convolutional neural networks with a weighted decision-level fusion strategy to generate unified predictions across multiple lung cancer subtypes. Clinical metadata is incorporated to dynamically adjust the contribution of each modality during prediction. Although both radiological imaging and histopathological examination pro- vide valuable diagnostic information, each modality presents inherent limitations when used independently. Computed tomography (CT) imaging enables non- invasive detection of pulmonary nodules and structural abnormalities; however, radiological features alone may not provide sufficient resolution to reliably differ- entiate between histological subtypes of lung cancer. Conversely, histopatholog- ical analysis using hematoxylin and eosin (H&E) stained tissue slides provides detailed cellular-level information that is essential for definitive diagnosis, but the process typically requires invasive biopsy procedures and time-consuming manual interpretation by expert pathologists. Furthermore, diagnostic systems based on a single imaging modality may suf- fer from limited robustness and reduced generalization due to modality-specific Explainable Dual-Modal AI for Lung Cancer Diagnosis5 Fig. 4: Independent diagnostic pipelines for lung cancer classification using ra- diological CT imaging and histopathological H&E microscopy. The left panel illustrates the CT-based classification system, where a convolutional neural net- work analyzes thoracic CT scans and highlights suspicious pulmonary nodules using explainable AI heatmaps. The right panel shows the H&E histopathology- based classification system, where microscopic tissue structures are analyzed to identify tumor morphology and cellular patterns. Each modality independently produces class probabilities for lung cancer subtypes including adenocarcinoma, squamous cell carcinoma, small cell lung cancer (SCLC), and normal tissue. biases and variability in imaging conditions. As a result, complementary diag- nostic information available from radiology and pathology is often underutilized when these modalities are analyzed independently. This limitation highlights the need for integrative artificial intelligence frameworks capable of combining het- erogeneous medical data sources to improve diagnostic reliability, interpretabil- ity, and clinical decision support. Furthermore, explainable AI techniques are integrated to provide visual ex- planations that align with clinically meaningful tumor regions. Methods such as Grad-CAM++, integrated gradients, occlusion analysis, and saliency mapping are employed to evaluate model interpretability and localization accuracy. Model performance is evaluated using standard metrics including accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (AUROC), while statistical significance is assessed using established techniques such as DeLong’s test for correlated ROC curves [16] and calibration metrics including Brier score [19]. The main contributions of this study can be summarized as follows: – A dual-modal deep learning framework that integrates CT radiology and H&E histopathology for lung cancer subtype classification. – A weighted decision-level fusion mechanism incorporating clinical metadata to improve diagnostic robustness. – An explainable AI module that provides interpretable visual explanations aligned with tumor regions. 6Baramee Sukumal and Aueaphum Aueawatthanaphisut – A comprehensive evaluation using publicly available datasets and statistical validation of model performance. Overall, the proposed framework aims to bridge the diagnostic gap between non-invasive radiological screening and definitive pathological diagnosis by com- bining complementary imaging modalities within an interpretable deep learning system. 2 Related Work Recent advances in artificial intelligence and deep learning have significantly influenced the development of automated diagnostic systems in medical imag- ing. In the context of lung cancer detection and classification, both radiological imaging and histopathological analysis have been extensively investigated using machine learning and deep neural network models. Existing studies can gener- ally be categorized into two main research directions: radiology-based diagnostic systems and histopathology-based diagnostic systems. 2.1 Radiology-Based Lung Cancer Detection Radiological imaging, particularly computed tomography (CT), has long been regarded as a fundamental modality for lung cancer screening and early de- tection. Large-scale datasets such as the Lung Image Database Consortium and Image Database Resource Initiative (LIDC-IDRI) have enabled the development of computer-aided diagnostic systems for pulmonary nodule analysis [1]. Further- more, the Cancer Imaging Archive (TCIA) has played a crucial role in facilitat- ing open-access medical imaging research by providing curated CT datasets and associated clinical information [3]. Traditional approaches in radiology-based cancer detection relied on hand- crafted radiomic features extracted from CT images. These features were de- signed to capture tumor shape, texture, and intensity characteristics. Radiomics studies have demonstrated that quantitative imaging features can reveal tumor phenotypes and potentially predict clinical outcomes and treatment responses [2]. However, handcrafted features are often limited in their ability to capture complex spatial patterns within medical images. More recently, deep convolutional neural networks (CNNs) have been widely adopted for automatic feature extraction and classification tasks in medical imag- ing. Modern CNN architectures such as EfficientNet have demonstrated strong performance in large-scale image recognition tasks and have been successfully adapted for medical imaging applications [9]. In addition, automated deep learn- ing frameworks such as nnU-Net have shown promising results in biomedical image segmentation and analysis by adapting network configurations to spe- cific datasets [10]. Despite these advancements, radiology-based approaches alone Explainable Dual-Modal AI for Lung Cancer Diagnosis7 may still face challenges in accurately distinguishing between different histologi- cal subtypes of lung cancer due to limited cellular-level information available in CT images. 2.2 Histopathology-Based Cancer Classification and Explainable AI Histopathological examination remains the gold standard for confirming lung cancer diagnosis and identifying tumor subtypes. Microscopic examination of tissue samples stained with hematoxylin and eosin (H&E) allows pathologists to analyze cellular morphology, tissue architecture, and tumor growth patterns. The availability of large digital pathology datasets, such as the Lung and Colon Cancer Histopathological Image Dataset (LC25000), has facilitated the applica- tion of deep learning models for automated histopathological classification [8]. Deep learning models have shown remarkable capability in analyzing histopatho- logical images by learning hierarchical features directly from raw pixel data. However, the adoption of such models in clinical settings has raised concerns regarding the interpretability of model predictions. Many deep neural networks operate as black-box systems, making it difficult for clinicians to understand the reasoning behind the model’s diagnostic decisions. To address this limitation, explainable artificial intelligence (XAI) techniques have been developed to visualize and interpret deep learning predictions. One of the most widely used methods is Gradient-weighted Class Activation Map- ping (Grad-CAM), which generates heatmaps indicating image regions that con- tribute most strongly to the model’s predictions [11]. An improved version, Grad- CAM++, has been proposed to enhance localization performance and produce more accurate visual explanations for convolutional neural networks [12]. These techniques have become increasingly important in medical imaging applications, as they allow clinicians to verify whether model decisions are consistent with clinically meaningful regions. Despite the progress achieved in both radiological and histopathological AI systems, most existing studies have focused on a single imaging modality. Conse- quently, valuable complementary information between radiology and pathology may remain underutilized. This limitation motivates the development of inte- grated diagnostic frameworks that combine multiple data modalities to improve diagnostic accuracy, interpretability, and clinical applicability. 3 Methodology 3.1 System Overview The proposed diagnostic framework integrates radiological imaging, histopatho- logical microscopy, and clinical metadata within a unified deep learning archi- tecture. As illustrated in Fig. 5, the system consists of three major components: 8Baramee Sukumal and Aueaphum Aueawatthanaphisut Fig. 5: System architecture of the proposed dual-modal diagnostic framework integrating CT radiology, H&E histopathology, and clinical metadata. Modality- specific CNN encoders extract features from each data source, and predictions are combined through a weighted decision-level fusion module to generate final subtype probabilities. (1) a CT-based convolutional neural network (CT-CNN) for radiological feature extraction, (2) a histopathology-based CNN (H&E-CNN) for microscopic tissue analysis, and (3) a clinical metadata encoder that incorporates patient-specific contextual information. Feature representations extracted from each modality are subsequently com- bined through a weighted decision-level fusion module to produce a unified pre- diction for lung cancer subtype classification. The final output represents the probability distribution across five target classes: adenocarcinoma, squamous cell carcinoma, large cell carcinoma, small cell lung cancer (SCLC), and normal tissue. 3.2 Dataset and Data Preparation To ensure reproducibility and generalizability, publicly available medical imaging datasets were utilized in this study. CT Radiology Dataset: Thoracic CT images were obtained from The Can- cer Imaging Archive (TCIA), including the LIDC-IDRI dataset and additional lung cancer collections. The combined dataset contained approximately 1,450 CT scans after quality control procedures [1,3]. CT scans were resampled to isotropic voxel spacing and normalized to Hounsfield unit ranges suitable for lung tissue analysis. Explainable Dual-Modal AI for Lung Cancer Diagnosis9 Table 1: Dataset distribution used in this study ClassCT Images H&E Slides Adenocarcinoma380240 Squamous Cell Carcinoma 320210 Large Cell Carcinoma210170 Small Cell Lung Cancer290160 Normal250160 Total1450940 Histopathology Dataset: Whole-slide histopathological images stained with hematoxylin and eosin (H&E) were obtained from TCGA-LUAD and TCGA- LUSC collections. Approximately 940 digital pathology slides were used after quality filtering. Additionally, the LC25000 dataset was incorporated to improve patch-level diversity [8]. Clinical Metadata: Clinical attributes including patient age, sex, and smok- ing history were incorporated as contextual variables to improve predictive ro- bustness. Data Splitting: The dataset was divided at the patient level to prevent data leakage: – Training set: 70% – Validation set: 10% – Test set: 20% Patient-level splitting was enforced to ensure that images from the same patient were not distributed across different subsets. 3.3 CT-CNN Branch The CT-CNN branch was designed to extract radiological features from thoracic CT scans. Each CT scan was first preprocessed using voxel resampling and in- tensity normalization. Pulmonary regions containing nodules were then cropped into image patches centered around suspicious lesions. A deep convolutional neural network backbone was employed to learn hierar- chical radiological features. In this study, EfficientNet and ResNet architectures were considered due to their strong performance in medical image classification tasks [9]. The CT encoder generates a feature representation: f CT = CN N CT (x CT ) where x CT represents the input CT image and f CT represents the extracted radiological feature vector. 10Baramee Sukumal and Aueaphum Aueawatthanaphisut The classification head subsequently produces logits: z CT = W CT f CT + b CT These logits represent modality-specific predictions for each lung cancer class. 3.4 H&E-CNN Branch The histopathology branch analyzes microscopic tissue structures extracted from H&E stained slides. Whole-slide images were divided into smaller image patches (tiles) of size 256× 256 pixels. To reduce staining variability, stain normalization was performed using the Macenko method [13]. Each tile was processed by a convolutional neural network encoder to extract histopathological features. f HE = CN N HE (x HE ) where x HE represents the histopathological image tile. Multiple Instance Learning (MIL) with attention pooling was used to aggre- gate tile-level representations into slide-level predictions. z HE = W HE f HE + b HE These logits correspond to pathology-based subtype predictions. 3.5 Clinical Metadata Encoder Clinical metadata were incorporated to provide patient-specific contextual infor- mation for adaptive modality weighting. Three clinical variables were included in the metadata vector: patient age, biological sex, and smoking history. Age was treated as a continuous variable and normalized using z-score standardization. Sex was encoded as a binary variable (0 = female, 1 = male), while smoking sta- tus was represented using a categorical encoding (0 = never smoker, 1 = former smoker, 2 = current smoker). The resulting metadata vector therefore consisted of three numerical features: x meta = [age, sex, smoking] Prior to model input, all metadata variables were standardized to zero mean and unit variance to ensure consistent scaling with image-derived features. The clinical metadata vector was processed using a multi-layer perceptron (MLP) consisting of two fully connected layers with 64 and 32 hidden units respectively. Each layer was followed by a Rectified Linear Unit (ReLU) activa- tion function and a dropout layer with a dropout probability of 0.3 to prevent overfitting. The transformation can be expressed as Explainable Dual-Modal AI for Lung Cancer Diagnosis11 f meta = M LP(x meta ) where f meta ∈R 32 represents the learned clinical context embedding. This embedding was subsequently used by the fusion module to dynamically generate modality weighting coefficients that regulate the relative contributions of the CT-CNN and H&E-CNN predictions during decision-level fusion. 3.6 Weighted Decision-Level Fusion The predictions from CT and histopathology branches were integrated through a weighted decision-level fusion mechanism. z = w CT (f meta )· z CT + w HE (f meta )· z HE where – z CT = logits from CT model – z HE = logits from histopathology model – w CT , w HE = dynamic modality weights derived from clinical metadata The final classification probabilities are obtained using the softmax function: ˆy = sof tmax(z) This fusion mechanism enables the system to adaptively prioritize radiologi- cal or pathological information depending on patient-specific context. 3.7 Explainable Artificial Intelligence (XAI) To improve interpretability, multiple explainable AI techniques were integrated into the diagnostic framework. Gradient-weighted Class Activation Mapping (Grad-CAM) and Grad-CAM++ were applied to visualize the image regions that contributed most strongly to the model predictions [11,12]. The quality of explanations was evaluated using two metrics: – Faithfulness (Insertion AUC) – Localization accuracy (Intersection over Union, IoU) Experimental results demonstrated that Grad-CAM++ achieved the highest explanation quality and localization accuracy across both imaging modalities. 12Baramee Sukumal and Aueaphum Aueawatthanaphisut 3.8 Training Strategy Model training was conducted using the TensorFlow and Keras deep learning frameworks. The AdamW optimizer was used for gradient-based optimization [22]. To improve generalization and reduce overfitting, several data augmentation techniques were applied during training, including MixUp and CutMix strategies [23,24]. Training was performed for multiple epochs with early stopping based on validation performance. 3.9 Hyperparameter Configuration The key hyperparameters used during model training are summarized in Table 2. Table 2: Training hyperparameters used in the proposed dual-modal framework ParameterValue OptimizerAdamW Learning rate1e-4 Batch size32 Epochs50 Image resolution256× 256 Dropout rate0.5 Weight decay10 −4 Fusion method Weighted decision-level fusion Loss functionCross-entropy 4 Results and Analysis The performance of the proposed dual-modal diagnostic framework was eval- uated using multiple quantitative metrics, including accuracy, area under the receiver operating characteristic curve (AUROC), and macro F1-score. A com- parative analysis was conducted against single-modality baseline models trained using CT images only and H&E histopathology images only. Table 3 summarizes the overall classification performance of the evaluated models. The CT-only model achieved an accuracy of 0.84 and an AUROC of 0.94, demonstrating strong performance in detecting structural abnormalities within thoracic CT scans. Similarly, the H&E-only model achieved an accuracy of 0.85 and an AUROC of 0.95, indicating the effectiveness of histopathological features for lung cancer subtype classification. Explainable Dual-Modal AI for Lung Cancer Diagnosis13 Fig. 6: Confusion matrix of the proposed dual-modal fusion model for lung cancer subtype classification. The model demonstrates strong classification performance across all five categories, with the highest accuracy observed for small cell lung cancer (SCLC) at 90%. The proposed dual-modal fusion framework achieved the highest overall per- formance, with an accuracy of 0.87, AUROC exceeding 0.97, and a macro F1- score of 0.88. These results indicate that integrating radiological and histopatho- logical information improves diagnostic performance compared to single-modality approaches. The confusion matrix shown in Fig. 6 illustrates the classification perfor- mance across five lung cancer categories: adenocarcinoma, squamous cell carci- noma, large cell carcinoma, small cell lung cancer (SCLC), and normal tissue. The model achieved the highest correct classification rate for SCLC (90%) and squamous cell carcinoma (88%), while slightly lower performance was observed Table 3: Performance comparison between single-modality models and the pro- posed dual-modal fusion framework MetricCT-only Model H&E-only Model Dual-Modal Fusion Accuracy0.840.850.87 AUROC0.940.950.97 Macro-F1 Score0.840.850.88 14Baramee Sukumal and Aueaphum Aueawatthanaphisut for large cell carcinoma (78%), which is consistent with the known difficulty in distinguishing this subtype from other non-small cell lung cancers. Statistical significance between models was evaluated using DeLong’s test for correlated ROC curves. The dual-modal fusion model demonstrated statistically significant improvement compared to single-modality models (p < 0.05). Although the proposed fusion framework achieved strong overall performance, classification accuracy for large cell carcinoma remained lower compared to other subtypes. This observation is consistent with previous studies reporting overlap- ping morphological characteristics between large cell carcinoma and other non- small cell lung cancers. Future work will explore additional multimodal features and larger datasets to further improve classification robustness. 5 Conclusion This study presented a dual-modal explainable artificial intelligence framework for lung cancer diagnosis and subtype classification by integrating radiologi- cal CT imaging, histopathological H&E microscopy, and clinical metadata. The proposed system employed modality-specific convolutional neural networks to extract complementary radiological and pathological features, while a weighted decision-level fusion mechanism dynamically combined predictions from both modalities using clinical context. Experimental evaluation demonstrated that the proposed fusion framework outperformed single-modality models. The CT-only model achieved an accuracy of 0.84, while the H&E-only model achieved an accuracy of 0.85. In contrast, the proposed dual-modal approach achieved the highest performance with an overall accuracy of 0.87, AUROC exceeding 0.97, and a macro F1-score of 0.88. These results indicate that combining radiological and histopathological information significantly improves diagnostic reliability and classification performance. In addition to improved predictive accuracy, the integration of explainable AI techniques provided interpretable visualizations that highlighted clinically rel- evant tumor regions. Methods such as Grad-CAM and Grad-CAM++ demon- strated strong correspondence between model attention and expert-annotated lesion areas, supporting the transparency and clinical interpretability of the pro- posed system. Overall, the results suggest that multimodal integration offers a promising direction for improving AI-assisted cancer diagnosis. By bridging the gap be- tween non-invasive radiological screening and definitive pathological assessment, the proposed framework provides a unified and interpretable diagnostic pipeline that may support future clinical decision-making systems in precision oncology. Future work will focus on expanding the dataset with multi-center clinical data, improving multimodal feature fusion strategies, and validating the frame- work in real-world clinical environments to further assess its robustness and clinical applicability. Explainable Dual-Modal AI for Lung Cancer Diagnosis15 References 1. Armato I, S.G., McLennan, G., Bidaut, L., McNitt-Gray, M.F., Meyer, C.R., Reeves, A.P., Clarke, L.P.: The Lung Image Database Consortium (LIDC) and Image Database Resource Initiative (IDRI): A completed reference database of lung nodules on CT scans. Medical Physics 38(2), 915–931 (2011) 2. Aerts, H.J.W.L., Velazquez, E.R., Leijenaar, R.T.H., Parmar, C., Grossmann, P., Carvalho, S., Lambin, P.: Decoding tumour phenotype by noninvasive imaging using a quantitative radiomics approach. Nature Communications 5, 4006 (2014) 3. Clark, K., Vendt, B., Smith, K., Freymann, J., Kirby, J., Koppel, P., Prior, F.: The Cancer Imaging Archive (TCIA): Maintaining and operating a public information repository. Journal of Digital Imaging 26(6), 1045–1057 (2013) 4. The Cancer Imaging Archive (TCIA): NSCLC-Radiomics Data Collection. Avail- able at: https://w.cancerimagingarchive.net 5. The Cancer Imaging Archive (TCIA): Small Cell Lung Cancer (SCLC) Radio- genomics Data Collection. Available at: https://w.cancerimagingarchive.net 6. The Cancer Genome Atlas Research Network: Comprehensive molecular profiling of lung adenocarcinoma. Nature 511(7511), 543–550 (2014) 7. The Cancer Genome Atlas Research Network: Comprehensive genomic character- ization of squamous cell lung cancers. Nature 489(7417), 519–525 (2012) 8. Borkowski, A.A., Bui, M.M., Thomas, L.B., Wilson, C.P., DeLand, L.A., Mas- torides, S.M.: Lung and Colon Cancer Histopathological Image Dataset (LC25000). arXiv:1912.12142 (2019) 9. Tan, M., Le, Q.: EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning (ICML) (2019) 10. Isensee, F., Jaeger, P.F., Kohl, S.A., Petersen, J., Maier-Hein, K.H.: nnU-Net: A self-configuring method for deep learning-based biomedical image segmentation. Nature Methods 18, 203–211 (2021) 11. Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D.: Grad- CAM: Visual explanations from deep networks via gradient-based localization. Proceedings of ICCV, 618–626 (2017) 12. Chattopadhay, A., Sarkar, A., Howlader, P., Balasubramanian, V.N.: Grad- CAM++: Improved visual explanations for deep convolutional networks. Proceed- ings of WACV, 839–847 (2018) 13. Macenko, M., Niethammer, M., Marron, J.S., Borland, D., Woosley, J.T., Guan, X., Thomas, N.E.: A method for normalizing histology slides for quantitative analysis. Proceedings of ISBI, 1107–1110 (2009) 14. Vahadane, A., Peng, T., Sethi, A., Albarqouni, S., Wang, L., Baust, M., Anand, D.: Structure-preserving color normalization and sparse stain separation for histo- logical images. IEEE Transactions on Medical Imaging 35(8), 1962–1971 (2016) 15. Guo, C., Pleiss, G., Sun, Y., Weinberger, K.Q.: On calibration of modern neural networks. Proceedings of ICML (2017) 16. DeLong, E.R., DeLong, D.M., Clarke-Pearson, D.L.: Comparing the areas under two or more correlated ROC curves. Biometrics 44(3), 837–845 (1988) 17. McNemar, Q.: Note on the sampling error of the difference between correlated proportions. Psychometrika 12(2), 153–157 (1947) 16Baramee Sukumal and Aueaphum Aueawatthanaphisut 18. Efron, B.: Better bootstrap confidence intervals. Journal of the American Statisti- cal Association 82(397), 171–185 (1987) 19. Brier, G.W.: Verification of forecasts expressed in terms of probability. Monthly Weather Review 78(1), 1–3 (1950) 20. Petsiuk, V., Das, A., Saenko, K.: RISE: Randomized Input Sampling for explana- tion of black-box models. BMVC (2018) 21. Hooker, S., Erhan, D., Kindermans, P.J., Kim, B.: A benchmark for interpretability methods in deep neural networks. NeurIPS (2019) 22. Loshchilov, I., Hutter, F.: Decoupled weight decay regularization (AdamW). ICLR (2019) 23. Zhang, H., Cisse, M., Dauphin, Y.N., Lopez-Paz, D.: mixup: Beyond empirical risk minimization. ICLR (2018) 24. Yun, S., Han, D., Oh, S.J., Chun, S., Choe, J., Yoo, Y.: CutMix: Regularization strategy to train strong classifiers with localizable features. Proceedings of ICCV, 6023–6032 (2019)