Paper deep dive
A Case Study on Concept Induction for Neuron-Level Interpretability in CNN
Moumita Sen Sarma, Samatha Ereshi Akkamahadevi, Pascal Hitzler
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/20/2026, 7:38:44 AM
Summary
This case study investigates the generalizability of a Concept Induction-based framework for neuron-level interpretability in Convolutional Neural Networks (CNNs). While prior work demonstrated effectiveness on the ADE20K dataset, this study applies the same methodology to the SUN2012 dataset. The authors fine-tuned multiple CNN architectures, selected InceptionV3 as the best performer, and extracted activations from the dense layer. Using the Efficient Concept Induction and Integration (ECII) system, they mapped neuron activations to semantic concepts from a Wikipedia-based ontology. The results confirmed that 32 neurons exhibited stable concept associations with high Target Level Activation (TLA) and statistical significance, demonstrating that the method reliably transfers across different datasets and architectures.
Entities (9)
Relation Signals (8)
InceptionV3 → bestperformeron → SUN2012
confidence 97% · in our experiments on SUN2012, InceptionV3 achieves the highest performance
Concept Induction → appliedto → ADE20K
confidence 95% · Prior work introduced a Concept Induction-based framework for hidden neuron analysis and demonstrated its effectiveness on the ADE20K dataset.
Concept Induction → appliedto → SUN2012
confidence 95% · In this case study, we investigate whether the approach generalizes by applying it to the SUN2012 dataset
ResNet50V2 → bestperformeron → ADE20K
confidence 95% · the prior study on ADE20K reported ResNet50V2 as the best-performing model
Neuron Activation → evaluatedby → Target Level Activation
confidence 95% · Label of a neuron is confirmed when Target Level Activation (TLA)≥80%
Concept Induction → uses → ECII
confidence 92% · The Efficient Concept Induction and Integration (ECII) system is employed to derive concepts from the neuron activation sets.
Neuron Activation → evaluatedby → Mann–Whitney U test
confidence 90% · For statistical validation, Mann–Whitney U test... is performed
ECII → uses → Wikipedia
confidence 90% · mapping them to exact lexical matches within a Wikipedia-based concept hierarchy
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Deep Neural Networks (DNNs) have advanced applications in domains such as healthcare, autonomous systems, and scene understanding, yet the internal semantics of their hidden neurons remain poorly understood. Prior work introduced a Concept Induction-based framework for hidden neuron analysis and demonstrated its effectiveness on the ADE20K dataset. In this case study, we investigate whether the approach generalizes by applying it to the SUN2012 dataset, a large-scale scene recognition benchmark. Using the same workflow, we assign interpretable semantic labels to neurons and validate them through web-sourced images and statistical testing. Our findings confirm that the method transfers to SUN2012, showing its broader applicability.
Tags
Links
- Source: https://arxiv.org/abs/2603.00197v1
- Canonical: https://arxiv.org/abs/2603.00197v1
Trouble viewing inline? Open PDF directly →
Full Text
14,310 characters extracted from source content.
Expand or collapse full text
A Case Study on Concept Induction for Neuron-Level Interpretability in CNN Moumita Sen Sarma 1,* , Samatha Ereshi Akkamahadevi 1 and Pascal Hitzler 1,* 1 Department of Computer Science, Kansas State University, Manhattan, Kansas, USA Abstract Deep Neural Networks (DNNs) have advanced applications in domains such as healthcare, autonomous systems, and scene understanding, yet the internal semantics of their hidden neurons remain poorly understood. Prior work introduced a Concept Induction–based framework for hidden neuron analysis and demonstrated its effectiveness on the ADE20K dataset [1]. In this case study, we investigate whether the approach generalizes by applying it to the SUN2012 dataset [2], a large-scale scene recognition benchmark. Using the same workflow, we assign interpretable semantic labels to neurons and validate them through web-sourced images and statistical testing. Our findings confirm that the method transfers to SUN2012, showing its broader applicability. Keywords Explainable AI (XAI), Concept Induction, Hidden Neuron Analysis, CNNs 1. Introduction Deep Neural Networks (DNNs), particularly Convolutional Neural Networks (CNNs), have achieved state-of-the-art performance in image classification and scene understanding. Yet, the hidden neurons remain opaque, limiting interpretability in domains where transparency is critical [3]. Explainable AI (XAI) techniques such as saliency maps and attribution methods (e.g., SHAP [4], LIME [5]) highlight input contributions but rarely capture what individual neurons semantically represent. Concept Induction [6] offers a neurosymbolic alternative by mapping neuron activations to high-level concepts grounded in knowledge graphs, which organize concepts into hierarchies of entities and relations. By contrasting positive and negative activation sets with a structured ontology, it induces logical class expressions for each neuron that yield candidate semantic labels. In previous work [1], this approach was applied to the ADE20K dataset [7] with the automation pipeline in [8], demonstrating that neurons could be systematically assigned interpretable labels with strong empirical support. Herein, we apply the same methodology to a different large-scale benchmark, SUN2012 [2]. Our objective is to assess whether the findings extend beyond ADE20K. By validating the approach on SUN2012, we indeed show that the method transfers and that robust neuron–concept associations emerge consistently across datasets. 2. Methodology The hidden neuron analysis follows a structured workflow that connects hidden neuron activations to human-understandable concepts. We follow the approach from [1], consisting of data preparation, model training, activation extraction, concept induction, and evaluation of the induced concepts. Data Selection & Preparation: SUN2012 contains 131,000 images across 908 scene categories with annotations for over 3,800 objects. For this study, the ten largest categories are selected: bathroom, bedroom, building facade, dining room, highway, kitchen, living room, mountain snowy, skyscraper, and street, yielding 3,157 images for training and validation and 793 for testing. Model Training:As in [1],multiple CNN architectures (VGG16/19,InceptionV3, ResNet50/101/152/50V2) are fine-tuned.The models were trained using their standard input K-CAP 2025 Posters and Demos: The 13 th International Conference on Knowledge Capture, December 10 - 12, 2025, Dayton, Ohio, USA * Corresponding author. $ moumita@ksu.edu (M. S. Sarma); samatha94@ksu.edu (S. E. Akkamahadevi); hitzler@ksu.edu (P. Hitzler) 0009-0003-6644-6531 (M. S. Sarma); 0009-0001-6333-8004 (S. E. Akkamahadevi); 0000-0001-6192-3472 (P. Hitzler) © 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (C BY 4.0). resolutions: 299× 299 for InceptionV3 and 224× 224 for the remaining architectures. Training was carried out with the Adam optimizer set to a learning rate of 0.001, using categorical cross-entropy as the loss function. A batch size of 32 was used, and training ran for up to 30 epochs. To prevent overfitting, early stopping was applied by monitoring the validation loss with a patience of three epochs, restoring the best-performing model weights.While the prior study on ADE20K reported ResNet50V2 as the best-performing model, in our experiments on SUN2012, InceptionV3 achieves the highest performance (96.83% training accuracy; 92.71% validation accuracy) and is therefore selected for further analysis. Neuron Activation Extraction: Following the process introduced in previous work [1], neuron activations are extracted from the dense layer of the trained network. Each test image is passed through the network, and activation values of all 64 dense-layer neurons are collected. To capture activation behavior, thresholds are defined relative to the maximum observed response: images at or above 80% form the positive set, while those at or below 20% form the negative set. These contrasting sets are then used in the Concept Induction stage to derive the semantic labels. Concept Induction: The Efficient Concept Induction and Integration (ECII) system [9] is employed to derive concepts from the neuron activation sets. For every image, a minimal ontology is constructed by considering only the annotated objects and mapping them to exact lexical matches within a Wikipedia-based concept hierarchy [10]. These image-specific ontologies are then integrated with the Wikipedia hierarchy through ECII to form a single, robust background ontology. ECII then made use of this background knowledge to generate logical class expressions evaluated using coverage score that distinguishes positive from negative sets and that is defined by coverage(퐸) = |푍 1 | +|푍 2 | |푃 ∪ 푁| , where푍 1 = 푝 ∈ 푃 | 퐾 |= 퐸(푝)and푍 2 = 푛 ∈ 푁 | 퐾 ̸|= 퐸(푛),푃and푁refer to positive set and negative set, and퐾represents the knowledge base. Coverage measures how well the induced concept aligns with a neuron’s activation pattern, with higher scores indicating a better match. This stage resulted in candidate semantic labels (e.g., snowy mountain, toilet tissue, crosswalk) for neurons. Concept Evaluation: We adopt the same evaluation procedure as in the ADE20K study, combining web-sourced image confirmation with statistical testing. For label confirmation of each neuron, up to 100 Google Images are retrieved and tested with 80% of the retrieved images. Label of a neuron is confirmed when Target Level Activation (TLA)≥80%, a measurement of the proportion of images associated with a concept that reliably activates the neuron. On the other hand, Non-Target Label Activation (Non-TLA) is defined as the proportion of those same images in which a different neuron reliably activates above the threshold. For statistical validation, Mann–Whitney U test, a non-parametric statistical test is performed on 20% of the retrieved images. Here, we require p < 0.05 with a negative z-score, which demonstrates that target images consistently trigger stronger activations than non-target images. This results in rejecting the null hypothesis, which states that there is no significant difference between activations for target images (retrieved with the neuron’s label) and non-target images (retrieved with other labels). 3. Results and Discussion The application of Concept Induction to the SUN2012 dataset produces a set of interpretable neuron labels that aligned strongly with semantic concepts. Out of the 64 dense-layer neurons analyzed, 32 are confirmed to exhibit stable concept associations with Target Label Activations (TLA) of at least 80%. Of these 32 neurons, 29 neurons also show statistically significant separation between target and non-target activations under the Mann–Whitney U test (p < 0.05), confirming that their responses were meaningfully stronger for concept-related images. Table 2 shows the result of statistical test, where confirmed labels include crosswalk, skyscraper, pillow, ceiling fan, and bidet, Compared to ADE20K, which yielded 19 confirmed neurons, SUN2012 produces 32 under the same evaluation procedure. This shows that despite dataset and architectural differences (ResNet50V2 for ADE20K and InceptionV3 for SUN2012), the Concept Induction framework reliably identifies Table 1 Label confirmation via Google Images results with confirmed neurons (TLA %≥ 80). Neuron IDECII ConceptsCoverage ScoreTLA %Non-TLA % 0snowy_mountain0.98695.0052.57 7snow, mountain0.99880.0039.55 9field0.97193.7574.96 13dish_rack0.89281.2531.12 15pillow, ceiling_fan0.99288.7557.15 16skyscraper, river_water0.99780.0042.05 18city0.99387.5068.62 19snowy_mountain0.97597.5046.58 21bedroom, duvet0.96785.0047.67 22dishwasher0.94596.2024.58 23fence0.96598.7577.85 24sink0.93996.2552.59 26toilet_tissue0.97995.0039.49 27toilet0.96195.0055.33 28cars0.90397.5072.79 31snowy_mountain0.94897.5067.00 32shell0.95080.0046.64 33bed, pillow0.94580.0034.41 36snowy_mountain0.97398.7568.33 40plant0.97283.7551.00 41pole, handrail0.99785.0070.35 42skyscraper0.96993.7549.47 43skyscraper0.978100.0066.88 47crosswalk0.98381.2523.42 48bidet0.96698.7557.35 49ironing_board, shoe0.96993.7563.09 50field0.97683.7566.54 54snowy_mountain0.96592.5040.80 58snowy_mountain0.96997.5055.05 60brocoli, cheese0.97993.7539.00 61cars0.97690.0045.51 62air_conditioning, chest_of_drawers0.98287.5061.18 semantically coherent neurons across benchmarks. Therefore, it provides fine-grained, human-readable, and verifiable neuron-level explanations, supporting transparent analysis, greater trust, and practical debugging of deep models. Acknowledgments The authors acknowledge partial funding through the Kansas State University Game-changing Research Initiation Program (GRIP). Declaration on Generative AI For this work, the authors used ChatGPT-5 for grammar and spelling checks. Subsequently, the authors reviewed and edited the content as needed and take full responsibility for the content. References [1]A. Dalal, R. Rayan, A. Barua, E. Y. Vasserman, M. K. Sarker, P. Hitzler, On the value of labeled data and symbolic methods for hidden neuron activation analysis, in: International Confer- ence on Neural-Symbolic Learning and Reasoning, Springer, 2024, p. 109–131. doi:10.1007/ 978-3-031-71170-1_12. [2]J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, A. Torralba, SUN database: Large-scale scene recognition from abbey to zoo, in: 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE, 2010, p. 3485–3492. doi:10.1109/CVPR.2010.5539970. Table 2 Statistical evaluation results. Bold rows represents neurons with p-value≥0.05, where we cannot reject the null hypothesis. Neuron ID ECII ConceptsTLA % Non-TLA % Target Median Non- Target Median Target Mean Non-Target Mean 푧-score푝-value 0snowy_mountain9553.025.990.235.581.22-6.18< 0.00001 7snow, mountain7039.682.100.001.740.68-3.040.00061 9field10075.484.381.434.031.97-3.640.00025 13dish_rack9032.621.650.001.680.56-4.66< 0.00001 15pillow, ceiling_fan9054.213.030.212.360.85-4.40< 0.00001 16skyscraper, river_water7043.411.540.001.440.45-3.150.00052 18city8071.112.170.822.031.10-2.850.00398 19snowy_mountain10045.955.910.005.391.09-6.55< 0.00001 21bedroom, duvet6545.950.860.001.910.57-2.700.00332 22dishwasher8026.111.600.001.870.29-5.04< 0.00001 23fence10078.733.071.612.961.80-3.400.00063 24sink9553.253.140.113.140.81-5.60< 0.00001 26toilet_tissue10038.652.470.002.360.51-6.29< 0.00001 27toilet9556.984.060.333.880.94-6.10< 0.00001 28cars9572.781.581.031.661.59-1.100.26788 31snowy_mountain9067.624.580.623.971.35-5.05< 0.00001 32shell8546.191.650.001.550.70-3.770.00004 33bed, pillow7535.873.190.002.600.43-4.62< 0.00001 36snowy_mountain10067.624.720.934.601.48-6.24< 0.00001 40plant7051.901.000.070.860.68-1.440.12808 41pole, handrail9069.212.231.052.051.51-2.180.02735 42skyscraper9550.792.910.052.781.01-4.49< 0.00001 43skyscraper10066.512.940.813.011.10-5.65< 0.00001 47crosswalk9522.943.200.003.200.29-6.74< 0.00001 48bidet10058.023.000.432.911.01-5.35< 0.00001 49ironing_board, shoe9062.061.980.552.201.02-3.370.00053 50field8566.830.980.731.151.11-0.890.36257 54snowy_mountain10039.684.530.004.010.72-6.58< 0.00001 58snowy_mountain9552.944.560.184.771.12-6.33< 0.00001 60brocoli, cheese9539.521.790.001.930.50-5.89< 0.00001 61cars8545.951.410.001.320.55-4.07< 0.00001 62air_conditioning, chest_of_drawers8562.222.120.502.590.95-3.920.00006 [3]P. P. Angelov, E. A. Soares, R. Jiang, N. I. Arnold, P. M. Atkinson, Explainable artificial intelligence: an analytical review, Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 11 (2021) e1424. doi:10.1002/widm.1424. [4] S. M. Lundberg, S.-I. Lee, A unified approach to interpreting model predictions, in: Advances in Neural Information Processing Systems (NeurIPS), volume 30, Curran Associates, Inc., 2017, p. 4765–4774. [5] M. T. Ribeiro, S. Singh, C. Guestrin, “Why Should I Trust You?” Explaining the predictions of any classifier, in: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, ACM, 2016, p. 1135–1144. doi:10.1145/2939672.2939778. [6]J. Lehmann, P. Hitzler, Concept learning in description logics using refinement operators, Machine Learning 78 (2010) 203–250. doi:10.1007/s10994-009-5146-2. [7] B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, A. Torralba, Semantic understanding of scenes through the ADE20K dataset, International Journal of Computer Vision 127 (2019) 302–321. doi:10.1007/s11263-018-1140-0. [8]S. E. Akkamahadevi, A. Dalal, P. Hitzler, Automating cnn neuron interpretation using concept induction, in: Proceedings of the ISWC, 2024, p. 11–15. [9] M. K. Sarker, P. Hitzler, Efficient concept induction for description logics, in: Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, AAAI Press, 2019, p. 3036–3043. doi:10.1609/aaai.v33i01.33013036. [10]M. K. Sarker, J. Schwartz, P. Hitzler, L. Zhou, S. Nadella, B. S. Minnery, I. Juvina, M. L. Raymer, W. R. Aue, Wikipedia knowledge graph for explainable AI, in: Knowledge Graphs and Semantic Web Conference, KGSWC, 2020, Proceedings, volume 1232 of Communications in Computer and Information Science, Springer, 2020, p. 72–87. doi:10.1007/978-3-030-65384-2\_6.