Paper deep dive
Beyond Predictive Fairness: Quantifying Attribution Consistency Across Demographic Groups in Diabetic Retinopathy Screening
Kerol Djoumessi, Philipp Berens
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/20/2026, 5:07:18 AM
Summary
This paper introduces the Explanation Consistency Score (ECS), a metric based on Jensen-Shannon divergence, to quantify the similarity of attribution maps across demographic groups in diabetic retinopathy screening. Using the EyePACS dataset and a ResNet-50 model, the study finds that while predictive performance (AUC, sensitivity) varies significantly across ethnicities, attribution consistency remains high and shows no significant association with performance disparities. This suggests that predictive fairness and explanation consistency are complementary dimensions of model behavior.
Entities (8)
Relation Signals (6)
Explanation Consistency Score â uses â Jensen-Shannon Divergence
confidence 98% · ECS is introduced as a JensenâShannon divergence-based metric to quantify attribution similarity across demographic groups.
ResNet-50 â trainedon â EyePACS
confidence 97% · Experiments were conducted on the EyePACS diabetic retinopathy dataset... A ResNet-50 [12] was trained for early diabetic retinopathy (DR) detection
ScoreCAM â usedfor â Attribution Map Generation
confidence 95% · Attribution maps were generated using SmoothGradCAM++ [21] and Score-CAM [24]
SmoothGradCAM++ â usedfor â Attribution Map Generation
confidence 95% · Attribution maps were generated using SmoothGradCAM++ [21] and Score-CAM [24]
Predictive Performance â variesacross â Ethnicity
confidence 92% · Performance differences were observed across demographic groups. Patients of Indian origin achieved the highest AUC... whereas Caucasian patients exhibited the lowest AUC
Explanation Consistency â showsnosignificantassociationwith â Predictive Performance Disparities
confidence 90% · explanation consistency remains relatively high and shows no significant association with performance disparities.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Fairness in medical imaging is commonly evaluated through subgroup performance metrics, yet it remains unclear whether models rely on consistent visual evidence across demographic groups. This work introduces the Explanation Consistency Score (ECS), a fairness-aware metric based on Jensen-Shannon divergence that quantifies the similarity of attribution maps across subgroups. Using diabetic retinopathy screening as a case study, ECS is evaluated globally and within disease severity. Experiments reveal that while predictive performance differs across ethnic groups, explanation consistency remains relatively high and shows no significant association with performance disparities. These findings suggest that predictive fairness and explanation consistency capture complementary dimensions of model behavior, motivating fairness evaluations that extend beyond predictive performance.
Tags
Links
- Source: https://arxiv.org/abs/2608.18759v1
- Canonical: https://arxiv.org/abs/2608.18759v1
Trouble viewing inline? Open PDF directly â
Full Text
28,013 characters extracted from source content.
Expand or collapse full text
Beyond Predictive Fairness: Quantifying Attribution Consistency Across Demographic Groups in Diabetic Retinopathy Screening Kerol Djoumessi(đ) OrcID: 0009-0004-1548-9758 Affiliation: Hertie Institute for AI in Brain Health, University of TĂŒbingen, Germany https://hertie.ai/ E-mail kerol.djoumessi-donteu@uni-tuebingen.de Philipp Berens OrcID: 0000-0002-0199-4727 Affiliation: Hertie Institute for AI in Brain Health, University of TĂŒbingen, Germany https://hertie.ai/ E-mail kerol.djoumessi-donteu@uni-tuebingen.de Affiliation: TĂŒbingen AI Center, University of TĂŒbingen, Germany Abstract Fairness in medical imaging is commonly evaluated through subgroup performance metrics, yet it remains unclear whether models rely on consistent visual evidence across demographic groups. This work introduces the Explanation Consistency Score (ECS), a fairness-aware metric based on JensenâShannon divergence that quantifies the similarity of attribution maps across subgroups. Using diabetic retinopathy screening as a case study, ECS is evaluated globally and within disease severity. Experiments reveal that while predictive performance differs across ethnic groups, explanation consistency remains relatively high and shows no significant association with performance disparities. These findings suggest that predictive fairness and explanation consistency capture complementary dimensions of model behavior, motivating fairness evaluations that extend beyond predictive performance. Keywords: Algorithmic Fairness Explainable AI Explanation Consistency Medical Image Analysis Diabetic Retinopathy Screening. 1 Introduction Deep learning systems have achieved strong performance in automated medical image diagnosis across a wide range of clinical tasks [23], including diabetic retinopathy (DR) screening from retinal fundus images [3, 9]. Alongside these advances, fairness and explainability have become increasingly important concerns in medical imaging due to the reported performance disparities across demographic subgroups [2, 23]. Existing fairness studies primarily evaluate subgroup disparities using predictive metrics such as sensitivity, specificity, and area under the receiver operating characteristic curve (AUC) [15, 17]. While such evaluations quantify differences in predictive outcomes, they provide limited insight into whether models rely on similar visual evidence across patient groups. Recent studies show that medical imaging models can encode demographic information despite being trained for clinical tasks [25, 20]âin ophthalmology, demographic attributes can even be inferred directly from fundus images [20, 14]. Consequently, two models may show similar subgroup performance while relying on different visual evidence, or vice versa, making predictive metrics alone insufficient for trustworthy deployment. Explainability methods such as Class Activation Mapping (CAM) are widely used in ophthalmology to improve interpretability by highlighting retinal regions contributing to model predictions [4, 21, 24]. These methods have been used to verify whether models attend to clinically relevant retinal structures and lesions, thereby supporting the interpretation of automated screening systems. Despite known limitations and sensitivity to model architecture [1, 24], attribution maps remain one of the most widely adopted approaches for understanding model behavior in medical imaging [5, 4]. However, existing studies have primarily leveraged explainability qualitatively to inspect subgroup failures, identify potential sources of bias, or visualize fairness-related concerns [26, 27]. Quantitative analyses of attribution differences across demographic groups remain limited [22], and their relationship with predictive fairness is not yet well understood. This raises an important question: do predictive disparities necessarily imply differences in model reasoning across demographic groups? In this work, we investigate the demographic consistency of visual attribution patterns in early diabetic retinopathy detection using retinal fundus images from the EyePACS dataset. An Explanation Consistency Score (ECS) is introduced as a JensenâShannon divergence-based metric to quantify attribution similarity across demographic groups. To account for potential confounding effects arising from differences in disease severity, a conditional formulation of ECS is additionally proposed. By jointly analyzing predictive performance and explanation consistency across demographic groups, this work provides a quantitative framework for fairness-aware attribution analysis in medical imaging. 2 Methods 2.1 Dataset and Demographic Subgroups Experiments were conducted on the EyePACS diabetic retinopathy dataset [6], which provides retinal fundus images, DR grades (0â4), and patient ethnicity metadata. Following the provided quality-control protocol, 90,595 good-quality images were retained (Tab. 1). Data were split at the patient level into 70%70\% training, 15%15\% validation, and 15%15\% test sets using iterative multi-label stratification based on ethnicity and DR grade to preserve subgroup and disease distributions. The fairness analysis focused on the five largest ethnicity groups: Latin American, African Descent, Indian Origin, Caucasian, and Asian. Samples with unspecified ethnicity or underrepresented groups were retained for training but excluded from subgroup analyses due to missing or limited data. Table 1: Demographic distribution and DR grade composition of the EyePACS cohort. Ethnicity # Samples DR0 DR1 DR2 DR3 DR4 DR1-4 Prev. Latin American 53,96053,960 42,15542,155 5,7875,787 5,5025,502 285285 231231 11,80511,805 21.9%21.9\% African descent 5,0975,097 3,5973,597 747747 663663 4040 5050 1,5001,500 29.4%29.4\% Indian origin 3,3413,341 1,8221,822 472472 930930 5151 6666 1,5191,519 45.5%45.5\% Caucasian 9,5429,542 7,8187,818 794794 884884 2424 2222 1,7241,724 18.1%18.1\% Asian 3,9953,995 3,3353,335 321321 314314 44 2121 660660 16.5%16.5\% Not specified 12,25312,253 9,8259,825 1,2281,228 1,1141,114 6060 2626 2,4282,428 19.8%19.8\% Underrepresented 2,4072,407 1,8161,816 244244 324324 1717 66 591591 24.5%24.5\% 2.2 DR classification model A ResNet-50 [12] was trained for early diabetic retinopathy (DR) detection by distinguishing healthy eyes (grade 0) from eyes exhibiting any stage of DR (grade 1â4). This binary setting was selected because early disease detection requires identifying subtle retinal lesions [9] and therefore provides a suitable benchmark for studying potential demographic differences in model explanations. 2.3 Attribution Map Generation To analyze the visual evidence underlying model predictions, attribution maps were generated using SmoothGradCAM++ [21] and Score-CAM [24], representing gradient-based and gradient-free attribution methods, respectively. Using both methods enables evaluating the robustness of the proposed explanation consistency analysis across different attribution mechanisms. Each map AââNĂMA ^NĂ M was normalized to sum to one (A~i,j=Ai,j/âu,vAu,v A_i,j=A_i,j/ _u,vA_u,v), yielding a spatial probability distribution. 2.4 Explanation Consistency Score To quantify attribution consistency across demographic groups, an Explanation Consistency Score (ECS) is introduced. For a demographic group g, the mean attribution distribution is defined as ÎŒg=xâŒPâĄ(x|g)â[A~â(x)]â1|Dg|ââxâDgA~â(x), _g=E_x P(x|g)[ A(x)]â 1|D_g| _xâ D_g A(x), where DgD_g denotes the set of samples belonging to group g. Attribution distributions are compared using JensenâShannon (JS) divergence [18], a symmetric and bounded measure of similarity between probability distributions. The pairwise ECS between demographic groups g1g_1 and g2g_2 is defined as ECS(g1,g2)=1âJS(ÎŒg1â„ÎŒg2).ECS(g_1,g_2)=1-JS( _g_1\| _g_2). ECS ranges from 0 to 1, where larger values indicate greater similarity between attribution distributions. To summarize the consistency of a demographic group relative to all remaining groups, a multi-group ECS is computed as AvgECSâĄ(g)=1||â1ââgâČâ gECSâĄ(g,gâČ),AvgECS(g)= 1|G|-1 _g â gECS(g,g ), where G denotes the set of demographic groups. Disease severity may act as a confounding factor because attribution patterns can vary across DR grades. To control this effect, ECS is additionally computed conditionally on DR grade. For demographic group g and DR grade y, the conditional attribution distribution is defined as ÎŒg(y)=xâŒPâĄ(x|g,y)â[A~â(x)]. _g^(y)=E_x P(x|g,y)[ A(x)]. The conditional ECS for grade y is then ECS(y)(g1,g2)=1âJS(ÎŒg1(y)â„ÎŒg2(y)),ECS^(y)(g_1,g_2)=1-JS ( _g_1^(y)\| _g_2^(y) ), and the final conditional score is obtained by averaging across grades: CondECSâĄ(g1,g2)=1||ââyâECS(y)â(g1,g2).CondECS(g_1,g_2)= 1|Y| _y ECS^(y)(g_1,g_2). In practice, conditional analyses were restricted to DR grades 1 and 2, corresponding to the early disease stages considered in this study. 3 Experiments and Results 3.1 Experimental Setup Data Preprocessing and Augmentation. Images were preprocessed by circular field-of-view cropping [19], resized to 512Ă512512Ă 512, and normalised using training set statistics. Augmentation followed [13], including random cropping, color jittering, and rotation. Model Training. Following prior work [13, 9], the model11 1 The code is available at: https://github.com/kdjoumessi/Fairness-ECS was initialized with ImageNet-pretrained weights and optimized using cross-entropy loss. Training used SGD with an initial learning rate of 10â310^-3 and a clipped cosine schedule, selecting the model with highest validation AUC after 100 epochs. Attribution Processing. Attribution maps were generated using the TorchCAM library [11] at the resulting native low resolution (16Ă1616Ă 16 pixels), normalized for all consistency analyses, and upsampled for qualitative visualization. 3.2 Predictive Performance Across Demographic groups Subgroup predictive performance was assessed using AUC, sensitivity, and specificity (Tab. 2). The latter two directly reflecting clinical screening relevance. Table 2: Predictive performance and attribution consistency across demographic groups. Results are reported as mean ± standard deviation over three random seeds. Predictive performance Explanation consistency Ethnicity AUC Sens. Spec. ECS-SC WG-ECS ECS-SG DR1-ECS Latin American .85±.002.85±.002 .57±.013.57±.013 .97±.004.97±.004 .92±.002.92±.002 .53±.08.53±.08 .88±.010.88±.010 .87±006.87± 006 African descent .88±.002.88±.002 .64±.023.64±.023 .97±.002.97±.002 .91±005.91± 005 .56±.08.56±.08 .88±.008.88±.008 .85±.014.85±.014 Indian origin .92±.001.92±.001 .77±.022.77±.022 .94±.007.94±.007 .90±006.90± 006 .57±.08.57±.08 .86±.010.86±.010 .85±009.85± 009 Caucasian .83±.003.83±.003 .54±.018.54±.018 .97±.004.97±.004 .91±004.91± 004 .57±.07.57±.07 .86±.014.86±.014 .85±008.85± 008 Asian .90±.006.90±.006 .68±.023.68±.023 .97±.005.97±.005 .89±.004.89±.004 .54±.08.54±.08 .85±.008.85±.008 .79±.011.79±.011 All .86±.001.86±.001 .59±015.59± 015 .97±004.97± 004 â- â- â- â- Performance differences were observed across demographic groups. Patients of Indian origin achieved the highest AUC (0.92±0.0010.92± 0.001) and sensitivity (0.83±0.0030.83± 0.003), whereas Caucasian patients exhibited the lowest AUC (0.83±0.0030.83± 0.003) and sensitivity (0.54±0.0180.54± 0.018), corresponding to a sensitivity gap of 23%23\%. In contrast, specificity remained relatively stable across groups (0.940.94â0.970.97), indicating that subgroup differences primarily arise from the detection of positive DR cases rather than healthy patients. Interestingly, predictive performance did not follow subgroup size (Tab. 1). Although Latin American patients represent the largest demographic group, they did not achieve the strongest performance. Conversely, Indian-origin and Asian patients, among the smallest groups considered, obtained the highest AUC and sensitivity values. Neither subgroup size nor disease prevalence alone explains the observed differences: the Asian subgroup achieved strong performance despite the lowest prevalence, and Indian-origin patients outperformed the much larger Latin American group. Low standard deviations across all metrics indicate that these trends were consistent across random initializations. These findings reveal measurable predictive disparities across ethnicities and motivate the subsequent analysis of whether such differences are reflected in model attribution behavior. 3.3 Attribution Consistency Across Demographic Groups To assess whether model explanations differ across demographic groups, the proposed Explanation Consistency Score (ECS) for a given model was computed from correctly classified test samples using ScoreCAM (ECS-SC) and smoothGradCAM++ (ECS-SG). This design choiceârestricting to correct predictionsâisolates explanation consistency under successful decision-making, as misclassified examples may exhibit heterogeneous failure modes that complicate cross-group comparisons. Multi-group ECS values are reported in Table 2, while average attribution maps are shown in Figure 1. High explanation consistency was observed across all demographic groups. ECS values ranged from 0.890.89 to 0.920.92 for ScoreCAM and from 0.850.85 to 0.880.88 for SmoothGradCAM++, indicating that attribution patterns remained largely similar across ethnicities. These findings are consistent with the qualitative visualizations (Fig. 1), where both attribution methods consistently focus on similar retinal regions despite minor differences in attribution concentration and spatial distribution. Figure 1: Explanation consistency across demographic groups. Average ScoreCAM and SmoothGradCAM++ attribution maps for correctly classified DR-positive samples across demographic groups. Similar attention patterns are observed across ethnicities, indicating high explanation consistency. Compared with predictive performance, attribution consistency exhibited substantially lower variability across groups. While sensitivity varies by 23%23\%, ECS varies by only 0.030.03 for both attribution methods. Notably, the subgroup achieving the highest predictive performance (Indian origin) did not exhibit the highest ECS, whereas groups with lower predictive performance, such as Caucasian patients, remained highly consistent with the attribution patterns observed in other groups. These findings suggest that predictive disparities are not necessarily accompanied by comparable differences in attribution behavior. To further investigate this relationship, Spearman correlations (Ï) were computed between subgroup performance and ECS. Moderate negative correlations were observed for both AUC and sensitivity (Ï=â0.80Ï=-0.80 for ScoreCAM, Ï=â0.60Ï=-0.60 for SmoothGradCAM++), though neither reached statistical significanceâunsurprising given the limited number of demographic groups (n=5n=5), preventing meaningful evidence. To assess whether group-level ECS values could be explained by differences in within-group attribution variability, the similarity between each attribution map and its corresponding group-average attribution map was additionally measured. Within-group consistency (WG.ECS) was comparable across demographic groups (ScoreCAM: 0.53âââ0.570.53 0.57; SmoothGradCAM++: 0.43âââ0.500.43 0.50), indicating comparable levels of attribution variability and supporting the validity of group-level ECS comparisons. This suggests that the high ECS values are unlikely to be driven by subgroup-specific differences in attribution variability. 3.4 Conditional Attribution Consistency Analysis Figure 2: Conditional attribution consistency across demographic groups. Average attribution maps for DR grades 1 (top two rows) and 2 (bottom two rows) across demographic groups using ScoreCAM and SmoothGradCAM++. Attribution patterns remain broadly consistent within each disease grade, supporting the robustness of the proposed ECS analysis after controlling for disease severity. Differences in disease prevalence and severity across demographic groups may influence attribution patterns and potentially confound the interpretation of explanation consistency. To control for this effect, conditional ECS was computed separately for correctly classified DR grades 1 and 2, thereby comparing attribution maps only among patients with the same disease severity. High explanation consistency was preserved after conditioning on disease severity (Fig. 2). For ScoreCAM, multi-group ECS ranged from 0.790.79 to 0.870.87 for grade 1 (DR1-ECS, Tab. 2) and from 0.870.87 to 0.900.90 for grade 2 across demographic groups. Similarly, trends were observed for SmoothGradCAM++ produced ECS values ranging between 0.770.77â0.820.82 for grade 1 and 0.810.81â0.860.86 for grade 2. The highest consistency was generally observed for Latin American patients, whereas Asian patients exhibited the lowest ECS values across both grades and attribution methods. Across all demographic groups and attribution methods, ECS was consistently higher for grade 2 than for grade 1. For example, the ScoreCAM-based ECS of Asian patients increased from 0.79±0,010.79± 0,01 (grade 1) to 0.87±0.0020.87± 0.002 (grade 2). This trend suggests that attribution patterns become more stable as diabetic retinopathy lesions become more visually apparent, whereas early-stage disease (grade 1) is associated with greater attribution variability. Importantly, high ECS values persisted after controlling for disease severity, indicating that the consistency observed in the global analysis (Fig. 1) is not solely explained by differences in disease prevalence or grade distribution across across demographic groups. Overall, these findings provide further evidence that the model relies on broadly similar visual evidence across demographic groups even when comparisons are restricted to patients with the same disease stage. 4 Discussion and Conclusion This work introduced the Explanation Consistency Score (ECS) to quantify the similarity of attribution patterns across demographic groups in diabetic retinopathy screening and investigated whether demographic differences in predictive performance can be explained by differences in model attribution patterns. Although predictive performance varied across ethnicities, attribution consistency remained uniformly high for both ScoreCAM and SmoothGradCAM++, indicating that predictive fairness and explanation consistency capture complementary rather than equivalent aspects of model behavior. In particular, performance disparities were substantially larger than attribution disparities, and no significant association was observed between ECS and subgroup predictive performance. Several observations may help explain this decoupling. First, subgroup performance did not appear to be determined solely by sample size: the Indian-origin subgroup achieved the highest AUC and sensitivity despite being among the smallest groups, whereas the largest subgroup did not achieve the strongest performance. Second, disease prevalence may partially contribute to these differences, as higher prevalence generally leads to higher sensitivity, although the strong performance of the Asian subgroup despite its low prevalence suggests that prevalence alone is insufficient to explain subgroup variation. These findings point toward additional factors, such as differences in disease manifestation, image characteristics, or subgroup-specific data distributions [20, 2, 17]. At the same time, attribution consistency remained high across all groups, including those with lower predictive performance, suggesting that the model relies on broadly similar visual evidence despite differences in predictive outcomes. The conditional analysis further supported this conclusion, as high ECS values persisted after controlling for disease severity. Interestingly, grade-2 images exhibited higher ECS than grade-1 images, suggesting that more advanced disease stages produce more stable attribution patterns. Several limitations should be acknowledged. First, ECS was computed on correctly classified samples to characterize attribution consistency under successful predictions. Consequently, subgroup-specific failure modes remain unexplored and should be investigated in future work. Second, ECS is a group-level measure based on mean attribution distributions. Although an additional within-group analysis revealed comparable attribution variability across demographic groups, suggesting that the observed ECS values are not merely averaging artifacts, ECS may not fully capture individual-level explanation variability. Future work could complement ECS with image-level analyses of attribution variability. Third, attribution maps were analyzed at their native CAM resolution, which may limit sensitivity to fine-grained spatial differences. Fourth, ECS relies on JensenâShannon divergence; alternative similarity measures such as cosine similarity, Wasserstein distance, or structural similarity may capture complementary notions of attribution similarity [16]. Finally, the study was conducted on a single dataset and focused exclusively on ethnicity, using posthoc methods, motivating future evaluation across additional modalities, datasets, protected attributes, and self-explainable models [7, 8, 10]. Acknowledgements This project was supported by the Hertie Foundation and by the Deutsche Forschungsgemeinschaft under Germanyâs Excellence Strategy with the Excellence Cluster 2064 âMachine Learning â New Perspectives for Scienceâ, and a regular grant (project number 571331899). Disclosure of Interests. The authors declare no competing interests. References [1] J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim (2018) Sanity checks for saliency maps. Advances in neural information processing systems 31. Cited by: §1. [2] A. Alloula, R. Mustafa, D. R. McGowan, and B. W. PapieĆŒ (2024) On biases in a uk biobank-based retinal image classification model. In MICCAI Workshop on Fairness of AI in Medical Imaging, p. 140â150. Cited by: §1, §4. [3] W. L. Alyoubi, W. M. Shalash, and M. F. Abulkhair (2020) Diabetic retinopathy detection through deep learning techniques: a review. Informatics in medicine unlocked 20, p. 100377. Cited by: §1. [4] M. S. Ayhan, L. B. Kuemmerle, L. Kuehlewein, W. Inhoffen, G. Aliyeva, F. Ziemssen, and P. Berens (2022) Clinical validation of saliency maps for understanding deep neural networks in ophthalmology. Medical Image Analysis 77, p. 102364. Cited by: §1. [5] D. Bhati, F. Neha, and M. Amiruzzaman (2024) A survey on explainable artificial intelligence (xai) techniques for visualizing deep learning models in medical imaging. Journal of Imaging 10 (10), p. 239. Cited by: §1. [6] J. Cuadros and G. Bresnick (2009) EyePACS: an adaptable telemedicine system for diabetic retinopathy screening. Journal of diabetes science and technology 3 (3), p. 509â516. Cited by: §2.1. [7] K. Djoumessi, B. Bah, L. KĂŒhlewein, P. Berens, and L. Koch (2024) This actually looks like that: proto-bagnets for local and global interpretability-by-design. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 718â728. Cited by: §4. [8] K. Djoumessi and P. Berens (2026) SoftCAM: making black box models self-explainable for medical image analysis. In Medical Imaging with Deep Learning, p. 433â467. Cited by: §4. [9] K. Djoumessi, Z. Huang, L. KĂŒhlewein, A. Rickmann, N. Simon, L. M. Koch, and P. Berens (2025) An inherently interpretable ai model improves screening speed and accuracy for early diabetic retinopathy. PLOS Digital Health 4 (5), p. e0000831. Cited by: §1, §2.2, §3.1. [10] K. R. D. Donteu, I. Ilanchezian, L. KĂŒhlewein, H. Faber, C. F. Baumgartner, B. Bah, P. Berens, and L. M. Koch (2023) Sparse activations for interpretable disease grading. In Medical Imaging with Deep Learning, Cited by: §4. [11] F. Fernandez (2020) TorchCAM: class activation explorer. GitHub. Note: https://github.com/frgfm/torch-cam Cited by: §3.1. [12] K. He, X. Zhang, S. Ren, and J. Sun (2016) Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 770â778. Cited by: §2.2. [13] Y. Huang, L. Lin, P. Cheng, J. Lyu, R. Tam, and X. Tang (2023) Identifying the key components in resnet-50 for diabetic retinopathy grading from fundus images: a systematic investigation. Diagnostics 13 (10), p. 1664. Cited by: §3.1, §3.1. [14] I. Ilanchezian, D. Kobak, H. Faber, F. Ziemssen, P. Berens, and M. S. Ayhan (2021) Interpretable gender classification from retinal fundus images using bagnets. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 477â487. Cited by: §1. [15] A. J. Larrazabal, N. Nieto, V. Peterson, D. H. Milone, and E. Ferrante (2020) Gender imbalance in medical imaging datasets produces biased classifiers for computer-aided diagnosis. Proceedings of the National Academy of Sciences 117 (23), p. 12592â12594. Cited by: §1. [16] A. Levy, B. R. Shalom, and M. Chalamish (2025) A guide to similarity measures and their data science applications. Journal of Big Data 12 (1), p. 188. Cited by: §4. [17] R. Mehta, C. Shui, and T. Arbel (2024) Evaluating the fairness of deep learning uncertainty estimates in medical image analysis. In Medical Imaging with Deep Learning, p. 1453â1492. Cited by: §1, §4. [18] M. L. MenĂ©ndez, J. A. Pardo, L. Pardo, and M. d. C. Pardo (1997) The jensen-shannon divergence. Journal of the Franklin Institute 334 (2), p. 307â318. Cited by: §2.4. [19] Fundus circle cropping External Links: Document, Link Cited by: §3.1. [20] S. MĂŒller, L. M. Koch, H. Lensch, and P. Berens (2022) A generative model reveals the influence of patient attributes on fundus images. In Medical Imaging with Deep Learning, Cited by: §1, §4. [21] D. Omeiza, S. Speakman, C. Cintas, and K. Weldermariam (2019) Smooth grad-cam++: an enhanced inference level visualization technique for deep convolutional neural network models. arXiv preprint arXiv:1908.01224. Cited by: §1, §2.3. [22] S. Pfohl, N. Harris, C. Nagpal, D. Madras, V. Mhasawade, O. Salaudeen, A. Dieng, S. Sequeira, S. Arciniegas, L. Sung, et al. (2026) Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness. Advances in Neural Information Processing Systems 38, p. 41887â41948. Cited by: §1. [23] A. Vrudhula, A. C. Kwan, D. Ouyang, and S. Cheng (2024) Machine learning and bias in medical imaging: opportunities and challenges. Circulation: Cardiovascular Imaging 17 (2), p. e015495. Cited by: §1. [24] H. Wang, Z. Wang, M. Du, F. Yang, Z. Zhang, S. Ding, P. Mardziel, and X. Hu (2020) Score-cam: score-weighted visual explanations for convolutional neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, p. 24â25. Cited by: §1, §2.3. [25] Y. Yang, H. Zhang, J. W. Gichoya, D. Katabi, and M. Ghassemi (2024) The limits of fair medical imaging ai in real-world generalization. Nature medicine 30 (10), p. 2838â2848. Cited by: §1. [26] Y. Zhao, Y. Wang, and T. Derr (2023) Fairness and explainability: bridging the gap towards fair model explanations. In Proceedings of the AAAI conference on artificial intelligence, Vol. 37, p. 11363â11371. Cited by: §1. [27] J. Zhou, F. Chen, and A. Holzinger (2020) Towards explainability for ai fairness. In International workshop on extending explainable AI beyond deep models and classifiers, p. 375â386. Cited by: §1.