Paper deep dive
A Practical Guide Towards Interpreting Time-Series Deep Clinical Predictive Models: A Reproducibility Study
Yongda Fan, John Wu, Andrea Fitzpatrick, Naveen Baskaran, Jimeng Sun, Adam Cross
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/27/2026, 1:17:01 AM
Summary
This paper presents a comprehensive reproducibility study and benchmark for interpretability methods in deep clinical time-series predictive models. The authors evaluate various attribution methods (including SHAP, LIME, Integrated Gradients, and Chefer) across multiple clinical tasks and model architectures (StageNet, Transformer, StageAttn) using the MIMIC-IV dataset. Key findings indicate that attention-based attribution (Chefer) is highly efficient and faithful, while black-box methods like SHAP and LIME are computationally infeasible for large-scale clinical data. The study provides an open-source implementation via the PyHealth framework.
Entities (6)
Relation Signals (3)
StageAttn â isvariantof â StageNet
confidence 95% · StageAttnâa modified StageNet with an additional multi-head attention layer
PyHealth â providesimplementationfor â Interpretability Benchmark
confidence 95% · To support reproducibility and extensibility, we provide our implementations via PyHealth
Chefer â outperforms â KernelSHAP
confidence 90% · Chefer emerges as a competitive method... while many of the other approaches move a tremendous degree depending on the task and model.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Clinical decisions are high-stakes and require explicit justification, making model interpretability essential for auditing deep clinical models prior to deployment. As the ecosystem of model architectures and explainability methods expands, critical questions remain: Do architectural features like attention improve explainability? Do interpretability approaches generalize across clinical tasks? While prior benchmarking efforts exist, they often lack extensibility and reproducibility, and critically, fail to systematically examine how interpretability varies across the interplay of clinical tasks and model architectures. To address these gaps, we present a comprehensive benchmark evaluating interpretability methods across diverse clinical prediction tasks and model architectures. Our analysis reveals that: (1) attention when leveraged properly is a highly efficient approach for faithfully interpreting model predictions; (2) black-box interpreters like KernelSHAP and LIME are computationally infeasible for time-series clinical prediction tasks; and (3) several interpretability approaches are too unreliable to be trustworthy. From our findings, we discuss several guidelines on improving interpretability within clinical predictive pipelines. To support reproducibility and extensibility, we provide our implementations via PyHealth, a well-documented open-source framework: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2603.24828v1
- Canonical: https://arxiv.org/abs/2603.24828v1
Trouble viewing inline? Open PDF directly â
Full Text
25,756 characters extracted from source content.
Expand or collapse full text
11institutetext: University of Illinois Urbana-Champaign, Champaign, IL 61801, USA 11email: yongdaf2,johnwu3@illinois.edu 22institutetext: PyHealth 33institutetext: University of Illinois College of Medicine, Chicago, IL 60612, USA A Practical Guide Towards Interpreting Time-Series Deep Clinical Predictive Models: A Reproducibility Study Yongda Fanâ John Wuâ Andrea Fitzpatrick Naveen Baskaran Jimeng Sun Adam Cross Abstract Clinical decisions are high-stakes and require explicit justification, making model interpretability essential for auditing deep clinical models prior to deployment. As the ecosystem of model architectures and explainability methods expands, critical questions remain: Do architectural features like attention improve explainability? Do interpretability approaches generalize across clinical tasks? While prior benchmarking efforts exist, they often lack extensibility and reproducibility, and critically, fail to systematically examine how interpretability varies across the interplay of clinical tasks and model architectures. To address these gaps, we present a comprehensive benchmark evaluating interpretability methods across diverse clinical prediction tasks and model architectures. Our analysis reveals that: (1) attention when leveraged properly is a highly efficient approach for faithfully interpreting model predictions; (2) black-box interpreters like KernelSHAP and LIME are computationally infeasible for time-series clinical prediction tasks; and (3) several interpretability approaches are too unreliable to be trustworthy. From our findings, we discuss several guidelines on improving interpretability within clinical predictive pipelines. To support reproducibility and extensibility, we provide our implementations via PyHealth, a well-documented open-source framework: https://github.com/sunlabuiuc/PyHealth. 1 Introduction One of the key barriers to deploying AI models in clinical settings is the need for explainability in deep clinical predictive models, as identified by practicing clinicians [12]. Beyond practical concerns such as model trustworthiness, regulations governing automated clinical systems legally require justification for each automated decision [1]. This has spurred development of interpretability methods ranging from white-box approaches like mechanistic interpretability [3] and gradient-based methods [2] to black-box approaches like SHAP [15]. While several studies have explored these approaches within the clinical domain [25], the best approach for interpreting deep clinical predictive models still remains unclear [13]. To better understand this problem, we devise an interpretability benchmark, and identify two practical concerns that guide our evaluation of interpretability approaches: Scalability across patient events and populations. Understanding model predictions extends beyond analyzing single samples. It requires exploring diverse feature combinations and characterizing the modelâs prediction space across entire populations and various modalities. Interpretability approaches must therefore scale to patient populations that contain hundreds of thousands of patients and millions of clinical events. Faithfulness to downstream predictions. While explanations may not always provide immediately useful qualitative insights [8], they must demonstrably influence the modelâs predictions. Explanations that fail to affect model outputs are fundamentally untrustworthy. With these criteria, we address two key questions in our reproducibility study: Do architectural changes such as attention improve explainability? Three model families dominate clinical time-series prediction [18][27]: state-based recurrent models like StageNet [10], attention-based models like Transformers [26], and hybrid architectures combining both approaches. Beyond improving downstream performance, attention mechanisms are often claimed to enhance model interpretability [21], motivating numerous attention-focused interpretability methods [6]. We directly compare attention-based and non-attention-based models to assess their impact on explanation faithfulness. Do interpretability approaches generalize across tasks? An interpretability method effective for one task may fail when input and output distributions differ. We evaluate how interpretability approaches perform across diverse clinical prediction tasks, including length of stay, mortality prediction, and condition-specific predictions such as diabetic ketoacidosis. In exploring these questions, our contributions are: (1) we demonstrate that attention, when leveraged properly, is an efficient and effective tool for interpretability across all tasks; (2) we show that black-box interpreters like KernelSHAP and LIME are computationally infeasible for time-series clinical prediction tasks; and (3) we reveal that several interpretability approaches are too unreliable to be trustworthy. From our findings, we discuss several guidelines on improving interpretability within clinical predictive pipelines. To support reproducibility and extensibility, we provide our implementations via PyHealth, a well-documented open-source framework. Table 1: Comparison of interpretability reproducibility studies in healthcare AI. This benchmark advances prior work by (1) evaluating recent interpretability methods, (2) comparing across multiple tasks and models, and (3) providing an accessible, open-source implementation in PyHealth that can directly extend to workflows beyond this study. Study Extensible to Other Workflows Public Code Available Explores Different Tasks Cross-compares Models & Tasks Explores New Approaches BenchXAI [17] â â â â â MIMIC-IF [16] â â â â â Zhou et al. [29] â â â â â Brankovic et al. [4] â â â â â Ours â â â â â 2 Related Works Existing interpretability benchmarks. While previous work has benchmarked interpretability approaches on clinical tasks [16][29][4][17], these efforts have notable limitations. First, they inadequately explore task diversity and architectural biases, often evaluating interpretability methods on a single task or model architecture [16][29][4]. Second, most focus primarily on older post-hoc techniques like LIME and SHAP, neglecting modern attention-based mechanisms [16][29][4][17]. Third, many lack publicly available code, hindering reproducibility [16][29][4]. In Table 1, our framework addresses these gaps in three key ways. Unlike existing benchmarks that focus purely on measuring interpretability performance, we provide an extensible toolkit that researchers can directly integrate into custom clinical workflows through PyHealth. We systematically evaluate interpretability across multiple tasks, model architectures, and both traditional post-hoc methods and modern attention mechanisms. Given recent advances in interpretability methods, we believe such an accessible framework for systematic evaluation is urgently needed. 3 Methodology There are a massive number of interpretability approaches [20] that have been developed, making extensive testing of all interpretability approaches out of scope for our work. Nonetheless, there are several key approaches that are classically used in many other interpretability evaluations. Setup. All interpretability methods evaluated in this work share a common goal: given an input ââdx ^d and model f, produce an attribution map ââdA ^d where each element AiA_i represents the importance of feature xix_i to the prediction fâ()f(x). This common output format enables fair comparison across diverse attribution strategiesâfrom Shapley values to gradient flows to attention mechanismsâwhich differ primarily in their theoretical justification and computational approach for assigning these importance weights. Black-box interpreters. Due to their model agnostic nature, black-box interpreters are a very popular first choice when attempting to interpret a clinical predictive model [11]. Of particular note, LIME [22] and SHAP [15] are well cited amongst the literature. Both SHAP and LIME are additive feature attribution methods that explain a prediction fâ(x)f(x) via a linear explanation model gâ(zâČ)=Ï0+âi=1MÏiâziâČg(z )= _0+ _i=1^M _iz _i, where zâČâ0,1Mz â\0,1\^M indicates feature presence and each Ïi _i is a featureâs contribution. LIME estimates the Ïi _i by fitting a weighted linear regression around the input with a heuristically chosen kernel ÏxâČ _x . SHAP instead computes Shapley values from cooperative game theory: Ïi=âSâFâi|S|!â(|F|â|S|â1)!|F|!â[fSâȘiâ(xSâȘi)âfSâ(xS)], _i= _S F \i\ |S|!\,(|F|-|S|-1)!|F|! [f_SâȘ\i\(x_SâȘ\i\)-f_S(x_S) ], which are the unique attributions satisfying local accuracy, missingness, and consistency. Here, we implement Kernel SHAP, which unifies the two by showing LIME recovers exact Shapley values under a specific kernel ÏxâČâ(zâČ)=(Mâ1)/[(M|zâČ|)â|zâČ|â(Mâ|zâČ|)] _x (z )=(M-1)/ [ M|z |\,|z |\,(M-|z |) ] with no regularization. Gradient-Based Counterfactual Attribution. A growing class of attribution methods explain predictions by measuring how the output changes as inputs move from a reference state to their observed values, differing primarily in how they compute and propagate these changes. Integrated Gradients [24] accumulates continuous gradients along a straight-line path from baseline x0x^0 to input x: IGiâ(x)=(xiâxi0)ââ«01âFâ(x0+αâ(xâx0))âxiâα,IG_i(x)=(x_i-x_i^0) _0^1 â F(x^0+α(x-x^0))â x_i\,dα, satisfying completeness (âiIGi=Fâ(x)âFâ(x0) _iIG_i=F(x)-F(x^0)) and implementation invariance. In practice, the integral is approximated via m interpolation steps, each requiring a gradient computation. DeepLIFT [23] achieves a similar summation-to-delta property in a single forward-backward pass by propagating discrete activation differences layer-by-layer. This is far cheaper, but because the chain rule does not hold for discrete gradients, DeepLIFT can yield different attributions for functionally equivalent networks, violating implementation invariance [24]. Both methods can underestimate component importance in transformers due to self-repair, where downstream components compensate for perturbations. [9] observe that this is especially prevalent in LLM attention mechanisms â softmax redistribution masks the true influence of attention scores â and propose GIM, which modifies gradient flow through softmax, layer normalization, and multiplicative interactions to account for these effects, yielding more faithful attributions. Attention-based Attribution. Finally, from the transformer architecture came a variety of works that claim to improve the interpretability of their models through the attention mechanism[21] [6] [7][19]. Of particular note, Chefer [6] shows that by simply aggregating and weighing the attention maps with its gradients, they can dramatically improve the faithfulness and explainability of transformer models compared to simply using only their attention scores. Ultimately, each approach has shown valid empirical evidence of their utility [28], making their exploration in a fair and reproducible manner a key priority. 4 Results Table 2: Interpretability Performance Matrix. Each cell shows the number of model-task pairs where the row method outperforms the column method, scored by ComprehensivenessĂ(1âSufficiency)ComprehensivenessĂ(1-Sufficiency). Darker orange indicates higher win rate when reading horizontally, and higher lose rates when reading vertically. Lose Chefer DeepLift GIM IG LIME SHAP Baseline Win Chefer - 5/6 6/6 4/6 5/6 5/6 6/6 DeepLift 1/6 - 4/9 1/9 5/9 3/9 6/9 GIM 0/6 5/9 - 1/9 5/9 3/9 7/9 IG 2/6 8/9 8/9 - 9/9 7/9 9/9 LIME 1/6 4/9 4/9 0/9 - 2/9 6/9 SHAP 1/6 6/9 6/9 2/9 7/9 - 9/9 Baseline 0/6 3/9 2/9 0/9 3/9 0/9 - Dataset. We evaluate interpretability approaches on three MIMIC-IV clinical tasks [14]: diabetic ketoacidosis (DKA) prediction, mortality prediction, and length-of-stay prediction. Following the StageNet implementation [10], we prepare patient-level data using ICD codes and lab events as features, including both time intervals and clinical measurements. The datasets contain 137,778 patients (mortality), 220,853 samples (length-of-stay), and 179,945 samples (DKA). Due to the computational complexity of SHAP and LIME, we interpret approximately 1,000 randomly selected samples for fair comparison across all methods and tasks. Dataset details are available in our codebase. Models. We train three models for each task: StageNet [10], Transformer [26], and StageAttnâa modified StageNet with an additional multi-head attention layer to examine attentionâs effect on interpretability. All models achieve comparable performance on mortality and DKA prediction. However, on length-of-stay prediction, the Transformer achieves only half the accuracy of StageNet and StageAttn. Performance metrics for all model-task pairs are in Table 3. Baselines. We evaluate six interpretability methodsâChefer, DeepLIFT, GIM (temperature 2.0), Integrated Gradients (50 steps), LIME (200 samples), and Kernel SHAPâacross all model-task pairs. A random baseline assesses whether attribution methods provide meaningful signal by randomly highlighting features. Metrics. We evaluate faithfulness using sufficiency and comprehensiveness [5], two widely-used metrics in the interpretability community. These metrics pose complementary questions: sufficiency measures the predicted softmax probability drop when removing features deemed irrelevant by an interpretability method, while comprehensiveness measures the drop when removing relevant features. Faithfulness across all tasks and models are shown in Figure 1 and runtime in Figure 2. For an overall comparison, Table 2 reports head-to-head win rates across all model-task pairs, scored by ComprehensivenessĂ(1âSufficiency)ComprehensivenessĂ(1-Sufficiency). Figure 1: Faithfulness Benchmark Results. For proper interpretation, (green arrows) higher comprehensiveness and lower sufficiency imply more faithful explanations generated. We observe that Integrated Gradients is consistently a top-performer in terms of interpretation and Chefer emerges as a competitive method for attention-based models, while many of the other approaches move a tremendous degree depending on the task and model. Top performers. Table 2 compares each interpretability method head-to-head across all model-task pairs. Integrated Gradients and Chefer emerge as the most faithful methods overall. Integrated Gradients consistently outperforms most approaches, and when it does not rank first, the margin to the leader is narrow. Chefer performs even more faithfully, outperforming Integrated Gradients in 4 of 6 pairs and all other methods in at least 5 of 6 pairs. Its only exception is the Transformer on length-of-stay prediction, likely due to that modelâs suboptimal training as shown in Table 3. Among black-box methods, SHAP is the most reliable, outperforming other methods in over half of cases and consistently surpassing the random baseline, though it falls behind the gradient-based approaches. Figure 2: Runtime Comparison. For proper interpretation, (green arrows) higher comprehensiveness and lower runtime imply more desirable trade-off. We observe that Shap tends to run significantly longer with minimal advantages compare to others, while Integrated Gradients and Chefer can typically interpret the models better with much shorter runtime. Unreliable methods for clinical time-series. Three methods prove unreliable in practice. DeepLIFT loses to the random baseline in 3 of 9 model-task pairs, all involving attention-based models. This volatility with attention mechanisms, observed by [9], likely stems from its layer-by-layer design [23] conflicting with attentionâs self-repair mechanism [9]. GIM, optimized for attention in language models [9], loses to the random baseline in 2 of 9 pairs and shows no advantage over other methods. It ranks first only on a non-attention model, where it practically reduces to GradientXInput [24], suggesting that substantially larger Transformers may be needed to realize its benefits [9]. LIME performs worst overall, losing to the random baseline in 3 of 9 pairs and consistently trailing other methods. Black-box methods are computationally infeasible at scale. Interpreting all 137,778 mortality prediction samples would require an estimated 300 hours for SHAP and 64 hours for LIME (Figure 2). While sampling fewer points could reduce runtime, this trades off faithfulnessâan unacceptable compromise given that SHAP and LIME already underperform gradient-based methods. We recommend against black-box interpreters for deep clinical predictive models. Integrated Gradients: faithful but costly for non-attention models. While 36% faster than LIME, Integrated Gradients remains computationally expensive for large patient populations. However, its superior faithfulness (Figure 1) justifies this cost for non-attention models. Gradient-weighed attention: the most efficient and faithful approach. The Chefer method [6], which interprets aggregated gradient-weighted attention maps, consistently produces the most faithful attributions across models and tasks. It is also remarkably efficientâapproximately 15 times faster than Integrated Gradients while achieving comparable faithfulness. Notably, adding attention layers to recurrent models does not harm predictive performance (Table 3). While the full Transformer underperforms on clinical time-series tasks, StageAttnâa hybrid combining StageNet with attentionâmatches the original StageNetâs performance. This suggests that incorporating attention layers may be a practical pathway to improving model interpretability without sacrificing accuracy. Table 3: Model Performance Across All Tasks on MIMIC-IV Dataset. Bold values indicate best performance within each task. Acc. = Accuracy, F1-W = F1-Weighted, F1-Ma = F1-Macro, F1-Mi = F1-Micro. Mortality Prediction DKA Prediction Length of Stay Model PR-AUC ROC-AUC Acc. F1 PR-AUC ROC-AUC Acc. F1 Acc. F1-W F1-Ma F1-Mi StageNet 0.6870 0.9576 0.9618 0.6188 0.0789 0.8604 0.9951 0.1682 0.5862 0.5809 0.5694 0.5862 StageAttn 0.6955 0.9441 0.9602 0.6102 0.0783 0.8448 0.9963 0.0571 0.5819 0.5791 0.5741 0.5819 Transformer 0.6134 0.9488 0.9504 0.5267 0.1075 0.8400 0.9968 0.0938 0.2434 0.1800 0.1511 0.2434 5 Future Work Our reproducibility study identifies two key directions for future work. Faithfulness across modalities. Patient profiles are inherently multimodal, containing time-series features, imaging, and clinical notes. While we evaluate interpretability methods for clinical time-series tasks, our findings may not generalize to other modalitiesâfor instance, GIM outperforms many baselines in language modeling [9]. Cross-comparing interpretability approaches across models, tasks, and modalities will help identify domain-specific biases and limitations. Developing improved interpretability approaches. Many interpretability methods are complementary rather than mutually exclusive. Combining these approaches may yield insights for building both more interpretable models and better interpretation methods. By releasing our benchmark in PyHealth, an extensible open-source framework, we enable others to apply and extend these techniques beyond small-scale reproducibility studies for their own tasks. credits 5.0.1 Acknowledgements This study was funded by Jump ARCHES endowment awarded by the Healthcare Engineering Systems Center at the University of Illinois Urbana-Champaign (UIUC) and made possible through the PyHealth Research Initiative. 5.0.2 There are no conflicts of interests here. References [1] G. Abgrall, A. L. Holder, Z. Chelly Dagdia, K. Zeitouni, and X. Monnet (2024) Should ai models be explainable to clinicians?. Critical Care 28 (1), p. 301. Cited by: §1. [2] M. Ancona, E. Ceolini, C. Ăztireli, and M. Gross (2019) Gradient-based attribution methods. In Explainable AI: Interpreting, explaining and visualizing deep learning, p. 169â191. Cited by: §1. [3] L. Bereska and E. Gavves (2024) Mechanistic interpretability for ai safety â a review. External Links: 2404.14082, Link Cited by: §1. [4] A. Brankovic, D. Cook, J. Rahman, S. Khanna, and W. Huang (2024) Benchmarking the most popular xai used for explaining clinical predictive models: untrustworthy but could be useful. Health Informatics Journal 30 (4), p. 14604582241304730. Cited by: Table 1, §2. [5] C. S. Chan, H. Kong, and G. Liang (2022) A comparative study of faithfulness metrics for model interpretability methods. External Links: 2204.05514, Link Cited by: §4. [6] H. Chefer, S. Gur, and L. Wolf (2021) Transformer interpretability beyond attention visualization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 782â791. Cited by: §1, §3, §4. [7] E. Choi, M. T. Bahadori, J. Sun, J. Kulas, A. Schuetz, and W. Stewart (2016) Retain: an interpretable predictive model for healthcare using reverse time attention mechanism. Advances in neural information processing systems 29. Cited by: §3. [8] G. CinĂ , T. E. Röber, R. Goedhart, and Ć. İ. Birbil (2025) Why we do need explainable ai for healthcare. Diagnostic and Prognostic Research 9 (1), p. 24. Cited by: §1. [9] J. Edin, R. CsordĂĄs, T. Ruotsalo, Z. Wu, M. Maistro, C. L. Christensen, J. Huang, and L. MaalĂže (2025) GIM: improved interpretability for large language models. External Links: 2505.17630, Link Cited by: §3, §4, §5. [10] J. Gao, C. Xiao, Y. Wang, W. Tang, L. M. Glass, and J. Sun (2020) Stagenet: stage-aware neural networks for health risk prediction. In Proceedings of the web conference 2020, p. 530â540. Cited by: §1, §4, §4. [11] R. Guidotti, A. Monreale, S. Ruggieri, F. Turini, F. Giannotti, and D. Pedreschi (2018) A survey of methods for explaining black box models. ACM computing surveys (CSUR) 51 (5), p. 1â42. Cited by: §3. [12] J. He, S. L. Baxter, J. Xu, J. Xu, X. Zhou, and K. Zhang (2019) The practical implementation of artificial intelligence technologies in medicine. Nature medicine 25 (1), p. 30â36. Cited by: §1. [13] A. Johannssen and N. Chukhrova (2025) The crucial role of explainable artificial intelligence (xai) in improving health care management. Health Care Management Science 28 (3), p. 565â570. Cited by: §1. [14] A. E. Johnson, L. Bulgarelli, L. Shen, A. Gayles, A. Shammout, S. Horng, T. J. Pollard, S. Hao, B. Moody, B. Gow, et al. (2023) MIMIC-iv, a freely accessible electronic health record dataset. Scientific data 10 (1), p. 1. Cited by: §4. [15] S. Lundberg and S. Lee (2017) A unified approach to interpreting model predictions. External Links: 1705.07874, Link Cited by: §1, §3. [16] C. Meng, L. Trinh, N. Xu, J. Enouen, and Y. Liu (2022) Interpretability and fairness evaluation of deep learning models on mimic-iv dataset. Scientific Reports 12 (1), p. 7166. Cited by: Table 1, §2. [17] J. M. Metsch and A. Hauschild (2025) BenchXAI: comprehensive benchmarking of post-hoc explainable ai methods on multi-modal biomedical data. Computers in Biology and Medicine 191, p. 110124. Cited by: Table 1, §2. [18] M. A. Morid, O. R. L. Sheng, and J. Dunbar (2023) Time series prediction using deep learning methods in healthcare. ACM Transactions on Management Information Systems 14 (1), p. 1â29. Cited by: §1. [19] J. Mullenbach, S. Wiegreffe, J. Duke, J. Sun, and J. Eisenstein (2018) Explainable prediction of medical codes from clinical text. External Links: 1802.05695, Link Cited by: §3. [20] S. Nazir, D. M. Dickson, and M. U. Akram (2023) Survey of explainable artificial intelligence techniques for biomedical imaging with deep neural networks. Computers in Biology and Medicine 156, p. 106668. Cited by: §3. [21] L. N. Pandey, R. Vashisht, and H. G. Ramaswamy (2023) On the interpretability of attention networks. External Links: 2212.14776, Link Cited by: §1, §3. [22] M. T. Ribeiro, S. Singh, and C. Guestrin (2016) Model-agnostic interpretability of machine learning. arXiv preprint arXiv:1606.05386. Cited by: §3. [23] A. Shrikumar, P. Greenside, and A. Kundaje (2017) Learning important features through propagating activation differences. In International conference on machine learning, p. 3145â3153. Cited by: §3, §4. [24] M. Sundararajan, A. Taly, and Q. Yan (2017) Axiomatic attribution for deep networks. External Links: 1703.01365, Link Cited by: §3, §3, §4. [25] Q. Teng, Z. Liu, Y. Song, K. Han, and Y. Lu (2022) A survey on the interpretability of deep learning in medical diagnosis. Multimedia Systems 28 (6), p. 2335â2355. Cited by: §1. [26] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ć. Kaiser, and I. Polosukhin (2017) Attention is all you need. Advances in neural information processing systems 30. Cited by: §1, §4. [27] J. Wang, J. Luo, M. Ye, X. Wang, Y. Zhong, A. Chang, G. Huang, Z. Yin, C. Xiao, J. Sun, et al. (2024) Recent advances in predictive modeling with electronic health records. In IJCAI: proceedings of the conference, Vol. 2024, p. 8272. Cited by: §1. [28] B. Xu and G. Yang (2025) Interpretability research of deep learning: a literature survey. Information Fusion 115, p. 102721. Cited by: §3. [29] P. Zhou, A. Takeuchi, F. Martinez-Lopez, M. Ehghaghi, A. K. Wong, and E. A. Lee (2025) Benchmarking interpretability in healthcare using pattern discovery and disentanglement. Bioengineering 12 (3), p. 308. Cited by: Table 1, §2.