Paper deep dive
Explainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer Models
Melkamu Abay Mersha, Jugal Kalita
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/21/2026, 1:24:58 AM
Summary
The paper introduces the Context-Aware Layer-wise Integrated Gradients (CA-LIG) framework, a unified hierarchical attribution method for explaining Transformer models. Unlike existing methods that rely on final-layer attributions or isolated attention patterns, CA-LIG computes layer-wise Integrated Gradients within each Transformer block and fuses them with class-specific attention gradients. This produces signed, context-sensitive attribution maps that capture supportive and opposing evidence, tracing the hierarchical flow of relevance through the model. The framework is evaluated across diverse tasks including sentiment analysis, document classification, hate speech detection, and image classification, demonstrating superior faithfulness and interpretability compared to established explainability methods.
Entities (10)
Relation Signals (10)
Transformer â hascomponent â Self-Attention
confidence 95% · Built upon multi-head self-attention... deep hierarchical representations
CA-LIG Framework â proposes â Layer-wise Integrated Gradients
confidence 95% · we proposed the Context-Aware Layer-wise Integrated Gradients (CA-LIG) Framework... computes layer-wise Integrated Gradients within each Transformer block
Integrated Gradients â satisfies â completeness property
confidence 94% · The Integrated Gradients method satisfies the completeness property
CA-LIG Framework â fuses â class-specific attention gradients
confidence 93% · fuses these token-level attributions with class-specific attention gradients
CA-LIG Framework â appliesto â Sentiment Analysis
confidence 92% · evaluate the CA-LIG Framework across diverse tasks... including sentiment analysis
CA-LIG Framework â appliesto â Hate Speech Detection
confidence 92% · evaluate the CA-LIG Framework... hate speech detection
CA-LIG Framework â evaluateson â BERT
confidence 90% · evaluate the CA-LIG Framework... with BERT
CA-LIG Framework â evaluateson â XLM-R
confidence 90% · evaluate the CA-LIG Framework... with XLM-R
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Transformer models achieve state-of-the-art performance across domains and tasks, yet their deeply layered representations make their predictions difficult to interpret. Existing explainability methods rely on final-layer attributions, capture either local token-level attributions or global attention patterns without unification, and lack context-awareness of inter-token dependencies and structural components. They also fail to capture how relevance evolves across layers and how structural components shape decision-making. To address these limitations, we proposed the \textbf{Context-Aware Layer-wise Integrated Gradients (CA-LIG) Framework}, a unified hierarchical attribution framework that computes layer-wise Integrated Gradients within each Transformer block and fuses these token-level attributions with class-specific attention gradients. This integration yields signed, context-sensitive attribution maps that capture supportive and opposing evidence while tracing the hierarchical flow of relevance through the Transformer layers. We evaluate the CA-LIG Framework across diverse tasks, domains, and transformer model families, including sentiment analysis and long and multi-class document classification with BERT, hate speech detection in a low-resource language setting with XLM-R and AfroLM, and image classification with Masked Autoencoder vision Transformer model. Across all tasks and architectures, CA-LIG provides more faithful attributions, shows stronger sensitivity to contextual dependencies, and produces clearer, more semantically coherent visualizations than established explainability methods. These results indicate that CA-LIG provides a more comprehensive, context-aware, and reliable explanation of Transformer decision-making, advancing both the practical interpretability and conceptual understanding of deep neural models.
Tags
Links
- Source: https://arxiv.org/abs/2602.16608v2
- Canonical: https://arxiv.org/abs/2602.16608v2
Trouble viewing inline? Open PDF directly â
Full Text
84,427 characters extracted from source content.
Expand or collapse full text
Explainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer Models Melkamu Abay Mersha Jugal Kalita Abstract Transformer models achieve state-of-the-art performance across domains and tasks, yet their deeply layered representations make their predictions difficult to interpret. Existing explainability methods rely on final-layer attributions, capture either local token-level attributions or global attention patterns without unification, and lack context-awareness of inter-token dependencies and structural components. They also fail to capture how relevance evolves across layers and how structural components shape decision-making. To address these limitations, we proposed the Context-Aware Layer-wise Integrated Gradients (CA-LIG) Framework, a unified hierarchical attribution framework that computes layer-wise Integrated Gradients within each Transformer block and fuses these token-level attributions with class-specific attention gradients. This integration yields signed, context-sensitive attribution maps that capture supportive and opposing evidence while tracing the hierarchical flow of relevance through the Transformer layers. We evaluate the CA-LIG Framework across diverse tasks, domains, and transformer model families, including sentiment analysis and long and multi-class document classification with BERT, hate speech detection in a low-resource language setting with XLM-R and AfroLM, and image classification with Masked Autoencoder vision Transformer model. Across all tasks and architectures, CA-LIG provides more faithful attributions, shows stronger sensitivity to contextual dependencies, and produces clearer, more semantically coherent visualizations than established explainability methods. These results indicate that CA-LIG provides a more comprehensive, context-aware, and reliable explanation of Transformer decision-making, advancing both the practical interpretability and conceptual understanding of deep neural models. The implementation code will be made publicly available at https://github.com/melkamumersha/Context-Aware-XAI upon acceptance of the paper. keywords: explainable artificial intelligence (XAI), interpretable deep learning, large language models (LLMs), natural language processing (NLP), explainability techniques, Vision Transformers, computer vision, black-box models. [1]organization=College of Engineering and Applied Science, University of Colorado Colorado Springs, addressline=Colorado Springs, postcode=80918, state=CO, country=USA 1 Introduction Transformer-based architectures, such as BERT [14], GPT [56], and T5 [57], have become foundational across modern natural language processing (NLP) and a wide range of artificial intelligence applications [46, 34, 72]. Built upon multi-head self-attention [73], deep hierarchical representations, and scalable parallel computation, Transformers model long-range dependencies and complex semantic interactions with exceptional effectiveness. Their architectural componentsâincluding self-attention, feedforward sublayers, residual connections, positional encodings, and normalization layersâjointly yield rich contextual representations that support tasks such as sentiment analysis, text classification, multimodal reasoning, and image recognition [60, 47]. Despite their success, Transformer models remain fundamentally opaque. Their layered and nonlinear structure makes it difficult to determine how token representations evolve across layers, how contextual dependencies influence predictions, and how evidence flows through the network. This inherent opacity has intensified the need for eXplainable AI (XAI) techniques capable of producing faithful and interpretable insights into Transformer decision-making. A wide spectrum of explanation approaches has been proposed, yet each exhibits notable limitations. Attention-based methods treat attention distributions as indicators of token importance [41, 76], but numerous studies show that raw attention weights do not reliably reflect model reasoning or provide faithful explanations [30, 63]. Extensions such as Attention Rollout or attention-flow [1] analyses improve traceability across layers but still rely solely on attention weights, omitting contributions from feedforward networks, value vectors, residual pathways, and normalization layers [4]. Gradient-based approaches, including Integrated Gradients (IG) [70] and Guided IG [33], provide theoretically sound input-level attributions but are typically computed only at the final layer, failing to capture how relevance evolves across intermediate Transformer blocks. They also struggle to incorporate structural relationships between tokensâan essential component of Transformer reasoning. Activation-based techniques such as Layer-wise Relevance Propagation (LRP) [6] propagate relevance through network layers but face difficulty preserving relevance conservation across multi-head attention, often producing coarse and context-insensitive attributions [38]. Overall, existing XAI methods suffer from three fundamental limitations. First, most techniques exhibit a final-layer bias, generating explanations only at the final prediction layer while overlooking how semantic information and contextual abstractions are progressively formed across earlier layers of the model. Second, current approaches lack unified localâglobal reasoning: they typically capture either local token-level salience, as in gradient-based methods such as Integrated Gradients, or global structural interactions, as seen in attention-based techniques like attention flow, but rarely integrate both perspectives into a single coherent explanatory representation. Third, prevailing methods offer insufficient context awareness, as they often fail to account for inter-token dependencies, residual connections, feedforward transformations, and cross-layer information flow, all of which are central to the Transformer architecture. These limitations underscore the need for more holistic, context-aware explanation frameworks that reflect the hierarchical and interconnected nature of Transformer language models. To address these limitations, we proposed the Context-Aware Layer-wise Integrated Gradients (CA-LIG) Framework, a unified, hierarchical framework that produces faithful, context-sensitive explanations for Transformer-based models. Prior studies have shown that raw attention weights do not reliably reflect model reasoning or yield faithful explanations. In CA-LIG, attention is not treated as a faithful explanation mechanism nor used in isolation to explain model decisions. Instead, Integrated Gradients serves as the primary attribution method, providing relevance scores that satisfy axiomatic properties such as completeness and sensitivity. Attention gradients are incorporated solely as a contextual interaction signal, capturing how relevance propagates across tokens and layers within the Transformer architecture. Faithfulness is grounded in IG, while attention gradients refine relevance distribution by accounting for inter-token and inter-layer dependencies. Rather than concentrating attribution solely at the modelâs final layer, the CA-LIG Framework computes Layer-wise Integrated Gradients (LIG) at every Transformer block, capturing how token relevance evolves as representations move through the model hierarchy. These layer-wise relevance signals are then fused with class-specific attention gradients, enabling the framework to reflect both local token contributions and global structural dependencies. The resulting attribution maps are signed, relevance-conserving, and sensitive to contextual interactions, capturing supportive and opposing evidence that aligns with the modelâs actual internal reasoning process. Our contributions are summarized as follows: âą We propose a unified and hierarchical XAI framework that captures how token relevance evolves across Transformer layers, enabling layer-wise interpretability rather than restricting explanations to the final output layer. âą We design an integrated gradientâattention attribution mechanism that fuses layer-wise gradients with attention-gradient structures, bridging local token relevance with global interaction patterns. âą We develop a context-aware XAI framework that enforces normalization and relevance preservation across multi-head attention pathways, improving interpretability. âą We conduct a comprehensive empirical evaluation that includes qualitative and quantitative assessments. âą We demonstrate the cross-domain and task generality of our framework by validating it across NLP tasks using BERT, XLM-R, AfroLM, and MAE Vision Transformers for the vision task. The remainder of this paper is organized as follows. Section 2 reviews related work on Transformer explainability. Section 3 introduces the CA-LIG framework, including its layer-wise decomposition and attentionâgradient fusion. Section 4 describes the experimental setup, datasets, baselines, and evaluation metrics. Section 5 presents the empirical results, including qualitative, quantitative, and causal analyses. Limitations and future work are discussed in Section 6. Finally, Section 7 concludes the paper. 2 Related Work 2.1 Explainability Techniques A wide range of explainability methods have been developed to interpret NLP models by identifying the importance of individual input tokens to a modelâs decision [45]. These methods are generally grouped into five main categories: perturbation-based, activation-based, gradient-based, attention-based, and hybrid approaches [50, 49]. Perturbation-based methods work by systematically modifying input features and observing the resulting change in the modelâs output [51]. This class of techniques is model-agnostic and offers intuitive insights. For example, SHAP (SHapley Additive exPlanations) applies game-theoretic principles to assign importance scores to input features [43]. LIME (Local Interpretable Model-Agnostic Explanations) builds a local surrogate model around a specific prediction to explain its rationale [58, 32]. Occlusion Sensitivity identifies important regions by masking parts of the input and analyzing the impact on predictions [79]. Activation-based approaches rely on analyzing internal neuron activations to trace how input features influence model predictions. Techniques such as Class Activation Mapping (CAM) highlight influential regions by combining activation maps with output weights [80]. Layer-wise Relevance Propagation (LRP) propagates relevance scores backward through the network to the input space [9]. Additionally, Concept Activation Vectors (CAVs) quantify how well human-understandable concepts are represented within learned features [35]. Gradient-based methods use gradients of the output with respect to the input to estimate feature contributions. Integrated Gradients (IG) is a gradient-based explainability approach widely used across diverse applications to attribute a modelâs prediction to its input features [64]. IG, for instance, averages gradients computed along a path from a baseline input to the actual input [70]. Integrated Hessians build on Integrated Gradients by characterizing pairwise feature interactions in deep networks under an axiomatic attribution framework [31]. GradientĂInput combines input values with their corresponding gradients to enhance interpretability [65]. FullGrad extends these methods by incorporating both input and bias gradients [68], while Guided IG introduces layer-specific propagation rules to improve attribution clarity [33]. Attention-based explainability techniques take advantage of attention mechanisms in models like Transformers to visualize and quantify input relevance. Attention Flow and Attention Rollout propagate attention scores through layers to trace token influence [81, 1]. These approaches help visualize the pathways through which information flows in deep networks [74]. Hybrid Methods improve faithfulness and by combining multiple explanation signals. Examples include combining gradient and activation maps [62], blending perturbation with activation or attention [10, 65], and integrating gradients with attention mechanisms [55, 77]. AttnLRP [2, 10], which extends Layer-wise Relevance Propagation to Transformer models by explicitly incorporating attention into relevance propagation. These fusion approaches are motivated by the complementary strengths of different explanation sources [50, 52]. 2.2 Explainability for Transformer-based Models Transformer-based architectures introduce substantial challenges for explainability due to their depth, self-attention mechanisms, and distributed contextual representations. Perturbation-based methods, such as SHAP [43] and LIME [58], while intuitive and model-agnostic, are computationally expensive and often generate inconsistent attributions when applied to models with deep and non-linear architectures [20, 63]. Their sampling-based approximations also disrupt the underlying dependencies within token sequences, leading to misleading explanations in Transformer-based models. Activation-based techniques, including Class Activation Mapping (CAM) [80], LRP [9], and Concept Activation Vectors (CAVs) [35], can highlight neuron responses but struggle to provide semantically grounded, fine-grained attributions. This limitation arises from their inability to isolate feature-level interactions or propagate relevance through architectural components like self-attention, residual connections, and normalization layers [11]. Gradient-based methodsâincluding Integrated Gradients (IG) [70] and SmoothGrad [66]âare efficient and theoretically sound. However, these methods typically generate token-level attributions at a final layer without capturing the contextual evolution of token semantics across the model hierarchy. As a result, they fail to faithfully explain how different layers contribute to the final decision, particularly in deeper Transformer stacks where abstract and distributed semantics emerge [29]. Attention-based explanations, such as attention weights, Attention Flow, and Attention Rollout [1, 21], focus on the modelâs internal alignment patterns but are limited in scope. They typically emphasize query-key interactions and ignore the contributions of value vectors, feedforward sublayers, residual connections, and LayerNorm operationsâcomponents that are central to the modelâs reasoning pipeline [10]. Recent Transformer explainability studies have moved beyond raw attention visualization. Flow-based approaches, such as Generalized Attention Flow [8], formulate attribution as a structured flow over attention and gradient signals to better capture cross-layer information routing. AttnLRP [2] extends Layer-wise Relevance Propagation to Transformers by incorporating attention mechanisms into relevance redistribution. Complementarily, Contrast-CAT [24] introduces a contrastive activation-based paradigm that suppresses class-irrelevant features to yield sharper token-level explanations. Furthermore, numerous studies have shown that attention weights alone are insufficient for reliable explanations [30, 75], as they do not necessarily correlate with a modelâs actual decision-making process. To address these limitations, recent efforts have explored relevance propagation extensions. Ali et al. [3] proposed Conservative Propagation, combining LRP and GradientĂInput to enhance stability across LayerNorm and attention heads. Chefer et al. [10] introduced a Taylor-decomposition-inspired relevance flow that integrates residual and feedforward components. Hou et al. [28] proposed Decoding Layer Saliency, identifying saliency-rich Transformer layers but without attributing contextual evolution of relevance across the entire network. Despite these advances, a persistent gap remains: existing techniques either aggregate relevance globally, losing layer-specific nuances, or fail to account for contextual shifts that emerge through hierarchical processing. To overcome this, we propose a hierarchical Context-Aware Layer-wise Integrated Gradient (CA-LIG) framework. Our approach decomposes relevance contributions across each Transformer block, while maintaining alignment with the evolution of contextual tokens. Specifically, we compute intermediate relevance scores at each layer, normalize them within context windows, and integrate them with directional signals from gradient-weighted attention scores. By combining hierarchical relevance decomposition, context-sensitive normalization, and tunable relevance aggregation, our CA-LIG framework advances the state of the art in Transformer explainability, providing more granular, interpretable explanations that address the limitations of existing XAI techniques. 3 Methodology The CA-LIG framework goes beyond final-layer attribution. It consists of four tightly coupled stages. First, Layer-wise Integrated Gradients are computed at each Transformer block to quantify how intermediate hidden representations contribute to the modelâs prediction. Second, class-specific attention gradients are derived to capture how information flows through attention heads and how token interactions influence the output. Third, these two signals are fused through a context-aware integration mechanism, producing unified relevance scores that reflect both local token importance and global structural dependencies. Finally, a context-aware attribution step stabilizes the scores and relevance conservation by rolling relevance across layers to produce the final signed attribution map. Figure 1 provides an overview of the CA-LIG framework. This design enables CA-LIG to capture how evidence forms, evolves, and interacts across the hierarchy of a Transformer layer, yielding faithful, interpretable, and context-sensitive explanations. The following subsections detail each component of the framework. Figure 1: Proposed architecture of the Context-Aware Layer-wise Integrated Gradients (CA-LIG) framework. 3.1 Layer-wise Token Relevance via Integrated Gradients Letâs consider a classifier model M with C output classes and a target class câ1,âŠ,Ccâ\1,âŠ,C\ for which an explanation is desired. The target class c does not need to correspond to the modelâs top prediction; the explanation framework allows generating interpretations for any class of interest. The transformer architecture comprises a stack of N layers or blocks, where x(n)x^(n) denotes the input to the nthn^th layer L(n)L^(n), for n=1,âŠ,Nn=1,âŠ,N. x(1)x^(1) represents the input embeddings of the model, while x(N)x^(N) corresponds to the output of the final encoder layer. Relevance and gradient signals are computed with respect to the target class score ycy_c, allowing the tracking of class-specific evidence throughout the model hierarchy. We employ a Layer-wise Integrated Gradients (LIG) approach to compute token-level attributions in a faithful and interpretable manner. Rather than attribution of feature importance only at the classifier layer, we extend the IG approach to each intermediate layer. For a given layer l, we extract the contextual hidden representation x(l)ââsĂdx^(l) ^sĂ d, where s denotes the sequence length and d is the hidden dimension size. A corresponding baseline representation xâČâŁ(l)ââsĂdx (l) ^sĂ d is defined and serves as a neutral reference point. We construct a trajectory of interpolated hidden states between a baseline and the actual representation of the model for the input to approximate the path integral used in IG. For a given layer l, let x(l)ââsĂdx^(l) ^sĂ d denote the actual hidden representation of the input sequence and xâČâŁ(l)ââsĂdx (l) ^sĂ d be the corresponding baseline. We then compute a sequence of interpolated hidden states using the following equation: x(l)â(αk)=xâČâŁ(l)+αkâ (x(l)âxâČâŁ(l)),where âαk=km,k=1,âŠ,m.x^(l)( _k)=x (l)+ _k· (x^(l)-x (l) ), _k= km, k=1,âŠ,m. (1) Here m is the number of interpolation steps, and αk _k defines the step size along the path from baseline to actual representation. We compute the gradient of the output score ycy_c with respect to each interpolated hidden state and then aggregate these gradients to compute the integrated gradients at layer l using a Riemann sum over m steps [70]: IG(l)=(x(l)âxâČâŁ(l))â1mââk=1mâycâx(l)â(αk),IG^(l)=(x^(l)-x (l)) 1m _k=1^m â y_câ x^(l)( _k), (2) where â denotes element-wise multiplication. We aggregate the feature-wise attributions across the hidden dimensions for each token to obtain token-level relevance scores. For token tâ1,âŠ,stâ\1,âŠ,s\, the relevance score at layer l is given by: Rt(l)=âj=1dIGt,j(l),R_t^(l)= _j=1^dIG_t,j^(l), (3) wherejâ1,âŠ,djâ\1,âŠ,d\ is the hidden dimension indices of a token representation. The use of this summation is grounded in attribution theory [70]. The Integrated Gradients method satisfies the completeness property, which ensures that the sum of all attributions equals the difference in the modelâs output between the input and the baseline. Summing the feature-wise attributions for each token preserves this property and provides a faithful measure of the total influence that a token exerts on the modelâs decision. The result is a signed, layer-wise attribution map, where positive values indicate supportive evidence and negative values indicate opposing influence on the target class decision. 3.2 Compute the Gradient of Attention per Transformer Block Based on token-level relevance scores obtained through LIG, we now examine how token-to-token contextual interactions, captured by the self-attention mechanism, contribute to the modelâs decision for a given target class ycy_c. This approach helps to move beyond the importance of isolated tokens and toward understanding how the model leverages structured dependencies between tokens to form and make its prediction. Each transformer block b contains a self-attention mechanism that produces an attention matrix A(b)ââhĂsĂsA^(b) ^hĂ sĂ s, where h is the number of attention heads and s is the sequence length. The attention score Ai,j,k(b)A_i,j,k^(b) represents how much attention token j pays to token k in the head i of block b, computed in 4. A(b)=softmaxâ(Q(b)â(K(b))â€dk),A^(b)=softmax ( Q^(b)(K^(b)) d_k ), (4) where Q(b)Q^(b) and K(b)K^(b) are the queries and keys at block b, respectively, computed from the input to the block X(b)X^(b), and dkd_k denotes the dimensionality of the key vectors [73]. To quantify the class-specific influence of these attention weights, we compute the gradient of the output score ycy_c with respect to the attention matrix using equation 5. âA(b)=âycâA(b)ââhĂsĂs.â A^(b)= â y_câ A^(b) ^hĂ sĂ s. (5) The âA(b)â A^(b) gradient tensor captures the sensitivity of the prediction of the model to changes in individual attention weights [10]. A high magnitude in âAi,j,k(b)â A_i,j,k^(b) implies that the connection of the attention of the token j to the token k in the head i has a strong influence on the output of the class ycy_c. These gradients provide a class-specific saliency map of the attention structure and allow us to identify token interactions in the input sequence that contribute to the modelâs reasoning. By interpreting this sensitivity map, we reveal the structural dependencies on which the model relies, highlighting not only which tokens are important but also how they interact and influence each other through attention mechanisms. 3.3 Layer-wise Relevance and Attention Gradient Fusion We combine token-level relevance scores obtained from LIG through Equation 3 with attention-based sensitivity maps computed from Equation 5 to create a context-aware explanation. This integration allows for capturing local importance (the individual contribution of each token) and contextual influence (how the relationships between tokens affect the modelâs prediction). For each transformer block b, we perform an element-wise combination of the attention gradients âA(b)ââhĂsĂsâ A^(b) ^hĂ sĂ s with a normalized form of token-level relevance Normâ(R(l))ââsNorm(R^(l)) ^s, where R(l)R^(l) is the relevance vector from layer l, which aligned to the attention mapâs sequence dimension. We apply the Symmetric Min-Max Normalization method. The fusion is defined as: Rcontext(b)=âA(b)âNormâ(R(l)),R_context^(b)=â A^(b) (R^(l)), (6) where â denotes the Hadamard (element-wise) product. This operation weights the attention gradients by the relative importance of each token, thereby enriching the attribution with contextual interactions while preserving the fidelity of individual token contributions. The element-wise fusion in Equation 6 and 7 operates as a relevance-gated sensitivity mechanism that ensures faithfulness in attentionârelevance fusion. Token-level relevance scores obtained via Layer-wise Integrated Gradients provide causally grounded attributions that indicate whether a token contributes to the modelâs prediction, while attention gradients capture how sensitive the prediction is to inter-token attention pathways. This design preserves token-level attribution faithfulness while extending explanations to capture context-aware, interaction-level effects within transformer representations. 3.4 Context-Aware Attribution via Attention-Relevance Fusion To obtain a unified and context-aware attribution map, we aggregate relevance information across attention heads and transformer layers. This aggregation captures both the sensitivity of attention pathways and the contextual contribution of input tokens. To achieve this, we introduce a tunable fusion coefficient λâ[0,1]λâ[0,1] that balances two complementary gradient-based relevance signals. Specifically, for each transformer block bâ1,âŠ,Bbâ\1,âŠ,B\, where B is the total number of blocks, we compute a fused attention relevance matrix as follows: AÂŻ(b)=hâ[λâ (âA(b)âNormâ(R(l)))+(1âλ)â Normâ(R(l))]+I, A^(b)=M_h [λ· (â A^(b) (R^(l)) )+(1-λ)·Norm(R^(l)) ]+I, (7) where, âA(b)ââhĂsĂsâ A^(b) ^hĂ sĂ s denotes the gradient of attention weights at block b with respect to the target score ycy_c, and R(l)ââsR^(l) ^s represents token-level relevance, normalized as Normâ(R(l))Norm(R^(l)), â denotes the element-wise product. The averaging operator hâ[â ]=1hââi=1h(â )M_h[·]= 1h _i=1^h(·) aggregates across heads, and the identity matrix IââsĂsI ^sĂ s preserves residual connections. The main advantage of introducing λ lies in its ability to modulate the degree of influence from the sensitivity of the attention weights and the relevance of the input level. When λ=1λ=1, the explanation balances attention-gradient information and token-level relevance. In contrast, λ=0λ=0, the explanation is purely driven by the input token. This controlled attribution method is useful for transformer architectures, where syntax and semantics are distributed across layers and heads. Each fused matrix AÂŻ(b) A^(b) is row-normalized before composition to maintain bounded influence scores. To trace the flow of information from input to deeper layers, we recursively multiply the normalized relevance-weighted attention matrices across blocks [1, 10]. C=AÂŻ(1)â AÂŻ(2)â âŻâ AÂŻ(B),C= A^(1)· A^(2)·âŠÂ· A^(B), (8) The resulting matrix CââsĂsC ^sĂ s constitutes the final context-aware attribution map, where each entry Ci,jC_i,j indicates the cumulative influence of token j on the representation of token i. This relevance map can be decomposed into positive and negative components to interpret supporting versus opposing contributions to the target prediction. C+=maxâĄ(0,C),Câ=minâĄ(0,C),C^+= (0,C), C^-= (0,C), (9) where C+C^+ reflects supportive evidence and CâC^- indicates inhibitory influence on the predicted class. This decomposition provides fine-grained insight into how different parts of the input promote or counteract the modelâs decision. 3.5 Methodological Comparison of CA-LIG with Prior Methods To understand how CA-LIG differs from existing attribution techniques, a clear methodological comparison across key aspects is helpful. Prior methods such as Integrated Gradients [70], Integrated Hessians [31], and Transformer-LRP [10] each capture different aspects of model reasoning. In contrast, CA-LIG integrates layer-wise Integrated Gradients with class-specific attention-gradient fusion, enabling hierarchical and context-aware explanations that track how evidence evolves across Transformer layers. Table 2 in A provides a consolidated comparison of these methods, highlighting the distinctions that motivate the design of the CA-LIG framework. 4 Experimental Setups To evaluate the effectiveness of our proposed approach, we conducted experiments on multiple encoder-based Transformer language models trained on diverse NLP tasks and datasets. We also evaluated it on the vision task for the domain reproducibility of our approach. We benchmark our approach against widely used explainability techniques. All experiments were conducted on Google Colab using an NVIDIA A100-SXM4 GPU (40GB VRAM), CUDA 12.2, PyTorch 2.0.1, and HuggingFace Transformers 4.35. Integrated Gradients is computed using a linear interpolation path between a zero embedding baseline and the input embedding. We use 50 interpolation steps because they provide sufficient numerical accuracy, stable attributions, and balanced computational efficiency. Token-level relevance scores are normalized using L1 normalization within each layer. Attention rollout depth includes all encoder layers, and attention gradients are aggregated via mean pooling across heads. Sensitivity analysis over 25, 50, and 100 steps showed stable attribution trends. For all experiments, random seeds were fixed to ensure reproducibility, and models were fine-tuned. Fine-tuning was performed for a fixed number of epochs under identical training conditions across all compared XAI methods to ensure fair evaluation. The total data scale is detailed in Section 4.1. Each experimental evaluation uses 20% of each dataset as the held-out test set, with the remaining data allocated to training. To ensure reproducibility, stability, and reliability, all experiments were conducted using independent repeated runs. For each datasetâmodel configuration, experiments were repeated 10 times under fixed random seeds and identical training conditions. Attribution maps were generated independently in each run, enabling a controlled, systematic evaluation of explanation consistency and robustness across repetitions rather than relying on isolated single-run outcomes. 4.1 Datasets We employ well-known NLP and vision datasets to ensure a comprehensive evaluation of our method. In NLP, we use the IMDB movie reviews dataset, which contains 50,000 reviews (25,000 for training and 25,000 for testing) for binary sentiment classification [44]. We also include the Amharic hate speech dataset, an under-resourced language resource consisting of 15,100 annotated samples for hate speech classification [7]. To further examine the modelâs ability to capture long-range contextual dependencies, we incorporate the 20 Newsgroups dataset [39], a collection of long-form documents containing 20 categories. In vision, we used the CIFAR-10 dataset, comprising 60,000 images in 10 classes [37]. We also include the ASIRRA dataset (cat vs dog) [18] for additional evaluation of the quality of the explanation, as its class-defining features are clear and visually interpretable. 4.2 Models For text classification, we use BERT-base [15], which processes sequences of up to 512 tokens, with the [CLS] token passed to a classification head for label prediction. To address low-resource settings, we additionally employ XLM-RoBERTa (XLM-R) [13] and AfroLM [17], a multilingual model effective across under-resourced languages. AfroLM is a self-active-learning multilingual language model for 23 African languages. In the visual task, we use a Masked Autoencoder (MAE) [25]. 4.3 Baseline XAI Methods We compare our proposed approach with well-established attribution techniques commonly used to interpret Transformer-based encoder models. Specifically, we benchmark our method against Input Ă Gradient (IxG) [65], Integrated Gradients [70], and Layer-wise Relevance Propagation (LRP) [9], which represent two primary explanation paradigms: gradient-based and relevance propagation-based. We also include Attention Rollout [1] and Attention Last (Attention-last) [27, 67] to capture cumulative attention flow across layers, providing a complementary perspective to gradient-based attributions. Our proposed approach can be instantiated in multiple variants to explore the effects of contextual and hierarchical attribution across Transformer layers. We experimented with the following three configurations. First, Context-Aware Integrated Gradients (CA-IG) (at last layer), (AÂŻ(B) A^(B) ), applies at the final layer, attributing relevance based on the output representation of the final transformer layer. Second, CA-IG (Layered) computes attributions independently at each layer, providing a detailed layer-wise view of token relevance and shedding light on how contextual representations evolve throughout the layers. Third, CA-LIG aggregates token-level attributions across layers by composing relevance scores using a rollout strategy. This variant captures the hierarchical flow of contextual importance by propagating relevance through the modelâs layered structure while maintaining the contributions of token interactions at each level. Furthermore, our approach supports attribution analysis at selective depths, such as earlier, middle, or deepest layers, highlighting the flexibility and strength of our approach in interpreting layer-wise model behavior. 4.4 Evaluation Settings We adopt a two-fold evaluation strategy for the linguistic domain, combining the ERASER benchmark [16] with the framework proposed by [48]. From ERASER, we use the Movie Reviews dataset for binary sentiment classification, which provides human-annotated rationales and thus enables the assessment of explanation quality against gold references. This evaluation allows us to measure not only the alignment of model explanations with human rationales but also their stability and coherence. Figure 2: CA-LIG token-level attributions for a document labeled Christian class from the 20 Newsgroups dataset using BERT-large. Brighter green tokens provide stronger positive evidence, lighter green indicates weaker support, red shows negative influence, and white denotes neutral relevance. Figure 3: CA-LIG token-level attributions for a document labeled atheist class from the 20 Newsgroups dataset using BERT-base. Brighter green tokens provide stronger positive evidence, lighter green indicates weaker support, red shows negative influence, and white denotes neutral relevance. Figure 4: CA-LIG token-level attributions for a negative IMDB review using BERT-Large. Brighter red indicates stronger negative evidence, green indicates positive relevance, and white denotes neutral tokens. 5 Results We present the results of the proposed CA-LIG framework, beginning with qualitative visualizations, followed by quantitative comparisons with baseline methods, and concluding with layer-wise attribution case analyses that highlight the effectiveness of our proposed framework across various Transformer models and tasks. All experiments are conducted with λ=1λ=1, which provides a balanced fusion of attention-gradient information and token-level relevance since our objective is to generate context-aware explanations. We examine the explanations generated by CA-LIG and the baseline methods for various tasks, datasets, models, and domains. Figures 4 and 4 present explanations for long-form, multi-class document classification of the 20 Newsgroups dataset using BERT-base and BERT-large models, which demonstrate CA-LIGâs ability to capture long-range contextual dependencies and propagate relevance coherently across long documents. Figure 4 presents the negative IMDB sentiment explanation generated by CA-LIG with the BERT-large model. Figures 11 show token-level relevance patterns for positive IMDB sentiment using BERT-base, comparing CA-LIG with baseline methods. Figures 6 and 6 evaluate CA-LIG in under-resourced language settings, showcasing Amharic hate-speech detection using XLM-R and AfroLM models. These results confirm that CA-LIG produces stable relevance scores even in morphologically rich, low-resource languages. Figure 9 provides explanations generated with CA-LIG using the MAE model, demonstrating the flexibility of CA-LIG beyond text and highlighting how the framework generalizes to vision tasks. We observe that attention-based methods, Figure 11, such as A-last and A-Rollout, tend to assign relatively uniform relevance scores across all tokens, with particular amplification of salient tokens like âamazingâ and âabsolutelyâ, yet they provide limited contrast in importance among the remaining tokens. IxG, LRP, and IG provide more selective relevance, with LRP focusing heavily on âamazingâ and IG assigning moderate negative relevance to âthisâ and âmovieâ. Interestingly, IG also overemphasizes the [CLS] token, which may dilute the interpretability of specific input tokens. Our proposed methods show clearer attribution patterns. CA-IG (last layer) assigns strong positive relevance to sentiment-bearing words (âabsolutelyâ and âamazingâ) while moderately suppressing less informative tokens. CA-LIG further refines this by progressively aggregating relevance across layers, resulting in sharp, focused attribution on the most sentiment-relevant terms. Unlike baseline methods, our CA-IG variant avoids over-weighting special tokens such as [CLS] by redistributing relevance across contextually interacting tokens, thereby preventing special-token dominance and offering a more intuitive, human-aligned explanation of model behavior. These qualitative observations support the effectiveness of our context-aware method in isolating semantically important words and capturing deeper attribution flow across the Transformerâs architecture. Furthermore, we visualize the head and layer-wise attribution heatmaps derived from C to identify patterns of structural and contextual importance, which aid in model interpretation, debugging, and trust calibration. Supplementary experiments are provided in B. Figure 5: CA-LIG token-level attributions for an Amharic hate speech sample using the XLM-R model. Brighter red indicates stronger negative evidence, green indicates positive relevance, and white denotes neutral tokens. Figure 6: CA-LIG token-level attributions for an Amharic not hate speech sample using the AfroLM model. Brighter green indicates stronger positive evidence, red indicates negative relevance, and white denotes neutral tokens. Figure 7: Token-F1 scores on the Movie Reviews reasoning task. CA-LIG and CA-IG(last layer) achieve consistently higher performance than baseline methods. Figure 8: Layer-wise sensitivity profiles for CA-LIG, classifier contribution, and mean attention to [CLS]. Transitional peaks at Layers 4 and 8 reflect key representational shifts. The final peak at Layer 12 corresponds to sentiment consolidation. Figure 9 shows a qualitative comparison for image classification. Baseline methods often highlight scattered or background regions, which limit interpretability. In contrast, our approach yields more coherent, concentrated visualizations, emphasizing areas directly relevant to the predicted class. Input Grad-CAM [61] LRP-ϔΔ [9] LRP(Attn+Grad) [10] IG [70] CA-LIG (Ours) Figure 9: Qualitative comparison of explanations generated by baseline XAI methods and CA-LIG using a MAE model. Warmer colors denote regions with higher positive relevance, while cooler colors indicate lower relevance. Images are taken from the ASIRRA dataset [18]. (a) Original Input (b) CA-LIG Explanation (c) Positive Attribution (helps prediction) (d) Negative Attribution (hinders prediction) Figure 10: Example of an explanation generated using CA-LIG for a prediction made by MAE model. (a) Original input image, (b) CA-LIG explanation heatmap, (c) positively attributed regions, and (d) negatively attributed regions. Warmer colors indicate stronger relevance. The CA-LIG attribution map highlights how BERT-Large captures semantically dependent tokens, as shown in Figure 4, even when these tokens appear far apart in the input sequence, demonstrating CA-LIGâs ability to capture long-range contextual structure. Strong positive attributions on tokens such as God, bible, timing, history, world, believers, christ, and close indicate that CA-LIG identifies the theological and eschatological semantics that characterize the âChristianâ class in the 20 Newsgroups dataset. Importantly, CA-LIG assigns high relevance not only to isolated words but also to concept pairs that co-occur across multiple clausesâfor example, evidence â bible, return â christ, day â lord, and end â world. These relationships show that CA-LIG does not rely solely on surface-level frequency but instead aggregates layer-wise signals that encode deep contextual interactions spanning entire sentences. For instance, the model links âevidence in the bibleâ with âtiming of the history of the worldâ, capturing a multi-clause reasoning chain that reflects Christian doctrinal themes, tokens that are not adjacent but are contextually unified. Similarly, negative relevance on words such as thief and night shows that CA-LIG appropriately down-weights metaphorical elements that do not directly characterize the Christian topic category. Figure 4 illustrates that CA-LIG assigns strong positive relevance (green) to multiple interacting topic-defining tokens such as âatheist,â âbelief,â âsystem,â âlack belief,â and âgods,â which together establish the documentâs ideological orientation. Instead of relying on a single lexical cue, CA-LIG captures how belief-related concepts and negation structures (e.g., âlack belief,â ânever seen convincing evidenceâ) collectively drive the prediction, while supporting tokens receive weaker relevance and semantically peripheral tokens remain neutral. Figure 4 presents CA-LIG attributions for a negative IMDB review using BERT-large, where strong negative relevance (red) is assigned to interacting sentiment-bearing expressions such as âworst,â âlame,â âpoor acting,â and âthings wrong.â The explanation shows that negative sentiment emerges from interactions across multiple evaluative tokens rather than a single word, with neutral or context-setting tokens remaining unhighlighted. Figures 6 and 6 illustrate CA-LIG attributions for an Amharic hate speech sample using XLM-R and AfroLM models, respectively. CA-LIG assigns strong negative relevance to explicitly abusive tokens (Figure 8), particularly those corresponding to âthese stupidâ and âshould be excluded,â which form the semantic core of the hate expression. Contextually reinforcing modifiers receive additional relevance, while syntactically necessary but semantically neutral tokens remain white. These results demonstrate that CA-LIG explains model predictions through contextual interactions among semantically aligned tokens, capturing compositional meaning across clauses, sentiment expressions, and targetâintent structures across models, domains, and languages. By combining layer-wise Integrated Gradients with attention-gradient fusion and context propagation, CA-LIG captures long-distance relevance patterns that reflect the hierarchical structure of the modelâs reasoning, highlighting global semantic coherence rather than relying on local tokens alone. In vision tasks, input images are resized to 224Ă224224Ă 224 and split into distinct 16Ă1616Ă 16 patches, which are then linearly embedded into a sequence that includes a [CLS] token; a classifier head employs its terminal representation. Context is as crucial for vision tasks as it is for language, yet many XAI attribution methods overlook it, leading to fragmented or noisy explanations. Figure 16 illustrates how the CA-LIG approach addresses this limitation by grounding explanations in semantically coherent features. Unlike baseline methods that highlight scattered pixel intensities, CA-LIG emphasizes relationally meaningful regions (e.g., the catâs ears, eyes, and muzzle) that are directly related to object identity. Positive attributions capture features that drive the prediction (eyes, nose, whiskers), as shown in Figure 16c, while negative attributions highlight distracting regions (background textures, fur patterns), as shown in Figure 16d, offering counterfactual insights into model uncertainty. This contextual coherence improves interpretability by aligning with human reasoning, enhances robustness against noise and adversarial perturbations, and aids generalization by disambiguating class-relevant from irrelevant features. Overall, CA-LIG demonstrates that context is a universal requirement for explainability, bridging NLP and vision under a unified framework for context-aware XAI. 5.1 Evaluation 5.1.1 Qualitative Evaluation Figure 11 presents a qualitative comparison of explanation visualizations for text classification. The baseline methods often emphasize irrelevant or dispersed tokens, resulting in noisy and less interpretable explanations. In contrast, our CA-LIG approach provides sharper and more focused attributions. XAI Explanation outputs A-last [27] A-Rollout [1] IxG [65] LRP [9] IG [70] CA-IGlast (Ours) CA-LIG (Ours) Figure 11: Qualitative comparison of explanations produced by baseline XAI methods and CA-LIG using a BERT-base model. Brighter green indicates stronger positive relevance, red indicates negative relevance, and white represents neutral tokens. 5.1.2 Quantitative Evaluation We evaluate explanation quality on the Movie Reviews rationale benchmark [78, 16], which provides human-annotated rationale tokens for assessment. We report token-F1, measuring the overlap between the top-ranked tokens and the gold rationales [10], using a length-normalized evaluation that selects the top-p% of tokens (pâ5,10,15,20,30,40,50pâ\5,10,15,20,30,40,50\), with a minimum of five tokens per instance. Figure 7 shows that while all methods improve as more tokens are included, our CA-LIG approach consistently achieves higher token-F1 than the baselines. In the vision task, we assess the faithfulness of explanations using perturbation-based AUC, evaluated through the insertion and deletion of the most important patches, results presented in Table 1. The image patches are ranked by importance and then progressively inserted from a blurred baseline or deleted from the original input image. The perturbation curves are computed using a length-normalized procedure that updates 5% of the patches per step, over 20 steps ranging 0â100%. A faithful explanation produces a rapid increase in confidence in insertion (larger AUC) and a decrease in confidence in deletion (smaller AUC). In the cat and dog examples, our method consistently emphasizes class-defining regions (e.g., eyes, muzzle), resulting in sharper insertion gains and stronger deletion declines compared to baselines. Perturbation Class GradCAM [61] LRP-ϔΔ [9] LRP (Attn+Grad) [10] IG [70] Ours Insertion â Predicted 0.424 0.448 0.567 0.473 0.617 Deletion â Predicted 0.374 0.437 0.237 0.346 0.215 Table 1: Perturbation-based AUC evaluation of explanation faithfulness. Higher insertion AUC (â) indicates faster confidence recovery when important patches are added to a blurred baseline, while lower deletion AUC (â) reflects faster confidence drop when critical patches are removed from the original image. 5.2 Case Study: Layer-wise Sensitivity Analysis of CA-LIG Transformer-based architectures, such as BERT, demonstrate a hierarchical progression of linguistic representations across their layers. Empirical analyses have shown that the early layers (1â4) primarily encode surface-level syntactic features and grammatical relations [71, 26, 23, 5, 22], the middle layers (5â8) capture semantic meaning and contextual dependencies [40, 12, 53, 42], and the deeper layers (9â12) consolidate sentence-level information to support task-specific decision making, particularly through refinement of the [CLS] representation [69, 19, 36]. To empirically investigate how our proposed CA-LIG framework aligns with this hierarchical organization, we conducted a case study on the input sentence âThis movie was absolutely amazingâ using a BERT-base model fine-tuned on the IMDB sentiment classification dataset. The goal of this study is to trace how token-level relevance evolves across layers and how it interacts with classifier alignment and attention dynamics. For each Transformer layer L, we extracted three complementary interpretability signals: (1) CA-LIG token-level relevance scores; (2) classifier contribution computed by projecting the hidden state hLh_L of each token onto the final classification layer as CL=hLâ Wc+bcC_L=h_L· W_c+b_c; and (3) mean attention to the [CLS] token, averaged across all heads, to assess shifts in the focus of aggregation. As shown in Figure 8, the layer-wise sensitivity patterns exhibit three distinct phases, aligned with the structural-functional hierarchy of BERT described in prior studies. In the surface feature extraction stage (Layers 1â4), CA-LIG relevance scores remain low, reflecting minimal token-level differentiation. A modest increase is observed at Layer 4, coinciding with an increase in the classifier contribution, suggesting the early stages of representational alignment with the decision space [59]. Attention to the [CLS] token remains relatively stable, indicating that attribution changes are driven primarily by representational transformation rather than attention reweighting. Attributions in this stage are uniformly distributed across tokens, with a limited emphasis on either syntactic or sentiment-bearing terms, consistent with the role of these layers in capturing POS and dependency structures [26, 23]. In the contextual reasoning stage (Layers 5â8), CA-LIG scores increase sharply, peaking at Layer 8, while classifier contributions remain elevated. Token-level relevance patterns reveal a de-emphasis on functional terms such as âwasâ and increased attribution toward sentiment-rich tokens like âabsolutelyâ and âamazingâ, as shown in Figure 18. This observation aligns with findings that these middle layers encode compositional semantics and contextual refinements [40, 12, 54]. Despite this semantic refinement, attention to [CLS] remains a kind of stable, further supporting that internal representational adjustmentsânot shifts in attention flowâaccount for the observed relevance dynamics. In the decision consolidation stage (Layers 9â12), CA-LIG scores briefly dip at Layer 10 before surging to their highest level at Layer 12. This peak co-occurs with the maximum classifier contribution, indicating the final integration of task-relevant evidence into the [CLS] embedding, a pattern in line with deep-layer specialization for decision-making [69, 36]. Here, relevance becomes sharply concentrated on the sentiment-bearing tokens, highlighting their role in determining model predictions. Despite this semantic convergence, attention to [CLS] remains relatively stable, suggesting that the model relies more on internal representational refinement than attention flow adjustments. Overall, this analysis reveals that CA-LIG effectively captures the hierarchical progression of interpretability across Transformer layers. The observed transition points at Layers 4 and 8 align with prior probing studies [71, 40], supporting the interpretive utility of CA-LIG in determining how contextual and task-relevant signals emerge within the model. Notably, CA-LIG token relevance trends closely mirror classifier contributions at key semantic layers, confirming its ability to track decision-aligned transformations. The weak alignment with [CLS] attention highlights that interpretability emerges from representational shifts rather than attention redistribution alone. These insights reinforce the view that interpretability techniques should account for the dynamic flow of information across layers, rather than relying solely on final-layer explanations. 6 Limitations and Future Work While the CA-LIG framework provides a comprehensive and context-aware approach to interpreting Transformer decoder-based models, some limitations remain. First, the framework is designed for encoder-only models and has not yet been experimented with decoder-based language models. Second, the fusion coefficient λ is manually tuned, requiring adaptive or learnable strategies. Third, while CA-LIG demonstrates strong performance on text-based tasks, its evaluation in the vision domain remains limited. More extensive experiments involving multiple vision models, tasks, and diverse datasets are required to fully assess its effectiveness in visual settings. Additionally, this study does not address multimodal Transformer explainability, in which cross-modal attention pathways introduce distinct interpretability challenges. In future work, we plan to address these limitations by extending CA-LIG to vision-based, decoder-based, and multimodal Transformer models, developing adaptive fusion mechanisms, and conducting broader evaluations across language, vision, and multimodal domains. 7 Conclusion Despite their strong performance, Transformer models remain difficult to interpret. Most existing XAI methods often rely on final-layer attributions, focus on either local gradients or global attention patterns, lack explicit context-awareness, and fail to preserve relevance consistently across layers. These limitations prevent faithful and structurally coherent explanations of Transformer decision-making. In this study, we introduced the CA-LIG Framework, a unified and hierarchical attribution framework designed to address these gaps. CA-LIG differs from conventional single-layer attribution by applying Layer-wise Integrated Gradients at every Transformer block, thereby revealing how token relevance emerges, evolves, and stabilizes as the network deepens. By incorporating class-specific attention gradients, CA-LIG captures the contextual dependencies and structural information flow that shape model predictions. The framework fuses these complementary signals through a context-aware integration mechanism and aggregates them using a relevance rollout procedure that preserves attribution stability and conserves the relevance across layers. Extensive experiments on sentiment analysis, low-resource hate speech detection, long-form and multiclass document classification, and image classification tasks demonstrate that CA-LIG consistently outperforms established XAI techniques. Both quantitative evaluations and qualitative analyses confirm that CA-LIG produces explanations that are not only more reliable but also more aligned with the hierarchical computation that defines Transformer architectures. CA-LIG enhances Transformer explainability by integrating completeness, context-awareness, and hierarchical fidelity. By tracing relevance across layers and unifying token-level and structural dependencies, it generates explanations that are more faithful, coherent, and aligned with human reasoning. This study presents a meaningful step toward building transparent, interpretable, and trustworthy Transformer-based models. Appendix A Methodological Comparison of CA-LIG with Prior Methods Methodologically, the CA-LIG Framework occupies a distinct space within the spectrum of Transformer interpretability approaches. Table 2, in A, presents a comparison of CA-LIG with the most widely used Transformer interpretability methods. Aspect IG [70] Integrated Hessians [31] Transformer-LRP [10] CA-LIG Framework (Ours) Core Principle Path-integrated gradients from baseline to input Second-order IG using Hessians for interaction effects LRP/Deep Taylor rules Layer-wise IG fused with class-specific attention gradients Primary Objective Token-level importance at input layer. Quantify pairwise feature interactions. Propagate and conserve relevance across attention and residuals. Produce hierarchical, context-aware token relevance across layers. Derivative Order First-order Second-order Rule-based First-order IG with gradient-weighted attention integration Granularity Input-level token salience Interaction matrix between feature pairs Single global input-level relevance map Per-layer token attribution + fused contextual map + multi-layer rollout Use of Attention None None Implicit via LRP propagation rules Explicit class-specific attention gradients. Layer-wise Tracking None (only final output layer) None (interaction-focused) Relevance flows through layers Explicit LIG at each layer to model relevance trajectories. Relevance Conservation Completeness guaranteed for IG No global conservation guarantees Strict conservation via LRP rules Approximate conservation across layers with normalized fusion Context Awareness Weak (only at final layer) Indirect via interaction structure Partially context-aware via attention propagation rules Explicit capturing of structural context via attention gradients + LIG Output Representation Single saliency map over tokens Interaction saliency matrix Single relevance heatmap Signed multi-layer relevance maps + fused final explanation Computational Cost Low (only final layer) High (mixed Hessian integrals) Moderate ModerateâHigh (per-layer IG + attention gradients) Best Use Cases General-purpose saliency Interaction analysis for important pairs Attention-heavy Transformer interpretability Hierarchical reasoning, context flow, fine-grained token relevance Table 2: Methodological comparison of IG, Integrated Hessians, Transformer-LRP, and the proposed CA-LIG framework, highlighting how CA-LIG combines layer-wise Integrated Gradients with attention-gradient fusion to produce hierarchical, context-aware explanations. Appendix B Supplementary Experimental Results Figure 12: CA-LIG token-level attributions for a positive IMDB review using BERT-Large. Brighter green indicates stronger positive evidence, red indicates negative relevance, and white denotes neutral tokens. Figure 13: CA-LIG token-level attributions for a document labeled politics class from the 20 Newsgroups dataset using BERT-base. Brighter green tokens provide stronger positive evidence, lighter green indicates weaker support, red shows negative influence, and white denotes neutral relevance. Figure 14: Token-level attribution visualization generated by the CA-LIG using the AfroLM model on an Amharic hate speech sample. Brighter Red tokens provide stronger negative evidence (hate); lighter red indicates weaker support; green shows positive influence; and white denotes neutral relevance. English translation: That religion is not respected, all its followers are destructive, and their behavior shows that they do not like peace anywhere. Figure 15: Token-level attribution visualization generated by the CA-LIG using the XLM-R model on an Amharic hate speech sample. Brighter green tokens provide stronger positive evidence (not hate), lighter green indicates weaker support, red shows negative influence, and white denotes neutral relevance. English translation: It is important to believe in the rule of law. Anyone who breaks the law must be held accountable. (a) Original Input (b) CA-LIG Explanation (c) Positive Attribution (helps prediction) (d) Negative Attribution (hinders prediction) Figure 16: Example of a CA-LIG explanation for a prediction made by the MAE model on the CIFAR-10 dataset. (a) Original input image, (b) CA-LIG explanation heatmap, (c) positively attributed regions, and (d) negatively attributed regions. Warmer colors indicate stronger relevance to the modelâs prediction. The CA-LIG explanation indicates that the MAE modelâs predictions are driven by contextually related and semantically meaningful object parts. As illustrated in (b) and (c), strong positive relevance is consistently assigned to the head, torso, and lower body/feet. In contrast, (d) indicates negative relevance to background and peripheral regions that do not contribute to the prediction. Input Grad-CAM [61] LRP-ϔΔ [9] LRP(Attn+Grad) [10] IG [70] CA-LIG (Ours) Figure 17: Qualitative comparison of explanations generated by various XAI methods using a Masked Autoencoder (MAE) fine-tuned on the CIFAR-10 dataset for image classification. Warmer colors denote regions with higher positive relevance to the predicted class, while cooler colors indicate lower relevance. Sample images are taken from the ASIRRA cat vs. dog dataset [18]. Figure 18: Layer-wise token relevance visualization generated by our proposed CA-IG method using a BERT-base model fine-tuned on IMDB sentiment analysis. Brighter green indicates stronger positive evidence, red indicates negative relevance, and white denotes neutral tokens. The progression shows early layers distributing relevance across tokens, middle layers emphasizing sentiment-bearing words (absolutely, amazing), and deeper layers consolidating decision-relevant cues. Appendix C Computational Complexity Analysis We analyze the computational cost of the proposed CA-LIG framework in comparison with standard Integrated Gradients. Integrated Gradients approximates the path integral between a baseline input and the actual input using m interpolation steps. At each step, a gradient of the model output with respect to the input is computed. The time complexity of IG is therefore: Oâ(mâ Cgrad),O(m· C_grad), (10) where m denotes the number of interpolation steps and CgradC_grad represents the cost of a single gradient computation, which is approximately equivalent to one forward and one backward pass through the model. Consequently, IG explanations are roughly m times more expensive than a single forward-pass inference. In practice, m is typically chosen between 20 and 300, with 50 being a common default that balances numerical accuracy and computational efficiency. The CA-IG (last-layer) variant extends IG by incorporating class-specific attention gradients as a contextual interaction signal at a single Transformer layer. While this introduces additional gradient computations for attention weights, the dominant cost remains the IG computation itself. As a result, the overall complexity remains: Oâ(mâ Cgrad),O(m· C_grad), (11) with a small constant-factor overhead due to attention-gradient extraction. The layer-wise approach computes IG attributions across all L encoder layers of the Transformer. Since IG is applied independently at each layer, the time complexity scales linearly with the number of layers: Oâ(Lâ mâ Cgrad).O(L· m· C_grad). (12) This formulation enables hierarchical attribution analysis at the cost of increased computation. The full CA-LIG framework further incorporates relevance rollout across layers by combining layer-wise IG relevance with attention-gradient-based contextual propagation. Although attention rollout itself is computationally inexpensive relative to gradient computation, CA-LIG requires computing attention gradients at each layer in addition to IG. As a result, the overall complexity remains dominated by IG and can be expressed as: Oâ(Lâ mâ Cgrad).O(L· m· C_grad). (13) The additional cost of relevance aggregation and rollout introduces only minor constant-factor overhead. Due to its linear dependence on both the number of layers L and interpolation steps m, CA-LIG is computationally more expensive than standard IG or other final-layer explainer methods. In practice, CA-LIG incurs approximately LĂmLĂ m forwardâbackward passes per input, making it most suitable for offline analysis, such as model auditing, explanation benchmarking, and error analysis. While CA-LIG introduces additional computational overhead, this cost is justified by its ability to provide faithful, context-aware, and hierarchical explanations for Transformer-based models, thereby overcoming the limitations of existing XAI methods. Despite a reasonable increase in computational cost, CA-LIG provides substantial advantages over conventional attribution methods. Context-aware explanations extracted from the final encoder layer are computationally efficient, as they rely on a single backward pass over already computed contextual representations, making them comparable in cost to standard Integrated Gradients. Note that IG is applied only at the final layer. Beyond this, CA-LIG enables layer-wise context-aware explanations, allowing relevance to be analyzed at different depths of the Transformer. This layered perspective reveals how token importance evolves from early lexical and syntactic representations to higher-level semantic and task-specific reasoning. As a result, CA-LIG offers richer, more faithful, and hierarchically structured explanations that better reflect the internal reasoning processes of Transformer-based models, justifying the additional computational overhead. References [1] S. Abnar and W. Zuidema (2020) Quantifying attention flow in transformers. arXiv preprint arXiv:2005.00928. Cited by: §1, §2.1, §2.2, §3.4, §4.3, Figure 11. [2] R. Achtibat, S. M. V. Hatefi, M. Dreyer, A. Jain, T. Wiegand, S. Lapuschkin, and W. Samek (2024) Attnlrp: attention-aware layer-wise relevance propagation for transformers. arXiv preprint arXiv:2402.05602. Cited by: §2.1, §2.2. [3] A. Ali and A. Kumar (2022) XAI methods for transformers via conservative propagation. In ICLR, Cited by: §2.2. [4] A. K. AlShami, R. Rabinowitz, K. Lam, Y. Shleibik, M. Mersha, T. Boult, and J. Kalita (2025) Smart-vision: survey of modern action recognition techniques in vision. Multimedia tools and applications 84 (27), p. 32705â32776. Cited by: §1. [5] T. Aoyama and N. Schneider (2022) Probe-less probing of BERTâs layer-wise linguistic knowledge with masked word prediction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop, p. 195â201. Cited by: §5.2. [6] L. Arras, G. Montavon, K. MĂŒller, and W. Samek (2017) Explaining recurrent neural network predictions in sentiment analysis. arXiv preprint arXiv:1706.07206. Cited by: §1. [7] A. A. Ayele, S. M. Yimam, T. D. Belay, T. Asfaw, and C. Biemann (2023) Exploring amharic hate speech data collection and classification approaches. In Proceedings of the 14th international conference on recent advances in natural language processing, p. 49â59. Cited by: §4.1. [8] B. Azarkhalili and M. W. Libbrecht (2025) Generalized attention flow: feature attribution for transformer models via maximum flow. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 19954â19974. Cited by: §2.2. [9] S. Bach, A. Binder, G. Montavon, F. Klauschen, K. MĂŒller, and W. Samek (2015) On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PloS one 10 (7), p. e0130140. Cited by: Figure 17, §2.1, §2.2, §4.3, Figure 11, Figure 9, Table 1. [10] H. Chefer, S. Gur, and L. Wolf (2021) Transformer interpretability beyond attention visualization. In CVPR, Cited by: Table 2, Figure 17, §2.1, §2.2, §2.2, §3.2, §3.4, §3.5, Figure 9, §5.1.2, Table 1. [11] Z. Chen, Y. Xie, Y. Wu, Y. Lin, S. Tomiya, and J. Lin (2024) An interpretable and transferrable vision transformer model for rapid materials spectra classification. Digital Discovery 3 (2), p. 369â380. Cited by: §2.2. [12] K. Clark, U. Khandelwal, O. Levy, and C. D. Manning (2019) What does BERT look at? an analysis of BERTâs attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, p. 276â286. Cited by: §5.2, §5.2. [13] A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. GuzmĂĄn, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov (2019) Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116. Cited by: §4.2. [14] J. Devlin, M. Chang, K. Lee, and K. Toutanova (2018) Bert: pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. Cited by: §1. [15] J. Devlin, M. Chang, K. Lee, and K. Toutanova (2019) Bert: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), p. 4171â4186. Cited by: §4.2. [16] J. DeYoung, S. Jain, N. F. Rajani, E. Lehman, C. Xiong, R. Socher, and B. C. Wallace (2019) ERASER: a benchmark to evaluate rationalized nlp models. arXiv preprint arXiv:1911.03429. Cited by: §4.4, §5.1.2. [17] B. F. Dossou, A. L. Tonja, O. Yousuf, S. Osei, A. Oppong, I. Shode, O. O. Awoyomi, and C. Emezue (2022) AfroLM: a self-active learning-based multilingual pretrained language model for 23 african languages. In Proceedings of The Third Workshop on Simple and Efficient Natural Language Processing (SustaiNLP), p. 52â64. Cited by: §4.2. [18] J. Elson, J. R. Douceur, J. Howell, and J. Saul (2007) Asirra: a captcha that exploits interest-aligned manual image categorization.. CCS 7 (366-374), p. 15. Cited by: Figure 17, Figure 17, §4.1, Figure 9, Figure 9. [19] K. Ethayarajh (2019) How contextual are contextualized word representations? comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), p. 55â65. Cited by: §5.2. [20] M. Fantozzi et al. (2024) Explainability in deep learning: challenges for transformers. Frontiers in Artificial Intelligence. Cited by: §2.2. [21] J. Ferrando, G. Sarti, A. Bisazza, and M. R. Costa-JussĂ (2024) A primer on the inner workings of transformer-based language models. arXiv preprint arXiv:2405.00208. Cited by: §2.2. [22] J. Ferrando (2022) Measuring the mixing of contextual information in the transformer. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Cited by: §5.2. [23] Y. Goldberg (2019) Assessing BERTâs syntactic abilities. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, p. 3623â3632. Cited by: §5.2, §5.2. [24] S. Han, J. Lee, and S. Lee (2025) Contrast-cat: contrasting activations for enhanced interpretability in transformer-based text classifiers. arXiv preprint arXiv:2507.21186. Cited by: §2.2. [25] K. He, X. Chen, S. Xie, Y. Li, P. DollĂĄr, and R. Girshick (2022) Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 16000â16009. Cited by: §4.2. [26] J. Hewitt and C. D. Manning (2019) A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), p. 4129â4138. Cited by: §5.2, §5.2. [27] N. Hollenstein and L. Beinborn (2021) Relative importance in sentence processing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), p. 141â150. Cited by: §4.3, Figure 11. [28] E. M. Hou and G. D. Castanon (2023) Decoding layer saliency in language transformers. In International Conference on Machine Learning, p. 13285â13308. Cited by: §2.2. [29] S. Jain et al. (2023) Inseq: a toolkit for sequence-level interpretability of nlp models. Note: https://github.com/penwang/inseq Cited by: §2.2. [30] S. Jain and B. C. Wallace (2019) Attention is not explanation. arXiv preprint arXiv:1902.10186. Cited by: §1, §2.2. [31] J. D. Janizek, P. Sturmfels, and S. Lee (2021) Explaining explanations: axiomatic feature interactions for deep networks. Journal of Machine Learning Research 22 (104), p. 1â54. Cited by: Table 2, §2.1, §3.5. [32] D. Kamen, M. A. Mersha, and J. Kalita (2025) Introducing semantic feature dependencies in nlp xai systems with suplime. In Recent Advances in Natural Language Processing, p. 47. Cited by: §2.1. [33] A. Kapishnikov, S. Venugopalan, B. Avci, B. Wedin, M. Terry, and T. Bolukbasi (2021) Guided integrated gradients: an adaptive path method for removing noise. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 5050â5058. Cited by: §1, §2.1. [34] S. Khapre, M. A. Mersha, H. Shakil, J. Baruah, and J. Kalita (2025) Toxicity in online platforms and ai systems: a survey of needs, challenges, mitigations, and future directions. Expert Systems with Applications, p. 129832. Cited by: §1. [35] B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas, et al. (2018) Interpretability beyond feature attribution: quantitative testing with concept activation vectors (tcav). In International conference on machine learning, p. 2668â2677. Cited by: §2.1, §2.2. [36] O. Kovaleva, A. Romanov, A. Rogers, and A. Rumshisky (2019) Revealing the dark secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), p. 4365â4374. Cited by: §5.2, §5.2. [37] A. Krizhevsky, G. Hinton, et al. (2009) Learning multiple layers of features from tiny images. Cited by: §4.1. [38] T. Lan, J. Xu, X. He, J. Hwang, and L. Li (2025) Attention consistency for llms explanation. arXiv preprint arXiv:2509.17178. Cited by: §1. [39] K. Lang (1995) Newsweeder: learning to filter netnews. In Machine learning proceedings 1995, p. 331â339. Cited by: §4.1. [40] N. F. Liu, M. Gardner, Y. Belinkov, M. Peters, and N. A. Smith (2019) Linguistic knowledge and transferability of contextual representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), p. 1073â1094. Cited by: §5.2, §5.2, §5.2. [41] S. Liu, F. Le, S. Chakraborty, and T. Abdelzaher (2021) On exploring attention-based explanation for transformer models in text classification. In 2021 IEEE International Conference on Big Data (Big Data), p. 1193â1203. Cited by: §1. [42] Z. Liu (2024) Cunliang kong, ying liu, and maosong sun. 2024. fantastic semantics and where to find them: investigating which layers of generative llms reflect lexical semantics. Findings of the Association for Computational Linguistics: ACL, p. 14551â14558. Cited by: §5.2. [43] S. Lundberg (2017) A unified approach to interpreting model predictions. arXiv preprint arXiv:1705.07874. Cited by: §2.1, §2.2. [44] A. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts (2011) Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies, p. 142â150. Cited by: §4.1. [45] M. A. Mersha, G. Y. Bade, J. Kalita, O. Kolesnikova, A. Gelbukh, et al. (2024) Ethio-fake: cutting-edge approaches to combat fake news in under-resourced languages using explainable ai. Procedia Computer Science 244, p. 133â142. Cited by: §2.1. [46] M. A. Mersha, J. Kalita, et al. (2024) Semantic-driven topic modeling using transformer-based embeddings and clustering algorithms. Procedia Computer Science 244, p. 121â132. Cited by: §1. [47] M. A. Mersha and J. Kalita (2025) Semantic-driven topic modeling for analyzing creativity in virtual brainstorming. arXiv preprint arXiv:2509.16835. Cited by: §1. [48] M. A. Mersha, M. G. Yigezu, and J. Kalita (2025) Evaluating the effectiveness of xai techniques for encoder-based language models. Knowledge-Based Systems 310, p. 113042. Cited by: §4.4. [49] M. A. Mersha, M. G. Yigezu, H. Shakil, A. K. AlShami, S. Byun, and J. Kalita (2025) A unified framework with novel metrics for evaluating the effectiveness of xai techniques in llms. arXiv preprint arXiv:2503.05050. Cited by: §2.1. [50] M. A. Mersha, M. G. Yigezu, A. L. Tonja, H. Shakil, S. Iskandar, O. Kolesnikova, and J. Kalita (2025) Explainable ai: xai-guided context-aware data augmentation. Expert Systems with Applications, p. 128364. Cited by: §2.1, §2.1. [51] M. Mersha, M. Bitewa, T. Abay, and J. Kalita (2024) Explainability in neural networks for natural language processing tasks. arXiv preprint arXiv:2412.18036. Cited by: §2.1. [52] M. Mersha, K. Lam, J. Wood, A. AlShami, and J. Kalita (2024) Explainable artificial intelligence: a survey of needs, techniques, applications, and future direction. Neurocomputing, p. 128111. Cited by: §2.1. [53] M. Nauta, J. Trienes, S. Pathak, E. Nguyen, M. Peters, Y. Schmitt, J. Schlötterer, M. van Keulen, and C. Seifert (2023) From anecdotal evidence to quantitative evaluation methods: a systematic review on evaluating explainable ai. ACM Computing Surveys 55 (13s), p. 1â42. Cited by: §5.2. [54] M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer (2018) Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p. 2227â2237. Cited by: §5.2. [55] Y. Qiang, D. Pan, C. Li, X. Li, R. Jang, and D. Zhu (2022) Attcat: explaining transformers via attentive class activation tokens. Advances in neural information processing systems 35, p. 5052â5064. Cited by: §2.1. [56] A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al. (2018) Improving language understanding by generative pre-training. Cited by: §1. [57] C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu (2020) Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21 (140), p. 1â67. Cited by: §1. [58] M. T. Ribeiro, S. Singh, and C. Guestrin (2016) â Why should i trust you?â explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, p. 1135â1144. Cited by: §2.1, §2.2. [59] A. Rogers, O. Kovaleva, and A. Rumshisky (2020) A primer in BERTology: what we know about how BERT works. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 1â17. Cited by: §5.2. [60] A. Rogers, O. Kovaleva, and A. Rumshisky (2020) A primer in bertology: what we know about how bert works. Transactions of the association for computational linguistics 8, p. 842â866. Cited by: §1. [61] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra (2017) Grad-cam: visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, p. 618â626. Cited by: Figure 17, Figure 9, Table 1. [62] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra (2020) Grad-cam: visual explanations from deep networks via gradient-based localization. International journal of computer vision 128, p. 336â359. Cited by: §2.1. [63] S. Serrano and N. A. Smith (2019) Is attention interpretable?. arXiv preprint arXiv:1906.03731. Cited by: §1, §2.2. [64] D. Shi, R. Jin, T. Shen, W. Dong, X. Wu, and D. Xiong (2024) Ircan: mitigating knowledge conflicts in llm generation via identifying and reweighting context-aware neurons. Advances in Neural Information Processing Systems 37, p. 4997â5024. Cited by: §2.1. [65] A. Shrikumar, P. Greenside, and A. Kundaje (2017) Learning important features through propagating activation differences. In International conference on machine learning, p. 3145â3153. Cited by: §2.1, §2.1, §4.3, Figure 11. [66] D. Smilkov et al. (2017) SmoothGrad: removing noise by adding noise. arXiv preprint arXiv:1706.03825. Cited by: §2.2. [67] E. Sood, S. Tannert, D. Frassinelli, A. Bulling, and N. T. Vu (2020) Interpreting attention models with human visual attention in machine reading comprehension. arXiv preprint arXiv:2010.06396. Cited by: §4.3. [68] S. Srinivas and F. Fleuret (2019) Full-gradient representation for neural network visualization. Advances in neural information processing systems 32. Cited by: §2.1. [69] C. Sun, X. Qiu, Y. Xu, and X. Huang (2019) Fine-tune BERT for extractive summarization. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), p. 3289â3299. Cited by: §5.2, §5.2. [70] M. Sundararajan, A. Taly, and Q. Yan (2017) Axiomatic attribution for deep networks. In ICML, Cited by: Table 2, Figure 17, §1, §2.1, §2.2, §3.1, §3.1, §3.5, §4.3, Figure 11, Figure 9, Table 1. [71] I. Tenney, D. Das, and E. Pavlick (2019) BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, p. 4593â4601. Cited by: §5.2, §5.2. [72] A. L. Tonja, M. Mersha, A. Kalita, O. Kolesnikova, and J. Kalita (2023) First attempt at building parallel corpora for machine translation of northeast indiaâs very low-resource languages. In Proceedings of the 20th International Conference on Natural Language Processing (ICON), p. 534â539. Cited by: §1. [73] A. Vaswani (2017) Attention is all you need. Advances in Neural Information Processing Systems. Cited by: §1, §3.2. [74] J. Vig (2019) A multiscale visualization of attention in the transformer model. arXiv preprint arXiv:1906.05714. Cited by: §2.1. [75] S. Wiegreffe and Y. Pinter (2019) Attention is not not explanation. arXiv preprint arXiv:1908.04626. Cited by: §2.2. [76] C. Yeh, Y. Chen, A. Wu, C. Chen, F. ViĂ©gas, and M. Wattenberg (2023) Attentionviz: a global view of transformer attention. IEEE Transactions on Visualization and Computer Graphics. Cited by: §1. [77] T. Yuan, X. Li, H. Xiong, H. Cao, and D. Dou (2021) Explaining information flow inside vision transformers using markov chain. In eXplainable AI approaches for debugging and diagnosis., Cited by: §2.1. [78] O. Zaidan, J. Eisner, and C. Piatko (2007) Using âannotator rationalesâ to improve machine learning for text categorization. In Human language technologies 2007: The conference of the North American chapter of the association for computational linguistics; proceedings of the main conference, p. 260â267. Cited by: §5.1.2. [79] M. Zeiler (2014) Visualizing and understanding convolutional networks. In European conference on computer vision/arXiv, Vol. 1311. Cited by: §2.1. [80] B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba (2016) Learning deep features for discriminative localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 2921â2929. Cited by: §2.1, §2.2. [81] H. Zhu, F. Wei, B. Qin, and T. Liu (2018) Hierarchical attention flow for multiple-choice reading comprehension. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32. Cited by: §2.1.