Paper deep dive
XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models
Yupei Li, Qiyang Sun, Xiaoliang Wu, Chenxi Wang, Berrak Sisman, Björn W. Schuller
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 6/20/2026, 7:52:12 AM
Summary
The paper proposes XGEG (XAI-Grounded Explanation Generation), a training-free framework designed to improve the reliability and specificity of explanations in Speech Deepfake Detection (SDD). Traditional XAI methods (like IG, LIME, and Saliency) provide low-level signals, while LLMs often produce ungrounded descriptions. XGEG integrates XAI evidence (spectrogram-based and SHAP-derived acoustic features) with multimodal LLMs (Qwen2.5-VL and Qwen3-Omni-30B) to generate grounded, natural language explanations. The authors also introduce a large-scale explainable SDD dataset based on PartialSpoof, containing approximately 65,000 instances. Experimental results show that XAI-guided methods significantly improve explanation accuracy, specificity, and faithfulness compared to audio-only LLM baselines.
Entities (11)
Relation Signals (5)
XGEG → integrates → XAI_Method
confidence 100% · our approach introduces cross-model XAI aggregation and fidelity-driven validation
PartialSpoof → isusedtoconstruct → XGEG_Dataset
confidence 100% · Using the PartialSpoof dataset, we construct a grounded explanation dataset
XGEG → uses → Qwen3-Omni-30B
confidence 100% · These time–frequency summaries are provided as input to the multimodal LLM Qwen3-Omni-30B
Qwen2.5-VL-7B → processes → spectrogram_evidence
confidence 90% · these heatmaps are then provided as input to the vision-LLM Qwen2.5-VL-7B
Integrated Gradients → providesevidencefor → XGEG
confidence 90% · To provide traditional XAI (IG, LIME, and Saliency, selected as representatives) evidence
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making. Existing explanation ways mainly fall into two categories. Traditional explainable AI (XAI), such as gradient-based attribution, produces low-level attribution signals tightly coupled with model decisions, and harder to be understood by human than natural language explanations. Meanwhile, large language model (LLM)-based explanation generation often produces generic and ungrounded descriptions due to the lack of heuristic evidence and task-specific supervision, stemming from limited grounded explanation datasets for SDD. We therefore propose a training-free explanation framework that integrates XAI evidence with multimodal LLMs to generate grounded and specific explanations. Using the PartialSpoof dataset, we construct a grounded explanation dataset and show that methods with XAI increase inside accuracy by over 45\%, verified through human evaluation and faithfulness checks.
Tags
Links
- Source: https://arxiv.org/abs/2606.16137v1
- Canonical: https://arxiv.org/abs/2606.16137v1
Trouble viewing inline? Open PDF directly →
Full Text
30,987 characters extracted from source content.
Expand or collapse full text
XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models Yupei Li 1,2 , Qiyang Sun 1 , Xiaoliang Wu 3 , Chenxi Wang 4 , Berrak Sisman 5 , Bj ̈ orn W. Schuller 1,2 1 Imperial College London, United Kingdom 2 Technical University of Munich, Germany 3 University of Southampton, United Kingdom 4 MBZUAI, United Arab Emirates 5 Johns Hopkins University, United States of America yl7622@ic.ac.uk, q.sun23@imperial.ac.uk, xiaoliang.wu@soton.ac.uk, Chenxi.Wang@mbzuai.ac.ae, sisman@jhu.edu, schuller@tum.de Abstract Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making. Existing explanation ways mainly fall into two categories. Traditional explainable AI (XAI), such as gradient-based attribution, produces low-level at- tribution signals tightly coupled with model decisions, and harder to be understood by human than natural language explanations. Meanwhile, large language model (LLM)-based explanation generation often produces generic and ungrounded descriptions due to the lack of heuristic evidence and task-specific supervi- sion, stemming from limited grounded explanation datasets for SDD. We therefore propose a training-free explanation frame- work that integrates XAI evidence with multimodal LLMs to generate grounded and specific explanations. Using the Partial- Spoof dataset, we construct a grounded explanation dataset and show that methods with XAI increase inside accuracy by over 45%, verified through human evaluation and faithfulness checks. Index Terms: Grounded explanation, Explainable artificial in- telligence, Multimodal large language models, Speech deepfake detection, Dataset creation 1. Introduction Speech deepfake detection (SDD) has been extensively investi- gated using both traditional deep learning models [1] and large language models (LLMs) [2,3]. Although often formulated as a binary classification task, SDD remains inherently challenging. This is because it does not merely require determining whether a speech sample is bona fide or spoofed; rather, it necessitates a principled justification of why a given sample is regarded as fake, so as to avoid arbitrary or ungrounded decisions. This explana- tory requirement may constitute a fundamental factor underlying the limited generalisability of current SDD systems. Many exist- ing models tend to rely on superficial statistical artefacts specific to particular synthesis techniques such as watermarking, rather than capturing the intrinsic generative mechanisms of spoofed speech [4]. Therefore, the development of trustworthy and in- terpretable explanations is not merely desirable but essential for achieving reliable, and responsible detection outcomes [5]. Explanations in current approaches are primarily derived from traditional explainable artificial intelligence (XAI) tech- niques and, more recently, from LLMs that generate natural language rationales. Traditional XAI methods take advantage of mathematically grounded attribution mechanisms, which pro- vide theoretical support for the validity and faithfulness of their outputs. For instance, Integrated Gradients (IG) [6] quantify feature importance by attributing prediction changes to input perturbations along a continuous path from a baseline; saliency maps [7] highlight the most influential input dimensions based on gradient sensitivity; Local Interpretable Model-agnostic Ex- planations (LIME) [8] approximate the model locally with an interpretable surrogate to estimate feature contributions; and SHapley Additive exPlanations (SHAP) estimate feature contri- butions based on Shapley values from cooperative game theory [9]. However, these methods predominantly rely on locality or linearity assumptions. Moreover, they require access to the original decision-making model, regardless of whether the XAI technique is model-specific or model-agnostic [10, 11]. On the other hand, LLM-based explanations can be gen- erated independently of the original detection model, without requiring gradient access or architectural transparency [12]. By articulating rationales in natural language rather than low-level spectral visualisations, they are generally more interpretable to human users. However, producing reliable explanations often demands either strong reasoning capabilities, which remain un- derdeveloped in current audio LLMs [13], or carefully curated training data to guide the model’s attention towards relevant acoustic cues. Without such constraints, LLM-generated expla- nations are prone to hallucination and may remain superficial, offering only high-level descriptions rather than faithful accounts of the underlying decision process [14]. Specifically within the SDD domain, approaches explicitly designed for explainable detection remain limited. Existing stud- ies predominantly rely on traditional XAI techniques [15,16,17], which have shown effectiveness in providing feature-level attri- butions. In contrast, LLM-based approaches typically follow a post hoc explanation pipeline similar to that discussed above [18], while novelly incorporating multiple LLMs for iterative self-consistency checking [19]. However, it still lacks princi- pled heuristic mechanisms for assessing explanation quality, and remains an absence of dedicated datasets tailored to SDD explanation. Also, validating the quality of LLM-generated ex- planations presents an additional challenge [20], as it is hard to objectively evaluate the flexible output with existing metrics to capture faithfulness and factual correctness. Therefore, to address the limitations of traditional XAI meth- ods, including their dependency on access to the underlying de- cision model and limited flexibility, as well as the unverifiability and lack of structural support in purely LLM-generated explana- tions, and to bridge the absence of dedicated explainable SDD datasets, we make the following contribution. First, Unlike prior post-hoc LLM explanation pipelines that rely solely on textual arXiv:2606.16137v1 [cs.CL] 15 Jun 2026 prompts or single-model attribution, our approach introduces cross-model XAI aggregation and fidelity-driven validation to produce semantically rich and specific explanations with less hallucination, while establishing a principled framework for ex- planation generation and verification. Second, we construct and publicly release a large-scale explainable SDD dataset based on the PartialSpoof dataset [21], comprising approximately 65,000 explanation instances, to support future research in explainable SDD. Codes and data are in Github Link 1 . 2. Methodology: XAI-Grounded Explanation Generation via LLMs (XGEG) As aforementioned, we aim to leverage traditional XAI methods to guide and provide supportive evidence for LLM-based gener- ation, thereby obtaining more trustworthy and specific explana- tions with less hallucinations. The overall pipeline is illustrated in Figure 1. Given our objective of proposing a generalizable pipeline, we primarily rely on a training-free framework to en- hance scalability across diverse dataset generation settings. To provide traditional XAI (IG, LIME, and Saliency, se- lected as representatives) evidence as guidance for LLM-based explanation generation, we employ three pre-trained SDD mod- els based on wav2vec 2.0, HuBERT [22], and WavLM [23] (checkpoint provided) as foundation models. These models are used directly without additional fine-tuning on the target dataset, ensuring an efficient and scalable pipeline without the need to retrain a dedicated detector for each new explanation task. In addition to spectrogram-based evidence, we provide SHAP-derived importance scores for acoustic features extracted using the openSMILE eGeMAPSv02 feature set. Specifically, we train a lightweight four-layer multilayer perceptron (MLP) [24] on the training split to estimate feature contributions, delib- erately keeping the classifier simple. As our goal is to generate high-quality explanations rather than optimise detection perfor- mance, potential data leakage in classification accuracy does not materially affect the validity of the explanatory analysis. Furthermore, we adopt the widely used PartialSpoof dataset [21] in the work, as it provides precise temporal annotated labels indicating which segments are spoofed. The dataset comprises approximately 25k, 25k, and 71k samples for training, develop- ment, and testing, respectively, which are sufficient to support our explanation generation pipeline. As a sanity check, we evaluate the pre-trained detection mod- els on the selected dataset, reporting Accuracy, F1 and Equal Error Rate (EER) in Table 1. The results show consistently strong performance across models, suggesting that their pre- dictions are sufficiently reliable to serve as a stable foundation for subsequent XAI analysis. Moreover, when constructing the explainable dataset, we retain only those samples that are cor- rectly classified by all four models. This reduces the risk of erroneous decisions propagating misleading attribution signals and degrading the quality of LLM-generated explanations. We focus on spoofed samples, as explanations for bona fide speech are generally uninformative and typically rely on the absence of anomalies (e.g., “no acoustic inconsistency detected”). After this filtering process, the dataset comprises around 15k, 15k, and 35k samples for training, development, and testing, respectively. After obtaining spectrogram-based explanations from IG, LIME, and Saliency for the three pre-trained models, these heatmaps are then provided as input to the vision-LLM Qwen2.5- 1 https://github.com/glam-imperial/ xai-grounded-speech-deepfake Table 1: Performance of Pre-trained Models and MLP Classifier on the PartialSpoof Dataset ModelSplitAccuracy↑F1-score↑EER↓ HuBERT Train.712.414.122 Validation.697.403.119 Test.715.419.108 Wav2Vec 2.0 Train.709.411.119 Validation.694.401.109 Test.717.421.097 WavLM Train.703.406.078 Validation.690.398.067 Test.700.408.074 MLP Train.981.912.024 Valid.961.813.055 Test.954.795.063 VL-7B [25], which has shown strong image understanding capa- bilities. Using a carefully designed prompt (full version provided in the released code), we instruct the model to summarise abnor- mal regions in terms of their temporal and frequency ranges, the two core dimensions of time–frequency speech representation. Together with the top three acoustic features ranked by SHAP importance scores, these time–frequency summaries are provided as input to the multimodal LLM Qwen3-Omni-30B [26]. This design follows recent research exploring LLM-based explanation generation [18]. Qwen3-Omni-30B has shown strong performance in audio captioning tasks with comparatively reduced hallucination, making it suitable for synthesising multi- modal evidence into coherent textual explanations. Importantly, our objective is not to merely restate the XAI outputs. Instead, we explicitly instruct the model to prioritise the acoustic content of the input and treat the XAI-derived evidence as supporting sig- nals rather than definitive conclusions. Furthermore, the model is prompted to critically analyse the XAI evidence instead of passively reproducing it. To further mitigate hallucination and encourage specificity, we constrain the output to a predefined structured format as follows, ensuring that the generated expla- nations remain focused, detailed, and evidence-grounded. The full version is provided in the released code. I. AUDIOABNORMALITY TIME RANGE: FREQRANGE: I. EXPLANATION (Free-text) I. XAI AGGREGATION (Indicate XAI contribution and whether it reflects cross- model agreement or single-model evidence.) 3. Results and Discussion 3.1. Generated samples We adopt the official Hugging Face implementations of all above models and verify their outputs remain consistent across multi- ple runs. To test whether the XAI signals provide meaningful guidance, rather than being merely paraphrased by the LLM, we design a series of comparative experiments: • Audio-only input, where the integrated LLM receives only the raw audio information without any XAI evidence; •Single-XAI input, where only one attribution method (IG, LIME, or Saliency) derived from a single model (wav2vec2) is provided, resulting in three separate experimental groups; •Full-XAI from one model, where all four XAI signals (IG, LIME, Saliency, and SHAP) from wav2vec 2.0 are supplied; Raw Waveform OpenSMILE Multimodal LLM Qwen2.5-VL-7B SHAP Analysis Comprehensive Human- Understandable Text Explanation The clip features a female voice with a neutral, synthetic tone and steady delivery. The audio is clean, but the lack of natural inflection suggests a text-to-speech or deepfake origin. A key anomaly appears at 0.5–1.0 seconds in the word “insisted,” where an abrupt pitch drop and a 5000– 6000 Hz frequency shift indicate possible voice synthesis artifacts...... Qualitative & Quantitative Analysis Foundation Model Integrated Gradient Saliency LIME XAI Attribution Maps Integrated Gradient Text Explanations Example:TIME_REGION: [0.75-1.25] s FREQUENCY_REGION: [3000-6000] Hz Saliency Text Explanations Example: TIME_REGION: [0.50-0.75]s FREQUENCY_REGION: [4000-6000]Hz LIME Text Explanations Example:TIME_REGION: [0.50-0.75]s Individual Text Explanations Integration LLM Qwen3-Omni-30B SHAP Text Explanations slopeUV0-500_sma3nz_amean: 6.97 equivalentSoundLevel_dBp: 2.39 StddevUnvoicedSegmentLength: 1.25 Figure 1: XGEG pipeline. Raw audio is processed by multiple deepfake detection models to generate attribution maps, which are interpreted by a multimodal LLM into notion-level explanations. Meanwhile, acoustic features are extracted with openSMILE, classified by an MLP, and analyzed using SHAP for feature-level attributions. These explanations and attributions are integrated by an LLM to produce the final explanation dataset for evaluation. • Cross-model XAI aggregation, where XAI evidence from all three pre-trained models is integrated. We do not conduct a SHAP-only experiment, as SHAP pro- vides feature-level importance without explicit time–frequency localisation, which offers limited guidance to the multimodal LLM in identifying specific abnormal temporal segments. The generated texts as published datasets are in Github Link. 3.2. Qualitative analysis Qualitative case studies show consistent trends in our model’s effectiveness. Raw audio alone often yields superficial or hallu- cinated explanations, while lightweight XAI guidance improves precision by identifying temporal and spectral anomalies. The LLM integrates acoustic evidence rather than merely restating XAI outputs, as supported by accurate ASR, precise localisation, and attention visualisations (Figures 2, 3), where audio tokens prepended to the text input receive attention weights and are effectively attended to during decoding. However, reasoning over aggregated XAI signals remains challenging due to LLM inherent reasoning bottlenecks [27], which we leave for future work. The model also shows emergent capability [28] in iden- tifying potential deepfake sources such as text-to-speech (TTS) origin. Although systematic evaluation is difficult due to the lack of structured outputs for such abilities, qualitative inspection suggests frequent correctness across many cases. 3.3. Human evaluation To assess the quality of the generated explanations, we conduct a human evaluation study. A total of 600 explanations corre- sponding to spoofed audio instances were randomly sampled for human evaluation. Twenty annotators were recruited to assess the explanations. The annotators consisted of 10 males and 10 females, aged between 20 and 30 years. All participants were fluent English users and held at least a bachelor’s degree, with five holding graduate degrees. In addition, five annotators had prior academic or practical experience in audio-related fields. Each annotator is assigned a balanced subset of samples cov- 050100150200250 Key 0 50 100 150 200 250 Query 0.0 0.2 0.4 0.6 0.8 1.0 Normalized Attention Figure 2: First generation step of the attention map for the first layer on average multiple heads 0100200300400500 Key 0.4 0.2 0.0 0.2 0.4 Query 0.2 0.4 0.6 0.8 1.0 Normalized Attention Figure 3: Last generation step of the attention map for the first layer on average multiple heads ering all six experimental conditions. In total, each annotator evaluates 30 samples, derived from five distinct original audio clips, with six different explanation formats generated from each clip using six distinct prompts. Participants are informed that all audio samples are spoofed and are instructed to listen carefully to each clip while reading the corresponding explanation. They then rate each explanation on a 5-point Likert scale (1 = lowest, 5 = highest), based solely on its alignment with the audio. As our goal is to produce explanations understandable to the public, no professional expertise was required. Annotators were recruited without restrictions on academic background, age, or nationality, and were instructed to focus on overall meaning after a brief introduction to the task. The assessment is conducted according to five criteria: correctness (C), measuring whether the explanation accurately identifies and justifies the abnormal acoustic region; evidence support (E), evaluating whether the explanation is grounded in observable audio cues rather than hallucinated content; specificity (S), reflecting the level of de- tail and precision in the description; missing explanation (M), assessing whether obvious abnormal regions are omitted; and overall preference (O), capturing the annotator’s comparative judgement of explanation quality (only selecting 1 sample out of 6). The scores are shown in Table 2. Table 2: Human Evaluation Results Under Different Settings SettingsC↑E↑S↑M↓O (average number)↑ Pure Audio (Baseline)3.151.751.902.750.35 IG3.603.503.452.200.70 Saliency3.753.653.302.150.75 LIME3.353.403.502.300.65 All XAI (Single Model)3.903.753.552.451.05 All XAI (Three Model)3.853.604.302.301.50 The results reveal a clear trend indicating that the XAI- guided versions achieve better alignment with human subjective assessments. Notably, the improvements in specificity and ev- idence support are particularly convincing, suggesting that the XAI-guided models produce fewer hallucinations and provide more detailed justifications. Furthermore, the findings demon- strate that incorporating a greater degree of XAI guidance leads to higher overall recognition. However, to ensure technical accu- racy, particularly in aspects such as whether the time period is correctly identified, rigorous quantitative analysis is required. 3.4. Quantitative analysis We therefore conducted a correctness evaluation using a subset of the generated explanations, based on the original development subset of the Partialspoof dataset. 3.4.1. Intersection over Union (IoU) and Inside Accuracy (IA) To evaluate time abnormal period detection, we adopt IoU and IA. IoU is defined as:IoU = |T p red∩ T g t|/|T p red∪ T g t|, whereT p redandT g tdenote the predicted and ground-truth abnormal time intervals, which localisation works often use as a metric [29,30]. IA is defined as the proportion of samples where the predicted time interval is fully contained within the ground-truth interval, which punishes hallucinated guess of the time abnormal period. The results are shown in Table 3. Pure Audio achieves the highest IoU due to over-expanded abnormal intervals, but its very low IA indicates poor localization precision. Single XAI methods, such as LIME, often produce unstable and overly narrow regions. Combining multiple XAI sources within one model provides a better balance between IoU and IA, whereas aggregating XAI from multiple models increases reasoning complexity and may reduce stability. 3.4.2. Area-Normalised Local Logit Sensitivity Additionally, we design a set of quantitative attribution exper- iments as one form of fidelity test [31]. Traditional XAI eval- uation methods are unsuitable for high-frequency continuous audio tasks. Such methods include direct masking, silencing, or frequency-domain cropping [32,33]. Aggressive time-frequency masking disrupts phase coherence. This artificially introduces strong digital artefacts. Consequently, this shifts the input out- of-distribution. It causes abnormal fluctuations in the model’s confidence for the ‘Fake’ class. It fails to reflect the true causal importance of the extracted features. To overcome this flaw, we propose the Area-Normalised Local Logit Sensitivity metric. We avoid destructive masking on salient regions. Instead, we apply a minimal multiplicative amplitude perturbation to the target time-frequency area. We set this perturbation toε = +1.0%in our experiments. This approach maintains the structural integrity and phase continuity of the audio. We calculate the absolute change in the model’s output logit. We then divide this change by the area of the Table 3: Time period detection performance under different explanation settings. SettingsIoUIA Pure Audio (Baseline)0.2690.049 IG0.2240.482 Saliency0.2390.489 LIME0.1570.811 All XAI (Single Model)0.2420.492 All XAI (Three Model)0.1340.295 time-frequency region (∆t× ∆f). A higher sensitivity density indicates greater model sensitivity to that specific information. We compare our multimodal fusion methods with a Pure Audio LLM baseline using localisation regions generated by IG, Saliency, LIME, and 4XAI. As shown in Table 4, multi- modal methods consistently identify more informative and influ- ential regions than the baseline, particularly with LIME guidance. This may be because LIME relies on a perturbation-based strat- egy to approximate the model’s local decision boundary, which aligns closely with our analytical framework. The 4XAI variant achieves the best performance, indicating that integrating multi- ple XAI signals improves localisation of key deepfake speech segments. Overall, the XAI (single-model) version achieves the highest quality as the evaluation metrics reported above; therefore, the released dataset follows this configuration. Table 4: Quantitative Evaluation of Explainability using Area- Normalised Logit Sensitivity (ε = +1.0%) MethodMean Sensitivity Density (×10 −6 )Ratio vs. Baseline Pure Audio (Baseline)0.14241.00× IG1.33829.40× Saliency2.925020.54× All XAI (Three Model)3.118621.90× All XAI (Single Model)3.146622.10× LIME46.5857327.20× 3.5. Bona fide sample analysis We mainly focus on fake audio explanation generation. We also tested some bona fide samples. However, explaining genuine audio is more challenging, as demonstrating the absence of abnormalities is inherently harder than identifying their presence, similar to the difficulty of proving innocence rather than guilt [34]. For bona fide cases, we use a similar prompt, except that the model is not asked to output abnormal time periods, but instead explain why no suspicious evidence is found. Part of one example is: “The audio sample is a high-fidelity recording of a female speaker saying... The voice is clear, natural, and exhibits subtle, human vocal characteristics such as breathiness...”. 4. Conclusion In this work, we proposed an XAI aggregation framework that guides training-free LLMs with signals from conventional XAI methods. The framework generates more specific and tempo- rally grounded SDD explanations with reduced hallucination. Experiments, with fidelity analysis and human evaluation, show improved localisation and semantic grounding. By providing interpretable evidence beyond binary decisions, our approach en- hances transparency and trust in SDD. Future work will focus on stronger XAI integration and explanation-aware LLM training. 5. Generative AI Use Disclosure We only used Generative AI for grammar check and proofreading of the manuscript. 6. References [1]M. S. Rana, M. N. Nobi, B. Murali, and A. H. Sung, “Deepfake detection: A systematic literature review,” IEEE access, vol. 10, p. 25 494–25 513, 2022. [2]Y. Li, L. Wang, Y. Wang, L. Wang, R. Cai, J. Shi, B. W. Schuller, and Z. Wu, “Dfallm: Achieving generalizable multitask deepfake detection by optimizing audio llm components,” arXiv preprint arXiv:2512.08403, 2025. [3]H. Gu, J. Yi, C. Wang, J. Tao, Z. Lian, J. He, Y. Ren, Y. Chen, and Z. Wen, “Allm4add: Unlocking the capabilities of audio large language models for audio deepfake detection,” in Proceedings of the 33rd ACM International Conference on Multimedia, 2025, p. 11 736–11 745. [4] W. Zong, Y.-W. Chow, W. Susilo, J. Baek, and S. Camtepe, “AudioMarkNet: Audio watermarking for deepfake speech de- tection,” in 34th USENIX Security Symposium (USENIX Security 25), 2025, p. 4663–4682. [5]L. Cirillo, A. Gervasio, and I. Amerini, “Explainability-driven ad- versarial robustness assessment for generalized deepfake detectors,” EURASIP Journal on Information Security, vol. 2025, no. 1, p. 23, 2025. [6]M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in International conference on machine learning. PMLR, 2017, p. 3319–3328. [7]K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep inside con- volutional networks: Visualising image classification models and saliency maps,” arXiv preprint arXiv:1312.6034, 2013. [8]M. T. Ribeiro, S. Singh, and C. Guestrin, “” why should i trust you?” explaining the predictions of any classifier,” in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, 2016, p. 1135–1144. [9]S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” Advances in neural information processing systems, vol. 30, 2017. [10] Y. Li, Q. Sun, A. Akman, and B. W. Schuller, “Explainable ai for healthcare,” Studies in health technology and informatics, vol. 330, p. 632–652, 2025. [11]Q. Sun, A. Akman, and B. W. Schuller, “Explainable artificial in- telligence for medical applications: A review,” ACM Transactions on Computing for Healthcare, vol. 6, no. 2, p. 1–31, 2025. [12] A. Pang, H. Jang, and S. Fang, “Generating descriptive expla- nations of machine learning models using llm,” in 2024 IEEE International Conference on Big Data (BigData).IEEE, 2024, p. 5369–5374. [13] J. Peng, Y. Wang, B. Li, Y. Guo, H. Wang, Y. Fang, Y. Xi, H. Li, X. Li, K. Zhang et al., “A survey on speech large language models for understanding,” IEEE Journal of Selected Topics in Signal Processing, 2025. [14]P. Sahoo, P. Meharia, A. Ghosh, S. Saha, V. Jain, and A. Chadha, “A comprehensive survey of hallucination in large language, image, video and audio foundation models,” Findings of the Association for Computational Linguistics: EMNLP 2024, p. 11 709–11 724, 2024. [15]A. Govindu, P. Kale, A. Hullur, A. Gurav, and P. Godse, “Deep- fake audio detection and justification with explainable artificial intelligence (xai),” 2023. [16]G. Channing, J. Sock, R. Clark, P. Torr, and C. S. de Witt, “Toward robust real-world audio deepfake detection: Closing the explain- ability gap,” arXiv preprint arXiv:2410.07436, 2024. [17]A. Akman, Q. Sun, and B. W. Schuller, “Audio explanation syn- thesis with generative foundation models,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, p. 1–5. [18]Y. Xie, X. Guo, J. Zhou, T. Wang, J. Liu, R. Fu, X. Wang, H. Cheng, and L. Ye, “Interpretable all-type audio deepfake detection with au- dio llms via frequency-time reinforcement learning,” arXiv preprint arXiv:2601.02983, 2026. [19]X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” arXiv preprint arXiv:2203.11171, 2022. [20]Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang et al., “A survey on evaluation of large language models,” ACM transactions on intelligent systems and technology, vol. 15, no. 3, p. 1–45, 2024. [21] L. Zhang, X. Wang, E. Cooper, N. Evans, and J. Yamagishi, “The partialspoof database and countermeasures for the detection of short fake speech segments embedded in an utterance,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, p. 813–825, 2022. [22]W. Ge, X. Wang, X. Liu, and J. Yamagishi, “Post-training for deep- fake speech detection,” in 2025 IEEE Automatic Speech Recogni- tion and Understanding Workshop (ASRU), 2025. [23]D. Combei, A. Stan, D. Oneata, and H. Cucu, “Wavlm model ensemble for audio deepfake detection,” in Proceedings of the Au- tomatic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, p. –. [24]D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature, vol. 323, no. 6088, p. 533–536, 1986. [25]Q. Team, “Qwen2.5-vl,” January 2025. [Online]. Available: https://qwenlm.github.io/blog/qwen2.5-vl/ [26]J. Xu, Z. Guo, H. Hu, Y. Chu, X. Wang, J. He, Y. Wang, X. Shi, T. He, X. Zhu, Y. Lv, Y. Wang, D. Guo, H. Wang, L. Ma, P. Zhang, X. Zhang, H. Hao, Z. Guo, B. Yang, B. Zhang, Z. Ma, X. Wei, S. Bai, K. Chen, X. Liu, P. Wang, M. Yang, D. Liu, X. Ren, B. Zheng, R. Men, F. Zhou, B. Yu, J. Yang, L. Yu, J. Zhou, and J. Lin, “Qwen3-omni technical report,” arXiv preprint arXiv:2509.17765, 2025. [27]Z. Li, Y. Cao, X. Xu, J. Jiang, X. Liu, Y. S. Teo, S.-W. Lin, and Y. Liu, “Llms for relational reasoning: How far are we?” in Pro- ceedings of the 1st international workshop on large language models for code, 2024, p. 119–126. [28]J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou, “Emergent abilities of large language models,” Transactions on Machine Learning Research, 2022. [29]Y. Li, L. Wang, Y. Wang, L. Wang, R. Cai, J. Shi, B. W. Schuller, and Z. Wu, “Dfallm: Achieving generalizable multitask deepfake detection by optimizing audio llm components,” 2025. [Online]. Available: https://arxiv.org/abs/2512.08403 [30] B. Zhang, Q. Yin, W. Lu, and X. Luo, “Deepfake detection and localization using multi-view inconsistency measurement,” IEEE Transactions on Dependable and Secure Computing, vol. 22, no. 2, p. 1796–1809, 2024. [31]M. Mir ́ o-Nicolau, A. Jaume-i Cap ́ o, and G. Moy ` a-Alcover, “A comprehensive study on fidelity metrics for xai,” Information Pro- cessing & Management, vol. 62, no. 1, p. 103900, 2025. [32]A. Akman, Q. Sun, and B. W. Schuller, “Improving audio expla- nations using audio language models,” IEEE Signal Processing Letters, vol. 32, p. 741–745, 2025. [33]Y. Li, Q. Sun, H. Li, L. Specia, and B. W. Schuller, “Detecting machine-generated music with explainability–a challenge and early benchmarks,” arXiv preprint arXiv:2412.13421, 2024. [34] S. Mumford, “Negative truth,” in Absence and Nothing: The Phi- losophy of What There Is Not, S. Mumford, Ed. Oxford University Press, 2021, p. 147–167.