Paper deep dive
MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection
Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual manipulation, addresses speech detection in isolation, or conflates speech and non-speech audio as a single undifferentiated audio stream, overlooking the distinct forensic challenges posed by background audio. This conflation is consequential: the two acoustic components arise from fundamentally different generative mechanisms, exhibit distinct artifact profiles, and pose different challenges to detection systems. We introduce MADBench, the first benchmark that treats speech and environmental audio as distinct acoustic components, enabling component-aware evaluation of audio deepfake detection across independently manipulated forgery sources. We benchmark representative state-of-the-art detectors and multimodal large language models under a unified protocol. Our experiments reveal that environmental audio manipulation is more detectable than synthetic speech across general-purpose encoders, while existing pretrained detectors fail on both acoustic components, and manipulated environmental audio asymmetrically degrades speech deepfake detection, findings entirely invisible under the single-label paradigm of prior benchmarks. MADBench establishes a rigorous foundation for future research into robust, component-aware audio deepfake detection.
Tags
Links
- Source: https://arxiv.org/abs/2608.09593v1
- Canonical: https://arxiv.org/abs/2608.09593v1
Trouble viewing inline? Open PDF directly →
Full Text
62,129 characters extracted from source content.
Expand or collapse full text
MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection Yanqiu Li1, Yang Xiao1, Jisheng Bai2, Bin Chen1, Hong Jia3, Ting Dang1 Abstract Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual manipulation, addresses speech detection in isolation, or conflates speech and non-speech audio as a single undifferentiated audio stream, overlooking the distinct forensic challenges posed by background audio. This conflation is consequential: the two acoustic components arise from fundamentally different generative mechanisms, exhibit distinct artifact profiles, and pose different challenges to detection systems. We introduce MADBench, the first benchmark that treats speech and environmental audio as distinct acoustic components, enabling component-aware evaluation of audio deepfake detection across independently manipulated forgery sources. We benchmark representative state-of-the-art detectors and multimodal large language models under a unified protocol. Our experiments reveal that environmental audio manipulation is more detectable than synthetic speech across general-purpose encoders, while existing pretrained detectors fail on both acoustic components, and manipulated environmental audio asymmetrically degrades speech deepfake detection, findings entirely invisible under the single-label paradigm of prior benchmarks. MADBench establishes a rigorous foundation for future research into robust, component-aware audio deepfake detection. 1 Introduction The rapid advancement of artificial intelligence has altered the threat landscape of synthetic media. Deepfakes, AI-manipulated content designed to deceive, have evolved from face-swap techniques into sophisticated, multimodal fabrications capable of fooling both human perception and automated detection systems alike (Dolhansky et al. 2020; Li et al. 2020). While significant research effort has been devoted to detecting visual forgeries (Khalid et al. 2021), a quieter and arguably more insidious manipulation vector has received far less scrutiny: the replacement or synthesis of audio in otherwise authentic video. Consider a realistic and increasingly common attack scenario. An adversary obtains a genuine video recording, a politician’s speech, a corporate announcement, a personal conversation, and leaves the visual content entirely intact. Instead, they replace or synthesize the speech to alter what is being said (Chen et al. 2025; Zhou et al. 2026), and manipulate the background audio to make the fabrication acoustically coherent and convincing (Kreuk et al. 2022; Liu et al. 2024b). The resulting deepfake is visually authentic by construction, yet entirely deceptive in meaning. This setting is operationally attractive precisely because visual forgery detection has matured considerably (Cai et al. 2022, 2023), making audio-domain manipulation a lower-cost, lower-risk avenue for bad actors. This attack surface exposes a critical and underappreciated distinction that the research community has largely failed to formalize: speech and environmental audio are distinct acoustic components. Speech carries linguistic and paralinguistic content, what is said and how it is said, and is generated by fundamentally different synthesis pipelines than non-speech audio (Wang et al. 2020; Liu et al. 2023). Background soundscapes, ambient acoustics, environmental noise, and foley-style effects constitute a separate generative and perceptual domain, one that plays an equally important role in establishing the perceived authenticity of a video. A fabricated crowd reaction, a synthetically generated room ambiance, or an artificially composed acoustic scene can all serve to reinforce a deceptive narrative (Yin et al. 2025b, a, 2026; Zhang et al. 2026b), yet none of these are captured by speech-centric detection models. Despite the growing importance of audio-only deepfakes, no existing benchmark is designed for authentic-video scenarios with independently manipulated speech and non-speech audio. Existing datasets either focus on visual manipulations, treat speech as representative of the entire acoustic stream (Rossler et al. 2019; Jiang et al. 2020), or collapse all audio manipulations into a single label (Frank and Schönherr 2021). As a result, the research community currently lacks the tools to answer even basic questions: How detectable are synthesized speech tracks when the video is genuine? How do non-speech audio artifacts differ from speech artifacts under detection? Does the presence of manipulated background audio confound or assist speech deepfake detectors? Without benchmarks that explicitly disentangle these acoustic components, it is impossible to systematically evaluate detector robustness or understand the limitations of current approaches. To address this gap, we introduce the first benchmark designed specifically for audio-component deepfake detection in authentic-video settings, where speech and non-speech audio are manipulated independently while the visual content remains unchanged. Unlike existing datasets, our benchmark explicitly annotates these two acoustic sources as separate forgery components, enabling fine-grained evaluation of component-specific detection performance as well as their interactions. This formulation provides a more realistic and challenging testbed for developing next-generation audiovisual deepfake detectors. Extensive experiments reveal a striking finding: although audio-visual encoders consistently detect environmental audio manipulation more reliably than speech, existing pretrained detectors fail on both acoustic components, and manipulated environmental audio degrades speech deepfake detection without the converse holding, a cross-component effect that remains invisible when speech and non-speech audio are collapsed into a single label, leading to an incomplete assessment of detector robustness. Therefore, our contributions are summarized as follows: • A new benchmark. We introduce MADBench, the first benchmark for audio-component deepfake detection in authentic-video settings, establishing a realistic evaluation task where speech and non-speech audio are manipulated independently. • A modality-aware dataset. We present a dataset with independently synthesized and annotated speech and non-speech audio, enabling fine-grained evaluation of acoustic deepfakes beyond speech-centric settings. • A comprehensive testbed. We benchmark representative state-of-the-art deepfake detection models under a unified evaluation protocol, providing strong baselines for future research. 2 Related Work 2.1 Audio-Visual Deepfake Benchmarks Early benchmarks such as FaceForensics++ and DeeperForensics (Rossler et al. 2019; Jiang et al. 2020) established face-forgery detection as the dominant paradigm, but their evaluation targets remain visual or binary real/fake. Recent audio-visual benchmarks extend the task beyond visual-only forgery. FakeAVCeleb (Khalid et al. 2021) combines fake faces with synthesized or cloned lip-synced audio, while LAV-DF (Cai et al. 2022) and AV-Deepfake1M (Cai et al. 2024) introduce content-driven audio-visual manipulation and temporal localization. These benchmarks are important for multimodal fake detection, but they generally treat audio manipulation as speech, identity, or content manipulation rather than separating speech and environmental audio as independent components. 2.2 Speech and Environmental Audio Deepfakes Speech-only spoofing benchmarks (Wang et al. 2024; Yi et al. 2023; Zhang et al. 2023) have supported research on detecting synthetic and converted speech, replay attacks, and partially spoofed utterances. However, these benchmarks focus on speech and do not evaluate non-speech background audio. Recent advances in text-to-audio generation have made background-sound manipulation increasingly realistic, creating a forensic challenge that speech-centric benchmarks are not designed to address. This has given rise to a nascent line of work on environmental audio deepfake detection: EnvSDD (Yin et al. 2025b) and Compspoof (Zhang et al. 2026a) focus on audio-only environmental sound deepfake detection, while VCapAV (Wang et al. 2025) studies audio-visual environmental manipulation. However, EnvSDD and Compspoof provide no visual grounding, while VCapAV does not independently control speech and environmental audio as separate forgery components. Neither benchmark supports the systematic evaluation of cross-component interactions. MADBench addresses these gaps by jointly controlling speech manipulation, environmental audio manipulation, and scene consistency within a single component-level audio-visual benchmark. 2.3 Deepfake Detection Models Deepfake detectors use different cues depending on modality. Speech anti-spoofing models (Tak et al. 2021) such as AASIST (Jung et al. 2022) focus on synthetic or converted speech artefacts, while audio-visual methods such as AVForensics (Zhu et al. 2023) compare cross-modal signals through lip-sync, speech-mouth alignment, or multimodal fusion. However, none of these methods can attribute a manipulation to a specific acoustic component, nor can they be directly evaluated on environmental audio deepfakes that MADBench is designed to expose. Beyond task-specific detectors, pretrained audio and multimodal encoders such as CLAP (Elizalde et al. 2023) and ImageBind (Girdhar et al. 2023) provide general-purpose representations that serve as non-task-specific baselines for probing whether component-level deepfake signals are recoverable without forensic fine-tuning. Recent omni models such as MiniCPM-o (Cui et al. 2026) also represent a qualitatively different detection paradigm: rather than relying on learned forensic features, they apply large-scale multimodal pretraining to reason about audio-visual content zero-shot. Evaluating these models on MADBench directly tests whether general multimodal intelligence transfers to component-level forensic detection, a question that existing benchmarks cannot address due to their lack of component-level annotation. 3 Dataset Construction 3.1 Overall Design Principles Figure 1: Overview of the MADBench dataset construction and benchmark evaluation pipeline. As illustrated in Figure 1, MADBench follows a modality-component hierarchy: each video comprises audio and visual modalities, while the audio modality is further decomposed into speech and environmental audio components. The visual stream is held fixed across all variants of a given source clip, while manipulation is applied exclusively to the speech component, the environmental audio component, or both. This design choice ensures that any difference in model behavior across variants is attributable to audio manipulation alone, rather than being confounded by visual artifacts. To further control evaluation difficulty and probe distinct model capabilities, we introduce a scene-consistency axis orthogonal to the manipulation type. Manipulated samples are categorized as either scene-matched, where the substituted environmental audio remains semantically plausible given the visible scene, or scene-mismatched, where it is intentionally inconsistent with the visual context. This distinction is critical because it separates two fundamentally different detection strategies: recognizing low-level synthesis artifacts or component-level manipulation cues, versus reasoning about high-level semantic coherence between the audio and visual streams. Together, these two axes, manipulation component and scene consistency form the structural foundation of the benchmark. 3.2 Source Preprocessing and Splitting We build MADBench from AVSpeech (Ephrat et al. 2018), which comprises real-world YouTube videos with visible speakers and naturally paired speech. For each source clip, the original video stream is retained, while MossFormer2 (Zhao et al. 2024) estimates a speech stem s s from the mixed waveform x. We define the non-speech residual as e^=x−s e=x- s, following the mixture-consistency principle that the separated components should reconstruct the original mixture (Wisdom et al. 2019). After quality control, s s and e e serve as the source-derived real speech and environmental-audio components. We validate the decomposition using independent speech-content and leakage checks. All speech stems are successfully transcribed with 95.9%95.9\% mean ASR coverage, while independent VAD and ASR checks detect no speech in any environmental residual. All residuals also satisfy a minimum environmental-energy criterion. These checks ensure that the retained clips contain intelligible speech and perceptible environmental audio without detectable speech leakage. Before fake-component generation, source clips are assessed for benchmark eligibility. We retain clips with valid media, usable English transcripts, sufficient spoken content, perceptible environmental audio, and no substantial music or media playback. The retained clips are approximately 44–1212 seconds long, making them suitable for controlled audio replacement and final audio-visual assembly. To reduce speaker and source leakage, all clips from the same detected speaker cluster or related source identifiers are assigned to the same training, validation, or test split. After source-level quality control, 1,892 clips remain eligible for subsequent four-way benchmark construction. These clips are partitioned approximately 7:1:27:1:2 into training, validation, and test sets. The split is fixed before fake-component generation, and all derived samples inherit the split assignment of their source clip. 3.3 Fake Speech Generation Fake speech is generated to match the duration of the source real speech and placed into the original speech regions, with non-speech regions left silent. We use three speech synthesis techniques to ensure broad coverage of diverse speaker manipulation strategies. Speech synthesis. We first use same-identity Text-to-Speech (TTS) models, which synthesises the source transcript using the source speaker as reference, preserving apparent speaker identity. F5-TTS (Chen et al. 2025) and IndexTTS2 (Zhou et al. 2026) are used for this branch. Cross-identity TTS synthesises the same transcript with a different target speaker. Voice conversion transforms the source speech toward a target speaker’s timbre while preserving source content and timing. SeedVC (Liu 2024) and Vevo-Timbre (Zhang et al. 2025) are used for this branch. Together, the three branches yield six generation settings, reducing dependence on any single manipulation mechanism. Speaker pairing. Cross-identity generation follows a fixed source-target pairing table, constrained to same-gender pairs from different source videos and different detected speaker clusters. Speaker embeddings are used to exclude near-duplicate voices, and target-speaker usage is balanced to prevent a small number of identities from dominating the pool. Each source clip is paired with four target speakers, providing redundancy for QC-based selection. ECAPA (Desplanques et al. 2020) similarity between source and target speakers is low (mean = 0.0268, max = 0.3040), confirming acoustic distinctiveness. Quality control. Generated speech components are assessed against automatic QC criteria covering generation duration, timeline alignment and ASR-based transcript preservation. Furthermore, same-identity TTS is checked for source-speaker preservation; cross-identity TTS and voice conversion are checked for target-speaker similarity and source-speaker leakage. 3.4 Fake Environmental Audio Generation Environmental-audio generation produces full-duration background components to replace the real environmental track. To construct scene-matched and scene-mismatched variants systematically, we require a shared definition of what sounds belong to each type of visible scene. We therefore build a visual-scene taxonomy to guide both audio generation and mismatch selection. The pipeline comprises three stages: taxonomy construction, audio generation, and QC. Taxonomy Construction. To ensure reproducible scene matching, we first define a taxonomy of visual scenes, where each label specifies (i) expected environmental sounds, (i) excluded sound categories, and (i) mismatch scene types. The taxonomy is constructed exclusively from the training and validation set. Qwen2.5-VL-32B (Bai et al. 2025) summarizes the visual scene, expected environmental sounds, and potential audio confounders (e.g., speech or music), providing standardized descriptions for taxonomy construction. The resulting taxonomy is refined through an LLM-assisted review followed by manual verification to remove redundant labels, identify missing scene categories, and validate mismatch relationships. The final taxonomy contains 10 coarse and 63 fine-grained scene categories spanning indoor, outdoor, transportation, industrial, educational, and natural environments. Since the environmental component represents background ambience only, all labels explicitly exclude speech, singing, music, and media playback. After the taxonomy is finalized, all source clips are automatically assigned a fine-grained scene label following the predefined taxonomy. These labels determine prompt construction, scene-matched generation, and valid mismatch selection throughout the pipeline. Environmental Audio Generation. We generate environmental audio using three complementary generation paradigms. The Text-to-Audio (TTA) branch synthesizes background audio from textual scene descriptions using AudioLDM2-TTA (Liu et al. 2024a) and AudioGen (Kreuk et al. 2022). The Video-to-Audio (VTA) branch conditions generation on both visual content and text prompts using MMAudio (Cheng et al. 2025) and FoleyCrafter (Zhang et al. 2024). The Audio-to-Audio (ATA) branch edits or regenerates the environmental component while conditioning on the source background audio using AudioX (Tian et al. 2026) and AudioLDM2-ATA (Liu et al. 2024a). Together, these branches span prompt-conditioned, video-conditioned, and source-audio-conditioned generation, reducing dependence on any single generation strategy. For scene-matched generation, all branches produce environmental audio consistent with the labeled visual scene. For scene-mismatched generation, the TTA branch directly synthesizes audio from prompts describing a different scene category. Since VTA and ATA are inherently conditioned on the source clip, mismatched samples are instead created by replacing the generated environmental track with one originating from a semantically incompatible scene. Replacement tracks are selected within the same data split while excluding the same source clip and related source groups, and reuse is limited to reduce identity leakage and shortcut learning. All prompts are generated automatically from the scene taxonomy using fixed templates that incorporate coarse and fine-grained scene labels, inferred environmental sound events, candidate mismatch scenes, and clip duration. This rule-based procedure ensures consistent prompt quality and enforces the exclusion of speech and music across all generated samples. Quality Control. Generated components are retained if they satisfy QC criteria. Checks cover duration alignment, severe silence or clipping, detectable speech or music content, and valid replacement pairing for mismatch variants. Voice activity detection (VAD) and ASR detect speech leakage; AST-based event checks and CLAP-style audio-prompt matching scores serve as diagnostic signals. 3.5 Sample Assembly and Dataset Statistics Statistic Count Component Construction Generation-eligible source clips ,1,892 Speech components (generated / QC-retained) 34,056/24,61434,056/24,614 Environmental components (generated / QC-retained) 22,987/21,38822,987/21,388 MADBench Dataset Source clips (train / val. / test) ,(986/143/249)1,378\;(986/143/249) Samples per protocol (train / val. / test) ,(3,944/572/996)5,512\;(3,944/572/996) Final audio-visual samples (across 3 protocols) ,16,536 Table 1: Statistics of the MADBench dataset. Final assembly combines the QC-passed speech and environmental components into complete audio-visual samples under a shared media construction policy. For each source clip, four samples are constructed: a reassembled real sample, a speech-only fake, an environment-only fake, and a joint fake. All four variants use the same real video. Within each source, the selected fake speech component is shared by the speech-only and joint-fake variants, while the selected fake environmental component is shared by the environment-only and joint-fake variants. This ensures that controlled comparisons differ in only one audio component. Importantly, the real sample is also reassembled from separated real speech and real environmental audio through the same mixing procedure as the fake samples, ensuring that all sample types share an identical construction process, so that labels reflect component differences rather than assembly differences. Component selection is balanced across generation settings to prevent any single generator or branch from dominating the benchmark. Speech components are distributed across the six speech generation settings. Environmental components are balanced across TTA, VTA, ATA branches and generation models. Final audio is standardized and mixed using source-relative RMS normalization, so that speech and environmental-audio levels remain tied to the original source audio. The final audio is then paired with the real source video to produce the audio-visual sample. Table 1 summarizes the core statistics of MADBench. 4 Evaluation Setup 4.1 Tasks and Protocols All samples in MADBench share an authentic visual stream and differ only in the states of their two audio components. We denote the four sample types as R (real speech and real environmental audio), S (manipulated speech and real environmental audio), E (real speech and manipulated environmental audio), and Q (both audio components manipulated). We evaluate three core tasks. Binary any-fake detection separates R from the three manipulated states S, E, and Q. Four-way classification directly predicts R/S/E/QR/S/E/Q. Component-level evaluation independently determines whether the speech and environmental audio have been manipulated, allowing both components to be identified as fake in a Q sample. The three tasks are evaluated under three core protocols. The Balanced protocol includes both scene-matched and scene-mismatched environmental manipulations. The Matched protocol contains only manipulated environmental audio that remains plausible for the visible scene, whereas the Mismatched protocol uses scene-incompatible environmental audio. All three core protocols retain the same manipulation labels and evaluate audio manipulation detection and attribution, supplemented by controlled analysis protocols for scene consistency, generator generalization, and input modality ablation. We report standard metrics, including ROC-AUC, EER, accuracy, and Macro-F1. Task-specific diagnostics are defined in the corresponding result subsections. 4.2 Model Groups and Evaluation Procedure We evaluate three complementary model groups to examine whether component-level audio forgery cues arise from task-specific forensic training, general-purpose audio-visual pretraining, or zero-shot multimodal reasoning. First, to assess the transferability of existing forensic systems, we evaluate the pretrained A-V deepfake detectors AVH-Align (Smeu et al. 2025), AV Anomaly (Feng et al. 2023), and BA-TFD+ (Cai et al. 2023). Their native scores are used for binary direct transfer, while lightweight prediction heads are fitted to frozen detector outputs for tasks not supported by the original models. Second, to determine whether general-purpose A-V representations contain recoverable manipulation and cross-modal correspondence signals, we evaluate ImageBind (Girdhar et al. 2023), PE-AV Base (Vyas et al. 2025), and CAV-MAE Sync (Araujo et al. 2025) as frozen representation encoders with lightweight task-specific heads. Audio-only encoders (Elizalde et al. 2023; Chen et al. 2023, 2022), and pretrained audio detectors (Jung et al. 2022; Xiao and Das 2024; Kulkarni et al. 2026; Huang et al. 2026) are included as unimodal reference baselines. Third, to test whether large-scale multimodal pretraining produces zero-shot manipulation detection, component attribution, and scene-consistency reasoning, we evaluate Qwen2.5-Omni-7B (Xu et al. 2025), MiniCPM-o 4.5 (Cui et al. 2026), Gemma-4-E4B-it (Gemma Team et al. 2026), and Baichuan-Omni-1.5 (Li et al. 2025) using fixed task-specific prompts without additional training. All pretrained detectors and encoders remain frozen. The prediction heads use the same logistic-regression architecture but are trained independently for each task on the training split, with thresholds calibrated on the validation split. Omni models are evaluated zero-shot using fixed task-specific prompts without additional training. The task-specific prediction targets, notation, and prompt assignments are introduced in the corresponding result setups. 5 Results and Analysis We organise our analysis around four diagnostic questions: (1) whether existing models can detect and identify the manipulations; (2) whether models respond differently to speech and environmental audio manipulation; (3) whether models leverage visual-audio inconsistency or rely purely on acoustic artifacts; and (4) whether visual input improves detection beyond audio alone. 5.1 Can Existing Models Detect Deepfakes? Model Balanced Matched Mismatched Mean (a) Binary Any-Fake Detection (AUC↑ / EER↓ ) Pretrained A-V Detectors (direct transfer) AVH-Align 0.525 / 0.491 0.532 / 0.487 0.515 / 0.498 0.524 / 0.492 AV Anomaly 0.518 / 0.481 0.512 / 0.492 0.524 / 0.481 0.518 / 0.484 BA-TFD+ 0.524 / 0.487 0.526 / 0.493 0.552 / 0.475 0.534 / 0.485 Frozen A-V Encoders (with HbinH_bin) ImageBind A–V 0.912 / 0.163 0.912 / 0.160 0.918 / 0.165 0.914 / 0.162 PE-AV Base 0.943 / 0.120 0.930 / 0.131 0.953 / 0.108 0.942 / 0.120 CAV-MAE Sync 0.954 / 0.123 0.949 / 0.133 0.958 / 0.112 0.954 / 0.122 Omni Models (zero-shot PbinP_bin) MiniCPM-o 4.5 0.621 / 0.418 0.600 / 0.430 0.634 / 0.410 0.618 / 0.419 Qwen2.5-Omni-7B 0.540 / 0.479 0.523 / 0.490 0.542 / 0.470 0.535 / 0.480 Gemma-4-E4B-it 0.557 / 0.449 0.558 / 0.444 0.560 / 0.450 0.558 / 0.448 Baichuan-Omni-1.5 0.553 / 0.456 0.465 / 0.523 0.462 / 0.523 0.494 / 0.501 (b) Four-Way Manipulation Classification (Acc.↑ / Macro-F1↑ ) Pretrained A-V Detectors (with H4wayH_4way) AVH-Align 0.271 / 0.201 0.268 / 0.192 0.268 / 0.199 0.269 / 0.197 AV Anomaly 0.269 / 0.187 0.254 / 0.177 0.265 / 0.188 0.263 / 0.184 BA-TFD+ 0.286 / 0.240 0.295 / 0.257 0.295 / 0.245 0.292 / 0.248 Frozen A-V Encoders (with H4wayH_4way) ImageBind A–V 0.590 / 0.588 0.594 / 0.593 0.597 / 0.595 0.594 / 0.592 PE-AV Base 0.745 / 0.745 0.711 / 0.709 0.756 / 0.756 0.737 / 0.737 CAV-MAE Sync 0.711 / 0.708 0.681 / 0.676 0.696 / 0.696 0.696 / 0.693 Omni Models (zero-shot P4wayP_4way) MiniCPM-o 4.5 0.249 / 0.106 0.250 / 0.105 0.256 / 0.118 0.252 / 0.110 Qwen2.5-Omni-7B 0.250 / 0.146 0.257 / 0.156 0.252 / 0.150 0.253 / 0.151 Gemma-4-E4B-it 0.250 / 0.100 0.250 / 0.100 0.250 / 0.100 0.250 / 0.100 Baichuan-Omni-1.5 0.250 / 0.117 0.254 / 0.125 0.250 / 0.117 0.251 / 0.120 Table 2: (a) Binary any-fake detection and (b) four-way R/S/E/QR/S/E/Q classification across the three core protocols. Mean denotes the average across protocols. Setup. We compare the three model groups at two levels of difficulty: binary detection tests whether they distinguish authentic from manipulated audio, while four-way classification tests whether the learned cues support fine-grained manipulation attribution. The binary head HbinH_bin distinguishes fully real samples R from S/E/QS/E/Q, whereas H4wayH_4way directly predicts R/S/E/QR/S/E/Q. Pretrained A-V detectors are evaluated using their native predictions to measure direct transfer. For four-way classification, detector features are frozen and a logistic-regression head H4wayH_4way is fitted on top. Frozen A-V encoders are similarly adapted with a binary and 4-way logistic-regression head. Omni models use the zero-shot Binary Detection Prompt (PbinP_bin) and Four-Way Classification Prompt (P4wayP_4way) respectively. Pretrained A-V detectors fail to transfer. The three pretrained detectors perform near chance in binary direct transfer (mean AUC: 0.518–0.534). To test whether this stems from a distribution shift between their training data and MADBench, we train a benchmark-specific binary classification head on their frozen features. Performance remains essentially unchanged, with only a 0.012 AUC gain for BA-TFD+, indicating that the frozen representations themselves lack transferable cues rather than being limited by the original decision heads. A likely reason is that existing A-V detectors are designed for face-centric videos, whereas MADBench contains diverse natural scenes beyond a single face. Four-way classification is likewise close to chance. The matched and mismatched protocols yield comparable performance across both tasks, suggesting that scene mismatch has little impact on either detection or manipulation identification. Frozen A-V encoders transfer strongly. In clear contrast, all three frozen encoders achieve over 0.91 mean AUC for binary classification and maintain strong performance on four-way classification. Compared with task-specific deepfake detectors, these broadly pretrained audio-visual encoders provide substantially more transferable representations. Performance consistently improves under the mismatched protocol, indicating that audio-visual scene inconsistency provides an additional discriminative signal. However, the improvement is modest across all encoders, suggesting that matched scenes already contain sufficient information for reliable detection. Omni models detect coarse anomalies but not components. Zero-shot omni models show limited binary sensitivity, with MiniCPM-o 4.5 achieving the best mean AUC of 0.618, but this does not extend to fine-grained four-way classification with all models remaining near chance. This gap reflects a fundamental mismatch between general multimodal pretraining and the forensic reasoning required to attribute manipulation to a specific acoustic source. Notably, performance does not improve under the mismatched protocol, confirming that visible scene-audio contradiction provides no additional attribution signal for these models. Takeaway. The three model groups establish a clear transfer hierarchy: pretrained A-V detectors fail to capture the required forgery structure, zero-shot omni models provide only coarse sensitivity, and frozen A-V encoders support both reliable detection and fine-grained manipulation attribution. 5.2 Do Models Respond Differently to Speech and Environmental Audio Manipulations? Model Comp. AUC↑ − _E-S , G_S,\,G_E Bal. Mat. Mis. Mean Pretrained A-V Detectors (with HSH_S and HEH_E) AVH-Align S 0.550 0.555 0.535 0.547 -0.034 0.008 E 0.517 0.508 0.513 0.513 -0.010 AV Anomaly S 0.516 0.510 0.506 0.511 0.033 0.012 E 0.549 0.533 0.547 0.543 -0.012 BA-TFD+ S 0.484 0.509 0.500 0.498 0.094 0.030 E 0.583 0.594 0.599 0.592 -0.047 Frozen A-V Encoders (with HSH_S and HEH_E) ImageBind A-V S 0.797 0.803 0.793 0.798 0.106 0.138 E 0.898 0.896 0.919 0.904 0.002 PE-AV Base S 0.895 0.886 0.896 0.892 0.044 0.046 E 0.940 0.914 0.954 0.936 -0.012 CAV-MAE Sync S 0.857 0.847 0.844 0.849 0.113 0.137 E 0.963 0.951 0.974 0.963 -0.017 Omni Models (zero-shot PcompP_comp) MiniCPM-o 4.5 S 0.470 0.463 0.489 0.474 0.022 0.036 E 0.499 0.486 0.504 0.496 0.031 Qwen2.5-Omni-7B S 0.500 0.502 0.492 0.498 -0.007 0.008 E 0.489 0.497 0.487 0.491 0.006 Gemma-4-E4B-it S 0.487 0.479 0.474 0.480 -0.005 0.034 E 0.477 0.470 0.479 0.475 0.034 Baichuan-Omni-1.5 S 0.486 0.485 0.494 0.488 +0.000 -0.006 E 0.497 0.483 0.485 0.488 -0.003 Table 3: Component-level binary detection performance for speech (S) and environmental-audio (E) manipulation across the balanced (Bal.), scene-matched (Mat.), and scene-mismatched (Mis.) protocols. GSG_S and GEG_E denote the speech and environmental interference gaps reported in the S and E rows, respectively. Setup. Although speech has dominated prior audio deepfake detection research, it remains unclear whether this produces a systematic advantage of speech deepfake detection over environmental audio deepfake detection. To test this, we decompose the four-way task into two binary detection problems: a speech detection head HSH_S that distinguishes samples with fake speech (S,QS,Q) from those without (R,ER,E), and an environmental detection head HEH_E that distinguishes samples with fake environmental audio (E,QE,Q) from those without (R,SR,S). The two heads are fitted independently on frozen detector outputs, whereas omni models use the zero-shot component detection prompt PcompP_comp. We summarize the relative detection advantage as ΔE−S=AUC¯E−AUC¯S _E-S= AUC_E- AUC_S, where a positive value indicates stronger sensitivity to environmental manipulation. We further measure component interference: whether detecting one manipulated component becomes harder when the other component is also manipulated. For speech, GS=AUC(S vs. R)−AUC(Q vs. E)G_S=AUC(S vs. R)-AUC(Q vs. E) compares speech detection when environmental audio is real versus fake. For environmental audio, GE=AUC(E vs. R)−AUC(Q vs. S)G_E=AUC(E vs. R)-AUC(Q vs. S) compares environmental detection when speech is real versus fake. A positive gap indicates that manipulating the other component degrades detection of the target component. Transferred detectors show no stable component preference. The pretrained A-V detectors do not exhibit the expected universal speech advantage. The comparable performance across balanced, matched, and mismatched cases across E and S confirms that there are no differences in detecting speech and audio deepfakes using existing detectors. This may still be due to the misalignment of the training objectives of these detectors, which are more face-centric and fails in both speech and environmental audio settings. Environmental manipulation is consistently easier for frozen A-V encoders. Environmental manipulation is detected more accurately than speech manipulation across all encoders and evaluation protocols. Notably, the environmental advantage persists under the matched protocol, indicating that it cannot be explained solely by audio-visual scene inconsistency. Instead, the consistent gains across all three encoders suggest that large-scale audio-visual pretraining yields transferable representations that are particularly effective for environmental audio manipulation, with PE-AV’s more diverse pretraining leading to the most balanced performance across the two components. Omni models do not separate the two components. Zero-shot omni models remain close to chance for both speech and environmental manipulation, with no consistent direction in ΔE−S _E-S. Their general multimodal perception capabilities therefore do not translate into component-specific forensic decisions without task supervision. Fake environmental audio obscures speech-specific deepfake detection. GSG_S and GEG_E are close to zero for most of the models, indicating limited cross-component interference. The clear exceptions are ImageBind and CAV-MAE Sync, whose GSG_S values reach 0.1380.138 and 0.1370.137, respectively. The positive GSG_S shows that fake environmental audio obscures speech-specific cues; however, environmental detection remains stable when speech is also manipulated. This reflects the component structure, as environmental manipulation spans the acoustic context of the clip, while speech manipulation is concentrated in speech-active regions and must be isolated from the surrounding soundscape. Takeaway. Environmental manipulation is consistently easier to detect than speech and also the dominant source of cross-component interference. Models pretrained on diverse sound events and audio-visual correspondence perform better than speech-focused detectors and zero-shot omni models. 5.3 Do Models Exploit Scene-Audio Consistency? D-R S-R D-S Model ↑P\! ↑Sh\! ↑ _pair\! ↑P\! ↑Sh\! ↑ _pair\! ↑P\! ↑Sh\! ↑ _pair\! Pretrained A-V Detectors (with HSCFMH_SC^FM) AVH-Align 0.500 0.502 -0.002 0.501 0.503 -0.002 0.500 0.505 -0.005 AV Anomaly 0.499 0.500 -0.001 0.491 0.492 -0.002 0.497 0.498 -0.001 BA-TFD+ 0.525 0.524 +0.001 0.507 0.509 -0.002 0.523 0.519 +0.004 Frozen A-V Encoders (with HSCFMH_SC^FM) ImageBind 0.744 0.719 +0.025 0.705 0.728 -0.023 0.551 0.531 +0.020 PE-AV Base 0.726 0.731 -0.005 0.728 0.712 +0.016 0.532 0.544 -0.012 CAV-MAE Sync 0.808 0.795 +0.013 0.789 0.782 +0.008 0.523 0.507 +0.016 Omni Models (zero-shot PSCP_SC) Qwen2.5-Omni-7B 0.500 0.490 +0.010 0.502 0.490 +0.012 0.495 0.498 -0.004 MiniCPM-o 4.5 0.510 0.525 -0.014 0.489 0.484 +0.005 0.484 0.520 -0.035 Gemma-4-E4B-it 0.502 0.498 +0.004 0.500 0.497 +0.002 0.505 0.505 +0.000 Baichuan-Omni-1.5 0.501 0.524 -0.024 0.505 0.516 -0.010 0.507 0.493 +0.015 Table 4: Audio-visual scene-consistency performance (AUC) with paired (P) and shuffled (ShSh) video inputs. Δpair _pair measures the benefit of the intended video and only positive gains are highlighted. Setup. When a model succeeds in the fixed-video setting, is it responding to acoustic cues in the audio, or to the semantic mismatch between what is seen and heard? A model driven by acoustic cues can make its decision without using the video, whereas detecting scene-audio mismatch requires exploiting the cross-modal relationship between vision and sound. To separate these two sources of evidence, we construct three controlled scene-consistency conditions. D-R uses real environmental audio from different coarse scenes; S-R uses real audio from the same coarse scene but different fine-grained labels; and D-S compares scene-consistent and scene-inconsistent samples in which both environmental tracks are synthetic. Comparing D-R with S-R tests whether models move beyond coarse-category matching to fine-grained scene consistency, while comparing D-R with D-S examines how synthetic acoustic characteristics affect consistency recognition. We formulate scene consistency as a binary task, where each sample is classified as scene-consistent or scene-inconsistent. Since this target differs from the authenticity and component labels defined above, supervised models use a dedicated pooled Full-Mix Scene-Consistency Head HSCFMH_SC^FM. The head is trained on the union of the paired D-R, S-R, and D-S samples with both speech and environmental audio retained, rather than fitting a separate head for each condition. Omni models make the same binary decision using the zero-shot Scene-Consistency Prompt (PSCP_SC). At test time, the Paired condition (P) uses the intended video, whereas the Shuffled condition (ShSh) keeps the audio and label unchanged and replaces only the video. We measure the contribution of the correct visual stream as Δpair=AUCP−AUCSh, _pair=AUC_P-AUC_Sh, where a positive value indicates that the intended video improves scene-consistency detection. Pretrained detectors do not capture scene consistency. All three pretrained detectors remain close to chance across D-R, S-R, and D-S. Their paired and shuffled results are nearly identical, indicating that the correct video contributes little to the decision. BA-TFD+ shows only a small improvement over chance on D-R and D-S, and this advantage largely remains after video shuffling, suggesting that it is driven by audio-side cues rather than scene grounding. Their objectives focus on speech-visual alignment rather than the consistency between environmental audio and the visual scene, limiting their transfer to environmental scene-consistency reasoning. Frozen encoders use both audio cues and visual context. Frozen A-V encoders perform strongly on both D-R and S-R. CAV-MAE Sync reaches paired AUCs of 0.8080.808 and 0.7890.789, respectively, while PE-AV maintains nearly identical performance across the two conditions. ImageBind shows the largest drop, from 0.7440.744 to 0.7050.705, indicating comparatively weaker fine-grained consistency recognition. Overall, the strong S-R results indicate that these representations capture fine-grained scene correspondence rather than relying only on coarse scene categories. Frozen A-V encoders exhibit the clearest paired-video benefit among the three model groups, showing that they make the most effective use of visual scene information. These gains are most apparent under D-R and D-S, whereas S-R yields smaller and model-dependent benefits, reflecting the greater difficulty of fine-grained correspondence. The high shuffled scores on D-R and S-R further show that audio-internal cues also remain an important source of evidence. Fine-grained alignment provides the most consistent visual benefit. Under D-S, models across all three groups perform close to chance. Because both consistency labels contain synthetic environmental audio, synthesis status and generic generation artifacts cannot directly determine the label. This broad performance drop exposes the difficulty of scene grounding when real-recording consistency cues are unavailable. Nevertheless, CAV-MAE Sync achieves positive Δpair _pair in all three conditions, making it the only encoder that consistently benefits from the correct video. This suggests that fine-grained audio-frame alignment preserves sample-specific scene correspondence more effectively than relying only on global cross-modal similarity. Omni models fail at zero-shot scene grounding. All four omni models remain close to chance across the three conditions, and their paired-shuffled differences are small and inconsistent. Their responses therefore provide no consistent evidence of zero-shot visual grounding. Takeaway. Pretrained detectors and zero-shot omni models remain near chance with negligible or inconsistent video pairing gains, whereas frozen A-V encoders capture both fine-grained consistency cues and modest visual grounding. The broad performance drop on D-S identifies synthetic scene correspondence as the most challenging setting. 5.4 Is Audio Alone Sufficient for Manipulation Detection? Model Input Binary AUC↑ 4-way Macro-F1↑ Speech AUC↑ Env. AUC↑ Frozen Audio Encoders CLAP A 0.969 0.775 0.915 0.966 BEATs A 0.926 0.659 0.815 0.958 WavLM A 0.904 0.623 0.862 0.858 Pretrained Audio Detectors AASIST A 0.766 0.497 0.725 0.845 XLSR-Mamba A 0.857 0.627 0.840 0.874 DF-Arena-1B A 0.967 0.826 0.991 0.945 AudioMosaic A 0.949 0.732 0.898 0.935 Frozen A-V Encoders ImageBind A 0.961 0.705 0.890 0.950 A-V 0.914 0.592 0.798 0.904 PE-AV Base A 0.963 0.787 0.937 0.968 A-V 0.942 0.737 0.892 0.936 CAV-MAE Sync A 0.975 0.763 0.904 0.985 A-V 0.954 0.693 0.849 0.963 Zero-Shot Omni Models MiniCPM-o 4.5 A 0.599 0.136 0.509 0.517 A-V 0.618 0.110 0.474 0.496 Qwen2.5-Omni-7B A 0.598 0.135 0.496 0.557 A-V 0.535 0.151 0.498 0.491 Gemma-4-E4B-it A 0.532 0.114 0.474 0.509 A-V 0.558 0.100 0.480 0.475 Baichuan-Omni-1.5 A 0.497 0.192 0.516 0.539 A-V 0.494 0.120 0.488 0.488 Table 5: Input-modality ablation for binary, four-way, speech, and environmental manipulation detection. A and A-V denote audio-only and audio-visual inputs, respectively. Setup. Unlike the preceding scene-consistency analysis, this input-modality ablation examines the input requirements of the manipulation-detection tasks introduced earlier. For the binary, four-way, speech, and environmental targets, A receives only the final mixed audio, whereas A-V receives the identical audio together with its authentic video. Audio encoders and detectors provide unimodal reference baselines. For each frozen A-V encoder, the two input modes use independently fitted versions of the same task-specific heads. Omni models receive the same task prompt and audio in both modes, with video as the only additional input. Results are averaged across the three core protocols. Broad acoustic pretraining provides strong unimodal baselines. Audio-only encoders and detectors perform strongly across all four tasks. They capture diverse acoustic event patterns beyond speech, supporting strong detection of both speech and environmental manipulations. Among the pretrained audio detectors, DF-Arena-1B and AudioMosaic achieve the strongest overall results. This advantage is consistent with their training coverage: DF-Arena-1B includes environmental deepfake data alongside speech and singing corpora, while the evaluated AudioMosaic checkpoint is fine-tuned on the EnvSDD TTA split. These results highlight the value of combining broad acoustic representations with direct exposure to environmental-sound spoofing, rather than relying on speech anti-spoofing alone. Frozen A-V encoders are consistently stronger with audio alone. For ImageBind, PE-AV, and CAV-MAE Sync, audio-only input outperforms A-V input on every reported task. This consistent pattern follows the benchmark target: the manipulation labels are determined by the states of the audio components, while the visual stream remains authentic across all classes. Audio therefore provides the most direct forensic evidence for these labels, whereas adding visual and relation features provides no further separation under the unified lightweight-head protocol. Although it does not improve manipulation classification, the visual stream remains essential for scene-audio consistency, providing the semantic reference needed to verify whether the audio matches the visible scene. Omni models show no systematic multimodal advantage. Adding video produces isolated binary improvements for some omni models, but these gains do not extend consistently to four-way or component-level detection. Audio-only input therefore provides the more stable zero-shot setting for forensic decisions. Takeaway. Across the manipulation-detection tasks evaluated here, audio provides the most reliable direct forensic signal. Broad audio representations and domain-relevant anti-spoofing models provide strong component cues, while frozen A-V encoders consistently perform best in audio-only mode and omni models gain no systematic benefit from video for these labels. The visual stream remains central to modality-aware evaluation, as it anchors scene-audio correspondence and enables models to be assessed beyond acoustic artefact detection alone. 6 Conclusion We introduced MADBench, a component-level audio-visual deepfake benchmark, where speech and environmental audio are independently manipulated over authentic video. Experiments show that existing pretrained A-V detectors transfer poorly and zero-shot omni models remain unreliable, whereas frozen A-V encoders support strong detection and attribution. Environmental manipulation is generally easier to detect and can obscure speech-specific cues, while video contributes mainly to scene-consistency reasoning rather than manipulation detection. These findings highlight the importance of component-aware audio-visual deepfake evaluation. References E. Araujo, A. Rouditchenko, Y. Gong, S. Bhati, S. Thomas, B. Kingsbury, L. Karlinsky, R. Feris, J. R. Glass, and H. Kuehne (2025) Cav-mae sync: improving contrastive audio-visual mask autoencoders via fine-grained alignment. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 18794–18803. Cited by: §4.2. S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin (2025) Qwen2.5-vl technical report. External Links: 2502.13923, Link Cited by: §3.4. Z. Cai, S. Ghosh, A. P. Adatia, M. Hayat, A. Dhall, T. Gedeon, and K. Stefanov (2024) AV-deepfake1m: a large-scale llm-driven audio-visual deepfake dataset. In Proceedings of the 32nd ACM International Conference on Multimedia, M ’24, New York, NY, USA, p. 7414–7423. External Links: ISBN 9798400706868, Link, Document Cited by: §2.1. Z. Cai, S. Ghosh, A. Dhall, T. Gedeon, K. Stefanov, and M. Hayat (2023) Glitch in the matrix: a large scale benchmark for content driven audio-visual forgery detection and localization. Computer Vision and Image Understanding 236, p. 103818. External Links: Document Cited by: §1, §4.2. Z. Cai, K. Stefanov, A. Dhall, and M. Hayat (2022) Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization. In 2022 International Conference on Digital Image Computing: Techniques and Applications (DICTA), p. 1–10. Cited by: §1, §2.1. S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, et al. (2022) Wavlm: large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing 16 (6), p. 1505–1518. Cited by: §4.2. S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei (2023) BEATs: audio pre-training with acoustic tokenizers. In Proceedings of the 40th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 202, p. 5178–5193. Cited by: §4.2. Y. Chen, Z. Niu, Z. Ma, K. Deng, C. Wang, J. Zhao, K. Yu, and X. Chen (2025) F5-TTS: a fairytaler that fakes fluent and faithful speech with flow matching. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, p. 6255–6271. External Links: Document, Link Cited by: §1, §3.3. H. K. Cheng, M. Ishii, A. Hayakawa, T. Shibuya, A. Schwing, and Y. Mitsufuji (2025) MMAudio: taming multimodal joint training for high-quality video-to-audio synthesis. External Links: 2412.15322, Link Cited by: §3.4. J. Cui, B. Xu, C. Wang, T. Yu, W. Sun, Y. Xu, T. Wang, Z. He, W. Ma, T. Cai, J. Gui, L. Zhang, X. Sun, F. Huang, M. Chen, Z. Lin, H. Liu, Q. Gui, Q. Han, Y. Wen, H. Liu, R. Wang, Y. Zhang, H. Wei, C. Chen, Y. Li, K. Fang, J. Zhou, Y. Li, G. Zeng, C. Xiao, Y. Lin, X. Han, M. Sun, Z. Liu, and Y. Yao (2026) MiniCPM-o 4.5: towards real-time full-duplex omni-modal interaction. External Links: 2604.27393, Link Cited by: §2.3, §4.2. B. Desplanques, J. Thienpondt, and K. Demuynck (2020) ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification. In Proceedings of Interspeech, p. 3830–3834. External Links: Document Cited by: §3.3. B. Dolhansky, J. Bitton, B. Pflaum, J. Lu, R. Howes, M. Wang, and C. C. Ferrer (2020) The deepfake detection challenge (dfdc) dataset. arXiv preprint arXiv:2006.07397. Cited by: §1. B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang (2023) Clap learning audio concepts from natural language supervision. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. Cited by: §2.3, §4.2. A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein (2018) Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation. ACM Trans. Graph. 37 (4). External Links: ISSN 0730-0301, Link, Document Cited by: §3.2. C. Feng, Z. Chen, and A. Owens (2023) Self-supervised video forensics by audio-visual anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 10491–10503. External Links: Document Cited by: §4.2. J. Frank and L. Schönherr (2021) Wavefake: a data set to facilitate audio deepfake detection. arXiv preprint arXiv:2111.02813. Cited by: §1. Gemma Team, S. E. El Abd, V. Aggarwal, R. Algayres, A. Andreev, O. Bachem, I. Ballantyne, C. Brick, V. Cărbune, M. Casbon, et al. (2026) Gemma 4 technical report. arXiv preprint arXiv:2607.02770. Cited by: §4.2. R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra (2023) ImageBind: one embedding space to bind them all. External Links: 2305.05665, Link Cited by: §2.3, §4.2. H. Huang, Q. Wang, X. Ma, C. Xie, C. Leckie, and S. Erfani (2026) AudioMosaic: contrastive masked audio representation learning. External Links: 2605.14231, Link Cited by: §4.2. L. Jiang, R. Li, W. Wu, C. Qian, and C. C. Loy (2020) DeeperForensics-1.0: a large-scale dataset for real-world face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §1, §2.1. J. Jung, H. Heo, H. Tak, H. Shim, J. S. Chung, B. Lee, H. Yu, and N. Evans (2022) Aasist: audio anti-spoofing using integrated spectro-temporal graph attention networks. In ICASSP 2022-2022 IEEE international conference on acoustics, speech and signal processing (ICASSP), p. 6367–6371. Cited by: §2.3, §4.2. H. Khalid, S. Tariq, M. Kim, and S. S. Woo (2021) FakeAVCeleb: a novel audio-video multimodal deepfake dataset. arXiv preprint arXiv:2108.05080. Cited by: §1, §2.1. F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi (2022) Audiogen: textually guided audio generation. arXiv preprint arXiv:2209.15352. Cited by: §1, §3.4. A. Kulkarni, S. Dowerah, A. Kulkarni, T. Alumäe, and M. M. Doss (2026) Do compact ssl backbones matter for audio deepfake detection? a controlled study with raptor. External Links: 2603.06164, Link Cited by: §4.2. Y. Li, J. Liu, T. Zhang, S. Chen, T. Li, Z. Li, L. Liu, L. Ming, G. Dong, D. Pan, et al. (2025) Baichuan-omni-1.5 technical report. arXiv preprint arXiv:2501.15368. Cited by: §4.2. Y. Li, X. Yang, P. Sun, H. Qi, and S. Lyu (2020) Celeb-df: a large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 3207–3216. Cited by: §1. H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley (2024a) Audioldm 2: learning holistic audio generation with self-supervised pretraining. IEEE/ACM Transactions on Audio, Speech, and Language Processing 32, p. 2871–2883. Cited by: §3.4. S. Liu (2024) Zero-shot voice conversion with diffusion transformers. External Links: 2411.09943, Link Cited by: §3.3. W. Liu, T. She, J. Liu, B. Li, D. Yao, Z. Liang, and R. Wang (2024b) Lips are lying: spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes. Advances in Neural Information Processing Systems 37, p. 91131–91155. Cited by: §1. X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch, et al. (2023) Asvspoof 2021: towards spoofed and deepfake speech detection in the wild. IEEE/ACM Transactions on Audio, Speech, and Language Processing 31, p. 2507–2522. Cited by: §1. A. Rossler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nießner (2019) Faceforensics++: learning to detect manipulated facial images. In Proceedings of the IEEE/CVF international conference on computer vision, p. 1–11. Cited by: §1, §2.1. S. Smeu, D. Boldisor, D. Oneata, and E. Oneata (2025) Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 18815–18825. Cited by: §4.2. H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher (2021) End-to-end anti-spoofing with rawnet2. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 6369–6373. Cited by: §2.3. Z. Tian, Z. Liu, Y. Jin, R. Yuan, L. Xue, X. Tan, Q. Chen, W. Xue, and Y. Guo (2026) AudioX: a unified framework for anything-to-audio generation. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §3.4. A. Vyas, H. Chang, C. Yang, P. Huang, L. Gao, J. Richter, S. Chen, M. Le, P. Dollár, C. Feichtenhofer, et al. (2025) Pushing the frontier of audiovisual perception with large-scale multimodal correspondence learning. arXiv preprint arXiv:2512.19687. Cited by: §4.2. X. Wang, H. Delgado, H. Tak, J. Jung, H. Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. Kinnunen, et al. (2024) ASVspoof 5: crowdsourced speech data, deepfakes, and adversarial attacks at scale. arXiv preprint arXiv:2408.08739. Cited by: §2.2. X. Wang, J. Yamagishi, M. Todisco, H. Delgado, A. Nautsch, N. Evans, M. Sahidullah, V. Vestman, T. Kinnunen, K. A. Lee, et al. (2020) ASVspoof 2019: a large-scale public database of synthesized, converted and replayed speech. Computer Speech & Language 64, p. 101114. Cited by: §1. Y. Wang, Y. Wang, Q. Zhang, H. Nishizaki, and M. Li (2025) VCapAV: a video-caption based audio-visual deepfake detection dataset. In Proceedings of Interspeech, p. 3908–3912. External Links: Document Cited by: §2.2. S. Wisdom, J. R. Hershey, K. Wilson, J. Thorpe, M. Chinen, B. Patton, and R. A. Saurous (2019) Differentiable consistency constraints for improved deep speech enhancement. In ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vol. , p. 900–904. External Links: Document Cited by: §3.2. Y. Xiao and R. K. Das (2024) XLSR-mamba: a dual-column bidirectional state space model for spoofing attack detection. External Links: 2411.10027, Link Cited by: §4.2. J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang, B. Zhang, X. Wang, Y. Chu, and J. Lin (2025) Qwen2.5-omni technical report. External Links: 2503.20215, Link Cited by: §4.2. J. Yi, J. Tao, R. Fu, X. Yan, C. Wang, T. Wang, C. Y. Zhang, X. Zhang, Y. Zhao, Y. Ren, et al. (2023) Add 2023: the second audio deepfake detection challenge. arXiv preprint arXiv:2305.13774. Cited by: §2.2. H. Yin, Y. Xiao, R. K. Das, J. Bai, and T. Dang (2025a) Environmental sound deepfake detection challenge: an overview. In ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Cited by: §1. H. Yin, Y. Xiao, R. K. Das, J. Bai, and T. Dang (2026) The first environmental sound deepfake detection challenge: benchmarking robustness, evaluation, and insights. In INTERSPEECH 2026, Cited by: §1. H. Yin, Y. Xiao, R. K. Das, J. Bai, H. Liu, W. Wang, and M. D. Plumbley (2025b) EnvSDD: benchmarking environmental sound deepfake detection. In Proceedings of Interspeech, p. 201–205. External Links: Document Cited by: §1, §2.2. L. Zhang, X. Wang, E. Cooper, N. Evans, and J. Yamagishi (2023) The partialspoof database and countermeasures for the detection of short fake speech segments embedded in an utterance. IEEE/ACM Transactions on Audio, Speech, and Language Processing 31, p. 813–825. Cited by: §2.2. X. Zhang, Y. Wang, L. Li, L. Jin, and M. Li (2026a) Compspoof: a dataset and joint learning framework for component-level audio anti-spoofing countermeasures. In ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 18067–18071. Cited by: §2.2. X. Zhang, H. Yin, Y. Xiao, L. Zhang, T. Dang, R. K. Das, and M. Li (2026b) Overview of esdd2: environment-aware speech and sound deepfake detection challenge. In 2026 IEEE International Conference on Multimedia and Expo (ICME), Cited by: §1. X. Zhang, X. Zhang, K. Peng, Z. Tang, V. Manohar, Y. Liu, J. Hwang, D. Li, Y. Wang, J. Chan, Y. Huang, Z. Wu, and M. Ma (2025) Vevo: controllable zero-shot voice imitation with self-supervised disentanglement. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §3.3. Y. Zhang, Y. Gu, Y. Zeng, Z. Xing, Y. Wang, Z. Wu, and K. Chen (2024) FoleyCrafter: bring silent videos to life with lifelike and synchronized sounds. External Links: 2407.01494, Link Cited by: §3.4. S. Zhao, Y. Ma, C. Ni, C. Zhang, H. Wang, T. H. Nguyen, K. Zhou, J. Yip, D. Ng, and B. Ma (2024) MossFormer2: combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 10356–10360. External Links: Document Cited by: §3.2. S. Zhou, Y. Zhou, Y. He, X. Zhou, J. Wang, W. Deng, and J. Shu (2026) IndexTTS2: a breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 35139–35148. External Links: Document Cited by: §1, §3.3. Y. Zhu, J. Gao, and X. Zhou (2023) AVForensics: audio-driven deepfake video detection with masking strategy in self-supervision. In Proceedings of the 2023 ACM International Conference on Multimedia Retrieval, p. 162–171. Cited by: §2.3.