Paper deep dive
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Aastha Sharma, Guangjing Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/18/2026, 2:56:33 PM
Summary
The paper introduces VoxENES 2026, a bilingual benchmark dataset designed to evaluate the generalization of speech spoofing detectors against modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems. The dataset comprises 53,628 audio samples generated using 10 contemporary synthesis methods and subjected to 10 standardized post-processing augmentations. Evaluation of eight pretrained detectors reveals a significant performance degradation, with the best model achieving only 28.98% Equal Error Rate (EER), indicating that current detectors rely on brittle artifacts and fail to generalize to modern synthetic speech and realistic transmission conditions.
Entities (16)
Relation Signals (12)
AST-ASVspoof5 → achievedeeron → VoxENES 2026
confidence 95% · AST-ASVspoof5 achieves an EER below 30% (i.e., 28.98%) ... on VoxENES 2026
AASIST2 → achievedeeron → VoxENES 2026
confidence 95% · AASIST2 (57.86%) ... on this benchmark
VoxENES 2026 → containssamplesgeneratedby → Seed-VC
confidence 95% · VoxENES 2026 ... generated using 10 contemporary speech synthesis methods ... Seed-VC
VoxENES 2026 → containssamplesgeneratedby → VoxCPM 1.5
confidence 95% · VoxENES 2026 ... generated using 10 contemporary speech synthesis methods ... VoxCPM 1.5
VoxENES 2026 → containssamplesgeneratedby → Qwen3-TTS
confidence 95% · VoxENES 2026 ... generated using 10 contemporary speech synthesis methods ... Qwen3-TTS
VoxPopuli → sourceofrealspeechfor → VoxENES 2026
confidence 90% · We drew 1,528 real Spanish speech samples from 96 speakers from VoxPopuli
LibriSpeech → sourceofrealspeechfor → VoxENES 2026
confidence 90% · we drew 1,500 real English speech samples from 40 speakers from LibriSpeech
VoxENES 2026 → usesaugmentation → AAC 128k
confidence 90% · VoxENES 2026 explicitly models the post-processing setting ... including codec compression ... AAC at 128 kbps
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish) benchmark of 53,628 audio samples generated using 10 contemporary speech synthesis methods and evaluated under 10 standardized post-processing conditions. Using VoxENES 2026, we benchmark eight pretrained detectors without fine-tuning and observe substantial performance degradation: the best model achieves 28.98\% EER overall, while most perform near or below random chance across modern generators and perturbations. Our results highlight the reliance on brittle artifacts in current detectors and establish VoxENES 2026 as a practical testbed for developing robust audio spoofing countermeasures.
Tags
Links
- Source: https://arxiv.org/abs/2607.11706v1
- Canonical: https://arxiv.org/abs/2607.11706v1
Trouble viewing inline? Open PDF directly →
Full Text
25,864 characters extracted from source content.
Expand or collapse full text
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion Aastha Sharma, Guangjing Wang University of South Florida, Tampa, FL, USA aasthasharma@usf.edu, guangjingwang@usf.edu Abstract Modern LLM-driven text-to-speech (TTS) and voice conver- sion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing bench- marks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post- processing conditions. We bridge this gap by introducing Vox- ENES 2026, a bilingual (English and Spanish) benchmark of 53,628 audio samples generated using 10 contemporary speech synthesis methods and evaluated under 10 standardized post- processing conditions.Using VoxENES 2026, we bench- mark eight pretrained detectors without fine-tuning and observe substantial performance degradation: the best model achieves 28.98% EER overall, while most perform near or below ran- dom chance across modern generators and perturbations. Our results highlight the reliance on brittle artifacts in current de- tectors and establish VoxENES 2026 as a practical testbed for developing robust audio spoofing countermeasures. Index Terms: audio deepfake detection, speech spoofing detec- tion, benchmark dataset 1. Introduction Robust speech spoofing and deepfake detection are essential to preserve trust in speech-based authentication [1, 2, 3, 4]. This need is growing as voice becomes a biometric and a con- trol channel for speech-driven agents and assistive technolo- gies. If synthetic speech becomes indistinguishable from gen- uine speech, detection failures not only compromise security protocols but also erode trust in audio evidence. Many benchmark datasets are proposed for speech spoofing and deepfake detection evaluation. For example, the ASVspoof challenge series has been the primary driver of spoofing coun- termeasure development. ASVspoof 2019 [5] introduces log- ical access (LA) with TTS and VC, and physical access (PA) with replay tracks. ASVspoof 2021 [6] adds the deepfake task targeting compressed manipulated speech, and ASVspoof 5 [7] introduces crowdsourced data with adversarial attacks at scale. In addition to ASVspoof, WaveFake [8] provides a multilingual dataset from six neural vocoder architectures. The In-the-Wild dataset [9] includes real-world deepfakes of celebrities. The MLAAD [10] dataset expands coverage to 23 languages and 54 TTS models. The VoiceWukong [11] benchmarks 12 de- tectors against 34 commercial and open-source tools with post- processing manipulations. Yet, existing benchmarks primarily rely on speech synthe- sis systems before 2024 and fail to capture artifact patterns pro- duced by the modern large language model (LLM)-driven gen- eration pipelines. For example, text-to-speech (TTS) designs include autoregressive language-model-based synthesis, such as VoxCPM [12] and Qwen3-TTS [13]; flow-matching mod- els, including GLM-TTS [14], CosyVoice 3 [15], and Chatter- box [16]; diffusion-based systems like FlashLabs Chroma [17]; and hybrid DiT architectures, such as VibeVoice [18]. For voice conversion (VC), zero-shot approaches such as Seed-VC [19], tone-color extraction methods such as OpenVoice v2 [20], and retrieval-based systems like RVC v2 [21] have substantially im- proved naturalness and speaker similarity. Evaluation on temporally stale benchmarks can overesti- mate real-world robustness as speech synthesis techniques im- prove. LLM-driven TTS and VC systems produce synthetic au- dio with acoustic characteristics that differ substantially from earlier spoofing corpora, creating a data drifting issue for exist- ing detectors. Deepfake detection models that perform well on older benchmarks may fail when deployed against newer deep- fake generators. This mismatch reflects a cat-and-mouse dy- namic in which detectors learn cues tied to previous synthesis artifacts, while generators and post-processing steps progres- sively suppress or conceal those cues. As a consequence, the rapid LLM-driven TTS and VC evolution motivates updated evaluation benchmarks. To study the generalization ability of deepfake detectors un- der the data drifting issue in the era of LLM, we introduce a modern bilingual benchmark dataset VoxENES 2026. We gen- erate synthetic audios from 10 LLM-driven synthesis methods (7 TTS, 3 VC), covering English and Spanish audios. In addi- tion, real-world audios are frequently subjected to transforma- tions such as compression, noise, and resampling, which can substantially affect detector behavior [6, 11, 10]. Therefore, VoxENES 2026 explicitly models the post-processing setting by applying a standardized set of audio augmentation methods that simulate common transmission and manipulation effects, including codec compression, additive noise, resampling, speed perturbation, and loudness normalization. With VoxENES 2026, we evaluate eight pretrained deepfake detectors without fine-tuning to measure out-of- distribution generalization. Our results show that existing de- tectors suffer from substantial degradation relative to reported performance on legacy benchmarks. The best detector achieves only 28.98% EER overall, while most detectors perform near or below random chance. These findings suggest that many current detectors rely on brittle artifact cues that may not gen- eralize across new TTS and VC generations and realistic post- processing. By providing a modern benchmark and a controlled out-of-distribution evaluation, our work establishes a necessary testbed for measuring real-world speech spoofing detector ro- bustness and for driving detector development that keeps pace with the evolving synthesis frontier. In summary, the main contributions of our work are: • We introduce a modern bilingual benchmark dataset Vox- arXiv:2607.11706v1 [cs.SD] 13 Jul 2026 Table 1: VoxENES 2026 dataset summary. ComponentCountNotes Real speech3,028LibriSpeech + VoxPopuli Original synthetic4,600TTS + VC originals Augmented synthetic46,00010× post-processing Total53,628EN + ES combined Table 2: Language and source breakdown in VoxENES 2026. CategoryENESTotal Real speech1,5001,5283,028 TTS (7 methods)11,0006,60017,600 VC (3 methods)16,50016,50033,000 Total29,00024,62853,628 ENES 2026 using LLM-driven TTS and VC systems with realistic post-processing to emulate deployment- time distribution shifts, totaling 53,628 audio samples available at https://w.kaggle.com/datasets/ interspeech2712/voxenes-2026. • We benchmark eight pretrained detectors and reveal a sub- stantial generalization gap, exposing detector-specific blind spots and robustness failures across synthesis methods. 2. VoxENES 2026 VoxENES 2026 consists of three components: (1) real speech from established corpora, (2) synthetic audio generated by mod- ern TTS and VC systems, and (3) post-processed augmented variants simulating real-world transmission conditions.As shown in Table 1, VoxENES includes 3,028 real speech sam- ples and 50,600 synthetic samples (original synthetic 4,600 and augmented synthetic 46,000). 2.1. Real Speech and Standardization As shown in Table 2, we drew 1,500 real English speech sam- ples from 40 speakers from LibriSpeech [22]. We drew 1,528 real Spanish speech samples from 96 speakers from VoxPop- uli [23]. All audio was standardized to 16 kHz mono WAV. For detector compatibility, samples were capped at 4 seconds with truncation and zero-padding. This fixed-length standard- ization was applied uniformly to both bonafide and synthetic samples so that any padding-induced cues are shared across classes rather than correlated with the spoofing label. We, there- fore, do not expect zero-padding to serve as a discriminative shortcut, though a detailed analysis of padding artifacts is left to future work. Note that for the downstream voice conversion tasks, the people whose voices appear in the real speech samples are not used as target speakers. 2.2. Synthesis Systems We selected seven TTS and three VC systems representing state-of-the-art LLM-driven synthesis techniques across diverse architectures as shown in Table 3. The TTS systems include: (i) VoxCPM 1.5 [12], an autoregressive language model from HKUST; (i) Qwen3-TTS [13], a streaming LM from Alibaba with native multilingual support; (i) GLM-TTS [14], a flow- matching model from Zhipu AI; (iv) FlashLabs Chroma [17], a diffusion-based system; (v) VibeVoice [18], combining DiT with flow matching; (vi) CosyVoice 3 [15], an LM plus flow- matching system from Alibaba; and (vii) Chatterbox ML [16], a flow-matching model from Resemble AI. The three VC systems represent distinct paradigms: (i) Seed-VC [19] uses diffusion- based zero-shot conversion with Whisper and WavLM features; (i) OpenVoice v2 [20] performs tone-color extraction and trans- fer; and (i) RVC v2 [21] uses HuBERT-based retrieval with HiFi-GAN vocoding. All TTS systems were released or up- dated in 2025, ensuring they represent modern threats that older detection models have never encountered. VC samples are gen- erated using disjoint source audio partitions and multiple target speaker references to maintain diversity. 2.3. Post-Processing Augmentations Each original synthetic sample was subjected to one post- processing operation in Table 4 that reflect common transmis- sion and editing effects: lossy codec compression (MP3 at 64 kbps and AAC at 128 kbps), additive noise perturbations (white noise at 10/20 dB SNR and babble noise at 15 dB SNR), band- width and sampling-rate distortions (downsampling to 8 kHz with restoration to 16 kHz, plus a 16 kHz resampling control), playback-rate changes (1.1× and 0.9× speed), and amplitude normalization (peak normalization to−3 dBFS). This pipeline expands the synthetic subset and enables systematic robustness analysis across realistic post-processing conditions. 3. Evaluation Setup 3.1. Detection Baselines We evaluated eight pretrained audio deepfake detection sys- tems spanning graph neural networks, raw-waveform CNNs, self-supervised learning (SSL) models, transformer-based de- tectors, and speaker-embedding anomaly detection as shown in Table 5. None were retrained or fine-tuned on VoxENES 2026, enabling an honest temporal generalization evaluation. The de- tectors differ in their training corpora (e.g., ASVspoof 2019 LA, ASVspoof 2021 DF, ASVspoof 5, and VoxCeleb), so absolute EER values should be read as out-of-distribution generalization indicators rather than as a strictly controlled head-to-head com- parison; the most directly comparable cases are detectors shar- ing the same training source. All detectors were evaluated in inference-only mode us- ing published pretrained weights. We report Equal Error Rate (EER) and accuracy in the results. EER was computed from the ROC operating point where the false acceptance rate (FAR) and the false rejection rate (FRR) are equal. In particular, ECAPA- TDNN [27] is a speaker verification model rather than a ded- icated spoof detector. Thus, we use an anomaly-based scor- ing scheme, where a reference centroid is computed from real speech audio, and the cosine distance from this centroid is used as the spoofing score. The system’s performance is then eval- uated using the EER, determined by the threshold at which the FAR and FRR are equivalent. 4. Evaluation Results 4.1. Qualitative Spectrogram Analysis Figures 1 and 2 visualize the spectral variability across bonafide (real) speech, synthetic speech, and post-processed synthetic samples in VoxENES 2026. The examples highlight that re- Table 3: TTS and VC systems used in VoxENES 2026. SystemTypeArchitectureDeveloperENESTotalYearES Mode VoxCPM 1.5 [12]TTSAutoregressive LMHKUST200–2002025– Qwen3-TTS [13]TTSStreaming LMAlibaba2002004002025Native GLM-TTS [14]TTSFlow matchingZhipu AI200–2002025– FlashLabs Chroma [17]TTSDiffusionFlashLabs200–2002025– VibeVoice [18]TTSDiT + flow matchingVibe AI200–2002025– CosyVoice 3 [15]TTSLM + flow matchingAlibaba–2002002025Native Chatterbox ML [16]TTSFlow matchingResemble AI–2002002025Native Seed-VC [19]VCDiffusion + zero-shotByteDance5005001,0002025– OpenVoice v2 [20]VCTone cloningMyShell AI5005001,0002024– RVC v2 [21]VCRetrieval + HiFi-GANRVC-Project5005001,0002023– 00.511.522.533.54 Time (s) 0.0 1 2 3 4 5 6 7 8 Frequency (kHz) (a) VoxCPM 1.5 + AAC 128k 00.511.522.533.54 Time (s) 0.0 1 2 3 4 5 6 7 8 Frequency (kHz) (b) GLM-TTS + White Noise (10 dB) 00.511.522.533.54 Time (s) 0.0 1 2 3 4 5 6 7 8 Frequency (kHz) (c) Chroma + White Noise (20 dB) 00.511.522.533.54 Time (s) 0.0 1 2 3 4 5 6 7 8 Frequency (kHz) (d) CosyVoice 3 Spanish (Original) 00.511.522.533.54 Time (s) 0.0 1 2 3 4 5 6 7 8 Frequency (kHz) (e) VibeVoice + Resample 16 kHz 00.511.522.533.54 Time (s) 0.0 1 2 3 4 5 6 7 8 Frequency (kHz) (f) Seed-VC + Speed 1.1× 00.511.522.533.54 Time (s) 0.0 1 2 3 4 5 6 7 8 Frequency (kHz) (g) OpenVoice v2 + Speed 0.9× 00.511.522.533.54 Time (s) 0.0 1 2 3 4 5 6 7 8 Frequency (kHz) (h) RVC v2 + Volume Norm 80 70 60 50 40 30 20 10 0 Power (dB) Figure 1: Representative spectrograms across multiple TTS and VC systems and augmentation conditions in VoxENES 2026. Table 4: Post-processing augmentation techniques. AugmentationDescription mp364kMP3 encoding at 64 kbps aac 128kAAC encoding at 128 kbps noisewhite10dbWhite Gaussian noise at 10 dB SNR noisewhite20dbWhite Gaussian noise at 20 dB SNR noisebabble15dbMulti-speaker babble noise at 15 dB SNR resample8kDownsample to 8 kHz, upsample to 16 kHz resample16kResample to 16 kHz (control) speedfast1.1× playback speed speedslow0.9× playback speed volumenormPeak normalization to−3 dBFS alistic perturbations, such as additive noise and lossy codec compression, can reshape time-frequency structure and atten- uate synthesis artifacts that many detectors implicitly rely on, providing a qualitative rationale for the observed performance changes under post-processing. 4.2. Impact of Post-Processing As shown in Table 6, post-processing impacts detector per- formance in a highly non-uniform and occasionally counter- intuitive manner. Adding white noise reduces EER for AST- ASVspoof5 (from 26.7% to 17.4%) and RawNet2 (from 51.3% Table 5: Detection baselines evaluated on VoxENES 2026. DetectorFamilyTraining Data AASIST2 [24]Graph NNASVspoof 2019 LA RawNet2 [25]Raw waveform CNNASVspoof 2021 DF Wav2Vec2-AASISTSSL + classifierASVspoof 2019 Wav2Vec2-DFSSL + classifierMixed deepfake Wav2Vec2-LargeSSL + classifierMixed deepfake AST-ASVspoof5 [26]Spectrogram TransformerASVspoof 5 Wav2Vec2-ASVspoof5SSL + classifierASVspoof 5 ECAPA-TDNN [27]Speaker embeddingsVoxCeleb to 27.9%).This is plausibly because noise perturbs syn- thetic and bonafide speech differently and attenuates generator- specific artifacts, thereby inducing alternative discriminative cues. In contrast, MP3 compression degrades AST-ASVspoof5 (from 26.7% to 48.4% EER), consistent with a reliance on fine- grained spectral structure, particularly in higher frequencies, that is suppressed by lossy codec encoding. 4.3. Overall Detection Performance As shown in Table 7, among the evaluated models, only AST- ASVspoof5 achieves an EER below 30% (i.e., 28.98%), a result that remains insufficient for reliable field deployment. In addi- tion, five of the eight detectors perform at or below the stochas- Table 6: Per-augmentation EER (%) across all detectors. AugmentationAASIST2RawNet2W2V-AASISTW2V-DFW2V-LargeAST-ASV5W2V-ASV5ECAPA original54.451.341.053.939.926.752.646.5 aac128k53.150.242.553.637.931.851.546.9 mp3 64k53.750.142.554.139.848.452.047.2 noisewhite10db67.527.932.369.561.217.452.533.3 noise white20db59.040.044.568.963.318.349.836.2 noisebabble15db65.633.022.950.845.125.746.344.5 resample 8k59.155.637.542.034.329.552.146.4 resample16k54.149.742.354.139.727.152.346.1 speedfast56.053.438.652.237.526.159.038.5 speedslow52.747.845.052.542.430.249.842.0 volume norm56.747.142.053.940.128.552.246.6 00.511.522.533.54 Time (s) 0.0 1 2 3 4 Frequency (kHz) (a) Bona fide Speech (English) 00.511.522.533.54 Time (s) 0.0 1 2 3 4 Frequency (kHz) (b) Original Qwen3-TTS Spanish (Synthetic) 00.511.522.533.54 Time (s) 0.0 1 2 3 4 Frequency (kHz) (c) Qwen3-TTS Spanish + Babble Noise (15 dB SNR) 00.511.522.533.54 Time (s) 0.0 1 2 3 4 Frequency (kHz) (d) Qwen3-TTS Spanish + MP3 64k Codec 80 70 60 50 40 30 20 10 0 Power (dB) Figure 2: Qualitative spectrogram comparison. Table 7: Overall detection results on VoxENES 2026. DetectorEER (%)Acc (%)TTS EER (%)VC EER (%) AASIST257.8642.1361.249.9 RawNet247.0352.9753.649.8 Wav2Vec2-AASIST39.1660.8543.738.4 Wav2Vec2-DF55.5144.4959.349.2 Wav2Vec2-Large44.3855.6240.039.7 AST-ASVspoof528.9875.9420.929.8 Wav2Vec2-ASVspoof551.5348.4950.653.5 ECAPA-TDNN43.2256.7828.552.5 tic baseline (EER≥ 47%), most notably AASIST2 (57.86%), which exhibits inverted prediction behavior on this benchmark and suggests a severe domain mismatch between the training distribution and the VoxENES 2026 benchmark. This phe- nomenon occurs when a classifier learns dataset-specific arti- facts, such as silence patterns or channel characteristics, that are reversed or absent in the target domain. Consequently, the model assigns higher bonafide scores to synthetic samples, mis- interpreting spoofing cues as genuine speech features. Across models, TTS-generated samples are marginally easier to detect than VC-generated samples for some detectors, but the relative difficulty varies by detector and does not hold uniformly. 4.4. Per-Method Detection Performance We evaluate different detectors on each speech synthesis method. As shown in Figure 3, Seed-VC is the most challeng- ing synthesis method overall: no evaluated detector achieves an EER below 41%. We hypothesize that its diffusion-based con- version pipeline, augmented with strong speech representations (e.g., Whisper- and WavLM-derived features), yields highly natural outputs that preserve many bonafide time–frequency characteristics, thereby reducing detector-accessible artifacts. More broadly, no single detector is consistently reliable across all synthesis methods. For example, ECAPA-TDNN performs well on several TTS systems (e.g., GLM-TTS at 10.2% and CosyVoice 3 at 11.5%) but degrades substantially on VC (e.g., OpenVoice v2 at 62.7%). The results suggest that speaker- embedding approaches may capture anomalies introduced by certain TTS pipelines yet struggle when VC more faithfully pre- serves speaker identity and natural speech structure. A promis- ing future direction is to explore alternative modalities, such as acoustic sensing data [28], for deepfake audio detection. VoxCPM 1.5Qwen3-TTSGLM-TTSChromaVibeVoice Method 0 10 20 30 40 50 60 70 EER (%) Detector AASIST2 RawNet2 W2V-AASIST W2V-DF W2V-Large AST-ASV5 W2V-ASV5 ECAPA CosyVoice 3Chatterbox MLSeed-VCOpenVoice v2RVC v2 Method 0 20 40 60 80 EER (%) Detector AASIST2 RawNet2 W2V-AASIST W2V-DF W2V-Large AST-ASV5 W2V-ASV5 ECAPA Figure 3: Per-method EER on Original Samples 5. Conclusion We presented VoxENES 2026, a modern bilingual benchmark for speech spoofing detection that reflects LLM-era TTS and VC generation and realistic post-processing. Using VoxENES 2026, we evaluated eight pretrained detectors without fine- tuning and observed a substantial generalization gap: the best model achieves only 28.98% EER overall, while many detectors perform near or below chance. These results suggest that many current countermeasures rely on brittle, benchmark-specific ar- tifacts and remain sensitive to modern generators and routine post-processing. We encourage future work on deployment- oriented spoofing detection that improves generalization under data drift and advances evaluation protocols that continuously track the fast-evolving TTS and VC frontier. 6. Use of Generative AI Disclosure Generative AI tools were used for language editing and manuscript polishing. All authors reviewed, verified, and take full responsibility for the content, experiments, and conclusions presented in this paper. 7. References [1] Y. Wang, B. Chen, H. Guo, G. Wang, W. Ding, and Q. Yan, “Clearmask: Noise-free and naturalness-preserving protection against voice deepfake attacks,” in Proceedings of the 20th ACM Asia Conference on Computer and Communications Security, 2025, p. 696–709. [2] H. Guo, G. Wang, B. Chen, Y. Wang, X. Zhang, X. Chen, Q. Yan, and L. Xiao, “Wavepurifier: Purifying audio adversarial exam- ples via hierarchical diffusion models,” in Proceedings of the 30th Annual International Conference on Mobile Computing and Net- working, 2024, p. 1268–1282. [3] H. Guo, G. Wang, Y. Wang, B. Chen, Q. Yan, and L. Xiao, “Phan- tomsound: Black-box, query-efficient audio adversarial attack via split-second phoneme injection,” in Proceedings of the 26th Inter- national Symposium on Research in Attacks, Intrusions and De- fenses, 2023, p. 366–380. [4] Y. Wang, H. Guo, G. Wang, B. Chen, and Q. Yan, “Vsmask: De- fending against voice synthesis attack via real-time predictive per- turbation,” in Proceedings of the 16th ACM Conference on Se- curity and Privacy in Wireless and Mobile Networks, 2023, p. 239–250. [5] M. Todisco, X. Wang, V. Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. Evans, T. Kinnunen, and K. A. Lee, “ASVspoof 2019: Future horizons in spoofed and fake audio de- tection,” in Proc. Interspeech, 2019, p. 1008–1012. [6] J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, X. Liu, K. A. Lee, T. Kinnunen, N. Evans, and H. Delgado, “ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection,” IEEE/ACM Trans. Audio, Speech, and Language Processing, 2024. [7] X. Wang, H. Delgado, H. Tak, J.-w. Jung, H.-j. Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. Kinnunen, N. Evans, K. A. Lee, and J. Yamagishi, “ASVspoof 5: Design, collection and val- idation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speech,” Computer Speech & Lan- guage, vol. 95, 2026. [8] J. Frank and L. Sch ̈ onherr, “WaveFake: A data set to facilitate audio deepfake detection,” in Proc. NeurIPS Datasets and Bench- marks Track, 2021. [9] N. M. M ̈ uller, P. Czempin, F. Dieckmann, A. Frober, and K. B ̈ ottinger, “Does audio deepfake detection generalize?” arXiv preprint arXiv:2203.16263, 2022. [10] N. M. M ̈ uller, P. Kawa, W. H. Choong, E. Casanova, E. G ̈ olge, T. M ̈ uller, P. Syga, P. Sperl, and K. B ̈ ottinger, “MLAAD: The multi-language audio anti-spoofing dataset,” in Proc. IJCNN, 2024. [11] Z. Yan, Y. Zhao, and H. Wang, “VoiceWukong: Benchmark- ing deepfake voice detection,” arXiv preprint arXiv:2409.06348, 2024. [12] Y. Zhou, G. Zeng, X. Liu, X. Li, R. Yu, Z. Wang, R. Ye, W. Sun, J. Gui, K. Li et al., “VoxCPM: Tokenizer-free TTS for context- aware speech generation and true-to-life voice cloning,” arXiv preprint arXiv:2509.24650, 2025. [13] H. Hu, X. Zhu, T. He, D. Guo, B. Zhang, X. Wang, Z. Guo, Z. Jiang, H. Hao, Z. Guo et al., “Qwen3-TTS technical report,” arXiv preprint arXiv:2601.15621, 2026. [14] J. Cui, Z. Yang, N. Li, J. Tian, X. Ma, Y. Zhang, G. Chen, R. Yang, Y. Cheng, Y. Zhou et al., “GLM-TTS technical report,” arXiv preprint arXiv:2512.14291, 2025. [15] Z. Du, C. Gao, Y. Wang, F. Yu, T. Zhao, H. Wang, X. Lv, H. Wang, X. Shi, K. An et al., “CosyVoice 3: Towards in-the- wild speech generation via scaling-up and post-training,” arXiv preprint arXiv:2505.17589, 2025. [16] Resemble AI, “Chatterbox: Open-source flow-matching TTS with emotion exaggeration control,” https://huggingface.co/ resemble-ai/chatterbox, 2025. [17] T. Chen, T. Chen, K. Shen, Z. Bao, Z. Zhang, M. Yuan, and Y. Shi, “FlashLabs Chroma 1.0: A real-time end-to-end spoken dialogue model with personalized voice cloning,” arXiv preprint arXiv:2601.11141, 2026. [18] Z. Peng, J. Yu, W. Wang, Y. Chang, Y. Sun, L. Dong, Y. Zhu, W. Xu, H. Bao, Z. Wang et al., “VibeVoice: Expressive podcast generation with next-token diffusion,” in Proc. ICLR, 2025. [19] S. Liu et al., “Zero-shot voice conversion with diffusion trans- formers,” arXiv preprint arXiv:2411.09943, 2024. [20] MyShell AI, “OpenVoice V2,” https://huggingface.co/myshell-ai/ OpenVoiceV2, 2024. [21] RVC-Project,“Retrieval-based-voice-conversion- webui(RVC),”https://github.com/RVC-Project/ Retrieval-based-Voice-Conversion-WebUI, 2024. [22] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- riSpeech: An ASR corpus based on public domain audio books,” in Proc. ICASSP, 2015, p. 5206–5210. [23] C. Wang, M. Riviere, A. Lee, A. Wu, C. Talber, J. Bhosale et al., “VoxPopuli: A large-scale multilingual speech corpus for repre- sentation learning, semi-supervised learning and interpretation,” in Proc. ACL, 2021, p. 993–1003. [24] J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. Evans, “AASIST: A new end-to-end anti- spoofing system using integrated spectro-temporal graph attention networks,” in Proc. ICASSP, 2022, p. 6367–6371. [25] H. Tak, J.-w. Jung, J. Patino, M. Kamble, M. Todisco, and N. Evans, “End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detec- tion,” in Proc. 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, 2021, p. 1–8. [26] Y. Gong, Y.-A. Chung, and J. Glass, “AST: Audio spectrogram transformer,” in Proc. Interspeech, 2021, p. 571–575. [27] B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA- TDNN: Emphasized channel attention, propagation and aggrega- tion in TDNN based speaker verification,” in Proc. Interspeech, 2020, p. 3830–3834. [28] G. Wang, Q. Yan, S. Patrarungrong, J. Wang, and H. Zeng, “Facer: Contrastive attention based expression recognition via smartphone earpiece speaker,” in IEEE INFOCOM 2023-IEEE conference on computer communications.IEEE, 2023, p. 1– 10.