Paper deep dive
What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection
Shree Harsha Bokkahalli Satish, Harm Lameris, Joakim Gustafson, Éva Székely
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 3/22/2026, 5:04:41 AM
Summary
The paper investigates the failure of binary audio anti-spoofing systems when faced with benign transformations like voice quality conversion (VQC) and speech restoration. These transformations introduce distributional shifts that cause models to misclassify authentic speech as spoofed. The authors propose a 4-way classification framework (bona fide, converted, spoofed, and converted-spoofed) to disentangle benign processing from malicious spoofing, demonstrating improved robustness and accuracy in cross-domain evaluations.
Entities (6)
Relation Signals (3)
4-way classification framework → improves → Robustness
confidence 95% · Reformulating anti-spoofing as a multi-class problem improves robustness to benign shifts
Voice Quality Conversion → induces → Distributional Shift
confidence 90% · The benign transformations induce a drift in the SSL space
Audio Anti-spoofing Systems → misclassifies → Benign Transformations
confidence 90% · benign transformations introduce distributional shifts that are misclassified as spoofing.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Audio anti-spoofing systems are typically formulated as binary classifiers distinguishing bona fide from spoofed speech. This assumption fails under layered generative processing, where benign transformations introduce distributional shifts that are misclassified as spoofing. We show that phonation-modifying voice conversion and speech restoration are treated as out-of-distribution despite preserving speaker authenticity. Using a multi-class setup separating bona fide, converted, spoofed, and converted-spoofed speech, we analyse model behaviour through self-supervised learning (SSL) embeddings and acoustic correlates. The benign transformations induce a drift in the SSL space, compressing bona fide and spoofed speech and reducing classifier separability. Reformulating anti-spoofing as a multi-class problem improves robustness to benign shifts while preserving spoof detection, suggesting binary systems model the distribution of raw speech rather than authenticity itself.
Tags
Links
- Source: https://arxiv.org/abs/2603.14033v1
- Canonical: https://arxiv.org/abs/2603.14033v1
Trouble viewing inline? Open PDF directly →
Full Text
28,573 characters extracted from source content.
Expand or collapse full text
What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection Shree Harsha Bokkahalli Satish ID , Harm Lameris ID , Joakim Gustafson ID , ́ Eva Sz ́ ekely ID Department of Speech, Music and Hearing, KTH Royal Institute of Technology, Sweden shbs, lameris, jkgu, szekely@kth.se Abstract Audio anti-spoofing systems are typically formulated as binary classifiers distinguishing bona fide from spoofed speech. This assumption fails under layered generative processing, where benign transformations introduce distributional shifts that are misclassified as spoofing. We show that phonation-modifying voice conversion and speech restoration are treated as out-of- distribution despite preserving speaker authenticity. Using a multi-class setup separating bona fide, converted, spoofed, and converted-spoofed speech, we analyse model behaviour through self-supervised learning (SSL) embeddings and acoustic cor- relates. The benign transformations induce a drift in the SSL space, compressing bona fide and spoofed speech and reducing classifier separability. Reformulating anti-spoofing as a multi- class problem improves robustness to benign shifts while pre- serving spoof detection, suggesting binary systems model the distribution of raw speech rather than authenticity itself. Index Terms: deepfake, human-computer interaction, compu- tational paralinguistics 1. Introduction The proliferation of highly realistic synthesised speech has ne- cessitated robust countermeasures, and yet the detection land- scape remains an ongoing adversarial cycle [1, 2, 3]. Attackers increasingly employ post-processing transformations, such as replay attacks, resulting in signals that are significantly harder to detect [4]. In response, modern audio spoofing evaluations have decisively shifted toward real-world variability, as seen in recent challenges [5] and comprehensive benchmarking plat- forms like DF Arena [6]. As the field moves beyond simple binary detection towards the attribution of specific sources of spoofed audio [7, 8], it often relies on an implicit assumption: that “authentic” audio remains a single, pristine distribution. High-fidelity media production relies heavily on signal process- ing chains, including modern speech enhancement and restora- tion [9]. Moreover, stylistic modifications like voice quality conversion to a creaky or breathy phonation – while some- times used adversarially to defeat acoustic fingerprinting [10] – can also be legitimately used to enhance paralinguistic ex- pression [11, 12]. In this work we demonstrate how current spoof detec- tion systems struggle under this layered generative processing, as benign distributional shifts are frequently misclassified as spoofing. Maintaining a rigid framework where any processed signal might be flagged as spoofed increases false-positive risk in practical deployments. In anti-spoofing deployments, the objective is typically to detect malicious impersonation rather 0 Under review at Interspeech 2026 60402002040 t-SNE Dimension 1 40 20 0 20 40 t-SNE Dimension 2 Bona fide Bona fide (Processed) Spoofed Spoofed (Processed) Figure 1: t-SNE plots of Wav2Vec2 embeddings before and after VQC on MLAAD matched dataset. than catch quality-of-life improvement artefacts added by the processing [13, 14]. To address this, we propose the term be- nign transformations to include intra-speaker variations that ro- bust spoof detectors should ignore, e.g. enhance audio quality, accessibility, or stylistic expressivity. We hypothesise that bi- nary spoof detectors conflate authenticity and benign transfor- mations and make the following contributions: • We show deficiencies with the binary framework of classi- fying recordings into spoofed and bona fide and introduce a 4-way classification framework to address it. • We analyse SSL representations and acoustic correlates un- der two benign transforms and examine drifts for bona fide and spoofed speech. • We release our dataset of transformed bona fide and spoofed audio to benchmark deepfake detectors against benign pro- cessing and also release additional details on our website. 2. Dataset To evaluate how benign transformations affect spoofing de- tection, we build upon and create a dataset that pairs bona fide audio with spoofed counterparts along with their benign- transformed versions. 2.1. Corpora and Synthetic speech (TTS) sources We used the paired utterances from the real M-AILABS cor- pus [15] and the deepfakes from the MLAAD corpus [16]. M-AILABS contains English audiobook recordings from Lib- arXiv:2603.14033v1 [cs.SD] 14 Mar 2026 riVox. The corpus includes both male and female speakers recorded in quiet conditions. MLAAD [16] contains synthetic data that is synthesised using M-AILABS as its bona fide data source. We selected 2,575 utterances that have matching TTS counterparts across 10 TTS systems, ensuring balanced repre- sentation across speakers and content. All audio was resampled to 16 kHz for consistency with SSL model requirements. We use 10 diverse TTS architectures: FireRedTTS-2.0 [17], Higgs-Audio-V2 [18], Index-TTS-2.0 [19], Llasa-1B [20], MiniCPM-o-2.6 [21], Openaudio-S1-Mini [22], OuteTTS [23], VoxCPM-0.5B [24], VoXtream [25], and ZipVoice [26]. These architectures were selected because they were annotated with reference speaker information. This allowed us to select only the utterances for which the reference speaker was identical to the source speaker from the original M-AILABS corpus. Each TTS system synthesises the same text content as the bona fide recordings, yielding 2,575 matched utterance pairs. This ensures that any observed differences between bona fide and spoofed embeddings are attributable to source characteristics rather than linguistic content or speaker conversion. 2.2. Benign Speech Transformations We evaluate our hypotheses using two benign transformations: (1) Voice Quality Conversion (VQC) and (2) Speech Restora- tion. Unlike generative attacks designed for impersonation, these transformations represent intra-speaker variations that a robust spoofing countermeasure should ignore. The models employed here report high speaker similarity metrics pre and post-transformation [9, 27]. In the case of VQC, the glottal source parameters are modified while preserving the semantic content and minimizing speaker identity changes. This allows for enhanced paralinguistic and pragmatic expression, for in- stance, using breathy voice to signal intimacy [11] or creaky voice to indicate a turn yield [28]. From the available VQC frameworks [29, 30], we chose [31] due to its support for mul- tiple phonation types. We converted 2,575 utterance pairs into four categories: modal, breathy, creaky, and end-creak. These types were selected because they account for the most com- mon phonation types in English [32] and serve diverse prag- matic functions, such as marking parenthetical comments [33] or signalling utterance termination [34]. An additional cross- transformation test is speech restoration. Speech restoration 3020100102030 t-SNE Dimension 1 40 20 0 20 40 t-SNE Dimension 2 Bona fide Bona fide (Processed) Spoofed Spoofed (Processed) Figure 2: t-SNE plots of Wav2Vec2 embeddings before and after Sidon enhancement on MLAAD matched dataset. models, such as Sidon [9] generate restored speech using repre- sentations from speech foundation models. Across both trans- formed scenarios, we then have the following subsets: • Bona fide: Unprocessed human speech recordings. • Spoofed: Unprocessed synthetic speech generated via TTS. • Processed Bona fide: Bona fide speech after benign trans- formations (voice quality conversion or speech restoration). • Processed Spoofed: Spoofed speech after the same benign transformations. For cross-domain evaluation on the ASVspoof5 dataset, we re- port results on a VQC class-balanced subset (2,000 utterances per class; 8,000 total) sampled without replacement using a fixed seed. For mixed-domain fine-tuning, we construct a dis- joint ASVspoof5 train/val/test partition by first sampling an equal number of utterances per class and then reserving a held- out test set of 2,000 per class; the remaining samples are split into train/validation (approximately 70/15). The ASVspoof5 test partition is never used for training, early stopping, or model selection. We used Sidon to restore the same 2,575 ut- terance pairs, resulting in bona fide enhanced/restored audio and spoofed enhanced/restored audio. Our processed datasets: https://zenodo.org/records/18803182 3. Methodology and Experiments We first conduct an analysis in SSL space using t-SNE plots, which guide model fine-tuning and classification experiments. Then, we present the results and conduct an acoustic analysis involving glottal source parameters and spectral tilt which offers interpretable explanations to the classification results. 3.1. Embedding analysis We use mean-pooled SSL embeddings from HuBERT- base [35], Whisper-small [36], and Wav2Vec2 XLS-R 1B [37] to analyse how our chosen transformations alter the represen- tation space used for spoof detection. To test whether VQC induces similar shifts regardless of source type, we compute di- rectional consistency between bona fide and spoofed shift vec- tors. For each source type s ∈ bona fide, spoofed, we com- pute the mean embedding shift vector: ∆ s = 1 N s X x∈s (Emb(VQC(x))− Emb(x)) .(1) Directional consistency is then measured as the cosine similar- ity between these mean shift vectors: cos(∆ bonafide , ∆ spoofed ) = ∆ bonafide · ∆ spoofed ∥∆ bonafide ∥∆ spoofed ∥ (2) A value near +1 indicates that VQC pushes bona fide and TTS embeddings in the same direction; a value near 0 indicates orthogonal shifts; a negative value indicates opposing direc- tions. We find that the Whisper features have variable direc- tional consistency under VQC as seen in Figure 3. HuBERT and Wav2Vec2 show high consistency across all VQC condi- tions, indicating VQC applies a similar transformation to both source types. However, cosine similarity does not account for the magnitude of shifts and the source/speaker-dependent shift directions which may average out when aggregated. We also plot t-SNE embeddings of the Wav2Vec2 features, used in the DF-Arena classifiers, before and after the benign transforma- tions in Figure 1 and Figure 2. Other features and additional plots can be found on our project website. ModalBreathyCreakyEnd Creak VC Condition 0.2 0.0 0.2 0.4 0.6 0.8 1.0 Cosine Similarity cos( bonafide , spoof ) Model HuBERT Whisper Wav2Vec2 Figure 3: Directional consistency of VQC-induced embed- ding shifts between mean bona fide and spoofed shift vectors. Whisper shows lower or negative values, suggesting potential source-dependent shift directions. 3.2. Experiments We compare two classification architectures:An MLP trained on concatenated mean-pooled SSL embeddings from Wav2Vec2, HuBERT, and Whisper (2816 Dimensions) in binary and 4-way configurations and the pretrained DF- Arena 1B [6], an open-source state-of-the-art binary anti- spoofing model built on a Wav2Vec2 XLS-R 1B backbone with a conformer head. Fine-tuning DF-Arena for multi-class detection: We replace DF-Arena’s binary head (fc5: 1280→2) with a 4-class head (1280→4), initialising classes 0 (bona fide) and 2 (spoofed) from the pretrained weights and classes 1 and 3 (processed vari- ants of bona fide and spoofed) from the spoof weights since the pretrained model already maps all processed audio near its spoof representation. The Wav2Vec2 backbone is frozen; only the final Conformer block and the new classification head are fine-tuned (AdamW , lr = 10 −4 , 20 epochs). A binary vari- ant uses the same protocol with a 2-class head. We select the checkpoint with the lowest validation loss and report results. Models fine-tuned solely on MLAAD achieved ∼99% in-domain accuracy but failed catastrophically cross-domain: bona fide accuracy on ASVspoof5 drops to 0.1%, as unseen bona fide speech collapses into the spoofed class.To re- cover cross-domain source detection, we continue from the MLAAD–only checkpoint with mixed-domain training, com- bining MLAAD and balanced ASVspoof5 data at a reduced learning rate (5×10 −5 ) with early stopping (patience = 8). Since speech restoration induces a distributional shift distinct from VQC, we further augment the training set with Sidon- enhanced utterances, yielding the final models reported in Sec- tion 4. Because accuracy depends on the decision threshold (argmax for 4-way; fixed threshold for binary), we interpret performance primarily through threshold-free metrics (EER src and, for 4-way models, EER proc ). 4. Results Table 1 summarises in-domain evaluation on MLAAD VQC (Seen/Unseen TTS) and out-of-domain evaluation on ASVspoof5, Sidon restoration. For 4-way models, we addition- ally report EER proc , which collapses predictions along the pro- cessed axis (unprocessed vs. processed) to isolate what gener- OriginalModalBreathyCreakyEnd Creak 10 0 10 20 30 40 dB H1-A3 Source Bonafide Spoofed OriginalModalBreathyCreakyEnd Creak 10 5 0 5 10 15 20 25 dB H1-H2 Figure 4: The acoustic feature shifts between the original M- AILABS/MLAAD recordings and the converted recordings by voice quality. alises across domains. We make the following observations re- garding model generalisation and the modelling of benign trans- formations: • Across architectures, we observe a consistent failure mode: models trained (or fine-tuned) to detect spoofing under a bi- nary objective treat benign processing as evidence of spoof- ing. This is visible both in the pretrained DF-Arena base- line and in the MLAAD-only fine-tuned models, which col- lapse to near-chance source performance on ASVspoof5 and sharply reduced bona fide accuracy. • However, when predictions are analysed along the pro- cessed axis, generalisation is strong: 4-way models achieve EER proc <0.2% on ASVspoof5, indicating that “processed vs. unprocessed” is a transferable cue even when “bona fide vs. spoofed” is not. Table 2 contains the confusion matrix of the mixed-domain DF-Arena 4-way model on the held- out ASVspoof5 test set. Source attribution within processed speech (Bona fide→Processed vs. Spoofed→Processed) re- mains difficult. • Mixed-domain fine-tuning resolves the domain shift on the source axis. Using the same backbone and training expo- sure (MLAAD VQC + ASVspoof5), 4-way supervision im- proves ASVspoof5 accuracy (86.8%) and bona fide accuracy (94.7%) compared to binary supervision (83.7% / 73.4%), showing that the gain is not only from additional data but from disentangling benign processing and source authen- ticity in the objective. • Finally, Sidon restoration induces a distinct shift from VQC: the mixed-domain 4-way model remains strong on ASVspoof5 but fails on Sidon–restored bona fide speech (9.2% bona fide accuracy). Adding Sidon augmented data recovers robustness (81.8%) with minimal in-domain degra- dation, suggesting that new benign transformations can be incorporated explicitly within the same 4-way framework. 4.1. Acoustic Analysis In order to investigate the glottal source parameters and the acoustic shifts by the transformations, we measured H1–A3, a spectral measure related to the abruptness of closure of the vo- cal folds and H1–H2, a spectral measure related to the open quotient, i.e. the fraction of time that the vocal folds allow the passage of air on the bona fide (M-AILABS) and spoofed (MLAAD) recordings, as well as the converted and enhanced versions. Results can be found in Figure 4. The measurements were analysed with a two-way ANOVA with Tukey HSD for each spectral measure with the main effects of source (bona fide or spoofed) and processing type (target voice quality), as Table 1: Summary of all experiments. MLAAD metrics use the full in-domain test sets (Seen = 9 TTS architectures; Unseen = held-out OuteTTS). ASVspoof5 is out-of-domain. Sidon columns evaluate all audio after passing through Sidon restoration [9]. For 4-way models, EER proc collapses predictions along the processed axis (unprocessed vs. processed). Acc bona is the accuracy on genuine bona fide speech. † Trained on a single voice quality condition; tested on the full MLAAD VQC test set. ∗ Trained on Sidon-enhanced data (train split only); Sidon columns use held-out test-split utterances. MLAAD VQCASVspoof5 VQCMLAAD Sidon AccEER src ModelTraining/Fine-tuning data# ClassesSeenUnseenAcc bona SeenUnseenAccEER src EER proc Acc bona AccEER src Acc bona SSL-embedding MLP; (HuBERT⊕ Whisper⊕ Wav2Vec2) BinaryMLAAD VQC299.190.2100.00.814.1848.054.7–6.288.65.6278.8 4-WayMLAAD VQC499.089.699.90.884.3843.552.32.424.076.710.0155.0 Unconverted binary † MLAAD (unconverted)275.275.3100.020.7528.0224.247.6–7.169.713.1939.7 DF-Arena 1B; Wav2Vec2 Baseline DF-Arena-1Bpretrained256.363.6100.051.0739.6361.544.2–69.956.727.8413.0 Binary fine-tunedMLAAD VQC299.097.199.51.172.5349.864.2–0.055.818.1811.2 4-Way fine-tunedMLAAD VQC498.998.1100.01.061.9549.446.30.150.155.819.4510.8 Binary fine-tunedMLAAD VQC + ASVspoof5298.497.599.01.362.3383.713.30–73.454.843.318.7 4-Way fine-tunedMLAAD VQC + ASVspoof5498.398.2100.01.361.95 86.811.80.0394.7 55.040.629.2 Binary fine-tuned ∗ MLAAD VQC + ASVspoof5 + Sidon298.697.799.51.432.1483.7 14.28–73.589.92.3879.6 4-Way fine-tuned ∗ MLAAD VQC + ASVspoof5 + Sidon498.096.898.81.832.3387.211.20.0580.590.32.3881.8 well as the interaction. For both H1–A3 and H1–H2, there was a strongly significant main effect of source (p < 0.0001) as well as processing type (p < 0.0001). Additionally, a highly significant interaction effect was observed for both mea- sures (p < 0.0001), demonstrating that VQC non-uniformly increases the acoustic differences. For the original bona fide and spoofed recordings, the sources exhibited no statistically significant differences in H1–A3 (p = 0.7403) or H1–H2 (p = 0.0548). Nevertheless, the VQC conversion widens the gap between spoofed and real and reaches its maximum in the end creak condition (interaction deltas of−0.99 dB for H1–A3 and−0.61 dB for H1–H2, p < 0.0001). The acoustic analysis of H1–A3 and H1–H2 for Sidon restored audio yielded signif- icant main effects for source authenticity and processing type, but no interaction effect, indicating that the enhancement pro- cess induced a consistent global shift for both bona fide and spoofed samples. These findings confirm that while the origi- nal recordings are acoustically similar, VQC exaggerates latent synthetic artifacts. While Sidon restoration does not interact Table 2: Confusion matrix of mixed-domain DF-Arena 4-way on ASVspoof5. B = Bona fide, S = Spoofed, S→P = Processed Spoofed, and B→P = Processed Bona fide True Pred. B→PSS→P B189401060 B→P013170683 S38019602 S→P022901771 Table 3: The acoustic differences of the voice quality transfor- mations between spoofed and bona fide audio Spoofed – Bona fideH1–A3H1–H2 MLAAD Unconverted Gap+0.36 dB −0.38 dB p = .7403 p = .0548 Modal Int. (∆)−0.02 dB −0.01 dB Breathy Int. (∆)−0.54 dB −0.36 dB Creaky Int. (∆)−0.81 dB −0.33 dB End Creak Int. (∆)−0.99 dB −0.61 dB Interaction (p-value)p < .0001 p < .0001 with voice quality features, our analysis offers interpretable ex- planations for why binary classifiers struggle with distinguish- ing out-of-domain processing and offers discriminatory features for improved detection of transformation source. 4.2. Discussion Although VQC offers the potential for enhanced expressiv- ity, the speaker’s intent can be obscured by common percep- tions of creak as having authoritative intent [38] and breathi- ness as having affiliative intent [11]. Similarly, enhancement techniques or replays may remove the background characteris- tics that are proxy signals for authenticity. Equating these pro- cessed speech with spoofing misrepresents the purpose of these transformations. In our experiments, we find that 4-way super- vision improves robustness over binary supervision in mixed- domain DF-Arena fine-tuning against out-of-domain test data. We expect this gap to widen as speech enhancement, VQC and other benign transformations improve with time. While multi- class deepfake detection of spoof sources is gaining popularity, our work addresses the complementary problem. Ultimately, while severe performance drops on unseen datasets are a well- documented challenge in-domain generalisation [39], spoof de- tection failures in the presence of benign transformations should not be viewed in just that light, but as a consequence of mod- elling authenticity and processed speech as a single variable. 5. Conclusion In this paper, we show that current binary anti-spoofing sys- tems conflate benign transformations, such as voice quality con- version and speech restoration, with spoofed generations. Our analysis reveals that benign authenticity-preserving transforma- tions also induce a collapse in baseline classifiers, causing pro- cessed bona fide speech to be systematically misclassified. To resolve this vulnerability, we introduced a 4-way classification framework that explicitly disentangles source from benign pro- cessing and open source a dataset of processed spoofed and bona fide recordings. By modelling processed speech authentic- ity as a distinct class, our approach better accommodates benign audio enhancements in the deepfake detection domain, provid- ing more reliable speech authentication in modern, processing- heavy media environments. 6. Generative AI Use Disclosure AI tools were used to assist with portions of coding the plots, website interface, and polishing text which were all reviewed, verified and modified by the authors. 7. References [1] K. Kamel, K. Sood, H. S. Dutta, and S. Aryal, “A survey of threats against voice authentication and anti-spoofing systems,” arXiv preprint arXiv:2508.16843, 2025. [2] M. Li, Y. Ahmadiadli, and X.-P. Zhang, “Audio anti-spoofing de- tection: A survey,” arXiv e-prints, p. arXiv–2404, 2024. [3] B. Zhang, H. Cui, V. Nguyen, and M. Whitty, “Audio deepfake detection: What has been achieved and what lies ahead,” Sensors, vol. 25, no. 7, p. 1989, 2025. [4] N. M ̈ uller, P. Kawa, W.-H. Choong, and et al., “Re- play attacks against audio deepfake detection,” arXiv preprint arXiv:2505.14862, 2025. [5] X. Wang, H. Delgado, H. Tak, J.-w. Jung, and et al., “ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale,” in Proc. of The ASVspoof Workshop, 2024. [6] S. Dowerah, A. Kulkarni, A. Kulkarni, H. M. Tran, J. Kalda, A. Fedorchenko, B. Fauve, D. Lolive, T. Alum ̈ ae, and M. M. Doss, “Speech DF arena: A leaderboard for speech deepfake detection models,” IEEE Open Journal of Signal Processing, 2026. [7] A. Stan, D. Combei, D. Oneata, and H. Cucu, “TADA: Training- free Attribution and Out-of-Domain Detection of Audio Deep- fakes,” in Proc. Interspeech 2025, 2025, p. 1543–1547. [8] J. Mishra, M. Chhibber, H.-j. Shim, and T. H. Kinnunen, “Towards explainable spoofed speech attribution and detection: A prob- abilistic approach for characterizing speech synthesizer compo- nents,” Computer Speech & Language, vol. 95, p. 101840, 2026. [9] W. Nakata, Y. Saito, Y. Ueda, and H. Saruwatari, “Sidon: Fast and robust open-source multilingual speech restoration for large-scale dataset cleansing,” arXiv preprint arXiv:2509.17052, 2025. [10] Y. ̈ Ozer, W. Ge, Z. Zhang, X. Wang, and J. Yamagishi, “Self Voice Conversion as an Attack against Neural Audio Watermarking,” arXiv preprint arXiv:2601.20432, 2026. [11] L. Tsvetanova, V. Auberg ́ e, and Y. Sasa, “Multimodal breathiness in interaction: from breathy voice quality to global breathy ”body behavior quality”,” in 1st International Workshop on Vocal Inter- activity in Humans, Animals and Robots (VIHAR), 2017. [12] L. Zimman, “Transgender voices: Insights on identity, embod- iment, and the gender of the voice,” Language and Linguistics Compass, vol. 11, no. 9, p. e12244, 2017. [13] Y. Zuo, M. Ge, J. Du, N. Jiang, Y. Hu, and H. Li, “CodecFake: Enhancing Anti-spoofing Models Against Deepfake Audios from Codec-based Speech Synthesis Systems,” in Proc. Interspeech 2024, 2024, p. 3155–3159. [14] J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, N. Evans, F. Delgado, N. Evans, and T. Kinnunen, “ASVspoof 2021: Towards Spoofed and Deepfake Speech Detec- tion in the Wild,” in Proc. ASVspoof 2021 Workshop, 2021, p. 1–10. [15] M-AILABS, “The M-AILABS Speech Dataset,” 2017, open- source multi-language speech database. [Online]. Available: https://w.caito.de/2019/01/the-m-ailabs-speech-dataset/ [16] N. M. M ̈ uller, P. Kawa, W. H. Choong, E. Casanova, E. G ̈ olge, T. M ̈ uller, P. Syga, P. Sperl, and K. B ̈ ottinger, “MLAAD: The Multi-Language Audio Anti-Spoofing Dataset,” arXiv preprint arXiv:2401.09512, 2024. [17] T. Zhou et al., “Fireredtts-2:Towards long conversational speech generation for podcast and chatbot,” arXiv preprint arXiv:2509.02020, 2025. [18] J. Liu et al., “Higgs-audio-v2: Redefining expressiveness in audio generation,” Boson AI Technical Report, 2025. [Online]. Available: https://github.com/boson-ai/higgs-audio [19] S. Zhou, Y. Zhou, Y. He, and et al., “IndexTTS2: A Break- through in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech,” arXiv preprint arXiv:2506.21619, 2025. [20] Y. Zuo et al., “Llasa: Training large language models for long- form speech synthesis,” arXiv preprint arXiv:2411.05963, 2024. [21] O. Team, “Minicpm-o 2.6:A comprehensive omni-modal foundation model,” GitHub Repository, 2025. [Online]. Available: https://github.com/OpenBMB/MiniCPM-o [22] F. A. Team, “Openaudio-s1: Advanced text-to-speech model se- ries,” https://openaudio.com/blogs/s1, 2025. [23] OuteAI, “Outetts-0.1: Purely language modeling based text-to- speech,” https://github.com/OuteAI/OuteTTS, 2024. [24] Y. Zhou et al., “VoxCPM: Tokenizer-Free TTS for Context- Aware Speech Generation and True-to-Life Voice Cloning,” arXiv preprint arXiv:2509.24650, 2025. [25] N. Torgashov, G. E. Henter, and G. Skantze, “VoXtream: Full- Stream Text-to-Speech with Extremely Low Latency,” arXiv preprint arXiv:2509.15969, 2025. [26] H. Zhu, W. Kang, Z. Yao, L. Guo, F. Kuang, Z. Li, W. Zhuang, L. Lin, and D. Povey, “ZipVoice:Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching,” arXiv preprint arXiv:2506.13053, 2025, accepted at ASRU 2025. [27] J. Li, W. Tu, and L. Xiao, “FreeVC: Towards high-quality text- free one-shot voice conversion,” in ICASSP. IEEE, 2023. [28] M. Heldner, M. Wlodarczak, ˇ S. Be ˇ nu ˇ s, and A. Gravano, “Voice quality as a turn-taking cue,” in Proc. Interspeech, 2019, p. 4165–4169. [29] F. Rautenberg, M. Kuhlmann, F. Seebauer, J. Wiechmann, P. Wag- ner, and R. Haeb-Umbach, “Speech synthesis along perceptual voice quality dimensions,” in ICASSP, 2025. [30] H. Lameris, J. Gustafson, and ́ E. Sz ́ ekely, “CreakVC: a voice con- version tool for modulating creaky voice.” in Proc. Interspeech, 2024. [31] H. Lameris, J. Gustafsson, and ́ E. Sz ́ ekely, “VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in Speech,” in Proc. Interspeech 2025, 2025, p. 2295–2299. [32] R. J. Podesva, “Gender and the social meaning of non-modal phonation types,” in Annual meeting of the Berkeley linguistics society, 2011, p. 427–448. [33] S. Lee, “Creaky voice as a phonational device marking parenthet- ical segments in talk,” Journal of Sociolinguistics, vol. 19, no. 3, p. 275–302, 2015. [34] C. G. Henton, “Sociophonetic aspects of creaky voice,” The Jour- nal of the Acoustical Society of America, vol. 86, no. S1, p. S26– S26, 1989. [35] W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdi- nov, and A. Mohamed, “Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,” IEEE/ACM TASLP, vol. 29, p. 3451–3460, 2021. [36] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak su- pervision,” in 40th ICML, vol. 202.PMLR, 2023, p. 28 448– 28 466. [37] A. Babu, C. Wang, A. Tjandra, and et al., “XLS-R: Self- supervised cross-lingual speech representation learning at scale,” in Proc. Interspeech, 2022, p. 2278–2282. [38] J. Laver, “The phonetic description of voice quality,” Cambridge Studies in Linguistics London, vol. 31, p. 1–186, 1980. [39] I. Gulrajani and D. Lopez-Paz, “In search of lost domain gener- alization,” in International Conference on Learning Representa- tions (ICLR), 2021.