Paper deep dive
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
Qinke Ni, Huan Liao, Dekun Chen, Yuxiang Wang, Zhizheng Wu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/22/2026, 5:20:43 AM
Summary
NV-Bench is a comprehensive, multi-lingual benchmark designed to evaluate nonverbal vocalization (NV) synthesis in text-to-speech (TTS) systems. It addresses the lack of standardized metrics by providing 1,651 in-the-wild utterances paired with human ground-truth audio, categorized into 14 NV types. The framework introduces a dual-dimensional evaluation protocol: Instruction Alignment (using Paralinguistic Character Error Rate, PCER) and Acoustic Fidelity (using FAD, FD, and speaker similarity), demonstrating strong correlation with human perception.
Entities (5)
Relation Signals (3)
NV-Bench → evaluates → TTS Systems
confidence 95% · NV-Bench, the first comprehensive evaluation framework for NV-capable TTS
NVASR → computes → PCER
confidence 90% · we utilize our NVASR (Section 2.1) to compute CER, OCER and PCER.
NV-CV3 → performson → NV-Bench
confidence 90% · NV-CV3 achieves the lowest PCER (27.69%) and OCER (4.90%) on the Mandarin single-label subset
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While recent text-to-speech (TTS) systems increasingly integrate nonverbal vocalizations (NVs), their evaluations lack standardized metrics and reliable ground-truth references. To bridge this gap, we propose NV-Bench, the first benchmark grounded in a functional taxonomy that treats NVs as communicative acts rather than acoustic artifacts. NV-Bench comprises 1,651 multi-lingual, in-the-wild utterances with paired human reference audio, balanced across 14 NV categories. We introduce a dual-dimensional evaluation protocol: (1) Instruction Alignment, utilizing the proposed paralinguistic character error rate (PCER) to assess controllability, (2) Acoustic Fidelity, measuring the distributional gap to real recordings to assess acoustic realism. We evaluate diverse TTS models and develop two baselines. Experimental results demonstrate a strong correlation between our objective metrics and human perception, establishing NV-Bench as a standardized evaluation framework.
Tags
Links
- Source: https://arxiv.org/abs/2603.15352v2
- Canonical: https://arxiv.org/abs/2603.15352v2
Trouble viewing inline? Open PDF directly →
Full Text
25,473 characters extracted from source content.
Expand or collapse full text
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation Qinke Ni, Huan Liao, Dekun Chen, Yuxiang Wang, Zhizheng Wu The Chinese University of Hong Kong, Shenzhen qinkeni@link.cuhk.edu.cn, huanliao@link.cuhk.edu.cn, dekunchen@link.cuhk.edu.cn, yuxiangwang1@link.cuhk.edu.cn, wuzhizheng@cuhk.edu.cn Abstract While recent text-to-speech (TTS) systems increasingly inte- grate nonverbal vocalizations (NVs), their evaluations lack stan- dardized metrics and reliable ground-truth references. To bridge this gap, we propose NV-Bench, the first benchmark grounded in a functional taxonomy that treats NVs as communicative acts rather than acoustic artifacts. NV-Bench comprises 1,651 multi-lingual, in-the-wild utterances with paired human refer- ence audio, balanced across 14 NV categories. We introduce a dual-dimensional evaluation protocol: (1) Instruction Align- ment, utilizing the proposed paralinguistic character error rate (PCER) to assess controllability, (2) Acoustic Fidelity, measur- ing the distributional gap to real recordings to assess acoustic realism. We evaluate diverse TTS models and develop two base- lines. Experimental results demonstrate a strong correlation be- tween our objective metrics and human perception, establishing NV-Bench as a standardized evaluation framework. Index Terms: Speech benchmark, Nonverbal vocalizations, Paralinguistic-aware ASR, Controllable TTS 1. Introduction Recent expressive text-to-speech (TTS) models [1, 2, 3] in- creasingly integrate nonverbal vocalizations (NVs) to enhance human-like communication. Current methods incorporate NVs either as discrete tokens (e.g., NVTTS [4]) or overlapping lay- ers (e.g., CapSpeech [5]). However, these approaches predomi- nantly treat NVs as generic ”sound effects” annexed to linguis- tic content. This mechanistic perspective overlooks their fun- damental nature: NVs are not merely acoustic textures, but in- trinsically communicative acts conveying physiological states, emotions, and interactional intent. Advancing this field requires moving beyond checking for acoustic presence to evaluating pragmatic appropriateness in context. To systematically model these phenomena, we adopt the functional taxonomy of Batliner et al. [6], categorizing NVs into three levels crucial for expressive TTS: (1) Vegetative sounds encompass biological reflexes such as breathing and coughing, which ground generated speech in physical realism. (2) Affect bursts consist of valenced vocalizations that succinctly con- vey the speaker’s emotional state or instantaneous reactions. (3) Conversational grunts function as interaction-management cues, including filled pauses and prosodic particles that disam- biguate communicative intent (e.g., confirmation or hesitation). This taxonomy reveals NVs are not merely acoustic events, they encompass a spectrum of non-lexical, pragmatically charged in- terjections essential for conveying emotion and discourse man- agement. Motivated by the need for expressive systems, several large- scale corpora with NVs have been proposed (Table 1), including Table 1: Comparison with nonverbal vocalizations datasets. DatasetLang.TestsetBalancePrompt SynParaSpeech [11] zh✗– NVS [10]zh/en✗– Emilia-NV [7]zh✗– NVTTS [4]en✓✗ SMIIP-NV [8]zh✓✗ NV-Benchzh/en✓ Emilia-NV [7], SMIIP-NV [8], NVTTS [4], DisfluencySpeech [9], NonverbalSpeech (NVS) [10], SynParaSpeech [11]. But scaling training data does not yield a reliable evaluation stan- dard. Beyond ambiguous task definitions, current evaluation practices rely on internal testsets or text-rewritten references rather than authentically paired human speech data. Without ground-truth (GT) NV recordings, evaluation tends to collapse to coarse checks (e.g., event presence/absence), making it im- possible to quantify the gap to real recordings. Furthermore, given that NV events are long-tailed, imbalanced testsets can bias aggregated metrics and hinder fair diagnosis. To address these gaps, we introduce NV-Bench 1 , a compre- hensive benchmark for NV-capable TTS. NV-Bench provides a public multi-lingual testset comprising 1,651 utterances which are curated from online audiovisual media published in 2025 to minimize the potential data leakage. To ensure fair comparison, we partition the benchmark into a strictly balanced single-label subset (50 utterances per category) and a relatively balanced multi-label subset with 14 NV types. All test samples contain in-the-wild NVs paired with GT audio. This paired design en- ables the reproducible assessment of instruction alignment via character error rate (CER) and its variants and acoustic fidelity via speaker similarity and fr ́ echet distance (FD). Our main contributions are threefold: (1) We propose NV- Bench, the first comprehensive evaluation framework for NV- capable TTS, featuring a public, in-the-wild dataset with paired human GT audio. (2) We establish a standardized and distri- butionally balanced evaluation protocol that supports fair and reproducible model comparisons. (3) We conduct extensive benchmarking of state-of-the-art (SOTA) TTS models, reveal- ing their controllability, intelligibility and acoustic fidelity. 2. Methods To construct NV-Bench, we introduce a two-phase pipeline bal- ancing acoustic diversity and label accuracy: (1) developing 1 Demo page: https://nvbench.github.io arXiv:2603.15352v2 [cs.SD] 18 Mar 2026 Data Processing Raw Data Internet data Emilia-Pipeline Standardize & Filter Mimo-Audio filter multi-speakers Clean Clips Clips duration < 30 s Multi-lingual NVASR Label Norm. Normalization & Taxonomy Open Data Open-source data Model training Multi-lingual NVASR Transcribe Rich Text Text with NVV labels Evaluation single-label ZH/EN I want to [Cough] ... multi-label ZH/EN [Laughter]我好开心[Uhm]... I want to [Cough] ... I want to [Laughter] ... D1: Instruction Alignment I want to [Cough] ... I want to [Laughter] ... D2: Acoustic Fidelity Human Judge Target text: Target text: Ground-truth TTS Synthesis Fréchet Distance Human Verification Speaker SIM DNSMOS NVASR Judge Align Figure 1: Overview of the NV-Bench. (1) Data Processing: Raw audio is filtered using the Emilia-Pipeline and MiMo-Audio. (2) Multi-lingual NVASR: We train a multi-lingual NVASR model on open-source data with a unified label taxonomy. (3) Evaluation: After human verification, the benchmark is evaluated in instruction alignment and acoustic fidelity dimensions and human subjective ratings. a robust multi-lingual NVASR to provide high-quality, event- aware transcriptions, (2) curating in-the-wild audio through a rigorous filter that culminates in human-verified GT. 2.1. Multi-lingual NVASR To facilitate efficient benchmark construction, we develop a robust multi-lingual NV-capable automatic speech recognition (NVASR) model. Following the methodology introduced in [7], we finetune the SenseVoice-Small [12] architecture and extend the framework to support multilingual settings. 2.1.1. Model architecture We select SenseVoice-Small [12] as our base model due to its pre-training on diverse audio understanding tasks, which allows it to capture rich features. The model is optimized by minimiz- ing the connectionist temporal classification (CTC) loss [13]: L CTC =− ln X π∈B −1 (y) P(π |x)(1) wherex represents the input acoustic features,y is the target sequence,B −1 (y) denotes all valid CTC alignment paths fory. 2.1.2. Data construction and label normalization To ensure broad generalization, we consolidate a comprehen- sive training corpus: Emilia-NV [7], NVTTS [4], Disfluen- cySpeech [9], NVS [10], SMIIP-NV [8], and MNV-17 [14]. To address the heterogeneity of source labels, we imple- ment a systematic normalization process. We map non-speech labels to Level 3 of the AudioSet Ontology [15], preserving dis- tinct events like “[Laughter]”. We adopt the Emilia-NV taxon- omy but conduct targeted manual annotations on the English subsets of NVTTS and DisfluencySpeech. This step is critical to distinguish nuanced pragmatic functions, such as “[Question- huh]” that were previously absent in English datasets. The final unified taxonomy is detailed in Table 2. 2.2. Benchmark dataset construction Constructing a reliable NV benchmark requires balancing in- the-wild realism with label distribution. We curate our dataset entirely from diverse web-sourced audio using a sophisticated filtering pipeline, as shown in Figure 1. 2.2.1. Data collection NV events naturally exhibit a long-tail distribution. To ensure adequate coverage across our taxonomy, we crawled a massive collection of audiovisual media uploaded within the most recent year. From a raw pool of 565,316 audio clips (≈ 1,560 hours), we filtered for candidate segments containing target NVs. 2.2.2. Filtering pipeline To guarantee acoustic fidelity and speaker purity, we subject the candidate segments to a rigorous refinement process: Audio Standardization.We first utilize the Emilia- Pipeline [16] for initial audio standardization and source sep- aration. While this includes speaker diarization, multi-speaker segments often persist. Single-Speaker Verification. To eliminate residual multi- speaker artifacts, we deploy MiMo-Audio-7B-Instruct [17], a SOTA audio large language model (LLM), prompting it to detect subtle speaker overlaps that escaped initial diarization. Only confirmed clean, single-speaker utterances are retained. Human Verification. In the final stage, ten expert annota- tors review and correct the NVASR-generated transcripts, val- idating the pragmatic appropriateness of NV labels. To ensure annotation consistency, 5% of the data was cross-annotated, achieving a high Cohen’s kappa above 0.85. This process yields a dataset of 1,651 prompt and GT pairs (7.9 hours), which are standardized to MP3 format at a 24 kHz sampling rate. Table 2: Unified label inventory categorized by communicative function. LanguageVegetative SoundsAffect BurstsConversational Grunts MandarinBreathing, Cough, SighLaughter, Surprise-ah, Surprise-oh, Dissatisfaction-hnn Uhm, Confirmation-en, Question-ei, Question-ah, Question-en, Question-oh EnglishBreathing, Cough, SighLaughter, Surprise-ohUhm, Question-huh Table 3: Comparison of CER and OCER (%) across testsets, where bracketed values are OCER and other values are CER. DatasetSVQwen2.5-OmniNVASR WS-net5.7720.145.55 LS-other12.7923.359.90 SMIIP-NV3.123.59 (4.17)1.29 (1.36) NVTTS 14.4521.69 (26.95)13.52 (16.10) 3. NV-Bench Despite the proliferation of NV-capable TTS systems, the field lacks a standardized benchmark to disentangle two distinct fail- ure modes: (i) failure to generate the intended event, and (i) generation of low-quality or unnatural audio. To address this, NV-Bench evaluates systems across two dimensions: Instruction Alignment: Evaluates the model’s ability to strictly follow textual prompts. Specifically, we assess whether the system generates target NV events at precise linguistic po- sitions without omissions or hallucinations, serving as a proxy for robust text-to-speech alignment. Acoustic Fidelity: Assesses realism relative to real record- ings. We quantify the distributional gap, timbre consistency and their perception quality to evaluate their acoustic fidelity. To probe these capabilities under varying complexity, NV- Bench is structured into two multi-lingual subsets: Single-label Subset: A strictly balanced testset where each utterance contains exactly one NV event (50 samples per cate- gory, 650 Mandarin and 350 English utterances in total), which isolates fundamental generation capabilities. Multi-label Subset: A more challenging set containing ut- terances with multiple (2+) NV events to test robustness un- der dense paralinguistic conditions. Acknowledging the long- tailed nature of natural co-occurrences, we ensure relative bal- ance (Mandarin: 41–91, English: 75–112 samples per label), providing sufficient coverage for all event types while reflect- ing real-world complexity. 4. Experiments 4.1. Experiments for multi-lingual NVASR As the backbone evaluator for NV-Bench, the performance of multi-lingual NVASR needs to be qualified in both standard au- tomatic speech recognition (ASR) testsets and NV testsets. 4.1.1. Experimental setup and baselines We compare our NVASR against two strong baselines: SenseVoice-Small (SV) [12]: The original ASR model be- fore finetuning, as a baseline for general ASR capability. Qwen2.5-Omni [18]: A 7B multimodal LLM capable of speech understanding. We use the finetuned checkpoint for NV recognition released in [14]. We evaluate on four testsets: WenetSpeech test-net (WS- net) [19] and LibriSpeech test-other (LS-other) [20] for general ASR performance; and SMIIP-NV [8], NVTTS [4] testset for NV-specific performance. 4.1.2. Evaluation metrics To jointly evaluate linguistic and paralinguistic accuracy, we ex- tend standard CER into Overall CER (OCER): OCER = S + D + I N text + N nvv × 100%(2) where S,D,I represent substitutions, deletions, and inser- tions computed over the full sequence including NV labels and N text ,N nvv denote the count of text characters and NV symbols. 4.1.3. Results As detailed in Table 3, our multi-lingual NVASR model demon- strates dual-capability.First, it strictly maintains and even marginally improves upon, the high-quality speech transcrip- tion of the original SenseVoice model on standard datasets. Second, it significantly outperforms the MNV-17 finetuned Qwen2.5-Omni in accurately detecting and classifying NV la- bels. Notably, on the SMIIP-NV testset, our NVASR achieves 1.29% CER and 1.36% OCER, confirming its reliability as an automated evaluator for NV-Bench. 4.2. Benchmarking NV-Capable TTS models We conduct a comprehensive evaluation on NV-Bench to assess the current performance of NV-capable synthesis in two dimen- sions: (1) Instruction Alignment, measures prompt control- lability and paralinguistic intelligibility; (2) Acoustic Fidelity, evaluates the distributional gap, speaker similarity, perceptual quality to assess the realism of the synthesized audio. 4.2.1. TTS systems Orpheus-TTS 2 : A single-speaker Llama-based 3B model with explicit NV control. SMIIP-NV-CV2 [8] & Emilia-NV-CV2 [7]: Two variants of the 0.5B CosyVoice2 [21], a zero-shot TTS model, finetuned respectively on the SMIIP-NV and Emilia-NV datasets. CosyVoice3 (CV3) [22]:The foundational zero-shot model, serving as a strong baseline for general speech synthesis. NV-FlexiVoice: We finetune a 0.5B FlexiVoice model [23] pretrained on the Emilia dataset [16] without NV events. NV-CV3: To establish a high-performance reference base- line, we finetune CV3 on a consolidated corpus. 2 https://orpheustts.net Table 4: Detailed performance on each subsets of NV-Bench. Bold indicates best in the column, underlinesecond-best. System Single-label SubsetMulti-label Subset AlignmentFidelityAlignmentFidelity CER (%) PCER (%) OCER (%)SIMDNSMOSCER (%) PCER (%) OCER (%)SIMDNSMOS Mandarin GT3.869.384.070.7813.123.7923.715.070.7943.12 Orpheus-TTS 11.3688.7713.91-3.4319.8384.8524.38-3.40 SMIIP-NV-CV28.8075.6411.340.7193.2210.6677.2014.790.7153.07 Emilia-NV-CV2 5.0540.006.640.7403.215.5448.748.090.7463.24 CosyVoice33.8557.695.860.7643.304.7561.948.260.7153.31 NV-FlexiVoice6.9831.088.150.7483.228.2039.3710.390.7503.07 NV-CV33.8027.694.900.7683.293.4430.044.840.7763.29 English GT6.738.316.900.7723.117.2621.418.620.7753.14 Orpheus-TTS9.0371.9210.63-3.338.6871.4611.89-3.34 SMIIP-NV-CV2 17.9256.8019.470.5832.9720.9354.4923.870.5802.97 Emilia-NV-CV212.5055.3013.210.6393.2111.7160.2814.630.6553.26 CosyVoice3 7.8762.759.060.7013.276.3957.8410.690.7153.31 NV-FlexiVoice11.8850.4313.210.6853.159.6051.3213.760.7083.07 NV-CV3 8.3346.139.440.6983.246.7047.1310.100.7213.30 Table 5: Evaluation on the full NV-Bench. SystemFADFDIMOSNMOS GT--4.39 ± 0.184.39 ± 0.15 Orpheus-TTS5.7124.493.27 ± 0.233.53 ± 0.22 SMIIP-NV-CV21.326.713.28 ± 0.223.28 ± 0.19 Emilia-NV-CV21.085.573.89 ± 0.183.99 ± 0.14 CosyVoice30.909.463.56 ± 0.223.94 ± 0.20 NV-FlexiVoice0.292.723.94 ± 0.234.00 ± 0.18 NV-CV30.863.943.95 ± 0.184.08 ± 0.16 4.2.2. Experimental setup Training Configuration. We finetune CV3 and FlexiVoice using a consolidated corpus comprising Emilia-NV, SMIIP-NV, NVTTS, Disfluency, and NVS. Following Section 2.1.2, nor- malized NV labels were injected into the text as special con- trol symbols. The model was optimized using AdamW [24] (lr = 1× 10 −5 ) on 4 NVIDIA A800 GPUs. Inference Procedures. All models follow the label nor- malization protocol defined in Section 2.1.2. For baselines with limited NV support, we evaluate only the intersection of sup- ported events. Unsupported interjections (e.g., [Question-huh]) are mapped to their nearest lexical equivalents with punctuation (e.g., “huh?”) to approximate the intended pragmatic function. 4.2.3. Evaluation metrics Instruction alignment. We utilize our NVASR (Section 2.1) to compute CER, OCER and PCER. To isolate NV gener- ation accuracy, PCER is calculated on extracted NV symbols: PCER = (S nvv + D nvv + I nvv )/N nvv , where S nvv ,D nvv ,I nvv are edit operations on NVs and N nvv is the target NV count. Acoustic Fidelity.Perceptual quality and timbre con- sistency are measured using DNSMOS and WavLM-based Speaker Similarity (SIM) [25]. To assess the distributional gap, we compute the Fr ́ echet Audio Distance (FAD) [26] and FD from PANNs [27]. Human Evaluation. Ten annotators rated 100 utterances per model on a 5-point scale for two dimensions: (1) NMOS (Naturalness), assessing acoustic fidelity, text-NV prosodic continuity and speaker consistency; (2) IMOS (Instruction Ac- curacy), evaluating precise NV execution without omissions, hallucinations and mispronunciations. 4.2.4. Results and analysis Table 4 presents performance across Mandarin and English subsets. In terms of instruction alignment, NV-CV3 achieves the lowest PCER (27.69%) and OCER (4.90%) on the Mandarin single-label subset, demonstrating superior controllability with large-scale, diversified training data. For acoustic fidelity, NV-CV3 maintains the strong tim- bre consistency of CV3, while Orpheus-TTS scores highest in DNSMOS perception. Table 5 shows NV-FlexiVoice achieves the lowest FAD and FD, indicating that its generated speech and NV events most closely match the real distribution. Subjective evaluations support these results, with NV-CV3 achieving the highest NMOS and IMOS. Crucially, IMOS ex- hibits a significant negative Spearman correlation with PCER (ρ = −0.65,p < 0.001) and NMOS aligns with FD, thereby validating the reliability of NV-Bench’s objective metrics. 5. Conclusion We introduce NV-Bench, the first comprehensive benchmark for NV-capable TTS systems. Grounded in a functional taxon- omy, the benchmark comprises 1,651 multi-lingual, in-the-wild utterances paired with human ground-truth, balanced across 14 categories into single- and multi-label subsets. NV-Bench evaluates acoustic fidelity and instruction alignment separately, thereby disentangling audio quality from paralinguistic control- lability. Extensive experiments confirm that our objective met- rics strongly correlate with human judgments, establishing NV- Bench as a standardized evaluation framework. 6. Generative AI Use Disclosure We utilized LLM solely for the purpose of refining the clarity and grammar of the text. The authors reviewed and revised the output to ensure accuracy. 7. References [1] Y. Wang, H. Zhan, L. Liu, R. Zeng, H. Guo, J. Zheng, Q. Zhang, X. Zhang, S. Zhang, and Z. Wu, “Maskgct: Zero-shot text-to- speech with masked generative codec transformer,” arXiv preprint arXiv:2409.00750, 2024. [2] Y. Chen, Z. Niu, Z. Ma, K. Deng, C. Wang, J. JianZhao, K. Yu, and X. Chen, “F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, p. 6255–6271. [3] C. Yan, B. Wu, P. Yang, P. Tan, G. Hu, L. Xie, Y. Zhang, F. Tian, X. Yang, X. Zhang et al., “Step-audio-editx technical report,” arXiv preprint arXiv:2511.03601, 2025. [4] M. Borisov, E. Spirin, and D. Diatlova, “Nonverbaltts:A public english corpus of text-aligned nonverbal vocalizations with emotion annotations for text-to-speech,” arXiv preprint arXiv:2507.13155, 2025. [5] H. Wang, J. Hai, D. Chong, K. Thakkar, T. Feng, D. Yang, J. Lee, T. Thebaud, L. M. Velazquez, J. Villalba et al., “Cap- speech: Enabling downstream applications in style-captioned text-to-speech,” arXiv preprint arXiv:2506.02863, 2025. [6] A. Batliner, S. Amiriparian, and B. W. Schuller, “Non-verbal vo- calisations and their challenges: Emotion, privacy, sparseness, and real life,” arXiv preprint arXiv:2508.01960, 2025. [7] H. Liao, Q. Ni, Y. Wang, Y. Lu, H. Zhan, P. Xie, Q. Zhang, and Z. Wu, “Nvspeech: An integrated and scalable pipeline for human-like speech modeling with paralinguistic vocalizations,” arXiv preprint arXiv:2508.04195, 2025. [8] Z. Wu, D. Liu, J. Liu, Y. Wang, L. Li, L. Jin, H. Bu, P. Zhang, and M. Li, “Smiip-nv: A multi-annotation non-verbal expressive speech corpus in mandarin for llm-based speech synthesis,” in Proceedings of the 33rd ACM International Conference on Multi- media, 2025, p. 12 564–12 570. [9] K. Wang and D. Herremans, “Disfluencyspeech – single-speaker conversational speech dataset with paralanguage,” 2024. [Online]. Available: https://arxiv.org/abs/2406.08820 [10] R. Ye, Y. Zhou, R. Yu, Z. Lin, K. Li, X. Li, X. Liu, G. Zeng, and Z. Wu, “A scalable pipeline for enabling non-verbal speech generation and understanding,” arXiv preprint arXiv:2508.05385, 2025. [11] B. Bai, Q. Lu, W. Yang, Z. Sun, Y. Hou, P. Jia, S. Pu, R. Fu, Y. Gao, Y. Li et al., “Synparaspeech: Automated synthesis of paralinguistic datasets for speech generation and understanding,” arXiv preprint arXiv:2509.14946, 2025. [12] K. An, Q. Chen, C. Deng, Z. Du, C. Gao, Z. Gao, Y. Gu, T. He, H. Hu, K. Hu et al., “Funaudiollm: Voice understanding and gen- eration foundation models for natural interaction between humans and llms,” arXiv preprint arXiv:2407.04051, 2024. [13] A. Graves, S. Fern ́ andez, F. Gomez, and J. Schmidhuber, “Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning, 2006, p. 369–376. [14] J. Mai, J. Ji, X. Xing, C. Yang, W. Chen, J. Xing, and X. Xu, “Mnv-17: A high-quality performative mandarin dataset for nonverbal vocalization recognition in speech,” arXiv preprint arXiv:2509.18196, 2025. [15] J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE inter- national conference on acoustics, speech and signal processing (ICASSP). IEEE, 2017, p. 776–780. [16] H. He, Z. Shang, C. Wang, X. Li, Y. Gu, H. Hua, L. Liu, C. Yang, J. Li, P. Shi et al., “Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,” in 2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2024, p. 885–890. [17] L.-C.-T. Xiaomi, “Mimo-audio: Audio language models are few- shot learners,” 2025. [Online]. Available: GitHub-XiaomiMiMo/ MiMo-Audio [18] J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang, B. Zhang, X. Wang, Y. Chu, and J. Lin, “Qwen2.5-omni technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2503.20215 [19] B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng et al., “Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, p. 6182–6186. [20] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- rispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015, p. 5206–5210. [21] Z. Du, Y. Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y. Yang, C. Gao, H. Wang et al., “Cosyvoice 2: Scalable stream- ing speech synthesis with large language models,” arXiv preprint arXiv:2412.10117, 2024. [22] Z. Du, C. Gao, Y. Wang, F. Yu, T. Zhao, H. Wang, X. Lv, H. Wang, C. Ni, X. Shi et al., “Cosyvoice 3: Towards in-the- wild speech generation via scaling-up and post-training,” arXiv preprint arXiv:2505.17589, 2025. [23] D. Chen, X. Zhang, Y. Wang, K. Dai, L. Ma, and Z. Wu, “Flex- ivoice: Enabling flexible style control in zero-shot tts with natural language instructions,” arXiv preprint arXiv:2601.04656, 2026. [24] I. Loshchilov and F. Hutter, “Decoupled weight decay regulariza- tion,” arXiv preprint arXiv:1711.05101, 2017. [25] P. Anastassiou, J. Chen, J. Chen, Y. Chen, Z. Chen, Z. Chen, J. Cong, L. Deng, C. Ding, L. Gao et al., “Seed-tts: A family of high-quality versatile speech generation models,” arXiv preprint arXiv:2406.02430, 2024. [26] K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fr\’echet audio distance: A metric for evaluating music enhancement algo- rithms,” arXiv preprint arXiv:1812.08466, 2018. [27] Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumb- ley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, p. 2880–2894, 2020.