Paper deep dive
Augmenting Dysarthric Speech Severity Assessment with MOS Supervision
Kaimeng Jia, Minzhu Tu, Zengrui Jin, Siyin Wang, Chao Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 6/21/2026, 2:44:25 AM
Summary
This paper proposes a novel data augmentation strategy for automatic dysarthric speech severity assessment by leveraging human-annotated Mean Opinion Score (MOS) data from the QualiSpeech speech synthesis corpus. The researchers demonstrate that the perceptual commonalities between speech synthesis artifacts and dysarthric speech (such as reduced intelligibility and unnaturalness) allow for effective cross-domain knowledge transfer. Using self-supervised learning (SSL) models like wav2vec 2.0 and HuBERT, the study compares two training paradigms: Fine-Tuning (FT) and Joint Training (JT). Results show that FT consistently improves both intelligibility and naturalness prediction, whereas JT primarily benefits naturalness prediction. The findings suggest that speech synthesis evaluation corpora can serve as a practical, scalable source to mitigate the scarcity of clinically annotated dysarthric speech.
Entities (8)
Relation Signals (4)
Wav2Vec 2.0 → usedasbackbonefor → Dysarthria Assessment
confidence 100% · self-supervised learning (SSL) pre-trained encoders are adopted as the backbone for feature extraction
Fine-Tuning → improves → Intelligibility Prediction
confidence 95% · Speech synthesis augmentation via FT consistently improves intelligibility prediction across all wav2vec encoders
Joint Training → improves → Naturalness Prediction
confidence 95% · while joint training yields gains primarily on naturalness
QualiSpeech → providesaugmentationfor → Dysarthria Assessment
confidence 90% · leveraging TTS assessment corpora, which are readily available and richly annotated, as an augmentation source for dysarthric speech assessment
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness. Automatic utterance-level assessment of dysarthric speech can support scalable speech monitoring and therapy-related analysis. Yet training such systems is bottlenecked by the scarcity of clinically annotated dysarthric speech. This work proposes to augment dysarthric speech assessment using data from speech synthesis evaluations, specifically human-annotated utterances with Mean Opinion Score (MOS) labels from the QualiSpeech corpus. Experiments show that fine-tuning on speech synthesis assessment data consistently improves performance on both intelligibility and naturalness prediction, while joint training yields gains primarily on naturalness. These results suggest that synthesis artifacts and dysarthric speech share perceptual commonalities, and speech synthesis evaluation corpora offer a practical augmentation source that reduces reliance on scarce clinical annotations.
Tags
Links
- Source: https://arxiv.org/abs/2606.18645v1
- Canonical: https://arxiv.org/abs/2606.18645v1
Trouble viewing inline? Open PDF directly →
Full Text
34,381 characters extracted from source content.
Expand or collapse full text
Augmenting Dysarthric Speech Severity Assessment with MOS Supervision Kaimeng Jia †1 , Minzhu Tu †2 , Zengrui Jin ∗1 , Siyin Wang 1 , Chao Zhang 1 1 Tsinghua University, 2 Beijing University of Posts and Telecommunications jkm23@mails.tsinghua.edu.cn, Epiphany 1104@bupt.edu.cn, wangsiyi23@mails.tsinghua.edu.cn,zrjin,cz277@tsinghua.edu.cn Abstract Dysarthria is a speech disorder marked by reduced intelligibil- ity and communicative effectiveness. Automatic utterance-level assessment of dysarthric speech can support scalable speech monitoring and therapy-related analysis. Yet training such sys- tems is bottlenecked by the scarcity of clinically annotated dysarthric speech. This work proposes to augment dysarthric speech assessment using data from speech synthesis evalua- tions, specifically human-annotated utterances with Mean Opin- ion Score (MOS) labels from the QualiSpeech corpus. Exper- iments show that fine-tuning on speech synthesis assessment data consistently improves performance on both intelligibility and naturalness prediction, while joint training yields gains pri- marily on naturalness. These results suggest that synthesis arti- facts and dysarthric speech share perceptual commonalities, and speech synthesis evaluation corpora offer a practical augmenta- tion source that reduces reliance on scarce clinical annotations. Index Terms:Dysarthria, Dysarthric Speech, Automatic Dysarthria Assessment, Mean Opinion Score 1. Introduction Dysarthria is a neuro-motor speech disorder resulting from neurological injuries or diseases, such as cerebral palsy, amy- otrophic lateral sclerosis, Parkinson’s disease, or stroke [1, 2]. These disruptions manifest in several perceptually salient char- acteristics, with reduced intelligibility and degraded naturalness being among the most prominent and impactful on daily com- munication. Despite recent advancement achieved in automatic speech recognition (ASR) [3, 4, 5, 6] and dysarthric speech recognition [7, 8, 9], accurate assessment of dysarthric speech remains a challenging problem with important clinical and technological implications. It serves as a sensitive biomarker for early detec- tion and tracking of neurological disease progression [10], pro- vides an objective measure to guide speech therapy and monitor rehabilitation outcomes [11], and is a key factor for improving downstream systems such as dysarthric speech recognition and disordered speech reconstruction [12, 13, 14, 15]. Despite its significant importance, clinical assessment of utterance-level intelligibility and naturalness still relies heav- ily on subjective evaluations by certified speech pathologists [16, 17], which are labor-intensive and difficult to scale. These limitations have motivated the development of automated, ob- jective assessment methods. However, patients with speech disorders often present with co-occurring physical disabili- ties, making large-scale data collection particularly challeng- ing. Furthermore, existing approaches are typically restricted † Equal contribution was made between the first two authors. * Corresponding author. to constrained lexicons and require matched control groups, re- sulting in an unnatural evaluation protocol. To the best of our knowledge, spontaneous utterances with unconstrained vocab- ulary have not been employed in prior studies. To address the data scarcity issue, self-supervised learn- ing (SSL) has emerged as a powerful paradigm in speech pro- cessing, offering robust and transferable representations from large-scale unlabeled corpora. Models such as wav2vec 2.0 [18] and HuBERT [19, 20] have been successfully applied to both detection and severity assessment of dysarthria [21, 22]. A wide spectrum of data augmentation techniques has also been explored, including signal processing [23], adversarial domain adaptation [24], generative approaches [25, 26, 27], reverse au- toencoders that transform healthy speech into dysarthric speech [28], and other pre-trained TTS systems [29, 30, 31]. However, these generative approaches are designed to model spectro- temporal characteristics of dysarthric speech for ASR or speech reconstruction purposes. The augmented data they produce lacks perceptual validation, as no clinical experts re-annotate the generated samples with severity labels, leaving it unclear whether the synthetic speech aligns with human perceptual judgments. As a result, such data cannot be directly used as augmentation for severity assessment, where label quality and perceptual alignment are critical. Meanwhile, automated assessment of Text-to-Speech (TTS) synthesis quality has been extensively investigated [32, 33, 34, 35, 36, 37, 38, 39], producing large quantities of human- annotated utterances with mean opinion score (MOS) labels measuring naturalness and intelligibility. Crucially, both syn- thesis artifacts and dysarthric speech represent deviations from natural human production, sharing perceptual characteristics such as reduced intelligibility, unnatural prosody, and degraded fluency. This acoustic and perceptual overlap motivates a novel data augmentation strategy, leveraging TTS assessment cor- pora, which are readily available and richly annotated, as an augmentation source for dysarthric speech assessment systems. The contributions of this work are threefold: (1) An em- pirical study of SSL-based automatic dysarthric speech assess- ment on utterances with unconstrained vocabulary, going be- yond the constrained protocols of prior work. (2) Demonstra- tion that synthesized speech with human-annotated MOS scores from TTS evaluation corpora can effectively augment dysarthric speech assessment systems. (3) Evidence that synthesis fail- ures and dysarthric manifestations share perceptual and acous- tic commonalities, offering a novel perspective on cross-domain knowledge transfer between TTS and articulatory disorders. The rest of this paper is organized as follows. Section 2 introduces the assessment task and describes the Speech Ac- cessibility Project (SAP) and QualiSpeech corpora. Section 3 presents the model architecture and the two training paradigms: arXiv:2606.18645v1 [eess.AS] 17 Jun 2026 joint training and fine-tuning. Section 4 reports experiments and results on intelligibility and naturalness prediction. The last section concludes this work. 2. Task Description 2.1. The Speech Accessibility Project Challenge 2025 1234567 Severity Level 0 1000 2000 3000 4000 # utterances (a) Intelligibility 1234567 Severity Level 0 250 500 750 1000 1250 1500 1750 2000 # utterances (b) Naturalness Figure 1: Distribution of utterances across severity levels for (a) Intelligibility and (b) Naturalness in the Speech Accessibil- ity Project (SAP) challenge. Intelligibility is heavily biased to- ward level 1 (minimal impairment), while Naturalness exhibits a more balanced spread across levels, reflecting the greater variability in perceived speech quality among dysarthric speak- ers. The SAP challenge [40] provides a large-scale, open- domain corpus of dysarthric speech. The corpus comprises utterances recorded from over 500 speakers diagnosed with Parkinson’s disease, Down syndrome, amyotrophic lateral scle- rosis, Cerebral palsy, or stroke, encompassing more than 400 hours of speech and over 190,000 utterances. The dataset was recorded through participants using their personal devices at home, capturing both read and spontaneously generated speech with unconstrained vocabulary to ensure naturalness and vari- ability. Fine-grained perceptual speech assessment across mul- tiple clinically relevant dimensions is provided in the corpus, including but not limited to, harsh voice, inappropriate silences, pitch level, naturalness, variable rate, and intelligibility. Each dimension is annotated by certified speech-language patholo- gists using a 7-point scale, where a higher score indicates more significant severity is manifested in the corresponding dimen- sion. Given the complexity among these perceptual dimensions and the primary objective of dysarthric speech assessment, the focus is specifically narrowed to two core utterance-level an- notations, Intelligibility and Naturalness. Distribution of utter- ances across severity levels can be seen from Figure 1. Due to the unavailability of the original SAP test labels, the development partition was used as the test set. The released SAP training and development partitions utilize a speaker-level data partitioning protocol, thus the test speakers are unseen dur- ing SAP training. For each target dimension, utterances an- notated for that dimension were selected from the SAP train- ing partition, yielding 5,046 utterances for Intelligibility and 5,040 for Naturalness before validation sampling. Validation sets consisting of 500 utterances per dimension were randomly sampled from the corresponding training subsets and used only for model selection. These validation utterances were removed from the corresponding training subsets, ensuring no utterance overlap across train, validation, and test. The resulting test sets include 716 utterances for Intelligibility and 714 utterances for Naturalness. 2.2. QualiSpeech: A Descriptive Corpus for Speech Quality Assessment QualiSpeech [41] is an English corpus for low-level speech quality assessment with comprehensive annotations and de- tailed descriptive comments. It comprises three speech cate- gories: synthetic speech (49%), primarily from the Blizzard Challenge and Voice Conversion Challenge (BVCC) corpus [42] and various TTS models with standardized MOS scores; simulated real speech (27%), sourced from the Non-Intrusive Speech Quality Assessment (NISQA) corpus [43] featuring distortions that emulate actual transmission; and real speech (24%), encompassing NISQA live recordings and in-the-wild samples from GigaSpeech [44]. To ensure balanced speech quality conditions, 20% of the synthetic speech samples are mixed with noise at signal-to-noise ratios ranging between 0 and 15 dB, while real speech samples are stratified into four quality groups based on UTMOS-predicted MOS scores [45]. Each utterance is evaluated across seven perceptual quality di- mensions, namely noise, distortion, speed, continuity, listen- ing effort, naturalness, and overall quality. This protocol re- sults in training, validation, and test splits of 10,558, 2,167, and 1,852 utterances, respectively, with balanced compositions of synthetic and real speech. Among the seven annotated dimensions, overall quality and naturalness are selected as they provide complementary yet comprehensive perspectives on perceptual evaluation. Over- all quality captures the aggregated listener impression across multiple degradations, while naturalness reflects the degree to which speech resembles genuine human production. Focus- ing on these two dimensions ensures both clinical relevance for pathological speech assessment and consistency with estab- lished practices in general speech quality evaluation. 3. Method 3.1. Model Architecture To leverage transferable representations from large-scale un- labeled speech, self-supervised learning (SSL) pre-trained en- coders are adopted as the backbone for feature extraction [46]. The core idea is to augment the limited in-domain dysarthric speech data with human-annotated TTS assessment data from QualiSpeech, using either joint training or fine-tuning to trans- fer perceptual supervision into the dysarthria assessment model. Given a raw waveform input, the SSL encoder produces a se- quence of frame-level contextual representations. These are ag- gregated via mean pooling over the time dimension to obtain a fixed-dimensional utterance-level embedding, which captures global perceptual characteristics without requiring explicit seg- mentation. The utterance embedding is then passed to a re- gression head consisting of a two-layer feed-forward network with ReLU activation and dropout regularization between lay- ers, producing a single continuous severity score. The entire model, including the SSL encoder, is fine-tuned in an end-to- end fashion during training. 3.2. Training Paradigms Two training paradigms are proposed to incorporate perceptual supervision from QualiSpeech into dysarthric speech assess- ment, as illustrated in Figure 2. Joint Training (JT). The JT paradigm trains a single model simultaneously on both the QualiSpeech and SAP corpora un- der a unified regression objective. To prevent the larger Qual- Overall Linear mapped to SAP scale Naturalness Qualispeech OR Intelli.NaturalnessOR Speech Accessi- bility Project 1:1 Pre-trained SSL Encoder Reg. Head Temp. Pooling Overall Linear mapped to SAP scale Naturalness Qualispeech OR Pre-trained SSL Encoder Reg. Head Temp. Pooling Intelli.NaturalnessOR Speech Accessi- bility Project Pre-trained SSL Encoder Reg. Head Temp. Pooling initialize Discard (a) Joint Training (JT)(b) Fine-Tuning (FT) Stage 1: Pre-train on QualispeechStage 2: Fine-tune on SAP Figure 2: An illustration of the proposed training paradigms. In Joint Training (JT), QualiSpeech and Speech Accessibility Project (SAP) utterances are mixed at a 1:1 ratio and fed jointly into a shared SSL encoder and regression head, with QualiSpeech MOS scores linearly mapped to the SAP severity scale. In Fine-Tuning (FT), the model is first trained on QualiSpeech, then fine-tuned on SAP without simultaneous cross-dataset optimization. iSpeech corpus from dominating training, utterances are ran- domly sampled from the QualiSpeech training set to match the size of the SAP training split, yielding a balanced 1:1 mixture ratio. Since the two datasets adopt different rating scales, with QualiSpeech using a MOS scale of 1–5 while SAP uses a sever- ity scale of 1–7, a linear transformation is applied to align Qual- iSpeech scores before joint optimization: ˆs = 1 + (5− s MOS )· 6 4 ,(1) where s MOS ∈ [1, 5] is the original QualiSpeech MOS and ˆs is the aligned score in the SAP scale. Under this mapping, a MOS of 5 (highest quality) corresponds to an SAP score of 1 (no dysarthric characteristics), and a MOS of 1 (lowest quality) maps to an SAP score of 7 (most severe). Fine-Tuning (FT). The FT paradigm decouples learning into two sequential stages. The model is first trained on Qual- iSpeech to predict the selected MOS dimension and acquire per- ceptual quality representations. The resulting weights then ini- tialize the SAP model, which is fine-tuned for dysarthric speech severity prediction. This design transfers perceptual supervision while avoiding simultaneous cross-dataset optimization. 3.3. Evaluation Metrics Model performance is evaluated using three complementary metrics: Mean Squared Error (MSE), Linear Correlation Co- efficient (LCC), and Spearman’s Rank Correlation Coefficient (SRCC). MSE quantifies the absolute deviation between pre- dicted scores and ground-truth ratings, with lower values indi- cating higher prediction accuracy. LCC measures the strength of the linear relationship between predictions and references, while SRCC assesses the consistency of relative ranking order, which is particularly relevant for severity-level discrimination. Together, these three metrics provide a comprehensive view of both regression accuracy and perceptual alignment with human judgments. 4. Experiments and Results To leverage human-perception-aligned supervision from Qual- iSpeech for dysarthria assessment, we evaluate two TTS-based data augmentation paradigms: joint training (JT) and fine- tuning (FT). All models evaluated are optimized using the Adam optimizer with mean squared error (MSE) as the train- ing objective, a learning rate of 1e− 5, and weight decay of 0.01. Details of the utilized self-supervised learning (SSL) pre- trained encoders are listed in Table 1. Table 1: Details of self-supervised learning (SSL) pre-trained encoders. Model# Param.Dataset wav2vec 2.0 Base94MLibrispeech [47] wav2vec 2.0 Large*315MLibri-Light [48] wav2vec 2.0 Large+315M Libri-Light [48] + CommonVoice [49] + Switchboard [50] + Fisher [51] HuBERT Base95MLibrispeech [47] HuBERT Large316MLibri-Light [48] All training paradigms share an identical experimental setup with parameters of SSL encoders kept trainable. Under the JT paradigm, 4,000 utterances are randomly sampled from the QualiSpeech training set with the original data distribution preserved, and then combined with the SAP training set for joint optimization. 4.1. Performance on the Intelligibility Dimension Table 2 (Row 1–5) reports intelligibility prediction on SAP, with Overall quality and Naturalness from QualiSpeech used as aux- iliary supervision under fine-tuning (FT) and joint training (JT) across five SSL encoders. Three key trends emerge: (i) Under in-domain training (IDT, Row 1), wav2vec 2.0 Base achieves the strongest performance, with a 36.4% relative MSE reduction over Large+, suggesting that smaller encoders gener- alize better to this task without additional supervision. (i) Speech synthesis augmentation via FT (Row 2 vs. 1) con- sistently improves intelligibility prediction across all wav2vec encoders, with the largest relative MSE reduction of 43.7% for Large+ and a 19.6% LCC improvement for Base, demonstrating that QualiSpeech overall quality provides effective perceptual transfer for intelligibility. (i) FT consistently outperforms JT across all encoder and di- mension combinations (Row 2 vs. 3, Row 4 vs. 5), with Over- all quality yielding the most robust MSE and LCC gains while https://dl.fbaipublicfiles.com/fairseq/ wav2vec/wav2vec_small.pt https://dl.fbaipublicfiles.com/fairseq/ wav2vec/wav2vec_vox_new.pt https://dl.fbaipublicfiles.com/fairseq/ wav2vec/w2v_large_lv_fsh_swbd_cv.pt https://dl.fbaipublicfiles.com/hubert/ hubert_base_ls960.pt https://dl.fbaipublicfiles.com/hubert/ hubert_large_l60k.pt Table 2: Results of self-supervised learning (SSL) pre-trained encoders under joint training (JT) and fine-tuning (FT) on the Speech Accessibility Project dysarthric speech corpus (SAP) and QualiSpeech. The two “Dimension” columns denote the SAP target di- mension (left) and the QualiSpeech auxiliary supervision dimension (right). “IDT” denotes in-domain training on SAP only, with no QualiSpeech augmentation (“–”). For each SSL encoder, Mean Squared Error (MSE↓), Linear Correlation Coefficient (LCC↑), and Spearman’s Rank Correlation Coefficient (SRCC ↑) are reported. Bold values indicate the best result per encoder within each SAP dimension group. ID Method Dimensionwav2vec 2.0 Basewav2vec 2.0 Large* wav2vec 2.0 Large+HuBERT BaseHuBERT Large SAPQualiSpeech MSE LCC SRCC MSE LCC SRCC MSE LCC SRCC MSE LCC SRCC MSE LCC SRCC 1IDTIntelligibility–0.348 0.628 0.482 0.421 0.523 0.368 0.547 0.471 0.322 0.461 0.540 0.406 0.475 0.485 0.404 2FT IntelligibilityOverall 0.272 0.751 0.427 0.351 0.612 0.393 0.308 0.534 0.285 0.487 0.594 0.368 0.431 0.596 0.356 3JT0.408 0.590 0.446 0.560 0.317 0.330 0.627 0.225 0.302 0.612 0.387 0.317 0.526 0.479 0.363 4FT Intelligibility Naturalness 0.303 0.648 0.464 0.387 0.572 0.401 0.495 0.566 0.443 0.379 0.451 0.295 0.379 0.608 0.464 5JT0.433 0.602 0.475 0.604 0.383 0.352 0.528 0.445 0.348 0.489 0.515 0.309 0.451 0.551 0.391 6IDTNaturalness–1.127 0.581 0.591 1.354 0.469 0.481 1.119 0.574 0.607 1.169 0.546 0.521 0.941 0.570 0.503 7FT NaturalnessOverall 1.053 0.718 0.723 1.033 0.695 0.706 1.000 0.686 0.696 0.847 0.644 0.587 0.819 0.701 0.690 8JT0.909 0.691 0.680 1.027 0.606 0.613 1.106 0.589 0.617 1.050 0.614 0.634 0.930 0.661 0.668 9FT NaturalnessNaturalness 0.717 0.717 0.657 0.800 0.709 0.713 0.880 0.692 0.690 1.075 0.646 0.635 0.823 0.678 0.675 10JT1.022 0.637 0.643 1.085 0.629 0.648 1.022 0.637 0.642 1.121 0.586 0.590 1.036 0.638 0.635 Naturalness further improves SRCC, particularly for larger en- coders, indicating that the choice of augmentation dimension interacts with model capacity. 4.2. Performance on Naturalness Dimension Table 2 (Row 6–10) reports naturalness prediction on SAP, uti- lizing Overall quality and Naturalness from QualiSpeech as auxiliary supervision under fine-tuning (FT) and joint training (JT) across five SSL encoders. Three key trends emerge: (i) Speech synthesis augmentation via FT (Row 7 vs. 6) consis- tently improves naturalness prediction across all encoders, with wav2vec 2.0 Large* showing the largest LCC (48.2% relative) and SRCC (46.8% relative) gains, confirming that overall qual- ity from QualiSpeech transfers effectively to dysarthric natural- ness assessment. (i) Unlike intelligibility, JT also yields consistent gains over IDT on naturalness (Row 8 vs. 6), with wav2vec 2.0 Base and Large* reducing MSE by 19.4% and 24.2% respectively, suggesting that the perceptual alignment between QualiSpeech quality ratings and dysarthric naturalness is sufficient to support joint optimization. (i) Using QualiSpeech Naturalness as auxiliary supervision under FT (Row 9) achieves the lowest MSE across all con- ditions, with wav2vec 2.0 Base and Large* achieving 36.4% and 40.9% relative MSE reductions, indicating that dimension- matched supervision yields the strongest transfer. Taken together, these results demonstrate that the percep- tual overlap between speech synthesis quality and dysarthric naturalness is sufficient to support effective cross-domain aug- mentation under both FT and JT, validating the use of speech synthesis evaluation corpora as a practical and label-efficient augmentation source for dysarthric speech assessment. 4.3. Analysis on the Performance Gap between FT and JT The consistent gap between FT and JT on intelligibility can be attributed to cross-domain label misalignment and negative transfer under joint optimization. The QualiSpeech supervision signals, overall quality and naturalness, are semantically mis- aligned with SAP intelligibility: while the latter reflects clini- cal judgments of phonemic clarity and word-level recoverabil- ity, the former capture aggregated listener impressions of multi- dimensional degradations and human-likeness, which are more closely related to the naturalness perceptual axis than to articu- latory intelligibility. Under JT, a shared regression head and a unified MSE ob- jective are applied over both corpora, implicitly treating the linearly mapped QualiSpeech scores as a proxy for SAP intel- ligibility severity. This assumption fails for intelligibility, as gradients from QualiSpeech bias the shared representations to- ward explaining MOS variance, suppressing the discriminative acoustic cues required for intelligibility assessment. FT sidesteps this conflict by decoupling the two learning stages: QualiSpeech pre-training initializes the encoder with perceptually informed representations, while subsequent fine- tuning on SAP allows the model to re-weight features toward the target domain without cross-dataset gradient interference, explaining its consistent advantage over JT. These findings sug- gest that speech synthesis evaluation data can serve as a viable augmentation source for dysarthric intelligibility assessment, provided that the transfer is mediated through sequential fine- tuning rather than joint optimization. 5. Conclusion This study demonstrated that leveraging perceptual annotations from the QualiSpeech corpus significantly enhances automatic dysarthria assessment. Fine-tuning consistently improved pre- diction accuracy over the baseline on both dimensions, while joint training yielded gains primarily on naturalness. Larger models derived greater benefit from additional supervision, and naturalness prediction consistently achieved stronger cor- relation with human judgments than intelligibility, likely at- tributable to the severe class imbalance in the intelligibility di- mension as shown in Figure 1. The effectiveness of synthetic speech augmentation underscores perceptual and acoustic com- monalities between synthesis failures and dysarthric speech, highlighting promising directions for future clinical research. 6. Generative AI Use Disclosure During the preparation of this manuscript, the authors used gen- erative AI to correct grammar mistakes and misspellings. Illus- trative icons in Figure 2 are AI generated. After using these tools, the authors carefully reviewed and edited the manuscript, and take full responsibility for the final content of the paper. 7. References [1] F. Rudzicz, “Articulatory Knowledge in the Recognition of Dysarthric Speech,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, no. 4, p. 947–960, 2011. [2] —, “Phonological Features in Discriminative Classification of Dysarthric Speech,” in Proc. IEEE ICASSP, Taipei, 2009. [3] H. Hadian, H. Sameti, D. Povey, and S. Khudanpur, “End-to- end Speech Recognition Using Lattice-free MMI,” in Proc. In- terspeech, Hyderabad, 2018. [4] A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Proc. Interspeech, Shanghai, 2020. [5] Z. Yao, L. Guo, X. Yang, W. Kang, F. Kuang, Y. Yang, Z. Jin, L. Lin, and D. Povey, “Zipformer: A faster and better encoder for automatic speech recognition,” in Proc. ICLR, Vienna, 2024. [6] Z. Yao, W. Kang, X. Yang, F. Kuang, L. Guo, H. Zhu, Z. Jin, Z. Li, L. Lin, and D. Povey, “CR-CTC: Consistency regularization on CTC for improved speech recognition,” in Proc. ICLR, Singapore, 2025. [7] Z. Jin, M. Geng, J. Deng, T. Wang, S. Hu, G. Li, and X. Liu, “Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, p. 413–429, 2023. [8] T. Wang, S. Hu, J. Deng, Z. Jin, M. Geng, Y. Wang, H. Meng, and X. Liu, “Hyper-parameter Adaptation of Conformer ASR Systems for Elderly and Dysarthric Speech Recognition,” in Proc. Inter- speech, Dublin, 2023. [9] H. Wang, Z. Jin, M. Geng, S. Hu, G. Li, T. Wang, H. Xu, and X. Liu, “Enhancing Pre-Trained ASR System Fine-Tuning for Dysarthric Speech Recognition Using Adversarial Data Augmen- tation,” in Proc. IEEE ICASSP, Seoul, 2024. [10] T. Bhattacharjee, Y. Belur, A. Nalini, R. Yadav, and P. K. Ghosh, “Exploring the Role of Fricatives in Classifying Healthy Subjects and Patients with Amyotrophic Lateral Sclerosis and Parkinson’s Disease,” in Proc. IEEE ICASSP, Rhodes Island, 2023. [11] C. Bhat, B. Vachhani, and S. K. Kopparapu, “Automatic Assess- ment of Dysarthria Severity Level Using Audio Descriptors,” in Proc. IEEE ICASSP, New Orleans, 2017. [12] M. Geng, Z. Jin, T. Wang, S. Hu, J. Deng, M. Cui, G. Li, J. Yu, X. Xie, and X. Liu, “Use of Speech Impairment Severity for Dysarthric Speech Recognition,” in Proc. Interspeech, Dublin, 2023. [13] Y. Jeon, S. Im, Y. Kim, and G. G. Lee, “Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning,” in Proc. Interspeech, Rotterdam, 2025. [14] X. Chen, Y. Wang, X. Wu, D. Wang, Z. Wu, X. Liu, and H. Meng, “Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction,” in Proc. IEEE ICASSP, Seoul, 2024. [15] Y. Wang, X. Wu, D. Wang, L. Meng, and H. Meng, “UNIT- DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization,” in Proc. IEEE ICASSP, Seoul, 2024. [16] F. Rudzicz, A. K. Namasivayam, and T. Wolff, “The TORGO database of acoustic and articulatory speech from speakers with dysarthria,” Lang. Resour. Eval., vol. 46, no. 4, p. 523–541, 2012. [17] H. Kim, M. Hasegawa-Johnson, A. Perlman, J. R. Gunderson, T. S. Huang, K. L. Watkin, S. Frame et al., “Dysarthric Speech Database for Universal Access Research,” in Proc. Interspeech, Brisbane, 2008. [18] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,” in Proc. NeurIPS, virtual, 2020. [19] W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdi- nov, and A. Mohamed, “HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, vol. 29, p. 3451–3460, 10 2021. [20] Y. Yang, J. Zhuo, Z. Jin, Z. Ma, X. Yang, Z. Yao, L. Guo, W. Kang, F. Kuang, L. Lin et al., “k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning,” in Proc. IEEE ICME, Nantes, 2025. [21] E. J. Yeo, K. Choi, S. Kim, and M. Chung, “Automatic Sever- ity Classification of Dysarthric Speech by Using Self-Supervised Model with Multi-Task Learning,” in Proc. IEEE ICASSP, Rhodes Island, 2023. [22] F. Javanmardi, S. Tirronen, M. Kodali, S. R. Kadiri, and P. Alku, “Wav2vec-Based Detection and Severity Level Classification of Dysarthria From Speech,” in Proc. IEEE ICASSP, Rhodes Island, 2023. [23] M. Geng, X. Xie, S. Liu, J. Yu, S. Hu, X. Liu, and H. Meng, “Investigation of Data Augmentation Techniques for Disordered Speech Recognition,” in Proc. Interspeech, Shanghai, 2020. [24] L. Stumpf, B. Kadirvelu, and A. A. Faisal, “Multilingual Speaker- Invariant Dysarthria Severity Assessment Using Adversarial Do- main Adaptation and Self-Supervised Learning,” in Proc. IEEE ICASSP, Hyderabad, 2025. [25] H. Wang, T. Thebaud, J. Villalba, M. Sydnor, B. Lammers, N. Dehak, and L. Moro-Velazquez, “DuTa-VC: A Duration-aware Typical-to-atypical Voice Conversion Approach with Diffusion Probabilistic Model,” in Proc. Interspeech, Dublin, 2023. [26] Z. Jin, M. Geng, X. Xie, J. Yu, S. Liu, X. Liu, and H. Meng, “Adversarial Data Augmentation for Disordered Speech Recogni- tion,” in Proc. Interspeech, Brno, 2021. [27] Z. Jin, X. Xie, M. Geng, T. Wang, S. Hu, J. Deng, G. Li, and X. Liu, “Adversarial Data Augmentation Using VAE-GAN for Disordered Speech Recognition,” in Proc. IEEE ICASSP, Rhodes Island, 2023. [28] C. Bhat, A. Panda, and H. Strik, “Improved ASR Performance for Dysarthric Speech Using Two-stage Data Augmentation,” in Proc. Interspeech, Incheon, 2022. [29] E. Hermann and M. Magimai-Doss, “Few-shot Dysarthric Speech Recognition with Text-to-Speech Data Augmentation,” in Proc. Interspeech, Dublin, 2023. [30] W.-Z. Leung, M. Cross, A. Ragni, and S. Goetze, “Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis,” in Proc. Interspeech, Kos, 2024. [31] M. Kim, M. Han, S. Hong, and M.-w. Koo, “Data Augmenta- tion using Speech Synthesis for Speaker-Independent Dysarthria Severity Classification,” in Proc. Interspeech, Rotterdam, 2025. [32] W.-C. Huang, E. Cooper, Y. Tsao, H.-M. Wang, T. Toda, and J. Ya- magishi, “The VoiceMOS Challenge 2022,” in Proc. Interspeech, Incheon, 2022. [33] C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H.-M. Wang, “MOSNet: Deep Learning-Based Ob- jective Assessment for Voice Conversion,” in Proc. Interspeech, Graz, 2019. [34] Y. Yang, H. Wang, B. Han, S. Liu, J. Li, Y. Qin, and X. Chen, “Position: Towards Responsible Evaluation for Text-to-Speech,” in Proc. ICML, Seoul, 2026. [35] Y. Yang, B. Han, H. Wang, W. Wang, Z. Ma, L. Zhou, Z. Jin, G. Yang, T. Wang, X. Tan et al., “Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training,” in Proc. ACL, San Diego, 2026. [36] Y. Yang, B. Han, H. Wang, L. Zhou, W. Wang, M. Cui, X. Tan, and X. Chen, “Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration,” in Proc. IEEE ICASSP, Barcelona, 2026. [37] S. Wang, W. Yu, Y. Yang, C. Tang, Y. Li, J. Zhuang, X. Chen, X. Tian, J. Zhang, G. Sun et al., “Enabling Auditory Large Lan- guage Models for Automatic Speech Quality Evaluation,” in Proc. ICASSP, Hyderabad, 2025. [38] C. Chen, Y. Hu, S. Wang, H. Wang, Z. Chen, C. Zhang, C.-H. H. Yang, and E. Chng, “Audio Large Language Models Can Be De- scriptive Speech Quality Evaluators,” in Proc. ICLR, Singapore, 2025. [39] S. Wang, Z. Jin, C. Tang, Q. Li, B. Li, C. Chen, Y. Hu, W. Yu, Y. Li, J. Zhuang et al., “Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking,” arXiv preprint arXiv:2511.01299, 2025. [40] X. Zheng, B. Phukon, J. Na, E. Cutrell, K. Han, M. Hasegawa- Johnson, P.-P. Jiang, A. Kuila, C. Lea, B. MacDonald et al., “The Interspeech 2025 Speech Accessibility Project Challenge,” in Proc. Interspeech, Rotterdam, 2025. [41] S. Wang, W. Yu, X. Chen, X. Tian, J. Zhang, L. Lu, Y. Tsao, J. Yamagishi, Y. Wang, and C. Zhang, “QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions,” in Proc. ACL, Vienna, 2025. [42] E. Cooper and J. Yamagishi, “How do Voices from Past Speech Synthesis Challenges Compare Today?” in Proc. ISCA SSW, Bu- dapest, 2021. [43] G. Mittag, B. Naderi, A. Chehadi, and S. M ̈ oller, “NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets,” in Proc. Inter- speech, Brno, 2021. [44] G. Chen, S. Chai, G. Wang, J. Du, W.-Q. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang et al., “GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio,” in Proc. Interspeech, Brno, 2021. [45] T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-SaruLab System for Voice- MOS Challenge 2022,” in Proc. Interspeech, Incheon, 2022. [46] E. Cooper, W.-C. Huang, T. Toda, and J. Yamagishi, “Gener- alization Ability of MOS Prediction Networks,” in Proc. IEEE ICASSP, Singapore, 2022. [47] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- riSpeech: An ASR Corpus Based on Public Domain Audio Books,” in Proc. IEEE ICASSP, South Brisbane, 2015. [48] J. Kahn, M. Riviere, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazar ́ e, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen et al., “Libri-Light: A Benchmark for ASR with Limited or No Supervision,” in Proc. IEEE ICASSP, Barcelona, 2020. [49] R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Hen- retty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common Voice: A Massively-Multilingual Speech Corpus,” in Proc. LREC, Marseille, 2020. [50] J. J. Godfrey, E. C. Holliman, and J. McDaniel, “SWITCH- BOARD: telephone speech corpus for research and development,” in Proc. IEEE ICASSP, San Francisco, 1992. [51] C. Cieri, D. Miller, and K. Walker, “The Fisher Corpus: a Re- source for the Next Generations of Speech-to-Text,” in Proc. LREC’04, Lisbon, 2004.