Paper deep dive
Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models
Paban Sapkota, Hemant Kumar Kathania, Mikko Kurimo, Sudarsana Reddy Kadiri, Shrikanth Narayanan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 6/21/2026, 4:35:53 AM
Summary
This paper presents a systematic investigation into optimizing Dysarthric Automatic Speech Recognition (DyASR) by evaluating various combinations of spectral features and acoustic models. Using the TORGO database, the researchers tested features such as FBANKs, MFCCs, and PLPCCs, often in combination with Pitch features, across models including HMM-GMM, SGMM, DNN, TDNN-LSTM, and F-TDNN. A key finding was that the F-TDNN model's performance could be significantly enhanced by optimizing the number of overlapping frames (finding 20 frames to be optimal) and selecting specific feature sets: FBANKs+MFCCs+Pitch for isolated word recognition and MFCCs alone for sentence recognition. The study achieved relative improvements of approximately 4.6% to 4.7% over previous research in both isolated word and sentence recognition tasks.
Entities (11)
Relation Signals (5)
FBANKs+MFCCs+Pitch → improves → Isolated Word Recognition
confidence 100% · the most effective feature combination for isolated word recognition was achieved by combining FBANKs with MFCCs and Pitch features.
MFCCs → improves → Sentence Recognition
confidence 100% · for the sentence recognition task, MFCCs emerged as the optimal feature set
TORGO → usedfor → Dysarthric Automatic Speech Recognition
confidence 100% · Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art F-TDNN model
20 frames → optimizes → F-TDNN
confidence 95% · 20 frames of overlap yielded the optimal performance for the dysarthric speech recognition task... in the F-TDNN architecture.
F-TDNN → uses → MFCCs
confidence 90% · In this investigation, we selected MFCCs as the baseline feature set [for F-TDNN].
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, offering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65\% relative improvement in isolated word recognition and a 4.63\% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improvement effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks.
Tags
Links
- Source: https://arxiv.org/abs/2606.19793v1
- Canonical: https://arxiv.org/abs/2606.19793v1
Trouble viewing inline? Open PDF directly →
Full Text
29,287 characters extracted from source content.
Expand or collapse full text
Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models Paban Sapkota ∗ , Hemant Kumar Kathania ∗ , Mikko Kurimo † , Sudarsana Reddy Kadiri ‡ , and Shrikanth Narayanan ‡ ∗ Department of Electronics and Communication Engineering,National Institute of Technology Sikkim, India. Emails: phec230006@nitsikkim.ac.in, hemant.ece@nitsikkim.ac.in † Department of Information and Communications Engineering, Aalto University, Finland. Email: mikko.kurimo@aalto.fi ‡ Signal Analysis and Interpretation Laboratory (SAIL), University of Southern California, Los Angeles, USA. Emails: skadiri@usc.edu, shri@usc.edu Abstract—The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, of- fering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time DelayNeural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65% relative improvement in isolated word recognition and a 4.63% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improve- ment effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks. Index Terms—Dysarthria, speech variability, isolated word recognition, sentence recognition, F-TDNN, overlapping frames. I. INTRODUCTION Dysarthria is a speech disorder characterized by difficulties in controlling the neuro-motor functions responsible for speech articulation [1]. Automatic Speech Recognition (ASR) systems are undergoing rapid development and are progressing toward the point where they have the potential to achieve human- level speech perception capabilities. It is essential to note that this progress primarily pertains to ASR systems designed for typical speech from control or healthy speakers. While a considerable amount of effort has been dedicated to the field of Dysarthric Automatic Speech Recognition (DyASR) [2], there remains a significant need to reach an acceptable level of ASR performance. Dysarthric speech differs from typical speech in terms of acoustic properties and pronunciation distinctions. Challenges in accurately articulating sounds and forming words resultin a notable lack of clarity and comprehensibility, leading to poor intelligibility. In DyASR, there are significant intra-speaker disparities and even greater inter-speaker disparities, constitut- ing two primary challenges. Collecting dysarthric speech data introduces another major dimension of challenge, primarily due to privacy concerns. Advancements in the DyASR system have the potential to substantially enhance individuals’ ability to effectively convey their intended communication, thereby improving overall comprehension. In the literature, researchers have explored various strategies to enhance the performance of DyASR systems. To address the limited availability of atypical speech data, some studies, like [3], have investigated the use of artificial data generated through non-linear speech tempo modification, demonstrating performance improvements. Additionally, domain adaptation of pre-trained ASR models, as discussed in [4], has shown promise in mitigating the scarcity of sufficient atypical data. In another study, presented in [5], it was demonstrated that improving speech quality through denoising techniques ap- plied to atypical speech can transform disordered speech into a form more akin to normal speech. The study in [6] compared traditional ASR with Neural Network-based ASR using features like Filterbanks (FBANKs) and Mel Frequency Cepstral Coefficients (MFCCs) for dysarthric speech. An investigation into raw magnitude spectra-based multi-stream acoustic modeling, which exhibited notable performance gains for dysarthric speech, was documented in [7]. In a recent study, researchers explored the combination of raw phase-based representations with raw magnitude spectra, as described in [8], employing single and multi-stream architectures composed of cascades of convolutional, recurrent, and fully-connected layers for acoustic modeling. Furthermore, a Deep Neural Network (DNN) model utilizing MFCC-based i-vectors was found to outperform other models, including convolutional neural networks, gated recurrent units, and long short-term memory networks, as reported in [9]. In [10], Hermann and Doss demonstrated that the Lattice Free - Maximum Mu- tual Information (LF-MMI) setup of F-TDNN surpassed the Time Delay Neural Network with Long Short Term Memory (TDNN-LSTM) model in recognizing dysarthric speech when using MFCC features. To the best of our knowledge, no studies have been con- arXiv:2606.19793v1 [eess.AS] 18 Jun 2026 ducted to determine the selection of acoustic features and acoustic modeling techniques for dysarthric speech. In this study, we propose a systematic investigation of DyASR sys- tems using spectral features with various acoustic models.The main contributions of our study includes: •A systematic investigation of spectral features and their combinations for acoustic models in the DyASR study. •An assessment of the incorporation of pitch features in different systems. •An evaluation of the impact of changing the number of overlapping frames in training the F-TDNN acoustic model. •The introduction of individualized speaker performance assessment and severity labeling for optimal configura- tion. I. EXPERIMENTAL SETUP Figure 1 depicts our experimental setup, which includes various features and acoustic models under investigation.We employed four types of acoustic features: FBANKs, MFCCs, Perceptual Linear Prediction Cepstral Coefficients (PLPCCs), as well as combinations with Pitch features. These features were evaluated using five different acoustic models: Hid- den Markov Model with Gaussian Mixture Model (HMM- GMM), Sub-space Gaussian Mixture Model (SGMM), Time Delay Neural Network with Long Short Term Memory layers (TDNN-LSTM), and Factorized-TDNN (F-TDNN). All experiments were conducted using the Kaldi speech processing toolkit [11], with the recipe available in [10]. It’s worth noting that increasing the frame shift duration for dysarthric speakers during feature extraction has demonstrated improved performance [6]. In this study, we maintained the same frame shift durations, i.e., 15 ms for dysarthric speakers and 10ms for control speakers. Fig. 1. Experimental setup with various spectral features and acoustic models investigated in this study. A. TORGO Dysarthric Speech Database We are utilizing the publicly available TORGO database [12], a crucial resource for dysarthric speech research. The freely accessible TORGO database consists of a total of 15 hours of speech, comprising 8 speakers with dysarthria contributing approximately 6 hours of speech data and 7 control speakers contributing approximately 9 hours of speech data. In total, there are 16,394 utterances, and the statistical distribution of unique utterances showed that out of 971 unique utterances, 615 utterances are single-worded and 356 utterances are sentenced. B. Selection of features For preliminary investigation, we selected three spectral features [13], [14]. They are: FBANKs [15], MFCCs [16], Per- ceptual Linear Prediction Cepstral Coefficients (PLPCCs) [17]. Prior research on dysarthric speech has revealed noticeable and consistent differences in pitch levels between dysarthricspeech and typical speech, particularly during spontaneous conversa- tions [18]. Consequently, we conducted an investigation into the impact of appending pitch features [19] to each of the three features, resulting in an additional three sets of features. Initially, we trained the F-TDNN system using FBANKs, MFCCs, and PLPCCs features independently, without com- bining them. We adhered to standard practices, selecting 40- dimensional high-resolution features for both FBANKs and MFCCs, as per the F-TDNN model’s requirements. When training with PLPCCs, we retained the original 13-dimensional configuration. Furthermore, in our quest for an optimal feature combi- nation, we opted for 40 high-resolution FBANKs features to complement the 13-dimensional features of MFCCs and PLPCCs with 3-dimensional Pitch features. This choice was made deliberately to strike a balance, ensuring that our feature set wouldn’t become excessively large. This considerationis especially crucial when working with models trained on lim- ited datasets, as overly complex features can pose challenges. C. Acoustic model configurations We investigated five distinct acoustic models to address the intricacies of DyASR. First, the HMM-GMM model utilized triphone modeling with context information and Gaussian Mixture Models comprising 400 components. This approach was complemented by forced alignment during training and subsequent decoding. Moving forward, the SGMM model fur- ther improved modeling capabilities through Subspace Gaus- sian Mixture Models [20], featuring 8000 leaves and 19000 sub-states. The DNN model [11] was configured with five hidden layers, employing mini-batch training with 5000 mix- ups and parameter tuning across 20 epochs. The Time-Delay Neural Network with Long Short-Term Memory (TDNN- LSTM) [11] model introduced temporal dependencies with LSTM layers and employed chunk-based training to enhance performance. The F-TDNN model was designed with a structure resem- bling that of TDNN-LSTM, with a particular focus on the in- tegration of online i-Vectors [21]. The corpus size is relatively limited, indicating a small-scale dataset. Furthermore, notice- able variations in speech tempo have been observed, primarily within the subset of dysarthric speakers. To address this, we applied a speed perturbation data augmentation technique [22]. This involved modifying temporal rates to 0.9 and 1.1 relative to the native speaking rate. Additionally, we employed variable frames-per-chunk during training. The training and evaluation followed a leave-one-out ap- proach, where fourteen speakers were included in the training set, and one speaker was reserved for evaluation. Two distinct language models (LMs) were employed based on the specific task requirements. For the isolated word recognition task, a unigram (1-gram) LM was utilized, while a bigram (2- gram) LM was applied for the sentence recognition task. The decoding grammar was constrained to generate single- word outputs for the 1-gram LM. This constraint has been previously shown to be effective in completely mitigating insertion errors in prior research efforts [10]. In the caseof the F-TDNN model, we conducted experiments varying the degree of overlapping frames between training example chunks. D. Overlapping frames between consecutive chunks Training a chain-based neural network acoustic model in Kaldi requires a focus on controlling contextual information [23]. We conducted experiments involving the manipulation of overlapping frames, ranging from 0 to 40 frames, to influ- ence context consideration during the training process. This parameter adjustment has notable implications for memory consumption and the model’s capacity to capture various aspects of speech variability, such as pronunciation, speaking rate, and acoustic characteristics. The incorporation of overlap- ping frames is crucial for mitigating speech variability within the chain-based model. In this model, each training instance consists of a chunk of frames containing extracted acoustic attributes. I. RESULTSANDDISCUSSION As discussed in Section I-C, for the dysarthric speech recognition task, we have divided it into isolated word and sentence recognition tasks based on their distinct applicability, aligning with previous research efforts [10]. In the dysarthric isolated word recognition task, Figure 2 illustrates the experi- mental results pertaining to three spectral features: FBANKs, MFCCs, and PLPCCs, along with their combinations with pitch features. These results were evaluated using the first four acoustic models (AMs) discussed in Section I-C, which include GMM-HMM, SGMM, DNN and TDNN-LSTM. Figure 2 provides valuable insights into the performance of different feature sets with various acoustic models. For the HMM-GMM acoustic model, PLPCCs features delivered the best performance with a Word Error Rate (WER) of 50.6%, closely followed by MFCCs features at 51.2%. Interestingly, the addition of Pitch features did not yield significant improve- ments for PLPCCs; however, for MFCCs showed a slight 0.7% relative improvement. The trend continued with the SGMM model, where PLPCCs features outperformed others with a WER of 46.8% without pitch features. Again, the combination of Pitch features followed a similar pattern, resulting in a noteworthy 4% reduction in WER for MFCCs, making it the top-performing feature for the SGMM model at 46.2%. Moving to the DNN model, PLPCCs features led the way with a WER of 50.6%, closely trailed by MFCCs with a 50.8% WER. Meanwhile, for the TDNN-LSTM model, MFCCs emerged as the clear winner with a WER of 43%. Importantly, Pitch features did not show significant improvements in the performance of neural network-based DNN and TDNN-LSTM AMs. Overall, the choice of features displayed distinctive trends across various acoustic models: HMM-GMM favored PLPCCs, SGMM performed best with MFCCs (with Pitch), DNN had a close competition between PLPCCs and MFCCs, and MFCCs excelled for the TDNN-LSTM model. Fig. 2. Isolated word recognition performances (in Word Error Rates (WERs)) for different acoustic models on dysarthric speech with various feature combination sets. Fig. 3. Sentence recognition performances (in Word Error Rates (WERs)) for different acoustic models on dysarthric speech with various feature combination sets. Figure 3 provides a comprehensive view of the results for sentence recognition task. The HMM-GMM model excelled with MFCCs features, achieving a WER of 47.4%. Notably, PLPCCs also performed impressively with a competitive score of 47.5%, when evaluated without Pitch features. However, the introduction of Pitch features resulted in a notable 2.9% improvement in the performance of PLPCCs, making it the top-performing feature set with a WER of 46.1%. When Pitch features were not combined, the SGMM model deliv- ered its best performance with MFCCs features, achieving a WER of 40.8%, closely followed by PLPCCs with a WER of 41.1%. Intriguingly, with the inclusion of Pitch features, the performance of PLPCCs showed a remarkable relative improvement of 4.9%, matching the performance of MFCCs features (39.1%) with Pitch. Moving to the DNN model, it demonstrated exceptional performance with MFCCs features, whether Pitch was included (WER of 35.7%) or not (WER of 35.9%). This result highlighted the robustness of MFCCs in this context. Interestingly, for the TDNN-LSTM model, the introduction of Pitch features did not significantly impact performance when used alongside PLPCCs. However, a slight relative improvement of 1.3% was observed when Pitch fea- tures were added to the MFCCs features, resulting in the best WER of 36.9% for the TDNN-LSTM model. Again in overall, the choice of features displayed distinct patterns across various acoustic models: HMM-GMM excelled with PLPCCs+Pitch, SGMM favored MFCCs+Pitch and PLPCCs+Pitch, DNN per- formed exceptionally with MFCCs+Pitch or without Pitch, and TDNN-LSTM showcased the best performance with MFCCs+Pitch. In the case of the F-TDNN acoustic model, we conducted a comprehensive analysis encompassing both dysarthric (dys) and control (ctl) speakers, addressing both isolated word and sentence recognition tasks. Our exploration into frame overlap commenced with Kaldi’s F-TDNN recipe, which, by default, assigns a null value for overlapping frames. In this investigation, we selected MFCCs as the baseline feature set. The baseline Word Error Rates (WERs) are conveniently summarized in Table I. TABLE I BASELINEWERS OFF-TDNN ACOUSTIC MODEL WITH DEFAULT VALUE OF0OVERLAPPING FRAMES,ON ISOLATED WORD AND SENTENCE RECOGNITION TASKS. FeatureIsol(dys)Sent(dys)Isol(ctl)Sent(ctl) MFCCs54.438.543.213.0 Subsequently, we extended our experimentation with the F- TDNN model, systematically varying the frame overlap from 0 to 40 frames, with a step size of 10 frames. The insights derived from Figure 4 reveal that 20 frames of overlap yielded the optimal performance for the dysarthric speech recognition task, whether it involved isolated words or sentences. Further- more, we also explored the effects of frame overlap variations on control speech. Remarkably, as illustrated in Figure 4, a30- frame overlap demonstrated superior performance comparedto other values in the case of control speech. However, it is essen- tial to emphasize that our primary focus remains on dysarthric speech recognition. As a result, we selected the configuration with 20 frames of overlap between example chunks for further experimentation within the F-TDNN architecture. Fig. 4. Effect of variation in overlapping frames on F-TDNN setup with MFCCs features. We thoroughly examined various feature combinations for the F-TDNN acoustic model, and the results are presented in Table I. Analyzing the findings from Table I for the dysarthric speech recognition task, it becomes evident that the most effective feature combination for isolated word recogni- tion was achieved by combining FBANKs with MFCCs and Pitch features. This combination exhibited a remarkable 4.7% relative improvement compared to [10], achieving a Word Er- ror Rate (WER) of 41%. Similarly, for the sentence recognition task, MFCCs emerged as the optimal feature set, showcasing a noteworthy improvement with a WER of 24.7%, representing a relative gain of 4.6%. Overall, our feature selection strategy for the F-TDNN model consisted of FBANKs+MFCCs+Pitch for isolated word recognition and MFCCs alone for sentence recognition. TABLE I COMPARISON OFWERS USING THE FRAME OVERLAP OF20IN THE F-TDNNACOUSTIC MODEL WITH THELF-MMISETUP[10]. THIS EVALUATION IS CONDUCTED FOR VARIOUS COMBINATIONS OF ACOUSTIC FEATURES,FOR BOTHDYSARTHRIC ANDCONTROL SPEECH DATASETS. FeaturesIsol (DYS)Sent(DYS)Isol (CTL)Sent(CTL) MFCCs (LF-MMI setup) [10]43.025.922.07.9 FBANKs43.727.324.16.7 MFCCs42.424.722.75.5 PLPCCs42.724.822.03.7 FBANKs+Pitch43.327.917.93.7 MFCCs+Pitch44.226.324.87.0 PLPCCs+Pitch42.425.822.05.6 FBANKs+MFCCs+Pitch41.026.021.75.2 FBANKs+PLPCCs+Pitch41.826.823.87.9 MFCCs+PLPCCs+Pitch44.327.624.07.5 FBANKs+MFCCs+PLPCCs+Pitch41.925.523.26.0 It is noteworthy that our approach, featuring 20 frames of overlap, has yielded superior results compared to the previous study [10] across various feature sets. Notably, when we combined FBANKs, MFCCs, and Pitch features, we observed a substantial 4.7% relative improvement in isolated word recognition tasks. Furthermore, in the context of sentence recognition tasks, the inclusion of MFCCs led to a remarkable 4.6% relative improvement. While our primary focus centered on dysarthric speakers, we also extended our evaluation to control speakers. The results, as presented in Table I, indicate that when the F-TDNN system was tested on control speakers, the WER demonstrated a noteworthy 18.6% relative reduction, particularly in the case of isolated word recognition, where the combination of FBANKs and Pitch features excelled. For the recognition of sentence utterances among control speakers, once again, the combination of FBANKs with Pitch features delivered the best performance, showcasing a significant 53.2% relative improvement. These outcomes underscore the adaptability of our approach, allowing us to select suitable feature and model combinations based on specific task requirements for further enhancing our study. TABLE I SPEAKER-SPECIFIC PERFORMANCE ANALYSIS CONDUCTED ON DYSARTHRIC SPEAKERS USINGF-TDNN ACOUSTIC MODEL FOR TWO TASKS: ISOLATED WORD ANDSENTENCE RECOGNITION. THE FIRST AND SECOND ROWS DISPLAYS SPEAKER-WISEWERS WITH THE BEST FEATURES FOR ISOLATED WORD RECOGNITION AND SENTENCE RECOGNITION TASKS. FeaturesTaskF01F03F04M01M02M03M04M05 FBANKs+MFCCs +Pitch Isol55.3239.7212.8550.6253.959.0259.6846.68 Sent50.5410.692.0830.0526.793.7038.6831.67 MFCCs Isol65.9638.8514.0652.4157.738.0357.8843.99 Sent47.557.601.2528.4231.302.6335.9925.82 In Table I, we present speaker-specific WERs for the two best-performing feature sets in the F-TDNN model, which achieved the highest scores in isolated and sentence recog- nition tasks, respectively. Among the speakers, those denoted as F04 and M03 exhibited the lowest WERs, warranting the classification of ’mildly severe.’ For speakers F03 and M05, their WERs fell within the range of 30-50%, leading to their categorization as ’moderately severe.’ Speakers with WERs exceeding 50% have been classified as ’highly severe.’ It’s important to note that these severity labels are derived from the WERs obtained in the isolated word recognition task, providing valuable insights into the individual performance characteristics. IV. CONCLUSIONS In this study, we investigated various acoustic features and their combinations with four different acoustic models (HMM-GMM, SGMM, DNN, and TDNN-LSTM) for DyASR, encompassing both isolated words and sentences. Among these models, TDNN-LSTM achieved the lowest WER of 44.2% in the isolated word recognition task, while DNN outperformed the others with a WER of 35.7% in sentence recognition. Our study underscores the importance of selecting suitable acoustic features tailored to specific acoustic models. Notably, pitch features played a vital role in enhancing sentence recognition, particularly for dysarthric speech. We also optimized a chain-based F-TDNN model by ad- justing the number of overlapping frames between consecu- tive training chunks, resulting in improved performance. Our focus on dysarthric speech recognition led us to choose an optimal 20-frame overlap between training examples, which outperformed previous research [10]. For isolated word recog- nition, the combination of FBANKs with MFCCs and Pitch achieved a WER of 41%, representing a relative improvement of 4.63%. In sentence recognition, MFCCs delivered the best performance with a WER of 24.7%, showcasing a relative improvement of 4.65%. Importantly, these improvements ex- tended to control speech datasets. Future work aims to further enhance DyASR by addressing speech variations, employing speech enhancement methods, and applying augmentation techniques to compensate for the limited dataset size. The preliminary investigation was conducted using a Wav2Vec2-based end-to-end ASR system, where various data augmentation techniques were explored for fine-tuning [24]. REFERENCES [1] Ray D Kent, Gary Weismer, Jane F Kent, Houri K Vorperian, and Joseph R Duffy, “Acoustic studies of dysarthric speech: Methods, progress, and potential,”Journal of communication disorders, vol. 32, no. 3, p. 141–186, 1999. [2] Kristin Rosen and Sasha Yampolsky, “Automatic speech recognition and a review of its functioning with dysarthric speech,”Augmentative and Alternative Communication, vol. 16, no. 1, p. 48–60, 2000. [3] Feifei Xiong, Jon Barker, and Heidi Christensen, “Phonetic analysis of dysarthric speech tempo and applications to robust personalised dysarthric speech recognition,” inInternational Conference on Acous- tics, Speech and Signal Processing (ICASSP). IEEE, 2019, p. 5836– 5840. [4] Shujie Hu, Xurong Xie, Zengrui Jin, Mengzhe Geng, Yi Wang, Mingyu Cui, Jiajun Deng, Xunying Liu, and Helen Meng, “Exploring self- supervised pre-trained asr models for dysarthric and elderly speech recognition,” inInternational Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, p. 1–5. [5] Chitralekha Bhat, Biswajit Das, Bhavik Vachhani, and Sunil Kumar Kopparapu, “Dysarthric speech recognition using time-delay neural network based denoising autoencoder.,” inINTERSPEECH, 2018, p. 451–455. [6] Cristina Espana-Bonet and Jos ́e AR Fonollosa, “Automatic speech recognition with deep neural networks for impaired speech,” inAdvances in Speech and Language Technologies for Iberian Languages:Third International Conference, IberSPEECH, Lisbon, Portugal,November 23-25. Springer, 2016, p. 97–107. [7] Zhengjun Yue, Erfan Loweimi, Heidi Christensen, Jon Barker, and Zoran Cvetkovic, “Acoustic modelling from raw source and filter components for dysarthric speech recognition,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, p. 2968–2980, 2022. [8] Zhengjun Yue, Erfan Loweimi, and Zoran Cvetkovic, “Dysarthric speech recognition, detection and classification using raw phase and magnitude spectra,” inProceedings of INTERSPEECH. ISCA-INST SPEECH COMMUNICATION ASSOC, 2023. [9] Amlu Anna Joshy and Rajeev Rajan, “Automated dysarthriaseverity classification: A study on acoustic features and deep learning tech- niques,”IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 30, p. 1147–1157, 2022. [10] Enno Hermann and Mathew Magimai Doss, “Dysarthric speech recog- nition with lattice-free mmi,” inInternational Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, p. 6109–6113. [11] Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, “The kaldi speech recognition toolkit,” inIEEE Workshop on Auto- matic Speech Recognition and Understanding. Dec. 2011, IEEE Signal Processing Society, IEEE Catalog No.: CFP11SRW-USB. [12] Frank Rudzicz, Aravind Kumar Namasivayam, and Talya Wolff, “The torgo database of acoustic and articulatory speech from speakers with dysarthria,”Language Resources and Evaluation, vol. 46, p. 523–541, 2012. [13] Bassam Ali Al-Qatab and Mumtaz Begum Mustafa, “Classification of dysarthric speech according to the severity of impairment:an analysis of acoustic features,”IEEE Access, vol. 9, p. 18183–18194, 2021. [14] Achraf Benba, Abdelilah Jilbab, Ahmed Hammouch, and Sara Sand- abad, “Voiceprints analysis using mfcc and svm for detecting patients with parkinson’s disease,” inInternational conference on electrical and information technologies (ICEIT). IEEE, 2015, p. 300–304. [15] Laxmi Priya Sahu and Gayadhar Pradhan, “Significance offilter- bank structure for capturing dysarthric information through cepstral coefficients,” inInternational Conference on Signal Processing and Communications (SPCOM). IEEE, 2022, p. 1–5. [16] Mittapalle Kiran Reddy and Paavo Alku, “A comparison ofcepstral features in the detection of pathological voices by varyingthe input and filterbank of the cepstrum computation,”IEEE Access, vol. 9, p. 135953–135963, 2021. [17] Veena Karjigi, N Sreedevi, et al., “Speech intelligibility assessment of dysarthria using fisher vector encoding,”Computer Speech & Language, vol. 77, p. 101411, 2023. [18] K-J Schlenck, R Bettrich, and K Willmes, “Aspects of disturbed prosody in dysarthria,”Clinical linguistics & phonetics, vol. 7, no. 2, p. 119– 128, 1993. [19] Pegah Ghahremani, Bagher BabaAli, Daniel Povey, Korbinian Riedham- mer, Jan Trmal, and Sanjeev Khudanpur, “A pitch extraction algorithm tuned for automatic speech recognition,” ininternational conference on acoustics, speech and signal processing (ICASSP). IEEE, 2014, p. 2494–2498. [20] Daniel Povey, Luk ́aˇs Burget, Mohit Agarwal, Pinar Akyazi, Feng Kai, Arnab Ghoshal, Ondˇrej Glembek, Nagendra Goel, Martin Karafi ́at, Ariya Rastrow, et al., “The subspace gaussian mixture model—a structured model for speech recognition,”Computer Speech & Language, vol. 25, no. 2, p. 404–439, 2011. [21] Daniel Povey, Gaofeng Cheng, Yiming Wang, Ke Li, HainanXu, Mahsa Yarmohammadi, and Sanjeev Khudanpur, “Semi-orthogonal low-rank matrix factorization for deep neural networks.,” inInterspeech, 2018, p. 3743–3747. [22] Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, “Audio augmentation for speech recognition,” inSixteenth annual conference of the international speech communication association, 2015. [23] Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, “Purely sequence-trained neural networks for asr based on lattice-free mmi.,” inInterspeech, 2016, p. 2751–2755. [24] Paban Sapkota, Hemant Kathania, Sudarsana Reddy Kadiri, and Shrikanth Narayanan, “Improving end-to-end speech recognition for dysarthric speech through in-domain data augmentation,” inAsilomar Conference on Signals, Systems and Computers. IEEE, 2024.