Paper deep dive
Demonstration of Adapt4Me: An Uncertainty-Aware Authoring Environment for Personalizing Automatic Speech Recognition to Non-normative Speech
Niclas Pokel, Yiming Zhao, Pehuén Moure, Yingqiang Gao, Roman Böhringer
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/23/2026, 12:12:02 PM
Summary
Adapt4Me is a web-based, uncertainty-aware authoring environment designed to personalize Automatic Speech Recognition (ASR) for non-normative speech. It utilizes Bayesian active learning, VI-LoRA fine-tuning, and a human-in-the-loop workflow to enable efficient, low-effort model adaptation for individuals with speech impairments, effectively reducing word error rates in data-scarce regimes.
Entities (5)
Relation Signals (3)
Adapt4Me â utilizes â VI-LoRA
confidence 100% · backend personalization using Variational Inference Low-Rank Adaptation (VI-LoRA)
Adapt4Me â addresses â Non-normative Speech
confidence 95% · personalizing Automatic Speech Recognition (ASR) to Non-normative Speech
Adapt4Me â implements â Bayesian Active Learning
confidence 95% · operationalizes Bayesian active learning to enable end-to-end personalization
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Personalizing Automatic Speech Recognition (ASR) for non-normative speech remains challenging because data collection is labor-intensive and model training is technically complex. To address these limitations, we propose Adapt4Me, a web-based decentralized environment that operationalizes Bayesian active learning to enable end-to-end personalization without expert supervision. The app exposes data selection, adaptation, and validation to lay users through a three-stage human-in-the-loop workflow: (1) rapid profiling via greedy phoneme sampling to capture speaker-specific acoustics; (2) backend personalization using Variational Inference Low-Rank Adaptation (VI-LoRA) to enable fast, incremental updates; and (3) continuous improvement, where users guide model refinement by resolving visualized model uncertainty via low-friction top-k corrections. By making epistemic uncertainty explicit, Adapt4Me reframes data efficiency as an interactive design feature rather than a purely algorithmic concern. We show how this enables users to personalize robust ASR models, transforming them from passive data sources into active authors of their own assistive technology.
Tags
Links
- Source: https://arxiv.org/abs/2603.20112v1
- Canonical: https://arxiv.org/abs/2603.20112v1
Trouble viewing inline? Open PDF directly â
Full Text
32,246 characters extracted from source content.
Expand or collapse full text
Demonstration of Adapt4Me: An Uncertainty-Aware Authoring Environment for Personalizing Automatic Speech Recognition to Non-normative Speech Niclas Pokel npokel@ethz.ch 0009-0005-6429-2019 Institute of Neuroinformatics, University of Zurich and ETH ZurichZurichSwitzerland , Yiming Zhao yimizhao@student.ethz.ch Department of Computer Science, ETH ZurichZurichSwitzerland , PehuĂ©n Moure pehuen@ini.ethz.ch 0009-0001-1631-8238 Institute of Neuroinformatics, University of Zurich and ETH ZurichZurichSwitzerland , Yingqiang Gao yingqiang.gao@cl.uzh.ch 0009-0000-0876-621X Department of Computational Linguistics, University of ZurichZurichSwitzerland and Roman Boehringer romaboeh@ethz.ch 0000-0003-2856-3262 Institute of Neuroinformatics, University of Zurich and ETH ZurichZurichSwitzerland (2026) Abstract. Personalizing Automatic Speech Recognition (ASR) for non-normative speech remains challenging because data collection is labor-intensive and model training is technically complex. To address these limitations, we propose Adapt4Me, a web-based decentralized environment that operationalizes Bayesian active learning to enable end-to-end personalization without expert supervision. The app exposes data selection, adaptation, and validation to lay users through a three-stage human-in-the-loop workflow: (1) rapid profiling via greedy phoneme sampling to capture speaker-specific acoustics; (2) backend personalization using Variational Inference Low-Rank Adaptation (VI-LoRA) to enable fast, incremental updates; and (3) continuous improvement, where users guide model refinement by resolving visualized model uncertainty via low-friction top-k corrections. By making epistemic uncertainty explicit, Adapt4Me reframes data efficiency as an interactive design feature rather than a purely algorithmic concern. We show how this enables users to personalize robust ASR models, transforming them from passive data sources into active authors of their own assistive technology. Accessibility, Dysarthria, Active Learning, Human-in-the-Loop, ASR Personalization, Uncertainty Visualization, Low-Friction Annotation â copyright: noneâ copyright: acmlicensedâ journalyear: 2026â doi: X.Xâ conference: CUI â26: Extended Abstracts of the 2026 Conference on Conversational User Interfaces ; July 21stâ24th, 2026; Bremen, Germanyâ isbn: 978-1-4503-X-X/2018/06â ccs: Human-centered computing Accessibility technologiesâ ccs: Human-centered computing Interactive systems and toolsâ ccs: Computing methodologies Active learning settingsâ ccs: Computing methodologies Speech recognition Figure 1. Adapt4Me: an end-to-end, human-in-the-loop ASR personalization system for non-normative speech. â The system prompts the user with high-information sentences for initial recordings; ⥠Word-level model uncertainty is visualized through color-coding, guiding users directly to likely identify and correct transcribing errors; âą Users resolve errors via low-effort top-k best corrections powered by a Bayesian backend, with validated edits continuously fed back into active learning and model fine-tuning. A three-panel screenshot of the web application. The left panel shows a microphone icon and a prompt. The center panel shows transcribed text with the word âPellikanâ highlighted in red. The right panel shows a dropdown menu suggesting âPelicanâ as a correction. 1. Introduction State-of-the-art ASR models like Whisper (Radford et al., 2023) or Wav2Vec (Schneider et al., 2019) have achieved remarkable performance on normative speech, yet they exclude individuals with motor speech disorders (e.g., dysarthria) (Maassen, 2007; Duffy, 2005; Ballati et al., 2018) or structural impairments of the speech apparatus (e.g., Apert syndrome) (Zimmerman et al., 2020). While previous work has explored fine-tuning foundation ASR models for non-normative speech, error rates remain substantially higher than those observed for normative speech (Singh et al., 2024; Kim and Sung, 2023; Liu et al., 2022). This lack of generalization to non-normative speech further reinforces the marginalization of individuals with speech impairments, as their specific communicative needs are rarely addressed by modern speech technologies, thereby limiting their access to the benefits of AI. Nevertheless, although personalizing ASR models for user-specific transcription is technically feasible, it remains largely inaccessible in practice. Conventional model personalization typically requires hours of voice recordings, which is often exhaustive for individuals with speech impairments and places a substantial burden on family members and caregivers (Cronin et al., 2020; Hitchcock et al., 2015; Page and Yorkston, 2022; Sarsenbayeva and et al., 2022). This process becomes even more burdensome when personalized models must be repeatedly optimized with newly recorded data to adapt to the evolving acoustic characteristics of individuals with speech impairments, who are often everyday ASR users without expert knowledge. Users often experience ASR systems as âblack-boxâ transcribers, with little insight into which phonemes cause recognition errors(Kuhn et al., 2025; Glazer et al., 2025). This opacity is largely due to the absence of token-level uncertainty visualizations that reveal the systemâs transcription difficulties, preventing users from providing targeted and effective feedback in subsequent fine-tuning iterations. Previous work in HCI suggests that well-designed uncertainty visualizations enable users to make better-informed decisions and act as effective supervisors of imperfect AI systems (Kay et al., 2016). We argue that effective ASR personalization requires a shift from passive data collection to interactive curating. This shift hinges on exposing the modelâs epistemic uncertainty, its awareness of what it does not yet model reliably, so users can guide personalization in a principled and efficient way. Recent work in Interactive Machine Learning has shown that users value transparency and agency when training assistive tools (Goodman et al., 2025; Kacorri et al., 2017). This shift is particularly critical for implementation of ASR personalization in real-world settings, where minimizing user effort is essential for enabling continuous adaptation to newly recorded speech data, but also for ensuring sustained user engagement. We present Adapt4Me, a web-based, accessible, and low-effort environment that operationalizes Bayesian Active Learning (AL) for personalizing ASR for non-normative speech. By integrating an uncertainty-aware VI-LoRA-based ASR model (Pokel et al., 2025c) into a user-friendly interface, Adapt4Me has the following advantages over existing work (Tobin et al., 2024; Martin et al., 2025): âą Efficient âCold Startâ: A profiling method using Greedy Biphone Coverage (Pokel et al., 2025a) that bootstraps personalization in minutes of training data rather than hours; âą Uncertainty-Guided Feedback: A visual interface that highlights confused phonemes, guiding users to provide personalization data for active learning through direct human-computer interaction; âą Longitudinal Resilience: A lifecycle management feature handling non-linear speech changes (e.g., post-surgery or puberty voice change). Adapt4Me is a proof-of-concept prototype that demonstrates how advances in ASR models can be converted into tangible solutions to improve communication accessibility for individuals with speech impairments. By bridging the gap between technological advances and real-world deployment, Adapt4Me enables efficient and personalized ASR adaptation that can support speech-impaired users over the long term. 2. Motivation Table 1. Comparison of hours of existing speech datasets (recording hours are reduced by redundant microphones). Domain Dataset Hours Normative Common Voice (EN) (Ardila et al., 2020) ⌠3,800 Normative LibriSpeech (EN) (Panayotov et al., 2015) ⌠1,000 Dysarthric SAP (EN) (Hasegawa-Johnson et al., 2024) ⌠415 Dysarthric UA-Speech (EN) (Kim et al., 2008) ⌠9 Dysarthric BF-Sprache (DE) (Pokel et al., 2025a) ⌠2.5 Dysarthric TORGO (EN) (Rudzicz et al., 2012) ⌠3 The rich diversity of speech impairments cannot be addressed simply by collecting more data, as variability across individuals is substantially higher than in normative speech populations (Xiong et al., 2019; Tröger et al., 2024; van Bemmel et al., 2023). This high inter-speaker variance poses significant challenges for training personalized ASR models. As shown in Table 1, available non-normative speech resources are orders of magnitude smaller than their normative counterparts. Consequently, this data-scarcity bottleneck limits the applicability of traditional personalization approaches, which are typically data-intensive and depend on extensive expert annotation, making them impractical for deployment in a home settings (Park et al., 2021). In addition, standard personalization approaches typically fine-tune models on small amounts of user data, making them prone to overfitting and catastrophic forgetting (McCloskey and Cohen, 1989; French, 1999). While Parameter-Efficient Fine-Tuning (PEFT (He et al., 2022)) methods such as LoRA (Hu et al., 2022) partially alleviate these issues, they do not address a more fundamental problem: the lack of an effective and accessible workflow for personalizing ASR models. Without human-in-the-loop (HITL (Wu et al., 2022; Mosqueira-Rey et al., 2023)) guidance, users often expend effort recording redundant data, namely utterances that the model already transcribes well, while failing to gather data that is most informative for further reducing recognition errors. This mismatch between user effort and model improvement underscores the necessity of systems that enable data-efficient personalization while minimizing cognitive and physical burden on end users. To address both data scarcity and the lack of effective personalization workflows in ASR, Adapt4Me is designed following the principle of active learning (Hakkani-TĂŒr et al., 2002; Settles, 2009). Rather than treating users as passive data providers, Adapt4Me positions them as active participants in the personalization loop, a paradigm aligned with interactive machine learning principles (Amershi et al., 2014). In this setting, users inspect transcription outputs, judge their quality, and correct erroneous predictions to directly guide model adaptation. This HITL process enables the system to focus data collection on informative samples that contribute most to performance improvement. Through Adapt4Me, we demonstrate that effective ASR personalization systems must go beyond algorithmic adaptation alone: they must reduce human effort through targeted interaction and provide actionable feedback that supports iterative model improvement in real-world, home-based settings. 3. System Architecture The architecture of Adapt4Me follows the three-step workflow visualized in Figure 2, starting from speech recording to continuous, uncertainty-guided adaptation. Figure 2. The Adapt4Me Personalization Lifecycle. Step 1 (Speech Recording): The system creates a âCold Startâ profile using a minimized set of 500 words selected via Greedy Biphone Coverage and LLM Prompting (Pokel et al., 2025a). Step 2 (Model Personalization): The backend adapts the model. It uses Semantic Re-chaining (SRC) (Pokel et al., 2025a) to generate context, calculates the Phoneme Difficulty Score (PhDScore) to search for highly informative training-samples (Pokel et al., 2025b), and fine-tunes the VI-LoRA adapters (Pokel et al., 2025c) (Top Right). Step 3 (Active Learning): The continuous interaction loop. The model transcribes free speech, highlights uncertain words (e.g., âPylikanâ), and generates context-aware N-best suggestions to minimize annotation effort. Step 1: Speech Profiling. To bootstrap personalization for a new user, we employ a Greedy Biphone Coverage algorithm (Pokel et al., 2025a). The system first selects a minimal set of initial words from the literature to maximize phonetic variance. These words are used as seeds for an LLM prompt to generate semantically and structurally coherent training samples. Step 2: End-to-End Personalization. Once the initial audio is captured, the backend personalizes the model using VI-LoRA (Pokel et al., 2025c). This parameter-efficient method stabilizes fine-tuning on small datasets and, crucially, quantifies epistemic uncertainty (Pokel et al., 2025b). We aggregate this uncertainty to calculate a Phoneme Difficulty Score, identifying the userâs specific articulatory struggles (Pokel et al., 2025b). This diagnosis drives the Semantic Re-chaining engine to synthesize new, targeted sentence prompts rich in these difficult phonemes, constructing a personalized training curriculum for the next phase. Step 3: Active Learning. Users correct the free speech predictions, guided by word-level uncertainty scores. A key challenge is that standard top-k alternatives from ASR models often lack semantic coherence. To address this, we employ a two-pass decoding strategy: a coherent pass, which produces semantically consistent transcriptions, and a variation pass, which selectively re-samples only high-uncertainty words while keeping their surrounding context fixed. 4. Human-Computer Interaction Design To guide user correction, Adapt4Me highlights low-confidence words in model transcriptions (e.g., âPylikanâ in Figure 2) using entropy-based uncertainty scores. These uncertainty scores serve two complementary purposes: (1) as a speech diagnostic tool, offering insights into how the model interprets user speech by revealing words or phonemes that consistently cause transcription errors; and (2) as a task management tool, directing user attention to the most probable errors, thereby eliminating the need to proofread correctly transcribed segments and reducing cognitive load. Motor impairments frequently co-occur with speech impairments in individuals with neurodevelopmental disorders (Lancioni et al., 2025), which can make fine-grained motor actions such as typing corrections difficult and fatiguing. As a result, manually editing ASR outputs can become a significant barrier for such users (Vertanen and Kristensson, 2009). To address this challenge, Adapt4Me replaces typing-based corrections with a context-aware top-k selection mechanism, allowing users to correct errors by simply selecting the appropriate word from a short list of alternatives. Manual insertion remains available as a fallback. By transforming error correction from a typing task into a lightweight selection task, the system substantially reduces physical effort and improves accessibility for users with motor limitations. To support lifelong ASR personalization, Adapt4Me is designed to accommodate the dynamic nature of speech impairments. The system incrementally adapts to gradual changes, such as a childâs speech development, as users continuously contribute corrected samples. In contrast, major physiological changes, including post-surgical recovery (Reddy and Reddy, 2023) or puberty-related voice changes (Pinheiro et al., 2024), may render an existing personalized acoustic model less effective. In such cases, users can re-initiate the system from the initial recording stage to establish a new acoustic baseline, while preserving previously learned semantic and lexical personalization. This design prevents a complete loss of progress and supports long-term, evolving use. When transcribing with Adapt4Me, ten forward passes produce an inference latency of â2â 2 s on a standard GPU (e.g. NVIDIA GeForce RTX 3090). This is acceptable for the home setting, as Adapt4Me was designed primarily to personalize ASR models rather than the real-time user experience. We employ a client-server setup, with a webapplication (Client) running on edge devices like tablets or smartphones, familiar to children. The computationally intensive personalization is offloaded to a secure cloud backend. This ensures that families do not need high-end GPUs for home deployment. Unlike sound-treated clinical environments, home settings are often acoustically noisy. Adapt4Me therefore incorporates a lightweight Signal-to-Noise Ratio (SNR) check that runs locally in the browser to prevent the submission of low-quality recordings, such as speech contaminated by background television noise, that could degrade system performance. In addition, Adapt4Me supports collaborative use by multiple family members. Prompts are presented in large, simple fonts for the speech-impaired child, while top-k suggestions and uncertainty visualizations are positioned to enable parental oversight, facilitating shared interaction and supervision during the personalization process. 5. Envisioned Conference Experience To effectively demonstrate this home-based, collaborative workflow in a conference setting, we will present Adapt4Me via a local web interface deployed on a provided laptop and tablet. Because conference attendees typically possess normative speech, and because live model fine-tuning exceeds feasible booth dwell times, the demonstration is designed as an interactive, simulated authoring scenario. Attendees will step into the role of the âhuman-in-the-loopâ (e.g., acting as the caregiver or user), tasked with curating pre-recorded, anonymized non-normative speech samples. This workflow allows attendees to directly experience the systemâs core HCI contributions: interpreting the token-level uncertainty visualizations and utilizing the low-friction top-k selection mechanism to correct transcription errors without typing. Following the correction phase, attendees will experience a live comparison, contrasting the semantic hallucinations of a baseline Whisper model with the phonetically accurate output of an Adapt4Me-personalized model. The physical setup natively supports the dyadic interaction described above, ensuring a fast, engaging 3-minute flow that accommodates high booth throughput while mitigating the challenges of the noisy conference exhibition hall. 6. Evaluation and Clinical Evidence Figure 3. Longitudinal user study with a teenage child with structural speech impairment. The system reduces word error rate (WER) from a non-functional ⌠70% to usable ⌠25% within under 90 minutes of total interaction. Uncertainty-aware active learning (orange/grey bars) eventually outperforms the full fine-tuning baseline (blue dashed line), demonstrating the effectiveness of VI-LoRA in data-scarce regimes. We deployed Adapt4Me in a home setting for a teenage user with structural speech impairment. As shown in Figure 3, only 75 minutes of cumulative interaction (recording and correction) reduced the Word Error Rate (WER) from a non-functional 70% of a Whisper Large baseline (Radford et al., 2023) to a usable level of approximately 25%. The proposed active learning pipeline with VI-LoRA fine-tuning outperformed a brute-force full-parameter fine-tuning baseline (blue dashed line) while using substantially less data, demonstrating the efficiency of targeting high-uncertainty phonemes. Qualitative analysis further shows that personalization restores communicative intent. Rather than producing semantic hallucinations, the adapted model makes phonetically plausible errors that remain intelligible to humans. For example, when dictating Swiss train stations (âWiedikon, Enge, Thalwil, Baarâ), the baseline hallucinated unrelated sentences, whereas Adapt4Me produced âVidikon, Enne, Talwil, Borgâ. Despite high WER, these phonetic spellings are understandable in context, making the system useful. Prior work (Pokel et al., 2025b) validated model-predicted phoneme difficulties against clinical logopedic assessments. The strong correlation between the modelâs internal uncertainty estimates and therapistsâ evaluations confirms that the visual feedback provided by Adapt4Me is clinically meaningful rather than a mere statistical artifact. 7. Future Work Adapt4Me offers a novel platform for large-scale speech science. By aggregating anonymized model states, specifically the trajectories of uncertainty heatmaps, researchers could perform meta-analysis on speech patterns without manual annotation. This could reveal latent population-level clusters, such as distinguishing between age-related phonetic drift and pathology-driven changes, effectively turning the ASR model into a diagnostic sensor. Acknowledgements.We would like to thank Dr. Corinne Mathys Zulauf for her logopedic consulting, Dr. Samra Hamzic for valuable discussion and Philipp Guldimann for the initial data collection and annotation. References S. Amershi, M. Cakmak, W. B. Knox, and T. Kulesza (2014) Power to the people: the role of humans in interactive machine learning. AI Magazine 35 (4), p. 105â120. External Links: Document, Link, https://onlinelibrary.wiley.com/doi/pdf/10.1609/aimag.v35i4.2513 Cited by: §2. R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber (2020) Common Voice: A Massively-Multilingual Speech Corpus. In Proceedings of the Twelfth Language Resources and Evaluation Conference, N. Calzolari, F. BĂ©chet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, and S. Piperidis (Eds.), Marseille, France, p. 4218â4222 (eng). External Links: Link, ISBN 979-10-95546-34-4 Cited by: Table 1. F. Ballati, F. Corno, and L. De Russis (2018) Assessing virtual assistant capabilities with italian dysarthric speech. In Proceedings of the 20th International ACM SIGACCESS Conference on Computers and Accessibility, p. 93â101. Cited by: §1. P. Cronin, R. Reeve, P. McCabe, R. Viney, and S. Goodall (2020) Academic Achievement and Productivity Losses Associated with Speech, Language and Communication Needs. International Journal of Language & Communication Disorders 55 (5), p. 734â750. External Links: Document Cited by: §1. J. R. Duffy (2005) Motor speech disorders: substrates, differential diagnosis, and management. 2nd edition, Elsevier Mosby, St. Louis, MO. Note: Includes bibliographical references and index. External Links: ISBN 978-0-323-02452-5 Cited by: §1. R. M. French (1999) Catastrophic forgetting in connectionist networks. Trends in Cognitive Sciences 3 (4), p. 128â135. External Links: ISSN 1364-6613, Document, Link Cited by: §2. N. Glazer, Y. Segal-Feldman, H. Segev, A. Shamsian, A. Buchnick, G. Hetz, E. Fetaya, J. Keshet, and A. Navon (2025) Beyond Transcription: Mechanistic Interpretability in ASR. arXiv preprint arXiv:2508.15882. Cited by: §1. S. M. Goodman, E. J. McDonnell, J. E. Froehlich, and L. Findlater (2025) SPECTRA: personalizable sound recognition for deaf and hard of hearing users through interactive machine learning. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Cited by: §1. D. Hakkani-TĂŒr, G. Riccardi, and A. Gorin (2002) Active Learning for Automatic Speech Recognition. In 2002 IEEE international conference on acoustics, speech, and signal processing, Vol. 4, p. IVâ3904. Cited by: §2. M. Hasegawa-Johnson, X. Zheng, H. Kim, C. Mendes, M. Dickinson, E. Hege, C. Zwilling, M. Channell, L. Mattie, H. Hodges, L. Ramig, M. Bellard, M. Shebanek, L. SarÎč, K. Kalgaonkar, D. Frerichs, J. Bigham, L. Findlater, C. Lea, and B. MacDonald (2024) Community-supported shared infrastructure in support of speech accessibility. Journal of Speech, Language, and Hearing Research 67, p. 4162â4175. External Links: Document Cited by: Table 1. J. He, C. Zhou, X. Ma, T. Berg-Kirkpatrick, and G. Neubig (2022) Towards a Unified View of Parameter-Efficient Transfer Learning. In Proceedings of the Tenth International Conference on Learning Representations, Cited by: §2. E. R. Hitchcock, D. Harel, and T. M. Byun (2015) Social, Emotional, and Academic Impact of Residual Speech Errors in School-Aged Children: A Survey Study. Seminars in Speech and Language 36 (4), p. 283â294. External Links: Document Cited by: §1. E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al. (2022) LoRA: Low-rank Adaptation of Large Language Models. ICLR 1 (2), p. 3. Cited by: §2. H. Kacorri, K. M. Kitani, J. P. Bigham, and C. Asakawa (2017) People with visual impairment training personal object recognizers: feasibility and challenges. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, p. 5839â5849. Cited by: §1. M. Kay, T. Kola, J. R. Hullman, and S. A. Munson (2016) When (ish) is my bus? user-centered visualizations of uncertainty in everyday, mobile predictive systems. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems, CHI â16, New York, NY, USA, p. 5092â5103. External Links: ISBN 9781450333627, Link, Document Cited by: §1. H. Kim, M. Hasegawa-Johnson, A. Perlman, J. Gunderson, T. S. Huang, K. Watkin, and S. Frame (2008) Dysarthric Speech Database for Universal Access Research. In Interspeech 2008, p. 1741â1744. External Links: Document, ISSN 2958-1796 Cited by: Table 1. M. Kim and J. E. Sung (2023) Exploring the Efficacy of OpenAI Whisper for the Recognition of Dysarthric Speech. Applied Sciences 13 (12), p. 7092. External Links: Document Cited by: §1. K. Kuhn, V. Kersken, and G. Zimmermann (2025) Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, p. 1â7. Cited by: §1. G. E. Lancioni, J. Navarro, N. N. Singh, M. F. OâReilly, J. Sigafoos, A. Mellino, P. Arcuri, G. Alberti, and V. Chiariello (2025) People with Neuro-Motor Impairment, Lack of Speech, and General Passivity Can Engage in Basic Forms of Activity and Communication with Technology Support. Advances in Neurodevelopmental Disorders 9 (1), p. 105â114. Cited by: §4. S. Liu, M. Geng, S. Hu, X. Xie, M. Cui, J. Yu, X. Liu, and H. Meng (2022) Recent Progress in the CUHK Dysarthric Speech Recognition System. arXiv preprint arXiv:2201.05845. Cited by: §1. B. Maassen (2007) Speech Motor Control: In Normal and Disordered Speech. Oxford University Press. Cited by: §1. A. Martin, R. L. MacDonald, P. Jiang, M. Ladewig, J. Cattiau, R. Heywood, R. Cave, J. Tobin, P. C. Nelson, and K. Tomanek (2025) Project Euphonia: Advancing Inclusive Speech Recognition through Expanded Data Collection and Evaluation. Frontiers in Language Sciences 4, p. 1569448. Cited by: §1. M. McCloskey and N. J. Cohen (1989) Catastrophic interference in connectionist networks: the sequential learning problem. G. H. Bower (Ed.), Psychology of Learning and Motivation, Vol. 24, p. 109â165. External Links: ISSN 0079-7421, Document, Link Cited by: §2. E. Mosqueira-Rey, E. HernĂĄndez-Pereira, D. Alonso-RĂos, J. Bobes-BascarĂĄn, and Ă. FernĂĄndez-Leal (2023) Human-in-the-Loop Machine Learning: A State of the Art. Artificial Intelligence Review 56 (4), p. 3005â3054. Cited by: §2. A. D. Page and K. M. Yorkston (2022) Communicative Participation in Dysarthria: Perspectives for Management. Brain sciences 12 (4), p. 420. Cited by: §1. V. Panayotov, G. Chen, D. Povey, and S. Khudanpur (2015) Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vol. , p. 5206â5210. External Links: Document Cited by: Table 1. J. S. Park, D. Bragg, E. Kamar, and M. R. Morris (2021) Designing an online infrastructure for collecting ai data from people with disabilities. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT â21, New York, NY, USA, p. 52â63. External Links: ISBN 9781450383097, Link, Document Cited by: §2. A. P. Pinheiro, J. Aucouturier, and S. A. Kotz (2024) Neural Adaptation to Changes in Self-voice during Puberty. Trends in Neurosciences. Cited by: §4. N. Pokel, P. Moure, R. Boehringer, and Y. Gao (2025a) Adapting foundation speech recognition models to impaired speech: a semantic re-chaining approach for personalization of german speech. In 12th edition of the Disfluency in Spontaneous Speech Workshop (DiSS 2025), p. 82â86. External Links: Document Cited by: 1st item, Table 1, Figure 2, §3. N. Pokel, P. Moure, R. Boehringer, and Y. Gao (2025b) Data-efficient asr personalization for non-normative speech using an uncertainty-based phoneme difficulty score for guided sampling. External Links: 2509.20396, Link Cited by: Figure 2, §3, §6. N. Pokel, P. Moure, R. Boehringer, S. Liu, and Y. Gao (2025c) Variational low-rank adaptation for personalized impaired speech recognition. External Links: 2509.20397, Link Cited by: §1, Figure 2, §3. A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever (2023) Robust Speech Recognition via Large-scale Weak Supervision. In International Conference on Machine Learning, p. 28492â28518. Cited by: §1, §6. S. N. B. Reddy and A. N. M. Reddy (2023) Automatic speech recognition for cleft lip and cleft palate patients using deep learning techniques. In Speech and Language Processing for Human-Machine Communications, External Links: Document, Link Cited by: §4. F. Rudzicz, A. K. Namasivayam, and T. Wolff (2012) The TORGO database of acoustic and articulatory speech from speakers with dysarthria. Language Resources and Evaluation 46 (4), p. 523â541. External Links: Document, Link, ISSN 1574-0218 Cited by: Table 1. Z. Sarsenbayeva and et al. (2022) Methodological standards in accessibility research on motor impairments: a survey. ACM Computing Surveys (CSUR) 55 (7), p. 1â35. Cited by: §1. S. Schneider, A. Baevski, R. Collobert, and M. Auli (2019) Wav2vec: Unsupervised Pre-Training for Speech Recognition. In Proc. Interspeech 2019, p. 3465â3469. Cited by: §1. B. Settles (2009) Active learning literature survey. Technical report Technical Report 1648, Computer Sciences Technical Report, University of Wisconsin-Madison, Department of Computer Sciences. Cited by: §2. A. Singh, A. Goyal, A. Kumar, and R. Kumar (2024) An Investigation of Whisper ASR for Dysarthric Speech Recognition. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 11711â11715. External Links: Document Cited by: §1. J. Tobin, P. Nelson, B. MacDonald, R. Heywood, R. Cave, K. Seaver, A. Desjardins, P. Jiang, and J. R. Green (2024) Automatic Speech Recognition of Conversational Speech in Individuals with Disordered Speech. Journal of Speech, Language, and Hearing Research 67 (11), p. 4176â4185. Cited by: §1. J. Tröger, F. Dör, L. Schwed, N. Linz, A. König, T. Thies, J. R. Orozco-Arroyave, and J. Rusz (2024) An Automatic Measure for Speech Intelligibility in Dysarthriasâvalidation Across Multiple Languages and Neurological Disorders. Frontiers in Digital Health 6, p. 1440986. Cited by: §2. L. van Bemmel, C. Pesenti, X. Wei, and H. Strik (2023) Automatic Assessments of Dysarthric Speech: the Usability of Acoustic-Phonetic Features. In Proc. Interspeech 2023, p. 141â145. Cited by: §2. K. Vertanen and P. O. Kristensson (2009) Parakeet: a continuous speech recognition system for mobile touch-screen devices. In Proceedings of the 14th International Conference on Intelligent User Interfaces, IUI â09, New York, NY, USA, p. 237â246. External Links: ISBN 9781605581682, Link, Document Cited by: §4. X. Wu, L. Xiao, Y. Sun, J. Zhang, T. Ma, and L. He (2022) A Survey of Human-in-the-Loop for Machine Learning. Future Generation Computer Systems 135, p. 364â381. Cited by: §2. F. Xiong, J. Barker, and H. Christensen (2019) Phonetic Analysis of Dysarthric Speech Tempo and Applications to Robust Personalised Dysarthric Speech Recognition. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 6465â6469. External Links: Document Cited by: §2. C. E. Zimmerman, J. Sun, A. M. Wes, G. H. Vu, C. L. Kalmar, L. S. Humphries, S. P. Bartlett, M. A. Cohen, J. W. Swanson, and J. A. Taylor (2020) Long term speech outcomes following midface advancement in syndromic craniosynostosis. The Journal of Craniofacial Surgery 31 (6), p. 1775â1779. External Links: Document Cited by: §1.