Paper deep dive
When Multiple Scripts Matter: Evaluating ASR in Clinical Settings
Jean Seo, Minkyu Kim, Jeonguk Lee, Jisoo Jung, Wooseok Han, Eunho Yang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 6/21/2026, 1:11:50 AM
Summary
The paper introduces MultiClin, a clinical ASR benchmark designed to address the challenges of multiscript variability in non-English clinical settings (specifically Korean-English). The authors argue that conventional single-reference evaluation metrics like Word Error Rate (WER) underestimate ASR performance by failing to account for multiple valid orthographic forms of the same term. Through experiments with models like Whisper, Qwen, and Gemini, the study demonstrates that multiscript-aware evaluation provides a fairer assessment. Furthermore, the research highlights that script unification (100% transliteration ratio) during training significantly improves model performance and reduces orthographic uncertainty compared to inconsistent script mapping.
Entities (7)
Relation Signals (4)
Whisper → isevaluatedby → MultiClin
confidence 100% · We evaluate ASR performance on the MultiClin benchmark... We consider three model families as baselines: (1) Whisper
script unification → improves → ASR performance
confidence 95% · script unification consistently yields the best ASR performance.
inconsistent script mapping → increases → orthographic uncertainty
confidence 95% · inconsistent script mappings increase orthographic uncertainty and hinder model convergence
MultiClin → addresses → multiscript variability
confidence 90% · MultiClin, a clinical ASR benchmark designed to evaluate robustness to multiscript variability.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automatic speech recognition (ASR) in non-English clinical settings is challenged by multiscript variability, where the same term may appear in multiple valid orthographic forms. Conventional string-matching evaluation metrics often underestimate ASR performance by treating orthographic variants as errors. To address this issue, we introduce MultiClin, a clinical ASR benchmark designed to evaluate robustness to multiscript variability. Experiments across diverse ASR models show that multiscript-aware evaluation provides a fairer assessment of recognition quality than conventional single-reference evaluation. We further investigate the impact of script consistency during training and find that inconsistent script mappings increase orthographic uncertainty and hinder model convergence, with a balanced 50% mapping ratio producing the highest entropy. In contrast, script unification consistently yields the best ASR performance. Our dataset and code are publicly available at: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2606.17826v1
- Canonical: https://arxiv.org/abs/2606.17826v1
Trouble viewing inline? Open PDF directly →
Full Text
25,997 characters extracted from source content.
Expand or collapse full text
When Multiple Scripts Matter: Evaluating ASR in Clinical Settings Jean Seo 1,2,∗ , Minkyu Kim 1,∗ , Jeonguk Lee 1 , Jisoo Jung 1 , Wooseok Han 1 , Eunho Yang 1,3,∗ 1 AITRICS, 2 University of Copenhagen, 3 KAIST jean.seo@di.ku.dk minkyu.kim@aitrics.com Abstract Automatic speech recognition (ASR) in non-English clini- cal settings is challenged by multiscript variability, where the same term may appear in multiple valid orthographic forms. Conventional string-matching evaluation metrics often under- estimate ASR performance by treating orthographic variants as errors. To address this issue, we introduce MultiClin, a clin- ical ASR benchmark designed to evaluate robustness to multi- script variability. Experiments across diverse ASR models show that multiscript-aware evaluation provides a fairer assessment of recognition quality than conventional single-reference eval- uation. We further investigate the impact of script consistency during training and find that inconsistent script mappings in- crease orthographic uncertainty and hinder model convergence, with a balanced 50% mapping ratio producing the highest en- tropy. In contrast, script unification consistently yields the best ASR performance. Our dataset and code are publicly avail- able at: https://github.com/aitrics-ronaldo/ Interspeech_MultiClin. Index Terms: automatic speech recognition, evaluation, multi- script variability, code-switching, healthcare 1. Introduction Automatic speech recognition (ASR) is increasingly adopted in clinical settings to improve workflow efficiency [1, 2, 3]. How- ever, domain-specific terminology and noisy environments con- tinue to challenge clinical ASR. These difficulties are further amplified in non-English settings, where English medical ter- minology frequently coexists with phonetic renderings in local scripts [4]. A central obstacle to reliable benchmarking in such environments is multiscript variability, where a single spoken term may correspond to multiple valid orthographic forms (e.g., English spelling or a phonetic rendering in the local script). Un- like conventional code-switching, which involves acoustic al- ternation between languages, multiscript variability arises from orthographic variation despite an identical acoustic realization. Conventional ASR evaluation assumes a single reference transcription per utterance. However, this assumption often breaks down in non-English clinical settings, where English- origin medical terms lack standardized localization guidelines and may be transcribed in multiple valid forms. This many- to-one mapping between orthography and speech invalidates strict string-based metrics such as word error rate (WER), sys- tematically penalizing outputs that are phonetically and seman- tically correct but orthographically different from the refer- ence [5, 6, 7]. Moreover, normalization-based solutions remain * Equal contribution. ** Corresponding author. Table 1: Example of the original, tagged, and translated dia- logue from the MultiClin dataset. Doctor-Patient Dialogue Original [patient] How long do I have to wear that brace? [doctor] You’re gonna be wearing the brace for about 6 weeks. [patient] 6 weeks? ... Tagged [patient] How long do I have to wear that <MEDICAL>brace</MEDICAL>? [doctor] You’re gonna be wearing the <MEDICAL>brace</MEDICAL> for about <NUMBER>6</NUMBER> weeks. [patient] <NUMBER>6</NUMBER> weeks? ... Translated [patient]그 <MEDICAL>brace,브레이스</MEDICAL>는얼마나오래착용 해야하나요? [doctor] <MEDICAL>brace,브레이스</MEDICAL>는약 <NUMBER>6,육</NUMBER>주정도착용하시게될거예요. [patient] <NUMBER>6,육</NUMBER>주요? ... impractical due to inconsistent clinical documentation prac- tices and the scarcity of standardized domain-specific corpora. While multilingual ASR research has extensively studied code- switching [8], prior work has largely focused on modeling and data augmentation [9, 10, 11] rather than evaluation. Existing benchmarks typically rely on a single ground-truth reference [12, 13], while transliteration-based approaches [14] and met- rics such as transliterated WER (T-WER) [15, 16] have pri- marily been evaluated on general-domain code-switching and dialectal variation, leaving clinical multiscript settings largely unexplored. To address this gap, we introduce MultiClin, a clinical ASR benchmark that provides multiple valid transcription vari- ants for multiscript terminology. Through a Korean clinical case study, we demonstrate that dynamic multi-reference evaluation yields a fairer assessment of ASR performance under ortho- graphic variability. 2. MultiClin dataset We construct the MultiClin dataset to reflect real-world clini- cal ASR challenges. Table 1 illustrates an example data corre- sponding to each phase of the annotation process. 2.1. Dataset construction 2.1.1. Collection We collect publicly available doctor–patient dialogues from ACIBench [17], Primock57 [18], and MTS-Dialog [19]. To arXiv:2606.17826v1 [cs.CL] 16 Jun 2026 Table 2: Statistics of the MultiClin dataset. A. Filtering Stages (Initial→ Final)B. Avg. tagged instances per dialogue ACI Bench126→ 116MEDICAL Tags44 Primock57186→ 9NUMBER Tags6 MTS-Dialog1, 175→ 191UNIT Tags1 Total Dialogues1, 487→ 316 C. Avg. Utterances Per DialogueD. # Speaker composition Doctor18.0Doctor, Patient268 Doctor 20.3Doctor, Patient, Guest Family22 Patient16.0Doctor, GuestFamily11 Guest Family0.9Doctor, Doctor 2, Patient9 GuestFamily 20.03Other6 ensure natural clinical conversations, we exclude dialogues in- volving virtual assistants and retain only interactions between doctors and patients, resulting in an initial corpus of 1,487 dia- logues. 2.1.2. Annotation The dataset undergoes three processing stages: tagging, transla- tion, and human annotation. We use gpt-5.2 1 to identify script- switching instances and assign them to three categories: MED- ICAL, UNIT, and NUMBER. MEDICAL tags denote English- origin medical terms appearing either in the Roman alphabet or as phonetic loanwords. UNIT tags represent measurement units expressed in native scripts or standardized symbols (e.g., %, cm), while NUMBER tags capture numerical expressions writ- ten in native scripts or Arabic numerals. We then translate the dialogues into Korean using the same model. Tagged spans pre- serve their original form and are augmented with Korean-script (Hangeul) renderings, separated by commas without spaces. For example, “You need an <medical>injection</medical>.” becomes “<medical>injection,인젝션</medical>이 필요합 니다.” In other words, tagged entities undergo translitera- tion, preserving lexical identity while changing only the script, whereas the remaining text undergoes full translation into Ko- rean. Finally, two annotators with nursing backgrounds re- view all dialogues for orthographic correctness, translation fi- delity, and naturalness. Any disagreements or errors are re- solved through consensus, resulting in the final curated dataset. 2.1.3. Speech generation To comply with the Health Insurance Portability and Account- ability Act (HIPAA) restrictions on releasing real-world clini- cal audio, we synthesize dialogues using gpt-4o-mini-tts. We map speaker roles to distinct speaking styles (e.g., professional tones for doctors and lethargic tones for patients) and apply accent-aware prompting to align multiscript spans with native intonation patterns. To reduce the acoustic mismatch between synthetic and real clinical speech, we incorporate human-like conversational dynamics, including overlaps and response la- tencies, and simulate clinical environments using a DSP chain 2 (e.g., reverb and HVAC noise). All audio is resampled to 16, kHz. 1 https://openai.com/ 2 https://github.com/spotify/pedalboard Table 3: Clinical specialty distribution in MultiClin. Other ∗ encompasses 8 minor fields (e.g., Pain Management, Dentistry, Plastic Surgery). Primary Specialty# DialoguesRatio Orthopedics950.301 Family Medicine / General Prac- tice 340.108 Neurology300.095 Urology190.060 Cardiology180.057 Pulmonology160.051 Gastroenterology120.038 Hematology / Oncology90.028 ENT (Otolaryngology)80.025 Rheumatology70.022 Dermatology70.022 Psychiatry70.022 Pediatrics60.019 Neurosurgery50.016 Emergency Medicine50.016 General Surgery50.016 Ophthalmology40.013 Obstetrics & Gynecology40.013 Allergy / Immunology40.013 Endocrinology30.009 Other ∗ 140.044 2.2. Statistics Dataset filtering. From the initial 1,487 dialogues, we retain 1,417 instances containing at least one MEDICAL, NUMBER, or UNIT tag. We then manually remove unnatural or hallucinated conversations, resulting in 316 final dialogues (Table 2A). Tag and dialogue statistics. MEDICAL terminology dominates script-switching instances (Table 2B). Each dialogue contains 34 turns and 68 sentences on average, with per-speaker utter- ance statistics reported in Table 2C. Speaker composition. Most dialogues involve a single doctor and a single patient (Table 2D). In cases where a patient is ab- sent, guardians (GUEST FAMILY) speak on their behalf. Clinical specialty distribution. All dialogues across the three sources are accompanied by structured clinical notes (e.g., SOAP notes). Using gpt-5.2, we infer the primary clinical spe- cialty of each dialogue from this metadata (Table 3). 3. Experiments We evaluate ASR performance on the MultiClin benchmark to quantify the impact of multiscript variability. We analyze zero- shot inference across diverse architectures and assess the effects of domain-specific fine-tuning under different labeling strate- gies. 3.1. Experimental setup 3.1.1. Baseline Models We consider three model families as baselines: (1) Whisper [20] (large-v3, v3-turbo), implemented via faster-whisper 3 ; (2) Qwen3 ASR [21] (0.6B, 1.7B); and (3) Gemini [22] (2.5 Flash, 2.5 Pro), representing frontier multimodal state-of-the-art mod- els. 3.1.2. Inference Configuration We detail the zero-shot inference configurations for our multi- modal baselines to ensure reproducibility. Gemini prompting strategy. We query the Gemini models us- ing a structured zero-shot prompt. We instruct the model to act as a professional medical stenographer and produce ver- batim transcriptions, explicitly prohibiting speaker diarization, speaker prefixes, and summarization. To ensure deterministic and parseable outputs, we set the sampling temperature to 0.0 and enforce a JSON output format, from which we extract the transcript as an array of sentences. Qwen inference setting. For Qwen3 ASR models, we accom- modate long clinical dialogues by setting the maximum gener- ation length to 65,536 tokens. To improve memory efficiency and avoid out-of-memory (OOM) errors during long-form au- dio processing, we limit the maximum inference batch size to 32. 3.1.3. Fine-tuning Configuration For the fine-tuning experiments, we train Whisper models using LoRA [23]. We split the MultiClin dataset into a 9:1 ratio to construct an independent test set. Importantly, we apply a 100% transliteration ratio, in which all tagged MEDICAL, NUMBER, and UNIT entities are consistently unified into the local script to maximize labeling consistency. This setup reduces ortho- graphic ambiguity during the learning phase. Finally, models are trained for 4 epochs with a batch size of 4. 3.1.4. Evaluation Protocol To enable more accurate evaluation of ASR performance under multiscript settings, we introduce a localized evaluation met- ric (Algorithm 1) that treats both the original English medical term and its phonetic rendering in the local script as valid refer- ences. Specifically, for each script-switching entity in the refer- ence transcript, we dynamically extract a 50-character window from the ASR prediction ˆy using a tracking cursor. To mitigate temporal misalignment, we apply Longest Common Substring (LCS) matching between the target entity and the correspond- ing predicted window. We then compute local CER and WER 3 https://github.com/SYSTRAN/faster-whisper Algorithm 1 Dynamic Multiscript Reference Resolution Input: Tagged reference y tag , ASR hypothesis ˆy, Window size W = 50, Mode mappingM∈original, both Output: Dynamically resolved reference y final 1: cursor ← 0 2: y final ← y tag 3: for each entity tuple (t, e orig , e tgt ) in y tag do 4:m←M[t]Fetch evaluation mode for tag type t 5:if m = original then 6:Replace tag with e orig in y final 7:else if m = both then 8:if cursor ≥|ˆy| then 9:Replace tag with e orig in y final 10:else 11:ˆy win ← ˆy[cursor : min(cursor + W,|ˆy|)] 12:cer orig , of f set orig ← LocalCER(e orig , ˆy win ) 13:cer tgt , of f set tgt ← LocalCER(e tgt , ˆy win ) 14: Priority selection based on minimal local error 15:if cer tgt < cer orig then 16:Replace tag with e tgt in y final 17:cursor ← cursor + of f set tgt 18:else 19:Replace tag with e orig in y final 20:cursor ← cursor + of f set orig 21:end if 22:end if 23:end if 24: end for 25: return y final 26: Function LocalCER(e, w) 27:LCS← FindLongestMatch(e, w) 28:w sub ← w[LCS start : LCS end ] 29:cer ← ComputeCER(e, w sub ) 30:return cer, LCS end within these aligned boundaries, reducing the influence of sur- rounding transcription errors and enabling a more robust com- parison of entity-level correctness across orthographic variants. 3.2. Inference results Table 4 presents zero-shot inference performance across differ- ent script evaluation settings. A consistent trend emerges: mov- ing from strict single-label matching (original) to multiscript- aware evaluation (both) yields substantial reductions in error rates across all models. For instance, Gemini 2.5 Pro’s WER decreases from 28.28% to 15.78% when medical terms are eval- uated with multiscript flexibility. These results empirically sup- port the claim that conventional string-based metrics systemati- cally underestimate ASR performance by failing to account for valid orthographic variation in clinical settings. Our proposed benchmark exposes this limitation by revealing the true capa- bilities of ASR models. While MEDICAL tags contribute most to the observed performance gap, model scale also plays a sig- nificant role; Qwen3 ASR 1.7B achieves a 37.01% WER under the full multiscript-aware setting (both). Among open-source systems, Whisper v3 Turbo demonstrates the strongest robust- ness, achieving a 23.00% WER. Overall, Gemini 2.5 Pro attains the best CER of 4.86%. By properly accounting for medical- domain orthographic variation, MultiClin provides a more fair and informative evaluation framework for multiscript clinical ASR. Table 4: Performance of baseline models. MedicalNumberUnit Whisper v3v3 TurboQwen3 0.6BQwen3 1.7BGemini 2.5 FlashGemini 2.5 Pro CERWERCERWERCERWERCERWERCERWERCERWER original original original29.6836.5726.6532.4459.0983.0929.7341.8324.8730.8324.2328.28 both29.5136.5626.4632.4258.8583.0629.3841.7924.6630.7924.0228.24 both original29.8336.7626.8032.6259.0583.0829.6941.8125.0131.0024.3728.47 both29.6536.7526.6032.6158.8083.0429.3441.7624.8030.9824.1628.44 both original original13.3727.788.7122.8948.7780.1415.1237.116.0119.165.0615.67 both13.1227.768.4422.8748.4680.1114.6837.065.7219.124.7615.64 both original13.4827.928.8223.0248.6980.1015.0437.066.1219.295.1515.81 both13.2427.918.5523.0048.3880.0714.6037.015.8319.274.8615.78 Table 5: Detailed CER (%) comparison between pre-trained and fine-tuned Whisper models. Parentheses indicate the ab- solute reduction in CER after fine-tuning on the MultiClin dataset. ModelsMedicalNumberUnit PretrainedFinetuned CER (%)CER (%) Large v3 original original original30.1827.84 (↓2.34) both30.0627.53 (↓2.53) both original30.3527.58 (↓2.77) both30.2427.28 (↓2.96) both original original14.018.39 (↓5.62) both13.858.00 (↓5.85) both original14.168.08 (↓6.08) both13.997.66 (↓6.33) Large v3 Turbo original original original28.7326.77 (↓1.96) both28.5926.47 (↓2.12) both original28.9126.55 (↓2.36) both28.7626.23 (↓2.53) both original original10.086.85 (↓3.23) both9.886.47 (↓3.41) both original10.196.58 (↓3.61) both9.996.16 (↓3.83) Table 6: Impact of transliteration ratio in the training dataset. Results show the performances of fine-tuned Whisper large v3 models on the held-out test set. Ratio (%)CER (%)WER (%) 069.1754.35 2527.4230.51 5057.4748.50 7513.5322.55 1007.6617.48 3.3. Fine-tuning results Table 5 summarizes the performance gains from fine-tuning on the independent test set. Training with a 100% transliteration ratio yields substantial improvements across all evaluation set- tings. Notably, Whisper-Large v3 Turbo achieves a best-in- class CER of 6.16%, corresponding to an absolute reduction of 3.83%p over its pre-trained baseline. Even larger gains are observed for the standard Whisper-Large v3 model, with CER decreasing by up to 6.33%p under multiscript-aware evaluation criteria. These consistent improvements across both architec- tures empirically demonstrate that full script unification is an effective strategy for mitigating orthographic ambiguity in clin- ical ASR. 3.4. Impact of labeling consistency We further investigate the impact of labeling consistency in the training data. As shown in Table 6, the 0% transliteration ra- tio—where all tagged entities are represented exclusively in the Roman alphabet or Arabic numerals while the rest of the utter- ance is written in Korean (Hangeul)—produces the highest error rates on the held-out test set (69.17% CER and 54.35% WER). As the transliteration ratio increases, meaning that a larger pro- portion of entities are represented in Hangeul, performance ex- hibits a non-monotonic pattern, with a secondary error peak at the 50% ratio (57.47% CER and 48.50% WER). This perfor- mance degradation confirms that inconsistent script mapping introduces orthographic ambiguity, maximizing the conditional entropy H(Y|X) for a given acoustic feature X : H(Y|X) =− X y∈Y P(y|X) log P(y|X) At the 50% ratio, the model faces maximum epistemic un- certainty between competing scripts, which disrupts internal alignment and prevents the decoder from forming stable deci- sion boundaries. Ultimately, the 100% ratio resolves the script alternation complexity, yielding the most robust performance (7.66% CER, 17.48% WER). This validates that full script uni- fication is essential for providing a deterministic learning signal. 4. Conclusion This work introduces the MultiClin dataset for fairer evalua- tion in non-English clinical ASR. Our experiments show that multiscript-aware criteria provide a fairer assessment than tra- ditional single-label metrics, which often underestimate true model performance. We further demonstrate that labeling con- sistency in the training data is essential for better performance. Future work should examine how these ASR improvements in- fluence downstream clinical tasks, such as entity extraction and SOAP note generation. 5. Generative AI Use Disclosure This work employs Generative AI tools including Google Gem- ini and OpenAI ChatGPT. Gemini is utilized for linguistic re- finement, including grammatical correction and improving the clarity of the initial manuscript. Furthermore, both Gemini and ChatGPT were integrated into our data construction process to generate synthetic clinical dialogues for the MultiClin dataset, addressing the inherent data scarcity and privacy constraints of the medical domain. We emphasize that the AI tools are used solely under human supervision. All AI-generated datasets are rigorously reviewed and validated by the authors for clinical ac- curacy and ethical compliance. We maintain full responsibility for the final content and the integrity of the published work. 6. References [1] Y. Xu, H. Jia, M. Wang, J. Feng, X. Xu, H. Wang, J. Chen, Z. Zheng, X. Yang, Y. Shen et al., “Enhancing clinical documen- tation with voice processing and large language models: a study on the laos system,” npj Digital Medicine, 2025. [2] A. Alboksmaty, R. Aldakhil, B. W. Hayhoe, H. Ashrafian, A. Darzi, and A.-L. Neves, “The impact of using ai-powered voice-to-text technology for clinical documentation on quality of care in primary care and outpatient settings: a systematic review,” EBiomedicine, vol. 118, 2025. [3] B. D. Tran, R. Mangu, M. Tai-Seale, J. E. Lafata, and K. Zheng, “Automatic speech recognition performance for digital scribes: a performance comparison between general-purpose and special- ized models tuned for patient-clinician conversations,” in AMIA Annual Symposium Proceedings, vol. 2022, 2023, p. 1072. [4] M. T. Agro, A. Kulkarni, K. Kadaoui, Z. Talat, and H. Aldarmaki, “Code-switching in end-to-end automatic speech recognition: A systematic literature review,” 2025. [Online]. Available: https://arxiv.org/abs/2507.07741 [5] M. B. Mustafa, M. A. M. Yusoof, H. K. Khalaf, A. A. R. M. Abushariah, M. L. M. Kiah, H. N. Ting, and S. Muthaiyah, “Code-switching in automatic speech recognition: The issues and future directions,” Applied Sciences, 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:252550241 [6] B. M. L. Srivastava and S. Sitaram, “Homophone iden- tification and merging for code-switched speech recog- nition,” in Interspeech, 2018. [Online]. Available:https: //api.semanticscholar.org/CorpusID:51937752 [7] S. A. Chowdhury, Y. Samih, M. Eldesouki, and A. M. Ali, “Effects of dialectal code-switching on speech modules: A study using egyptian arabic broadcast speech,” in Interspeech, 2020. [Online]. Available:https://api.semanticscholar.org/CorpusID: 226205255 [8] S. Nakayama, A. Tjandra, S. Sakti, and S. Nakamura, “Zero- shot code-switching asr and tts with multilingual machine speech chain,” 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), p. 964–971, 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:211243868 [9] M. G. Kumar, J. Kuriakose, A. Thyagachandran, A. Seth, L. D. Prasad, S. Jaiswal, A. Prakash, H. Murthy et al., “Dual script e2e framework for multilingual and code-switching asr,” arXiv preprint arXiv:2106.01400, 2021. [10] K. Li, J. Li, G. Ye, R. Zhao, and Y. Gong, “Towards code- switching asr for end-to-end ctc models,” ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 6076–6080, 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:145994388 [11] E.Yilmaz,M.McLaren,H.vandenHeuvel,and D. A. van Leeuwen,“Language diarization for semi- supervised bilingual acoustic model training,” 2017 IEEE Automatic Speech Recognition and Understanding Work- shop(ASRU),p.91–96,2017.[Online].Available: https://api.semanticscholar.org/CorpusID:27208838 [12] I. Hamed, A. Hussein, O. Chellah, S. A. Chowdhury, H. Mubarak, S. Sitaram, N. Habash, and A. M. Ali, “Benchmarking evaluation metrics for code-switching automatic speech recognition,” 2022 IEEE Spoken Language Technology Workshop (SLT), p. 999–1005, 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:254070055 [13] G. Paik, Y. Kim, S. Lee, S. Ahn, and C. Kim, “Hike: Hierarchical evaluation framework for korean-english code-switching speech recognition,” ArXiv, vol. abs/2509.24613, 2025. [Online]. Available: https://api.semanticscholar.org/CorpusID:281674977 [14] J. Emond, B. Ramabhadran, B. Roark, P. J. Moreno, and M. Ma, “Transliteration based approaches to improve code-switched speech recognition performance,” 2018 IEEE Spoken Language Technology Workshop (SLT), p. 448–455, 2018. [Online]. Available: https://api.semanticscholar.org/CorpusID:61809382 [15] S. A. Chowdhury, A. Hussein, A. Abdelali, and A. Ali, “Towards one model to rule all: Multilingual strategy for dialectal code- switching arabic asr,” ArXiv, vol. abs/2105.14779, 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:235254012 [16] A. M. Ali, W. Magdy, and S. Renals, “Multi-reference evaluation for dialectal speech recognition system: A study for egyptian asr,” in ANLP@ACL, 2015. [Online]. Available: https://api.semanticscholar.org/CorpusID:13338981 [17] W. wai Yim, Y. Fu, A. B. Abacha, N. Snider, T. Lin, and M. Yetisgen, “Aci-bench: a novel ambient clinical intelligence dataset for benchmarking automatic visit note generation,” 2023. [Online]. Available: https://arxiv.org/abs/2306.02022 [18] A. Papadopoulos Korfiatis, F. Moramarco, R. Sarac, and A. Savkov, “PriMock57:A dataset of primary care mock consultations,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), S. Muresan, P. Nakov, and A. Villavicencio, Eds. Dublin, Ireland: Association for Computational Linguistics, May 2022, p. 588–598. [Online]. Available: https://aclanthology.org/ 2022.acl-short.65/ [19] A. Ben Abacha, W.-w. Yim, Y. Fan, and T. Lin, “An empirical study of clinical note generation from doctor- patient encounters,” in Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. Dubrovnik, Croatia: Association for Computational Linguistics, May 2023, p. 2291–2302. [Online]. Available: https://aclanthology.org/2023.eacl-main.168 [20] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning. PMLR, 2023, p. 28 492–28 518. [21] X. Shi, X. Wang, Z. Guo, Y. Wang, P. Zhang, X. Zhang, Z. Guo, H. Hao, Y. Xi, B. Yang et al., “Qwen3-asr technical report,” arXiv preprint arXiv:2601.21337, 2026. [22] G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen et al., “Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabil- ities,” arXiv preprint arXiv:2507.06261, 2025. [23] E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations, 2022. [Online]. Available:https: //openreview.net/forum?id=nZeVKeeFYf9