Paper deep dive
Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi, Kaushal Bhogale, Mitesh M. Khapra
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcriptions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large language models. Indic DiarBench is released as an open-access resource to advance inclusive, multilingual speech technology research for Indian languages.
Tags
Links
- Source: https://arxiv.org/abs/2607.23808v1
- Canonical: https://arxiv.org/abs/2607.23808v1
Trouble viewing inline? Open PDF directly →
Full Text
27,374 characters extracted from source content.
Expand or collapse full text
Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages Deovrat Mehendale 1,∗ , Aditya Mehndiratta 1,∗ , Dhruv Rathi 1,∗ Kaushal Bhogale 2 , Mitesh M. Khapra 2 1 Sarvam AI, India 2 AI4Bharat, IIT Madras, India deovrat.mehendale@gmail.com, adityam0309@gmail.com, dhruvsubhashrathi@gmail.com Abstract In this work, we introduce Indic DiarBench, a speaker diariza- tion and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcrip- tions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large lan- guage models. Indic DiarBench is released as an open-access resource 1 to advance inclusive, multilingual speech technology research for Indian languages. Index Terms: speaker diarization, Indian languages, multilin- gual benchmark, speaker-attributed ASR, code-mixing 1. Introduction Recent years have seen significant progress in automatic speech recognition (ASR) for Indian languages. Large scale data collection efforts such as IndicVoices [1] have enabled multilingual ASR systems that begin to cover India’s linguis- tic diversity. However, most of this progress has focused on single speaker speech, while many real world scenarios such as meetings, interviews, panel discussions, and casual conversa- tions involve multiple interacting speakers. In practical multi speaker transcription pipelines, speaker diarization is first used to identify and segment speakers, after which ASR is applied to each segment. Despite substantial advances in speaker diarization re- search [2], existing datasets and benchmarks remain heavily concentrated in English and a few other high resource lan- guages, leaving no standardized benchmark for multi speaker speech processing in Indian languages. This gap is particu- larly consequential because diarization and ASR must operate jointly in real deployments. Evaluating them independently masks important failure modes: ASR performance often de- grades sharply when applied to the short, fragmented, and over- lapping segments produced by diarization systems. As a re- sult, progress in real world conversational transcription requires benchmarks that evaluate speaker attributed ASR under realis- tic multi speaker conditions, a capability that remains largely unexplored for Indian languages spanning 22 constitutionally recognized languages and four major language families. * These authors contributed equally. 1 https://huggingface.co/datasets/sarvamai/ indic-diarbench Figure 1: Speaker Data Collection by District and Language To address this gap, we introduce Indic DiarBench, an open-access benchmark for evaluating speaker-attributed ASR in multilingual Indian conversational speech. Our contributions can be summarised as follows: • A corpus of 108 hours of conversational speech covering all 22 scheduled Indian languages, collected across near-field meetings, far-field recordings, and in-the-wild YouTube con- versations. • Meeting recordings include 485 unique speakers from 189 districts across India (Figure 1). The in-the-wild data further includes≈ 750 speakers across the 10 most widely spoken Indian languages. • Human-corrected,speaker-attributed,segment-level transcriptions paired with aligned speaker time annotations (RTTM format), enabling evaluation of joint diarization and ASR performance. • Baseline results from multiple state-of-the-art commercial APIs, multimodal LLMs, and Indic-specialized systems, establishing reference performance across languages and acoustic conditions. Indic DiarBench, along with all annotations, evaluation pro- tocols, and baseline systems, is publicly released to support reproducible research and advancement of speaker-attributed ASR and diarization for Indian languages. 2. Related Work Early diarization benchmarks such as the AMI [3] and ICSI [4] meeting corpora provided multi-channel English arXiv:2607.23808v1 [cs.CL] 26 Jul 2026 Table 1: Comparison of diarization evaluation datasets. In- dic DiarBench is the first to cover all 22 Indian scheduled lan- guages with joint ASR + diarization labels. DatasetLang.Hrs DomainSpkrs ASR CALLHOME6 langs20Telephone2–6✓ AMIEnglish100 Meeting3–5✓ DIHARD IIIMulti33Mixed1–10– VoxConverseMulti70YouTube1–21– LibriCSSEnglish10Simulated8✓ AliMeetingMandarin 118 Meeting2–4✓ NOTSOFAR-1 English28Distant mtg.4–8✓ DISPLACE ’24 5 Indic38Conversation 3–5– Ours22 Indic∼108 Mixed2–9✓ recordings in controlled rooms. CALLHOME [5] extended coverage to six languages but is limited to telephonic speech with two to six speakers. The DIHARD challenge series [6] pushed evaluation toward degraded and diverse domains, while VoxConverse [7] introduced in-the-wild YouTube audio with audio-visual annotation. For Mandarin, AliMeeting [8] and AISHELL-4 [9] provide meeting-style corpora. LibriCSS [10] targets continuous speech separation with controlled overlap ra- tios. More recently, SDBench [11] unified 13 datasets under a single evaluation framework, and NOTSOFAR-1 [12] intro- duced realistic meeting data with tcpWER evaluation. Despite this progress, none of these benchmarks provide substantial coverage of Indian languages. Table 1 summarizes key datasets. At a broader multilingual ASR level, Common Voice [13] and FLEURS [14] significantly expanded language coverage, but both are primarily designed for single-speaker recognition and do not provide speaker-attributed conversational diarization la- bels. For Indian languages, the DISPLACE challenges [15, 16] represent the first attempts at diarization benchmarks. DIS- PLACE 2023 provided 32 hours of conversational audio across seven languages, while the 2024 edition expanded to 158 hours (38 hours labelled). However, the ASR track in 2024 was eval- uated on a separate 12-hour subset of cleaner near-field, single- speaker audio. By decoupling diarization from ASR and cover- ing only a subset of Indian Languages, DISPLACE leaves crit- ical gaps for evaluation of speaker-attributed transcriptions for Indian languages in multi-speaker conversational settings. 3. The Indic DiarBench Corpus Indic DiarBench is a multilingual conversational speech benchmark designed to evaluate speaker attributed ASR in re- alistic multi speaker settings for Indian languages. The follow- ing subsections describe the data collection process, annotation pipeline, and dataset statistics. 3.1. Data Collection We now describe the recording conditions represented in the corpus, followed by the criteria used for audio selection and the characteristics of the speaker population. 3.1.1. Recording Conditions • Near-field meetings (∼ 53 hrs): Recorded using one close- proximity microphone per speaker and designed to capture spontaneous interaction, frequent interjections, and overlap- ping speech. Participants were not co-located; they joined virtually via an online meeting platform. This setup allowed accurate capture of speaker turns by combining individual mi- crophone streams. This subset covers all 22 scheduled Indian languages. The top eight languages by native-speaker popu- lation (Hindi, Bengali, Marathi, Telugu, Tamil, Gujarati, Pun- jabi, Kannada) contribute∼ 4 hours each, while the remain- ing 14 languages contribute∼ 1.5 hours each. • Far-field meetings (∼ 27 hrs): Recorded using distant mi- crophones, introducing reverberation, background noise, and variable speaker-to-microphone distances. This subset in- cludes approximately 1.4 to 4.2 hours per language for the top eight languages, with sessions involving 2–8 speakers. • In-the-wild audio (∼ 28 hrs): Curated from publicly avail- able YouTube videos to capture unconstrained acoustic envi- ronments. Approximately 2 hours per language are collected for the ten most widely spoken Indian languages. Selected clips emphasize sustained multi-speaker interaction and avoid broadcast-style content dominated by a single speaker. 3.1.2. Data Curation and Quality Control To encourage natural conversation and avoid scripted turn- taking, participants were provided debate topics and quiz ques- tionnaires in advance and asked to speak freely. An initial warm-up portion of each recording is discarded to retain only spontaneous interaction. Sessions with same-gender speakers are included to increase speaker confusability. After collection, all recordings undergo a curation and quality control process to ensure suitability for multi-speaker evaluation. For meeting recordings, trained language experts review sessions for recording quality, vocabulary diversity, speaker overlap, and conversational spontaneity.Sessions that are overly scripted or exhibit poor audio quality are discarded. Noise cancellation is kept at “low” whenever possible during recording to preserve natural ambient sound and conversational artifacts. For in-the-wild audio, clips are curated from publicly avail- able videos that (i) contain sustained multi-speaker interaction, (i) avoid advertisements and non-speech-dominant music, and (i) provide a clear visual indication of the active speaker, en- abling reliable speaker identification during annotation. 3.1.3. Speaker Population Following the data collection principles established in In- dicVoices [1], we recruit speakers from diverse geographic, de- mographic, and dialectal backgrounds across India. Figure 1 illustrates the geographic distribution of speaker data collection across districts and languages. Meeting recordings across near- and far-field conditions include 485 unique speakers drawn from urban and rural populations across 22 languages and 189 districts, capturing substantial accent and dialectal variation. Speaker identities and associated metadata (e.g., language, re- gion, and demographics) are tracked using unique speaker IDs and session-level metadata files that will be released with the meetings dataset. The speaker pool is gender-balanced and spans a range of educational backgrounds. For in-the-wild audio, speaker uniqueness is enforced by restricting the dataset to one video per channel. Speaker em- beddings are subsequently clustered to identify potential over- laps across videos, followed by manual verification to ensure distinct speaker identities. 3.2. Annotation Pipeline All recordings are annotated using a unified human-in- the-loop pipeline designed to produce high-quality speaker- attributed transcriptions. The pipeline combines automatic tran- scription with three stages of human verification to ensure ac- curate transcriptions, timestamps, and speaker labels. 1. Bootstrap Transcription. Initial transcripts are generated using multiple independent ASR systems, including publicly available models and closed source APIs. Annotators are presented with these hypotheses as editable drafts, reducing annotation effort while limiting bias from any single system. 2. Human Transcription and Speaker Attribution. Profes- sional annotators produce time-aligned, speaker-attributed transcriptions by verifying and correcting word sequences, timestamps, and speaker labels. No machine-generated an- notation is retained without human validation. For meeting recordings, where the number of speakers is known and indi- vidual microphone channels provide reliable speaker timing, annotators may adjust timestamps or speaker labels but can- not introduce new speakers. For in-the-wild audio, where visual cues are available, annotators may add, merge, or re- move speaker labels when necessary. 3. Code-Mixed Transcription. Indian conversational speech has frequent code-mixing with English. To accommodate this, annotators produce two transcription formats. The first is a native-script form in which all text, including English words, is rendered in the Indic script. The second is a nor- malized form in which English words are written using Ro- man script and numerals using Arabic digits. Both formats are accepted during word error rate computation to avoid pe- nalizing models that differ in output conventions. 4. Quality Control.Dedicated quality checkers (2–3 per language) verify transcription consistency, code-mixing conventions, numeral representations, non-speech tags (<laughter>, <noise>, <cry>, etc.), and speaker timestamps and labels.Particular attention is paid to overlapping speech segments, which often require multiple rounds of review. Minor corrections are applied at this stage, while major issues are returned to annotators for revision. 5. Expert Review. Finally, each file undergoes language spe- cific quality checks by an in-house expert (superchecker). Rather than correcting instances directly, the expert identi- fies systematic issues and returns substandard annotations for revision until quality standards are met. 3.3. Dataset Statistics Indic DiarBench contains approximately 108 hours of an- notated audio across all 22 scheduled Indian languages. Table 2 presents per-language statistics including duration by acoustic condition, speaker counts, and overlap ratios. Limitations. The in-the-wild subset covers 10 of the 22 lan- guages; extending coverage to all languages is planned. We do not provide speaker IDs for the in-the-wild subset. The dataset is designed for evaluation rather than training. 4. Evaluation Setup Metrics. We report evaluation metrics along two complemen- tary axes: acoustic diarization and word-level speaker attribu- tion. For acoustic segmentation, we report Diarization Error Rate (DER) computed without a forgiveness collar and includ- ing overlapping speech. To jointly evaluate ASR and diarization performance, we Table 2: Per-language statistics for Indic DiarBench. All dura- tions are in hours. Hours are broken down by acoustic condition (NF = near-field, F = far-field, ITW = in-the-wild). Overlap % is averaged over all conditions. “–” indicates that the con- dition is not available for that language. Language FamilyNFFF ITW Total Overlap % AssameseIndo-Aryan1.5–1.513.9 BengaliIndo-Aryan4.44.14.112.67.8 BodoSino-Tibetan1.6–1.615.2 DogriIndo-Aryan1.4–1.424.2 GujaratiIndo-Aryan4.14.22.811.17.6 HindiIndo-Aryan4.242.510.716.6 KannadaDravidian3.71.43.38.515.2 KashmiriIndo-Aryan1.1–1.121.2 KonkaniIndo-Aryan1.6–1.614.0 MaithiliIndo-Aryan1.3–1.324.7 Malayalam Dravidian1.3–2.43.712.9 ManipuriSino-Tibetan1.5–1.520.6 MarathiIndo-Aryan4.23.52.710.411.3 NepaliIndo-Aryan1.3–1.322.9 OdiaIndo-Aryan1.5–1.63.111.1 PunjabiIndo-Aryan4.342.410.66.1 SanskritIndo-Aryan1.6–1.621.4 SantaliAustroasiatic 1.6–1.66.5 SindhiIndo-Aryan1.5–1.516.0 TamilDravidian4.22.53.21012.4 TeluguDravidian3.832.59.320.4 UrduIndo-Aryan1.6–1.612.5 Total4 families53.2 26.8 27.6∼10812.8 Table 3: Duration-weighted aggregate metrics across all three acoustic conditions (%). Best values in bold. CategoryModelDER cpWER WDER Miss FA Conf Indic-Spec.Sarvam16.038.833.16.3 3.9 5.9 Comm. APIs AWS Transcribe23.543.734.313.1 3.1 7.4 ElevenLabs Scribe 35.058.340.713.6 6.2 15.3 Azure STT34.860.839.524.4 1.7 8.7 Deepgram Nova-3 32.063.239.318.3 5.4 8.3 AssemblyAI40.588.643.725.5 5.5 9.6 Multimod. LLMs GPT-4o36.283.140.417.7 3.8 14.7 Gemini 3 Pro74.058.933.041.7 5.8 26.5 use Concatenated minimum-permutation Word Error Rate (cp- WER) [17] and Word Diarization Error Rate (WDER) [18]. cpWER primarily reflects transcription accuracy after resolv- ing speaker permutations, whereas WDER explicitly penalizes speaker attribution errors and therefore more directly captures diarization quality. Models: We benchmark several prominent ASR–diarization systems, including commercial speech APIs (Sarvam AI [19], Deepgram Nova-3 [20], ElevenLabs Scribe [21], AssemblyAI Universal-2 [22], Azure STT [23], AWS Transcribe [24]) and multimodal large language models (Gemini 3 Pro [25] and GPT- 4o Transcribe [26]). Evaluation is restricted to systems capable of producing joint ASR and diarization outputs, as assessing diarization in- dependently of transcription is increasingly misaligned with the requirements of modern speech applications. Consequently, diarization-only models (e.g., Pyannote [27]) are excluded. Language coverage varies across systems, as not all commer- cial providers support all 22 scheduled Indian languages. (a) cpWER (%) AsBnBrxDoiGuHiKnKsKokMaiMlMniMrNeOrPaSaSatSdTaTeUr Assembly9999–1005098–95–9297–10095–92909951 AWS–37–373753–56–39–5435–6357– Azure 10050–574969–70–591006566–6366100 Deepgram–52–4979–61–9063– ElevenLabs5358–474470–67–54656358–75706793 Gemini 65547781493966103677856116396470715899786167100 Sarvam46355953273151584165385634493734512844485126 < 3030–4445–5960–74≥ 75N/A (b) WDER (%) AsBnBrxDoiGuHiKnKsKokMaiMlMniMrNeOrPaSaSatSdTaTeUr 2026–154049–51–5242–3435–33544028 –28–314127–35–30–2737–3544– 5230–364240–40–43423145–384827 –37–3636–40–4845– 3438–334442–40–41473640–29454923 32284441273230492941355633423032393428354129 28313731243934272136343332412435301826344428 < 1515–2425–3435–49≥ 50N/A Figure 2: Per-language (a) cpWER and (b) WDER (%) heatmap. Figure 3: Metrics variation vs overlap ratio 5. Results and Analysis For evaluation, all systems were provided the same single- channel mixed audio to ensure fairness. Figure 2 presents per-language cpWER and WDER for all evaluated systems across 22 Indic languages. Grey cells indicate unsupported lan- guages; consequently, global averages over languages are inher- ently skewed, and we focus instead on duration-weighted aggre- gates and category-wise trends. Table 3 summarizes duration- weighted performance across all three acoustic conditions. Model Comparison. We first compare overall performance across model categories. Among all evaluated systems, the Indic-specialized Sarvam pipeline consistently achieves the strongest results, obtaining the lowest DER (16.0%) and cp- WER (38.8%) in Table 3. Other model APIs exhibit moder- ate performance, with AWS Transcribe performing best in this group (23.5% DER, 43.7% cpWER), while other APIs show substantially higher error rates.Multimodal LLMs present a contrasting trade-off: Gemini 3 Pro achieves competitive WDER (33.0%) due to strong ASR quality on detected speaker segments, but suffers from very high DER (74.0%), whereas GPT-4o achieves better diarization accuracy but exhibits poor cpWER, particularly on lower-resource languages. Error Analysis. To better understand system behavior beyond aggregate metrics, we decompose DER into missed detection, false alarm, and speaker confusion errors (see Table 3). This breakdown reveals distinct failure patterns across model classes. Indic-specialized pipelines exhibit a balanced error profile: Sar- vam’s 16.0% DER is composed of 6.3% missed detection, 3.9% false alarms, and 5.9% speaker confusion, indicating no single dominant failure mode. In contrast, multimodal LLMs are dom- inated by missed detection errors. Gemini 3 Pro’s 74.0% DER includes 41.7% missed detection and 26.5% confusion, largely due to unreliable timestamping and missed detection on small utterances like affirmations and interjections. Despite this, its cpWER of 58.9% remains competitive with commercial ASR APIs, indicating strong transcription quality when speaker seg- ments are correctly identified. Generic Commercial APIs ex- hibit more model-specific behavior: Azure STT achieves low false alarm rates (1.7%) due to conservative voice-activity de- tection, but incurs high missed detection (24.4%), while AWS Transcribe has relatively more balanced error decomposition. Effects of overlap In the Indic-specific model (Sarvam), across the 22 languages overlap ratio is strongly correlated with DER and cpWER (Figure 3). Performance across languages. We provide language-wise and near-field vs far-field results in the supplementary mate- rial (DatasetSummary.csv) and discuss them briefly here. In the near-field condition, Telugu emerges as the most challeng- ing language, exhibiting the highest overlap (24.7%) and 4.5 speakers per recording on average, yielding a DER of 27.7%. Maithili (24.7% overlap, 4.7 speakers, 22.7% DER) and Do- gri (24.2%, 5.3 speakers, 23.2% DER) follow closely. At the other end, Santali (6.5% overlap, 4.2 speakers, 9.7% DER) and Urdu (12.5%, 3.8 speakers, 12.8% DER) are among the easi- est; these trends are consistently reflected across all evaluated systems. In-the-wild YouTube recordings, which have the low- est overlap (6.5%) and cleaner turn-taking, yield the best per- formance. Lower-resource languages compound these effects. The 12 languages available only in near-field recordings show higher DER and cpWER compared to the 10 higher-resource languages that span multiple acoustic conditions. Across lan- guage families, Dravidian languages (Kannada, Malayalam, Tamil, Telugu) show near-field cpWER roughly 5 percentage points above Indo-Aryan languages at comparable DER. 6. Conclusion We present Indic DiarBench, the first open benchmark for joint diarization and speaker-attributed ASR spanning all 22 scheduled Indian languages. By unifying near-field meetings, far-field recordings, and in-the-wild conversations, the bench- mark captures realistic variation in speaker counts, overlap ra- tios, and acoustic conditions, establishing a standardized eval- uation framework for multi-speaker conversational speech in Indian languages. Our results underline that robust speaker- attributed recognition remains challenging in short-utterance, high-overlap, and low-resource settings, motivating continued research on tightly coupled diarization and ASR systems. 7. Acknowledgments We thank Sshubam, Sadakopa, and Vamsi from Sarvam AI for generously giving their time and helping with the YouTube data collection effort. We also thank the language experts at Sarvam AI and AI4Bharat for their excellent work; this effort would not have been possible without their contributions. 8. Generative AI use disclosure Generative AI tools were used only for limited language editing and polishing of parts of the manuscript. All technical content, analyses, results, and conclusions were produced and verified by the authors. 9. References [1] T. Javed, J. Nawale, E. I. George, S. Joshi, K. S. Bhogale et al., “IndicVoices: Towards building an inclusive multilingual speech dataset for Indian languages,” in Findings of ACL, 2024, p. 10 740–10 782. [2] T. J. Park, N. Kanda, D. Dimitriadis, K. J. Han, S. Watanabe, and S. Narayanan, “A review of speaker diarization: Recent advances with deep learning,” Computer Speech & Language, vol. 72, p. 101317, 2022. [3] J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain et al., “The AMI meeting corpus: A pre-announcement,” in Proc. Machine Learning for Multimodal Interaction (MLMI), 2005, p. 28–39. [4] A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke, and C. Wooters, “The ICSI meeting corpus,” in Proc. ICASSP, 2003, p. 364–367. [5] A. Canavan, D. Graff, and G. Zipperlen, “CALLHOME American English speech,” 1997, lDC97S42. [6] N. Ryant, P. Singh, V. Krishnamohan, R. Varma, K. Church, C. Cieri, J. Du, S. Ganapathy, and M. Liberman, “The third DI- HARD diarization challenge,” in Proc. Interspeech, 2021, p. 3570–3574. [7] J. S. Chung, J. Huh, A. Nagrani, T. Afouras, and A. Zisserman, “Spot the conversation: Speaker diarisation in the wild,” in Proc. Interspeech, 2020, p. 299–303. [8] F. Yu, S. Zhang, Y. Fu, L. Xie et al., “M2MeT: The ICASSP 2022 multi-channel multi-party meeting transcription challenge,” in Proc. ICASSP, 2022, p. 6167–6171. [9] Y. Fu, L. Cheng, S. Lv, Y. Jv, Y. Kong, Z. Chen, Y. Hu, L. Xie, J. Wu, H. Bu, X. Xu, J. Du, and J. Chen, “AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,” in Proc. Inter- speech, 2021, p. 3665–3669. [10] Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y. Luo, X. Xiao, J. Li, and J. Wu, “Continuous speech separation: Dataset and anal- ysis,” in Proc. ICASSP, 2020, p. 7284–7288. [11] B. Durmus, B. Munyampirwa, E. Pacheco, A. Orhon, and A. Leonov, “SDBench: A comprehensive benchmark suite for speaker diarization,” in Proc. Interspeech, 2025, p. 1598–1602. [12] A. Vinnikov, A. Ivry, A. Hurvitz, I. Abramovski, S. Koubi, I. Gur- vich, S. Peer, X. Xiao, B. M. Elizalde, N. Kanda, X. Wang, S. Shaer, S. Yagev, Y. Asher, S. Sivasankaran, Y. Gong, M. Tang, H. Wang, and E. Krupka, “NOTSOFAR-1 challenge: New datasets, baseline, and tasks for distant meeting transcription,” in Proc. Interspeech, 2024, p. 5003–5007. [13] R. Ardila et al., “Common voice: A massively-multilingual speech corpus,” in Proc. LREC, 2020, p. 4218–4222. [14] A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “FLEURS: Few-shot learning evaluation of universal representations of speech,” in 2022 IEEE Spoken Language Technology Workshop (SLT), 2023, p. 798– 805. [15] S. Baghel, S. Ramoji, Sidharth, R. H, P. Singh, S. Jain, P. R. Chowdhuri, K. Kulkarni, S. Padhi, D. Vijayasenan, and S. Gana- pathy, “The DISPLACE challenge 2023 – DIarization of SPeaker and LAnguage in Conversational Environments,” in Proc. Inter- speech, 2023, p. 3562–3566. [16] S. B. Kalluri, P. Singh, P. R. Chowdhuri, A. Kulkarni, S. Baghel, P. Hegde, S. Sontakke, D. K T, S. R. M. Prasanna, D. Vijayasenan, and S. Ganapathy, “The second DISPLACE challenge: DIariza- tion of SPeaker and LAnguage in Conversational Environments,” in Proc. Interspeech, 2024, p. 1630–1634. [17] S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Araki, X. Chang et al., “CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,” in Proc. CHiME Workshop, 2020. [18] L. El Shafey, H. Soltau, and I. Shafran, “Joint speech recogni- tion and speaker diarization via sequence transduction,” in Proc. Interspeech, 2019, p. 396–400. [19] Sarvam AI, “Sarvam ASR,” https://w.sarvam.ai/blogs/asr/, ac- cessed: 2026-03-05. [20] Deepgram, “Introducing Nova-3 Speech-to-Text API,” https:// deepgram.com/learn/introducing-nova-3-speech-to-text-api, ac- cessed: 2026-03-05. [21] ElevenLabs, “Speech to Text Capabilities,” https://elevenlabs.io/ docs/overview/capabilities/speech-to-text, accessed: 2026-03-05. [22] AssemblyAI, “Universal-2 Speech Recognition Model,” https:// w.assemblyai.com/universal-2, accessed: 2026-03-05. [23] Microsoft Azure, “Azure Speech-to-Text,” https://learn.microsoft. com/en-us/azure/ai-services/speech-service/speech-to-text,ac- cessed: 2026-03-05. [24] Amazon Web Services, “Amazon Transcribe,” https://aws. amazon.com/transcribe/, accessed: 2026-03-05. [25] Google, “Gemini 3,” https://blog.google/products-and-platforms/ products/gemini/gemini-3/, accessed: 2026-03-05. [26] OpenAI, “GPT-4o Transcribe Model Documentation,” https: //developers.openai.com/api/docs/models/gpt-4o-transcribe, ac- cessed: 2026-03-05. [27] H. Bredin, “pyannote.audio 2.1 speaker diarization pipeline: Prin- ciple, benchmark, and recipe,” in Proc. Interspeech, 2023, p. 1983–1987.