Paper deep dive
Tagarela - A Portuguese speech dataset from podcasts
Frederico Santos de Oliveira, Lucas Rafael Stefanel Gris, Alef Iury Siqueira Ferreira, Augusto Seben da Rosa, Alexandre Costa Ferro Filho, Edresson Casanova, Christopher Dane Shulby, Rafael Teixeira Sousa, Diogo Fernandes Costa Silva, Anderson da Silva Soares, Arlindo Rodrigues Galvão Filho
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 5:20:28 AM
Summary
TAGARELA is a large-scale, high-quality Portuguese speech dataset containing over 8,972 hours of podcast audio, designed to address the resource gap in Portuguese speech technology. The dataset includes a multi-stage preprocessing pipeline (standardization, diarization, overlap detection, and denoising) and was transcribed using a bootstrap approach with Whisper large-v3. It supports both ASR and TTS tasks, with models trained on the dataset demonstrating state-of-the-art performance for Portuguese.
Entities (6)
Relation Signals (4)
TAGARELA → supportstask → Automatic Speech Recognition
confidence 98% · composed of over 8,972 hours of podcast audio, specifically curated for training automatic speech recognition (ASR)
TAGARELA → supportstask → Text-to-Speech
confidence 98% · composed of over 8,972 hours of podcast audio, specifically curated for training... text-to-speech (TTS) models.
TAGARELA → derivedfrom → Cem Mil Podcasts
confidence 95% · The TAGARELA dataset is a large-scale Portuguese audio corpus derived from the “Cem Mil Podcasts” collection
Parakeet v2 → trainedon → TAGARELA
confidence 95% · the finetuned Parakeet v2 yielded the best performance on the TAGARELA test set.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Despite significant advances in speech processing, Portuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over 8,972 hours of podcast audio, specifically curated for training automatic speech recognition (ASR) and text-to-speech (TTS) models. Notably, its scale rivals English's GigaSpeech (10kh), enabling state-of-the-art Portuguese models. To ensure data quality, the corpus was subjected to an audio pre-processing pipeline and subsequently transcribed using a mixed strategy: we applied ASR models that were previously trained on high-fidelity transcriptions generated by proprietary APIs, ensuring a high level of initial accuracy. Finally, to validate the effectiveness of this new resource, we present ASR and TTS models trained exclusively on our dataset and evaluate their performance, demonstrating its potential to drive the development of more robust and natural speech technologies for Portuguese. The dataset is released publicly, available at this https URL, to foster the development of robust speech technologies.
Tags
Links
- Source: https://arxiv.org/abs/2603.15326v1
- Canonical: https://arxiv.org/abs/2603.15326v1
Trouble viewing inline? Open PDF directly →
Full Text
23,999 characters extracted from source content.
Expand or collapse full text
TAGARELA - A PORTUGUESE SPEECH DATASET FROM PODCASTS Frederico Santos de Oliveira ⋆ Lucas Rafael Stefanel Gris † Alef Iury Siqueira Ferreira † Augusto Seben da Rosa ‡ Alexandre Costa Ferro Filho † Edresson Casanova § Christopher Dane Shulby ¶ Rafael Teixeira Sousa ⋆ Diogo Fernandes Costa Silva † Anderson da Silva Soares † Arlindo Rodrigues Galv ̃ ao Filho † ⋆ Federal University of Mato Grosso (UFMT) † Federal University of Goias (UFG) ‡ Paulista State University (UNESP) § NVIDIA ¶ Elsa Speak ABSTRACT Despite significant advances in speech processing, Por- tuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over 8,972 hours of podcast audio, specifically curated for train- ing automatic speech recognition (ASR) and text-to-speech (TTS) models. Notably, its scale rivals English’s GigaSpeech (10kh), enabling state-of-the-art Portuguese models. To en- sure data quality, the corpus was subjected to an audio pre- processing pipeline and subsequently transcribed using a mixed strategy: we applied ASR models that were previously trained on high-fidelity transcriptions generated by propri- etary APIs, ensuring a high level of initial accuracy. Finally, to validate the effectiveness of this new resource, we present ASR and TTS models trained exclusively on our dataset and evaluate their performance, demonstrating its potential to drive the development of more robust and natural speech technologies for Portuguese. The dataset is released publicly 1 to foster the development of robust speech technologies. Index Terms— speech processing, text-to-speech, dataset, automatic-speech-recognition 1. INTRODUCTION Portuguese ranks among the most widely spoken languages globally, with hundreds of millions of speakers across several continents. Despite this global prominence, it remains sig- nificantly under-resourced in the field of speech technology when compared to English. Recent advances in deep learning have propelled the fields of Automatic Speech Recognition (ASR) and Text-to-Speech (TTS), but their progress is funda- mentally driven by the availability of large-scale, high-quality speech datasets. High-resource languages like English benefit from extensive corpora such as LibriSpeech (1000 hours) [1], GigaSpeech (10k hours) [2] and Emilia (139k hours) [3], which comprise tens of thousands of hours of transcribed 1 Available at https://freds0.github.io/TAGARELA/ audio. This resource gap creates a significant bottleneck that hinders the development of robust and natural-sounding speech technologies tailored to the linguistic nuances of Por- tuguese. To address this disparity, the community has released Portuguese corpora focused on spontaneous speech. Land- mark initiatives include CORAA [4] (290h), NURC-SP [5] (239h), MuPe [6] (365h), and VoxCeleb-PT [7] (18h). Al- though these datasets provide high-quality data crucial for ASR, their spontaneous nature — often characterized by dis- fluencies, background noise, and interruptions — makes them ill-suited for training TTS models, which conventionally re- quire clean, high-quality audio to achieve natural synthesis. Furthermore, being orders of magnitude smaller than their English counterparts, these corpora limit the performance of data-hungry state-of-the-art models. A significant opportunity to bridge this data gap emerged with the release of the “Cem Mil Podcasts” collection [8], a massive corpus offering over 76,000 hours of diverse, multi-dialect Portuguese audio. However, the dataset was provided as raw, unprocessed audio with automatically gen- erated transcripts of varying quality. The absence of essential pre-processing — speaker diarization, noise reduction, and overlapping speech removal — rendered it impractical for training high-performance ASR and TTS models, which re- quire clean, single-speaker segments. To unlock this resource, we introduce TAGARELA (/ta.ga."RE.lA/), a large-scale Portuguese speech dataset cu- rated from “Cem Mil Podcasts”. We developed a pipeline — including standardization, diarization, overlap detection, and denoising — to transform raw audio into a high-quality cor- pus. For transcription, we employed a bootstrap approach: a 1,000-hour seed corpus, transcribed by commercial ASR, was used to fine-tune Whisper large-v3 [9], generating pseudo- labels for the remaining data. The corpus comprises two subsets: a full 8,972-hour set with disfluencies for robust ASR, and a 2,800-hour clean-speech subset for speech gener- ation. The main contributions of this work are: (1) the release arXiv:2603.15326v1 [cs.CL] 16 Mar 2026 Overlapping Detector OverlapNo Overlap 1 2 3 4 Denoiser ASR 5 6 Standarlization Seg. 1Seg. 2Seg. 3 Segmentation Speaker Diarization Speaker 1Speaker 2 Speaker 3 Transcription Speech Enhancement Overlapping Speech Detection Fig. 1. Overview of the TAGARELA preprocessing pipeline. of TAGARELA, a curated Portuguese speech corpus exceed- ing 8,972 hours for ASR and TTS; (2) a detailed description and evaluation of our multi-stage processing pipeline; and (3) training and evaluation of open-source models exclusively on TAGARELA, demonstrating its effectiveness. 2. TAGARELA DATASET The TAGARELA dataset is a large-scale Portuguese audio corpus derived from the “Cem Mil Podcasts” collection [8] and released exclusively for research purposes to address the lack of non-English podcast resources. It consists of roughly 16,806 episodes from 2,094 shows, totaling over 8,972 hours of audio. The data includes both Brazilian (8,130 hours) and European Portuguese (842 hours) dialects. In terms of gen- der, 70% of the audio (6,368.34 hours) is attributed to male speakers and 30% (2,604.37 hours) to female speakers. The dataset’s audio segments have an average duration of 9.30± 5.49 seconds and contain an average of 27.69± 17.06 words. Figure 2 presents the distribution of audio duration according to gender and accent. Fig. 2. Violin plots showing the distribution of audio segment duration in seconds. The left plot compares accents (pt br vs. pt pt), and the right plot compares genders (male vs. female). 3. TAGARELA PIPELINE The creation of the TAGARELA dataset involved a multi- stage pipeline designed to ensure high quality and consistency for both ASR and TTS tasks. Each stage was planned to han- dle the challenges inherent to podcast audio, such as multiple speakers, background noise, and the need for accurate, large- scale transcriptions. An overview of the pipeline is shown in Figure 1, and we detail each component below. 3.1. Audio Standardization and Segmentation All audio files were converted to a uniform format (FLAC, 16kHz, 16-bit, mono) to ensure training consistency. Long- form recordings were then segmented into 5–20 second clips, with the algorithm prioritizing splits at natural silences to pre- serve speech cohesiveness. 3.2. Diarization and Speaker Separation A common feature in podcasts is the presence of multiple speakers, which poses a challenge for creating clean datasets. To address this, we applied a diarization process using the pyannote framework [10]. This stage identifies and labels the speech segments for each speaker individually. By separat- ing the different speakers, we ensure that each final sample in the dataset contains the voice of only one person. This step is particularly crucial for training TTS models, which require single-speaker data to generate consistent voices. 3.3. Overlapping Speech Detection Although diarization separates speakers, some segments may still contain overlapping speech where multiple speakers talk simultaneously.This is highly detrimental to the quality of TTS models. To mitigate this issue, we trained a dedi- cated classification model based on Wav2vec2-XLS-R [11] to specifically identify these instances. All audio samples that were flagged by the model as containing overlapping speech were subsequently discarded from the dataset, ensuring that the final clips consist of clean, single-speaker utterances. To ensure the reproducibility of this filtering step, the trained model and its checkpoint are made publicly available for download. 3.4. Two-Stage Bootstrap Transcription We employed a bootstrap strategy for transcription. First, a high-fidelity “seed corpus” of approximately 1000 hours was generated with a commercial ASR service, ElevenLabs Scribe v1 2 , to fine-tune a Whisper large-v3 model for pseudo- labeling. To filter Whisper’s potential hallucinations [12], we also trained a Wav2vec2-XLS-R model 3 [13] on the same seed data. We then calculated the Word and Character Error Rates (WER/CER) between the outputs of both models, mak- ing it possible to select only samples with a high agreement score to ensure the training dataset’s quality. 3.5. Quality Enhancement For the final audio enhancement stage, we repurposed a Vo- cos vocoder [14] to act as a denoiser. The model was trained specifically for this task on a private dataset, optimizing it to remove common podcast artifacts like background noise, hiss, and light reverberation. This process significantly im- proves the clarity and quality of the finalized segments. To ensure the reproducibility of this pipeline, we also provide a version of the denoiser trained exclusively on public datasets. 3.6. Speaker and Dialect Labeling To enrich the dataset, we implemented a multi-stage labeling process. First, given that the original data lacks speaker la- bels, we performed a cluster-based speaker labeling step. We extracted embeddings for each audio segment using the Red- imNet B6 model [15] and grouped them with the HDBSCAN algorithm [16]. This process was performed independently for each podcast to avoid merging speaker identities, result- ing in approximately 13,368 distinct speaker labels. Furthermore, we developed a model to classify each seg- ment’s dialect as either Brazilian or European Portuguese. This was achieved by first pre-training a wav2vec-base model on all segmented audio from our dataset.Subsequently, this model was fine-tuned on a balanced combination of the CORAA, CommonVoice [17], and CML-TTS [18] datasets to create the final accent classifier thus adding valuable dialectal metadata to the corpus. 4. EXPERIMENTS AND EVALUATION In this section, we validate the quality and effectiveness of the TAGARELA dataset by using it to train state-of-the-art models for ASR and TTS. Our goal is to demonstrate that the data curated through our pipeline can produce models with competitive or state-of-the-art performance for Portuguese. 4.1. Objective Metrics We assess audio quality using three objective metrics: Short- Time Objective Intelligibility (STOI) [19] for speech in- telligibility, the wideband version of Perceptual Evaluation 2 https://elevenlabs.io/docs/models#scribe-v1, run in June 2025. 3 https://huggingface.co/facebook/wav2vec2-xls-r-1b of Speech Quality (PESQ) [20, 21] for perceived quality, and Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) [22] for signal fidelity in decibels (dB). For all metrics, higher values indicate better quality. Since these traditionally re- quire a clean reference signal, we employ the TorchAudio- Squim [23] framework to obtain reference-free estimates. The results are presented in Figure 3. Fig. 3. Violin plots showing STOI, PESQ and SI-SDR. 4.2. Speech Recognition (ASR) Experiments To evaluate the potential of TAGARELA for ASR tasks, we assessed a diverse set of model architectures covering differ- ent sizes and capabilities, including Parakeet TDT [24, 25], Wav2Vec Large [11], Distil-Whisper [26], Whisper Large V3 [9], and Parakeet v3 [27]. Training Setup: We fine-tuned Distil-Whisper, Parakeet TDT v2, and Wav2Vec Large using the full 8,972-hour TAGARELA dataset. In contrast, Whisper Large V3 and Parakeet v3 were evaluated as pre-trained baselines without further training. All experiments were conducted using either NVIDIA A100 or B200 GPUs, subject to availability. Evaluation: We evaluated model performance on the TAGARELA test set manually transcribed, using WER and CER, calculated after text normalization. Results: As detailed in Table 1, the finetuned Parakeet v2 yielded the best performance on the TAGARELA test set. Achieving a WER of 15.18% and CER of 7.09%, it outper- formed all other models, including Wav2Vec Large, Distil- Whisper, and Whisper Large V3, proving to be highly effec- tive for Portuguese speech recognition. Table 1. WER results on the TAGARELA test set. ModelWER (%)↓ CER (%)↓ Whisper Large V320.9112.42 Wav2Vec Large FT21.858.55 Distil-Whisper FT20.0211.18 Parakeet v323.3014.86 Parakeet v2 FT15.187.09 4.3. Text-to-Speech (TTS) Experiments Training and Evaluation Setup: For the TTS task, we trained the Orpheus-TTS 4 and Chatterbox 5 models using the 2,800- 4 https://github.com/canopyai/Orpheus-TTS 5 https://github.com/resemble-ai/chatterbox hour clean-speech subset of the TAGARELA dataset. We assessed intelligibility using WER/CER (Whisper Large V3) and perceptual quality using a Mean Opinion Score (MOS), for which 50 evaluators rated 40 samples. To ensure a ro- bust evaluation, we applied a two-stage outlier removal pro- cess, filtering samples shorter than five seconds, and applying a quartile-based method to remove statistical outliers. Table 2. TTS model performance. The values for CER, WER, and MOS are presented as mean± standard deviation. ModelWER (%)↓CER (%)↓MOS↑ Chatterbox 0.3111± 0.442 0.268± 0.423 4.176± 0.983 Orpheus-TTS 0.095± 0.100 0.046± 0.051 4.155± 1.001 Ground Truth 0.010± 0.033 0.006± 0.018 4.231± 1.001 Results and Analysis: As shown in Table 2, Orpheus-TTS achieved superior intelligibility (9.5% WER), while Chatter- box attained slightly higher naturalness (MOS 4.176) despite significant errors. This reveals a trade-off between Orpheus- TTS’s linguistic precision and Chatterbox’s focus on prosody. These experiments validate the TAGARELA dataset’s critical role in advancing Portuguese TTS. Despite the dataset’s imperfect text-audio alignment, the results are highly encour- aging and provide a solid foundation for developing robust, high-quality TTS systems for Portuguese. 5. CONCLUSION To address the resource gap in Portuguese speech technology, we introduce TAGARELA, a new large-scale dataset with over 8,972 hours of podcast audio. We presented a compre- hensive pipeline using diarization, denoising, and a scalable transcription strategy to create a high-quality corpus suitable for both ASR and TTS. The public release of this dataset is a significant contribution, offering the community a resource on a scale previously unavailable for the Portuguese language. The effectiveness of TAGARELA was validated by train- ing ASR and TTS models exclusively on our data, achieving highly competitive performance. This confirms the dataset’s potential to drive significant advancements in Portuguese speech processing.While there is room for refinements, such as improving text-audio alignment, TAGARELA of- fers a robust foundation for future innovations. We believe this resource will foster the development of more accurate and natural speech technologies, benefiting millions of Por- tuguese speakers. Acknowledgements: This work has been fully funded by the project Research and Development of Algorithms for Construction of Digital Human Technological Components supported by the Advanced Knowledge Center in Immersive Technologies (AKCIT), with financial resources from the PPI IoT of the MCTI grant number 057/2023, signed with EM- BRAPII. 6. REFERENCES [1] Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, “Librispeech: An asr corpus based on public do- main audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, p. 5206–5210. [2] Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu, Chen Yang, Jiaqi Li, Peiyang Shi, Yuancheng Wang, Kai Chen, Pengyuan Zhang, and Zhizheng Wu, “Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,” in 2024 IEEE Spoken Language Technology Workshop (SLT), 2024, p. 885–890. [3] Guoguo Chen, Shuzhou Chai, Guan-Bo Wang, Jiayu Du, Wei- Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, Mingjie Jin, Sanjeev Khudanpur, Shinji Watan- abe, Shuaijiang Zhao, Wei Zou, Xiangang Li, Xuchen Yao, Yongqing Wang, Zhao You, and Zhiyong Yan, “Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,” in Interspeech 2021. Aug. 2021, inter- speech 2021, ISCA. [4] Arnaldo Candido Junior, Edresson Casanova, Anderson Soares, Frederico Santos de Oliveira, Lucas Oliveira, Ri- cardo Corso Fernandes Junior, Daniel Peixoto Pinto da Silva, Fernando Gorgulho Fayet, Bruno Baldissera Carlotto, Lucas Rafael Stefanel Gris, and Sandra Maria Alu ́ ısio, “Coraa asr: a large corpus of spontaneous and prepared speech manually validated for speech recognition in brazilian portuguese,” Lan- guage Resources and Evaluation, vol. 57, no. 3, p. 1139– 1171, 2023. [5] Rodrigo Lima, Sidney Evaldo Leal, Arnaldo Candido Junior, and Sandra Maria Aluisio, “A large dataset of spontaneous speech with the accent spoken in s ̃ ao paulo for automatic speech recognition evaluation,” in Proceedings of 34th Brazil- ian Conference on Intelligent Systems (BRACIS), 2024. [6] “MuPe life stories dataset: Spontaneous speech in Brazilian Portuguese with a case study evaluation on ASR bias against speakers groups and topic modeling,” in Proceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al- Khalifa, Barbara Di Eugenio, and Steven Schockaert, Eds., Abu Dhabi, UAE, Jan. 2025, p. 6076–6087, Association for Computational Linguistics. [7] John Mendonc ̧a and Isabel Trancoso, “Voxceleb-pt–a dataset for a speech processing course,” Proc. IberSPEECH, vol. 2022, p. 71–75, 2022. [8] Ekaterina Garmash, Edgar Tanaka, Ann Clifton, Joana Cor- reia, Sharmistha Jat, Winstead Zhu, Rosie Jones, and Jussi Karlgren, “Cem mil podcasts: A spoken portuguese docu- ment corpus for multi-modal, multi-lingual and multi-dialect information access research,” in Experimental IR Meets Mul- tilinguality, Multimodality, and Interaction, Cham, 2023, p. 48–59, Springer Nature Switzerland. [9] Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever,“Robust speech recognition via large-scale weak supervision,” in International conference on machine learning. PMLR, 2023, p. 28492– 28518. [10] Herv ́ e Bredin, Ruiqing Yin, Juan Manuel Coria, Gregory Gelly, Pavel Korshunov, Marvin Lavechin, Diego Fustes, Hadrien Titeux, Wassim Bouaziz, and Marie-Philippe Gill, “Pyannote.audio: Neural building blocks for speaker diariza- tion,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, p. 7124–7128. [11] Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakho- tia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Miguel Pino, Alexei Baevski, Alexis Conneau, and Michael Auli, “XLS-R: Self-supervised cross-lingual speech representation learning at scale,” in Proc. Interspeech 2022, 2022, p. 2278–2282. [12] Mateusz Bara ́ nski, Jan Jasi ́ nski, Julitta Bartolewska, Stanisław Kacprzak, Marcin Witkowski, and Konrad Kowalczyk, “Inves- tigation of whisper asr hallucinations induced by non-speech audio,” 04 2025, p. 1–5. [13] Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdel- rahman Mohamed, and Michael Auli, “Unsupervised cross- lingual representation learning for speech recognition,”in Proc. Interspeech 2021, 2021, p. 2426–2430. [14] Hubert Siuzdak,“Vocos: Closing the gap between time- domain and fourier-based neural vocoders for high-quality au- dio synthesis,” in International Conference on Representa- tion Learning, B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun, Eds., 2024, vol. 2024, p. 25719–25733. [15] Ivan Yakovlev, Rostislav Makarov, Andrei Balykin, Pavel Malov, Anton Okhotnikov, and Nikita Torgashov, “Reshape dimensions network for speaker recognition,” in Interspeech 2024, 2024, p. 3235–3239. [16] Claudia Malzer and Marcus Baum, “A hybrid approach to hi- erarchical density-based cluster selection,” in 2020 IEEE In- ternational Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI), 2020, p. 223–228. [17] Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saun- ders, Francis Tyers, and Gregor Weber,“Common voice: A massively-multilingual speech corpus,” in Proceedings of the Twelfth Language Resources and Evaluation Conference, Nicoletta Calzolari, Fr ́ ed ́ eric B ́ echet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hi- toshi Isahara, Bente Maegaard, Joseph Mariani, H ́ el ` ene Mazo, Asuncion Moreno, Jan Odijk, and Stelios Piperidis, Eds., Mar- seille, France, May 2020, p. 4218–4222, European Language Resources Association. [18] Frederico S Oliveira, Edresson Casanova, Arnaldo Candido Junior, Anderson S Soares, and Arlindo R Galv ̃ ao Filho, “CML-TTS: A multilingual dataset for speech synthesis in low-resource languages,” in International Conference on Text, Speech, and Dialogue. Springer, 2023, p. 188–199. [19] Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jes- per Jensen, “A short-time objective intelligibility measure for time-frequency weighted noisy speech,” in 2010 IEEE inter- national conference on acoustics, speech and signal process- ing. IEEE, 2010, p. 4214–4217. [20] Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of tele- phone networks and codecs,” in 2001 IEEE international con- ference on acoustics, speech, and signal processing. Proceed- ings (Cat. No. 01CH37221). IEEE, 2001, vol. 2, p. 749–752. [21] I. Rec, “P.862.2: Wideband extension to recommendation P.862 for the assessment of wideband telephone networks and speech codecs,” Recommendation P.862.2, International Telecommunication Union, Geneva, Switzerland, 2005. [22] Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey, “Sdr–half-baked or well done?,” in ICASSP 2019- 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, p. 626–630. [23] Anurag Kumar, Ke Tan, Zhaoheng Ni, Pranay Manocha, Xi- aohui Zhang, Ethan Henderson, and Buye Xu, “Torchaudio- squim: Reference-less speech quality and intelligibility mea- sures in torchaudio,” in ICASSP 2023-2023 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, p. 1–5. [24] Dima Rekesh, Nithin Rao Koluguri, Samuel Kriman, Somshubra Majumdar, Vahid Noroozi, He Huang, Oleksii Hrinchuk, Krishna Puvvada, Ankur Kumar, Jagadeesh Balam, and Boris Ginsburg, “Fast conformer with linearly scalable attention for efficient speech recognition,”in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2023, p. 1–8. [25] Hainan Xu, Fei Jia, Somshubra Majumdar, He Huang, Shinji Watanabe, and Boris Ginsburg, “Efficient sequence transduc- tion by jointly predicting tokens and durations,” in Proceed- ings of the 40th International Conference on Machine Learn- ing. 2023, ICML’23, JMLR.org. [26] Sanchit Gandhi, Patrick Von Platen, and Alexander M Rush, “Distil-whisper: Robust knowledge distillation via large-scale pseudo labelling,” arXiv preprint arXiv:2311.00430, 2023. [27] Monica Sekoyan, Nithin Rao Koluguri, Nune Tadevosyan, Piotr Zelasko, Travis Bartley, Nikolay Karpov, Jagadeesh Balam, and Boris Ginsburg, “Canary-1b-v2 & parakeet-tdt-0.6 b-v3: Efficient and high-performance models for multilingual asr and ast,” arXiv preprint arXiv:2509.14128, 2025.