Paper deep dive
Make It Hard to Hear, Easy to Learn: Long-Form Bengali ASR and Speaker Diarization via Extreme Augmentation and Perfect Alignment
Sanjid Hasan, Risalat Labib, A H M Fuad, Bayazid Hasan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 9:26:05 AM
Summary
This paper addresses the scarcity of joint Automatic Speech Recognition (ASR) and speaker diarization resources for Bengali by introducing Lipi-Ghor-882, an 882-hour multi-speaker dataset. The authors detail their methodology for the DL Sprint 4.0 competition, demonstrating that for ASR, targeted fine-tuning with synthetic acoustic degradation (noise/reverberation) and perfect alignment is superior to raw data scaling or ensembling. For speaker diarization, they found that state-of-the-art models like Diarizen performed poorly, and retraining yielded negligible gains; instead, heuristic post-processing of baseline model outputs (specifically Pyannote Community-1) was the primary driver for accuracy. The optimized dual pipeline achieved a Real-Time Factor (RTF) of approximately 0.019.
Entities (19)
Relation Signals (10)
Lipi-Ghor-882 â hasduration â 882 hours
confidence 98% ¡ a comprehensive 882-hour multi-speaker Bengali dataset
Lipi-Ghor-882 â createdfor â Bengali ASR and Speaker Diarization
confidence 95% ¡ To address the severe scarcity of joint ASR and diarization resources for this language, we introduce Lipi-Ghor-882
Pyannote Community-1 â usedwith â heuristic post-processing
confidence 93% ¡ Pyannote Community-1 proved more robust... paired with a strict custom post-processing algorithm
Whisper Medium â achievedmetric â 0.019 RTF
confidence 92% ¡ dropping Whisperâs inference time from 4 hours to just 26 minutes (an RTF of âź0.019)
Synthetic Acoustic Degradation â improves â Whisper-Medium ASR accuracy
confidence 91% ¡ targeted fine-tuning utilizing perfectly aligned annotations paired with synthetic acoustic degradation (noise and reverberation) emerges as the singular most effective approach
Whisper Medium â optimizedby â Faster-Whisper
confidence 90% ¡ utilized CTranslate2 (CT2) and faster-whisper with parallel processing
Diarizen â performedpoorlyon â Bengali long-form audio
confidence 90% ¡ Diarizen-the supposed open-source state-of-the-art performed poorly on both the public and private leaderboards
Contextual Biasing â â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Although Automatic Speech Recognition (ASR) in Bengali has seen significant progress, processing long-duration audio and performing robust speaker diarization remain critical research gaps. To address the severe scarcity of joint ASR and diarization resources for this language, we introduce Lipi-Ghor-882, a comprehensive 882-hour multi-speaker Bengali dataset. In this paper, detailing our submission to the DL Sprint 4.0 competition, we systematically evaluate various architectures and approaches for long-form Bengali speech. For ASR, we demonstrate that raw data scaling is ineffective; instead, targeted fine-tuning utilizing perfectly aligned annotations paired with synthetic acoustic degradation (noise and reverberation) emerges as the singular most effective approach. Conversely, for speaker diarization, we observed that global open-source state-of-the-art models (such as Diarizen) performed surprisingly poorly on this complex dataset. Extensive model retraining yielded negligible improvements; instead, strategic, heuristic post-processing of baseline model outputs proved to be the primary driver for increasing accuracy. Ultimately, this work outlines a highly optimized dual pipeline achieving a $\sim$0.019 Real-Time Factor (RTF), establishing a practical, empirically backed benchmark for low-resource, long-form speech processing.
Tags
Links
- Source: https://arxiv.org/abs/2602.23070v1
- Canonical: https://arxiv.org/abs/2602.23070v1
Trouble viewing inline? Open PDF directly â
Full Text
17,861 characters extracted from source content.
Expand or collapse full text
Make It Hard to Hear, Easy to Learn: Long-Form Bengali ASR and Speaker Diarization via Extreme Augmentation and Perfect Alignment Sanjid Hasan E, KUET sanjid9657@gmail.com Risalat Labib CSE, BUET risalatlabib@gmail.com A H M Fuad CSE, BUET ahmfuad9@gmail.com Bayazid Hasan CSE, BUET bhsarf@gmail.com AbstractâAlthough Automatic Speech Recognition (ASR) in Bengali has seen significant progress, processing long-duration audio and performing robust speaker diarization remain critical research gaps. To address the severe scarcity of joint ASR and diarization resources for this language, we introduce Lipi-Ghor- 882, a comprehensive 882-hour multi-speaker Bengali dataset. In this paper, detailing our submission to the DL Sprint 4.0 competition, we systematically evaluate various architectures and approaches for long-form Bengali speech. For ASR, we demonstrate that raw data scaling is ineffective; instead, targeted fine-tuning utilizing perfectly aligned annotations paired with synthetic acoustic degradation (noise and reverberation) emerges as the singular most effective approach. Conversely, for speaker diarization, we observed that global open-source state-of-the-art models (such as Diarizen) performed surprisingly poorly on this complex dataset. Extensive model retraining yielded negligible improvements; instead, strategic, heuristic post-processing of baseline model outputs proved to be the primary driver for increasing accuracy. Ultimately, this work outlines a highly optimized dual pipeline achieving aâź0.019 Real-Time Factor (RTF), establishing a practical, empirically backed benchmark for low-resource, long-form speech processing. Index TermsâBengali ASR, Speaker Diarization, Long-form Audio, Noise Augmentation, Pyannote, Faster-Whisper I. INTRODUCTION The transcription and segmentation of multi-speaker Bengali audio is a highly complex task, historically constrained by a lack of large-scale, temporally aligned conversational datasets. In the DL Sprint 4.0 (BUET CSE FEST 2026) competition, models were rigorously evaluated on a massive 22-hour hidden test set consisting of long-form audio. This paper details the journey of Team Villagers a process of exhaustive trial and error that forced us to discard conventional wisdom, leading to our core methodology: Make it hard to hear, easy to learn. Our journey began with a straightforward evaluation of pre- trained Bengali Automatic Speech Recognition (ASR) mod- els. We tested Moonshine, Hishab-Titu-Bn-Conformer-Large, Wav2Vec2, and Whisper-Medium (fine-tuned by Tugstugi). The initial results presented a stark trade-off between speed and accuracy. Moonshine was blisteringly fast (5 min- utes for the 22-hour set) but performed very poorly. Titu- FastConformer was exceptionally fast (17 minutes) and competitive. Wav2Vec2 took 2 hours with chunking, while Whisper-Medium delivered the best accuracy but took an agonizing 4 hours to run. Because the Real-Time Factor (RTF) and inference time were critical competition metrics, our next phase was pure op- timization. We aggressively optimized Wav2Vec2 with batch- ing to drop its time to 1 hour, but Titu-FastConformer re- mained superior in speed. However, because Whisper-Medium held the highest accuracy potential, we focused our engineer- ing efforts there. Transitioning to Kaggleâs dual T4 environ- ment, we utilized CTranslate2 (CT2) and faster-whisper with parallel processing. To handle long-form context, we experimented with Silero VAD and Pyannote VAD for chunk- ing; Silero VAD proved far more reliable. This engineering phase was a massive success, dropping Whisperâs inference time from 4 hours to just 26 minutes (an RTF of âź0.019). We also discovered that faster-whisper outperformed whisperX for our use case, though we had to manually adjust the default VAD parameters, which were initially truncating the first words of speech chunks. With inference optimized, we entered the training phase, which quickly became a graveyard of failed hypotheses. We attempted Parameter-Efficient Fine-Tuning (PEFT) with a frozen decoder, which yielded negligible gains. We integrated Contextual Biasing [1], which initially improved our scores but began degrading performance catastrophically after a certain training threshold, completely failing on out-of-distribution data. We tried applying Demucs for music and noise removal; while it improved the ASR score, the computational overhead inflated inference time beyond acceptable limits, forcing us to abandon it. Finally, we attempted to ensemble all our best model outputs using ROVER, but to our surprise, the ensembled predictions performed worse than our standalone models. Our breakthrough in ASR finally came when we shifted our focus from the models to the data. We trained Whisper on a small, perfectly aligned subset of data where 20% of the audio was artificially corrupted with random noise and reverberation. This counter-intuitive approach forced the model to rely on deep phonetic features rather than acoustic memorization, leading to a significant score boost. We continued this training until we hit a ceiling caused by the poor annotation quality of the broader dataset. Simultaneously, our speaker diarization journey was fraught with similar challenges. We tested a suite of models in- arXiv:2602.23070v1 [cs.SD] 26 Feb 2026 cluding ECAPA-TDNN, Pyannote 3.1, Pyannote Community- 1, and Diarizen. Surprisingly, Diarizen-the supposed open- source state-of-the-art performed poorly on both the public and private leaderboards. Integrating a Bengali-trained WavLM with Diarizen or using VBx with a Pyannote pipeline yielded no changes. In stark contrast to ASR, using Demucs for noise removal actually worsened our diarization scores. Hoping to force robustness, we trained the segmentation model on muf- fled audio for 20 epochs, but observed negligible improvement in the Diarization Error Rate (DER). Ultimately, we realized that for Bengali diarization, model retraining was a dead end. Success relied entirely on heuristic post-processing. While Pyannote 3.1 topped the public leader- board, Pyannote Community-1 proved more robust for the private leaderboard. We paired this base model with a strict custom post-processing algorithm that forced inter-speaker gaps, merged same-speaker micro-segments, and ruthlessly mitigated overlaps. This exhaustive journey from the failure of global SOTAs to the success of noisy training and algorithmic post-processing highlighted a severe lack of joint ASR and diarization data in the Bengali domain. To bridge this gap, and as the final piece of our contribution, we created Lipi-Ghor-882, an 882-hour dataset engineered utilizing yt-dlp and the Pyannote API, providing the foundation for our methodology. I. THE LIPI-GHOR-882 DATASET A significant research gap exists in the joint optimization of ASR and diarization for Bengali. To address this, we cu- rated Lipi-Ghor-882. We collected diverse long, medium, and short audio files from YouTube spanning multiple domains. Transcriptions were extracted utilizing the yt-dlp library. To establish speaker boundaries, we utilized the current SOTA Pyannote API, carefully merging these temporal annotations with the speech segments. [2] TABLE I LIPI-GHOR DATASET STATISTICS AttributeValueAttributeValue Total Hoursâź882Unique Channels596 Annotated Hoursâź856Categorized Domains150+ Total Videos1,019Annotation FormatSSTT I. METHODOLOGY: AUTOMATIC SPEECH RECOGNITION A. Base Model Evaluation and Inference Profiling We initially evaluated multiple pre-trained Bengali base models on a 22-hour validation set. Inference time and the Real-Time Factor (RTF) were critical constraints. ⢠Moonshine: Highly efficient (5 minutes inference) but exhibited very poor transcription accuracy. ⢠Hishab-Titu-Bn-Conformer-Large: Exceptionally fast (17 minutes) with competitive accuracy. ⢠Wav2Vec2: Slower (2 hours with chunking) and moder- ate accuracy. ⢠Whisper-Medium (Tugstugi): Longest base inference (4 hours) but yielded the highest transcription fidelity. For sequence chunking, we compared Silero VAD and Pyan- note VAD; Silero VAD demonstrated superior boundary de- tection for our specific use case. B. Optimization and Pipeline Engineering Given Whisperâs superior accuracy and Titu-Conformerâs speed, we focused on optimizing their inference. By con- verting the Whisper-Medium model to CTranslate2 (CT2) format and deploying it via faster-whisper on dual T4 GPUs (parallel processing), we reduced the 4-hour inference time to 26 minutes (RTF âź0.019). We observed that the default VAD parameters in faster-whisper aggressively truncated initial words in speech chunks; manually fine-tuning these parameters was necessary to prevent word-deletion er- rors. C. Training Ablations and the Noise Paradigm Extensive training ablations were conducted to maximize the Whisper-Medium architecture: ⢠Frozen Decoder: Training with a frozen decoder using adapters yielded improvements but not impressive. ⢠Contextual Biasing: Implementing contextual biasing [1] improved initial scores but caused catastrophic degrada- tion after a training threshold. ⢠Audio Separation (Demucs): While applying Demucs to remove background music improved the WER, it drastically inflated inference time and was discarded. ⢠Ensembling (ROVER): Attempting to ensemble our best predictions via ROVER degraded the final score. ⢠Final Strategy (Corrupted Audio): Training the model on a small, perfectly aligned dataset where 20% of the audio was artificially corrupted (noise and reverberation) forced the model to learn robust acoustic features, leading to our best score. IV. METHODOLOGY: SPEAKER DIARIZATION A. Evaluation of Global SOTAs We evaluated several standard and SOTA diarization mod- els, including ECAPA-TDNN, Pyannote 3.1, Pyannote Com- munity 1, and Diarizen. Despite Diarizen being recognized as a leading open-source SOTA [3], it performed poorly on both leaderboards for this dataset. Integrating a Bengali-trained WavLM into Diarizen or testing VBx [4] with a Pyannote pipeline resulted in no performance shift. Furthermore, unlike in ASR, utilizing Demucs for noise removal worsened the diarization score. B. The Failure of Retraining We attempted to fine-tune the Pyannote segmentation model utilizing artificially muffled data to improve robustness. How- ever, even after 20 epochs, no improvement in DER was observed. Fig. 1. Duration Classification of Audios C. Algorithmic Post-Processing (The Strategy) Because model retraining failed, our final solution relied entirely on Pyannote Community-1 paired with a highly struc- tured overlap mitigation and segment merging strategy. We converted our pipeline logic into the following heuristic post- processing algorithm to force strict inter-speaker gaps and filter micro-segments. Algorithm 1: Strict Gap Post-Processing 1) Input: Set of detected segments S, thresholds θ merge = 3.79s, θ gap = 0.17s, θ seg = 0.75s, θ spk = 9.0s. 2) Sort: Order S chronologically by start time. 3) Strict Rename: Reassign speaker IDs serially (e.g., SPEAKER 0, SPEAKER1) based on first appearance. 4) Overlap Resolution & Gap Enforcement: ⢠Initialize timeline end = 0.0, lastspk = Null ⢠For each segment in S: ⢠If starttime < timelineend, set starttime = timelineend. ⢠If speaker id̸= lastspk: ⢠reqstart = timelineend + θ gap ⢠If start time < reqstart, set starttime = reqstart. ⢠Update timeline end = endtime, lastspk = speakerid. 5) Segment Merging: Merge adjacent segments of the same speaker if the temporal gap is < θ merge . 6) Duration Filtering: Discard any individual segment where (end timeâ starttime) < θ seg . 7) Speaker Pruning: Discard all segments belonging to any speaker with a total speaking duration < θ spk . 8) Output: Cleaned and rounded temporal segments. V. EXPERIMENTS AND RESULTS Trainings were done using L40S GPU, 48GB VRAM. All inference processes were evaluated under strict hardware constraints (2x T4 GPUs). The separation of the ASR and Diarization pipelines allowed for maximum computational efficiency without relying on complex, end-to-end joint mod- eling. TABLE I ASR PERFORMANCE BENCHMARKS ON 22-HOURS TEST SET ModelRTFPub WERPriv WER Whisper-Medium (Tugstugi)0.1820.444430.45064 Wav2Vec20.0450.513890.52338 Titu-STT-BN-Conformer-Large0.0050.444300.44550 FT Whisper-Med (Adapter)0.0190.310730.32145 FT Titu-STT-BN-Conformer-Large0.0050.406620.41434 FT Whisper-Med (Contextual)0.0190.355780.38231 ROVER Ensemble0.0190.337700.35762 FT Whisper-Med (Faster Whisper)*0.0190.308420.31070 FT Whisper-Med (Faster Whip. + Demucs)0.0680.303540.30612 FT Whisper-Med (WhisperX)0.0190.304060.31479 TABLE I SPEAKER DIARIZATION BENCHMARKS ModelRetrainedPost-ProcPub DERPriv DERRTF DiarizenYes (WavLM)Yes0.280770.278930.020 Pyannote 3.1YesYes0.202830.281450.019 Pyannote Comm-1YesYes0.245220.266400.019 ECAPA-TDNNNoYes0.253220.316400.012 VI. CHALLENGES AND LIMITATIONS During the competition, we encountered several challenges and limitations that impacted the overall performance and experimentation process: A. Compute Resource Constraints Our best model was trained on a single L40S GPU with 48 GB VRAM, which was limited to only 5 free hours per month from Lightening AI. This constraint made extensive experimentation and hyperparameter tuning difficult, limiting our ability to fully optimize the model. More compute hours could have significantly improved model performance. B. Data-Related Challenges The dataset consisted of multiple sources with varying quality, including missing or partial transcriptions. Handling these inconsistencies required significant preprocessing effort. Additionally, audio diarization issues made it difficult to accurately separate speakers in multi-speaker samples. C. Model Limitations Pre-trained models such as Titu fastConformer and Moon- shine were not fully optimized for Bangla features. Fine-tuning was required, and low-epoch fine-tuning sometimes led to a decrease in performance. Overfitting was observed on smaller or incomplete subsets of the dataset. D. Evaluation Challenges Comparing different models was non-trivial due to differ- ences in dataset coverage, preprocessing methods, and evalua- tion splits. Limited test set size further affected the reliability of evaluation metrics. E. Time and Experimentation Constraints The combination of limited GPU hours, large model sizes, and time-consuming preprocessing constrained the number of experiments we could conduct. This made it challenging to systematically explore model architectures, hyperparameters, and training strategies to achieve optimal performance. VII. CONCLUSION This study highlights a critical dichotomy in low-resource speech processing. For Bengali ASR, raw data scaling and ensembling techniques are inferior to targeted fine-tuning on noisy audio paired with immaculate textual alignment. Conversely, for Bengali speaker diarization, state-of-the-art architectures and extensive fine-tuning fail to overcome the inherent acoustic complexities of long-form audio; instead, rigid algorithmic post-processing is the most effective method for minimizing the Diarization Error Rate. Our optimized pipelines, achieving âź0.019 RTF, alongside the open-source release of the 882-hour Lipi-Ghor dataset, provide a robust foundation for the future of conversational AI in Bengali. ACKNOWLEDGMENT Team Villagers would like to thank the organizers of DL Sprint 4.0 and BUET CSE FEST 2026 for providing the computational resources and dataset that made this research possible. REFERENCES [1] V. Lall and Y. Liu, âContextual biasing to improve domain- specificcustomvocabularyaudiotranscriptionwithoutexplicit fine-tuningofwhispermodel,âin20247thInternational Conference on Machine Learning and Natural Language Processing (MLNLP).IEEE,Oct.2024,p.1â6.[Online].Available: http://dx.doi.org/10.1109/MLNLP63328.2024.10800265 [2] S. Hasan, A. H. M. Fuad, R. Labib, and B. Hasan, âLipi-ghor: A large-scale bengali speech dataset with speaker diarization and transcription,â 2026, dL Sprint 4.0, Team Villagers. [Online]. Available: https://huggingface.co/datasets/Sanjidh090/Lipi-Ghor-bn-882-SSTT [3] L. A. Lanzend Ě orfer, F. Gr Ě otschla, C. Blaser, and R. Wattenhofer, âBenchmarkingdiarizationmodels,â2025.[Online].Available: https://arxiv.org/abs/2509.26177 [4] P. P Ě alka, J. Han, M. Delcroix, N. Tawara, and L. Burget, âVbx for end-to-end neural and clustering-based diarization,â 2025. [Online]. Available: https://arxiv.org/abs/2510.19572 [5] H. M. S. Tabib, I. A. Rifti, A. M. A. Ehsan, S. Dasgupta, M. Z. M. S. Sowdha, A. J. Sarker, M. R. I. Nijamy, T. Hossain, M. M. Khatun, M. Mahmood, R. Debnath, G. Biswas, A. Karim, W. A. A. Navid, M. Muztahid, F. A. Udoy, S. S. Rahman, M. T. R. Shifat, M. S. Khatun, M. Rahman, M. M. Hasan, A. Saha, M. N. M. Nobo, S. Bhattacharjee, T. Bhomik, A. N. Swapnil, and S. Kabir, âBengali-loop: Community benchmarks for long-form bangla asr and speaker diarization,â 2026. [Online]. Available: https://arxiv.org/abs/2602.14291 [6] R. N. Nandi, M. Menon, T. Muntasir, S. Sarker, Q. S. Muhtaseem, M. T. Islam, S. Chowdhury, and F. Alam, âPseudo-labeling for domain- agnostic Bangla automatic speech recognition,â in Proceedings of the First Workshop on Bangla Language Processing (BLP-2023), F. Alam, S. Kar, S. A. Chowdhury, F. Sadeque, and R. Amin, Eds.Singapore: Association for Computational Linguistics, Dec. 2023, p. 152â162. [Online]. Available: https://aclanthology.org/2023.banglalp-1.16/ [7] S.Hasan,A.H.M.Fuad,R.Labib,and B.Hasan,ââtrainingnotebooksâ.â[Online].Available: https://github.com/ahmfuad/villagerstrainingnotebooks [8] J. Han, F. Landini, J. Rohdin, A. Silnova, M. Diez, and L. Burget, âLeveraging self-supervised learning for speaker diarization,â 2024. [Online]. Available: https://arxiv.org/abs/2409.09408