Paper deep dive
Bulbul: A Dataset for Dialectal Arabic Speech Recognition
Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas, Nada Almarwani, Samah Aloufi, Saad Ezzini, Maged S. Al-Shaibani, Doaa Dalaq, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Mohamed Mehdi Trigui, Dania Refai, Layan Refai, Mohamed Akrout, Mustafa Jarrar, Wasfi G. Al-Khatib, Alaa Dalaq, Darin El-Nakla, Samir Abdaljalil, Abdulrahman Al-Fakih, Nour El Imane Zeghib, Moussa Redah, Salmane Chafik, Mohamed El-Attar, Rima Grati, Sarah Kohail, Malak Alkhorasani, Khadijah Al Safwan, Ismail M. Mudhaffar, Ali Altam, Ahmed Al-Shaikh, Adnan Saeed, Hamzah Luqman
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/25/2026, 7:46:07 AM
Summary
The paper introduces BULBUL, a community-driven Arabic Automatic Speech Recognition (ASR) dataset designed to address challenges in dialectal diversity and accent-aware modeling. Collected from 275 speakers across 11 Arab countries, the dataset includes 71.59 hours of dialectal speech and 10.42 hours of accented Modern Standard Arabic (MSA) and Classical Arabic (CA). It features structured dialect/sub-dialect coverage, human-verified transcriptions, and demographic metadata. The authors benchmark several state-of-the-art ASR models (e.g., Whisper, SeamlessM4T) on this dataset to establish baselines for dialectal and accented Arabic recognition.
Entities (11)
Relation Signals (9)
BULBUL → containsspeechfrom → Saudi Arabia
confidence 95% · BULBUL includes structured dialect and sub-dialect coverage... covering 11 countries... Saudi Arabia
BULBUL → containsspeechfrom → Yemen
confidence 95% · BULBUL includes structured dialect and sub-dialect coverage... covering 11 countries... Yemen.
BULBUL → includeslanguagevariety → Classical Arabic
confidence 95% · BULBUL includes... recordings of classical Arabic and modern standard Arabic
BULBUL → includeslanguagevariety → Modern Standard Arabic
confidence 95% · BULBUL includes... recordings of classical Arabic and modern standard Arabic
BULBUL → supportstask → Arabic ASR
confidence 95% · We present BULBUL, a multi-dialect Arabic ASR dataset
BULBUL → containsdialect → Hijazi
confidence 90% · The Saudi dialect is divided into six sub-dialects: Hijazi (HJ)
BULBUL → containsdialect → Ta’izzi
confidence 90% · For Yemen, we collected data from two sub-dialects: ... Ta’izzi (TZ)
Whisper large-v3 → evaluatedon → BULBUL
confidence 90% · we benchmark a range of recent ASR systems... including Whisper Large-V3... on BULBUL
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets often focus on single dialects or large-scale broadcast/web data, leading to trade-offs between linguistic diversity and annotation quality. We present BULBUL, a multi-dialect Arabic ASR dataset collected from 275 speakers in 11 Arab countries. BULBUL includes structured dialect and sub-dialect coverage, as well as recordings of classical Arabic and modern standard Arabic spoken by participants in their native dialectal accents to support accent-aware modeling. The quality of the recordings was ensured through a two-level human verification process. We further benchmark a range of recent ASR systems, establishing strong baselines for modern dialectal and accented Arabic ASR.
Tags
Links
- Source: https://arxiv.org/abs/2608.21950v1
- Canonical: https://arxiv.org/abs/2608.21950v1
Trouble viewing inline? Open PDF directly →
Full Text
82,116 characters extracted from source content.
Expand or collapse full text
Bulbul: A Dataset for Dialectal Arabic Speech Recognition Ahmed Ashraf 1 , Aisha Alansari 1 , Fadel Al Abbas 1 , Nada Almarwani 2 , Samah Aloufi 2 , Saad Ezzini 1 Maged S. Al-Shaibani 1 , Doaa Dalaq 1 , AbdelRahim A. Elmadany 3 , Muhammad Abdul-Mageed 3 Mohamed Mehdi Trigui 1 , Dania Refai 1 , Layan Refai 4 , Mohamed Akrout 1 , Mustafa Jarrar 5,6 Wasfi G. Al-Khatib 1 , Alaa Dalaq 1 , Darin El-Nakla 1 , Samir Abdaljalil 7 , Abdulrahman Al-Fakih 1 Nour El Imane Zeghib 1 , Moussa REDAH 1 , Salmane Chafik 8 , Mohamed El-Attar 9 , Rima Grati 9 Sarah Kohail 9 , Malak Alkhorasani 10 , Khadijah Al Safwan 1 , Ismail M. Mudhaffar 1 , Ali Altam 11 Ahmed Al-Shaikh 1 , Adnan Saeed 12 , Hamzah Luqman 1 1 KFUPM, 2 Taibah University, 3 UBC, 4 PSUT, 5 Birzeit University, 6 HBKU 7 Texas A&M University, 8 Mohammed VI Polytechnic University, 9 Zayed University, 10 IAU 11 Symbiosis International University, 12 Taiz University g202411740,aisha.ansari,hluqman@kfupm.edu.sa فك الباب أبغى أكلمك شوي. راح استلم راتبي من هاظ الاشي. شو جاي عبالكم تعملوا؟ شو الي غير الي براسك؟ البطاقة بتاعتي لسه موصلتش لحد دلوقتي ديجا تعدت نص ساعة. غادي نديرلو هذا وش يدير؟ البطاقة ديالي ما وصلتش لحد دايا. وين سرتوا أمس؟ لو سمحت أشتي خدمة الغرف افتح الباب داير اتكلم معاك Figure 1: Overview of BULBUL, covering 11 Arab countries and presenting representative examples from each dialect (English translations provided in Appendix B.2). Abstract Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive re- gional dialect variation, and limited speech re- sources. Existing speech datasets often focus on single dialects or large-scale broadcast/web data, leading to trade-offs between linguistic diversity and annotation quality. We present BULBUL, a multi-dialect Arabic ASR dataset collected from 275 speakers in 11 Arab countries. BULBUL in- cludes structured dialect and sub-dialect coverage, as well as recordings of classical Arabic and mod- ern standard Arabic spoken by participants in their native dialectal accents to support accent-aware modeling. The quality of the recordings was en- sured through a two-level human verification pro- cess. We further benchmark a range of recent ASR systems, establishing strong baselines for modern dialectal and accented Arabic ASR. 1 Introduction Recent advances in automatic speech recogni- tion (ASR) have revolutionized spoken language technologies, demonstrating remarkable capabil- arXiv:2608.21950v1 [cs.CL] 22 Aug 2026 ities across diverse languages and domains. In Arabic, ASR systems must navigate a uniquely complex linguistic landscape characterized by im- mense phonetic and lexical diversity in dialects and subdialects (Alqadasi et al., 2025; Alshargi et al., 2019). This variety results in limited and unbal- anced speech resources, heavily skewed toward a few dialects, leading to poor generalization for un- derrepresented dialects (Djanibekov et al., 2025). Despite progress in Arabic ASR datasets, most of these datasets cover only a few dialects, pri- marily modern standard Arabic (MSA), Egyptian, broad Levantine, Gulf, and Moroccan (Djanibekov et al., 2025). On the other hand, Algerian, Su- danese, Yemeni, and Tunisian Arabic, as well as parts of Palestinian and Jordanian Arabic, remain comparatively low-resource (Sullivan et al., 2026; Alqadasi et al., 2025; Talafha et al., 2024). Beyond dialect coverage, existing research has also largely overlooked accented MSA and classical Arabic (CA). MSA is the formal variety used in modern communication, while CA represents the historical form found in religious and classical texts (Ryd- ing, 2005; Versteegh, 2014). To our knowledge, no dedicated speech dataset exists to systematically evaluate accented MSA and CA across diverse di- alectal backgrounds in ASR. To address these limitations, we introduce BUL- BUL, a community-driven Arabic speech dataset designed to improve dialectal coverage in ASR, covering 11 countries: Algeria, Egypt, Jordan, Morocco, Palestine, Saudi Arabia, Sudan, Syria, Tunisia, United Arab Emirates, and Yemen. To capture finer linguistic variation, we further divide the Saudi and Yemeni dialects into fine-grained sub-dialects. The dataset also includes accented MSA and CA, enabling accent-aware evaluation. BULBUL is constructed from text collected across multiple sources and domains to ensure broad lex- ical and topical diversity. We employ a two-level human verification process to ensure transcription accuracy, audio quality, and dialect authenticity to improve the dataset’s reliability. Moreover, we benchmark a diverse set of existing ASR systems on BULBUL, establishing comprehensive baselines across dialectal and formal Arabic varieties. We also release demographic metadata, including age and gender, to support demographic analysis and fairness-aware evaluation of ASR models. 2 Related Work Arabic speech resources span multilingual, uni- dialect, and multi-dialect datasets. Multilingual datasets such as GlobalPhone (Schultz, 2002) and Mozilla Common Voice (Ardila et al., 2020) in- clude Arabic and support cross-lingual learning, but typically provide limited dialectal diversity. Uni-dialect corpora, including Saudi, Algerian, and Tunisian resources, offer deeper coverage of specific dialects or code-switching scenarios (Al- ghamdi et al., 2008; Droua-Hamdani et al., 2010; Masmoudi et al., 2017; Messaoudi et al., 2021), but their narrow geographic focus and controlled recording conditions limit generalizability across the broader Arabic-speaking world. Multi-dialect Arabic datasets provide broader lin- guistic coverage by capturing speech from multiple regions, as seen in OrienTel (Siemund et al., 2002), Fisher Levantine (Maamouri et al., 2004), QASR (Mubarak et al., 2021), MGB (Ali et al., 2019a), and MASC (Al-Fetyani et al., 2021). However, many of these datasets rely heavily on broadcast, telephone, or web-sourced recordings, introducing domain bias, inconsistent annotation quality, and limited representation of spontaneous community speech. These limitations highlight the need for more diverse and dialect-rich Arabic speech cor- pora for robust ASR development. Table 1 com- pares BULBUL with existing datasets. Appendix A provides a detailed literature review. BULBUL in Comparison. Unlike many prior re- sources that focus on a single dialect, broadcast speech, or controlled scripted recordings, BUL- BUL captures naturally diverse speech collected from real speakers across multiple dialect fami- lies, with explicit sub-dialect annotations for some countries to reflect fine-grained regional variation. Furthermore, BULBUL distinguishes itself by in- cluding accented MSA and CA recordings from 11 dialect speakers, enabling accent-aware ASR anal- ysis under standardized linguistic content. Com- bined with human-verified transcriptions, rich de- mographic metadata, and balanced domain diver- sity, this makes BULBUL a more realistic and com- prehensive benchmark for developing robust Ara- bic ASR systems. 3BULBUL Dataset The BULBUL dataset is a community-driven project conducted from January 2025 to April 2026. The participants comprise 275 members from 11 Arab Table 1: Comparison of the BULBUL dataset with existing Arabic ASR datasets. DatasetHrsSpkDialMicroText SrcRecEnvDomHVMetaAcc Multilingual GlobalPhone (Schultz, 2002)351701✗ScriptRecCtrl1✓✗ Common Voice (Ardila et al., 2020)152251✗ScriptRecUnctrl3✓ Uni-Dialect FACST (Djegdjiga et al., 2018)7.5201✗ScriptRecCtrl1✓✗ ArzEn (Hamed et al., 2021)11.4401✗Conv.RecCtrl2✓ TUN-SWITCH (Abdallah et al., 2023)∼9–1✗BroadcastTV/YTMixed2✓✗✓ SAAVB (Alghamdi et al., 2008)∼96.41,0331✓ScriptTelMixed∼9✓ ALGASD (Droua-Hamdani et al., 2010)–3001✗ScriptRecCtrl1✓ TARIC (Masmoudi et al., 2017)∼101081✗Conv.RecUnctrl1✓✗ STAC (Zribi et al., 2014)5∼701✗Conv.TVCtrl5✓✗ TunSpeech (Messaoudi et al., 2021)∼10.5101✗MixedYT/TV/RecMixed2✓✗ Multi-Dialect ESCWA.CS (Chowdhury et al., 2021)2.8–4✗Conv.RecUnctrl–✓✗ OrienTel (Siemund et al., 2002)–6✗Conv.TelMixed–✗✓ Fisher Lev. (Maamouri et al., 2004)45–4✗Conv.TelCtrl–✗ QASR (Mubarak et al., 2021)2,00011,0925✗ScriptTVCtrl5✗✓✗ MASC (Al-Fetyani et al., 2023)1,000–23✗SocialYTUnctrl4✓✗ MGB-2 (Ali et al., 2019a)1,200–4✗ScriptTVCtrl12✓✗ MGB-5 (Ali et al., 2019b)14–17✗SocialYTUnctrl7✓✗ BULBUL (Ours)71.5927511✓MixedRecUnctrl11✓ Abbreviations: Hrs = total hours with transcription; Spk = number of speakers; Dial = number of major dialect groups; Micro = micro-dialect coverage; Text Src = text prompt source (Script: scripted text; Conv.: conversational speech; Broadcast; Social: social media); Rec = recording style (Rec: recorded; Tel: telephone conversations; YT: YouTube; TV); Env = environment (Ctrl: controlled; Unctrl: uncontrolled; Mixed); Dom = number of domains; HV = human verification; Meta = metadata availability; Acc = accented speech diversity.✓ = yes,✗ = no, – = not reported. countries, of whom 160 were male and 115 were female. We organized the participants into country- specific teams, each headed by a leader. The lead- ers were responsible for collecting and verifying the raw text, providing guidance to participants dur- ing recording, and verifying the recorded audio. To ensure steady progress, we maintained continuous communication through a dedicated Slack group for discussions and held weekly meetings to review progress, present statistics, and address any chal- lenges. Figure 2 illustrates the pipeline for curating the BULBUL dataset. 3.1 Data Collection The dataset collection process was conducted in- crementally. It initially began with two countries, Saudi Arabia and Egypt. After establishing the collection pipeline and validation procedures, addi- tional countries were gradually incorporated, bring- ing the total to 11: Algeria (AG), Egypt (EG), Jor- dan (JO), Morocco (MA), Palestine (PS), Saudi Arabia (KSA), Sudan (SD), Syria (SY), Tunisia (TN), United Arab Emirates (UAE), and Yemen (YE). The resulting dataset includes dialectal Arabic record- ings and formal Arabic speech, comprising CA and MSA spoken in participants’ native dialectal ac- cents. Therefore, BULBUL enables the study of dialectal and accent-aware speech recognition. We used the white dialectZ A íJ J . À @ Èj . Í À @for seven countries: Algeria, Egypt, Jordan, Morocco, Pales- tine, Sudan, and Tunisia. The white dialect is a neutralized variety that blends regional dialectal features with widely understood MSA while avoid- ing highly localized vocabulary (Alkhamees, 2023). It is commonly used in media, cross-regional com- munication, and formal social contexts within each country. For Syria and the United Arab Emirates, we focused on the Levantine ÈJ ” A ÇÀ @and Abu Dhabi ÈJ K AJ J . ¢À @ dialects, respectively. For Saudi and Yemeni dialects, we collected data at the sub-dialect 1 level to better capture intra- country variation. These dialects were selected based on the availability of sub-dialectal data and speaker contributions from these regions. The Saudi dialect is divided into six sub-dialects: Hijazi (HJ) ÈK P Am . k , Qatifi (QT) ÈJ ÆJ ¢ Ø, Southern (ST) ÈJ K . Ò Jk . , Najdi (NJ) ÈK Ym . ' , Hijazi-Badwi (HJ-BD) ÈK YK . ÈK P Am . k , and Hassawi (HS) ÈK AÇk. We combined QT and HS to represent the Eastern (ES) dialect. For Yemen, we collected data from two sub-dialects: Sana’ani (SN) ̇ G A™ Jì and Ta’izzi (TZ) ̄ Q™ K. Text Collection. The text used in our dataset was collected separately for each Arabic dialect to en- sure adequate representation of linguistic diver- 1 Sub-dialect: a regional variety within a country. Domain Classification 1. Dialects and Subdialects2. Data Collection3. Audio Recording and Revision Dialectal Text Collection Accented MSA and Classical Arabic Preprocessing Sanaa Taaz Hijazi Najdi Qatifi Hijazi- Badawi Hasawi Janoobi Others Subdialects Name, Age, Gender, Dialect, Subdialect Recording and Self-verification Training, video recording, guidelines Cross-verification 1. Clear sound. 2. Matching text. 3. Dialect correct 1. Noise. 2. Unmatched text. 3. Wrong dialect 4. Long pauses Sentences Preparation Dialects Selection Sanaa Taaz Southern Hijazi Najdi Hijazi-Bedouin Eastern KSA Yemen Dialects Subdialects Social media News Broadcasts Conversations Filtration & Preprocessing Dialect Verification Text Quality Check Preprocessing Domain Classification Recording Process Audio Recording Self-Verification Cross-Verification BULBUL Dataset Text Collection Filtration & Processing Domain Classification Guidelines Development and Training DURATION ~72 HOURS DIALECTS 11 Countries 16 Dialects SPEAKERS 275 Whisper SeamlessM4T Benchmarking Figure 2: BULBUL dataset construction pipeline. sity across regions. For each dialect, texts were collected from multiple sources, including pub- licly available datasets (e.g., Casablanca (Talafha et al., 2024), SADA (Alharbi et al., 2024), and MADAR (Bouamor et al., 2018)), social media plat- forms (e.g., Facebook, X, and YouTube comments), and private communication via platforms, such as WhatsApp and Telegram. This multi-source col- lection strategy was designed to capture naturalis- tic, regionally grounded linguistic variation. For accented MSA and CA, we used texts from the Targamat dataset (Kadaoui et al., 2023), consist- ing of 200 CA samples and 200 MSA samples. In Appendix B.4, we provide a detailed lexical and semantic analysis of the collected textual data. Text Revision. After the initial text collection stage, we manually reviewed the text to remove offensive and slang content, politically sensitive material, and texts that are not aligned with the targeted dialects. For the tarjamat dataset, the filtration resulted in 129 MSA and 200 CA sen- tences. The reviewed texts were then passed to a pre-processing pipeline before being uploaded to the recording tool, as described in Appendix B.1. Domain Classification. We employed GPT-5-mini to categorize the collected texts into thematic do- mains. The model was prompted with a struc- tured domain taxonomy, consisting of 11 prede- fined categories (Social, Politics, Economy, Re- ligion, Health, Education, Technology, Entertain- ment, Sports, Crime, and Other), along with con- cise definitions for each domain. Appendix B.3 provides details about the domain classification process and human review. 3.2 Recording Process The collected texts from all dialects were uploaded to a recording tool developed using Gradio 2 . We organized the text by country and, if applicable, sub-dialect. We provided participants with written guidelines (Appendix C) to guide them through the recording stage. We also recorded an instruc- tional video to help participants become familiar with the recording tool and to minimize the risk of errors. Participants were recruited using different strategies depending on the country. In Saudi Ara- bia, volunteers were recruited through an online volunteering opportunity published on the Saudi National Volunteer Platform 3 . For every 10 min- utes of recorded audio, participants were credited with one hour of voluntary work. For other con- tributing countries, the participants were selected by country leaders. 3.3 Human Sanity Check We used two levels of revision. The first level is self-verification, in which the speaker optionally listens to their own recording before submission to ensure clarity and alignment with the text. The second level is external verification, in which a na- tive member of the corresponding country reviews the recordings. We provided the reviewers with written guidelines and a video recording explaining the verification strategy. The detailed verification rules are outlined in Appendix C. 2 https://w.gradio.app/ 3 Saudi NVG: https://nvg.gov.sa/home 4BULBUL Analysis Overall Dataset Distribution. Tables 2 and 3 show the statistics of the BULBUL dataset. BULBUL consists of 36,949 dialectal utterances recorded over 61.46 hours alongside 2,325 ac- cented MSA/CA recordings totaling 10.42 hours. The longest recordings in the dialectal subset are for the Yemeni (18.97 hours), Saudi (11.89 hours), and Syrian (8.21 hours) dialects. In contrast, the Su- danese dialect remains comparatively underrepre- sented, with less than one hour of recorded speech. This lower representation is mainly due to the lim- ited availability of native Sudanese participants dur- ing the data collection period. In the accented MSA/CA subset, the highest durations are observed for Yemeni (2.82 hours) and Egyptian (2.02 hours) dialects, while the lowest are observed for UAE (0.02 hours) and Tunisian (0.11 hours) dialects. This imbalance highlights both the practical challenges of large-scale dialectal data collection and the need for continued efforts to expand coverage of lower-resourced varieties. Sub-dialectal Distribution. As shown in Table 2, the two countries, further divided into sub-dialects, are Saudi Arabia and Yemen. In Saudi Arabia, the largest number of recordings is for the Hijazi dialect, with 7,120 utterances totaling 10,73 hours. In Yemen, the Ta’izzi dialect contains nearly twice as many utterances and approximately 1.6 times as many speech hours as the Sana’ani dialect. Cross-dialectal Speech Tempo. To directly as- sess speaking tempo, we analyze the WPS, which normalizes the utterance length by duration. Clear cross-dialectal variation is observed in the dialec- tal subset, with Egyptian dialect exhibiting the fastest speech tempo (1.94 WPS), followed by Jordanian (1.83 WPS), while Yemeni shows the slowest tempo (1.09 WPS). Interestingly, accented MSA/CA speech exhibits a higher average tempo (1.65 WPS) than dialectal speech (1.46 WPS), suggesting faster, more fluent articulation in the scripted formal setting. This contrast is particu- larly pronounced for Yemeni speakers, whose di- alectal speech has the slowest tempo yet becomes among the fastest in the accented subset (1.75 WPS). These findings suggest that speaking style, recording conditions, and linguistic structure all influence temporal speech characteristics. The ob- served differences between dialectal speech and MSA/CA should be interpreted with caution, as the MSA/CA subset was recorded under different Table 2: Descriptive statistics BULBUL across countries. We report the total number of utterances (Utts), total hours, minimum duration (Min) in seconds, average du- ration (AvgDur) in seconds, maximum duration (Max) in seconds, average words per utterance (WPU), and average words per second (WPS). Saudi Arabia and Yemen are further divided into subdialects. DialectUttsHrsMinAvgDurMaxWPU WPS AG1,172 2.21 2.10 6.78 120.42 10.06 1.48 EG2,034 3.09 1.74 5.46 34.38 10.58 1.94 JO3,284 4.92 1.72 5.39 23.46 9.86 1.83 KSA7,872 11.89 0.83 5.44 152.34 9.54 1.75 HJ7,120 10.73 1.02 5.43 152.34 9.51 1.75 HJ-BD255 0.49 2.91 6.91 30.40 11.35 1.64 ST307 0.56 2.35 6.56 29.55 9.66 1.47 NJ761 0.85 1.82 4.02 10.26 8.72 2.17 ES449 1.01 0.83 8.13 25.26 14.31 1.76 MA1,169 2.30 2.52 7.10 38.80 11.22 1.58 PS2,444 3.63 1.14 5.35 47.76 8.27 1.55 SD577 0.80 1.92 4.98 16.74 6.68 1.34 SY5,449 8.21 1.50 5.43 50.82 7.28 1.34 TN1,959 3.34 1.18 6.14 22.98 8.32 1.36 UAE1,518 2.10 1.72 4.97 39.81 8.02 1.61 YE9,470 18.97 1.01 7.21 348.96 7.85 1.09 SN3,145 7.38 1.38 8.45 348.96 9.97 1.18 TZ6,322 11.58 1.01 6.60 300.60 6.79 1.03 Summary 36,949 61.46 0.8346.00348.9608.801.46 Table 3: Descriptive statistics of the accented MSA and CA subset across dialect groups. Dial denotes dialects, and Utts denotes the number of utterances. DialUttsHrsMinAvgMaxWPU WPS AG68 0.35 8.28 18.72 40.32 32.68 1.75 EG447 2.02 4.80 16.27 61.32 26.52 1.63 JO67 0.28 1.92 15.02 40.43 28.96 1.93 KSA102 0.40 1.26 14.13 37.49 24.47 1.73 MA55 0.29 17.58 19.18 48.00 27.82 1.45 PS418 1.78 4.50 15.36 47.39 26.06 1.70 SD80 0.41 7.20 18.66 41.52 28.38 1.52 SY389 1.94 6.00 17.91 46.62 26.44 1.48 TN22 0.11 11.70 17.75 26.70 29.18 1.64 UAE6 0.02 6.24 10.42 16.14 16.50 1.58 YE671 2.82 3.78 15.15 48.78 26.47 1.75 Summary 2,325 10.421.26016.23461.320 26.68 1.65 conditions, including read speech, which may influ- ence speaking rate. From an ASR perspective, such tempo variation is important, as faster speech may challenge temporal alignment and acoustic model- ing, whereas slower speech may introduce longer temporal dependencies (Talafha et al., 2024). Outlier Analysis. As shown in Table 2, the dataset exhibits noticeable variability in maximum utter- ance duration across dialects, with some recordings substantially exceeding the typical range. For ex- ample, Yemeni Arabic has a maximum recording duration of 348,96 seconds, while the Saudi dialect reaches 152.340 seconds, both significantly higher than the average segment duration across dialects. Such long recordings likely represent long scripts. Gender Distribution. BULBUL exhibits variabil- Figure 3: Normalized distribution of male and female speakers across the represented countries. ity in gender representation across countries, as shown in Figure 3. Approximately balanced gen- der distributions are observed in Algeria (53.3% female), Jordan (56.0%female), Saudi Arabia (52.8%male), and Tunisia (55.6%female). Other countries, such as Sudan, Yemen, Algeria, and the UAE, show skewed distributions. Sudan includes only male participants, while the UAE shows a strong female majority (∼89.5%). Although the recruitment strategy aimed to encourage diversity, these were recruitment objectives rather than en- forced demographic quotas. As BULBUL is a community-driven dataset relying on voluntary par- ticipation and participant availability, achieving an equal gender distribution was not always possible across all countries. Domain Distribution. To ensure topical diversity, we analyzed the domain composition of the col- lected corpus across the 11 participating countries. Figure 4 shows the overall domain distribution, re- vealing a strong concentration in the Social domain, which constitutes the largest portion of the dataset (65.4%), followed by Economy (15.7%). The re- maining domains are more sparsely represented, each contributing a smaller proportion. This im- balance reflects the natural prevalence of conversa- tional and socially oriented content in community- contributed speech data, while still maintaining broad coverage across topics. Data Splits. To ensure a robust and fair evalu- ation of ASR models’ performance on the BUL- BUL dataset, we rely on the human-verified subset for development and testing. The total duration of verified data varies across countries. Thus, in- stead of enforcing a fixed absolute duration per split, we divide the verified data within each coun- try into 50% development and 50% test splits that Domains Social: 12118 (65.4%) Economy: 2916 (15.7%) Health: 726 (3.9%) Religion: 625 (3.4%) Other: 492 (2.7%) Education: 441 (2.4%) Entertainment: 354 (1.9%) Crime: 303 (1.6%) Technology: 216 (1.2%) Politics: 203 (1.1%) Sports: 147 (0.8%) Figure 4: Overall domain distribution across the 11 participating countries. are speaker-disjoint. Table 9 reports the statistics of the development and test splits for each country, including the number of utterances, total duration, and number of speakers. 5 Evaluation Evaluated Models To benchmark BULBUL, we evaluate a diverse set of state-of-the-art multilin- gual speech recognition models spanning a wide range of model scales, including Whisper Large-V3 (Radford et al., 2023), Whisper Large-V3-Turbo (Radford et al., 2023), SeamlessM4T (Barrault et al., 2023), massively multilingual speech (MMS) (Pratap et al., 2024), OmniLLM (team et al., 2025) (300M, 1B, 3B, and 7B), and OmniCTC (team et al., 2025) (300M, 1B, 3B, and 7B). Evaluation Setup To evaluate the models on BUL- BUL, we used a zero-shot evaluation setup using the test set without fine-tuning on the BULBUL dataset or any country-specific subset. The evaluation is conducted using each model’s default decoding configuration. These settings allow us to measure out-of-the-box generalization performance and as- sess both cross-lingual and cross-accent robustness. More details are present in Appendix F. Evaluation Protocol To ensure that the reported character error rate (CER) and word error rate (WER) accurately reflect transcription quality, the same text normalization pipeline was applied to both the reference and predicted transcripts. The details are described in Section F.1. 6 Results and Discussion We evaluate model performance on the BULBUL test set using WER and CER. Tables 4 and 5 report the results across all evaluated systems tested on di- alectal subsets and accented MSA/CA, respectively. For analysis, we group countries into broader re- gional dialect categories: Northwest African, Ara- bian Peninsula, Levantine, and Nile Valley, based on geographic proximity and established Arabic dialect classifications. Overall Results. Overall, Omni-based models consistently outperform the baseline multilingual ASR models across regional dialect clusters, with OmniLLM-7B achieving the best overall perfor- mance in dialectal speech, with a mean WER of 41.7 and a CER of 12.7. This is because OmniLLM- 7B processes acoustic features through an LLM decoder, enabling stronger contextual understand- ing during transcription. This architecture likely improves the model’s ability to select contextually appropriate words and generalize more effectively to previously unseen regional Arabic dialects in zero-shot settings. In the MSA/CA data, performance improves substantially across all models, and OmniLLM-7B again achieves the best results. This suggests that dialectal variation poses a greater challenge to mul- tilingual ASR systems than standardized speech. Per-dialect Results. The performance varies sub- stantially across dialects. The Palestinian and Saudi dialects consistently achieve the lowest error rates across most models, indicating comparatively stronger recognition performance. In particular, the Palestinian subset yields the lowest WER for 9 out of the 12 evaluated systems. This is likely driven by the limited domain diversity of the Palestinian subset, which reduces lexical and semantic vari- ability, resulting in samples that is easier for ASR models to transcribe. In contrast, Algeria and Sudan exhibit the high- est error rates across most models, reflecting greater recognition difficulty in these dialects. For Sudanese Arabic, this may be partly attributed to the limited availability of public speech datasets. On the other hand, Algerian Arabic likely poses greater recognition difficulties due to its substan- tial lexical and phonological divergence from MSA, highly non-standard orthographic conven- tions, and the spontaneous conversational nature of the speech, which leads to a stronger mismatch with multilingual ASR training distributions. More- over, the relatively long average utterance duration in the Algerian subset may further contribute to in- creased transcription difficulty, as longer sequences are more susceptible to cumulative decoding errors AGTNMAUAEKSAYEJOSYPSEGSD Country / Dialect 0 10 20 30 40 50 60 WER Gap (%) WhisperV3 Seamless MMS1B OmniLLM7B OmniCTC1B Figure 5: Country-level gap between dialectal Arabic and accented MSA/CA speech. The gap is computed as dialectal WER minus accented WER; larger values indicate weaker cross-variety robustness and greater sensitivity to dialectal variation. and context drift. Accented MSA and CA. Table 5 and Figure 5 show that ASR performance improves substan- tially on accented MSA/CA compared to dialectal speech, suggesting that the formal linguistic struc- ture of MSA/CA is generally better aligned with the models’ training distributions, even when spoken with regional accents. OmniLLM-7B achieves the best overall performance, with a mean WER/CER of 13.3/4.8, followed closely by OmniLLM-3B (13.7/5.1) and OmniLLM-1B (14.2/5.3), indicating strong robustness across accent variations. Among general multilingual foundation models, Seamless remains competitive (14.4/5.4), while MMS1B per- forms worst (33.8/9.7). Performance still varies by accent. Levantine-accented speech yields the lowest error rates overall, particularly for Syrian- accented speech, where OmniLLM-7B achieves 8.1/1.7 WER/CER. In contrast, Sudanese-accented speech remains consistently challenging across all models, with even the strongest systems exceeding 20% WER. Figure 5 further shows that the robust- ness gap between dialectal and accented speech is largest for North African and Yemeni dialects, indicating that spontaneous dialectal variation intro- duces substantially greater difficulty than accented formal speech. Model Scale. Within the OmniASR-LLM fam- ily, performance improves almost monotonically with increasing model size, with larger models con- sistently yielding lower error rates in both met- rics. This indicates that increased model capac- ity enhances robustness to phonetic, lexical, and acoustic variation. In contrast, the scaling behav- ior within the OmniASR-CTC family is less stable. Increasing parameter size does not yield system- Table 4: Performance of ASR models on dialectal Arabic speech of BULBUL dataset grouped by regional dialect clusters. Each entry is reported as WER/CER (%).is the lowest andis the highest WER/CER per dialect. ModelNorthwest AfricanArabian PeninsulaLevantineNile ValleyMean AlgeriaTunisiaMaghribUAEKSAYemenJordanSyriaPalestineEgyptSudan WhisperV380.5/42.148.4/11.771.7/33.146.8/13.926.0/9.767.9/31.734.9/8.658.9/17.929.6/9.361.6/29.865.6/38.553.8/22.9 WhisperV3T88.2/46.051.2/12.767.5/25.260.2/19.933.3/12.079.8/35.339.6/10.775.2/28.434.7/13.081.6/44.383.8/41.763.2/26.8 Seamless69.5/28.045.2/10.656.1/21.645.9/12.727.4/10.362.9/22.534.4/8.160.1/17.926.1/7.243.3/16.370.9/30.149.8/16.6 MMS1B88.5 /31.770.7 /18.883.8 /29.572.9 /21.066.8/22.382.3 /30.271.9/20.884.4 /28.162.4/18.884.9 /34.079.1/27.577.1 /25.1 OmniLLM300M66.4/23.846.7/11.454.0/16.353.3/16.727.0/8.762.5/22.035.5/8.663.2/20.924.8/7.451.2/20.168.5/27.050.8/17.0 OmniLLM1B60.0/19.243.8 /10.549.1/14.546.5/13.322.3/7.057.1/17.733.3/7.758.4/18.319.8/5.445.1/17.061.0/22.545.1/13.9 OmniLLM3B58.1 /19.142.7/8.948.0/13.943.6/11.819.7 /6.556.9 /16.831.6 /6.856.2/16.820.0/5.543.1 /16.859.3/20.543.6/13.0 OmniLLM7B56.1 /18.442.7 /10.045.3/12.243.0/10.919.6/6.554.5 /17.830.8/6.852.1/14.917.8/4.840.2/14.356.5/19.141.7/12.7 OmniCTC300M76.0/28.657.7/18.681.5/53.466.0/25.949.3/21.675.9/36.351.5/15.076.2/30.139.0/12.069.2/30.484.0 /44.265.7/28.8 OmniCTC1B58.5/15.946.2/12.260.0/20.647.2/13.726.9/7.962.6/23.839.0/9.760.2/20.226.0/7.652.6/20.365.8/28.749.5/16.6 OmniCTC3B71.1/42.152.4/19.568.5/39.851.5/19.630.0/10.368.4/32.938.8/10.965.2/27.524.3/6.551.6/20.170.3/32.253.3/23.8 OmniCTC7B66.3/30.451.1/19.865.9/38.349.4/17.428.4/10.167.3/30.136.3/10.261.8/24.423.9/6.650.1/20.571.4/37.151.8/22.1 Table 5: Performance of ASR models on accented MSA/CA speech of BULBUL dataset grouped by regional accent clusters. Each entry is reported as WER/CER (%).is the lowest andis the highest WER/CER per dialect. ModelNorthwest AfricanArabian PeninsulaLevantineNile ValleyMean AlgeriaTunisiaMaghribUAEKSAYemenJordanSyriaPalestineEgyptSudan WhisperV315.4/6.024.1/6.718.4/7.113.3/3.015.2/4.414.3/5.517.0/4.512.4/5.815.3/6.419.0/6.934.2/15.319.2/8.0 WhisperV3T 16.4/6.425.9/7.219.7/7.311.2/2.616.1/4.715.8/5.816.1/3.911.8/5.114.9/6.219.4/6.836.2/15.220.0/7.5 MMS1B30.6/7.545.1/12.336.6/9.621.4/5.827.1/7.325.6/6.336.2/8.923.6/5.130.7/7.334.4/8.844.8 /13.733.8/9.7 Seamless 11.2/2.917.1 /5.112.1 /3.708.2/1.916.9/5.110.7/3.016.5/4.19.2/2.310.9/3.213.1/3.623.2/7.714.4/5.4 OmniLLM300M 12.8/3.223.3/7.114.5/4.012.2/3.414.5/3.913.0/3.818.3/4.79.1/2.213.4/3.817.7/5.027.6/9.017.0/6.1 OmniLLM1B10.1/2.419.9/5.712.8/3.210.2/2.413.5/4.011.9/3.413.5/3.17.8 /1.812.3/3.414.4/3.923.4/7.814.2/5.3 OmniLLM3B10.2/2.718.0/5.013.2/3.011.2/2.613.5/4.110.2/2.913.1/3.17.9/2.010.6/3.114.4/3.922.1/7.113.7/5.1 OmniLLM7B08.9/2.118.5/5.212.9/3.012.2/2.813.0/3.79.4/2.512.3/2.78.1/1.710.2 / 2.512.7/3.521.6/6.313.3/4.8 OmniCTC300M 21.3/5.037.6/9.927.9/7.517.4/3.920.9/5.319.3/4.726.4/6.315.5/3.222.5/5.329.9/7.939.1/12.526.9/7.8 OmniCTC1B13.4/3.226.1/7.918.7/4.514.3/3.614.3/3.913.1/3.016.7/3.810.3/2.314.1/3.918.4/4.828.5/8.718.3/6.0 OmniCTC3B11.8/2.722.4/6.014.9/3.711.2/2.814.2/3.911.6/2.814.6/3.29.1/2.011.7/2.815.9/4.024.5/7.215.3/5.2 OmniCTC7B10.6/2.521.7/5.314.8/3.514.3/3.414.3/3.711.0/2.512.0/2.79.0/2.010.9/2.414.3/3.924.3/6.815.4/5.1 atic improvements; in some cases, performance degrades slightly. Notably, omniASR-CTC-1B out- performs the larger 3B and 7B variants, indicating that scaling alone does not guarantee gains under the CTC-based architecture. This implies optimiza- tion challenges when applying CTC-based decod- ing at higher parameter counts. 6.1 Error Analysis We observe recurring error patterns including phonetically similar substitutions, dialect-driven phonological mismatches, and distortions of rare or dialect-specific lexical items. Character-level er- rors often reflect systematic phoneme-to-grapheme mismatches, whereas word-level errors are more common for proper nouns and low-frequency vo- cabulary. To better understand these failure modes, we conduct a qualitative analysis of the best- performing model (OmniASR-LLM-7B) on the test split, revealing three dominant sources of error. Phonetically Similar Substitutions. The model frequently confuses acoustically similar phonemes, particularly when dialectal realizations diverge from standard pronunciations. For example, em- phatic/sˇ/(ê) is confused with plain/s/(Ä), pro- ducing HQÂïinstead of HQÂÖ. Likewise, pharyngeal /Q/(®) may weaken in connected speech, leading to®↔ @substitutions. In dialects where/q/( Ü ) is realized as a glottal stop/P/, the model often alter- nates between Ü and @, as in Å Y Ø↔ Å X @. Dialect-driven Phonological Variation. Errors also arise when dialect-specific pronunciation pat- terns do not align with expected orthographic forms. For instance, the imperfective prefix may appear as/b-/or/bi-/(... í . vs.... íJ K . ), leading to inconsis- tent transcription. Reduced vowels in fluent speech can cause deletions, such as transcribingºP AJ . K asº Q . K . Similarly, variable realizations ofh . , such as/j/or /tS/, result in forms like QÂ Ñ AK . instead of Qk . AK . . 7 Conclusions In this work, we introduce BULBUL, a large-scale, community-driven Arabic speech corpus designed to better reflect the linguistic diversity of the Arab world beyond dominant dialects and MSA. BUL- BUL provides a challenging and realistic bench- mark for Arabic speech technologies under diverse regional and demographic conditions by cover- ing 11 countries and fine-grained sub-dialect anno- tations, resulting in 16 dialects. Moreover, BUL- BUL includes accented MSA and CA recordings, enabling analysis of accent transfer and pronunci- ation variation across standardized speech forms. Our benchmark evaluation across multiple SOTA ASR systems reveals substantial performance dis- parities across countries and dialects, highlighting that current speech models remain uneven in their support for underrepresented Arabic varieties. The lexical overlap and speaking-rate analyses further demonstrate the significant heterogeneity across Arabic speech communities, emphasizing the lim- itations of evaluating models solely on coarse di- alect categories. By releasing BULBUL, we aim to support more inclusive Arabic speech research, enabling future work in dialect-aware ASR, dialect identification, accent adaptation, and speech gener- ation for low-resource Arabic varieties. 8 Limitations Emotionally Neutral Recordings. Due to the ab- sence of surrounding paragraph-level context, par- ticipants delivered the sentences in a formal, neutral tone, as the intended emotional expression could not be clearly determined. In addition, the partici- pants were requested to avoid environments with excessive background noise or overlapping speech from other speakers. Consequently, the BULBUL dataset is limited in its applicability to emotion recognition tasks or speech processing scenarios involving noisy and unconstrained environments. Since the collected data were controlled for emo- tional expression and background noise, the eval- uation of the proposed models is more closely as- sociated with dialectical understanding than with speaker emotion or robustness to environmental noise. Country and Gender Imbalance. Some countries are underrepresented in the dataset, and the gen- der distribution across dialects is not perfectly bal- anced. This imbalance is primarily due to the vol- untary nature of the data collection process. While this limitation cannot be fully addressed retrospec- tively, a detailed statistical analysis of the dataset is provided to characterize these imbalances and support a more informed interpretation of the ex- perimental findings. Sub-Dialect Coverage Limitation. Another limi- tation is the restricted coverage of sub-dialects. The Arab region is characterized by substantial dialec- tal variation, not only between countries but also within the same country. Comprehensive coverage of all sub-dialects would require collecting speech samples from rural and hard-to-access areas, which presents significant logistical challenges. Domain Coverage Imbalance. Some country sub- sets exhibit limited domain diversity due to the nature of the data collection. For example, the Palestinian subset is primarily concentrated in the economy domain, resulting in a narrower lexical and topical distribution compared to other coun- try subsets. This limited variation may make the corresponding test data more homogeneous and po- tentially easier for ASR models, which should be considered when interpreting cross-country perfor- mance differences. Imbalance in CA and MSA Coverage. MSA and CA subsets are comparatively limited in size, with recordings contributed by only a small number of participants. This results in both speaker imbalance and reduced diversity in speaking styles, accents, and acoustic conditions. Consequently, findings related to CA and MSA should be interpreted with caution, as the observed performance may not fully generalize to broader speaker populations. 9 Ethics and Data Statement The speech corpus used in this work was collected from speakers across 11 Arab countries to improve dialectal diversity in ASR. All participants pro- vided informed consent for the use of their record- ings and transcripts for research purposes. The dataset includes both prompted speech and natu- rally occurring conversational content, with a sub- set derived from private chat conversations (What- sApp and Telegram) voluntarily contributed by their owners. Only conversations for which ex- plicit permission was obtained were included in the dataset. To protect participant privacy, all conversational texts were carefully anonymized before recording and release. Personally identifiable information (e.g., names, phone numbers, addresses, usernames, email addresses, organizations, and other sensitive references) was removed. Only non-identifying de- mographic metadata, specifically age and gender, are retained to support research on speaker diver- sity. No original private chat logs are distributed. Only the anonymized speech recordings and corre- sponding anonymized transcripts are released. The dataset is released under a research-only license for non-commercial academic use upon contact with the corresponding author. Users are expected to comply with applicable privacy regula- tions and ethical guidelines, and the dataset must not be used to identify individuals, reconstruct per- sonal information, or support surveillance or other harmful applications. Although the dataset covers multiple Arabic di- alects and accents, it may still contain demographic and regional imbalances that could affect model performance across different speaker groups. We therefore encourage future work on fairness evalu- ation and dialect-aware benchmarking. The dataset is intended solely for academic and research use, and we discourage applications that may compro- mise user privacy or enable harmful surveillance. Overall, this work aims to support more inclu- sive and representative Arabic ASR systems while adhering to standard ethical practices commonly adopted in NLP and speech research. References Ahmed Amine Ben Abdallah, Ata Kabboudi, Amir Ka- noun, and Salah Zaiem. 2023. Leveraging data col- lection and unsupervised learning for code-switched tunisian arabic automatic speech recognition. Mohammad Al-Fetyani, Muhammad Al-Barham, Gheith Abandah, Adham Alsharkawi, and Maha Dawas. 2021. Masc: Massive arabic speech corpus. Mohammad Al-Fetyani, Muhammad Al-Barham, Gheith Abandah, Adham Alsharkawi, and Maha Dawas. 2023. Masc: Massive arabic speech cor- pus. In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 1006–1013. Mansour Alghamdi, Fayez Alhargan, Mohammed Alkanhal, Ashraf Alkhairy, Munir Eldesouki, and Ammar Alenazi. 2008. Saudi accented arabic voice bank. Journal of King Saud University - Computer and Information Sciences, 20:45–64. Sadeen Alharbi, Areeb Alowisheq, Zoltán Tüske, Ka- reem Darwish, Abdullah Alrajeh, Abdulmajeed Al- rowithi, Aljawharah Bin Tamran, Asma Ibrahim, Raghad Aloraini, Raneem Alnajim, et al. 2024. Sada: Saudi audio dataset for arabic. In ICASSP 2024-2024 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP), pages 10286–10290. IEEE. Abbas Raza Ali. 2020. Multi-dialect arabic speech recognition. In 2020 International Joint Conference on Neural Networks (IJCNN), page 1–7. IEEE. Ahmed Ali, Peter Bell, James Glass, Yacine Messaoui, Hamdy Mubarak, Steve Renals, and Yifan Zhang. 2019a. The mgb-2 challenge: Arabic multi-dialect broadcast media recognition. Ahmed Ali, Suwon Shon, Younes Samih, Hamdy Mubarak, Ahmed Abdelali, James Glass, Steve Re- nals, and Khalid Choukri. 2019b. The mgb-5 chal- lenge: Recognition and dialect identification of di- alectal arabic speech.In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1026–1033. Ahmed Ali, Stephan Vogel, and Steve Renals. 2017. Speech recognition challenge in the wild: Arabic mgb-3. Bushra Alkhamees. 2023. The “white dialect” of young Arabic speakers from Qassim (Saudi Arabia). Nether- lands Graduate School of Linguistics. Ammar Mohammed Ali Alqadasi, Akram M Zeki, Mohd Shahrizal Sunar, Siti Zaiton Mohd Hashim, Md Sah hj Salam, and Rawad Abdulghafor. 2025. Arabic dialects speech corpora: A systematic review. Speech Communication, page 103322. Faisal Alshargi, Shahd Dibas, Sakhar Alkhereyf, Reem Faraj, Basmah Abdulkareem, Sane Yagi, Ouafaa Kacha, Nizar Habash, and Owen Rambow. 2019. Morphologically annotated corpora for seven arabic dialects: Taizi, sanaani, najdi, jordanian, syrian, iraqi and moroccan. In Proceedings of the Fourth Ara- bic Natural Language Processing Workshop, pages 137–147. Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gre- gor Weber. 2020. Common voice: A massively- multilingual speech corpus. In Proceedings of the Twelfth Language Resources and Evaluation Confer- ence, pages 4218–4222, Marseille, France. European Language Resources Association. Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, et al. 2023. Seamlessm4t: Massively mul- tilingual & multimodal machine translation. arXiv preprint arXiv:2308.11596. Fatma Zahra Besdouri, Inès Zribi, and Lamia Hadrich Belguith. 2024. Arabic automatic speech recognition: Challenges and progress. Speech Commun., 163(C). Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexan- der Erdmann, et al. 2018. The madar arabic dialect corpus and lexicon. In Proceedings of the eleventh international conference on language resources and evaluation (LREC 2018). Shammur Absar Chowdhury, Amir Hussein, Ahmed Abdelali, and Ahmed Ali. 2021. Towards one model to rule all: Multilingual strategy for dialectal code- switching arabic asr. Amirbek Djanibekov, Hawau Olamide Toyin, Raghad Alshalan, Abdullah Alatir, and Hanan Aldarmaki. 2025. Dialectal coverage and generalization in ara- bic speech recognition. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 29490– 29502. Amazouz Djegdjiga, Martine Adda-Decker, and Lori Lamel. 2018. The French-Algerian code-switching triggered audio corpus (FACST). In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA). Ghania Droua-Hamdani, Sid Ahmed Selouani, and Ma- lika Boudraa. 2010. Algerian arabic speech database (algasd): Corpus design and automatic speech recog- nition application. Arabian Journal for Science and Engineering, 35(2C):157–166. Injy Hamed, Pavel Denisov, Chia-Yu Li, Mohamed Elmahdy, Slim Abdennadher, and Ngoc Thang Vu. 2021.Investigations on speech recognition sys- tems for low-resource dialectal arabic-english code- switching speech. Karima Kadaoui, Samar Magdy, Abdul Waheed, Md Tawkat Islam Khondaker, Ahmed El-Shangiti, El-Moatez-Billah Nagoudi, and Muhammad Abdul- Mageed. 2023. Tarjamat: Evaluation of bard and chatgpt on machine translation of ten arabic varieties. In Proceedings of ArabicNLP 2023, pages 52–75. Mohamed Maamouri, Tim Buckwalter, and Christopher Cieri. 2004. Dialectal arabic telephone speech cor- pus: Principles, tool design, and transcription conven- tions. In Proceedings of the NEMLAR International Conference on Arabic Language Resources and Tools, pages 22–23, Cairo, Egypt. Abir Masmoudi, Fethi Bougares, Mariem Ellouze, Y. Es- tève, and Lamia Hadrich Belguith. 2017. Automatic speech recognition system for tunisian dialect. Lan- guage Resources and Evaluation, 52:249 – 267. Abir Messaoudi, Hatem Haddad, Chayma Fourati, Moez Ben Haj Hmida, Aymen Ben Elhaj Mabrouk, and Mohamed Graiet. 2021. Tunisian dialectal end- to-end speech recognition based on deepspeech. In International Conference on Arabic Computational Linguistics. Hamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, and Ahmed Ali. 2021. QASR: QCRI aljazeera speech resource a large scale annotated Ara- bic speech corpus. In Proceedings of the 59th Annual Meeting of the Association for Computational Lin- guistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2274–2285, Online. Association for Computational Linguistics. Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, et al. 2024. Scaling speech technology to 1,000+ languages. Journal of Machine Learning Research, 25(97):1–52. Alec Radford, Jong Wook Kim, Tao Xu, Greg Brock- man, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak su- pervision. In International conference on machine learning, pages 28492–28518. PMLR. Karin C Ryding. 2005. A reference grammar of modern standard Arabic. Cambridge university press. Tanja Schultz. 2002. Globalphone: a multilingual speech and text database developed at karlsruhe uni- versity. In 7th International Conference on Spoken Language Processing (ICSLP 2002), pages 345–348. Rainer Siemund, Barbara Heuft, Khalid Choukri, Os- sama Emam, Emmanuel Maragoudakis, Herbert Tropf, Oren Gedge, Sherrie Shammass, Asuncion Moreno, Albino Nogueiras Rodriguez, Imed Zitouni, and Dorota Iskra. 2002. OrienTel - multilingual ac- cess to interactive communication services for the mediterranean and the Middle East. In Proceed- ings of the Third International Conference on Lan- guage Resources and Evaluation (LREC’02), Las Palmas, Canary Islands - Spain. European Language Resources Association (ELRA). Peter Sullivan, AbdelRahim Elmadany, Alcides Al- coba Inciarte, and Muhammad Abdul-Mageed. 2026.Arab voices: Mapping standard and di- alectal arabic speech technology. arXiv preprint arXiv:2601.13319. Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Mo- hamedou Cheikh Tourad, Rahaf Alhamouri, Rwaa Assi, et al. 2024. Casablanca: Data and models for multidialectal arabic speech recognition. In Proceed- ings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21745–21758. Omnilingual ASR team,Gil Keren,Artyom Kozhevnikov, Yen Meng, Christophe Ropers, Matthew Setzler, Skyler Wang, Ife Adebara, Michael Auli, Can Balioglu, Kevin Chan, Chierh Cheng, Joe Chuang, Caley Droof, Mark Duppenthaler, Paul-AmbroiseDuquenne,AlexanderErben, Cynthia Gao, Gabriel Mejia Gonzalez, Kehan Lyu, Sagar Miglani, Vineel Pratap, Kaushik Ram Sadagopan, Safiyyah Saleem, Arina Turkatenko, Albert Ventayol-Boada, Zheng-Xin Yong, Yu-An Chung, Jean Maillard, Rashel Moritz, Alexandre Mourachko, Mary Williamson, and Shireen Yates. 2025. Omnilingual asr: Open-source multilingual speech recognition for 1600+ languages. Kees Versteegh. 2014. Arabic language. Edinburgh University Press. Inès Zribi, Rahma Boujelbane, Abir Masmoudi, Mariem Ellouze, Lamia Belguith, and Nizar Habash. 2014. A conventional orthography for Tunisian Arabic. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), pages 2355–2361, Reykjavik, Iceland. European Lan- guage Resources Association (ELRA). Appendix This appendix provides supplementary material to support the main findings of this work. It is orga- nized as follows: • Appendix A: Related Work Additional background on multilingual, uni- dialect, and multi-dialect speech datasets. • Appendix B: Transcription Content Further details of the textual samples recorded in the BULBUL dataset. • Appendix C: Speech Recording Details of the recording tools and guidelines used to collect the BULBUL dataset. • Appendix D: Speech Verification Detailed explanation of the speech verifica- tion tools and guidelines used to verify the recorded speech data. • Appendix E: Dataset Split More information on the BULBUL dataset split. •Appendix F: Benchmarking Experiment Setup Detailed explanation of model configurations, training settings, computational resources, and reproducibility details. A Related Work Despite significant advances in Arabic ASR mod- eling techniques, the availability of diverse, high- quality annotated speech resources remains limited (Besdouri et al., 2024). Arabic presents unique challenges due to diglossia, extensive dialectal vari- ation, and frequent code-switching with foreign lan- guages. Moreover, existing datasets differ substan- tially in linguistic coverage, domain focus, annota- tion standards, and recording conditions, resulting in fragmented resource landscapes. To contextual- ize these limitations, we review prior Arabic speech across multilingual, uni-dialect, and multi-dialect datasets. Multilingual Speech Datasets. Multilingual datasets containing Arabic are essential for lan- guage identification, mixed-language decoding, and cross-lingual transfer learning. However, rel- atively few multilingual datasets incorporate Ara- bic. The GlobalPhone project (Schultz, 2002) in- troduced a multilingual speech and text dataset en- compassing Arabic, collected from Tunisian speak- ers through prompted newspaper readings. While the dataset offers clean transcriptions and con- trolled recordings, it is limited to reading text rather than spontaneous conversational data. Similarly, Mozilla Common Voice (Ardila et al., 2020) is a crowdsourced multilingual dataset containing ap- proximately 15 hours of Arabic speech with demo- graphic metadata including age, gender, and accent, supporting fairness-aware modeling. However, its dialect distribution depended on volunteer partici- pation and therefore lacked systematic balance. In general, while multilingual datasets address cross- lingual modeling challenges, they tend to prioritize language interaction over dialectal diversity within Arabic, limiting their utility for dialect-aware ASR development. Uni-DialectSpeechDatasets. Uni-dialect datasets target a single regional dialect or accented Arabic variety.Some datasets offer a single Arabic dialect mixed with another language to address the Arabic code-switching problem. For example, FACST corpus (Djegdjiga et al., 2018) covers Algerian Arabic mixed with French, ArzEn corpus (Hamed et al., 2021) targets Egyptian Ara- bic–English code-switching, and TUN-SWITCH (Abdallah et al., 2023) proposes Tunisian Arabic speech mixed with French and English. Beyond code-switching, dialect-specific efforts include the SAAVB dataset (Alghamdi et al., 2008), which provides 96 hours of Saudi-accented MSA from 1,033 speakers, and the ALGASD dataset (Droua-Hamdani et al., 2010), which captures regional Algerian diversity from 300 speakers across 11 regions. For Tunisian Arabic, TARIC (Masmoudi et al., 2017), STAC (Zribi et al., 2014), and TunSpeech (Messaoudi et al., 2021) collectively cover spontaneous, broadcast, parliamentary, and read speech genres. While these Uni-dialect corpora offer valuable depth for dialect-specific modeling, their narrow geographic scope limits generalizability across broader Arabic dialect families. Moreover, most of these datasets are collected from broadcast sources, where the voice is recorded in a controlled environment. Multi-Dialect Arabic Speech Datasets. Multi- dialect corpora aim to capture variation across mul- tiple Arabic dialects, enabling ASR systems that generalize across regions. One of the earliest ef- forts, OrienTel (Siemund et al., 2002), contains telephone recordings from six Arabic countries with balanced gender representation and diverse acoustic conditions. The Fisher Levantine Arabic corpus (Maamouri et al., 2004) further contributes approximately 45 hours of spontaneous telephone conversations within the Levantine dialect contin- uum, covering speakers from Jordan, Lebanon, and Palestine. Large-scale broadcast datasets have significantly expanded multi-dialect coverage.The QASR dataset (Mubarak et al., 2021) provides approx- imately 2,000 hours of multi-dialect broadcast speech crawled from Al Jazeera news channel. The MGB challenge series (Ali et al., 2017, 2019a; Ali, 2020) introduced standardized benchmarks draw- ing from broadcast and online media, with later ver- sions expanding dialect coverage and incorporating web-sourced recordings. Similarly, the MASC (Al- Fetyani et al., 2023) offers approximately 1,000 hours of YouTube-sourced speech, covering multi- ple dialects and topics, though at the cost of vari- able recording quality and annotation consistency. The ESCWA.The CS dataset (Chowdhury et al., 2021) contains Arabic–English code-switching in a formal setting, drawn from speech collected at official meetings in Algeria, Tunisia, and Morocco, alternating between Arabic and French. Overall, multi-dialect datasets offer broader lin- guistic coverage than uni-dialect datasets; however, domain biases toward broadcast speech, uncon- trolled acoustic variability in web-sourced data, and inconsistent annotation standards remain signifi- cant challenges to robust and generalized Arabic ASR development. B Transcription Content In this section, we present further details of the textual samples recorded in BULBUL. B.1 Text Pre-processing To ensure high-quality and consistent textual data, we developed an automated Arabic text clean- ing tool tailored for preprocessing raw corpora. The tool operates on plain text files and applies a multi-stage filtering pipeline. The tool is deployed through an interactive Gradio interface, enabling team members to upload datasets, preview cleaned outputs, and download processed files, thereby streamlining the data preparation workflow. The preprocessing pipeline starts with Uni- code normalization, removal of Arabic diacritics, tatweel, RTL marks, HTML tags, parenthetical text, emojis, and ASCII emoticons, while retaining Ara- bic letters, relevant punctuation, and Arabic numer- als. It then normalizes repeated characters, removes excessive laughter tokens, standardizes whitespace, and trims leading/trailing punctuation. For the Tar- jamat MSA/CA subset, additional preprocessing is applied to remove numerical digits, as speakers may verbalize the same number differently, intro- ducing transcription inconsistencies. B.2 Transcription Samples from BULBUL To illustrate the linguistic diversity captured in BULBUL, Table 6 presents representative transcrip- tion samples from each participating country, along with their English translations. B.3 Domain Classification Schema Domain Taxonomy Definition. To ensure con- sistent thematic categorization, we defined an 11- category domain taxonomy covering common top- ical areas observed in the collected Arabic texts: Social, Politics, Economy, Religion, Health, Edu- cation, Technology, Entertainment, Sports, Crime, and Other. Each domain was associated with a concise operational definition to guide classifica- tion and reduce ambiguity between semantically overlapping categories. The Other category was reserved for texts that did not clearly align with any predefined domain and required explicit justifica- tion. Table 7 presents the full domain definitions used during annotation. The highest number of samples were categorized as social. The large pro- portion of the social domain is primarily due to the nature of the text collection strategy. A substantial portion of the collected texts consists of everyday conversational messages from communication plat- forms, such as WhatsApp and Telegram, as well as social media content. We intentionally retained such texts because everyday social communication more naturally reflects dialect-specific vocabulary, expressions, and sentence structures used by na- tive speakers. In contrast, texts from domains such as politics, health, and technology often exhibit a stronger influence from MSA and therefore provide less dialectal variation. Prompt strategy. The prompt instructed the model to assign a single primary domain label, allow- ing up to two labels only when two domains were equally related to the text. If the label Other was assigned, the model was required to provide an ex- Table 6: Representative transcription samples from BULBUL across participating dialects, with English translations. DialectText AGAJ J ” A Ø @ÒÎ @P ̄ P Ò Ø ̇ Œ K . ΩÀ AK . C ́ You know! Fouzi is a family. EG! È J Ø Y g AK ËÒJ . É ! È ́ A‘ g . AK Guys! Let him take his time! JO ̇ Ê Ö B @ † AÎ ·” ̇ Ê . K @P ’Œ JÉ @ h @P I will receive my salary from this thing. KSA ̄ Ò É Ω“ ø @ ̆ ™K . @ H . AJ . À @ Ω Ø Open the door I want to talk to you a bit. MAAK @ X Ym à Šì A” ̇ Õ AK X È Ø A¢J . À @ My card hasn’t arrived yet. TN È ́ AÉ ë H Y™ K Am . ' X It’s been more than half an hour. PS? ΩÉ @QK . ̇ Õ @ Q ́ ̇ Õ @ Ò É What changed your mind? SDI . K Q Ø » Òj À È” CÇÀ @ H . P X Wishing you a safe journy. SY? @Ò “™ K ’∫À AJ . ́ ̄ Ag . Ò É What would you like to do? UAE? Å” @ @Ò KQÂÖ ·K Where did you go yesterday? YE ̈Q ™À @ È” Y g ̇ Ê É @ Ij÷ fi Ö ÒÀ Excuse me, I want room service. plicit justification. The model was also prompted to give a short reasoning statement and a confidence score between 0 and 1, indicating its certainty. Fig- ure 6 illustrates the prompt used for domain classi- fication. We manually reviewed and verified all instances with a confidence score below 0.70 to ensure consistent labeling and reduce potential misclas- sification. The threshold is selected based on an initial manual revision of the classification results. Domain Distribution Statistics. Figure 7 further examines domain coverage across countries. Most major domains, including Social, Economy, Health, You are an expert in Arabic domain classification. Task: Classify the given Arabic text into one or two domains from the following list: Social, Politics, Economy, Religion, Health, Education, Technology, Entertainment, Sports, Crime, Other Domain Definitions: ●Social: Everyday life, family, relationships, social norms, interpersonal interactions. ●Politics: Government, elections, policy, governance, international relations. ●Economy: Employment, business, inflation, taxes, finance, cost of living. ●Religion: Religious teachings, worship, theological discussion, faith-based guidance. ●Health: Illness, treatment, hospitals, medical or mental health topics. ●Education: Schools, universities, exams, scholarships, academic life. ●Technology: AI, software, programming, apps, cybersecurity, digital systems. ●Entertainment: Movies, music, celebrities, TV, influencers, media content. ●Sports: Teams, matches, players, competitions, athletic events. ●Crime: Violence, theft, legal disputes, police or criminal incidents. ●Other: Use only if none of the above apply. Rules: ●Select the primary domain of the text. ●Assign one label by default. ●Assign two labels only if both are equally central. ●If Other is selected, explain why no listed domain applies. Domain Classification Prompt Figure 6: Prompt template used for GPT5-mini domain annotation. Figure 7: Cross-country domain coverage matrix. Filled markers indicate domain presence, and the rightmost column reports domain coverage across the 11 countries. Religion, Education, and Entertainment, are repre- sented in 10 out of 11 countries, indicating broad topical consistency across regional subsets. More specialized domains, such as Crime and Technol- ogy, appear in 8 countries, while Politics and Sports are the least represented, each appearing in 7 coun- tries. Palestine differs from other subsets as it con- sists exclusively of economically oriented utter- ances. Overall, these findings demonstrate that de- spite domain imbalance, the dataset achieves broad cross-country topical coverage, supporting robust evaluation under diverse semantic conditions. Table 7: Domain taxonomy used for text classification. DomainDefinition SocialEveryday life, family, relationships, gender issues, social norms/traditions, gatherings, interpersonal interactions, and personal opinions about society. PoliticsGovernment, political leaders, public policy, elections, governance, corruption, interna- tional relations, and political debate. EconomyMoney, employment, inflation, poverty, business, investment, taxes, cost of living, and financial hardship. ReligionReligious teachings, doctrine, worship, religious rulings, theological discussion, and faith-based guidance. Metaphorical mention of “Allah” alone is not sufficient. HealthPhysical health, illness, mental health, medical advice, hospitals, treatment, and public health. EducationSchools, universities, exams, scholarships, academic life, and learning/teaching pro- cesses. Technology AI, programming, software, platforms/apps, gadgets, cybersecurity, and digital systems. EntertainmentMovies, series, music, celebrities, influencers, and TV/media content. SportsTeams, matches, competitions, players, and athletic events. CrimeTheft, assault, violence, police/court cases, legal disputes, and criminal incidents. OtherUsed only when none of the predefined domains apply; a justification is required. B.4 Lexical and Semantic Analysis To better capture the linguistic diversity in BUL- BUL, we conduct a lexical and semantic analysis across the 11 countries. In addition to cross-country comparisons, we analyze intra-country variation for Saudi Arabia and Yemen, which are further subdivided into sub-dialects. Table 8: Lexical diversity statistics of the BULBUL dataset across countries. The table reports the num- ber of sentences (# Sent.), total number of tokens (Tot. Tokens), vocabulary size (Size; number of unique to- kens), type-token ratio (TTR; ratio of unique tokens to total tokens, measuring lexical diversity), and hapax proportion (H Prop.; proportion of tokens that occur only once, indicating lexical richness and sparsity). Country# Sent.Tot. TokensSizeTTRH Prop. Algeria6326,0552,8110.4640.750 Egypt1,27113,2834,6810.3520.728 Jordan8478,2283,2070.3900.707 Morocco1,09612,1654,1250.3390.707 Palestine1,70014,0312,5810.1840.550 Saudi3,08230,9179,4500.3570.676 Sudan1036604990.7560.860 Syria1,0066,6443,6370.5470.777 Tunisia4293,5031,9490.5560.796 UAE483061910.6240.812 Yemen4,07836,28511,5930.3200.624 Lexical Diversity. Table 8 presents lexical di- versity statistics across countries in the BUL- BUL dataset. Yemen (4,078 sentences; 36,285 to- kens) and Saudi (3,082 sentences; 30,917 tokens) sets contain the largest corpora but show moder- ate type-token ratio (TTR) values of 0.320 and 0.357, respectively. This reflects expected repeti- tion in larger datasets despite their large vocabulary sizes (11,593 and 9,450). In contrast, smaller sub- sets, such as Sudan (103 sentences) and UAE (48 sentences), exhibit very high TTR values (0.756 and 0.624) and high hapax proportions (0.860 and 0.812), likely inflated by limited data. Countries with broader thematic coverage, such as Syria (TTR = 0.547; hapax prop. = 0.777) and Tunisia (TTR = 0.556; hapax prop. = 0.796), demonstrate strong lexical richness. Notably, Palestine shows the lowest TTR (0.184) and a relatively lower ha- pax proportion (0.550), as expected, since its text is restricted to a single domain: the Economy. Over- all, lexical diversity in the dataset is influenced by both corpus scale and domain heterogeneity. Cross-Country Lexical Overlap.Figure 8 presents the cosine similarity heatmap based on TF-IDF representations over the top 100 highest- variance terms. Overall, the heatmap highlights identifiable regional clustering patterns along- side strong cross-dialect distinctiveness across the dataset. The highest similarity is observed between Saudi Arabia and Yemen (0.57), which suggests strong Gulf lexical alignment. Similarly, Jordan shows high similarity with Palestine (0.52) and Syria (0.49), forming a clear Levantine cluster. Some similarity patterns deviate from linguistic ex- pectations due to dataset composition. For example, the low similarity between Syrian and Palestinian dialects is likely due to the domain imbalance. Semantic Similarity. Figure 9 illustrates cross- country semantic similarity in the BULBUL dataset. We used the Arabic-arabert-all-nli-triplet sentence embedding model to encode each sentence and computed country-level centroids by averaging sen- tence embeddings. Because the number of avail- dzegjoksamapssdsytnuaeye dz eg jo ksa ma ps sd sy tn uae ye 1.000.050.040.160.310.020.010.040.260.050.20 0.051.000.300.240.010.150.190.070.090.070.18 0.040.301.000.360.020.520.060.490.170.220.37 0.160.240.361.000.070.210.060.280.240.260.57 0.310.010.020.071.000.030.000.030.100.020.07 0.020.150.520.210.031.000.020.300.150.120.26 0.010.190.060.060.000.021.000.070.030.040.03 0.040.070.490.280.030.300.071.000.160.170.21 0.260.090.170.240.100.150.030.161.000.040.29 0.050.070.220.260.020.120.040.170.041.000.23 0.200.180.370.570.070.260.030.210.290.231.00 0.2 0.4 0.6 0.8 1.0 Cosine Similarity Figure 8: Lexical similarity heatmap across Arabic di- alectal corpora in the BULBUL dataset, computed using cosine similarity between country-level TF–IDF repre- sentations. able sentences varied substantially across countries, we limited our computation by randomly sampling n sentences per country to avoid corpus-size bias, where n=48, which is the minimum number of available sentences across all countries. We re- peated the procedure five times with different ran- dom seeds and computed cosine similarity between country centroids in each run. Then, we reported the mean similarity matrix across the five runs to ensure robustness. The results indicate generally high semantic sim- ilarity across dialects, mostly greater than 0.75, which reflects a strong shared semantic backbone despite dialectal variation. This can also be at- tributed to the small subset of samples used in our analysis. A prominent similarity cluster emerges among Jordan, Egypt, Saudi Arabia, Syria, Tunisia, and Yemen. However, Morocco shows slightly lower similarity to the Eastern dialects, consis- tent with regional divergence. The Palestinian di- alect exhibits comparatively lower similarity scores (0.45–0.69), which reflect the topical or lexical characteristics specific to the dataset. Overall, the findings suggest that Arabic dialects, while lexi- cally diverse, remain semantically aligned at the sentence level. Sub-dialect Analysis. In the Yemeni dialect, we identified 1,464 shared vocabulary items between the Sana’ani and Ta’izzi sub-dialects, including terms such as "what do you want?" ̇ Ê Ç A”, "ours" A J Æk , and "leave"h Q£ @. We also observed sys- tematic lexical differences between the two sub- ag eg jo ksa ma pssd sy tu uae ye ag eg jo ksa ma ps sd sy tu uae ye 1.000.880.840.850.850.570.840.870.870.640.83 0.881.000.920.900.810.630.780.870.850.640.85 0.840.921.000.910.800.630.750.870.820.650.86 0.850.900.911.000.800.660.770.870.850.680.91 0.850.810.800.801.000.590.740.820.790.590.76 0.570.630.630.660.591.000.490.580.530.450.69 0.840.780.750.770.740.491.000.880.850.690.81 0.870.870.870.870.820.580.881.000.880.730.87 0.870.850.820.850.790.530.850.881.000.700.85 0.640.640.650.680.590.450.690.730.701.000.72 0.830.850.860.910.760.690.810.870.850.721.00 0.0 0.2 0.4 0.6 0.8 1.0 Cosine Similarity Figure 9: Mean country-level semantic similarity heatmap across the countries from the BULBUL dataset, computed using sentence embeddings from the Arabic NLI model Arabic-arabert-all-nli-triplet. dialects. For example, Ta’izzi speakers commonly use ‡ A Ç ́, whereas Sana’ani speakers prefer ‡ A Ç ́ for "because". Similarly, Ta’izzi often uses Å ́ for "why", while Sana’ani typically uses Å À. In the Saudi dialect, many differences between sub-dialects are not lexical substitutions but pho- netic or phonological variations. For example, the word "coffee" ËÒÍ Ø is commonly realized as gahwa (/gahwa/) in Najdi speech. In contrast, southern Saudi speakers pronounce it as ghawa (/ghawa/), reflecting dialect-specific phonological restructur- ing that simplifies the consonant sequence. Al- though both forms correspond to the same ortho- graphic word, they exhibit distinct phonotactic pat- terns, illustrating how Saudi micro-dialects can di- verge in pronunciation even when vocabulary is shared. C Speech Recording This section presents details of the recording tools and guidelines used to record speech data for di- alectal and accented formal Arabic. C.1 Recording Tools We employed Gradio to develop a lightweight web- based annotation interface for dataset collection and validation. The interface enabled participants from different countries to record the text corre- sponding to their country. To ensure high-quality recording, we added the guidelines within the in- 3 https://w.gradio.app/ terface. We customized the tool to show each par- ticipant’s progress and the country’s leaderboard to motivate the participants. We also allow par- ticipants to review, trim, and repeat the recording before saving it and moving to the next sample. A Skip button is also added to enable participants to skip the text and move to another sample. C.2 Recording Guidelines We provided the participants with guidelines to ensure a controlled environment. The following guidelines were provided to the participants: •Environment: Record in a quiet place and make sure no other human voices appear in the recording. Also, avoid loud background noise that can make the voice inaudible. • Microphone: It is preferable to use an exter- nal microphone or headset, as it is typically clearer than a built-in laptop microphone. If using a mobile phone, ensure the recording quality before proceeding. • Speaking Style: Read each sentence clearly and naturally in your own voice. Do not alter or replace any words unless they reflect natu- ral pronunciation variations (e.g., differences in local pronunciation). If you cannot record a specific sentence or encounter difficulty pro- nouncing it, you may use the Skip option to move to the next sentence. • Editing: You may modify the sentence before starting the recording. •Saving: After completing a recording, click Save and Next to save your audio. To re- record, delete the current recording using the Delete (X) button. To move to the next sen- tence without recording, click on Skip. C.3 Recording Agreement Letter All participants who volunteered to join this study were required to agree to the terms outlined in the consent form presented in Figure 10. The original consent form was presented to participants in Ara- bic, their native language; the English translation provided here is intended solely for clarity. D Speech Verification This section presents details of the speech verifica- tion tools and guidelines used to verify the recorded speech data for dialectal Arabic. D.1 Verification Tool To ensure that all recordings adhere to the estab- lished guidelines and that the evaluation samples remain well-grounded, we developed a verifica- tion tool that enables secondary-speaker oversight during the recording process. The tool allows re- viewers to either accept or reject each recording sample. In cases of rejection, the reviewer must se- lect from a predefined set of rejection reasons. This structured process helps ensure that evaluation de- cisions are justified, consistent, and less susceptible to reviewer bias. D.2 Verification Guidelines Recordings that did not meet any of these crite- ria were rejected. The reviewer could select one or more of the following primary reasons for re- jection: unclear recording, text–audio mismatch, prolonged silence, dialect mismatch, or other. If unclear recording was selected, the reviewer was additionally required to choose exactly one of the following sub-reasons: the presence of background noise (e.g., music, urban noise, or audible conver- sations), more than one clearly audible speaker, no- ticeable echo, or low volume that made the record- ing difficult to hear. Additionally, if text–audio mismatch was selected, the reviewer was addition- ally required to choose exactly one of the follow- ing sub-reasons: complete mismatch, partial mis- match, or minor discrepancies (e.g., one or two mispronounced words). Recordings rejected due to unclear recording will be published as a separate noise set for researchers interested in evaluating their ASR systems on noisy speech samples. This tool streamlined the review process by al- lowing reviewers to systematically evaluate each sample and select from a predefined set of rejection reasons. Such a structured workflow helps reduce bias and improve consistency during evaluation. If none of the predefined reasons adequately de- scribes the issue, reviewers can select Other and provide a free-text comment to justify their deci- sion and support their assessment. E Dataset Split Statistics The experimental results presented in Tables 4 and 5 report the normalized WER and CER obtained from evaluating different ASR models. In this sec- tion, we further describe the distribution and com- position of the evaluation samples to provide a more comprehensive understanding of the reported Table 9: Statistics of the dataset splits across dialectal and accented datasets. We report the number of utter- ances (Utts), total minutes (Mns), and speakers (Spk). Dialectal ArabicAccented Arabic dial DevTestTest Utts Mns SpkUtts Mns SpkUtts Mns Spk Algeria33632.2430332.355517.521 Egypt38534.31431334.8165515.861 Jordan 60756.9772657.07275.091 Morocco15121.2220120.435517.581 Palestine 78269.1475068.955516.131 Saudi1219.681239.685512.401 Sudan 13913.6516213.253010.171 Syria996 105.881317 105.385518.111 Tunisia 62662.2274472.23205.761 UAE19314.9253838.0661.041 Yemen 46050.01747649.9175514.981 metrics. Table 9 summarizes the number of utterances, total duration in minutes, and number of speakers for the development and test splits across the evalu- ated dialect datasets. The development splits were primarily used in the error analysis discussed in Section 4. Table 9 also presents the statistics of the accented MSA/CA dataset. Due to the limited availability of speakers capable of recording ac- cented MSA/CA samples, all recordings for each country were collected from a single speaker. F Benchmarking Experiment Setup Due to computational resource availability, experi- ments were conducted across multiple computing environments, including Google Colab and insti- tutional server machines equipped with NVIDIA GPUs. The hardware resources utilized in this work included NVIDIA T4, A100, and H100 GPUs on Google Colab, as well as RTX 3090, RTX A6000, and A100 GPUs on servers provided by the SDAIA–KFUPM Joint Research Center. The selection of hardware depended on the experiment scale and resource availability. All experiments were implemented in Python 3. Since the evaluated systems consisted of three dis- tinct ASR model families with different software requirements and dependencies, separate execution environments were maintained for each model fam- ily. On the institutional servers, isolated Anaconda environments were used to manage dependencies and ensure compatibility across models, while sep- arate Google Colab notebooks were configured for corresponding experimental setups. The experiments in this work focused exclu- sively on inference using pre-trained ASR models without additional training or fine-tuning. There- fore, variations in hardware configurations primar- ily affected execution time and resource utilization rather than transcription quality or evaluation out- comes. To ensure fair and reproducible compar- isons, all inference experiments were conducted using consistent model weights, inference settings, and software configurations across environments. F.1 Text Normalization To ensure that the reported CER and WER accu- rately reflect transcription quality rather than super- ficial formatting differences, the same text normal- ization pipeline was applied to both the reference and predicted transcripts. First, Unicode normaliza- tion (NFKC) was performed to ensure a consistent character representation. Next, Tatweel (-) charac- ters and Arabic diacritics (Tashkeel) were removed, as they generally do not contribute to the lexical meaning of words in Arabic NLP tasks. Arabic- Indic numerals were then converted to their ASCII equivalents, and all Alef variants (e.g., @ , @ , @) were normalized to the standard Alef (@). Finally, punc- tuation marks and special symbols were removed, and consecutive whitespace characters were col- lapsed into a single space after trimming leading and trailing whitespace. Consent Form for Data Collection and Usage This agreement is entered into by and between the participant and the research team from King Fahd University of Petroleum and Minerals (KFUPM) and Taibah University (hereinafter referred to as “the Universities”). The purpose of this agreement is to collect, use, and distribute audio recordings to support audio deepfake detection research and other non-commercial academic pursuits. 1. Purpose of Data Collection. The research team is collecting audio recordings to construct a dataset dedicated to detecting synthetic voices generated via text-to-speech (TTS), voice conversion (VC), and other generative techniques. These data will be utilized in scientific and academic research to develop advanced methodologies for deepfake detection and related non-commercial research. 2. Nature of Collected Data. The participant agrees to provide: • Audio Recordings: Speech samples in their natural voice or generated by reading specified texts/sentences. • Optional Metadata: Demographic details such as gender, age group, dialect, and other relevant metadata. • Modification Consent: Permission to modify, alter, or synthesize their voice using artificial intelligence and speech processing techniques. 3. Granted Rights. The participant grants the research team the full, royalty-free, and unrestricted right to: • Record, process, and utilize both their natural voice and any synthetic variants derived from it. •Distribute the resulting dataset (comprising both natural and synthetic audio) to the broader scientific community strictly for non-commercial research purposes. •Publish brief audio samples on professional or academic platforms (such as LinkedIn, X/Twitter, and YouTube) to promote deepfake research awareness or announce dataset availability. 4. Data Availability and Licensing. The audio dataset (both natural and synthetic components) will be publicly released under a Creative Commons Attribution-NonCommercial 4.0 International (C BY-NC 4.0) license, allowing researchers to utilize and share the data for non-commercial academic purposes. 5. Privacy and Confidentiality. • The participant’s name and any directly identifying personal information will remain strictly confidential and will not be published without explicit written consent. •To ensure anonymity, each participant will be assigned a unique pseudonymized identifier (ID) within the dataset. 6. Voluntary Participation and Withdrawal. • Participation is entirely voluntary (100%). • Participants retain the right to withdraw from the study or request the deletion of their recordings at any point prior to the public release of the dataset. •Once the dataset is publicly released, removal of the data will no longer be possible due to the decentralized nature of open-source data distribution. 7. Compensation. The participant acknowledges that participation is voluntary and carries no financial compensation. The contribution is made solely to support and advance scientific research. *By creating an account, you explicitly acknowledge that you have read, understood, and agreed to all the terms and conditions outlined above. Figure 10: Consent form for data collection and usage.