Paper deep dive
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
K. M. Jubair Sami, Dipto Sumit, Ariyan Hossain, Farig Sadeque
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/26/2026, 2:27:03 AM
Summary
The paper presents a two-phase framework to evaluate and benchmark dialectal bias in Large Language Models (LLMs) across nine Bengali dialects. It utilizes a RAG-based translation pipeline to generate gold-labeled question sets and an RLAIF-based evaluation framework with multi-judge validation and human fallback to quantify performance disparities. The study introduces the Critical Bias Sensitivity (CBS) metric and reveals that increased model scale does not consistently mitigate dialectal bias.
Entities (5)
Relation Signals (3)
RLAIF → evaluates → LLM
confidence 95% · We benchmark 19 LLMs across these gold-labeled sets, running 68,395 RLAIF evaluations
Gemma 3-27b-it → usedin → RAG Translation Pipeline
confidence 95% · For translation generation, we used Gemma-3-27B-IT, the best-performing mid-weight open-source model
Critical Bias Sensitivity → measures → Dialectal Bias
confidence 90% · We contribute... a Critical Bias Sensitivity (CBS) metric for safety-critical applications
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) frequently exhibit performance biases against regional dialects of low-resource languages. However, frameworks to quantify these disparities remain scarce. We propose a two-phase framework to evaluate dialectal bias in LLM question-answering across nine Bengali dialects. First, we translate and gold-label standard Bengali questions into dialectal variants adopting a retrieval-augmented generation (RAG) pipeline to prepare 4,000 question sets. Since traditional translation quality evaluation metrics fail on unstandardized dialects, we evaluate fidelity using an LLM-as-a-judge, which human correlation confirms outperforms legacy metrics. Second, we benchmark 19 LLMs across these gold-labeled sets, running 68,395 RLAIF evaluations validated through multi-judge agreement and human fallback. Our findings reveal severe performance drops linked to linguistic divergence. For instance, responses to the highly divergent Chittagong dialect score 5.44/10, compared to 7.68/10 for Tangail. Furthermore, increased model scale does not consistently mitigate this bias. We contribute a validated translation quality evaluation method, a rigorous benchmark dataset, and a Critical Bias Sensitivity (CBS) metric for safety-critical applications.
Tags
Links
- Source: https://arxiv.org/abs/2603.21359v1
- Canonical: https://arxiv.org/abs/2603.21359v1
Trouble viewing inline? Open PDF directly →
Full Text
51,571 characters extracted from source content.
Expand or collapse full text
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF K. M. Jubair Sami, Dipto Sumit, Ariyan Hossain, Farig Sadeque Department of Computer Science and Engineering BRAC University, Dhaka, Bangladesh km.jubair.sami,dipto.sumit@g.bracu.ac.bd ariyan.hossain,farig.sadeque@bracu.ac.bd Abstract Large language models (LLMs) frequently exhibit performance biases against regional dialects of low-resource languages. How- ever, frameworks to quantify these dispari- ties remain scarce. We propose a two-phase framework to evaluate dialectal bias in LLM question-answering across nine Bengali di- alects. First, we translate and gold-label standard Bengali questions into dialectal vari- ants adopting a retrieval-augmented genera- tion (RAG) pipeline to prepare 4,000 ques- tion sets. Since traditional translation quality evaluation metrics fail on unstandardized di- alects, we evaluate fidelity using an LLM-as-a- judge, which human correlation confirms out- performs legacy metrics. Second, we bench- mark 19 LLMs across these gold-labeled sets, running 68,395 RLAIF evaluations validated through multi-judge agreement and human fall- back. Our findings reveal severe performance drops linked to linguistic divergence. For instance, responses to the highly divergent Chittagong dialect score 5.44/10, compared to 7.68/10 for Tangail. Furthermore, increased model scale does not consistently mitigate this bias. We contribute a validated translation quality evaluation method, a rigorous bench- mark dataset, and a Critical Bias Sensitivity (CBS) metric for safety-critical applications. 1 Introduction Large Language Models (LLMs) have achieved re- markable performance across diverse NLP tasks, yet their behavior on dialectal variants of low- resource languages remains poorly understood (Fleisig et al.,2024;Hofmann et al.,2024). This gap is critical because dialectal variations in low- resource settings create severe digital divides, marginalizing vast speaker populations. We ex- plore this broader challenge using Bengali as a rep- resentative case study, as its regional dialects spo- ken by millions diverge substantially from the stan- dardized written form (Wasi et al.,2025). Such dialectal variations, whether in Bengali (e.g., Chittagong, Sylhet) or other low-resource languages like Arabic, exhibit distinct phonolog- ical, lexical, and syntactic features that confuse LLMs trained predominantly on standard forms (Sami et al.,2025;Jawad et al.,2025). Unlike stan- dardized language that benefits from large training corpora, dialectal variants face severe data scarcity, creating potential disparities in model comprehen- sion and response quality ( Chang et al.,2024;Sind- hujan et al. ,2025). We address this challenge through a two-stage framework:(1)Adopting a high-performance RAG-based translation pipeline (Sami et al.,2025) that translates standard Bengali questions into dialectal variants for benchmark construction, and(2)An RLAIF-inspired evaluation frame- work, with human fallback and multi-judge valida- tion that quantifies LLM performance disparities across dialects using validated scoring rubrics. Our contributions are: •A human-validated translation evaluation methodology for standard-to-dialect Bengali, demonstrating the catastrophic failure of tra- ditional metrics •A gold-standard benchmark dataset of 4,000 questions across 9 Bengali dialects for bias evaluation in LLM question-answering •An RLAIF bias evaluation framework with Chain-of-Thought enabled rubrics, validated through multi-judge agreement analysis (Lin (1989)’s Concordance Correlation Coeffi- cient (C) = 0.861), and human inspection •A comprehensive benchmark of 19 open- weight LLMs across 9 dialects (68,395 evalu- 1 arXiv:2603.21359v1 [cs.CL] 22 Mar 2026 ations), revealing systematic bias patterns •A novel Critical Bias Sensitivity (CBS) met- ric for safety-critical applications requiring high judge agreement on critical bias cases 2 Related Works 2.1 Bias in Large Language Models Bias in LLMs manifests across multiple dimen- sions including gender, race, religion, and so- cioeconomic status (Gallegos et al.,2024). Re- cent work established frameworks for systematic bias evaluation (Liang et al.,2023), though dialec- tal bias remains understudied compared to demo- graphic dimensions. Fleisig et al.(2024) demonstrated that ChatGPT exhibits linguistic bias, providing lower-quality re- sponses to users of non-standard English dialects. Hofmann et al.(2024) found that dialect prejudice in LLMs predicts discriminatory decisions about character, employability, and criminality. These findings motivate our investigation into dialectal bias for Bengali. 2.2 Bengali NLP and Dialectal Variation Bengali NLP research has expanded significantly, with benchmarks like BenLLMEval ( Kabir et al., 2024) evaluating LLM capabilities. While new dialectal resources are emerging, such as Vashan- tor (Faria et al.,2025) for translation, BanglaD- ial (Mahi et al.,2025) for identification, and DIALTSA-BN ( Jawad et al.,2025) for down- stream benchmarks, dialectal variation remains broadly underexplored. Alongside resource cre- ation, bias auditing has revealed systematic reli- gious dialect disparities ( Wasi et al.,2025) and broader socio-cultural biases (Sadhu et al.,2025, 2024) in Bengali LLMs. Our work extends this line by specifically focusing on regional dialectal bias. Furthermore, while recent RAG-based di- alect translation models ( Sami et al.,2025) show promise, their evaluation relied heavily on tradi- tional token-matching metrics (BLEU (Papineni et al. ,2002), WER, ChrF (Popović,2015), and BERTScore (Zhang* et al.,2020)). Because these metrics fail to capture true semantic equivalence in highly agglutinative languages like Bengali (Re- iter,2018;Lee et al.,2023), we investigate more robust embedding-based ( Rei et al.,2020;Sellam et al.,2020;Lo,2019) and LLM-as-judge (Sind- hujan et al.,2025) evaluation methods for dialect translation quality. 2.3 LLM-as-Judge Evaluation LLM-based evaluation has emerged as a scalable alternative to human annotation (Zheng et al., 2023). Recent work improves judge alignment with humans via rubric-style prompting and Chain- of-Thought guided evaluation (Liu et al.,2023). While concerns about self-enhancement bias ex- ist (Panickssery et al.,2024;Xu et al.,2024), multi-judge validation can ensure reliability.Sind- hujan et al.(2025) specifically highlighted the challenges of reference-less evaluation for low- resource languages, proposing refined prompt- based approaches. Broader surveys also sys- tematize known judge failure modes (e.g., bias, leakage, inconsistency) and mitigation strategies (Li et al.,2025;Gu et al.,2025). Our RLAIF framework extends this paradigm with Chain-of- Thought enabled rubrics and multi-judge valida- tion protocols. 3 Methodology Figure1illustrates the complete architecture of our framework. 3.1 Translation Pipeline Construction & Evaluation To generate dialectal translations of the stan- dard Bengali questions for bias evaluation, we adopted the optimizedStructured Sentence-Pair RAGpipeline (Pipeline 2) from Sami et al.(2025). For translation generation, we used Gemma-3- 27B-IT, the best-performing mid-weight open- source model identified in that study, operating via Pipeline 2. 3.1.1 Indexing and Datasets To construct the indexes for the RAG based trans- lation pipeline, we utilized 2 datasets conain- ing parallel standard_bengali:dialectal_translation sentence pairs: Dataset: Standardized Parallel Corpus (Has- san et al.,2025;Dipto et al.,2025):20,635 structured sentence pairs from existing Bengali dialect (Chittagong, Habiganj, Rangpur, Kishore- ganj, Tangail) corpora, providing aligned dialectal and standard Bengali variants. Dataset: Vashantor Benchmark (Faria et al., 2025):12,500 Bengali sentence pairs paired with standard Bengali and five regional dialects (Chittagong, Noakhali, Sylhet, Barishal, My- mensingh) containing 2,500 sentence pairs each. 2 Methodology: Benchmarking Bengali Dialectal Bias in Multilingual LLMs Translation Pipeline Construction & Evaluation Parallel Corpora Std. Corpus: 20,635 pairs Vashantor: 11,250 (RAG) + 1,250 (val.) RAG Translation Pipeline a) Vector Index Creation FAISS (dense) + BM25 (sparse) b) Adaptive Hybrid Retrieval Dense + Sparse + Fallback Search c) LLM Translation Gemma-3-27B-IT Evaluation Protocols a) Legacy & Embedding Metrics BLEU, WER, ChrF, BERTScore (L3Cube) Gemini Embedding (3072-dim) b) LLM-as-a-Judge Gemma-3-27B + GPT-OSS-120B c) Human Annotation n=125 (25/dialect) Pearson, Spearman, C ⭐ Finding LLM-as-Judge = Superior Metric (C=0.506 with Human Judge) Pipeline Selected RAG + Gemma-3-27B-IT → Proceed to next Phase Bias Benchmark Generation & Gold Labeling Question Set Generation 400 Questions Across 4 Types Dialect Translation Using RAG Pipeline Standard → 9 Dialectal Variants Dialectal Benchmark 400 Standard + 3,600 Dialectal (9 dialects) Human Gold Labeling Native Speaker Review & Correction (36 Human Annotators) Final Cross- Checking Gold Standard Dataset 4,000 Question Pairs (Standard + 9 Dialects) Bias Measurement (RLAIF Framework with Multi-Judge Validation) Gold Standard Dataset (from prev. Phase) 19 LLMs Under Test (Open-weight) Standard Questions Responses Dialectal Questions Responses RLAIF Primary Judge: Gemini 2.5 Flash CoT-Enabled Rubric 5 Categories 6-Point Likert Scale (Norm. 10) Response Pair Comparison Analysis Bias Score + Confidence Per Category Human Fallback When Confidence ≤ 3 Re-examined by humans Gold-Labeled Bias Scores 68,395 Evaluations (19 models × 3,600 pairs) Multi-Judge Validation Secondary Judge 1: GPT-OSS-20B Secondary Judge 2: Gemma 3 27B IT Correlation Analysis (Lin's C, CBS) Comprehensive Bias Analysis Report • Per-Model Bias Scores (19 LLMs) • Per-Dialect Analysis (9 Bengali Dialects) • Per-Question Type Breakdown (4 Types) • Multi-Judge Validation Results Key Statistics • Validated Translation Pipeline • 19 LLMs for Bias Testing • 400 Questions, 4 Types, 6 Domains • 4,000 Gold-Labeled Pairs (9 Dialects) • 68,395 Final Bias Evaluations Figure 1: Overview of the dialectal bias measurement framework. The pipeline translates standard Bengali ques- tions into dialectal variants via Retrival-Augmented Generation, which are then used to probe LLMs with RLAIF- based scoring. The training and testing splits were combined to build the RAG retrieval indexes (11,250 pairs), while the validation splits (1,250 pairs) were strictly reserved for the translation evaluation phase. 3.1.2 Retrieval Module To construct the few-shot context for translation generation, we relied on the hybrid vector-based retrieval system introduced by Sami et al.(2025). Rather than utilizing a static retrieval approach, this module employs dynamic weighting to handle standard and fragmented inputs effectively. The process consists of three core stages: Input Normalization and Tagging:The stan- dard Bengali query undergoes thorough normal- ization (e.g., Unicode standardization and numeral conversion). Queries containing fewer than four tokens are explicitly appended with a[[SHORT]] tag to isolate them during the lexical matching phase. Adaptive Hybrid Retrieval:The system iden- tifies relevant sentence pairs by fusing dense and sparse retrieval methods. Dense retrieval captures semantic equivalence using a sentence transformer and FAISS cosine similarity search, while BM25 sparse retrieval captures exact lexical overlaps. The module applies adaptive weighting based on the query length: standard queries favor dense re- trieval, whereas short queries prioritize sparse re- trieval and expand the candidate pool to ensure suf- ficient contextual matches. Fallback Search and Blended Scoring:If the initial retrieval lacks diversity (yielding fewer than two unique examples), a token-level “Deep Search” fallback is triggered. Finally, all retrieved candidates are ranked using a blended score that aggregates the hybrid similarity metrics alongside bonuses for target district matching and character- level similarity. The top-ranked standard-dialect pairs are then formatted as few-shot examples to guide the language model. 3.1.3 Translation Quality Evaluation WhileSami et al.(2025) validated their RAG pipeline using BLEU, WER, ChrF, and BERTScore (via L3Cube (Deode et al.,2023) embeddings), we identified critical limitations in these metrics when applied to Bengali dialects. Bengali is a highly agglutinative language, and in informal or dialectal contexts, word spacing is highly inconsistent (e.g., ‘ভালা লােগ না’ vs ‘ভালালােগনা’, meaning ‘does not feel good’). Consequently, traditional n-gram/word bound- ary metrics (BLEU, WER) often completely fail due to tokenization artifacts, even when sentences are semantically identical. Furthermore, we found that subword-based BERT models severely penal- ize cases like spatial inconsistencies, dropping sim- ilarity scores significantly despite human equiva- lence. To conduct a robust assessment of translation accuracy, we proposed two complementary ap- proaches: semantic similarity using a higher di- mentional, proprietary embedding model, and an LLM-as-a-judge scoring protocol. For the embedding-based evaluation, using 1,238 vali- dation pairs from the Vashantor dataset across all five dialects, we computed cosine similarity and BERTScore between the generated transla- 3 tions and human gold references using the 3072- dimensional Gemini Embedding-001 embedding model. We additionally evaluated BERTScore (Zhang* et al.,2020) using the L3Cube Bengali sentence- similarity model (Deode et al.,2023) as contex- tual embedding baselines alongside the legacy lex- ical metrics BLEU (Papineni et al.,2002), ChrF (Popović,2015), and WER. LLM-as-a-Judge for Translation FidelityFol- lowing the same Chain-of-Thought-first paradigm used in our RLAIF bias evaluation (§3.3), we de- veloped a LLM-as-a-judge approach specifically for translation quality assessment. The judge LLM assumes the persona of a native speaker of the target dialect and scores the machine translation against the human reference on a 0–10 integer scale, prioritizingphonetic equivalenceover sur- face orthography to account for non-standardized Bengali dialectal spelling. The prompt enforces a three-step CoT:Step 1 exempts phonetically equivalent spellings (e.g., খরইন/কেরাইন, meaning ‘does’), digit–word al- ternations, whitespace variants (ভালা লােগ নাvs. ভালালােগনা, meaning ‘does not feel good’), and terminal punctuation;Step 2counts genuinely in- accurate or meaning-shifted words;Step 3maps that count to a strict integer score with hard ceil- ings (one inaccuracy⇒score≤7; two⇒score ≤6). The judge returns structured JSON in which reasoning is generatedbeforethe integer score, pre- venting post-hoc rationalization. Each evaluation receives four inputs: the stan- dard Bengali source, an English translation, the hu- man reference dialect translation, and the machine translation. Two judges were employed: Gemma- 3-27B-IT and GPT-OSS-120B across the complete 1,238 successful translations of the Vashantor val- idation split. Human Annotation for Metric ValidationTo determine which automated metric best reflects genuine translation quality, we conducted a row- level correlation study. A stratified random sam- ple of 25 translation pairs per dialect (N=125 to- tal) was drawn from the Vashantor validation split. Native speaker annotators (Appendix C) scored each pair on the same 0–10 scale as the LLM judge, judging how closely the machine translation matched the human reference present in the dataset. All automated metrics were normalized to [0,1] prior to correlation analysis. Row-level Pearsonr, Spearmanρ, andLin(1989)’s Concordance Corre- lation Coefficient (C) were then computed be- tween each automated metric and the normalized human scores. 3.2 Question Generation & Gold-Labeling We generated evaluation questions across four types designed to probe different comprehension aspects: •Type 1: Definitional Questions: Framework:[িবষয ় ] কােক বেল? / [িবষয ় ] বলেত কী েবাঝায ় ?(Translation: “What is [Topic]? / What is meant by [Topic]?”) •Type 2: Contrasting Questions: Framework:[বস্তু -১] এবং [বস্তু -২]-এর মেধয প্ ę ধান পাথর্কয কী?(Translation: “What is the main difference between [Object-1] and [Object-2]?”) •Type 3: Factual Identification & Enumer- ation Questions: Framework:[েপ্ ę ক্ষাপট]-এর [িবষয ় ]-িঢট র নাম কী? / [িবষয ় ]-এর সংখযা কত?(Translation: “What is the name of the [Topic] in [Context]? / What is the number of [Topic]?”) •Type 4: Functional/Purpose-Based Ques- tions: Framework:[বস্তু ]-িঢট কী কােজ বযবহ ধ Śত হয ় ? / [িবষয ় ]-এর প্ ę ধান কাজ কী?(Translation: “What is the [Object] used for? / What is the main function of the [Topic]?”) Questions spanned six knowledge domains: Technology (count=85/400), Social Sciences (85), Health & Sports (41), Physical & Natural Sci- ences (115), Arts & Humanities (34), and Busi- ness & Economics (40), enabling analysis of genre- specific dialectal effects across both technical and cultural topics. After preparing this 400 base question sets in Standard Bengali, we used the translation pipeline to generate a total of 4,000 question sets across 9 dialectal variations (dialects not supported by the pipeline were translated manually). To ensure fair- ness, the dialectal translations were entirely cor- rected and gold-labeled by human annotators (Ap- pendix C) native to each dialect region. Using these 4,000 question sets benchmark, we generated responses using 19 open-weight LLMs (details deferred to Appendix D), totaling 76,000 responses. We prompted the LLMs to generate the 4 responses in standard Bengali for fairer bias assess- ment. Example Prompt (Sylhet): তলর ফশ্ ন াটার উত্তর খািঢট বাংলাত েদইন। [Answer the following question in stan- dard Bengali.] প্ ę শ্ ন া:[Question: ] (খািল ফশ্ ন াটার উেত্তার িদবা।)[(Only pro- vide the answer to the question.)] 3.3 RLAIF Evaluation Framework To evaluate the bias present in the generated re- sponses, we employed a proprietary LLM as the primary judge. The judge LLM was given both the standard and dialectal questions, and their gen- erated responses. A detailed evaluation rubric, guidelines, confidence score generation (of judge) guidelines were also provided. Theoretical FoundationInspired by Reinforce- ment Learning from AI Feedback (Bai et al., 2022), we designed a structured evaluation frame- work grounded in recent advances in LLM-based evaluation reliability. Tian et al.(2023) demon- strated that raw scalar values suffer from cali- bration gaps due to false precision, necessitating verbally-anchored discrete scales.Zheng et al. (2023) established that Chain-of-Thought (CoT) reasoningbeforescore assignment is mandatory for alignment with human judges, preventing hal- lucinated scores. LikertScaleBasedJudgmentsThe judge LLM was asked to express their agreements using a Lik- ert scale on 5 different statements as part of the evaluation (Table 1). We implemented a 6-point Likert scale ranging from 0 (Strongly Disagree) to 5 (Strongly Agree), with natural language anchors as suggested byTian et al.(2023) for improved cal- ibration. WeightSelectionOur designed statements were based on five weighted categories (Table1): Weights were normalized such that the maxi- mum possible score is 10.0, calculated as: Score f inal = N X i=1 w i · L i L max (1) where w i is the weight for category i , L i is the as- signed Likert score (0–5),L max is the maximum possible Likert value (5), andNis the total num- ber of evaluated categories (5). Script Validity and CoT-First ScoringTo en- sure evaluation integrity, we implemented a strict Bengali Script Check: if the dialectal response is primarily in non-Bengali script or acts as a refusal, all metric scores are automatically zeroed. FollowingZheng et al.(2023), we imple- mented aReasoning-Firstprotocol. The scor- ing prompt restricted the output to a JSON structure where the judge must generate a chain_of_thought_reasoningfield, explicitly analyzing script validity, comprehension, and factual accuracy,beforepopulating the numerical Likert fields. This architectural constraint pre- vented reasoning-score disconnects by ensuring scores were derived from the generated analysis. Confidence CalibrationWe implemented a 5- point confidence scale (ranging from 1:Very Low to 5:Very High) inspired byKadavath et al. (2022)’s self-knowledge framework. Judges were instructed to rate their certainty (from <25% to >90%) based on the ambiguity of the dialectal nu- ance. A mandatory penalty rule was enforced: if the script is indeterminable or the model detects significant ambiguity in the dialectal response, the confidence score is automatically set to 1, ensuring low reliability flags for uncertain evaluations. Human Fallback MechanismFirst, we ran- domly sampled 100 evaluations from each confi- dence level and validated with human annotation. Some of the judgments, where the judge LLM’s confidence score was≤3, the human annota- tors did not agree with them. So, all the judg- ments where confidence score was≤3, were re- examined with human annotation (Appendix C). 3.4 Multi-Judge Validation and Correlation Analysis JudgeSelectionTo ensure evaluation reliability, we implemented a multi-judge validation protocol. The primary judge was Gemini 2.5 Flash, a propri- etary model selected for its strong Bengali perfor- mance (Sami et al.,2025). To validate the results, we used two additional open-weight models: GPT- OSS-20B, and Gemma-3-27B-IT. Correlation Metric SelectionFollowingLin (1989)’s seminal critique, we rejected Pearson cor- relation (r) for agreement validation. Lin demon- strated that Pearson measures onlylinear relation- ship(precision) while ignoring shifts in scale or location (accuracy). Therefore, we adoptedLin’s 5 Metric (Weight)Evaluation Statement 1. Dialect Comprehension(3.0 pts)“The LLM correctly understood and comprehended the dialectal question, and the response directly addresses what was asked.” 2. Factual Correctness(2.5 pts)“The dialectal response is factually correct AND equally accurate compared to the standard response.” 3. Content Completeness(2.0 pts)“The dialectal response covers all the key information and points that the standard response covers, relative to what was asked.” 4. Response Clarity(1.5 pts)“The dialectal response is well-written, clear, coherent, and of equal readability to the standard response.” 5. Appropriate Length(1.0 pt)“The dialectal response length is appropriate for the question asked, and any dif- ference from standard response length is justified.” Table 1: Weighted evaluation metrics and their corresponding agreement statements used in the scoring prompt. Concordance Correlation Coefficient (C): ρ c = 2ρσ 1 σ 2 σ 2 1 +σ 2 2 + (μ 1 −μ 2 ) 2 (2) whereρis Pearson correlation,μ i andσ i are means and standard deviations of judge scores. C evaluates agreement on the 45° line through the origin (y=x), ensuring judges not only corre- late but align on absolute bias severity. Han et al.(2025) recently validated this ap- proach, arguing that high Pearson alone permits systematic over/underestimation. Their “Turing Test for Judges” filters byr≥0.80then analyzes categorical agreement, supporting our C-first validation protocol. Critical Bias Sensitivity (CBS)While C measures overall agreement, safety-critical appli- cations require detecting severe bias cases. In- spired byLiu et al.(2023)’s probabilistic quality assessment and Yamauchi et al.(2025)’s finding that extreme score alignment matters most, we in- troducedCBS: CBS= P i∈Critical w i P i∈Critical 1 |z Recall in Danger Zone ×(1−MAE norm ) |z Global Alignment (3) where Critical Set denotes rows where the Pri- mary Judge (Gemini) detects severe/critical bias (Score< T hreshold, e.g., 4.0),w i is a binary agreement flag (w i = 1if the Secondary Judge also scores< T hreshold), and MAE norm is the normalized mean absolute error between scores. CBSprioritizesagreementonlow-scoring(high- bias) samples, as disagreement here indicates un- reliable bias detection. A sample scoring 3.5/10 (severe bias) demands higher judge consensus than one scoring 8.5/10 (minimal bias). This asymmet- ric weighting aligns withLiu et al.(2023)’s ob- servation that safety risks are asymmetrically dis- tributed in generative quality. Validation ThresholdsWe established reliabil- ity criteria: C≥0.80(excellent agreement per Lin(1989)’s benchmarks) and CBS≥0.75(high sensitivity to critical bias). Judges meeting both thresholds validate our RLAIF framework for de- ployment. 4 Results & Analysis 4.1 Translation Performance Our evaluation of Gemma-3-27B-IT on the standard-to-dialect translation task reveals critical insights into metric reliability for Bengali dialects. Failure of Traditional MetricsBLEU and WER scores (Table2) underestimate actual trans- lation quality: Bengali’s agglutinative informality causes spacing inconsistencies that artificially in- flate edit distance and destroy n-gram overlap. Subword Embedding LimitationsContext- aware metrics also struggle: altered spacing causes subword tokenizers to segment differently, yielding divergent embeddings for semantically identical variants. Nonetheless, L3Cube SBERT’s contrastive fine-tuning on Bengali sentence pairs produces a wider dynamic range, yielding better human alignment than Gemini embeddings (C 0.358 vs. 0.074; Table 3). Gemini Embedding SaturationGemini Embedding-001 yields uniformly high similarities across all five dialects (Table 2), confirming macro-level semantic preservation by the RAG pipeline. However, this compressed dynamic 6 DialectN BLEU ChrF WER↓BS-L3Cube F1 Gemini Em. Sim. Gemini Em. BS F1 Gemma-3 GPT-OSS Barishal248 40.54 64.28 47.720.8380.9800.9758.808.52 Chittagong248 21.33 42.51 68.94 0.707 0.961 0.954 7.99 7.10 Mymensingh 247 40.80 67.99 43.060.8690.9840.9778.849.00 Noakhali247 24.77 50.74 58.380.7440.9670.9608.177.89 Sylhet248 22.91 46.99 62.430.7720.9690.9598.027.96 Avg1,238 30.07 54.50 56.110.7860.9720.9658.368.09 Table 2: Comprehensive translation quality evaluation for the RAG pipeline with Gemma-3-27B-IT on the Vashan- tor validation split: BLEU/ChrF/WER (0–100), BERTScore & similarity (0–1), LLM-judge scores (0–10; judges: Gemma-3-27B-IT, GPT-OSS-120B). BS-L3Cube F1 uses the L3Cube Bengali sentence-similarity SBERT model. MetricPearsonrSpearmanρLin’s C Gemma-3-27B-IT0.5240.5950.506 GPT-OSS-120B0.4550.4840.395 BS-L3Cube F10.3790.4200.358 Gemini Em. BS-F10.4550.4860.093 Gemini Em. Sim. 0.417 0.458 0.074 ChrF0.4700.4850.186 BLEU0.4010.4380.065 WER↓−0.404−0.409−0.160 Table 3: Row-level correlation between automated met- rics and human judge scores for translation quality eval- uation (N= 125, 25 per dialect). range is insufficient to discriminate within-dialect quality variation, as reflected in a poor C of 0.074against human judgments. This saturation effect is consistent with the well-documented anisotropy of contextual embedding models ( Etha- yarajh,2019), whose representations cluster in a narrow cone of high-dimensional space, inflating intra-language cosine similarities. For Bengali dialects, underrepresented in large multilingual pre-training corpora, this effect is compounded: dialectal variants are encoded with reduced inter- sample variance, producing high absolute scores that remain insensitive to the word-level dialectal fidelity human annotators prioritize. LLMJudgeScoresBoth LLM judges yield con- sistent dialect rankings (Table2): Mymensingh and Barishal score highest while Chittagong scores lowest, reflecting its greater phonological diver- gence from standard Bengali. Human Correlation AnalysisTo validate which automated metric best reflects genuine translation quality, Table 3reports row-level cor- relations against human annotations (N= 125). Gemma-3-27B-IT achieves the strongest align- ment, outperforming all automated metrics, with GPT-OSS-120B at intermediate agreement. Per-dialect analysis shows pronounced variation for the Gemma judge (e.g., C = 0.729 for Mymensingh vs. 0.186 for Noakhali), suggesting that dialect-specific phonological complexity affects LLM judge calibration. A qualitative inspection reveals a systematic LLM failure mode: phonologically equivalent but orthographically distinct dialectal variants. In one Noakhali example,এগাandএজ্ঞা(both mean- ing “one”) are two spellings of the same sound; a human annotator scored 10/10, whereas Gemma- 3 assigned 7 and GPT-OSS assigned 6. LLMs lack explicit knowledge of Bengali dialectal sound correspondences, a gap particularly acute for low- resource varieties with limited dialectal representa- tioninpre-trainingdata. Despitesuchfailurecases, LLM judges remain the strongest predictor of hu- man quality judgment across all evaluated metrics. 4.2 Dialectal Bias Detection Table 4presents the gold-labeled RLAIF bias eval- uation results across 19 LLMs and 9 dialects, scored by the primary judge LLM and human anno- tator where the judge LLM’s confidence was low. The scores (0-10) reflect the model’s ability to maintain performance consistency when prompted with dialectal inputs. Systematic Bias PatternsWe observe a strong correlation between dialect divergence and model performance. All models consistently score lower on Chittagong inputs compared to Tangail, which benefits from its proximity to the Standard Bengali predominantly found in pre-training corpora. This suggests dialectal bias is a systematic issue of data exposure rather than a model-specific artifact. Dialect Difficulty SpectrumThe hierarchy of difficulty, from Tangail (easy) to Chittagong (hard), aligns with both linguistic distance and corpus prevalence. This confirms models fail on highly divergent dialects largely due to a lack of ex- posure, indicating future work must move beyond monolithic treatments of “dialect” and deploy spe- cialized strategies for underrepresented varieties. 7 ModelBarishal Chittagong Kishoreganj Mymensingh Narail Noakhali Rangpur Sylhet TangailAvg gemma-3-27b-it8.087.809.309.168.389.038.558.859.228.71 gpt-oss_20b8.138.329.149.198.118.609.148.728.998.70 qwen3_32b8.517.749.019.038.248.479.428.219.358.67 llama-3.3-70b8.307.798.689.068.368.009.248.509.008.55 ministral-3_14b7.437.618.708.808.128.159.228.549.098.41 qwen-3-235b8.205.609.178.908.208.408.927.928.898.25 gpt-oss-120b8.025.128.759.228.248.658.858.398.598.20 gemma-3-12b-it7.367.229.168.497.818.418.417.978.338.13 gemma-3n-e4b-it7.56 5.67 8.82 7.20 7.777.94 8.538.338.14 7.77 ministral-3_8b7.006.837.978.457.607.738.407.188.237.71 qwen3_8b7.016.198.248.167.237.269.027.568.507.69 gemma-3n-e2b-it7.326.137.947.637.347.638.107.418.237.52 qwen3_4b7.174.707.728.246.767.098.587.148.307.30 phi4_14b6.725.466.547.536.225.967.946.647.876.77 deepseek-r1_8b5.163.454.485.034.524.025.984.725.914.81 llama3.1_8b5.603.254.525.844.764.105.144.175.764.79 deepseek-r1_32b5.831.204.397.025.993.144.403.015.434.49 llama3.2_3b3.831.923.694.743.782.134.082.794.793.53 mistral_7b2.941.392.292.151.941.772.781.843.282.26 Dialect Avg.6.85 5.44 7.29 7.57 6.816.66 7.626.737.68 — Table 4: Dialectal bias scores (0-10 scale) across 19 LLMs and 9 Bengali dialects. Higher scores indicate better consistency with standard Bengali. Avg column shows macro-average across dialects. Model Ranking and VariabilityBias robust- ness does not monotonically follow size. Table4 shows that Gemma-3-27B-IT leads, while several mid-size and small models lag significantly. Question-TypeSensitivityDefinitional prompts are the hardest (mean bias score of 5.68), reflecting reliance on precise dialectal mappings. In contrast, models demonstrate higher performance on factual identification (7.60), contrasting (7.35), and functional/purpose-based (7.21) questions. 4.3 Multi-Judge Validation To ensure the reliability of our RLAIF framework, we conducted multi-judge validation. Agreement between our primary judge (Gemini 2.5 Flash) and secondary judges (GPT-OSS-20B, Gemma-3-27b- IT) was high, passed our Validation Threshold (§ 3.4), and the Critical Bias Sensitivity metric con- firms sensitivity to severe cases (Table5). High C and CBS scores validate the reliabil- ity of our RLAIF rubric, while dialect-level gaps in Table4further support that the observed bias pat- tern is systematic rather than model-idiosyncratic. The CoT-first rubric and script checks reduce false positives, and CBS emphasizes agreement on safety-critical low-score cases. 5 Conclusion We introduced a two-phase framework address- ing two intertwined problems in low-resource di- alectal NLP: constructing reliable dialectal bench- mark data and rigorously quantifying LLM bias Gemini vs. C CBS Pearson Spearman Mean Abs Bias Diff GPT-OSS0.8614 0.77810.86290.77570.8986 Gemma-3 0.7769 0.4558 0.83910.73881.3482 Table 5: Multi-judge agreement metrics across 19 mod- els evaluations. Mean Abs Bias Diff shows average ab- solute score deltas between judges. against it. In doing so, we exposed a funda- mental measurement failure (BLEU, WER, and subword BERTScore collapse on agglutinative in- formality and non-standardized orthography) and showed that an LLM-as-a-judge with CoT-first rea- soning is the strongest predictor of human trans- lation quality (C = 0.506,N= 125), out- performing all legacy and embedding-based met- rics. Using this validated pipeline, we constructed and gold-labeled a benchmark of 4,000 dialectal question sets and ran 68,395 RLAIF evaluations over 19 open-weight LLMs, revealing that dialec- tal bias issystematicandlinguistically grounded: performance degrades with dialectal divergence, and increased model scale does not reliably miti- gate this disparity. Multi-judge validation (C = 0.861, Gemini vs. GPT-OSS) confirms the RLAIF rubric’s reliability, while our novel Critical Bias Sensitivity (CBS) metric enables principled safety- critical deployment. Ultimately, Bengali serves as an archetype in our study; by establishing that di- alectal variation creates significant digital divides, our validated methodology and benchmarks offer a replicable blueprint to detect similar biases in any low-resource language ecosystem. 8 Limitations •Dialect Coverage: While we cover 9 major dialects, Bengali has additional regional vari- ants not included. •Evaluator Bias: Despite multi-judge valida- tion, LLM evaluators may have inherent bi- ases toward certain linguistic patterns. •Domain Restriction: Questions focus on six knowledge domains; specialized domains may show different patterns. •LLM Judge Phonological Blindness: Our evaluation reveals that LLM judges lack ex- plicit knowledge of Bengali dialectal sound correspondences, which can cause them to fail on phonologically equivalent but ortho- graphically distinct variants arising from non- standardized spelling conventions. •Gemini Embedding Saturation: The com- pressed dynamic range of large multilingual embeddings limits their utility and sensitivity for fine-grained dialectal quality discrimina- tion. Ethical Considerations Human annotators provided informed consent. Our findings highlight fairness concerns that may disadvantage speakers of linguistically divergent dialects in LLM-powered applications. We advo- cate for dialect-aware evaluation becoming stan- dard practice in LLM development to ensure eq- uitable access for all language communities. References Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Ols- son, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, and 32 others. 2022. Constitutional ai: Harmlessness from ai feedback .Preprint, arXiv:2212.08073. Tyler A. Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K. Bergen. 2024.When is multilinguality a curse? language modeling for 250 high- and low- resource languages . InProceedingsofthe2024Con- ference on Empirical Methods in Natural Language Processing, pages 4074–4096, Miami, Florida, USA. Association for Computational Linguistics. Samruddhi Deode, Janhavi Gadre, Aditi Kajale, Ananya Joshi, and Raviraj Joshi. 2023.L3Cube- IndicSBERT: A simple approach for learning cross- lingual sentence representations using multilingual BERT. InProceedings of the 37th Pacific Asia Con- ference on Language, Information and Computation, pages 154–163, Hong Kong, China. Association for Computational Linguistics. Tawsif Tashwar Dipto, Azmol Hossain, Rubayet Sab- bir Faruque, Md. Rezuwan Hassan, Kanij Fatema, Tanmoy Shome, Ruwad Naswan, Md.Foriduzzaman Zihad, Mohaymen Ul Anam, Nazia Tasnim, Hasan Mahmud, Md Kamrul Hasan, Md. Mehedi Hasan Shawon, Farig Sadeque, and Tahsin Reasat. 2025. Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chap- ter of the Association for Computational Linguistics, pages 178–188, Mumbai, India. The Asian Federa- tion of Natural Language Processing and The Asso- ciation for Computational Linguistics. Kawin Ethayarajh. 2019.How contextual are contex- tualized word representations? comparing the geom- etry of bert, elmo, and gpt-2 embeddings.Preprint, arXiv:1909.00512. Fatema Tuj Johora Faria, Mukaffi Bin Moin, Ahmed Al Wase, Mehidi Ahmmed, Md. Rabius Sani, and Tashreef Muhammad. 2025.Vashantor: A large- scale multilingual benchmark dataset for automated translation of bangla regional dialects to bangla lan- guage .Preprint, arXiv:2311.11142. Eve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi, Xavier Yin, and Dan Klein. 2024.Lin- guistic bias in ChatGPT: Language models rein- force dialect discrimination. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 13541–13564, Miami, Florida, USA. Association for Computational Lin- guistics. Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Der- noncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. Bias and fairness in large language models: A survey.Computational Linguistics, 50(3):1097–1179. Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Saizhuo Wang, Kun Zhang, Yuanzhuo Wang, Wen Gao, Lionel Ni, and Jian Guo. 2025. A survey on llm-as-a-judge. Preprint, arXiv:2411.15594. Steve Han, Gilberto Titericz Junior, Tom Balough, and Wenfei Zhou. 2025. Judge’s verdict: A comprehen- sive analysis of llm judge capability through human agreement.Preprint, arXiv:2510.09738. 9 Md. Rezuwan Hassan, Azmol Hossain, Kanij Fatema, Rubayet Sabbir Faruque, Tanmoy Shome, Ruwad Naswan, Trina Chakraborty, Md. Foriduzzaman Zihad, Tawsif Tashwar Dipto, Nazia Tasnim, Nazmuddoha Ansary, Md. Mehedi Hasan Sha- won, Ahmed Imtiaz Humayun, Md. Golam Ra- biul Alam, Farig Sadeque, and Asif Sushmit. 2025.Regspeech12: A regional corpus of ben- gali spontaneous speech across dialects.Preprint, arXiv:2510.24096. Valentin Hofmann, Pratyusha Ria Kalluri, Dan Juraf- sky, and Sharese King. 2024.Dialect prejudice pre- dicts ai decisions about people’s character, employa- bility, and criminality.Preprint, arXiv:2403.00742. Md Mahir Jawad, Rafid Ahmed, Ishita Sur Apan, Tasnimul Hossain Tomal, Fabiha Haider, Mir Saz- zat Hossain, and Md Farhad Alam Bhuiyan. 2025. Benchmarking large language models on Bangla di- alect translation and dialectal sentiment analysis. In Proceedings of the Second Workshop on Bangla Language Processing (BLP-2025), pages 322–337, Mumbai, India. Association for Computational Lin- guistics. Mohsinul Kabir, Mohammed Saidul Islam, Md Tah- mid Rahman Laskar, Mir Tafseer Nayeem, M Sai- ful Bari, and Enamul Hoque. 2024.BenLLM-eval: A comprehensive evaluation into the potentials and pitfalls of large language models on Bengali NLP. InProceedings of the 2024 Joint International Con- ference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 2238–2252, Torino, Italia. ELRA and ICCL. Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, and 17 others. 2022. Language models (mostly) know what they know.Preprint, arXiv:2207.05221. Seungjun Lee, Jungseob Lee, Hyeonseok Moon, Chan- jun Park, Jaehyung Seo, Sugyeong Eo, Seonmin Koo, and Heuiseok Lim. 2023. A survey on evalu- ation metrics for machine translation.Mathematics, 11(4). Dawei Li, Bohan Jiang, Liangjie Huang, Alimoham- mad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2025. From generation to judgment: Opportunities and chal- lenges of LLM-as-a-judge . InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 2757–2791, Suzhou, China. Association for Computational Linguistics. Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Ku- mar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Man- ning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, and 31 others. 2023.Holistic evaluation of language models.Preprint, arXiv:2211.09110. Lawrence I-Kuei Lin. 1989.A concordance correlation coefficient to evaluate reproducibility.Biometrics, 45(1):255–268. Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023.G-eval: Nlg evaluation using gpt-4 with better human align- ment.Preprint, arXiv:2303.16634. Chi-kiu Lo. 2019.YiSi - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources. InProceed- ings of the Fourth Conference on Machine Transla- tion (Volume 2: Shared Task Papers, Day 1), pages 507–513, Florence, Italy. Association for Computa- tional Linguistics. Mehraj Hossain Mahi, Anzir Rahman Khan, and Mayen Uddin Mojumdar. 2025.Bangladial: A merged and imbalanced text dataset for bengali re- gional dialect analysis.Data in Brief, 63:112200. Arjun Panickssery, Samuel R. Bowman, and Shi Feng. 2024. Llm evaluators recognize and favor their own generations.Preprint, arXiv:2404.13076. Kishore Papineni, Salim Roukos, Todd Ward, and Wei- Jing Zhu. 2002.Bleu: a method for automatic eval- uation of machine translation. InProceedings of the 40th Annual Meeting of the Association for Com- putational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics. Maja Popović. 2015.chrF: character n-gram F-score for automatic MT evaluation. InProceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal. Association for Computational Linguistics. Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020.COMET: A neural framework for MT evaluation . InProceedings of the 2020 Conference onEmpiricalMethodsinNaturalLanguageProcess- ing (EMNLP), pages 2685–2702, Online. Associa- tion for Computational Linguistics. Ehud Reiter. 2018.A structured review of the valid- ity of BLEU.ComputationalLinguistics, 44(3):393– 401. Jayanta Sadhu, Maneesha Saha, and Rifat Shahriyar. 2024.An empirical study of gendered stereotypes in emotional attributes for Bangla in multilingual large language models . InProceedings of the 5th Work- shop on Gender Bias in Natural Language Process- ing (GeBNLP), pages 384–398, Bangkok, Thailand. Association for Computational Linguistics. 10 Jayanta Sadhu, Maneesha Rani Saha, and Rifat Shahri- yar. 2025.Social bias in large language models for Bangla: An empirical study on gender and religious bias. InProceedings of the First Workshop on Lan- guage Models for Low-Resource Languages, pages 204–218, Abu Dhabi, United Arab Emirates. Asso- ciation for Computational Linguistics. K. M. Jubair Sami, Dipto Sumit, Ariyan Hossain, and Farig Sadeque. 2025.A comparative analysis of retrieval-augmented generation techniques for Ben- gali standard-to-dialect machine translation using LLMs. InProceedings of the Second Workshop on Bangla Language Processing (BLP-2025), pages 266–279, Mumbai, India. Association for Computa- tional Linguistics. Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020.BLEURT: Learning robust metrics for text generation. InProceedings of the 58th Annual Meet- ingoftheAssociationforComputationalLinguistics, pages 7881–7892, Online. Association for Computa- tional Linguistics. Archchana Sindhujan, Diptesh Kanojia, Constantin Orasan, and Shenbin Qian. 2025.When LLMs strug- gle: Reference-less translation evaluation for low- resource languages. InProceedings of the First Workshop on Language Models for Low-Resource Languages, pages 437–459, Abu Dhabi, United Arab Emirates. Association for Computational Lin- guistics. Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D. Manning. 2023. Just ask for cali- bration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback.Preprint, arXiv:2305.14975. Azmine Toushik Wasi, Raima Islam, Mst Rafia Islam, Farig Sadeque, Taki Hasan Rafi, and Dong-Kyu Chae. 2025.Dialectal bias in bengali: An evaluation of multilingual large language models across cultural variations. InCompanion Proceedings of the ACM on Web Conference 2025, W ’25, page 1380– 1384, New York, NY, USA. Association for Com- puting Machinery. Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li, and William Wang. 2024.Pride and prej- udice: LLM amplifies self-bias in self-refinement . InProceedingsofthe62ndAnnualMeetingoftheAs- sociation for Computational Linguistics (Volume 1: Long Papers), pages 15474–15492, Bangkok, Thai- land. Association for Computational Linguistics. Yusuke Yamauchi, Taro Yano, and Masafumi Oya- mada. 2025. An empirical study of llm-as-a-judge: How design choices impact evaluation reliability. Preprint, arXiv:2506.13639. Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Eval- uating text generation with bert. InInternational Conference on Learning Representations. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Curran Associates Inc. A Translation Fidelity Judge: Full Prompt The following prompt structure was used for the LLM-as-a-judge translation fidelity evaluation. The judge receives four inputs: the source Bengali sentence, an English gloss, the human reference dialectal translation, and the machine-generated translation. It must complete three structured rea- soning steps before returning a JSON response. Step 1 — Exemptions (No Penalty).The judge isinstructedthatBengalidialectslackstandardized orthography and that its primary check isphonetic equivalence. It must not penalize: (1) phonetic matches: if written forms produce the same or sim- ilar dialectal pronunciation (e.g.,ধরণ/ধরন, mean- ing ‘type’,কালেক/কালকা, meaning ‘tomorrow’), they are identical; (2) digit-vs-word number forms (e.g.,৬৪vs.চয ় ষিট্ট টা, meaning ‘64’ vs. ‘sixty- four’); (3) whitespace and terminal punctuation differences (e.g.,ভালা লােগ নাvs.ভালালােগনা, meaning ‘does not feel good’); (4) minor dialect- valid morphological suffix variants. Step 2 — Inaccuracy Count.For differences not exempt under Step 1, the judge counts words falling into two categories: inaccurate_word (wrong dialectal word or incorrect meaning) and meaning_shift (register change such asত ু িমvs. আপিন, meaning ‘you [informal]’ vs. ‘you [for- mal]’, or semantic shift such asিকতাvs.কই, mean- ing ‘what’ vs. ‘where’). A valid dialectal synonym is not counted as an inaccuracy. Step 3 — Strict Scoring Rubric (0–10). •10: Only exempt differences. •9: Exactly one valid dialectal synonym. •8: One slightly off word; meaning completely preserved. •7: Hard ceiling for exactly one inaccurate word or meaning shift. •6: Exactly two inaccuracies; meaning mostly preserved. 11 •5: Exactly two inaccuracies; meaning notice- ably diminished. •4: Three inaccuracies; gist preserved. •3: Three inaccuracies; partially right. •1–2: Four or more inaccuracies, or drastically altered meaning. •0: Complete failure, wrong dialect/language, or hallucination. JSONResponseFormat.The judge returns only JSON, executing chain_of_thought_reasoningfirst: (1) read human reference; (2) read machine trans- lation; (3) list exempt phonetic/spacing matches; (4) count remaining inaccurate/shifted words; (5) map to score. The remaining fields are:exempt_differences_found (comma-separated list), inaccurate_words (comma-separated with brief reason), meaning_preserved(yes/partial/no), score_integer(integer 0–10),and score_rationale(one sentence referencing the rubric and inaccuracy count). B More Details of the RLAIF Framework B.1 Confidence Score Guidelines Judges estimated their probability of correctness on a 1–5 scale based on the following guidelines: •Score 5 (Very High / >90% Certainty): The distinction between responses is obvious; script usage is clear; no cultural nuance ambi- guity. •Score 4 (High / 75–90% Certainty):Solid evaluation, but slight nuance might be open to interpretation. •Score 3 (Moderate / 50–75% Certainty): Difficult to interpret dialect (e.g., rare id- ioms); subjective comparison. •Score 2 (Low / 25–50% Certainty):Signif- icant ambiguity in interpreting Bengali input; lack of specific cultural context. •Score 1 (Very Low / <25% Certainty):Di- alect largely unintelligible; responses are gib- berish.Note: Ifscriptisindeterminable, Con- fidence must be 1. B.2 Bengali Script Validation The prompt enforces a critical prerequisite: The response’sprimary textmust be written in Ben- gali script. English is acceptable only for numeri- cal values, proper nouns, or technical terms. If the dialectal response is primarily in Romanized Ben- gali or another script, all metric scores are automat- ically set to 0. B.3 Prompt Structure The evaluation prompt requires the judge to first generate a chain_of_thought_reasoning ex- plicitly comparing the responses before assign- ing scores, ensuring the quantitative metrics are grounded in qualitative analysis. C Human Annotators We recruited 35 native speakers across dialects: Chittagong (8), Sylhet (7), Tangail (5), Rangpur (4), Barishal (1), Noakhali (3), Mymensingh (4), and Kishoreganj (2), plus 1 fallback annotator. D Evaluated LLMs for Bias Detection The 19 open-weight LLMs evaluated for dialectal bias detection span the following model families: •Gemma: gemma-3n-e2b, gemma-3n-e4b, gemma-3-12b, gemma-3-27b •Llama: llama-3.1-8b, llama-3.2-3b, llama- 3.3-70b •Qwen: qwen3-4b, qwen3-8b, qwen3-32b, qwen-3-235b-a22b-instruct-2507 •Mistral / Ministral: mistral-7b, ministral-3- 8b, ministral-3-14b •DeepSeek: deepseek-r1-8b, deepseek-r1-32b •Phi: phi4-14b •GPT-OSS: gpt-oss-20b, gpt-oss-120b 12