Paper deep dive
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
Samar M. Magdy, Fakhraddin Alwajih, Abdellah El Mekki, Wesam El-Sayed, Muhammad Abdul-Mageed
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 99%
Last extracted: 6/21/2026, 11:35:03 AM
Summary
The paper introduces LQM (Linguistically Motivated Multidimensional Quality Metrics), a hierarchical error taxonomy designed to improve Machine Translation (MT) evaluation for diglossic and dialectal languages, specifically targeting Arabic dialects. Unlike the standard MQM framework, which often uses broad categories like 'Mistranslation' to capture heterogeneous errors, LQM provides a six-level linguistic hierarchy: Sociolinguistics, Pragmatics, Semantics, Morphosyntax, Orthography, and Graphetics. The authors constructed a bidirectional parallel corpus of 3,850 sentences across seven Arabic dialects (Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni) and evaluated six LLMs (including Gemini and Command models) using expert span-level human annotation. The results demonstrate that LQM provides more granular diagnostic capabilities for identifying failures in variety choice, register, and cultural appropriateness.
Entities (12)
Relation Signals (8)
LQM â containslevel â Graphetics
confidence 100% · LQM... spans six linguistically grounded levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics.
LQM â containslevel â Sociolinguistics
confidence 100% · LQM... spans six linguistically grounded levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics.
LQM â containslevel â Pragmatics
confidence 100% · LQM... spans six linguistically grounded levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics.
LQM â containslevel â Semantics
confidence 100% · LQM... spans six linguistically grounded levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics.
LQM â containslevel â Morphosyntax
confidence 100% · LQM... spans six linguistically grounded levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics.
LQM â containslevel â Orthography
confidence 100% · LQM... spans six linguistically grounded levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics.
Gemini 2.5 Pro â evaluatedby â LQM
confidence 100% · We evaluate six LLMs... including Gemini-2.5-Pro... using LQM.
LQM â â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Existing MT evaluation frameworks, including automatic metrics and human evaluation schemes such as Multidimensional Quality Metrics (MQM), are largely language-agnostic. However, they often fail to capture dialect- and culture-specific errors in diglossic languages (e.g., Arabic), where translation failures stem from mismatches in language variety, content coverage, and pragmatic appropriateness rather than surface form this http URL introduce LQM: Linguistically Motivated Multidimensional Quality Metrics for MT. LQM is a hierarchical error taxonomy for diagnosing MT errors through six linguistically grounded levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics (Figure 1). We construct a bidirectional parallel corpus of 3,850 sentences (550 per variety) spanning seven Arabic dialects (Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni), derived from conversational, culturally rich content. We evaluate six LLMs in a zero-shot setting and conduct expert span-level human annotation using LQM, producing 6,113 labeled error spans across 3,495 unique erroneous sentences, along with severity-weighted quality scores. We complement this analysis with an automatic metric (spBLEU). Though validated here on Arabic, LQM is a language-agnostic framework designed to be easily applied to or adapted for other languages. LQM annotated errors data, prompts, and annotation guidelines are publicly available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2604.18490v1
- Canonical: https://arxiv.org/abs/2604.18490v1
Trouble viewing inline? Open PDF directly â
Full Text
87,766 characters extracted from source content.
Expand or collapse full text
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation Samar M. Magdy λ Fakhraddin Alwajih λ * Abdellah El Mekki λ * Wesam El Sayed Ο Muhammad Abdul-Mageed λ,Îł λ The University of British Columbia Ο Minia University Îł Canada Research Chair in NLP and ML samar.ahmad, muhammad.mageed@ubc.ca SociolinguisticsGraphetics Pragmatics SemanticsMorphosyntax Orthography LQM La mine Ă©tait trĂšs sombre. The mine was very dark. The expression was very gloomy. SRC TRG MT Polysemy Failure The meeting had already ended. äŒèźźæŁćšç»æă äŒèźźæŁćšç»æă SRC TRG MT Mandarin GrammarâVerbal Features I am thirty years old. T engo treinta años. Tengo treinta años. SRC TRG MT Character EncodingSpanish I like drinking coffee. ì ë 컀íŒë„Œ ë§ìë êČì ìąìí©ë ë€ ì ë 컀íŒë„Œ ë§ìë êČì ìąìí©ë ë€ SRC TRG MT Typo Korean Komm mal her. Could you come here for a sec? Come here. SRC TRG MT Illocutionary ForceGerman French Where will I meet my sister Ruqaya? I am busy this period. ŰŁÙۧ ۱ÙÙŰ©Ű ŰźŰȘÙ Ù Űč ŰŁÙۧ ŰșÙŰȘÙۧÙÙ ÙÙÙ .ۧÙÙŰ§Ù Ű§ŰȘ Ùۧۯ Ù ŰŽŰșÙÙ Ùۧۯ Ù ŰŽŰșÙÙ ŰŁÙۧ ۱ÙÙŰ©Ű ÙÙÙۧ ŰșŰ§ŰŻÙ ÙÙÙ .ۧÙÙŰȘ۱۩ SRC TRG MT Code & Register SelectionMoroccan Figure 1: Cross-lingual examples illustrating the proposed LQM frameworkâs linguistic levels, demonstrating its language-agnostic design and broad applicability beyond Arabic. Abstract Existing MT evaluation frameworks, includ- ing automatic metrics and human evaluation schemes such as Multidimensional Quality Metrics (MQM), are largely language-agnostic. However, they often fail to capture dialect- and culture-specific errors in diglossic languages (e.g., Arabic), where translation failures stem from mismatches in language variety, con- tent coverage, and pragmatic appropriateness rather than surface form alone. We introduce LQM: Linguistically Motivated Multidimen- sional Quality Metrics for MT. LQM is a hi- erarchical error taxonomy for diagnosing MT errors through six linguistically grounded lev- els: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics * Equal contribution (Figure 1). We construct a bidirectional par- allel corpus of3,850sentences (550per vari- ety) spanning seven Arabic dialects (Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni), derived from con- versational, culturally rich content. We eval- uate six LLMs in a zero-shot setting and con- duct expert span-level human annotation us- ing LQM, producing 6,113 labeled error spans across 3,495 unique erroneous sentences, along with severity-weighted quality scores. We com- plement this analysis with an automatic metric (spBLEU). Though validated here on Arabic, LQM is a language-agnostic framework de- signed to be easily applied to or adapted for other languages. LQM annotated errors data, prompts, and annotation guidelines are pub- licly available athttps://github.com/ UBC-NLP/LQM_MT. 1 arXiv:2604.18490v1 [cs.CL] 20 Apr 2026 1 Introduction Evaluating MT remains challenging, particularly when systems must preserve sociolinguistic and pragmatic constraints in addition to semantic con- tent. While automatic metrics such as BLEU (Pap- ineni et al., 2002), COMET (Rei et al., 2020), and chrF++ (Popovi Ì c, 2017) provide rapid, scalable feedback, they primarily capture surface overlap or embedding similarity and can diverge from hu- man judgments in cases where translation quality depends on variety choice, register, or cultural ap- propriateness (Yao et al., 2024; Chao, 2025). This has motivated increased use of human-in-the-loop evaluation to characterize failure modes (Brewster et al., 2025). Among fine-grained human evaluation frame- works, MQM (Lommel et al., 2014) has become a widely adopted standard, offering a flexible error taxonomy for identifying translation issues. How- ever, in complex diglossic languages such as Ara- bic (Ferguson, 1959; Bassiouney, 2020), we find that MQM-style categorizations can under-specify dialect- and culture-conditioned errors. In particu- lar, when errors reflect mismatches in dialect, reg- ister, or pragmatic intent, surface-oriented error labels may not provide enough structure to consis- tently localize the linguistic level at which a model fails. In this work, we analyze the limitations that arise when applying MQM to bidirectional MT involv- ing Arabic dialects, and highlight three recurring challenges. (i) MQM provides an operational in- ventory of error types, but it does not explicitly index errors to linguistic levels, which can make it difficult to separate similar surface manifestations with different underlying explanations. (i) The taxonomy offers limited built-in support for the systematic treatment of variation, including dialect choice, register, and culturally conditioned prag- matic constraints, which are central in dialectal Arabic. (i) As a result, MQM annotations can be harder to translate into targeted corrective signals for improving dialectal Arabic MT, since distinct phenomena (e.g., pragmatic infelicity vs. seman- tic reference errors) may be grouped under broad categories. To address these gaps, we introduce Linguisti- cally Motivated Multidimensional Quality Metrics, dubbed LQM, a linguistically grounded taxonomy designed to diagnose MT errors in a way that is aligned with linguistic theory and practical anno- tation. We develop LQM through a two-pronged process aimed at balancing theoretical coverage with empirical adequacy: Top-down: We applied MQM to a pilot subset of our MT data and an- alyzed where existing categories and guidelines were insufficient to consistently capture dialectal and sociopragmatic phenomena. Bottom-up: We then performed iterative, data-driven refinement over observed errors (span-level), consolidating recurring patterns into categories organized by lin- guistic level. We synthesize these perspectives into a hierarchical framework spanning Sociolinguis- tics, Pragmatics, Semantics, Morphosyntax, Or- thography, and Graphetics. Our contributions are: (i) LQM framework: We propose LQM, a linguisti- cally grounded taxonomy that explicitly separates sociolinguistic and pragmatic phenomena from se- mantic and form-level errors, enabling more tar- geted diagnosis in diglossic and dialectal settings. The hierarchy supports two complementary anno- tation settings: a lightweight version for assigning broad error categories and a diagnostic version for labeling specific error types. (i) Parallel dataset: We introduce a conversational, culturally rich paral- lel corpus covering seven Arabic varieties to stress- test models under dialect-sensitive conditions. (i) Multi-model evaluation: We evaluate six LLMs with expert span-level human annotations (includ- ing severity-weighted scores) and complement this analysis with standard automatic metrics. 2 Related works MQM has become a widely used framework for diagnostic MT evaluation and error taxonomy anal- ysis (Lommel et al., 2014; Freitag et al., 2021). Recent work has pushed MQM-style evaluation toward finer granularity by localizing errors at the span level, including metric-based approaches such as xCOMET (Guerreiro et al., 2023) and prompted LLM evaluators such as GEMBA-MQM (Kocmi and Federmann, 2023; Fernandes et al., 2023). Re- lated directions include agentic or multi-step eval- uators and refinement pipelines (He et al., 2024; Wang et al., 2025a), as well as efforts to improve re- liability by filtering or validating annotated spans via post-editing signals (Lu et al., 2025; Kocmi et al., 2024; Kreutzer et al., 2020). In parallel, research in Automatic Post-Editing (APE) has evolved from classical formulations (Simard et al., 2007) to LLM-assisted pipelines (Bhattacharyya et al., 2023; Raunak et al., 2023). 2 Recent resources and protocols increasingly em- phasize human-centered corrections and explain- ability, providing structured annotations or ratio- nales that can support targeted diagnosis (Wasti et al., 2025; Jung et al., 2024; Alves et al., 2024). Large Reasoning Models (LRMs) further enable multi-step explanations for ambiguity resolution (Liu et al., 2025), motivating work that constrains or structures LLM-based MQM annotation, for ex- ample, via compressed taxonomies (ThinMQM) or tagged span annotation (Zhan et al., 2025; Yeom et al., 2025a; He et al., 2025; Wang et al., 2025b; Feng et al., 2025). Despite these advances, applying general- purpose MQM-style frameworks to languages with diverse varieties such as Arabic remains challeng- ing due to diglossia and dialectal variation (Fer- guson, 1959; Bassiouney, 2020), as well as non- standardized orthographies across dialects (Habash et al., 2012). Dialectal MT research has lever- aged resources such as Alexandria (Mekki et al., 2026) and MADAR (Bouamor et al., 2018), has increasingly studied LLM-based systems in spe- cific dialectal settings (Yakhni and Chehab, 2025; Fernandes et al., 2023). Yet, morphological and syntactic variations continue to complicate error identification and interpretation (Zbib et al., 2012; Sajjad et al., 2020). Existing benchmarks such as Tarjamat (Kadaoui et al., 2023) and NADI 2024 (Abdul-Mageed et al., 2024) underscore persistent gaps for low-resource dialects, but they do not provide a linguistically leveled, dialect-sensitive diagnostic taxonomy for isolating how and where translations lose dialectal identity or pragmatic ap- propriateness. These limitations motivate LQM, which orga- nizes MQM-style error annotation by linguistic level (from sociolinguistics and pragmatics to form- level phenomena), enabling a more consistent di- agnosis in dialectal Arabic settings. 3 Pitfalls of MQM MQM (Lommel et al., 2014) is widely used for human diagnostic evaluation and is intentionally designed to be broadly applicable across languages. In our setting, bidirectional MT involving multiple Arabic dialects, we find that MQMâs generic cate- gories can under-specify error patterns that are pri- marily driven by variety choice and sociopragmatic constraints rather than surface form alone (Abdul- Nabi et al., 2024). MQM CategoryMQM SubcategoryCountRate (%) AccuracyMistranslation3,67360.09 AccuracyAddition4537.41 AccuracyMissing3575.84 FluencyGrammar3255.32 StyleUnidiomatic style3225.27 StyleAwkward style2924.78 TerminologyWrong term1923.14 AccuracyUndertranslation1302.13 FluencySpelling1252.04 AccuracyOvertranslation661.08 AccuracyUntranslated570.93 TerminologyInconsistent with term resource470.77 StyleLanguage register310.51 StyleInconsistent style150.25 Locale conv.Currency format90.15 Fluency/Ling. conv.Punctuation80.13 FluencyInconsistency/unintelligible50.08 FluencyCharacter encoding40.07 Locale conventionNumber format1< 0.02 TerminologyInconsistent use of terminology1< 0.02 Total6,113100.00 Table 1: MQM error counts and rates aggregated over annotated spans. TheMistranslationrow is highlighted to indicate its frequent use as a catch-all label in our setting. Setup. We applied MQM to translations from six Arabic-aware LLMs on a bidirectional trans- lation dataset covering English and seven Arabic dialects. Six annotatorsâtwo senior linguists and four trained annotatorsâannotated a stratified sam- ple of 50 sentences per model and translation di- rection. Each dialect was evaluated by annotators who are native speakers of that dialect. Full experi- mental details are provided in Section 5. Findings.Table 1 reports aggregated MQM span annotations. A prominent pattern is the dominance of the Mistranslation label (60.09% of annotated errors). Based on qualitative inspection and anno- tator feedback, Mistranslation frequently functions as a catch-all label for heterogeneous phenomena that do not fit cleanly under other MQM tags, re- ducing diagnostic specificity regarding where the model fails. Annotators reported that many such cases re- flect three recurring sources of error: (i) Pragmatic mismatches: illocutionary force, discourse mark- ers, vocatives, and honorifics. (i) Variety consis- tency: defaulting to MSA or drifting into a differ- ent dialect. (i) Idiomatic usage: mistranslation of proverbs and other fixed expressions. 4 LQM Framework Motivated by the MQM analysis above, we intro- duce LQM, a linguistically grounded taxonomy 3 7 dialects Reference Translation Casablanca Human Translation LQM Guidelines + Training + Platform MT with Errors (EnglishâDialects) Human Annotation Using LQM MT Labeled Errors (Fine-Grained Categorized) 6 Models Figure 2: Data and annotation workflow. Casablanca dataset of seven Arabic dialects is translated by humans and evaluated by six LLM models; creating LQM guidelines and training support human annotation to produce fine-grained MT error labels. for span-level human MT evaluation. LQM is de- signed to improve diagnostic precision in diglossic and dialectal settings by distinguishing sociolin- guistic and pragmatic failures from semantic and form-level errors. Specifically, LQM organizes errors into six linguistic levels: Sociolinguistics, Pragmatics, Semantics, Morphosyntax, Orthogra- phy, and Graphetics. We summarize these levels below; full definitions and additional examples are provided in Appendix §A. (i) Sociolinguistics. We place sociolinguistics at the top of the hierarchy to reflect communicative competence (Hymes et al., 1972): a translation may be grammatically well-formed yet inappropriate for the social context. This is especially consequen- tial in Arabic, where diglossia makes variety choice functional rather than merely stylistic (Ferguson, 1959). LQM therefore introduces code & regis- ter selection errors with three subcategories: (a) standardization interference (vertical mismatch), (b) wrong dialect (horizontal mismatch), and (c) register mismatch (tone/formality). (i) Pragmatics. While sociolinguistics targets broader social norms, the pragmatics level cap- tures failures of communicative intent and implied meaning (Levinson, 1983), i.e., mismatches be- tween sentence meaning and speaker meaning. We group these under use, context, and cultural ap- propriateness and include targeted subcategories: (a) speech acts/illocutionary force, (b) code switch- ing, (c) MWEs/proverbs, (d) discourse marker mis- match, and (e) vocatives/honorifics /titles (Farwell and Helmreich, 1999). (i) Semantics. This level evaluates mean- ing transfer, and preservation of propositional content (Cruse, 1986).We distinguish: (a) lexical semantics (word meaning and lexical relations), (b) propositional semantics (truth- conditional content; avoiding unintended addi- tions/omissions) (Soames, 1987), and (c) discourse semantics (cohesion and reference across sen- tences) (Halliday and Hasan, 2014; Kamp and Reyle, 2013).Example subcategories include named entity, wrong term, polysemy failure, and cross-variety interference; see Appendix§A for complete definitions. (iv) Morphosyntax. This level captures viola- tions of target-language structural constraints at the morphology-syntax interface (Radford, 2004). We separate: (a) grammar (e.g., agreement and inflectional features) (Corbett, 2006), including verbal features (tense, aspect, voice, mood, per- son) (Palmer, 2001) and nominal features (num- ber, gender, case, definiteness, state); and (b) con- stituent order (Greenberg et al., 1963), including locale-sensitive reordering such as address format and date format. (v) Orthography/Writing Conventions. This level evaluates written-form conventions and me- chanical correctness (Derwing, 1992). Since di- alectal Arabic lacks fully standardized orthogra- phy, we accept dialectal spellings recognized by native annotators as conventional in informal digi- tal contexts (e.g., texting and social media). 1 We define five error types: (a) spelling (including ty- 1 Given limited standardized dialect orthographies, we rely on annotator judgments for acceptability within emerg- ing, socially shared informal conventions. 4 pos/slips), (b) inconsistent spelling, (c) unconven- tional Spelling, (d) surface mechanics (number, currency, time, telephone formats), and (e) punctu- ation. (vi) Graphetics. The lowest level captures fail- ures in the technical realization of the text code. We include character encoding errors, where the output is garbled due to encoding/decoding issues. A comprehensive breakdown of LQM, includ- ing both the lightweight and diagnostic versions and their subcategories, is provided in Table A.1 (Appendix). We also report an external validation of LQM conducted by two linguists specializing in linguistics and translation studies in Appendix §B. 5 Experimental Setup We conduct a case study to evaluate LQM in a stress-test setting for dialectal Arabic MT with Arabic-aware LLMs. Our experimental design mir- rors the MQM analysis in Section 3, but replaces MQM with the LQM annotation protocol described in Section 4. Below, we describe the dataset, the evaluated models, and our human and automatic evaluation procedures. Figure 2 summarizes the data construction and LQM-based annotation work- flow used in our experiments. 5.1 Dataset As part of our contribution, we construct a new bidi- rectional parallel corpus of3,850sentences based on the manually transcribed conversational speech data introduced by Talafha et al. (2024). The cor- pus covers seven Arabic varieties: Egyptian, Emi- rati (UAE), Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni. The dialectal transcripts were then professionally translated into MSA with support from native speakers of each dialect and subsequently translated into English. The dataset is balanced with550sentences per variety.It is organized for bidirec- tional MT: DialectâEnglish (DAâEN) and EnglishâDialect (ENâDA). We report DAâEN results for all seven dialects. While for ENâDA, the Jordanian, and Yemeni dialects were excluded due to the lack of available native-speaker transla- tors for target-side validation. 5.2 Evaluated LLMs We select six Arabic-aware LLMs spanning both closed and open-weight systems.For closed- source models, we evaluateGemini-2.5-Pro andGemini-2.5-Flash(Teametal., 2023).Foropen-weightmodels,we evaluateFanar-9B(Team et al., 2025), Gemma-27B(Team et al., 2024),Command-A (111B)(CohereLabs,2024),and Command-R7B (7B). All models are evaluated in a zero-shot setting using a fixed prompt (Appendix§C). 5.3 Human Labeling via LQM Annotation guidelines. LQM relies on expert human annotation to yield both (i) fine-grained error diagnostics and (i) severity-weighted qual- ity scores. We developed annotation guidelines that define LQM categories and provide Arabic examples to support consistent span selection and labeling. Comprehensive annotation guidelines are available in the previously mentioned repository. Annotation scope. Annotators applied LQM to a random sample of50translations per direc- tion, drawn from the full set of outputs generated by the six models across dialects and directions. Across this annotated sample, annotators identi- fied3,495unique translations containing at least one error and labeled6,113error spans. 2 Each labeled error includes: (i) span boundaries, (i) LQM category/subcategory/sub-subcategory, (i) severity (minor/major/critical), and (iv) an optional free-text explanation. Appendix §D describes the annotation process, annotator profiles, and quality assurance workflow. Figure 3 presents represen- tative examples from our dataset across different dialects, highlighting the error spans for each level. 5.4 Evaluation Metrics 5.4.1 LQM score derivation Our LQM score equation is inspired by the MQM original scoring equation (Lommel et al., 2014), we follow the same severity weightss i : minor=1, major=5, and critical=25. Then, we fix a weight of 1across all the error types. For a translation with lengthL(in words) and annotated errorsiwith severity weights s i , we compute: LQM_Score = max 0, 100 1â P i s i L This yields a normalized score in[0, 100], where higher values indicate better translation quality. 2 A sentence may receive multiple error labels, and the full set of3,850sentence pairs also includes outputs judged error-free. 5 Sociolinguistics Standardization MAJOR ENGâ EGY GM SRCOh my God, youâve gone so high, youâre around sixteen floors high by now. ,â ÌÂw yq , ©w Âwf rV  , Ìh A§ ............... . ÌtÂw C€ rK TtF Pragmatics Speech Acts MAJOR MORâ ENG C SRC ÂA Ahyl ¹l ÂCAb Âwy TbyÂĄ ÂyÂĄr ÌF Ahyl Âw` Cdq € wW ,â Mr. Brahim is a respectable man today, thank God, he was excellent and you can count on him. Semantics Named Entity MAJOR JORâ ENG C SRC € Ly wy HyAb` wA§ yl yl Âdn Ìlbs ÂąlyOfl ,âOkay, okay, Mr. Grumpy , bro, what time is it for the group? Count it for me. Morphosyntactic Verbal Feature MAJOR MAUâ ENG F SRC Âslsm YlÂ Ì r§ ÂÂz Âyl wf§ Ìhy ÌK à € ,â Donât let this time passyou by without watching the series and that thing. Orthography Typo MINOR ENGâ UAE G SRCOMG! Mira, your father who worked hard for you and raised you. ,â ,Âry !Âąl A§w .AC€ yl ` Ìl Graphetics Character Encoding MAJOR JORâ ENG C SRCbO ÂwyÂ Ì Âr A A ÌnÂd Ew ,âFE€believe me, I only found out this morning. Figure 3: LQM-annotated examples across the six linguistic levels. Severity and fine-grained error types are indicated in each example. Error spans are highlighted in red. Icons indicate model families: G Gemini, C Command, F Fanar, GM Gemma. LQM scores are based on human judgments (span selection and severity assignment) rather than au- tomatic matching metrics. 5.4.2 Automatic surface metrics We report automatic scores using spBLEU (Goyal et al., 2022), a sentence-level variant of BLEU (Pa- pineni et al., 2002) to compute the correlation be- tween LQM human annotated errors and automatic scores. We do not report model-based metrics such as COMET (Rei et al., 2020), since their perfor- mance and calibration may be less reliable for di- alectal Arabic varieties that are underrepresented in common training and evaluation resources. 6 Results of Error Analysis 6.1 LQM Error Counts and Rates Table 2 summarizes aggregated error counts and rates under LQM. Compared to MQM, which fre- quently assigned heterogeneous phenomena to the broad Mistranslation label, LQM reallocates these cases into linguistically interpretable categories. In particular, many instances previously collapsed under the MQM Mistranslation are separated into (i) sociolinguistic failures (e.g., code/register se- lection), (i) pragmatic infelicities (e.g., discourse markers, vocatives, MWEs/proverbs), and (i) se- mantic failures (e.g., named entities, polysemy, lexical coverage), enabling a more targeted di- agnosis of where models break down. The pri- mary limitations of current LLMs in dialectal ArabicâEnglish translation appear to be linguis- tic rather than merely computational. We there- fore provide additional linguistic analysis in Ap- pendix§E. 6.2 Severity-weighted LQM Scores Table 3 reports severity-weighted LQM scores (0â100) across models and directions. Overall, Gemini-2.5-Proachieves the strongest perfor- mance, ranking first in9of the12evaluated direc- tions. Performance is generally higher and more stable in DAâEN than in ENâDA, consistent with the additional constraint in ENâDA of main- taining the target dialectâs sociolinguistic identity. Across directions, ENâUAE yields the highest scores for most models, whereas ENâMOR is the most challenging setting, with most models scoring below40(with theGeminivariants as notable ex- ceptions). At the model level,Command-R7B exhibits the weakest overall performance, with notably low scores on ENâUAE (21.94) and ENâMOR (17.00).Fanar-9Bshows substan- tial variance across dialects, performing strongly on ENâUAE (78.32) but dropping sharply on ENâMOR (6.46). Finally, several other mod- els lead in specific directions (e.g., Command- A on JORâEN,Gemini-2.5-Flashon 6 LQM CategoryLQM SubcategoryLQM SubsubcategoryCountRate (%) sociolinguisticscode & register selectionstandardization interference (vertical mismatch)90414.8 semanticslexical semanticsnamed entity5729.4 sociolinguisticscode & register selectionwrong dialect (horizontal mismatch)5398.8 semanticslexical semanticscoverage: unknown term/dialect3656.0 semanticspropositional semanticsomission3575.8 semanticslexical semanticsunnatural/ unidiomatic style3225.3 morphosyntaxgrammarverbal features2954.8 semanticslexical semanticsawkward style2924.8 pragmaticsuse, context, cultural appropriatenessmwes, proverbs & metaphors2894.7 semanticsdiscourse semanticspronouns2684.4 semanticspropositional semanticsaddition2634.3 semanticslexical semanticsdisambiguation: cross-variety interference2464.0 semanticslexical semanticswrong term1923.1 semanticspropositional semanticshallucination1903.1 semanticslexical semanticsundertranslation1302.1 orthography/ writing conventionsspellingtypo / slip1252.0 semanticslexical semanticsdisambiguation: polysemy failure1202.0 semanticslexical semanticstransliteration1151.9 pragmaticsuse, context, cultural appropriatenessspeech acts mismatch1041.7 pragmaticsuse, context, cultural appropriatenessforms of address (vocatives/honorifics/titles)971.6 semanticslexical semanticsovertranslated661.1 semanticslexical semanticsuntranslated570.9 semanticsdiscourse semanticsinconsistent with terminology resource470.8 pragmaticsuse, context, cultural appropriatenessdiscourse marker mismatch350.6 sociolinguisticscode & register selectionregister mismatch310.5 morphosyntaxgrammarnominal features300.5 pragmaticsuse, context, cultural appropriatenesscode switching190.3 semanticsdiscourse semanticsinconsistent style150.2 orthography/ writing conventionssurface mechanicscurrency format90.1 orthography/ writing conventionspunctuationâ80.1 semanticslexical semanticsunintelligible5<0.1 grapheticscharacter encodingâ4<0.1 orthography/ writing conventionssurface mechanicsnumber format1<0.1 semanticsdiscourse semanticsinconsistent use of terminology1<0.1 Total6,113100.0 Table 2: LQM error counts and rates utilizing the LQM Diagnostic layer for fine-grained analysis. MORâEN, andFanar-9Bon ENâMAU), sug- gesting that relative strengths depend on both di- rection and dialect. 6.3 Correlations Between Automatic and Human Evaluation Table 4 reports spBLEU scores per direction and model. We compute the correlation between sp- BLEU and human-derived LQM scores to quan- tify the agreement between surface-based metrics and expert judgments. We observe a weak pos- itive association (Pearsonr = 0.289, Spearman Ï = 0.322,p < 0.001), indicating that spBLEU captures only a limited portion of the variance in severity-weighted human assessments. This gap is expected in our setting, where major quality degradations often stem from sociolinguistic and pragmatic phenomena (e.g., dialect/register mis- matches, vocatives/honorifics, discourse markers) that are not well-modeled by n-gram overlap. In Appendix F, we provide additional experiments ex- amining the robustness of LQM to sentence length. 6.4 Inter-Annotator Agreement (IAA) We report IAA using three complementary types of scores: span-detection scores, label-agreement scores on overlapping spans, and chance-corrected agreement measured with Cohenâs kappa (Îș). The analysis was conducted on 377 doubly annotated items from both translation directions. Follow- ing span-based MT evaluation work (Yeom et al., 2025b), we use overlap-based span agreement as the primary detection metric and exact span F1 as a stricter secondary metric. Agreement was strong under overlap-based matching (character-level F1 = 0.760; overlap span F1 = 0.821) but lower un- der exact span matching (F1 = 0.440), suggesting that annotators usually identified the same error regions while differing in precise span boundaries. On overlapping spans, agreement was highest for coarse error category (F1 = 0.662;Îș= 0.681), followed by severity (F1 = 0.630;Îș= 0.427) 7 DirectionFanar-9BCommand-ACommand-R7BGemma-27bGemini-2.5-flashGemini-2.5-pro ENGLISHâ DIALECT ENGâEGY49.9071.1435.8954.4465.2372.31 ENGâMAU 45.2641.8737.4019.6238.2840.90 ENGâMOR6.4638.6017.0019.2951.5368.60 ENGâPAL38.9659.8943.2951.9166.56 67.14 ENGâUAE78.3272.8921.9463.3982.1083.07 DIALECTâ ENGLISH EGYâENG53.5760.8845.5668.6974.4575.89 JORâENG65.27 73.2658.2269.7672.2766.19 MAUâENG40.7656.3843.9061.4559.2663.88 MORâENG40.9864.4551.5062.34 72.1570.32 PALâENG64.7873.1862.6272.8973.4279.47 UAEâENG43.8566.6145.0854.0662.65 67.09 YEMâENG61.2862.9058.0767.0070.7673.41 Table 3: LQM scores (severity-weighted, 0â100) by model and direction. Higher is better. Open-Weight ModelsProprietary Models DirectionFanar-9BCommand-ACommand-R7BGemma-27BGemini-2.5-FlashGemini-2.5-Pro Englishâ Dialect ENGâ EGY13.9723.3813.1519.8424.1126.09 ENGâ MAU1.421.752.614.244.34 5.88 ENGâ MOR2.6410.186.979.2815.3918.30 ENGâ PAL9.6515.8412.2016.4821.0723.26 ENGâ UAE5.6713.215.8911.1717.7819.77 Dialectâ English EGYâ ENG28.8931.6226.7027.5432.1831.47 JORâ ENG29.1931.7826.5628.9932.2231.88 MAUâ ENG10.1312.828.9611.1916.03 16.59 MORâ ENG16.9923.1917.6419.27 24.0023.34 PALâ ENG25.6731.1723.4725.8730.5627.12 UAEâ ENG20.9127.1819.8323.36 27.8926.90 YEMâ ENG22.5423.9820.1722.0925.9726.26 Table 4: spBLEU scores across translation directions and model families. Higher is better. The best result in each row is highlighted in green and boldfaced. and category+severity (F1 = 0.517;Îș= 0.509). Fine-grained error types were more variable (F1 = 0.484), and the strictest criterionâjoint agreement on span, error type, and severityâyielded F1 = 0.388. Overall, the results indicate reliable error detection and coarse categorization, with greater variability in fine-grained labeling. 6.5 Analysis of Error Distributions and Model Attribution In Table 5, we analyze the human-annotated cor- pus along two axes: (i) model-wise contribution to the total error mass, and (i) the distribution of error types by translation direction. Figure A.1 (Appendix) summarizes model-wise error contri- butions and presents the overall distribution across LQM categories and fine-grained subcategories. Model-wise error contribution. In DAâEN, Command-RandFanarcontribute the largest share of errors, withCommand-Rpeaking in Egyptian (26.7%) andFanarpeaking in Mo- roccan (23.3%) and Emirati (23.8%). In con- trast, Gemini-2.5-Pro consistently contributes the smallest error mass across dialects (e.g.,8.5% in JORâENG). In ENâDA,Fanar,Command-R, andGemmadominate the error mass. Peaks are observed forCommand-Rin ENâUAE (34.4%), Fanarin ENâMOR (33.0%), andGemmain ENâMAU (22.7%). Gemini-2.5-Pro continues to contribute substantially less, reaching as low as 2.1% of the error mass in ENâMOR and7.0% in 8 Direction PART I: Model Error Contribution (%) (Who Failed?) PART I: Error Type Distribution (%) (Why?) Cmd-A Cmd-R Fanar Gemma FlashProSocPrag Sem Morph OrthGra Dialect â English EGYâ ENG8.026.726.614.812.211.82.424.367.75.40.2â JORâ ENG12.326.28.925.518.68.50.414.882.61.60.40.2 MAUâ ENG18.118.120.019.514.010.30.73.492.33.40.2â MORâ ENG15.521.6 23.319.28.412.0â10.485.13.90.6â PALâ ENG15.023.318.418.815.09.5â10.987.01.40.6â UAEâ ENG13.423.823.814.913.011.2â10.884.54.10.40.2 YEMâ ENG16.318.818.613.517.415.40.711.984.42.80.2â English â Dialect ENGâ EGY19.730.816.814.411.37.040.17.239.98.44.30.2 ENGâ MAU12.516.218.322.716.713.781.54.911.81.9â ENGâ MOR14.420.433.023.96.22.139.60.635.919.24.7â ENGâ PAL15.922.524.718.39.59.158.52.420.56.012.40.2 ENGâ UAE15.134.417.313.810.09.370.64.515.45.04.5â Table 5: Diagnostic dashboard of dialectal failures. Model error contribution (Part I) and error-type dis- tribution (Part I) by translation direction. In DAâEN (top), Semantic errors dominate across dialects. In ENâDA (bottom), Sociolinguistic errors dominate overall, especially for Mauritanian, Palestinian, and Emi- rati, while Egyptian and Moroccan show a more mixed pattern, with Semantic errors remaining prominent and Morphosyntax also notable in Moroccan. (Soc=Sociolinguistics, Prag=Pragmatics, Sem=Semantics, Morph=Morphosyntax, Orth=Orthography, Gra=Graphetics). ENâEGY. Error typology by direction. We observe a direction-dependent shift in the dominant failure mode. In DAâEN, errors are primarily seman- tic: semantic categories account for67.7% (Egyp- tian) up to 92.3% (Mauritanian) of the error mass, with particularly high concentrations in Maghrebi dialects (Mauritanian92.3%, Moroccan85.1%). Pragmatic errors form a notable secondary cluster in dialects such as Egyptian (24.3%) and Jorda- nian (14.8%), indicating difficulties with commu- nicative intent even when the literal meaning is partially recovered. Morphosyntactic and ortho- graphic errors are comparatively rare in DAâEN (often†5.4%), consistent with target-side normal- ization when generating English. In ENâDA, the distribution shifts significantly toward sociolinguistic failures. This shift is most pronounced in ENâMAU, where sociolinguistic errors account for81.5% of the error mass. Similar patterns are observed in ENâUAE (70.6%) and ENâPAL (58.5%). However, in Egyptian and Mo- roccan, semantic errors remain a high secondary failure mode (39.9% and35.9%, respectively). We also observe increases in morphosyntactic and or- thographic errors in generation (e.g., morphosyn- tax up to19.2% for Moroccan; orthography12.4% for Palestinian). Appendix G provides a more de- tailed analysis of the fine-grained error distribu- tions per dialect. 7 Conclusion We presented LQM, a linguistically grounded framework for MT evaluation that organizes errors by linguistic level, enabling diagnostic analysis in diglossic and dialect-rich settings. Crucially, while our empirical evaluation centers on Arabic, the un- derlying taxonomy is inherently language-agnostic and readily adaptable to other linguistic contexts. In a case study of seven Arabic dialects, we ob- served a direction-dependent shift in failure modes: DAâEN is largely driven by semantic breakdowns (67.7â92.3%), reflecting challenges in lexical cov- erage and meaning transfer from dialectal input, whereas ENâDA is dominated by sociolinguistic failures, with systems often defaulting to dialect- faithful varieties instead of maintaining the target dialect. Improving dialectal MT requires optimiz- ing both semantic adequacy and sociolinguistic fidelity. We also found that spBLEU aligns only weakly with severity-weighted LQM scores (Pearsonr = 0.289, SpearmanÏ = 0.322), consistent with n-gram overlapâs insensitivity to pragmatic and dialect-identity errors emphasized by human anno- tation. 9 Acknowledgments We acknowledge support from Canada Research Chairs (CRC), the Natural Sciences and Engi- neering Research Council of Canada (NSERC; RGPIN-2018-04267), the Social Sciences and Humanities Research Council of Canada (SSHRC; 895-2020-1004), the Canadian Foundation for Innovation (CFI; 37771), the Digital Research Alliance of Canada, 3 and UBC ARC-Sockeye. 4 We thank Aisha Alraeesi for annotating the Emi- rati Arabic data translated from English and Na- jla Hassan for annotating the Palestinian Arabic data translated from English. We are also grate- ful to Yayhay Mohamed Elhaj and Sidi Ebidi for annotating the English-to-Mauritanian direction, to Alcides Alcoba for support with the annota- tion platform, and to Abderahim Elmadany for his feedback. Finally, we thank the external linguists Ranada Hassan and Saudi Sadiq for independently validating the LQM framework. 8 Limitations Our study has several limitations. âąDialect and direction coverage: Although our study covers seven Arabic varieties over- all, we report ENâDA evaluation for only five dialectsâEgyptian, Emirati, Maurita- nian, Moroccan, and Palestinian. We exclude Jordanian and Yemeni because we were un- able to secure enough qualified native-speaker translators to validate outputs in those target dialects. Therefore, our conclusions about dialect preservation in the ENâDA setting apply only to the five evaluated varieties. âą Domain specificity: The corpus is derived from transcribed TV dialog, which is conver- sational and culturally grounded. Although this setting is well-suited for eliciting dialect- and pragmatics-related errors, results may dif- fer in more formal or specialized domains (e.g., legal, medical, or technical translation) with distinct terminology and register con- straints. âą Non-standardized orthography: Arabic di- alect writing lacks fully standardized conven- tions. We accept spellings judged by native 3 https://alliancecan.ca 4 https://arc.ubc.ca/ubc-arc-sockeye annotators to be common in informal digi- tal contexts, but orthographic variation can complicate consistent span-level localization and make automatic evaluation less straight- forward. âą Prompting and inference settings: All mod- els are evaluated in a zero-shot setting with a fixed prompt. We do not systematically study the effects of alternative prompting strategies (e.g., few-shot exemplars, constrained output formats) or inference-time controls, which may change both overall quality and the dis- tribution of LQM error types. Ethical Considerations This work studies MT quality for Arabic dialects and introduces LQM, a linguistically grounded framework for span-level error annotation. Be- cause dialectal data can reflect speakersâ regional and social identities, we minimize the risk of sensi- tive attribute inference by (i) reporting results at the dialect/variety level rather than attempting to infer or annotate personal attributes such as gender, age, socioeconomic status, or education, and (i) restrict- ing annotations to translation errors and linguistic phenomena observable in text (e.g., code/register selection, pragmatic appropriateness, semantics, and form-level issues). Our corpus is derived from previously released material and contains conversational content; we acknowledge that such data may include cultur- ally specific expressions or potentially sensitive topics. Annotators were instructed to focus on translation quality rather than judge speakers or communities, and to provide brief explanations only when needed for clarity. We also recognize that resources for dialectal MT can be misused to generate targeted or stereotyped content; to miti- gate this, we provide documentation emphasizing appropriate use and the limitations of automatic metrics for dialect identity and pragmatics, and we avoid presenting LQM as a tool for profiling individuals. References Muhammad Abdul-Mageed, Amr Keleg, AbdelRahim Elmadany, Chiyu Zhang, Injy Hamed, Walid Magdy, Houda Bouamor, and Nizar Habash. 2024. NADI 2024: The fifth nuanced Arabic dialect identifica- tion shared task. In Proceedings of the Second Ara- bic Natural Language Processing Conference, pages 10 709â728, Bangkok, Thailand. Association for Com- putational Linguistics. Razan Abdul-Nabi, Rasha Obeidat, and Anas Bsoul. 2024. A survey on machine translation of low- resource arabic dialects. In 2024 15th International Conference on Information and Communication Sys- tems (ICICS), pages 1â6. Duarte M Alves, JosĂ© P Pombal, Nuno M Guerreiro, and AndrĂ© FT Martins. 2024. xtower: Multilingual trans- lation error explanation and correction with large language models. arXiv preprint arXiv:2406.19482. John Langshaw Austin. 1975. How to do things with words. Harvard university press. Reem Bassiouney. 2020.Arabic sociolinguistics: Topics in diglossia, gender, identity, and politics. Georgetown University Press. Pushpak Bhattacharyya, Rajen Chatterjee, Markus Fre- itag, Diptesh Kanojia, Matteo Negri, and Marco Turchi. 2023. Findings of the wmt 2023 shared task on automatic post-editing. In Conference on Machine Translation. Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexan- der Erdmann, and 1 others. 2018. The madar arabic dialect corpus and lexicon. In Proceedings of the eleventh international conference on language re- sources and evaluation (LREC 2018). Ryan CL Brewster, Gabriel Tse, Angela L Fan, Marwa Elborki, Maiah Newell, Priscilla Gonzalez, Amitra Hoq, Crystal Chang, Maksud Chowdhury, Adiba Geeti, and 1 others. 2025. Evaluating human-in- the-loop strategies for artificial intelligence-enabled translation of patient discharge instructions: a multidisciplinary analysis. NPJ digital medicine, 8(1):629. Dunstan Brown, Marina Chumakina, and Greville G Corbett. 2013. Canonical morphology and syntax. Oxford University Press. Penelope Brown and Stephen C Levinson. 1987. Polite- ness: Some universals in language usage, volume 4. Cambridge university press. Yang Chao. 2025. Natural language processing and deep learning in cross-cultural language acquisition: From machine translation to cultural context under- standing. Theoretical and Natural Science, 92:13â 18. CohereLabs. 2024. Command R+: A 104b parameter open-weight model for rag and tool use. Greville G Corbett. 2006. Introduction: Canonical agreement. In Agreement, pages 1â34. Cambridge University Press. D Alan Cruse. 1986. Lexical semantics. Cambridge university press. Bruce L Derwing. 1992.Orthographic aspects of linguistic competence. The linguistics of literacy, 21:193â211. David Farwell and Stephen Helmreich. 1999. Prag- matics and translation. Procesamiento del lenguaje natural, nÂș 24 (mayo 1999); p. 18-39. Zhaopeng Feng, Shaosheng Cao, Jiahan Ren, Jiayuan Su, Ruizhe Chen, Yan Zhang, Zhe Xu, Yao Hu, Jian Wu, and Zuozhu Liu. 2025. Mt-r1-zero: Advanc- ing llm-based machine translation via r1-zero-like reinforcement learning. ArXiv, abs/2504.10160. Charles A Ferguson. 1959. Diglossia. word, 15(2):325â 340. Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, AndrĂ© F. T. Martins, Graham Neubig, Ankush Garg, J. Clark, Markus Freitag, and Orhan Firat. 2023. The devil is in the errors: Leverag- ing large language models for fine-grained machine translation evaluation. In Conference on Machine Translation. Bruce Fraser. 1999. What are discourse markers? Jour- nal of pragmatics, 31(7):931â952. Bruce Fraser. 2009. An account of discourse markers. International review of Pragmatics, 1(2):293â320. Markus Freitag, George F. Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021. Experts, errors, and context: A large-scale study of human evaluation for machine translation. Transactions of the Association for Computational Linguistics, 9:1460â1474. Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng- Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Kr- ishnan, MarcâAurelio Ranzato, Francisco GuzmĂĄn, and Angela Fan. 2022. The Flores-101 evaluation benchmark for low-resource and multilingual ma- chine translation. Transactions of the Association for Computational Linguistics, 10:522â538. Joseph H Greenberg and 1 others. 1963. Some uni- versals of grammar with particular reference to the order of meaningful elements. Universals of lan- guage, 2:73â113. Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, LuĂsa Coheur, Pierre Colombo, and AndrĂ© Martins. 2023. xcomet: Transparent machine translation evaluation through fine-grained error detection. Transactions of the Association for Computational Linguistics, 12:979â995. John J Gumperz. 1982. Discourse strategies. 1. Cam- bridge University Press. Nizar Habash, Mona T Diab, and Owen Rambow. 2012. Conventional orthography for dialectal arabic. In LREC, pages 711â718. Michael AK Halliday. 1978. Language as social semi- otic. London Arnold. 11 Michael Alexander Kirkwood Halliday and Ruqaiya Hasan. 2014. Cohesion in english. Routledge. Minggui He, Yilun Liu, Shimin Tao, Yuanchang Luo, Hongyong Zeng, Chang Su, Li Zhang, Hongxia Ma, Daimeng Wei, Weibin Meng, Hao Yang, Boxing Chen, and Osamu Yoshie. 2025. R1-t1: Fully incen- tivizing translation capability in llms via reasoning learning. ArXiv, abs/2502.19735. Zhiwei He, Xing Wang, Wenxiang Jiao, Zhuosheng Zhang, Rui Wang, Shuming Shi, and Zhaopeng Tu. 2024. Improving machine translation with human feedback: An exploration of quality estimation as a reward model. ArXiv, abs/2401.12873. Dell Hymes, JB Pride, and Janet Holmes. 1972. On communicative competence. sociolinguistics. Eds. Pride, JB y J. Holmes, pages 269â293. Dahyun Jung, Sugyeong Eo, Chanjun Park, and Heui- Seok Lim. 2024. Explainable ced: A dataset for explainable critical error detection in machine trans- lation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies (Volume 4: Student Research Workshop), pages 25â35. Karima Kadaoui, Samar M. Magdy, Abdul Waheed, Md Tawkat Islam Khondaker, Ahmed Oumar El- Shangiti, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023. TARJAMAT: Evaluation of bard and ChatGPT on machine translation of ten Arabic varieties. In Proceedings of ArabicNLP 2023, pages 52â75, Singapore (Hybrid). Association for Computational Linguistics. Hans Kamp and Uwe Reyle. 2013. From discourse to logic: Introduction to modeltheoretic semantics of natural language, formal logic and discourse rep- resentation theory, volume 42. Springer Science & Business Media. Tom Kocmi and Christian Federmann. 2023. Gemba- mqm: Detecting translation quality error spans with gpt-4. ArXiv, abs/2310.13988. Tom Kocmi, VilĂ©m Zouhar, Eleftherios Avramidis, Roman Grundkiewicz, Marzena Karpinska, Maja Popoviâc, Mrinmaya Sachan, and Mariya Shmatova. 2024. Error span annotation: A balanced approach for human evaluation of machine translation. ArXiv, abs/2406.11580. Julia Kreutzer, Nathaniel Berger, and Stefan Riezler. 2020. Correct me if you can: Learning from error corrections and markings. In European Association for Machine Translation Conferences/Workshops. Stephen C Levinson. 1983. Pragmatics. Cambridge university press. Sinuo Liu, Chenyang Lyu, Minghao Wu, Longyue Wang, Weihua Luo, Kaifu Zhang, and Zifu Shang. 2025.New trends for modern machine translation with large reasoning models.ArXiv, abs/2503.10351. Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014. Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics. In Proceedings of the Ninth Inter- national Conference on Language Resources and Evaluation (LRECâ14). Qingyu Lu, Liang Ding, Kanjian Zhang, Jinxia Zhang, and Dacheng Tao. 2025. Mqm-ape: toward high- quality error annotation predictors with automatic post-editing in llm translation evaluators. In Pro- ceedings of the 31st International Conference on Computational Linguistics, pages 5570â5587. Elin McCready. 2019. The semantics and pragmat- ics of honorification: Register and social meaning, volume 11. Oxford University Press. Abdellah El Mekki, Samar M Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-hibiri, Razan Saadie, Hamzah Alsayadi, Nadia Ghezaiel Hammouda, and 1 others. 2026. Alexandria: A multi-domain dialectal arabic ma- chine translation dataset for culturally inclusive and linguistically diverse llms.arXiv preprint arXiv:2601.13099. Geoffrey Nunberg, Ivan A Sag, and Thomas Wasow. 1994. Idioms. Language, 70(3):491â538. Frank Robert Palmer. 2001. Mood and modality. Cam- bridge university press. Kishore Papineni, Salim Roukos, Todd Ward, and Wei- Jing Zhu. 2002. Bleu: a method for automatic evalu- ation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computa- tional Linguistics, ACL â02, page 311â318, USA. Association for Computational Linguistics. Maja Popovi Ì c. 2017. chrF++: words helping character n-grams. In Proceedings of the Second Conference on Machine Translation, pages 612â618, Copen- hagen, Denmark. Association for Computational Lin- guistics. Andrew Radford. 2004. English syntax: An introduc- tion. Cambridge University Press. Vikas Raunak, Amr Sharaf, Hany Hassan Awadallah, and Arul Menezes. 2023. Leveraging gpt-4 for au- tomatic translation post-editing. In Conference on Empirical Methods in Natural Language Processing. Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Process- ing (EMNLP), pages 2685â2702, Online. Associa- tion for Computational Linguistics. 12 Hassan Sajjad, Ahmed Abdelali, Nadir Durrani, and Fahim Dalvi. 2020.AraBench: Benchmarking dialectal Arabic-English machine translation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5094â5107, Barcelona, Spain (Online). International Committee on Computational Linguistics. Deborah Schiffrin. 1987. Discourse markers. 5. Cam- bridge University Press. John R Searle. 1969. Speech acts: An essay in the philosophy of language. Cambridge university press. Michel Simard, Cyril Goutte, and Pierre Isabelle. 2007. Statistical phrase-based post-editing. In North Amer- ican Chapter of the Association for Computational Linguistics. Scott Soames. 1987. Direct reference, propositional attitudes, and semantic content. Philosophical topics, 15(1):47â87. Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Mohamedou Cheikh Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdelrahman Mohamed, Ab- dellah El Mekki, El Moatez Billah Nagoudi, Benel- hadj Djelloul Mama Saadia, Hamzah A. Alsayadi, Walid Al-Dhabyani, and 8 others. 2024. Casablanca: Data and models for multidialectal Arabic speech recognition. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Pro- cessing, pages 21745â21758, Miami, Florida, USA. Association for Computational Linguistics. Fanar Team, Ummar Abbas, Mohammad Shahmeer Ah- mad, Firoj Alam, Enes Altinisik, Ehsannedin Asgari, Yazan Boshmaf, Sabri Boughorbel, Sanjay Chawla, Shammur Chowdhury, and 1 others. 2025. Fanar: An arabic-centric multimodal generative ai platform. arXiv preprint arXiv:2501.13944. Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean- Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Mil- lican, and 1 others. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805. Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupati- raju, LĂ©onard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre RamĂ©, and 1 others. 2024. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118. Jiaan Wang, Fandong Meng, Yunlong Liang, and Jie Zhou. 2025a. Drt: Deep reasoning translation via long chain-of-thought. In Findings of the Associa- tion for Computational Linguistics: ACL 2025, pages 6770â6782. Jiaan Wang, Fandong Meng, and Jie Zhou. 2025b. Deep reasoning translation via reinforcement learn- ing. arXiv preprint arXiv:2504.10187. Adnan Wasti, Matthew Lee, Tausifa Alam, Sreyasi Ghosh, and Marine Carpuat. 2025. Translationcor- rect: A human-centered post-editing framework for error-aware machine translation. In Proceedings of ACL 2025 (to appear). Malak Yakhni and Jeanine Chehab. 2025. Fine-tuning arabic llms for lebanese dialect translation and eval- uation. arXiv preprint arXiv:2405.12534. Binwei Yao, Ming Jiang, Tara Bobinac, Diyi Yang, and Junjie Hu. 2024. Benchmarking machine translation with cultural awareness. In Findings of the Associa- tion for Computational Linguistics: EMNLP 2024, pages 13078â13096, Miami, Florida, USA. Associa- tion for Computational Linguistics. Taemin Yeom, Yonghyun Ryu, Yoonjung Choi, and Jinyeong Bak. 2025a. Tagged span annotation for detecting translation errors in reasoning llms. Pro- ceedings of the Tenth Conference on Machine Trans- lation. Taemin Yeom, Yonghyun Ryu, Yoonjung Choi, and JinYeong Bak. 2025b. Tagged span annotation for detecting translation errors in reasoning llms. In Proceedings of the Tenth Conference on Machine Translation, pages 878â886. Rabih Zbib, Spyros Matsoukas, Richard Schwartz, John Makhoul, Enrique Jimenez, and Chad Malarkey. 2012. Machine translation of arabic dialects. In Proceedings of the 2012 Conference of the North American Chapter of the Association for Computa- tional Linguistics (NAACL). Runzhe Zhan, Zhihong Huang, Xinyi Yang, Lidia S. Chao, Min Yang, and Derek F. Wong. 2025. Are large reasoning models good translation evalua- tors?analysis and performance boost.ArXiv, abs/2510.20780. 13 Appendices We offer an additional structure as follows: âą LQM Framework §A âą LQM External Validation §B âą Prompts §C âą Data Annotation §D âą Linguistic Insights on LLMs Performance §E âą Robustness of LQM to Sentence Length §F âą LQM Fine-Grained Error Distribution §G A LQM Framework Structured along the hierarchy of linguistic analysis, the LQM framework targets six distinct layers: sociolinguistic, pragmatic, semantic, morphosyntactic, orthographic/writing conven- tions, and graphetic. (i) Sociolinguistic: LQM introduces the code & register selection type of error under the soci- olinguistic level, explicitly penalizing three main subcategories: (a) standardization interference (vertical mismatch), where the model reverts to use the standardized, "high," or prestige variety (e.g., Standard German, MSA) when a specific vernacular or "low" variety is requested. For example, the use of MSA appearing in Emirati data as Ìn d`t ÂwÂ(âGet up, stay away from meâ), where Ìn d`tis an MSA expression rather than the Emirati one. (b) wrong dialect (horizontal mismatch): Where the model uses features specific to a different regional or social variety than the target (e.g., using Mexican Spanish slang in a translation for Spain, or Egyptian idioms in a Levantine text. For example, in the Jordanian sentenceAnÂĄ A rFA . bF wyl ¹y dy¥€ ,Âr Â rtÂ(âJasser came here more than once, and this has a million reasons.â), the model mixes Jordanian with other dialects, using the Lebanese formdyÂĄand the Gulf formA. (Halliday, 1978)âs conceptualization of language as a âsocial semiotic,â wherein register is defined as the specific language variety linked to the field, tenor, and mode of a situation. Accordingly, we include (c) register mismatch (tone/formality) as a distinct error type to ensure translations are evaluated not just on propositional content, but on their adherence to the tenor of discourse (e.g., formal vs. informal), a critical dimension often lost in NMT outputs that default to a neutral register (e.g., using "Tu" instead of "Vous" in French, or casual slang in a legal document). (i) Pragmatics: While sociolinguistics governs the broad social norms of interaction, the prag- matic level addresses failures in communicative intent and implied meaning, as well as the gap between sentence meaning and speaker meaning. We classify errors at this level under the error types of use, context, and cultural appropriateness, capturing instances where the translation is grammatically valid but functionally or culturally anomalous (Farwell and Helmreich, 1999). In high-context languages like Arabic, literal translation often yields semantically accurate but pragmatically hollow outputs. To capture these nuances, LQM introduces five targeted subcategories: (a) speech acts/illocutionary force: We evaluate whether the translation preserves the illocutionary force of the source based on Speech Act Theory (Austin, 1975; Searle, 1969). NMT models can flatten the pragmatic force of an utterance (e.g., converting a polite request into a direct command). It covers errors in interpreting speech acts (e.g., requests, threats, advice, jokes). For instance, the phraseÂąl ÂșAJ literally means âGod willingâ, but its pragmatic function can vary significantly to mean âeven ifâ, âmaybeâ, or âhopefullyâ depending on use. (b) code switching:This category identifies failures where the model either ignores embedded foreign lexical items or forces them into Arabic morphology inappropriately, signaling a failure to recognize the shift in linguistic code.We categorize CS under pragmatics because lan- guage alternation serves specific communicative functions (Gumperz, 1982), often includes both situational and metaphorical switching. CS serves as a communicative strategy, linking it to social meaning and interactional goals. for example, the model translated âI want you to bring me twelve bananasâ into Arabic asTÂAn21 Ì y Ây. This represents a code-switching error, where the English root is morphologically integrated into an Arabic verb form. (c)MWEs/proverbs:Multi-wordexpres- sions (MWEs) often function as single non- compositional units. Consistent with theories of semantic non-compositionality (Nunberg et al., 14 1994), this subcategory targets "literalism errors," where models translate the constituent parts of an idiom rather than its figurative meaning. For instance, the proverbÂls ÂŁr Â L € ÂŁrÂwas translated literally as âand the jar doesnât stay safe every time.â This fails to convey the underlying warning (i.e., âyou wonât get lucky every timeâ), resulting in a nonsensical output for the target audience. (d) discourse marker mismatch: An inconsistency between the semantic/pragmatic meaning of a discourse marker and the context in which it is used (Schiffrin, 1987; Fraser, 2009). This can lead to a lack of coherence or an unintended meaning in a conversation or text. Cohesion in Arabic relies on discourse markers (e.g., wa, fa) that signal logical relations distinct from English (Fraser, 1999). We include this subcategory to capture errors where the model misinterprets these procedural cues, disrupting the textâs argumentative structure or flow.For example, the model mistranslated the phraseÂąl€in the sentence Ì wÂĄ ÂblV Âąl€ dyÂĄ as âGod willing, your request is what will determine.â This represents a pragmatic failure; the model interpretedÂąl€as a literal religious oath (or confused it withÂąl ÂșAJ ), failing to recognize its function here as a discourse marker used for emphasis/hedging. (e) vocatives/honorifics/titles: This category eval- uates the translation of forms of address, which serve crucial social functions in communication. We penalize failures to map honorifics or vocative particles (e.g., Ya) to their target equivalents, often resulting in a loss of the intended politeness level or social hierarchy (Brown and Levinson, 1987; McCready, 2019). This error frequently occurs when models treat titles as literal nouns rather than pragmatic markers of respect. In this example, the model failed to identify the vocative address in the sentenceTÂA  AK Âs ,¹§w AÂ, mistranslating it as âIâm his father, itâs better for hygiene.â The correct interpretation should be a direct address: âDad, I prefer it for the sake of hygiene. (i) Semantics: The Semantics level evaluates the fidelity of meaning transfer (Cruse, 1986), focusing on the preservation of propositional content. We classify these errors into three primary domains: (i) lexical semantics, (i)propositional semantics, and (i) discourse semantics.(a) lexical semantics: Under this level, we propose subcategories specifically tailored for MT in general, and for Arabic dialectal in particular. These include: named entity, failure of the model to translate proper nouns, geographical locations, or organizations correctly. For example,©r  where (Um Fakhri) is translated into (Um Fakhr). wrong term: This refers to the incorrect use of a term that violates domain-expert usage (e.g., legal or medical, etc.). For instance, the model translates the word comediansâ in the phrase break with comedians is very differentâ into ÂyĂrhmÂinstead ofÂylmÂ; overtranslation, where the model uses a specific word (hyponym) when the source used a general word (hypernym) (e.g., âCarâ vs.âFerrariâ); undertranslation, where the model uses a general word when the source uses a specific word (e.g., âFerrariâ vs. âVehicleâ); transliteration: This category applies when the model provides a phonetic rendering of the source text rather than a semantic translation. For example, translatingdm ¹l dm w§ Âhm ¹l as "Ayou al-hamd lillah al-hamd lillah" instead of its English equivalent (e.g., Yes, thank God, anywayâ); unidiomatic/unnatural style, where the output is grammatically correct but sounds unnatural; awkward style, where the style involves excessive wordiness or overly embedded clauses; unintelligible, where the output is garbled or incomprehensible; measurement units: use of an inappropriate measurement format for its locale; for example, converting units (lbsâkg) is a lexical change required to preserve semantic reality; coverage: unknown term/dialect: This category includes cases where the model fails to recognize a specific dialectal lemma.For example, in the sentenceA§ Âm ÌÂwbyF ÂąylVwÂ, the model mistranslated the Egyptian termÂąylVwÂ(âunemployedâ) as âyou crazy ones.â This indicates a coverage failure where the model lacked the necessary lexical representation for this dialect-specific term, resulting in incorrect English output; disambiguation: polysemy failure, where the model knows the word but selects the wrong meaning (your âSignâ vs âKnock downâ example); and disambiguation: cross-variety interference, where the model knows the word but assigns it the Standard meaning instead of the Dialect meaning. For example, the model misinterpreted the homographÂyrJin the phraseÂyrJ rm € rf Ìm§ras âdrinks,â rendering it as âTwo girly yellow and red drinks.â The correct translation is âtwo pairs of womenâs socks.â This represents a failure to distinguish the dialectal 15 term for socks (shurrab) from the standard term for drinks (sharab). (b) Propositional Semantics: Truth-conditional semantics (Soames, 1987) defines the meaning of a sentence by specifying the conditions under which it is true or false. This category addresses the conservation of information, ensuring that the target text neither adds nor subtracts from the message of the source. We categorize these errors as follows: addition, where the translation contains information not found in the source; omission, where source content is missing from the target; and untranslated, where segments remain in the source language in the final output; and finally hallucination: This error occurs when the model generates content that is not supported by the source text, often producing output that is fluent but semantically unrelated. For instance, the model translatedryb AÂĄrt Ât °¥into âNow, your sister is very spoiled.â The concept of âspoiledâ is a hallucination; it has no basis in the source text, which actually refers to a âlarge delayâ or âarrearsâ (AÂĄrtÂ) (c) Discourse Semantics: This domain includes four subcategories: inconsistent use of terminol- ogy, where the model uses multiple terms for the same concept in contexts where consistency is desirable; inconsistent with terminology resource, when the model uses a term that differs from term usage required by a specified termbase or other resource; pronouns, when the model fails to correctly render pronouns, causing changes in speaker or referent, or flipping the intended gender; and inconsistent style, when the style varies inconsistently throughout the text. (iv) Morphosyntax: Evaluates adherence to the structural rules of the target language, focusing on the interface between morphology and syntax (Radford, 2004). We classify errors into two types: grammar: Addresses violations of morphological agreement (Brown et al., 2013; Corbett, 2006), capturing errors in wrong number, gender, and verb tense. We include two specific subcategories under this type: verbal features, where the translation violates target rules regarding tense, aspect, voice, mood, and person (Palmer, 2001).For example, the Moroccan phrasem A A At was mistranslated as âyouâve really cooked up something.âThis indicates a mistranslation of syntactic form, where the model rendered the intended imperative mood as a declarative statement; and nominal features, which covers errors in number, gender, case, definiteness, and state (Idafa). For instance, the model failed to translate the subject in this Mauritanian sentence and changed it into something else ÌÂC Ì Ay A§ ÂA ÂkÂto be "Oh, my brothers, Iâm taking a break". constituent order: Focuses on the linear arrangement of elements (Greenberg et al., 1963). We explicitly include address format and date format here, as these require specific reordering to match the target locale conventions rather than direct translation. 16 Category (Lightweight LQM) Error Type (Lightweight LQM) Subcategory (Diagnostic LQM) Definition SociolinguisticsLanguage Register Standardization InterferenceUse of MSA instead of the target variety. Wrong DialectOutput overlaps with or uses an incorrect dialect. Register MismatchFormality level higher or lower than required. Pragmatics Use, Context, Cultural Appropriateness MWEs/ ProverbsFails to deliver idiomatic translation; misuse of expression. Code SwitchingFailure to handle foreign words or recognize code shift. Speech Acts/ Illocutionary Force Intended illocutionary force or speaker intention not conveyed. Discourse Marker Mis- match Misuse of discourse markers affecting cohesion. Semantics Lexical Semantics Named EntityFailing to map the proper noun to the correct referent. Wrong termThe term is invalid for the domain or creates a conceptual mismatch. OvertranslationUsing a specific word (Hyponym) for a general source word (Hypernym). UndertranslationUsing a general word (Hypernym) for a specific source word (Hyponym). TransliterationIncorrect phonetic rendering into the target language. Unidiomatic StyleThe style is grammatical but unnatural. Awkward StyleExcessive wordiness or overly embedded clauses. UnintelligibleThe text is garbled or incomprehensible. Measurement UnitsThe measurement format is inappropriate for the locale. CoverageThe model fails to recognize a specific dialectal lemma. Polysemy FailureThe model picks the wrong meaning for a polysemous word. Cross-Variety InterferenceStandard meaning assigned instead of the Dialect meaning. Propositional Semantics AdditionThe target includes information that is not present in the source. OmissionContent present in the source is missing from the target. UntranslatedThe source segment is carried over without translation. HallucinationAdding new information that changes the facts. Discourse Semantics Inconsistent use of terminol- ogy Multiple terms used for the same concept. Inconsistent with terminol- ogy resource The usage of terms differs from the specified resource. PronounsIncorrect pronoun causing a change in speaker/gender. Inconsistent StyleThe style varies inconsistently throughout the text. Morphosyntax Grammar (wrong number, gender, verb tense) Verbal FeaturesViolates grammatical rules (Tense, Aspect, Mood, etc). Nominal FeaturesViolates nominal rules (Number, Gender, Case, etc). constituent order Address FormatInappropriate address format for locale. Date formatInappropriate date format for locale. Orthography/ Writing conventions Spelling Typo / SlipObvious mechanical error (e.g., typos). Inconsistent Spelling -Same word spelled differently within text. Unconventional Spelling -Spelling hard to read, even if phonetic. Surface Mechanics Number FormatInappropriate number format for locale. CurrencyIncorrect currency format for locale. Time formatIncorrect time format for locale. TelephoneInappropriate telephone number format. Punctuation -Incorrect according to target conventions. Graphetics -Character EncodingCharacters garbled due to incorrect encoding. Table A.1: Hierarchical classification of the LQM. Categories and error types represent the Lightweight LQM; subcategories represent Fine-grained analysis (Diagnostic LQM). 17 Figure A.1: Distribution of LQM error categories in our dataset, showing that semantic and lexical- semantic errors constitute the majority of labeled errors. (v) Orthography/Writing Conventions: In this level, we evaluate the mechanical correctness of the written output, focusing on the visual repre- sentation of language (Derwing, 1992). We clas- sify errors into five specific types: spelling: This category addresses deviations from standard ortho- graphic norms. It includes the typos/slips subcat- egory for obvious mechanical errors (e.g., char- acter insertions or deletions) that violate the fixed spelling rules of the target language. We also distin- guish between inconsistent spelling and unconven- tional spelling. The category of surface mechanics governs non-lexical formatting conventions and comprises four subcategories: number format, cur- rency, time format, and telephone. Finally, punc- tuation addresses cases where punctuation marks are missing, misused, or inconsistent with target language conventions. (vi) Graphetics: The lowest level of the hierarchy addresses the technical realization of the text code. We identify one primary failure mode, which is character encoding, where the output text is gar- bled due to incorrect encoding or decoding pro- cesses. For example, in this model outputE A§±. Illustrative examples of LQM categories and sub- categories are in Table G.2. B LQM External Review To obtain external expert feedback on the proposed framework, we invited two linguists specializing in Arabic linguistics and translation to assess its conceptual soundness and practical applicability. Overall, both reviewers viewed the proposed LQM as a valid and promising linguistically motivated framework for machine translation evaluation. In particular, they highlighted its relevance for diglos- sic language settings, where variation across stan- dard and non-standard varieties, register, and soci- olinguistic meaning is central to translation quality. They identified this as a key strength of the frame- work, especially for Arabic and related contexts. At the same time, the reviewers noted that some category boundaries, particularly those span- ning semantic, pragmatic, and sociolinguistic di- mensions, would benefit from further clarifica- tion. They also suggested streamlining some fine- grained subcategories to improve annotation con- sistency and usability. Taken together, their feed- back supports the relevance of the framework for MT assessment in diglossic languages while also identifying areas for refinement. In response, we revised the category boundaries to reduce potential overlap and strengthened the definitions and illus- trative examples associated with each category. C Prompts D Data Annotation Annotator selection went through multiple stages of quality assurance. First, we developed detailed annotation guidelines that included illustrative ex- amples tailored to the seven Arabic dialects cov- ered in our study. These examples were drawn from a wide range of regional varieties to ensure clarity and consistency for annotators from differ- ent Arab countries. Next, we designed an anno- tation interface that incorporated all LQM cate- gories along with the subcategories introduced in our framework. We uploaded all model outputs to this interface to facilitate the annotation pro- cess. Before launching the full annotation round, we conducted an internal pilot in which we sam- pled a subset of the data. This allowed us to verify that the tool functioned properly and that all nec- essary features were available to support accurate and efficient annotation. During the pilot, we iden- tified several model-generated outputs that were ambiguous or did not fit neatly into the existing error categories. In response, we added a comment box to the interface so that annotators could pro- vide feedback on model behavior or flag unlisted error types, such as pragmatic errors. After refin- 18 Prompt Translation Prompt Generator (Both Directions ENGâDialect) def create_prompt(src, trg): if src == âENGâ: prompt = fâTranslate the following English phrase into langs_map[trg] Arabic dialect written in Arabic script. Your response must only contain the translated text, with no additional explanations or labels.â else: prompt = fâTranslate the following langs_map[src] Arabic dialect phrase into English. Your response must only contain the translated text, with no additional explanations or labels.â return prompt Figure C.1: Prompt used to generate all dialectal variants in our dataset (ENâDialect; DialectâEN). ing the tool and guidelines based on these insights, we onboarded all annotators by conducting live demonstrations that walked them through both the annotation procedures and the interface for label- ing span-level errors. All annotators were invited to share feedback and comments, which we incor- porated as part of our iterative quality assurance process. Annotator Profiles. To ensure linguistic ex- pertise and dialectal authenticity, we employed direction-specific annotation teams: âąDialectâEnglish: This direction was annotated by two senior linguists (Ph.D. holders) special- izing in translation studies and linguistics, both of whom are native Egyptian Arabic speakers. To mitigate source-side ambiguity, annotators were provided with both the original dialectal text and its MSA equivalent as a reference. Ap- proximately40% of the dataset was finalized through collaborative live sessions to ensure full agreement, leveraging the MSA context to re- solve nuanced dialectal expressions and improve overall comprehension. âąEnglishâDialect: This task involved four na- tive speakers (Moroccan, Palestinian, Emirati, and Mauritanian), each holding an MA or Ph.D. in related fields. Each annotator labeled only their native dialect to ensure authentic judgments of naturalness. For the English-to-Egyptian di- rection, the annotation was performed by the same two linguists mentioned previously. For each of the two directions,40% of the data were labeled by two annotators in live sessions with full agreement, once full agreement was reached, the rest of the data was annotated by a single annotator. E Linguistic insights on the LLMs performance The primary limitations of current LLMs in dialectal ArabicâEnglish translation are linguistic rather than merely computational. Our analysis shows that the most persistent errors stem from in- terpreting dialect-specific semantics, the dominant failure mode in the DialectâEnglish direction. Across dialects, Semantic errors account for 67.7%â92.3% of total error mass, indicating that the main bottleneck is weak lexical and conceptual mapping for idiomatic and culturally situated expressions. These often encode illocutionary force or social alignment that surface-level lexical correspondence cannot recover, producing translations that are formally plausible but pragmatically deficient. This is especially visible in high-resource dialects such as Egyptian, where Pragmatic errors reach24.3%. Models also show systematic weakness in dialectal morphosyntax; even in comprehension, Morphosyntactic failures remain notable in dialects such as Mauritanian (3.4%), where non-canonical constructions diverge from MSA norms. These challenges intensify and change struc- turally in the EnglishâDialect direction, where failure shifts from semantic decoding to deficient sociolinguistic authenticity. Models display a pro- nounced âidentity crisis,â often defaulting to stan- dardization (MSA-vertical mismatch) or producing hybrid forms (wrong dialect-horizontal mismatch). This is reflected in the rise of Sociolinguistic errors, peaking at70.6% for UAE and58.5% for Pales- tinian generation. The problem is especially pro- nounced in open-source architectures: whilePro contributes as little as2.1% to the error mass in 19 Model Selection Direction of Translation Models Output User name Source Sentence Dialectal Reference Error Severity Error Subcategory Error TypeError Category LQM Categories MSA Reference Figure D.1: Screenshot of our LQM annotation tool for error labeling. Moroccan generation, open-source models such as FanarandCommand-Rcontribute up to33.0% and20.4%, respectively. This suggests a data- poverty effect in smaller or open-source models, yielding overgeneralized constructions and fewer culturally grounded pragmatic markers. Overall, these findings show that existing LLMs inade- quately model the interaction of morphology, syn- tax, and sociolinguistic variation, motivating the more granular, linguistically informed evaluation provided by LQM. FRobustness of LQM to Sentence Length Per-sentence normalization can inflate or dilute scores for very short or very long segments. To verify that our findings are not artifacts of such effects, we conduct three complementary analy- ses: (i) corpus-level micro-averaged LQM, (i) a robustness check across length buckets, and (i) a rank-stability test using Spearman Ï. Micro-averaged LQM.Instead of averaging per- segment scores, we accumulate all error mass and all token counts across the corpus before dividing once: LQM ÎŒ = max 0, 100â 100· P s E s P s L s (1) This âsum first, divide onceâ principle eliminates the sensitivity to segment length. Table F.1 reports micro-averaged scores for selected directions. Dir.Fanar Cmd-A Cmd-R7B Gemma FlashPro EgyâEn 49.4 62.5 39.7 73.0 77.477.5 EnâEgy 53.6 70.8 9.9 56.6 68.171.9 PalâEn 67.1 78.0 64.3 74.7 76.181.0 UaeâEn 24.0 66.2 36.8 53.2 55.366.5 YemâEn 64.0 57.8 52.9 69.3 72.175.0 Avg. 51.6 67.1 40.7 65.4 69.874.4 Table F.1: Micro-averaged LQM (LQM ÎŒ ) for selected directions. The model ranking is consistent with the per-sentence results in Table 3. The model ranking under micro-averaged scor- ing is consistent with the per-sentence results re- ported in Table 3.Gemini-2.5-Proranks first in the majority of directions under both formu- lations, andCommand-R7Bremains the weak- est model overall. The direction-level patternâ ENâMOR as the most challengingâis also pre- served. Length-bucket analysis. We split all segments into short, medium, and long using the 33rd and 66th percentiles of target-side token count (quan- tile bucketing ensures balanced counts per bucket). The cut-offs are: short†14tokens (n=1,165), medium15â22tokens (n=1,233), and long> 22 tokens (n=1,080). We recompute micro-averaged LQM within each bucket for every for each mfor each modelâdirection combination. Tables F.2 and F.3 show results for two representative direc- 20 ModelLength bucket ShortMediumLong Gemini-2.5-Pro66.173.580.0 Gemini-2.5-Flash66.071.181.3 Gemma-27b58.965.078.3 Command-A37.464.965.9 Fanar-9B37.843.253.9 Command-R7B0.045.348.3 Table F.2: Micro-averaged LQM by length bucket for EgyâEng. ModelShortMediumLong Gemini-2.5-Pro76.075.282.3 Gemma-27b65.667.276.8 Gemini-2.5-Flash58.766.879.5 Command-A50.060.9 82.8 Fanar-9B69.858.369.7 Command-R7B48.144.968.7 Table F.3: Micro-averaged LQM by length bucket for PalâEng. tions. The top-tier models (Gemini-2.5-Pro, Gemini-2.5-Flash) and the weakest model (Command-R7B) maintain their relative positions across all three length buckets in the vast majority of directions. Rank stability. To quantify stability formally, we compute the Spearman rank correlation of model scores between pairs of length buckets for each direction (Table F.4). Averaged across all 12 directions, SpearmanÏ = 0.71(short vs. medium),0.71(medium vs. long), and0.62 (short vs. long), confirming that model rankings are substantially preserved across length strata. The exceptionsâUaeâEng (Ïâ 0.06â0.43) and EngâMau (Ï â â0.15â0.70)âreflect genuine variation in error profiles within those dialect pairs rather than scoring artifacts, consistent with the higher cross-model variability observed for these directions in Table 3. Summary.Both the micro-averaged formulation and the length-bucket analysis confirm that our main conclusionsâGemini-2.5-Proleading overall, ENâMOR as the most challenging direc- tion, and direction-dependent shifts in error typeâ are robust to sentence length effects and are not driven by score distortion in short segments. DirectionS vs MÏM vs LÏS vs LÏ EgyâEng0.829 â 0.886 â 0.886 â EngâEgy0.943 â 0.7710.829 â EngâMau â0.696 â0.029 â0.145 EngâMor1.000 â 0.7910.791 EngâPal0.943 â 0.886 â 0.771 EngâUae0.6000.886 â 0.829 â JorâEng0.6570.7140.829 â MauâEng0.829 â 0.886 â 0.886 â MorâEng0.6570.7710.829 â PalâEng0.6570.6000.200 UaeâEng0.1450.4290.058 YemâEng0.6000.886 â 0.714 Mean0.710.710.62 Table F.4: SpearmanÏof model rankings across length buckets. S = Short, M = Medium, L = Long. â p < 0.05; â p < 0.01. GLQM Fine-Grained Error Distribution Across Models and Per Direction Table G.1 shows the error distribution across the 6, 113samples. It reveals a sharp divide between translation directions, with Dialect-to-English ac- counting for58.6% of failures compared to41.4% in English-to-Dialect. The most critical gen- erational hurdle in English-to-Dialect is related to standardization interference, which represents 35.62% of errors in this direction. This indicates a pervasive MSA bias, where models default to formal or prestige varieties rather than maintaining the requested dialectal authenticity. This sociolin- guistic failure is further complicated by horizontal dialectal bleed, where features from different re- gional varieties overlap, leading to a21.29% error rate in dialect target selection (wrong dialect). In the Dialect-to-English direction, failures are predominantly semantic and pragmatic. Named Entity recognition is the primary bottleneck at 13.49%, likely driven by the lack of standardized dialectal orthography, which makes proper noun recovery inconsistent. Furthermore, the model fre- quently fails to capture the speakerâs social intent, with speech acts (2.54%) and vocatives (2.49%) to- gether accounting for over5% of errors in Dialect- to-English. While these stylistic and lexical is- sues vary by direction, morphosyntactic strug- glesâspecifically verbal featuresâremain a per- sistent technical barrier across the board, appearing at significant rates in both decoding (3.04%) and generation (7.35%) tasks. 21 Error Sub-categoryDialectâ Eng (%)Engâ Dialect (%) Standardization Interference0.0035.62 Wrong Dialect (Horiz. Mismatch)0.0021.29 Socioling. Register Mismatch0.590.39 MWEs/Proverbs6.282.53 Speech Acts & Illocutionary Force2.540.51 Vocatives/Honorifics/Titles2.490.32 Discourse Marker Mismatch0.890.12 Pragmatic Code Switching0.110.59 NE (Named Entities)13.493.52 Coverage: Unknown Term/Dialect9.331.22 Omission8.941.46 Unnatural/Unidiomatic Style6.653.32 Awkward Style4.615.02 Addition5.862.09 Pronouns6.840.91 Disambiguation: Cross-Variety6.670.28 Wrong Term3.183.08 Hallucination4.471.18 Undertranslation2.901.03 Disambiguation: Polysemy Failure3.020.47 Transliteration2.960.36 Overtranslation1.650.28 Untranslated1.420.24 Inconsistent Term. Resource1.230.12 Inconsistent Style0.030.55 Unintelligible0.030.16 Semantic Inconsistent Use of Terminology0.030.00 Verbal Features3.047.35 Morph. Nominal Features0.280.79 Typo/Slip0.144.74 Currency0.140.16 Punctuation0.080.20 Ortho. Number Format0.000.04 Graph. Character Encoding0.060.08 Table G.1: Fine-grained LQM error distribution. Bar lengths are proportional to the percentage of total errors per translation direction. 22 CategoryError TypeSubcategoryDirSourceTargetModel Sociolinguistics Code & Register Selection Standardization Interference ENâMO My phone was silent so I did not see your calls. A ÌÂwfl ÂALtfJ A MA .AmÂAkm ÂA§ ÂÂA§ ⊠Gemma Wrong DialectUAâEN May Allah protect you. You are good and blessed.  . § Âąl§E .ÂCAb€â Gemini Register MismatchENâMOWhat are you doing, man? r§ A M An â Gemini Pragmatics Use, Context, Cultural Approp MWEs/ Proverbs EGâEN Âąmq ÂÂA Ìy A ¹§ ÂÂwq ?l € Ly Yqb§ AK wF Tell you what, why donât we grab a bite together so thereâs bread and salt between us ? â Gemini Code SwitchingENâEG Itâs not called a casino, itâs called a nightclub, maâam. §A ¹mF ,wn§EA ¹mF L .Âd A§ wl â Gemini Speech Acts/ Illoc Force MOâEN € Ìt`m € ÌtyK At A ?ÂA§ .Âyrt AÂĄ Ìt`mF Right? Now you went, gathered, and listened to this nonsense. ⊠Gemma Discourse Marker Mismatch MOâEN € ÂA w ÌÂdJ A A TrO Ât ©dh wÂEw€ yÂwlb ÂÂwm To be honest, dad, Iâm really embar- rassed about what happened with that girl and the photos, andhonestly, Mehdi saved the situation. ⊠Gemma Vocatives /Hon- orifics/Titles YEâEN wms A QÂź , ÂŁA A§ L§ Ìn`§ .ÂŁrAF AÂCw Âw € , Cw€ , So what,man, just toughen up, inherit [this], and make our affairs difficult. ⊠Gemma Semantics Lexical Semantics Named EntityENâPA A§ AW w A§ T`yF€ ÂtÂĂ Âąl€ !r By God, your heart is wide, Abu Atta, yourooster ! âČ Fanar Wrong termENâPA € Ây r € Âr w w ¹l€ € wÂwW§ € w`s§ wlS € Âź ÂĂ ÂhÂrfÂn A wÂwW§ € w`s§ dlb Ì€ Ì £wlm Ìl d` by God, even if they go for Hajj once, twice, three times, and keep circumambulating and walking, we wonât forgive them for their sin after what they did to the children of the town. âČ Fanar OvertranslationMAâEN wn Â§CA€ wn ÂyK§A ©wJ Ă ÂrZ Any  ÂÂAV ÂÂA Ă€ C  §rK ÂwW` You see how we live, we are barely surviving, and you want everything, youthinkwe will give you twenty thousand? âČ Fanar UndertranslationYEâEN Ìbl ¹ÂA ,H Âąl Yl Ahyl Âąyl £dÂd dÂd Leave it to God, but honestly, my heart isrestless with worry about it. âČ Fanar TransliterationEGâEN A ©E CwÂd ÂA ÂCA d`F XzÂA ÂąnyÂCA An SaadArafaNasif Madkor, as we all know. â¶ Command-R Unnatural StyleMOâEN ÂÂA r§d A T§A Ârm . fF ÌA TC r§dÂA Next time Iâl do like you, Iâl take it easy like Iâm weak or tired. â¶ Command-R Awkward StylePAâEN € C€ dlb Â ÂÂźF A§ ÂÂźF A§ ÂwÂK yq A H !Âd TÂÂźF ÂCA n Peace, peace, the whole country is behind and safety is in front! But Iâve become busy, I donât know. â¶ Command-R Unknown Term/ Dialect UAâEN dy € ,? ÂŁwtm  y§ A wJ ? ÂŁwtm  y§ A wJ What did he bring forthe guest? What did he bring for the guest? â¶ Command-R Polysemy FailureJOâEN ÂrK` ÂÂw €  ÂÂw ©A AÂĄ A Ârk Ìn`§ wJ Am Âl ? Ìn`§ §w Give it to me, I will sign it for you and ten people like you. Do you think I am afraid? ⊠Gemma Cross-Variety Interference JOâEN Âąs A§ €dnO Yl ÂÂA ©r§ Ây € ÂȘwb\ ÌJ  ©d Yly Take care ofthe box, Miss Leila. I want everything to be accurate and precise. â¶ Command-A Propositional Semantics AdditionMAâEN w Â wty ©dn ÌÂAyF ! €rhZ  wty ÂŁw  wty I have a problem with my mother, my brother, and myfather . âČ Fanar OmissionUAâEN §E¹§w wOm AnÂr §€ ©ryF ,Ahy ÂAtK Ìy A¥€rÂE wm Ah§rÂE but where is our private space? I miss it, go, show it to me, uncle. âČ Fanar UntranslatedJOâEN Ârkb Ahyl Ìn`§ ÂFw§ A§ yV .bO , ÌakÂFw§€C i H' , .niÂC âČ Fanar HallucinationEGâEN Ìl Ìt Ìtn L Ìt wRr A . ÂdÂA I also thought you werenât the one sitting there, you stupid woman. â¶ Command-A Discourse Seman- tics PronounsMAâEN €  l ÌÂĂ Ì Âyl ÂwJ .Âąy Ìw Ì Âwq Ìhy ÌÂĂ Look, stay with me, I told you, and stay with her â¶ Command-A Morphosyntax Grammar (number, gender, tense) Verbal FeaturesMOâEN ÌRr§ Âąl ryF Aw ÂwJ A Ì xAn Â€ dn Âyl A§w AÂĄ AnÂw€AO§ may God be pleased with you. Look, there are people whoare helping us with these needs âČ Fanar Nominal FeaturesENâEG I came to check on you, my love, because we heard womenâs feet going up the stairs. I was afraid about you, Kuka. , Ìtbyb A§ Ìkyl ÂmV y An`mF AK ÌlCT`ÂAV §r A§ Ìkyl Tf§A n .Âls Yl .AÂw ⶠCommand-A Continued on next page... 23 Table G.2 â continued from previous page CategoryError TypeSubcategoryDirSourceTargetModel Orthography/ Writing conventions Unconventional Spelling ENâPA What do I know, uncle? Is Maârouf the imposter better than us!? Âwq` ? ÌĂm A§ ,A Âr`t wJ ĂyÂĂdÂ!?AĂn Âs €r` ⊠Gemma Surface Mechanics CurrencyENâUA The fellow will give me ten rubies, father. I told you I went and talked to him, and now she has agreed. rK ÌnyW`y ÂArÂwÂA§.Âąb§ , Ây € ,Âątml€ C  l . q€ ⊠Gemma Table G.2: LQM error types across all linguistic levels examples from covered dialects. 24