Paper deep dive
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu, David Tan, Doreen Osmelak, Ona de Gibert, Ariun-Erdene Tumurchuluun, Ashok Urlana, Fedor Sizov, Hale Sirin, Jesujoba Alabi, Karrar Talib Abed, Mateusz Klimaszewski, Nikolay Bogoychev, Niyati Bafna, Patricia Schmidtova, Preksha Manjunath Shanbhag, Sherrie Shen, Vilem Zouhar, Vivek Iyer, Yasser Hamidullah, Yusser Al Ghussin, Zheng Zhao
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.
Tags
Links
- Source: https://arxiv.org/abs/2608.09766v1
- Canonical: https://arxiv.org/abs/2608.09766v1
Trouble viewing inline? Open PDF directly â
Full Text
57,049 characters extracted from source content.
Expand or collapse full text
CULTIVAR: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness Pinzhen Chen 1 Koel Dutta Chowdhury 2 Xiaoya Xu 1 David Tan 3 Doreen Osmelak 3,4 Ona de Gibert 5 Ariun-Erdene Tumurchuluun Ashok Urlana 6,7 Fedor Sizov 3 Hale Sirin 8 Jesujoba O. Alabi 3 Karrar Talib Abed 9 Mateusz Klimaszewski 10 Nikolay Bogoychev 11 Niyati Bafna 8 PatrĂcia SchmidtovĂĄ 12 Preksha Manjunath Shanbhag 1 Sherrie Shen 13 VilĂ©m Zouhar 14 Vivek Iyer 13 Yasser Hamidullah 15 Yusser Al Ghussin 3 Zheng Zhao 13 1 Queenâs University Belfast, 2 University of Technology Nuremberg, 3 Saarland University 4 University of Melbourne, 5 University of Helsinki, 6 IIIT Hyderabad, 7 TCS Research 8 Johns Hopkins University, 9 Imam Jaâafar Al-Sadiq University, 10 Cohere, 11 Last Token 12 Charles University, 13 University of Edinburgh, 14 ETH Zurich, 15 University of Zurich p.chen@qub.ac.uk Abstract Multilingual translation benchmarks are typ- ically sourced in English and translated into other languages, treating language pairs as the unit of evaluationâa design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore ad- vocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific transla- tion evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisa- tion robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US con- tent better than that of other locales, regardless of language. 1 Introduction Benchmarks are an essential element for develop- ing translation systems. As machine translation technology reaches maturity and is deployed for more languages, recent years have seen a shift from English-centric evaluation, like the early WMT (Bojar et al., 2014) and IWSLT (Cettolo et al., 2014) evaluation campaigns, towards multilingual test suites such as FLORES (Goyal et al., 2022; Team et al., 2024), NTREX (Federmann et al., 2022), and WMT24++ (Deutsch et al., 2025), en- abling evaluation for hundreds of languages in ar- bitrary directions. Cultivar is potted athttps://hf.co/datasets/pinzhen chen/Cultivar-flores, harvestable under C BY-SA 4.0. Yet, massively multilingual test suites share a common limitation due to construction: regardless of language, each instance is anchored to a single information origin, which is typically English, and then translated. While this facilitates multiway par- allelism, it introduces two vulnerabilities. First, because the original English content is publicly available, test sets become susceptible to data con- tamination over time (Magar and Schwartz, 2022; Sainz et al., 2023; Yao et al., 2024; Tan et al., 2026). Second, using such test sets to evaluate translation into English creates a language-content mismatch (Chen et al., 2024), which diverges from real-world use cases, where the input content is natively pro- duced in the input language. BOUQuET (Andrews et al., 2025) alleviates this problem by starting with original texts in eight distinct languages and then translating into more. A separate line of effort trades away multiway parallelism for naturally writ- ten, culturally grounded text, such as WMT since 2019 (Barrault et al., 2019) and multilingual bench- marks like TyDi QA (Clark et al., 2020), which improves the authenticity of evaluation from a mul- ticultural perspective but loses comparability. Nonetheless, we argue that source-origin con- tent alone is insufficient. Existing test sets are fundamentally language-oriented, assuming that each language corresponds to a single cultural con- text. In practice, however, a language can often be used in multiple regions, with diaspora com- munities adding an extra layer of variation. Con- sequently, evaluating a model on a single test for each language either fails to capture other locales 1 arXiv:2608.09766v1 [cs.CL] 10 Aug 2026 or receives an arbitrary, mixed-origin result. Thus, we argue that translation and multilingual evalua- tion should move beyond language-oriented tests towards locale-oriented tests, where multiple lo- calisations of the same language are separately as- sessed to better reflect real-world use. We therefore advocate a source-contrastive paradigm that grounds an input in distinct con- texts (Bawden et al., 2018; Futeral et al., 2023; Alam et al., 2024). Instead of testing a single static sentence, an evaluation suite should feature paired instances that preserve sentence structure while systematically varying locale-specific con- tent based on desired locations. Such a controlled design isolates a modelâs sensitivity to varied con- tent from general translation capability. Crucially, it enables analyses that are impossible with conven- tional single-version tests: diagnosing benchmark contamination and quantifying true localisation ro- bustness through paired performance discrepancies. To instantiate this paradigm, we present the Cul- tivar benchmark, grounded in 27 locales, created through an LLM-human co-creation pipeline start- ing from FLORES. We benchmark 32 language and translation models and share our findings. To summarise, our contributions are: âąResource: Cultivar, a multilingual translation test set grounded in locales, produced via a human-LLM co-creation pipeline. We preserve the comparability of massively multilingual test sets while incorporating native source content. âąEvaluation: We evaluate 32 models on Culti- var using metrics like BLEU, chrF, and word recall. We report standalone locale-level scores and conduct a contrastive study with FLORES, quantifying localisation robustness and examin- ing overfitting. âą Findings: Performance of MT-specialised and smaller models tends to drop on localised con- tent. Significant gaps between original and lo- calised tests reveal potential FLORES overfit- ting. Two smaller-scale studies also show that models perform better on content grounded in the US than on localised content in the official language of a particular region. 2 Cultivar We construct Cultivar through a three-stage pipeline that localises FLORES: 1) named entity- rich instance selection; 2) LLM-assisted paraphras- ing; 3) human post-editing to ensure quality. Al- though we create Cultivar by localising FLORES, technically, the pipeline can be applied to any ex- isting benchmark. Unlike efforts that extend FLORES to new lan- guages by translating from scratch (Dale et al., 2025), we formulate localisation as creative para- phrasing. Our design preserves the sentence struc- ture of original instances while adapting cultural content to a specific location, thereby producing sentence pairs that remain directly comparable. An LLM-assisted annotation workflow is par- ticularly suitable for this setting because localisa- tion requires both creativity and diversity beyond a translation task, making creation from scratch dif- ficult for humans. Our use of human post-editing ensures quality and cultural appropriateness. In this study, we only localise any-to-English translation directions. The resulting English target sentences across locales are no longer the same, but they remain comparable. Namely, it is still meaningful to compare error counts or even word- level metrics like BLEU between into-English di- rections. 2.1 Named entity-rich instance identification Localisation primarily affects named entities such as people, organisations, locations, events, etc. We thus first identify named entities in each instance of the English FLORES dev set using spaCy (Hon- nibal et al., 2020) and manually correct missing or erroneous annotations. This process finds 658 sentences (approximately 65% of the dev set) con- taining at least one named entity. From these, we manually select 200 sentences having at least 2 named entities. We exclude in- stances dominated by numerals or topics such as medicine and science, where named entities are globally shared. The resulting named entity-rich subset forms the input to the LLM localisation step. 2.2 LLM paraphrasing Given a source language of interest, for each se- lected English instance, an LLM generates a lo- calised source-English pair from the original FLO- RES source-English pair and a target location. The model is instructed to preserve the sentence struc- ture whenever possible while replacing content with appropriate alternatives for the new location. Our full prompt is provided in Figure 4 in Sec- tion A. We use GPT-5.5 (gpt-5.5-2026-0423). 2 deu_Latn_Germany_0 confirmedGerman diff: 14English diff: 17 ORIGINAL GERMAN Am Montag haben die Wisenschaftler der Stanford University School of Medicine die Erfindung eines neuen Diagnosetools bekanntgegeben, mit dem Zellen nach ihrem Typ sortiert werden können: ein winziger, ausdruckbarer Chip, der fĂŒr jeweils etwa einen US-Cent mit Standard- Tintenstrahldruckern hergestellt werden kann. REPHRASED GERMAN Am Montag haben die Wissenschaftler der CharitĂ© â UniversitĂ€tsmedizin Berlin die Erfindung eines neuen Diagnosetools bekanntgegeben, mit dem Zellen nach ihrem Typ sortiert werden können: ein winziger, ausdruckbarer Chip, der fĂŒr jeweils etwa einen Euro-Cent mit Standard- Tintenstrahldruckern hergestellt werden kann. ORIGINAL ENGLISH On Monday, scientists from the Stanford University School of Medicine announced the invention of a new diagnostic tool that can sort cells by type: a tiny printable chip that can be manufactured using standard inkjet printers for possibly about one U.S. cent each. REPHRASED ENGLISH On Monday, scientists from CharitĂ© â UniversitĂ€tsmedizin Berlin announced the invention of a new diagnostic tool that can sort cells by type: a tiny printable chip that can be manufactured using standard inkjet printers for possibly about one euro cent each. ANNOTATION Confirm final version Revised rephrased German sentence î· î· Am Montag haben die Wissenschaftler der CharitĂ© â UniversitĂ€tsmedizin Berlin die Erfindung eines neuen Diagnosetools bekanntgegeben, mit dem Zellen nach ihrem Typ sortiert werden können: ein winziger, ausdruckbarer Chip, der fĂŒr jeweils etwa einen Euro-Cent mit Standard-Tintenstrahldruckern Revised rephrased English sentence On Monday, scientists from CharitĂ© â UniversitĂ€tsmedizin Berlin announced the invention of a new diagnostic tool that can sort cells by type: a tiny printable chip that can be manufactured using standard inkjet printers for possibly about one euro cent each. deu_Latn_Germany_2 pending confirmationGerman diff: 6English diff: 6 ORIGINAL GERMAN Der JAS 39C Gripen stĂŒrzte gegen 9:30 Uhr Ortszeit (02:30 UTC) auf eine Startbahn und explodierte, sodass der Flughafen fĂŒr kommerzielle FlĂŒge geschlossen werden musste. REPHRASED GERMAN Der Eurofighter Typhoon stĂŒrzte gegen 9:30 Uhr Ortszeit (02:30 UTC) auf eine Startbahn und explodierte, sodass der Flughafen fĂŒr kommerzielle FlĂŒge geschlossen werden musste. ORIGINAL ENGLISH The JAS 39C Gripen crashed onto a runway at around 9:30 am local time (0230 UTC) and exploded, closing the airport to commercial flights. REPHRASED ENGLISH The Eurofighter Typhoon crashed onto a runway at around 9:30 am local time (0230 UTC) and exploded, closing the airport to commercial flights. ANNOTATION Confirm final version Revised rephrased German sentence Der Eurofighter Typhoon stĂŒrzte gegen 9:30 Uhr Ortszeit (02:30 UTC) auf eine Startbahn und explodierte, sodass der Flughafen fĂŒr kommerzielle FlĂŒge geschlossen werden musste. Revised rephrased English sentence The Eurofighter Typhoon crashed onto a runway at around 9:30 am local time (0230 UTC) and exploded, closing the airport to commercial flights. deu_Latn_Germany_3 pending confirmationGerman diff: 4English diff: 4 ORIGINAL GERMAN Der Pilot wurde als StaffelfĂŒhrer Dilokrit Pattavee identifiziert. REPHRASED GERMAN Der Pilot wurde als StaffelfĂŒhrer Werner Mölders identifiziert. ORIGINAL ENGLISH The pilot was identified as Squadron Leader Dilokrit Pattavee. REPHRASED ENGLISH The pilot was identified as Squadron Leader Werner Mölders. ANNOTATION Confirm final version Revised rephrased German sentence Der Pilot wurde als StaffelfĂŒhrer Werner Mölders identifiziert. Revised rephrased English sentence The pilot was identified as Squadron Leader Werner Mölders. deu_Latn_Germany_5 pending confirmationGerman diff: 6English diff: 6 ORIGINAL GERMAN Der 28-jĂ€hrige Vidal war vor drei Spielzeiten von Sevilla zu Barça gekommen. REPHRASED GERMAN Der 28-jĂ€hrige Goretzka war vor drei Spielzeiten von Schalke zu Bayern gekommen. ORIGINAL ENGLISH 28-year-old Vidal had joined Barça three seasons ago, from Sevilla. REPHRASED ENGLISH 28-year-old Goretzka had joined Bayern three seasons ago, from Schalke. ANNOTATION Confirm final version Revised rephrased German sentence Der 28-jĂ€hrige Goretzka war vor drei Spielzeiten von Schalke zu Bayern gekommen. Revised rephrased English sentence Download JSONLoad JSONReset draft 1 / 200reviewed0revised pairs Loaded saved draft from this browser. Download the annotation JSON when the review is done. Figure 1: An illustration of the annotation interface and the content before and after localising German written in Latin script (deu_Latn) to Germany, paired with English. The difference between theoriginalandreplaced content is highlighted. The annotation panel on the right is, by default, populated with the LLM-rephrased content, which the annotator can edit; as they edit, the highlights update accordingly. Once done, annotators check âConfirm final versionâ to freeze the edit box to finalise this sentence pair. 2.3 Human annotation Annotator profileWe use three criteria to select an annotator: 1) proficient in the source language; 2) proficient in the target language (English); and 3) they live(d) in the location. These requirements ensure both language competence and familiarity with local cultural conventions. The majority of annotators are NLP researchers with some understanding of NLP evaluation. All annotators are included as authors, as their contri- butions directly shape the benchmark and analyses. Annotation processWe ask annotators to inspect each instance pair before and after LLM localisa- tion and perform one of the three actions below: 1. accept the LLM paraphrase; 2. post-edit the paraphrase; 3. revert to the original FLORES entry. Similar to the LLM prompting stage, annotators are instructed to preserve the original sentence struc- ture whenever possible to maintain comparability between the original and localised benchmarks. The first two options produce approved localised instances. Annotators were also allowed to use op- tion 3 sparingly, when localisation is unnecessary, culturally inappropriate, or difficult. Before annotation starts, annotators are in- structed to read through a brief motivation and description of the project, annotation guidelines, and a few examples of what they do and do not need to edit. Detailed annotation instructions are presented in Figure 5 in Section B. Annotation time ranges from 4 hours to 12 hours. 3 LangScriptLocation Kept LLM Paraphrase Reverted to FLORES Human EditedAverage Edits Source OnlyTarget OnlyBothSourceEnglish acmArabIraq176014194.804.28 apcArabSyria1550142294.784.38 benBengIndia1460410137.674.19 bulCyrlBulgaria161020374.414.14 catLatnSpain136811544.704.43 cesLatnCzechia10901512644.614.44 cmnHansChina168127224.014.27 Singapore153175344.474.72 UK â 1707100134.374.22 US â 180401154.344.28 deuLatnGermany117032784.484.46 Switzerland154161384.664.50 filLatnPhilippines19314024.244.15 gomDevaIndia174800186.143.99 hinDevaGermany â 11001986.904.32 India112707747.214.53 US â 1133480457.444.47 jpnJpanJapan892732703.634.47 khkCyrlMongolia6805314655.254.46 pltLatnMadagascar18705084.123.92 polLatnPoland176301924.284.12 rusCyrlRussia1368260304.904.33 slkLatnSlovakia144131514.684.29 spaLatnSpain117623724.914.44 telTeluIndia173052207.304.36 turLatnTurkey1476131334.214.19 yorLatnNigeria1270561165.284.15 Average140.073.6711.004.1541.115.104.31 Table 1: Summary of human edits to LLM-localised FLORES instances: LLM paraphrases are kept as is, reverted to FLORES, or edited by the annotators. Localisation targets regions where a language holds official or de facto status, except for the four entries marked with aâ . Annotation interface Annotation is performed through a custom web interface shown as Figure 1. It displays the original and localised sentence pairs side by side while highlighting changed spans. The edit box is initialised with the LLM paraphrase, allowing annotators to efficiently post-edit the text. Differences and highlights are updated as edits are made. After completing each entry, annotators finalise it by checking a confirmation box. This deliberate step acts as a check on attention and responsibility, as it encourages careful review and reduces the likelihood of inadvertent skipping. 2.4 Annotated data statistics The final localised Cultivar benchmark consists of 27 language-script-location combinationsâ referred to as locales. We represented them with the same language and writing script codes as FLO- RES, together with the country name where the data is grounded. There are 21 unique languages or dialects, 8 unique writing scripts, and 21 unique countries or regions. Each locale has 200 instances contributed by a single annotator. A breakdown of annotation statistics for each locale is in Table 1. On average, annotators accept 70% of LLM-generated localisations without mod- ification, post-edit roughly 28%, and revert fewer than 2% to the original FLORES entries. Conse- quently, over 98% of Cultivar is newly localised content. A few outliers exist, for example: all but one of the LLM paraphrases are human-edited when Hindi is localised to Germany; only 7 and 13 are changed when Filipino and Malagasy are localised, respectively. At the source-target pair level, we count a mod- ified continuous string surrounded by original strings as an edit. We only count edits to an in- stance if the LLM paraphrase is revisedâeither reverted to FLORES or edited by the annotator. On average, there are 4 edits per English instance and 5 edits per source instance. However, the latter is likely inflated by the Indian languages in our dataset (ben,gom,hin,tel), all of which average around 7 edits per source. 4 3 Experimental Setup 3.1 Models We evaluate 32 open-weight models that can be de- ployed on no more than two NVIDIA A100 80GB GPUs. Our model selection considers four dimen- sions: 1) translation-specific or general-purpose; 2) 0.8â122B in size; 3) English-focused or multilin- gual; 4) made in or targeting diverse locations. We list them under three groups: âą Pre-trained machine translation models (MT): nllb-200-distilled-600M,1.3B(Team et al., 2024) andmadlad400-3b,7b-mt (Kudugunta et al., 2023). âąLLM-based translation models (LLM+MT): Seed-X-PPO-7B(Cheng et al., 2025) andHy- MT2-1.8B,7B,30B-A3B (Zheng et al., 2026). âąLLMs:tiny-aya-global(Salamanca et al., 2026),aya-expanse-8b(Dang et al., 2024), SmolLM-1.7B-Instruct(Allal et al., 2024), emma-500-llama3.1-8b-mono,bi(Ji et al., 2026),Qwen3.5-0.8B,2B,4B,9B,27B,35B- A3B,122B-A10B-FP8(Qwen Team, 2026), Olmo-3-7B-Instruct(Olmo Team et al., 2025),gemma-3-1b,4b-it(Gemma Team,2025),Llama-3.1-8B-instruct, Llama-3.2-1B,3B-instruct,Llama-3.3- 70B-Instruct (Grattafiori et al., 2024),Phi- 4-mini-instruct(Abouelenin et al., 2025), Ministral-3-3B,8B-Instruct-2512(Liu et al., 2026), and EuroLLM-1.7B-Instruct, 9B-instruct-2512 (Martins et al., 2025). 3.2 Evaluation BLEU and chrF We compute BLEU (Papineni et al., 2002) and chrF (Popovi Ì c, 2015), two stan- dard automatic metrics for machine translation eval- uation. Specifically, we use spBLEU withflo- res200tokenizer and chrF with 0 word order, as implemented in sacrebleu (Post, 2018). This way, the metrics operate on subwords and characters, re- spectively, without complications due to tokeniza- tion for certain languages. â BLEU andâ chrF as measures of robustness Our primary research objective is not to discover the best translation models, but to quantify their robustness via results before and after localisation. Therefore, we define a performance discrepancy between the localised and original tests as: â BLEU = BLEU localised â BLEU original â chrF = chrF localised â chrF original Asâbecomes increasingly negative, the model ex- hibits greater performance degradation when trans- lating localised source texts. Importantly,âis largely invariant to translation qualityâa model that performs poorly or strongly on both test sets will receive a ââ 0 for its consistent behaviour. â BLEU andâ chrF To aggregate over multiple locales after localisation, we pair all locales with their language, forming pairsi = 1, 2,· ,nand compute a mean pairwiseâ metric applicable to both BLEU and chrF: â metric = 1 n n X i=1 (score localised,i â score original,i ) Likewise, the meanâ metric can be averaged across models but without requiring pairing. Locale-specific word recallStandard metrics re- veal overall quality (difference) but provide little focus on locale-specific content. Since both LLMs and human annotators were instructed to avoid changing the sentence structure, most word edits would be directly related to the effort of language- location grounding. We therefore introduce a sim- ple lexical analysis that looks at the translation of a localised input, its original FLORES reference, and the localised Cultivar reference: âąFalse positives (FP): unique terms in the origi- nal FLORES reference that appear in the trans- lation of a localised instance. These should no longer appear after localisation. âąFalse negatives (FN): unique terms in the lo- calised Cultivar reference that do not appear in the translation. Since the localisation process introduced such words, they should appear. A high false positive indicates that a model con- tinues to generate terms in the original FLORES despite their removal during localisation, suggest- ing memorization. Conversely, a high false nega- tive means that models fail to generate newly in- troduced locale-specific content. Nonetheless, as multiple valid translations may exist for a localised concept, false negatives should be interpreted as an upper bound on localisation failures rather than an exact error rate. 5 OriginalLocalised BLEUchrF r s Ï b r s Ï b acm_ArabIraq0.970.890.970.90 apc_ArabSyria0.970.880.980.91 ben_BengIndia0.980.920.980.92 bul_CyrlBulgaria0.930.800.940.83 cat_LatnSpain 0.970.880.980.91 ces_LatnCzechia0.980.890.980.90 cmn_HansChina0.950.840.960.85 cmn_HansSingapore 0.940.800.950.84 cmn_HansUK0.950.850.980.90 cmn_HansUS0.980.910.990.92 deu_LatnGermany0.980.900.990.93 deu_LatnSwitzerland 0.970.880.980.90 fil_LatnPhilippines0.970.920.980.93 gom_DevaIndia0.950.850.970.89 hin_DevaGermany0.960.870.970.89 hin_DevaIndia0.950.850.960.88 hin_DevaUS 0.980.920.970.90 jpn_JpanJapan0.940.820.950.83 khk_CyrlMongolia0.970.860.980.90 plt_LatnMadagascar0.980.920.990.94 pol_LatnPoland0.930.820.960.85 rus_CyrlRussia 0.960.870.970.88 slk_LatnSlovakia 0.970.890.970.89 spa_LatnSpain0.970.880.980.91 tel_TeluIndia0.960.900.970.90 tur_LatnTurkey0.970.870.980.90 yor_LatnNigeria0.990.920.990.96 Average0.980.910.980.89 Table 2: Spearmanâsr s and KendallâsÏ b for BLEU and chrF, measuring the alignment between rankings of 32 models based on the original FLORES and the localised Cultivar when translating into English. 4 Empirical Results and Analysis 4.1Do Cultivar and FLORES yield consistent model rankings? We use the 32 models to individually translate FLO- RES and Cultivar test sets into English. We then compute BLEU and chrF scores for the translations using their respective references. Finally, we com- pute Spearmanâs and Kendallâs rank correlation coefficients,r s andÏ b , between model rankings for each language or locale, as determined by BLEU and chrF respectively. Table 2 showsr s â„ 0.93,Ï b â„ 0.80for all transla- tion directions andr s â„ 0.98,Ï b â„ 0.89on average. These indicate that Cultivar and FLORES counter- parts induce highly consistent outcomes in ranking models. 4.2 Do models perform better or worse with localised source content? Our localised Cultivar demonstrates a more natural and aligned use case of source-to-English transla- tion, since for each locale, all source instances are localised to relevant locations. To study the modelsâ behaviour on the localised source, we report previ- ously introduced metric scores in Table 3: 1) mean pairwiseâ BLEU andâ chrF ; 2) BLEU and chrF on FLORES; 3) BLEU and chrF on Cultivar; and 4) false positive and false negative rates for unique word recall. It is worth noting that a mean pairwise âscore does not equal the difference between the averaged scores for FLORES and Cultivar, because one language can be paired with multiple locales. As Table 3 shows, for most models,â BLEU ranges betweenâ2and+2, andâ chrF ranges be- tweenâ2and+0.7, with averages ofâ0.49and â1.1, respectively. Specifically,SmolLM-1.7B- Instructacts as a ârandomâ baselineâas the model is not good at translation, reflected by its BLEU and chrFâwithâscores close to0. How- ever, two models at the bottom see significantly lowerâscores when the input is localised, with a > 8 BLEU drop and a > 5 chrF drop. Most modelsâ false negatives fall within 10â30%, with a few higher numbers associated with weaker translation models showing low BLEU and chrF scores, which is largely expected. If we treat these weaker models withFN>44% as random baselines, we can conclude that most models can adapt to the localised source to a certain extent. Model typeConsidering how models have been trained, MT-optimised ones rank lower byâ scores, including both pre-trained translation mod- els (NLLB,MADLAD) and LLMs further optimised for translation (Hy-MT2,Seed-X-PPO-7B). Even though some of these have a small magnitude, they are consistently clustered in the negative region. Model size We then plotâ BLEU and average â chrf against model size in Figure 2 using differ- ent marks and colours to represent model families. In general, we seeâ BLEU andâ chrf increase as the model gets larger for 6 out of 8 families. Two families do not follow this pattern:Hy-MT2slightly drops as it gets larger, andLlama 3.Xpeaks at 8B followed by a sharp decline at 70Bâinterpreted with reservations, becauseLlamaspans version 3.1â 3.3. Sinceâmeasures the performance difference between FLORES and our Cultivar, an in-family scaling ofâmeans that the BLEU scores on the two test sets diverge as models get larger. This could imply that FLORES may underestimate the true multilingual translation capabilities of stronger models. 6 TypeModel PairwiseFLORESCultivar â BLEU ââ chrF BLEUchrFBLEUchrFF PosF Neg LLMQwen3.5-35B-A3B1.700.2041.4664.3143.3464.860.00690.1186 LLMLlama-3.1-8B-instruct1.420.3135.3659.3937.5860.300.00760.1717 LLMMinistral-3-8B-Instruct-25121.40-0.0435.2560.1037.4060.750.00870.1690 LLMEuroLLM-9B-instruct-25121.36-0.0336.9859.3839.6160.630.00770.1753 LLMPhi-4-mini-instruct1.07-0.5329.2554.9631.5755.860.00920.2579 LLMgemma-3-4b-it 0.96-0.2435.9560.4737.6361.010.00800.1908 LLMQwen3.5-122B-A10B-FP80.92-0.1743.2665.7144.3865.800.00670.1039 LLMLlama-3.2-3B-instruct0.89-0.3532.2256.0034.3356.850.00940.2394 LLMQwen3.5-27B0.80-0.2542.7065.1443.7965.250.00690.1149 LLMOlmo-3-7B-Instruct0.61-0.4827.0553.4228.7554.110.01140.3205 LLMQwen3.5-9B 0.61-0.3739.2962.5140.3862.680.00770.1520 LLMQwen3.5-2B0.56-0.4027.4652.4628.5852.650.00980.2928 LLMMinistral-3-3B-Instruct-25120.52-0.3829.7255.7131.0756.200.01000.2291 LLMEuroLLM-1.7B-Instruct0.48-0.4327.8349.9930.4551.990.00820.2978 LLMQwen3.5-4B 0.42-0.6435.3359.8236.3259.780.00820.1916 LLMLlama-3.2-1B-instruct0.30-0.9321.8244.8923.2145.320.00980.4400 LLMgemma-3-1b-it0.27-0.7925.3550.8426.6351.190.00970.3675 LLMtiny-aya-global0.03-0.8832.9056.8833.3756.520.01060.2780 LLMSmolLM-1.7B-Instruct -0.020.073.9220.154.3821.830.00850.6930 LLMemma-500-llama3.1-8b-mono-0.03-0.9636.5859.6437.3259.250.00920.2155 LLMemma-500-llama3.1-8b-bi-0.10-0.7543.2065.1843.5764.830.00740.1569 MTmadlad400-7b-mt-0.53-1.3442.3864.3142.6763.780.00890.2001 LLMLlama-3.3-70B-Instruct-0.58-0.9844.1665.8444.4065.490.00730.1214 LLMQwen3.5-0.8B-0.70-1.4119.7043.8019.6642.830.01060.4403 MTnllb-200-distilled-1.3B-0.79-1.4341.4463.7640.7162.750.00940.2382 MTmadlad400-3b-mt-0.90-1.5040.8863.2540.8462.520.00920.2266 MTnllb-200-distilled-600M-1.21-2.0537.4060.9536.6059.510.00980.2924 LLM+MTHy-MT2-1.8B -2.06-2.0938.5461.4537.7760.680.00980.2243 LLM+MTHy-MT2-7B-2.16-2.0644.2965.5442.9864.380.00920.1661 LLM+MTHy-MT2-30B-A3B-2.28-2.2147.4668.1745.7766.580.00660.1217 LLMaya-expanse-8b-8.87-5.8450.5167.3744.5063.810.00970.2115 LLM+MTSeed-X-PPO-7B-9.68-6.3253.1871.9945.6467.190.01220.1665 Average-0.49-1.1036.6659.4336.1058.790.00890.2370 Table 3: Aggregated results for each model: mean pairwiseâscores; FLORES scores averaged across languages; Cultivar scores averaged across locales; and locale-specific word recalls averaged across locales. 4.3 Are models FLORES contaminated? Looking at BLEU and chrF scores,aya-expanse- 8bandSeed-X-PPO-7Bare strong on both FLORES and Cultivar tests, beating much larger general-purpose LLMs likeLlama-3.3-70B- InstructandQwen3.5-122B-A10B-FP8. How- ever, they exhibit a much larger performance de- cline (â BLEU < â8) than other models (â2 < â BLEU < +2 ) when moving from FLORES to Cul- tivar. A further score breakdown in Section C Fig- ure 6 shows that, for the two models, a substan- tial BLEU or chrF drop occurs in most source-to- English directions, implying a model artifact. We further test for text memorizationâwhether verbatim FLORES English words are undesirably recalled when translating counterpart Cultivar in- stances. We probe this via false positives as defined in Section 3.2 and reported in Table 3. Most mod- els fall below 1%. Looking at the models at the bottom of the table with negativeâscores, they do not record significantly higher false positives com- pared to models at the top. Hence, there is no direct evidence that these models have memorised the ex- act FLORES text. Nonetheless, disproportionate drops in BLEU and chrF could hint at a lack of robustness to localisation or some degree of overfit- ting to FLORES, since localisation still retains the original sentence structure. This is corroborated by an error analysis in Section 5.2 later. 4.4 Which locales are harder? Table 4 listsâ BLEU andâ chrF scores again, but this time averaged across models for each locale. Most â BLEU range fromâ2to+2, and most chrF range fromâ2to+1. According to BLEU differences, models perform better on localisedplt_Latn_- Madagascarandgom_Deva_India; according to chrF differences, models perform better onhin_- Deva_US ,cmn_Hans_US, andplt_Latn_Madagas- car . At the bottom of the table,cmn_Hans_Sin- 7 110100 â2 â1 0 1 2 Llama 3.X Qwen 3.5 Gemma 3 Ministral EuroLLM MADLAD Hy-MT2 NLLB Total Parameters (B) â BLEU 110100 â2 â1.5 â1 â0.5 0 0.5 Llama 3.X Qwen 3.5 Gemma 3 Ministral EuroLLM MADLAD Hy-MT2 NLLB Total Parameters (B) â chrF Llama 3.XQwen 3.5Gemma 3Ministral EuroLLM MADLADHy-MT2 NLLB Figure 2: Plots of â BLEU (left) and â chrf (right) averaged across locales against model size (logarithmic scale). US IN DE UK CN SG â4 â2 0 â BLEU by Location Hindi Chinese US IN DE UK CN SG â2 â1 0 1 â chrF by Location Figure 3: Plots ofâ BLEU andâ chrF averaged across models for translating localised Hindi and Chinese to English. gaporereceives significantly lowâ, and another five localised test sets also haveâ BLEU andâ chrF aroundâ2. These results demonstrate that the difficulty of a test set could change after being localisedâsome appear easier while others become harder. Nonethe- less, it is noteworthy that localisations were pro- duced by an LLM and edited by individual annota- tors, who may bring in their personal biases, so the difficulty of individual locales could be attributed to a combination of model artifact, e.g. training data, as well as data artifact relative to the original FLORES instances. 4.5 Same language, different localisations In our test set, Hindi and Chinese have been lo- calised to locations where they are not the official or de facto language, allowing us to conduct a small- scale ablation study on how models respond to lo- calisation while controlling the language. In detail, Hindi has been localised to the US and Germany in addition to India; Chinese has been localised to the US and UK in addition to China and Singapore. In Figure 3, we visualise theâ BLEU andâ chrF for both languages, averaged across models but separated by location. We see negativeâfor Ger- many, China, and Singapore, mixed results for In- dia, and positiveâs for the US. Moreover, since the âscores are relative to FLORES, we can compare the performance between two localisations directly by comparingâs. For example, when translating from Chinese to English, by definition: BLEU cmn_US â BLEU cmn_SG = â BLEU cmn_US â â BLEU cmn_SG . Apositiveâ BLEU cmn_US andanegative â BLEU cmn_SG in Figure 3 suggests that the difference betweenBLEU cmn_US andBLEU cmn_SG is significant. The same applies to chrF scores too. Overall, the results indicate that, relative to the original FLORES, models translate US content bet- ter in both Hindi and Chinese, but translate content about China and Singapore worse in Chinese. The positiveâfor the US implies that, compared to FLORES, Cultivarâs US localisation is easier for models, even though the source languages, Hindi and Chinese, are not de facto in the US. On the other hand, the negativeâs are not ideal, because they indicate that models are less capable of trans- lating source-localised content into English, which 8 LangScriptLocationâ BLEU ââ chrF pltLatnMadagascar2.530.86 gomDevaIndia2.05-0.06 yorLatnNigeria 1.480.09 hinDevaUS1.381.36 telTeluIndia1.14-0.27 hinDevaIndia1.00-0.76 benBengIndia 0.96-0.38 cmnHansUS0.921.11 polLatnPoland0.38-1.89 spaLatnSpain0.05-0.76 deuLatnGermany 0.02-1.25 bulCyrlBulgaria-0.21-0.65 hinDevaGermany-0.28-1.31 slkLatnSlovakia-0.30-1.55 filLatnPhilippines-0.68-0.57 turLatnTurkey -0.96-1.91 cmnHansUK-1.03-0.52 catLatnSpain-1.24-1.26 cesLatnCzechia-1.30-2.16 cmnHansChina -1.51-1.06 rusCyrlRussia-1.70-1.80 deuLatnSwitzerland-1.97-2.31 acmArabIraq-2.06-2.32 khkCyrlMongolia -2.13-2.88 jpnJpanJapan -2.65-2.81 apcArabSyria -2.77-2.49 cmnHansSingapore-4.29-2.22 Average-0.27-1.05 Table 4: Average â BLEU and â chrF for each locale. is the main use case of source-to-English trans- lation. From a model building perspective, this contrastive evaluation reflects a US-centric bias in training data, and Cultivar has provided a system- atic way to probe such locale-centric bias. 5 Manual Error Analysis 5.1 Setup In addition to analysis based on automatic evalu- ation results, we conduct a manual inspection of translations of regionally localised content. In par- ticular, a Chinese native speaker who is fluent in English inspected the Chinese-to-English outputs from selected models together with sources and references. Nine error types are pre-defined and grouped into two categories: âąFour locale-specific errors: named entities, id- ioms and culturally specific expressions, number or unit, as well as Pinyin, a deterministic pho- netic conversion system for transcribing Man- darin Chinese names into the Latin alphabet. âąFive general errors: semantic mistranslation, omitted information, hallucinated or added in- formation, original language retained, and un- wanted reasoning or formatting tokens. We classify errors based on their manifestation rather than underlying cause, with errors directly concerning localised content categorised as locale- specific. An observed error may nevertheless arise from either locale-specific or general difficulties. Repeated errors of the same type within an instance are counted only once, while different types are counted separately. 5.2 Observations corroborate earlier findings Table 5 presents the results of our human error analysis, and we also include BLEU for reference. Across the three regional variants, Chinese content grounded in the US yields the fewest errors, with this difference becoming more pronounced among smaller models with higher error rates. Across models, we observe a clear scaling trend: locale-specific (and general) errors decrease as the modelâs total parameter count increases, while the number of active parameters plays a minor role. This pattern is evident in the Qwen3.5 mixture-of- experts models: both35B-A3Band122B-A10Bex- hibit substantially fewer errors than the2B,4B,9B dense models, despite activating a smaller or com- parable number of parameters. We observe two interesting patterns forSeed- X-PPO-7B: 1) Its US split BLEU is much higher than the UK and China splits, whereas BLEU dif- ferences across locations are smaller for the other models. 2) We find a moderate-to-strong negative correlation between BLEU and total error count for Chinese, withSeed-X-PPO-7Bbeing an outlierâ Pearsonâsr incl =â0.64andr excl = â0.86when including and excluding it. Despite having an er- ror count comparable toQwen3.5-9Band higher than larger models, it obtains a substantially higher BLEU than all other models in Table 5. This mis- match between error counts and BLEU hints at the possibility of FLORES-like data contamination, whereby overfitting to domain, style, or sentence structure may inflate automatic scores despite the presence of errors. Broadly, our observed error patterns in model size and locale difficulty, despite a small scale, are consistent with, and provide qualitative support for, our earlier findings based on BLEU, chrF, andâ scores. 5.3 Locale-specific error patterns Locale-specific errors substantially outnumber gen- eral translation errors. In aggregate, named entities 9 Model Chinese_CNâEnglishChinese_UKâEnglishChinese_USâEnglish Locale General Total BLEU Locale General Total BLEU Locale General Total BLEU NE PY IC NUNE PY IC NUNE PY IC NU Seed-X-PPO-7B871142141.871200182144.941301141947.31 Llama-3.1-8B-Instruct15712123733.4523015124134.001600172435.36 Ministral-3-8B-Instruct-2512 27723115032.3224003144130.921800172633.63 EuroLLM-9B-Instruct-2512 16102253537.8116002123037.221300051838.82 Qwen3.5-2B33734136032.7939005155931.8923002204534.02 Qwen3.5-4B1260172635.5924002103635.411700062335.74 Qwen3.5-9B1031072137.321400152037.23900041338.10 Qwen3.5-27B 42002838.30900221338.0150004939.44 Qwen3.5-35B-A3B530021038.00600141137.29800031138.28 Qwen3.5-122B-A10B-FP832013938.4840003737.59900021138.63 Total133 54 10 1466 277171 0 1 2285 279131 0 1 562 199 Table 5: Error counts and BLEU by model and country. We list locale-specific, general, and total errors separately. Locale-specific errors include named entities (NE), Pinyin (PY), idiom and culture (IC), and number and unit (NU). (NE) errors are the most frequent, accounting for > 80%of locale-specific errors and> 50%of all recorded errors. This pattern is consistent across models and locations, indicating that the accurate translation of localised entities remains a primary source of errors even when the general translation is successful. The remaining locale-specific error types exhibit more pronounced regional variation. Errors involv- ing idioms and culturally specific expressions (IC) are more frequent in the China split. As expected, Pinyin (PY) transcription errors are exclusive to China, making up 20% of all errors. These errors arise when Chinese names require romanisation rather than translation. The asymmetry with the UK and US splits is consequential: names in those splits originate from English and are represented in Chinese before being translated back into En- glish, whereas the China split requires the hard-to- generalise Chinese-to-Latin script mapping ability. Thus, the China split presents a unique difficulty of script- and language-specific entity transformation that is absent from other localisations. Number and unit errors (NU) also show substan- tial regional variation. The UK and China splits contain 22 and 14 such errors, respectively, com- pared with only 5 in the US split. These errors involve region-specific measurements and format- ting, such as jin (weight), British pounds (weight and currency), and chronological conversions (cen- tury and decade to year). The concentration of such errors in the China and UK splits suggests that models are more likely to struggle when the source text contains numbers or units that differ from those in general-purpose translation data. Overall, the error analysis indicates that locali- sation affects the difficulty of translation, reflected in the nature of the errors that models make, which vary for different locales. This underscores the value of a locale-oriented test set with a unique source of locale-specific difficulties. 6 Conclusion This work introduces Cultivar, a source-contrastive, locale-oriented test set that complements the cur- rent massive multilingual translation evaluation. Cultivar isolates a modelâs robustness to cultural content from general translation capability and serves as a diagnostic tool for data contamination. Evaluating 32 models reveals that smaller and MT- specialised systems perform worse on localised content. Furthermore, our contrastive design ex- poses potential FLORES overfitting and a perva- sive US-centric bias, where models translate US- grounded content better than content from native regions. Ultimately, Cultivar demonstrates that true multilingual proficiency requires locale-aware eval- uation to accurately reflect real-world use cases. Acknowledgements AI assistants have been used for coding, visual- isations, and paper editing. Conceptualisation, methodology, experiments and analysis, and writ- ing were done by the authors. We thank Teresa Sy Ortin and Bouazza Laracha for data contribution. We are grateful for the com- puting resources from the Northern Ireland High Performance Computing (NI-HPC) service funded by EPSRC (EP/T022175). 10 References Abdelrahman Abouelenin, Atabak Ashfaq, Adam Atkin- son, Hany Awadalla, Nguyen Bach, Jianmin Bao, Alon Benhaim, Martin Cai, Vishrav Chaudhary, Con- gcong Chen, and others. 2025. Phi-4-mini techni- cal report: Compact yet powerful multimodal lan- guage models via mixture-of-loras. arXiv preprint arXiv:2503.01743. Md Mahfuz Ibn Alam, Sina Ahmadi, and Antonios Anastasopoulos. 2024. CODET: A benchmark for contrastive dialectal evaluation of machine transla- tion. In Findings of the Association for Computa- tional Linguistics: EACL 2024. Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Le- andro von Werra, and Thomas Wolf. 2024. SmolLM - blazingly fast and remarkably powerful. Pierre Andrews, Mikel Artetxe, Mariano Coria Megli- oli, Marta R. Costa-jussĂ , Joe Chuang, David Dale, Mark Duppenthaler, Nathanial Paul Ekberg, Cyn- thia Gao, Daniel Edward Licht, Jean Maillard, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Eduardo SĂĄnchez, Ioannis Tsiamas, Arina Turkatenko, Albert Ventayol-Boada, and Shireen Yates. 2025. BOUQuET : dataset, benchmark and open initiative for universal quality evaluation in translation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Process- ing. LoĂŻc Barrault, Ond Ë rej Bojar, Marta R. Costa-jussĂ , Christian Federmann, Mark Fishel, Yvette Gra- ham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias MĂŒller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. Findings of the 2019 conference on machine trans- lation (WMT19). In Proceedings of the Fourth Con- ference on Machine Translation (Volume 2: Shared Task Papers, Day 1). Rachel Bawden, Rico Sennrich, Alexandra Birch, and Barry Haddow. 2018. Evaluating discourse phenom- ena in neural machine translation. In Proceedings of the 2018 Conference of the North American Chap- ter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Pa- pers). Ond Ë rej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint- Amand, Radu Soricut, Lucia Specia, and AleĆĄ Tam- chyna. 2014. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the Ninth Workshop on Statistical Machine Translation. Mauro Cettolo, Jan Niehues, Sebastian StĂŒker, Luisa Bentivogli, and Marcello Federico. 2014. Report on the 11th IWSLT evaluation campaign. In Proceed- ings of the 11th International Workshop on Spoken Language Translation: Evaluation Campaign. Pinzhen Chen, Simon Yu, Zhicheng Guo, and Barry Haddow. 2024. Is it good data for multilingual in- struction tuning or just bad multilingual evaluation for large language models? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Shanbo Cheng, Yu Bao, Qian Cao, Luyang Huang, Liyan Kang, Zhicheng Liu, Yu Lu, Wenhao Zhu, Jing- wen Chen, Zhichao Huang, and others. 2025. Seed- X: Building strong multilingual translation LLM with 7B parameters. arXiv preprint arXiv:2507.13618. Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. TyDi QA: A bench- mark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8. David Dale, Laurie Burchell, Jean Maillard, Idris Ab- dulmumin, Antonios Anastasopoulos, Isaac Caswell, and Philipp Koehn. 2025. Findings of the WMT 2025 shared task of the open language data initiative. In Proceedings of the Tenth Conference on Machine Translation. John Dang, Shivalika Singh, Daniel Dâsouza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, and others. 2024. Aya Expanse: Com- bining research breakthroughs for a new multilingual frontier. arXiv preprint arXiv:2412.04261. Daniel Deutsch, Eleftheria Briakou, Isaac Rayburn Caswell, Mara Finkelstein, Rebecca Galor, Juraj Juraska, Geza Kovacs, Alison Lui, Ricardo Rei, Ja- son Riesa, Shruti Rijhwani, Parker Riley, Elizabeth Salesky, Firas Trabelsi, Stephanie Winkler, Biao Zhang, and Markus Freitag. 2025. WMT24++: Ex- panding the language coverage of WMT24 to 55 languages & dialects. In Findings of the Association for Computational Linguistics: ACL 2025. Christian Federmann, Tom Kocmi, and Ying Xin. 2022. NTREX-128 â news test references for MT evalua- tion of 128 languages. In Proceedings of the First Workshop on Scaling Up Multilingual Evaluation. Matthieu Futeral, Cordelia Schmid, Ivan Laptev, BenoĂźt Sagot, and Rachel Bawden. 2023. Tackling ambi- guity with images: Improved multimodal machine translation and contrastive evaluation. In Proceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Gemma Team. 2025. Gemma 3. Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng- Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Kr- ishnan, MarcâAurelio Ranzato, Francisco GuzmĂĄn, and Angela Fan. 2022. The Flores-101 evaluation benchmark for low-resource and multilingual ma- chine translation. Transactions of the Association for Computational Linguistics, 10. 11 Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al- Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and others. 2024. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783. Matthew Honnibal, Ines Montani, Sofie Van Lan- deghem, and Adriane Boyd. 2020. spaCy: Industrial- strength natural language processing in Python. Shaoxiong Ji, Zihao Li, Jaakko Paavola, Hengyu Luo, and Jörg Tiedemann. 2026. Data-centric continual pre-training for 500+ languages: A new bilingual translation corpus and multilingual models. In Find- ings of the Association for Computational Linguistics: ACL 2026. Sneha Kudugunta, Isaac Rayburn Caswell, Biao Zhang, Xavier Garcia, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2023. MADLAD-400: A multilingual and document-level large audited dataset. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Alexander H Liu, Kartik Khandelwal, Sandeep Sub- ramanian, Victor Jouault, Abhinav Rastogi, Adrien SadĂ©, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, and others. 2026. Ministral 3. arXiv preprint arXiv:2601.08584. Inbal Magar and Roy Schwartz. 2022. Data contamina- tion: From memorization to exploitation. In Proceed- ings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Pedro Henrique Martins, Patrick Fernandes, JoĂŁo Alves, Nuno M Guerreiro, Ricardo Rei, Duarte M Alves, JosĂ© Pombal, Amin Farajian, Manuel Faysse, Ma- teusz Klimaszewski, and others. 2025. EuroLLM: Multilingual language models for Europe. Procedia Computer Science, 255:53â62. Olmo Team, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groen- eveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, and others. 2025. Olmo 3. arXiv preprint arXiv:2512.13961. Kishore Papineni, Salim Roukos, Todd Ward, and Wei- Jing Zhu. 2002. Bleu: a method for automatic evalu- ation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Compu- tational Linguistics. Maja Popovi Ì c. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation. Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers. Qwen Team. 2026. Qwen3.5: Towards native multi- modal agents. Oscar Sainz, Jon Campos, Iker GarcĂa-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. 2023. NLP evaluation in trouble: On the need to mea- sure LLM data contamination for each benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2023. Alejandro R Salamanca, Diana Abagyan, Daniel Dâsouza, Ammar Khairi, David Mora, Saurabh Dash, Viraat Aryabumi, Sara Rajaee, Mehrnaz Mofakhami, Ananya Sahu, and others. 2026. Tiny Aya: Bridg- ing scale and multilingual depth. arXiv preprint arXiv:2603.11510. David Tan, Pinzhen Chen, Josef van Genabith, and Koel Dutta Chowdhury. 2026. When Flores bloomz wrong: Cross-direction contamination in machine translation evaluation. In Proceedings of the 19th Conference of the European Chapter of the Associa- tion for Computational Linguistics (Volume 2: Short Papers). NLLB Team and others. 2024. Scaling neural machine translation to 200 languages. Nature, 630(8018):841â 846. Feng Yao, Yufan Zhuang, Zihao Sun, Sunan Xu, Ani- mesh Kumar, and Jingbo Shang. 2024. Data contam- ination can cross language barriers. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Mao Zheng, Zheng Li, Tao Chen, Bo Lv, Mingrui Sun, Mingyang Song, Jinlong Song, Hong Huang, Decheng Wu, Hai Wang, and others. 2026. Hy-MT2: A family of fast, efficient and powerful multilin- gual translation models in the wild. arXiv preprint arXiv:2605.22064. 12 A LLM Localisation Prompts System Prompt: I will provide a pair ofsrc_name-tgt_nameparallel sentences. Your task is to localize both sentences to src_locationby rephrasing the named entities. This includes but is not limited to culturally, historically, geographically, and politically relevant entities. The rephrasedsrc_nameandtgt_namesentences must remain parallel. Maintain the original grammatical structure, tone, and sentence length as closely as the new entities allow. Output in the exact same format without any formatting or explanatory text: src_name: ... tgt_name: ... User Prompt: src_name: src_text tgt_name: tgt_text Figure 4: System and user prompts used for LLM localisation B Instructions Presented to Annotators Background and Motivation âą FLORES+ is an MT evaluation set, translated from English into hundreds of languages, resulting in multiway parallelism. However, this data creation process made its content English-centric. In practice, when translating non-English content, the source content would rarely be English-native. For example, as a Hindi speaker in India, you wonât need to translate "the Stanford University School of Medicine announced [...]" into English. âą Therefore, this project intends to transcreate a variant of FLORES+ by replacing named entities in the original non- English entries to make each test direction grounded to the source language, culturally, historically, geographically, politically, etc. Annotation Instructions âą For each review item, you will see: ⊠FLORES+âs original source and English (as target). ⊠automatically rephrased source and English, based on the original instances. âąYour task is to review the rephrased texts and either accept them as they are or edit them to make them better. Rephrases do not need to be factual (and many are not), but should be grounded, reasonable, and plausible. When editing, please try your best to only edit the named entities and maintain the original sentence structure in both languages, even if you think they are not ideal. âą Checklist: ⊠Are the replaced named entities reasonable? Do they ground the translations in the source language and its region, with minimal changes to sentence structure? ⊠Are the rephrased source and English sentences still natural and parallel? âą Your action for each instance can be either: ⊠Accept the current rephrased pair as is and tick Confirm final version. ⊠Edit one or often both sentences, then tick Confirm final version. ⊠In rare cases, you can delete the rephrased sentences and copy the original sentences to the text box. As a result, this instance will remain unchanged at all. Then tick Confirm final version. This should be used sparingly when 1) the original sentences are already localized in the language/country; 2) if the rephrased sentences are completely off and you cannot come up with a good version/edit either (e.g. this particular instance isnât suitable for localization). âą Your progress should be autosaved in your browser, but itâs not sent back to the server. To be safe, you can save a partly done annotation by clicking Download JSON; you can load it via Load JSON to continue the review. Once you are done with the entire annotation, download the JSON file and send it back. Figure 5: Annotation instruction given to human annotators 13 Câ Breakdown by Model and Locale plt_Latn_Madagascar gom_Deva_India yor_Latn_Nigeria hin_Deva_US tel_Telu_India hin_Deva_India ben_Beng_India cmn_Hans_US pol_Latn_Poland spa_Latn_Spain deu_Latn_Germany bul_Cyrl_Bulgaria hin_Deva_Germany slk_Latn_Slovakia fil_Latn_Philippines tur_Latn_Turkey cmn_Hans_UK cat_Latn_Spain ces_Latn_Czechia cmn_Hans_China rus_Cyrl_Russia deu_Latn_Switzerland acm_Arab_Iraq khk_Cyrl_Mongolia jpn_Jpan_Japan apc_Arab_Syria cmn_Hans_Singapore Qwen3.5-35B-A3B Llama-3.1-8B-instruct Ministral-3-8B-Instruct-2512 EuroLLM-9B-instruct-2512 Phi-4-mini-instruct gemma-3-4b-it Qwen3.5-122B-A10B-FP8 Llama-3.2-3B-instruct Qwen3.5-27B Olmo-3-7B-Instruct Qwen3.5-9B Qwen3.5-2B Ministral-3-3B-Instruct-2512 EuroLLM-1.7B-Instruct Qwen3.5-4B Llama-3.2-1B-instruct gemma-3-1b-it tiny-aya-global SmolLM-1.7B-Instruct emma-500-llama3.1-8b-mono emma-500-llama3.1-8b-bi madlad400-7b-mt Llama-3.3-70B-Instruct Qwen3.5-0.8B nllb-200-distilled-1.3B madlad400-3b-mt nllb-200-distilled-600M Hy-MT2-1.8B Hy-MT2-7B Hy-MT2-30B-A3B aya-expanse-8b Seed-X-PPO-7B â15â10â505 â BLEU hin_Deva_US cmn_Hans_US plt_Latn_Madagascar yor_Latn_Nigeria gom_Deva_India tel_Telu_India ben_Beng_India cmn_Hans_UK fil_Latn_Philippines bul_Cyrl_Bulgaria hin_Deva_India spa_Latn_Spain cmn_Hans_China deu_Latn_Germany cat_Latn_Spain hin_Deva_Germany slk_Latn_Slovakia rus_Cyrl_Russia pol_Latn_Poland tur_Latn_Turkey ces_Latn_Czechia cmn_Hans_Singapore deu_Latn_Switzerland acm_Arab_Iraq apc_Arab_Syria jpn_Jpan_Japan khk_Cyrl_Mongolia Llama-3.1-8B-instruct Qwen3.5-35B-A3B SmolLM-1.7B-Instruct EuroLLM-9B-instruct-2512 Ministral-3-8B-Instruct-2512 Qwen3.5-122B-A10B-FP8 gemma-3-4b-it Qwen3.5-27B Llama-3.2-3B-instruct Qwen3.5-9B Ministral-3-3B-Instruct-2512 Qwen3.5-2B EuroLLM-1.7B-Instruct Olmo-3-7B-Instruct Phi-4-mini-instruct Qwen3.5-4B emma-500-llama3.1-8b-bi gemma-3-1b-it tiny-aya-global Llama-3.2-1B-instruct emma-500-llama3.1-8b-mono Llama-3.3-70B-Instruct madlad400-7b-mt Qwen3.5-0.8B nllb-200-distilled-1.3B madlad400-3b-mt nllb-200-distilled-600M Hy-MT2-7B Hy-MT2-1.8B Hy-MT2-30B-A3B aya-expanse-8b Seed-X-PPO-7B â10â8â6â4â2024 â chrF Figure 6: Heatmaps ofâ BLEU (top) andâ chrF (bottom). Models (y-axis) are sorted by meanâacross locales, and locales (x-axis) are sorted by mean â across models, with more negative â towards the bottom and right. 14