Paper deep dive
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
Daryna Dementieva, Nikolay Babakov, Kathy Hämmerl, Ilseyar Alimova, Jindřich Libovický, Shu Okabe, Miras Baisbay, Lukas Edman, Abrorkhon Inomkhujaev, Antonia Karamolegkou, Mateusz Lango, Volkan Özer, Nikola Selic, Subhankar Swain, Tsedeniya Kinfe Temesgen, Galit Bary Weisberg, Alexander Fraser
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/8/2026, 4:28:04 AM
Summary
This paper introduces PLURAMATH, a multilingual mathematical reasoning benchmark that extends the PolyMath dataset to 18 underrepresented languages across 6 language families. The authors construct the dataset using a rigorous human-curated translation pipeline and benchmark 27 reasoning LLMs across varying scales. Their analysis reveals a persistent performance gap between high-resource and underrepresented languages, with stronger results primarily correlating with models' instruction-following abilities rather than translation capabilities. The dataset, pipeline, and evaluation framework are fully open-sourced to lower barriers for multilingual benchmark development.
Entities (10)
Relation Signals (7)
PLURAMATH → extends → PolyMath
confidence 95% · PLURAMATH, an extension of PolyMath to 18 additional underrepresented languages
18 Underrepresented Languages → belongto → 6 Language Families
confidence 94% · 18 additional underrepresented languages spanning 6 language families
PLURAMATH → benchmarks → 27 Reasoning LLMs
confidence 92% · Using PLURAMATH, we then benchmark 27 reasoning LLMs across four model scales
Technical University of Munich → affiliatedwith → PLURAMATH
confidence 90% · Daryna Dementieva 1,2 ... 1 Technical University of Munich (TUM)
PLURAMATH → usesmetric → Difficulty-Weighted Accuracy
confidence 89% · aggregated score per language overall, we used the difficulty-weighted accuracy across all levels
Mathematical Reasoning → correlateswith → Instruction-Following Ability
confidence 88% · stronger results largely associated with better instruction-following ability
27 Reasoning LLMs → evaluatedwith → Human-Curated Translation Pipeline
confidence 85% · constructed the dataset through a human-curated pipeline... benchmark 27 reasoning LLMs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales -- small, mid-size, large, and closed-source ensembles -- probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.
Tags
Links
- Source: https://arxiv.org/abs/2607.05992v1
- Canonical: https://arxiv.org/abs/2607.05992v1
Trouble viewing inline? Open PDF directly →
Full Text
169,410 characters extracted from source content.
Expand or collapse full text
PLURAMATH: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages Daryna Dementieva 1,2 * Nikolay Babakov 4 Kathy Hämmerl 1,2 Ilseyar Alimova 5 Jind ˇ rich Libovický 6 Shu Okabe 1,2 Miras Baisbay 7 Lukas Edman 1,2 Abrorkhon Inomkhujaev 1 Antonia Karamolegkou 8 Mateusz Lango 6 Volkan Özer 1,2 Nikola Selic 1 Subhankar Swain 9 Tsedeniya Kinfe Temesgen 1,2 Galit Bary Weisberg 10 Alexander Fraser 1,2,3 1 Technical University of Munich (TUM) 2 Munich Center for Machine Learning (MCML) 3 Munich Data Science Institute (MDSI) 4 Independent Researcher 5 Applied AI Institute 6 Charles University 7 Nazarbayev University 8 Inria 9 Indian Institute of Technology, Kharagpur (IIT Kharagpur) 10 German University of Digital Science Abstract Mathematical reasoning has become a cen- tral task for evaluating and tuning reasoning Large Language Models (LLMs), yet exist- ing benchmarks remain heavily biased toward high-resource languages, with English and Chi- nese dominating both pre-training corpora and evaluation suites. The recently released Poly- Math (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PLURAMATH, an extension of PolyMath to 18 additional un- derrepresented languages spanning 6 language families—ranging from mid-resource to ex- treme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre- computed translations. Using PLURAMATH, we then benchmark 27 reasoning LLMs across four model scales—small, mid-size, large, and closed-source ensembles—probing the multi- lingual mathematical reasoning capabilities of state-of-the-art models under diverse linguis- tic conditions. Our fine-grained analysis con- firms a persistent gap in mathematical reason- ing performance between high-resource and underrepresented languages, with stronger re- sults largely associated with better instruction- following ability. We fully open-source our dataset, data acquisition pipeline, and evalua- tion framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities. 1 Introduction Both the pretraining and the evaluation of large language models (LLMs) remain heavily skewed towards English and set of high-resource languages (Joshi et al., 2020a; Ghosh et al., 2025). Mathemat- ical reasoning is no exception: the popular bench- marks are either English-only—GSM8K (Cobbe * Correspondence: daryna.dementieva@tum.de Figure 1: Plura—from Latin, “more”: our extension of PolyMath (Wang et al., 2025) 4-level mathemati- cal reasoning benchmark to 18 underrepresented lan- guages through comprehensive human annotation of pre-computed translations. et al., 2021), AIME, MathArena (Balunovic et al., 2025), OMEGA (Sun et al., 2025)—or Chinese- only (Wei et al., 2023; Zhong et al., 2024). Re- cent multilingual efforts such as PolyMath (Wang et al., 2025), MGSM8KInstruct (Chen et al., 2024), and MathNet (Alshammari et al., 2026) broaden coverage but still concentrate on the high-resource languages, leaving underrepresented ones untested. At the same time, a growing amount of recent work shows that mathematical reasoning ability does not transfer easily across languages, even with translation or pivot-language prompting. For in- stance, models exhibit reasoning-answer misalign- ment that is more severe in non-Latin scripts as in Latin scripts (Ovalle et al., 2025); latent reasoning pathways remain English-centred even across typo- logically diverse languages (Liu et al., 2026). Thus, truly democratized for multiple languages LLMs development progress therefore requires evaluation benchmarks that stress-test reasoning models in genuinely diverse linguistic settings, including the long tail. arXiv:2607.05992v1 [cs.CL] 7 Jul 2026 CodeLanguage L1Lang Language Family MathTrans.Source## Spk. (M)ClassInstr.ModelLanguageAnn.Rev. hiHindi600.04IE – Indo-Aryan✓SarvamaiEnglish1<10 trTurkish 80.04Turkic✓GeminiEnglish2179 plPolish 45.04IE – Slavic✓DeepLEnglish1427 ukUkrainian40.03IE – Slavic✓DeepLRussian214 uzUzbek35.03TurkicDeepLEnglish1500 orOdia35.01IE – Indo-Aryan✓SarvamaiEnglish1<10 amAmharic32.02A – Semitic✓GeminiEnglish1126 elGreek 13.03IE – Hellenic✓DeepLEnglish227 kkKazakh 13.03Turkic✓DeepLRussian2500 csCzech 10.04IE – Slavic✓DeepLGerman1173 heHebrew9.03A – Semitic✓DeepLEnglish234 srSerbian8.24IE – Slavic✓DeepLEnglish1370 ttTatar5.51Turkic✓DeepLRussian2390 skSlovak5.03IE – Slavic✓DeepLEnglish1305 caCatalan 4.04IE – Romance✓SalamandraTASpanish1222 cvChuvash 1.01TurkicGeminiRussian1496 hsbUpper Sorbian 0.0131IE – SlavicTartuNLPGerman1379 dsbLower Sorbian0.0071IE – SlavicTartuNLPGerman1440 Number of L1 speakersLanguage class (Joshi et al., 2020b)Language family > 100M4 Underdogs: much unlabeled, less labeledIE – Slavic 10–100M3 Rising Stars: strong web, few labelsIE – Indo-Aryan 1–10M2 Hopefuls: some labeled data, supportIE – Hellenic < 1M1 Scraping-Bys: little data, few labelsIE – Romance Turkic A – Semitic Table 1: Overview of languages in PLURAMATH. L1 speaker counts (in millions) are approximate, sourced from Ethnologue. Lang Class follows the resource taxonomy of Joshi et al. (2020b) (1–4).✓indicates the language is used for mathematics teaching in a region;indicates if not for the university level. IE = Indo-European; A = Afro-Asiatic. We cover a diversity of underrepresented languages including quite low-resource ones. We address this gap by extending PolyMath (Wang et al., 2025) to 18 underrepresented lan- guages from 6 language families. Crucially, rather than relying on machine-translation-only pipelines or with LLM-based post-editing, we conduct full human rigorous check of pre-computed translations by native speakers for our benchmark. Our main contributions are as follows: • We release PLURAMATH, a novel multilin- gual mathematical reasoning benchmark that extends PolyMath to 18 underrepresented lan- guages from 6 language families (Table 1). •We open-source the full data acquisition and validation pipeline, including annotators guidelines and quality-control procedures. • We then evaluate 27 modern reasoning models under three prompting setups, spanning small, mid-sized, large open-weight, and closed- source systems, to assess multilingual math- ematical reasoning across diverse linguistic conditions and compare performance against high-resource base languages. •We conduct a comprehensive analysis of mod- els’ performance, reasoning behavior, and error patterns, identifying correlations with models’ families, language categories, and translation-related capabilities, complemented by human evaluation. Our findings confirm a persistent gap in math- ematical reasoning performance between high- resource and underrepresented languages.Al- though reasoning behavior varies across models, including reasoning in either English or the target language, performance gains are only weakly asso- ciated with translation capability and are not sub- stantially improved by prompting in high-resource languages. Instead, stronger results primarily cor- relate with general instruction-following ability, with larger models consistently achieving better performance with shorter reasoning. The dataset, prompts, analysis code, and annotation instructions are publicly released. 1,2,3 We hope that our ob- tained dataset, extension methodology, and error analysis will support further progress on multilin- gual mathematical reasoning. 1 Website: tum-nlp.github.io/pluramath 2 03.07.26, 17:32 Page 1 of 1https://huggingface.co/front/assets/huggingface_logo-noborder.svg PLURAMATH: tum-nlp/PluraMath 3 § Code and details: TUM-NLP/pluramath 2 Related Work 2.1 Mathematical Reasoning Benchmarks Early benchmarks for evaluating mathematical rea- soning in language models have been overwhelm- ingly English-centric.GSM8K (Cobbe et al., 2021) established the standard for grade-school word problems, while OpenWebMath (Paster et al., 2024) curated 14.7B tokens of English mathemat- ical web text. Beyond grade school, English-only resources have probed competition-level mathemat- ics through AIME (the American Invitational Math- ematics Examination), MathArena (Balunovic et al., 2025), and OMEGA (Sun et al., 2025). Ad- jacent reasoning evaluations follow the same pat- tern: HumanEval (Chen et al., 2021) for English code generation and MME-Reasoning (Yuan et al., 2025) for multimodal logical reasoning are sim- ilarly English-centred. The principal exceptions have come from Chinese CMATH (Wei et al., 2023) and AGIEval (Zhong et al., 2024). Multilingual mathematical benchmarks have only recently begun to emerge, but their lan- guage coverage remains skewed towards mostly high-resource languages. PolyMath (Wang et al., 2025) covers 18 languages across four difficulty tiers, from K-12 to Olympiad, and is among the most diverse benchmarks of its kind. Neverthe- less, the great majority of its languages are high- resource and already well-served by pretraining corpora. MGSM8KInstruct and the contemporane- ous MSVAMP test set (Chen et al., 2024) translate GSM8K and SVAMP into ten languages with a similar high-resource bias. MathNet (Alshammari et al., 2026) is the most ambitious effort to date, providing over 30K Olympiad problems across 47 countries and 17 languages. Despite this progress, the long tail of the world’s languages—particularly low-resource languages—remains absent from the reasoning evaluation Ghosh et al. (2025). 2.2 Multilingual and Cross-Lingual Reasoning A growing amount of work examines how reason- ing quality varies across languages and whether observed cross-lingual gaps reflect reasoning com- petence or surface artefacts. K et al. (2021) provide an early findings by category-annotating multilin- gual NLI and showing that transfer performance correlates with reasoning type and language simi- larity. More recent diagnostic work establishes that high task accuracy can mask reasoning that fails to support the model’s own conclusions: Ovalle et al. (2025) introduce a human-validated frame- work over 65K traces from Global-MMLU and find that non-Latin-script outputs exhibit at least twice the reasoning-answer misalignment of Latin- script outputs, with low-resource languages fre- quently arriving at correct answers through inco- herent traces. Liu et al. (2026) confirm this show- ing that latent reasoning in large reasoning models is uneven across 11 languages and broadly aligns with an English-centred internal pathway. Several studies already attempt to close these gaps through cross-lingual reasoning transfer. A common strategy routes reasoning through a high- resource pivot, either by translating inputs—like MathOctopus (Chen et al., 2024)—or by aggregat- ing traces across languages. Thus, Rajaee et al. (2026a) propose Best-of-L, a cross-lingual out- come reward model that ranks reasoning candidates across languages and improves accuracy on MGSM over single-language reward modeling. Some work propose edits directly on parameters: MAEC (Chen et al., 2025) decomposes language-agnostic ability- related weights from LLMs and recombines them with language-specific weights via simple addition and subtraction, transferring mathematical and sci- entific competencies into low-resource languages without additional training. 2.3 Translation as a Proxy for Multilingual Benchmarks Given the high cost of native data collection, trans- lation has become the dominant approach for scal- ing multilingual evaluation. Recent benchmarks such as BenchMAX (Huang et al., 2025) and MMLU-ProX (Xuan et al., 2025) combine ma- chine translation with LLM-based selection and native-speaker post-editing. For reasoning, Issaka et al. (2026) show that LLMs translation qual- ity strongly correlates with downstream bench- mark performance, though this relationship weak- ens on reasoning-intensive tasks such as MGSM. Similarly, Rajaee et al. (2026b) demonstrate that naively applying chain-of-thought reasoning dur- ing translation degrades quality, suggesting that translation and reasoning are distinct capabilities. Machine-translated benchmarks often lose cultural and pragmatic nuance, especially in low-resource languages underrepresented in pretraining data. In contrast, we rely on a thorough human control of pre-computer translations by native speakers for all mid- and low-resource languages. 3PLURAMATH Collection We build upon the PolyMath (Wang et al., 2025) dataset, which consists of four difficulty levels— low, medium, high, and top—with 125 samples per level, for a total of 500 tasks per language. Our data acquisition pipeline consists of three stages: (i) producing a first-draft translation through auto- matic systems; (i) manual verification by native speakers; and (i) automated and manual L A T E X checking together with final error analysis. All an- notators were fully informed about the goals of the project and were provided with written instructions, the full text of which is reproduced in Appendix B. 3.1 Language Stakeholders Selection For every selected language we recruited native speakers (i.e., speakers whose mother tongue is the target language) who additionally hold a higher-education degree in Computer Science or Mathematics—at least at the Master’s level, and in most cases at the PhD level. For languages requir- ing more checks, the primary stakeholder recruited additional native speakers (Bachelor’s or Master’s level in Computer Science or Mathematics) to con- firm the annotations and perform a second pass over the data. 3.2 Translation Pre-computation Translation Models and Language Pair Selec- tion We asked each language stakeholder to rec- ommend the source language and translation sys- tem that, in their empirical experience and after a small pilot on a handful of PolyMath tasks, yielded the most accurate translations into their target lan- guage. The resulting language pairs and systems are summarised in Table 1. Automated Translation For most languages, proprietary systems such as DeepL or Gemini were selected. For a smaller set of languages, stakehold- ers preferred open-weights LLMs that are specif- ically tuned for a given language pair and that ex- hibit very few translation artefacts in their domain— for instance,salamandraTA-7b-instructfor Spanish–Catalan andsarvam-mfor English–Hindi and English–Odia. For closed-source systems, we report the budget spent on obtained translations in Appendix G. 3.3 Manual Translation Verification Annotators were asked to verify three properties of every translated task: (i) the fluency and adequacy of the natural-language content; (i) the correct- ness of the mathematical terminology with respect to both the target language and its mathematical conventions; and (i) the strict equivalence of the L A T E X code to that of the source. 3.4 Post-polishing We performed an additional automated and manual pass to verify the L A T E X code and the completeness of every translation. The automated check verifies (i) that the L A T E X code compiles—i.e. that all$de- limiters are correctly opened and closed—and (i) that all command names and reserved keywords re- main in English and unaltered. The corresponding scripts are released for public use. 4 3.5 Final Dataset 3.5.1PLURAMATH Statistics Table 1 summarizes the PLURAMATH data creation process with dataset examples in Appendix C. The extent of required revisions varied substantially across languages, largely depending on language family and the maturity of available NLP and ma- chine translation technologies. To ensure quality, we conducted a second annotation pass for nearly half of the languages. We also compare task lengths between high- resource PolyMath languages and our target lan- guages in Figure 2, with a complete analysis pro- vided in Appendix D. Tokenization was performed using theQwen3-4Bbase model. Length differ- ences are strongly influenced by language family: Hindi, Odia, and Amharic exhibit notably longer sequences, whereas Latin- and Cyrillic-based lan- guages remain closer to the distributions observed in high-resource counterparts. 3.5.2 Encountered Errors Typical Translation Problems. The required corrections across languages fall into four main categories: (i) L A T E X corruption, including trans- lated command names, identifiers, and malformed $-delimited expressions; (i) incorrect mathemati- cal terminology, often involving generic substitu- tions or cross-linguistic leakage between related languages; (i) morphological and syntactic issues in morphologically rich languages, such as agree- ment, word order, and overly literal phrasing; and (iv) residual hallucinations, including untranslated tokens and incorrect entity substitutions. Correc- tion effort varied substantially and correlated with 4 § TUM-NLP/pluramath/blob/main/check_latex.ipynb LowMedHighTop 0 100 200 300 Length (tokens) English LowMedHighTop Russian LowMedHighTop Hindi LowMedHighTop Ukrainian LowMedHighTop Amharic LowMedHighTop Catalan LowMedHighTop Chuvash Figure 2: Comparison of math task lengths across high-resource languages from PolyMath (English and Russian) and five newly added languages in PLURAMATH, usingQwen3-4Btokenization. Task lengths vary substantially across language families. A complete comparison for all languages is provided in Appendix D. both linguistic distance from the source language and target-language resource availability, ranging from minimal editing (e.g., Ukrainian, Greek, He- brew) to extensive revision (e.g., Slovak, Tatar, Kazakh). In several cases, specialized language- pair models, includingTartuNLP,sarvamai, and salamandraTA, outperformed closed-source sys- tems due to the languages targeted adaptation. Problems in the Original PolyMathDuring an- notation and subsequent evaluation, we identified several inconsistencies in the original PolyMath dataset, including translation errors in non-English subsets and incorrect answers in the English data. We summarize the main issues in Appendix K and will report them to the PolyMath HuggingFace repository via a pull request. 4 Reasoning LLMs Benchmarking To evaluate a broad range of reasoning capabil- ities comparing4high-resource and our18lan- guages performance, we benchmark 27 state-of- the-art reasoning-oriented LLMs spanning multi- ple model families, training paradigms, parameter scales from sub-billion open-weight models to fron- tier proprietary APIs. The selection covers diverse architectural designs, post-training strategies, and levels of model accessibility. Models details, direct links to all checkpoints, and their corresponding licenses are provided in Appendix A. 4.1 Prompt Design We evaluate each model under three prompting set- tings: Base (problem and solution in the target lan- guage), Base+EN-CoT (target-language problem with reasoning instructed in English), and Back- translated (problem translated back with NLLB model (NLLB Team et al., 2024) into the original high-resource language with instructions in that language). The Base setting follows the original PolyMath setup (Wang et al., 2025), while the lat- ter two probe whether shifting reasoning to a high- resource language reduces performance gaps. The Base prompt is shown below; all prompt templates, language-specific instructions, and full examples for the Ukrainian case are provided in Appendix E. Base Prompt Template Problem in the target language Closing instruction in the target language (Original: Note: Please put the final answer in the $ $.) 4.2 Evaluation Metric For per-level scores, we used exact match accuracy of the text in the . For the aggregated score per language overall, we used the difficulty- weighted accuracy across all levels as defined in the original PolyMath benchmark (Wang et al., 2025). Specifically, leta i 4 i=1 denote the per-level accu- racies for the four difficulty levels (low, medium, high, top), and let the weights be defined asw 1 = 1 andw i = 2,w i−1 fori = 2, 3, 4. The aggregated benchmark score per language is then given by: DW-ACC = P 4 i=1 w i a i P 4 i=1 w i = 4 X i=1 2 i−1 15 a i . (1) 4.3 Hyperparameters Selection Before running the full evaluation, we con- ducted a hyperparameter search on the high- resource PolyMath languages using a subset of large reasoning models (gpt-oss-120b, Qwen3-235B,andQwen3-30B).Weex- ploreddifferentreasoning_effortand temperaturesettings,ultimatelyselecting reasoning_effort=medium temperature=0.1, andmax_completion_tokens=2kas the best overall configuration across languages. This setup was used for all subsequent experiments. Full hyperparameters search details and results are provided in Appendix F. High-resource Summary (HR) Our target languages Summary (target) Model en de ru es Avg. Len Lang hi tr pl uk uz or am el k cs he sr t sk ca cv hsb dsb Avg. Len Lang Open-weight — small ( ≤ 4B) Qwen3.5-0.8B 4.7 38 3.1 53 0.7 15 3.7 53 3.1 2288 ± 1432 TL 0.4 37 0.4 15 0.5 19 1.4 39 0.1 9 1.2 58 0.1 10 1.7 41 0.4 24 1.8 46 2.4 39 0.5 19 0.7 22 1.3 34 0.3 22 0.3 8 0.4 21 0.8 45 0.8 2238 ± 780 EN LFM2.5-1.2B 0.0 0 0.0 0 0.0 0 0.0 0 0.0 3485 ± 1213 EN 0.0 0 1.0 15 3.6 23 0.0 0 0.5 84 0.0 0 0.9 26 0.0 0 0.0 0 0.0 0 0.0 0 1.0 24 0.0 0 0.0 0 2.1 19 0.0 0 0.0 0 0.0 0 0.5 2660 ± 835 EN Ouro-1.4B 7.3 26 6.9 25 5.7 21 7.2 26 6.8 1477 ± 532 TL 2.9 15 3.3 17 5.2 22 1.4 10 0.5 15 0.4 7 0.6 34 2.1 13 0.4 8 2.6 16 1.7 17 1.4 10 1.4 11 2.3 15 4.8 21 0.3 14 0.7 15 1.0 14 1.8 1608 ± 372 EN R1-Distill-Qwen-1.5B 4.0 25 0.2 24 0.2 16 0.5 29 1.2 2691 ± 1058 EN 0.1 24 0.3 17 0.1 22 0.0 26 0.5 26 0.1 38 1.0 12 0.2 28 0.2 23 1.1 24 0.0 24 0.9 12 0.5 27 0.1 15 0.5 19 0.8 29 0.2 23 0.1 30 0.4 3032 ± 832 EN Qwen3.5-2B 2.0 15 2.1 23 2.0 23 2.1 20 2.1 1932 ± 151 TL 0.5 21 1.8 20 2.6 27 2.3 22 1.4 23 0.7 22 0.3 7 2.3 20 1.1 16 2.1 26 2.0 24 2.1 28 0.8 18 2.2 23 2.3 22 0.0 33 0.4 23 0.1 39 1.4 1982 ± 146 TL Ouro-2.6B 8.2 28 6.7 26 7.5 26 7.8 27 7.5 1587 ± 544 EN 4.4 20 5.7 23 0.9 1 5.8 22 1.8 16 1.8 15 0.3 19 5.5 22 0.6 13 5.3 22 3.0 16 4.2 20 1.2 18 4.7 20 5.7 23 1.1 20 2.7 20 2.0 18 3.1 1600 ± 453 EN Ministral-3-3B 9.1 38 7.9 39 9.9 46 14.5 71 10.3 1421 ± 761 TL 7.8 37 4.0 16 8.4 36 8.7 51 1.2 8 1.3 31 0.0 20 8.2 34 6.1 52 7.3 37 7.2 36 2.2 9 3.8 34 6.7 32 9.7 57 2.2 51 3.9 40 2.8 37 5.1 1634 ± 667 EN Gemma-3-4B 13.3 88 10.8 82 6.7 61 11.6 79 10.6 1262 ± 616 TL 8.4 73 6.4 68 9.0 79 10.5 81 7.6 71 2.5 59 5.3 66 10.7 79 3.6 35 9.5 78 7.4 76 9.9 75 4.1 44 8.2 83 10.3 84 2.1 39 3.2 75 3.3 64 6.8 1512 ± 841 TL Qwen3.5-4B 2.9 23 3.5 26 3.5 26 3.2 25 3.3 1912 ± 188 TL 2.6 39 3.1 26 3.8 28 3.1 26 3.0 31 1.8 37 2.8 35 3.6 27 3.2 28 3.1 27 4.0 28 3.6 29 3.0 35 3.7 26 3.7 27 1.4 41 1.2 47 1.4 40 2.9 1942 ± 187 TL Open-weight — mid (7–35B) OLMo-3-7B-Think 5.1 19 4.7 19 2.7 11 5.6 22 4.5 1891 ± 318 EN 1.8 11 3.3 15 4.6 20 3.8 17 0.1 22 1.4 9 0.1 10 2.8 12 1.3 12 2.5 14 2.4 12 0.2 3 0.3 40 2.3 12 4.0 17 0.1 46 0.5 23 0.2 24 1.8 1960 ± 187 EN R1-0528-Qwen3-8B 4.1 19 3.1 19 0.0 2 2.2 20 2.4 3181 ± 489 TL 0.0 53 0.6 14 0.7 13 0.1 51 0.3 24 0.0 43 0.0 42 0.0 54 0.0 54 0.4 37 0.0 53 0.0 46 0.0 49 0.2 29 0.8 22 0.0 45 0.1 30 0.1 38 0.2 3177 ± 429 EN Ministral-3-8B 10.7 36 8.7 38 9.6 37 10.9 42 10.0 1514 ± 769 EN 9.3 45 9.7 39 11.8 55 9.4 38 8.1 52 1.5 25 2.7 23 9.3 46 5.9 42 8.7 35 10.5 41 9.7 37 7.0 35 7.9 36 9.7 40 1.4 32 5.7 38 5.9 37 7.5 1544 ± 747 EN Qwen3.5-9B 3.7 24 3.6 31 5.0 30 4.5 27 4.2 3617 ± 782 TL 4.9 44 3.5 27 3.7 28 5.4 32 4.1 30 3.4 39 3.4 34 5.0 32 5.1 35 4.1 32 5.2 32 3.7 27 4.6 46 4.7 31 3.6 29 3.3 58 2.5 56 2.5 59 4.0 3083 ± 581 TL Ministral-3-14B 9.4 34 11.7 56 4.8 18 15.5 68 10.3 1559 ± 663 TL 8.5 37 4.7 20 5.9 23 3.9 18 3.5 19 0.1 7 0.0 6 9.1 38 8.0 33 9.3 43 10.5 51 3.3 15 2.3 23 3.1 40 4.4 17 2.2 36 5.5 32 4.3 28 4.9 1810 ± 549 EN gpt-oss-20b 17.6 50 15.8 51 19.2 50 18.0 50 17.7 2847 ± 1665 EN 15.8 51 12.4 38 12.1 37 18.6 51 10.2 36 15.2 73 9.3 34 16.7 48 16.0 48 17.3 50 16.6 49 11.8 37 16.3 48 14.9 47 12.7 38 4.4 68 13.6 50 12.6 55 13.7 2585 ± 1277 EN Nemotron3-Nano-30B 9.5 42 7.7 36 7.8 34 7.7 39 8.2 2207 ± 1796 EN 7.4 38 7.9 46 8.6 42 6.5 44 3.9 61 0.9 61 2.0 48 8.9 41 4.3 42 7.6 47 9.7 48 6.6 51 4.3 48 5.3 35 6.6 39 1.9 40 3.5 34 2.5 31 5.5 1608 ± 1461 EN Gemma-4-31B 6.6 23 7.4 28 8.7 29 8.6 29 7.8 1664 ± 534 TL 8.4 30 9.1 30 8.7 29 8.5 29 7.5 28 6.4 26 7.6 28 7.1 27 7.6 29 6.9 28 8.5 29 8.6 29 6.8 27 7.4 27 7.6 28 3.3 23 5.3 26 5.0 26 7.2 1724 ± 485 TL Qwen3.5-35B-A3B 3.7 24 3.6 28 3.9 26 4.2 26 3.9 1880 ± 241 TL 2.9 32 3.7 27 3.8 27 4.4 28 4.7 32 2.3 37 4.1 40 4.1 28 3.7 30 3.2 26 3.8 28 3.7 28 3.3 40 3.8 29 3.8 28 1.9 60 2.8 61 2.8 78 3.5 1919 ± 224 TL Open-weight — large / API R1-Distill-Llama-70B 6.8 28 5.9 29 9.4 33 6.5 29 7.2 1678 ± 667 TL 6.7 30 6.5 28 7.7 29 8.6 35 6.2 29 5.2 25 1.3 31 7.4 38 6.8 30 6.4 29 6.3 29 8.3 34 6.9 34 6.6 29 6.1 29 6.6 35 5.1 30 6.1 31 6.4 1703 ± 626 EN gpt-oss-120b 24.6 64 22.7 64 22.0 64 21.5 64 22.7 2590 ± 1607 TL 19.6 62 12.8 44 12.2 43 21.0 62 13.6 44 14.6 76 12.9 44 19.1 63 21.3 66 18.7 66 21.5 64 13.0 43 20.5 63 16.6 78 12.5 43 8.1 80 17.9 70 17.6 70 16.3 2328 ± 1267 EN Qwen3.5-122B-A10B 4.7 26 4.9 31 5.4 31 4.2 29 4.8 3447 ± 971 TL 4.0 35 4.2 27 3.9 27 5.1 31 4.3 28 5.2 41 4.1 33 5.8 34 4.9 33 4.3 32 4.8 31 4.3 28 5.0 41 5.1 33 4.4 28 3.9 61 4.2 63 4.3 67 4.5 2991 ± 666 TL Qwen3-235B-A22B 6.6 26 7.6 28 8.6 30 7.7 29 7.6 3269 ± 1264 TL 7.0 27 6.1 23 5.8 23 8.1 29 5.5 22 6.3 25 4.5 16 7.6 28 8.0 29 7.1 28 9.0 30 6.2 23 8.2 29 8.0 29 6.2 23 3.3 18 7.9 30 7.2 27 6.8 2888 ± 846 TL DeepSeek-V3.2 10.6 34 10.4 32 9.3 32 10.3 34 10.1 1751 ± 675 TL 10.0 38 9.8 35 11.0 39 8.8 32 8.5 31 9.2 30 8.7 36 10.3 41 8.9 31 9.4 34 10.4 39 10.3 38 7.8 28 10.1 33 9.7 36 9.0 33 9.2 42 9.8 35 9.5 1806 ± 605 TL Kimi-K2.5 6.9 26 4.9 27 4.5 29 5.8 25 5.5 1804 ± 580 TL 3.7 36 6.7 32 5.9 36 5.2 35 4.4 29 5.0 30 4.4 37 6.1 32 4.9 35 4.0 33 5.0 31 5.0 36 3.9 42 5.1 30 4.9 28 2.7 45 4.9 44 5.0 46 4.8 1878 ± 499 TL Closed-source Claude-Haiku-4.5 33.5 100 27.7 100 25.3 100 28.2 100 28.7 848 ± 401 TL 28.1 99 29.1 100 26.3 100 32.3 100 26.3 100 18.6 97 28.6 99 29.8 100 28.9 100 27.6 100 28.4 100 27.1 100 26.3 100 18.1 100 30.2 100 15.3 96 27.6 100 27.0 100 26.4 934 ± 414 TL Gemini-2.5-Flash 6.2 25 6.7 26 5.4 26 6.9 26 6.3 3636 ± 4226 TL 6.0 26 6.7 27 6.9 27 6.5 26 7.0 25 5.7 25 6.9 27 5.3 25 6.8 27 5.5 26 6.3 26 6.3 27 6.2 26 6.1 25 5.7 26 6.3 26 5.3 26 4.5 26 6.1 3891 ± 4364 TL GPT-5.4 17.3 100 16.3 100 15.4 100 15.8 100 16.2 317 ± 249 TL 16.1 100 16.5 100 16.2 100 16.5 100 17.3 100 13.4 100 14.2 100 15.3 100 16.3 100 15.6 100 17.5 100 16.0 100 14.8 100 14.5 100 16.8 100 14.7 100 15.2 100 15.0 100 15.7 371 ± 261 TL Table 2: Base prompting aggregated across levels results on our P LURA M ATH languages vs high-resource ones from PolyMath. Each cell shows difficulty-weighted accuracy (DW-Acc, %; large) over the answer-format compliance rate (%; small). Cell shading encodes DW-Acc (pale → deep teal, 0 → 40 + ); a numeral is printed in black on light (low-score) cells and in white on dark (high-score) cells solely for legibility—the text colour carries no extra meaning. Best and second-best per column are highlighted. After each language block we report the macro-average DW-Acc, the mean ± std generation length (tokens, reasoning+answer), and the dominant answer language coded as EN when the model answers predominantly in English or TL when it answers in the requested target language. 0102030 DW-ACC (%) en de es ru hi tr pl uk uz or am el k cs he sr t sk ca chv hsb dsb Aggregate 0255075100 acc. (%) en de es ru hi tr pl uk uz or am el k cs he sr t sk ca chv hsb dsb low 02040 acc. (%) medium 010203040 acc. (%) en de es ru hi tr pl uk uz or am el k cs he sr t sk ca chv hsb dsb high 01020 acc. (%) top 0102030 DW-ACC (%) Qwen3.5-0.8B LFM2.5-1.2B Ouro-1.4B DS-R1-Qwen-1.5B Qwen3.5-2B Ouro-2.6B Ministral-3-3B Gemma-3-4B Qwen3.5-4B Olmo-3-7B DS-R1-Qwen3-8B Ministral-3-8B Qwen3.5-9B Ministral-3-14B gpt-oss-20b Nemotron3-Nano-30B Gemma-4-31B Qwen3.5-35B-A3B DS-R1-Llama-70B gpt-oss-120b Qwen3.5-122B-A10B Qwen3-235B DeepSeek-V3.2 Kimi-K2.5 claude-haiku-4.5 gemini-2.5-flash gpt-5.4 Aggregate 0255075100 acc. (%) Qwen3.5-0.8B LFM2.5-1.2B Ouro-1.4B DS-R1-Qwen-1.5B Qwen3.5-2B Ouro-2.6B Ministral-3-3B Gemma-3-4B Qwen3.5-4B Olmo-3-7B DS-R1-Qwen3-8B Ministral-3-8B Qwen3.5-9B Ministral-3-14B gpt-oss-20b Nemotron3-Nano-30B Gemma-4-31B Qwen3.5-35B-A3B DS-R1-Llama-70B gpt-oss-120b Qwen3.5-122B-A10B Qwen3-235B DeepSeek-V3.2 Kimi-K2.5 claude-haiku-4.5 gemini-2.5-flash gpt-5.4 low 02040 acc. (%) medium 010203040 acc. (%) Qwen3.5-0.8B LFM2.5-1.2B Ouro-1.4B DS-R1-Qwen-1.5B Qwen3.5-2B Ouro-2.6B Ministral-3-3B Gemma-3-4B Qwen3.5-4B Olmo-3-7B DS-R1-Qwen3-8B Ministral-3-8B Qwen3.5-9B Ministral-3-14B gpt-oss-20b Nemotron3-Nano-30B Gemma-4-31B Qwen3.5-35B-A3B DS-R1-Llama-70B gpt-oss-120b Qwen3.5-122B-A10B Qwen3-235B DeepSeek-V3.2 Kimi-K2.5 claude-haiku-4.5 gemini-2.5-flash gpt-5.4 high 01020 acc. (%) top (a) Per-language score distributions (across models)(b) Per-model score distributions (across languages) basebase+enCoTbt high-resource (en, de, es, ru)target languages Figure 3: Math answer correctness distributions per language (a) and per model (b). We compare high-resource and target PLURAMATH languages across all prompting settings; the same for models, we analyze performance stability for thebaseprompting and differences between languages groups. Extended versions: Figures 11, 12 in Appendix. 4.4PLURAMATH Benchmarking Results AggregatedDW-ACCscores across all four difficulty levels are reported in Table 2, with per-level results in Appendix H.1. In addition to task accuracy, we report the proportion of responses adhering to the required format, aggregated reasoning length and the dominant language used in models’ outputs for both high-resource and our target lan- guages. We also report fully API costs and budget spent on our experiments in Appendix G. Math Answers CorrectnessThe distribution of results across languages (Figure 3) and the com- parison between high-resource and PLURAMATH languages confirm a persistent gap in mathemat- ical reasoning performance across language cate- gories. We compute the Spearman correlation be- tween languages’ class (Table 1) and math reason- ing benchmark rankings, obtainingρ = 0.646with p = 0.0038. This result indicates that the level of language support in current language tech- nologies strongly correlates with downstream mathematical reasoning performance. The magnitude of the gap varies substantially across languages. Greek and Polish show the small- est differences relative to high-resource languages (+0.67), whereas the average gap is +2.15, reach- ing+4.86for Chuvash and Ahmaric. Smaller mod- els often exhibit unstable performance across lan- guage groups, while larger recent systems, and especially proprietary models such asGPT-5.4and Claude-Haiku-4.5, remain considerably more stable across all evaluated languages. Reasoning Length Comparison A compari- son of model reasoning lengths is shown in Ap- pendix H.2. The gap between language groups is relatively modest, with models often generating shorter outputs for PLURAMATH languages due to failure to reach correct solutions. Notably, the best-performing models,Claude-Haiku-4.5and GPT-5.4, often produce correct answers with sub- stantially shorter reasoning traces, indicating that longer reasoning does not necessarily correspond to better performance. Prompts Choice Comparison Figure 3 also compares the effects of different prompting strate- gies across languages; full results are in Ap- pendix H.3. Overall, alternative prompting designs yield limited improvements for most models. 4.5 Correlation with Translation Capabilities To examine whether gains in reasoning are related to underlying translation capabilities, we conduct a ModelQ1Q2Q3Q4Q5Q6 High-resource (EN, RU) gemma-3-4B09994774577 Ministral-3-8B323100100031 Nemotron3_30B04245100327 gpt-oss-120b217931001869 DeepSeek-V3.23100100100633 PLURAMATH (11 underrepresented languages) gemma-3-4B2894966673 Ministral-3-8B3256563557 Nemotron3_30B192647256 gpt-oss-120b161171831264 DeepSeek-V3.233556641130 Figure 4: (A): Correlations with translations capabilities analysis. (B): Human reasoning assessment results. case study on a subset of models using both reason- ing and non-thinking modes for translation tasks on the FLORES+ (NLLB Team et al., 2024)dev and Sorbian shared-task (Okabe et al., 2025) splits. Correlation analyses are shown in Figure 4 (a), with detailed translation results in Appendix J. Overall, we observe a negative correlation be- tween CHRF++scores and output length (r = −0.16,p < 0.01), further confirming that longer reasoning traces do not necessarily improve transla- tion quality. Translation quality is moderately cor- related with both math task accuracy (r = +0.45, p < 10 −8 ) and instruction-following performance (r = +0.35,p < 10 −4 ), suggesting that some models are generally more reliable at adhering to task requirements. In contrast, gains from EN- COT prompting (r = +0.11,p = 0.21), the ten- dency to answer in English versus the target lan- guage (r pb = +0.07,p = 0.36), and language class (ρ = +0.08,p = 0.36) show little corre- lation with translation capability, indicating that other cross-lingual reasoning mechanisms play a larger role. 4.6 Human Assessment of Reasoning Finally, we conduct a human evaluation of se- lected models using 12 instances per difficulty level across two high-resource and (N) target languages. Annotators assessed whether correct answers ap- peared despite extraction failures (Q1), reasoning language compliance with the target (Q2), coher- ence and step-by-step validity of reasoning (Q3), consistency between reasoning and final answers (Q4), if a correct answer appeared at some point in reasoning (Q5), and if the model finished reason- ing at all (Q6). Aggregated results are reported in Figure 4 (b), with details provided in Appendix I. Reasoning quality declines for underrepre- sented languages, with models less consistently reasoning in the target language and producing less coherent explanations. With the exception of gemma, models frequently switch to English dur- ing reasoning. Moreover, all models—particularly Nemotron—often generate disfluent or incoherent reasoning and responses, including cases when the final answer was never produced during the rea- soning. Across all language groups, models also frequently fail to complete their reasoning within the2ktoken limit of our setup. These findings indi- cate that substantial improvements in multilingual reasoning capabilities are still required. 5 Conclusion We introduced PLURAMATH, an extension of a four-level mathematical reasoning benchmark to 18 underrepresented languages spanning six lan- guage families. We presented a multilingual bench- mark construction pipeline combining automatic translation for initial drafts with rigorous human evaluation and correction. Using this benchmark, we conducted a large-scale evaluation of 27 mod- ern reasoning LLMs. The performance gap persists between high-resource and our languages. We fur- ther compared standard prompting with EN-CoT and backtranslation strategies, finding that these interventions rarely yield consistent gains. Our analyses also indicate that stronger reasoning per- formance is not substantially correlated with trans- lation capability, but is more closely associated with general instruction-following ability. Human evaluation further reveals that many models gener- ate repetitive, incomplete, or non-fluent reasoning traces, whereas top-performing proprietary systems achieve substantially better results with shorter and concise reasoning. We hope that PLURAMATH and its accompanying analyses will support future re- search on multilingual and cross-lingual reasoning. Limitations Our work has several limitations. First, despite covering 18 underrepresented languages, the bench- mark remains far from exhaustive. Nevertheless, we hope the presented pipeline can serve as a practi- cal framework for extending reasoning benchmarks to additional digitally represented languages. Sec- ond, our approach depends on the availability of at least minimal machine translation resources, such as WMT-supported systems. For many extremely low-resource languages, benchmark construction would require fully manual translation and annota- tion, substantially increasing the cost and complex- ity of data creation. Third, advanced mathematical concepts cov- ered by higher difficulty levels may not be com- monly taught in all target-language educational contexts, raising questions about the practical rele- vance of such evaluations for some languages. At the same time, multilingual mathematical bench- marks may help future LLM systems support ac- cess to advanced education in underrepresented languages. Fourth, our evaluation uses a single- generation pass@1 setup rather than multiple gen- erations strategies and such scores aggregation as pass@k, which may underestimate the capabilities of some models. For some high-resource languages and difficulty levels, our experiments yielded lower results than those reported in PolyMath (Wang et al., 2025), likely due to the only pass@1 over the benchmark. However, the primary objective of this work is to highlight the performance gap between high-resource and underrepresented lan- guages and demonstrate the diversity of model be- haviors across different model families and scales. Finally, we do not investigate continued pretrain- ing or domain adaptation on multilingual mathe- matical corpora. Future work could explore how to collect and balance mathematical training data for underrepresented languages in order to improve multilingual reasoning performance. Ethics Statement Annotator Well-being We took explicit mea- sures to support annotator well-being, both in work- load management and recognition of contributions. All annotation was conducted on a non-profit basis by academic researchers and independent contrib- utors motivated to support their native languages. Most contributors are acknowledged through co- authorship. To reduce annotation burden, we first generated high-quality machine translations for each language pair, allowing annotators to focus on verification and correction rather than translation from scratch. Depending on translation quality, an- notators were given two to four weeks to complete their assignments with flexible scheduling. This process resulted in a manageable workload and a positive collaborative environment. LLM Usage Our evaluation involved extensive testing of dozens of LLMs across 22 languages, which likely incurred substantial computational and environmental costs. To maximize the long- term value of these experiments, we plan to re- lease the generated reasoning traces and model outputs, subject to license compliance. We hope these resources will support future research, includ- ing knowledge distillation into smaller and more accessible models that may better serve underrep- resented language communities and have a smaller environmental cost. Benchmark Data Contamination Finally, we do not systematically investigate potential bench- mark contamination in large-scale models. Since PolyMath (Wang et al., 2025) was released in mid 2025, some newer flagship LLMs may already have been exposed to portions of the benchmark dur- ing training. In addition, certain tasks originate from widely used mathematics textbooks and well- known olympiad problem sets, making incidental memorization possible across multiple languages. A deeper analysis of contamination effects remains an important direction for future work, and we hope that the release of our multilingual benchmark will facilitate such studies beyond high-resource lan- guages. Acknowledgments We express an enormous gratitude for all an- notators and supporters of the project. Firstly, we are grateful for our collaboration with the WITAJ-Sprachzentrum and thank Anita Hendri- chowa, Marko M ˇ eškank, and Kryštof Peršín, in particular, for their annotations of the Upper Sor- bian and Lower Sorbian splits. Secondly, the trans- lation to Catalan has been promoted by the Aina Project. We are also grateful to Šimon Kapusta for the help in Slovak translations check. The work of authors on Czech and Slovak splits was supported by the project CZ.02.01.01/00/23_020/0008518 of the Czech Ministry of Education, Youth and Sports. Finally, we warmly thank Alexander Antonov for the annotation of the Chuvash split. This work was co-funded by the European Union (ERC, EPICAL, 101141712 and ERC, NG-NLG, 101039303). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council. Neither the European Union nor the granting authority can be held responsible for them. References Shaden Alshammari, Kevin Wen, Abrar Zainal, Mark Hamilton, Navid Safaei, Sultan Albarakati, William T. Freeman, and Antonio Torralba. 2026. MathNet: A global multimodal benchmark for math- ematical reasoning and retrieval. In International Conference on Learning Representations (ICLR). ArXiv:2604.18584. Alexander Amini, Anna Banaszak, Harold Benoit, Arthur Bök, Tarek Dakhran, Song Duong, Alfred Eng, Fernando Fernandes, Marc Härkönen, Anne Harrington, Ramin M. Hasani, Saniya Karwa, Yuri Khrustalev, Maxime Labonne, Mathias Lechner, Valentine Lechner, Simon Lee, Zetian Li, Noel Loo, and 14 others. 2025. LFM2 technical report. CoRR, abs/2511.23404. Anthropic. 2025. System Card: Claude Opus 4 and Claude Sonnet 4. Technical report, Anthropic. Ac- cessed: May 2026. Mislav Balunovic, Jasper Dekoninck, Ivo Petrov, Nikola Jovanovic, and Martin T. Vechev. 2025. Matharena: Evaluating llms on uncontaminated math competi- tions. CoRR, abs/2505.23281. Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, and 39 others. 2021. Evaluating large language models trained on code. CoRR, abs/2107.03374. Nuo Chen, Zinan Zheng, Ning Wu, Ming Gong, Dong- mei Zhang, and Jia Li. 2024. Breaking language bar- riers in multilingual mathematical reasoning: Insights and observations. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Findings of ACL, pages 7001–7016. Association for Computa- tional Linguistics. Zhipeng Chen, Kun Zhou, Liang Song, Wayne Xin Zhao, Bingning Wang, Weipeng Chen, and Ji-Rong Wen. 2025. Extracting and combining abilities for building multi-lingual ability-enhanced large lan- guage models. In Proceedings of the 2025 Confer- ence on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, Novem- ber 4-9, 2025, pages 17563–17580. Association for Computational Linguistics. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word prob- lems. CoRR, abs/2110.14168. DeepSeek-AI. 2025. Deepseek-r1: Incentivizing rea- soning capability in llms via reinforcement learning. CoRR, abs/2501.12948. Akash Ghosh, Debayan Datta, Sriparna Saha, and Chi- rag Agarwal. 2025. A survey of multilingual reason- ing in language models. In Findings of the Associ- ation for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, pages 8920– 8936. Association for Computational Linguistics. Xu Huang, Wenhao Zhu, Hanxu Hu, Conghui He, Lei Li, Shujian Huang, and Fei Yuan. 2025. Benchmax: A comprehensive multilingual evaluation suite for large language models. In Findings of the Associa- tion for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, pages 16751– 16774. Association for Computational Linguistics. Sheriff Issaka, Erick Rosas Gonzalez, Lieqi Liu, Evans Kofi Agyei, Lucas Bandarkar, Nanyun Peng, David Ifeoluwa Adelani, Francisco Guzmán, and Saa- dia Gabriel. 2026. Translation as a scalable proxy for multilingual evaluation. CoRR, abs/2601.11778. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Men- sch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Re- nard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timo- thée Lacroix, and William El Sayed. 2023. Mistral 7b. CoRR, abs/2310.06825. Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020a. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 6282–6293. Association for Computational Linguistics. Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020b. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics. Karthikeyan K, Aalok Sathe, Somak Aditya, and Mono- jit Choudhury. 2021. Analyzing the effects of rea- soning types on cross-lingual transfer performance. In Proceedings of the 1st Workshop on Multilingual Representation Learning, pages 86–95, Punta Cana, Dominican Republic. Association for Computational Linguistics. Yihong Liu, Raoyuan Zhao, Hinrich Schütze, and Michael A. Hedderich. 2026. Large reasoning mod- els are (not yet) multilingual latent reasoners. CoRR, abs/2601.02996. NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Hef- fernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, and 20 others. 2024. Scaling neural machine translation to 200 languages. Nature, 630(8018):841–846. NVIDIA. 2026.Nemotron 3 nano omni:Effi- cient and open multimodal intelligence.CoRR, abs/2604.24954. Shu Okabe, Daryna Dementieva, Marion Di Marco, Lukas Edman,Katharina Haemmerl,Marko M ˇ eškank, Anita Hendrichowa, and Alexander Fraser. 2025. Findings of the WMT 2025 shared task LLMs with limited resources for Slavic languages: MT and QA. In Proceedings of the Tenth Conference on Ma- chine Translation, pages 503–519, Suzhou, China. Association for Computational Linguistics. Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groen- eveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, Nathan Lambert, Dustin Schwenk, Oyvind Tafjord, Taira Anderson, David Atkinson, Faeze Brahman, Christopher Clark, Pradeep Dasigi, Nouha Dziri, and 21 others. 2025. 2 olmo 2 furious. CoRR, abs/2501.00656. OpenAI. 2025. gpt-oss-120b & gpt-oss-20b model card. CoRR, abs/2508.10925. OpenAI. 2026. Openai GPT-5 system card. CoRR, abs/2601.03267. Anaelia Ovalle, Candace Ross, Sebastian Ruder, Ad- ina Williams, Karen Ullrich, Mark Ibrahim, and Levent Sagun. 2025. Beg to differ: Understand- ing reasoning-answer misalignment across languages. CoRR, abs/2512.22712. Keiran Paster, Marco Dos Santos, Zhangir Azerbayev, and Jimmy Ba. 2024. Openwebmath: An open dataset of high-quality mathematical web text. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. Sara Rajaee, Rochelle Choenni, Ekaterina Shutova, and Christof Monz. 2026a. Best-of-l: Cross-lingual re- ward modeling for mathematical reasoning. In Find- ings of the Association for Computational Linguistics: EACL 2026, Rabat, Morocco, March 24-29, 2026, Findings of ACL, pages 1930–1939. Association for Computational Linguistics. Sara Rajaee, Sebastian Vincent, Alexandre Berard, Marzieh Fadaee, Kelly Marchisio, and Tom Kocmi. 2026b.Unlocking reasoning capability on ma- chine translation in large language models. CoRR, abs/2602.14763. Yiyou Sun, Shawn Hu, Georgia Zhou, Ken Zheng, Han- naneh Hajishirzi, Nouha Dziri, and Dawn Song. 2025. OMEGA: can llms reason outside the box in math? evaluating exploratory, compositional, and transfor- mative generalization. CoRR, abs/2506.18880. Gemini Team. 2023. Gemini: A family of highly capa- ble multimodal models. CoRR, abs/2312.11805. Gemma Team. 2025a.Gemma 3 technical report. CoRR, abs/2503.19786. Kimi Team. 2025b. Kimi K2: open agentic intelligence. CoRR, abs/2507.20534. Qwen Team. 2025c. Qwen3 technical report. CoRR, abs/2505.09388. Yiming Wang, Pei Zhang, Jialong Tang, Haoran Wei, Baosong Yang, Rui Wang, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang, Junyang Lin, Fei Huang, and Jingren Zhou. 2025. Polymath: Evaluating math- ematical reasoning in multilingual contexts. CoRR, abs/2504.18428. Tianwen Wei, Jian Luan, Wei Liu, Shuang Dong, and Bin Wang. 2023. CMATH: can your language model pass chinese elementary school math test? CoRR, abs/2306.16636. Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Jun- jue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, and 13 others. 2025. Mmlu- prox: A multilingual benchmark for advanced large language model evaluation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, pages 1513–1532. Association for Computational Linguistics. Jiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu, Renrui Zhang, Kaituo Feng, Chaoyou Fu, Tao Chen, Lei Bai, Bo Zhang, and Xiangyu Yue. 2025. Mme- reasoning: A comprehensive benchmark for logical reasoning in mllms. CoRR, abs/2505.21327. Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2024. Agieval: A human-centric benchmark for evaluating foundation models. In Findings of the Association for Computational Lin- guistics: NAACL 2024, Mexico City, Mexico, June 16-21, 2024, Findings of ACL, pages 2299–2314. Association for Computational Linguistics. Rui-Jie Zhu, Zixuan Wang, Kai Hua, Tianyu Zhang, Ziniu Li, Haoran Que, Boyi Wei, Zixin Wen, Fan Yin, He Xing, Lu Li, Jiajun Shi, Kaijing Ma, Shanda Li, Taylor Kergan, Andrew Smith, Xingwei Qu, Mude Hui, Bohong Wu, and 14 others. 2025. Scaling la- tent reasoning via looped language models. CoRR, abs/2510.25741. A Details and Licensing of Resources In this section, we provide details about used LLMs for the experiments as well as all details, direct links, and licenses of all used resources. A.1 Short Summary of LLMs Families Used for the Experiments Qwen The Qwen series from Alibaba (Team, 2025c) has become one of the most widely adopted open-weight backbones for reasoning research, owing to its strong multilingual coverage, competitive math and code performance, and an actively maintained range of dense and Mixture-of-Experts variants. We include Qwen3.5 checkpoints at 0.8B, 2B, 4B, 9B, and the 35B-A3B, 122B-A10B MoE configurations as well as the older version of Qwen3 with the biggest 235B-A22B MoE variant. LiquidAI LFM The Liquid Foundation Models (Amini et al., 2025) depart from the standard Trans- former template by combining attention with structured state-space and convolutional blocks, yielding favourable memory and latency characteristics at small scales. The LFM2.5-1.2B-Thinking variant is interesting as a low-resource reasoning baseline. ByteDance Ouro The Ouro family (Zhu et al., 2025) introduces looped language models, in which a shared block of parameters is iterated multiple times during inference to amortise reasoning depth without increasing the parameter count. We test the 1.4B and 2.6B Thinking checkpoints to measure how this recurrence-style compute trades off against conventional depth-scaling. DeepSeekDeepSeek-R1 (DeepSeek-AI, 2025) pioneered large-scale reinforcement learning with veri- fiable rewards (RLVR) for open-weight reasoning, and its distilled checkpoints have become standard reference points. We evaluate the R1-Distill-Qwen-1.5B, R1-0528-Qwen3-8B, and R1-Distill-LLama-70B distillates as well as full V3.2 MoE version, providing both small-scale distillation evidence and a frontier open MoE comparison. Ministral / Mistral Mistral’s Ministral line (Jiang et al., 2023) targets efficient on-device and edge inference while inheriting the reasoning post-training recipe from the larger Mistral models. We include the recent 3B, 8B, and 14B Reasoning-2512 releases. Google Gemma Gemma (Team, 2025a) is Google DeepMind’s open-weight family derived from the same research stack as Gemini. The Gemma-3 4B and Gemma-4 31B variants allow us to compare a Gemini-aligned training recipe against other lineages at matched parameter budgets. Nemotron 3 Nano Omni is a multimodal large language model developed by NVIDIA (NVIDIA, 2026) that jointly processes video, audio, images, and text for tasks such as question answering, summarization, transcription, OCR, and document understanding. We use the Reasoning variant of the Nemotron series Nemotron-3-Nano-Omni-30B-A3B-Reasoning. AllenAI OLMo OLMo (OLMo et al., 2025) is, to our knowledge, the most fully open frontier-style series, releasing not only weights but also data, training code, and intermediate checkpoints. The OLMo- 3-7B-Think allow examining whether full transparency comes at a measurable reasoning cost. OpenAI gpt-oss The gpt-oss series (OpenAI, 2025) are OpenAI’s open-weight Mixture-of-Experts re asoning models, released under a permissive license and explicitly tuned for chain-of-thought workloads. The 20B and 120B checkpoints anchor the open-weight side of the comparison at the upper end. Moonshot KimiKimi-K2.5-Thinking (Team, 2025b) is Moonshot AI’s frontier MoE reasoning model, notable for its very long native context window and aggressive RL post-training. Closed-source frontier modelsTo upper-bound the comparison, we include API-only models from the three major frontier labs: Anthropic’s Claude-Haiku-4.5 (Anthropic, 2025), OpenAI’s GPT-5.4 (OpenAI, 2026), and Google’s Gemini-2.5-Flash (Team, 2023). A.2 Direct Links and Licensing Information for All Resources Used in This Work Below is an overview of the licenses and direct repositories of every resource used in this work (Table 3). The licenses associated with the models and datasets are consistent with the intended use of conducting academic research on multilingual mathematical reasoning for positive societal impact. Closed-source models are accessed exclusively through their respective official APIs in compliance with each provider’s terms of service. ResourceLicenseResource Link Datasets PLURAMATHApache-2.0hf.co/datasets/tum-nlp/PluraMath PolyMathApache-2.0hf.co/datasets/Qwen/PolyMath github.com/QwenLM/PolyMath FLORES+C-BY-SA-4.0hf.co/datasets/openlanguagedata/flores_plus Sorbian MT Dev SetCC-BY-NC-SA-4.0github.com/TUM-NLP/llms-limited-resources2025/tree/main/Sorbian Open-weight models — small (≤4B) Qwen3.5-0.8BApache-2.0hf.co/Qwen/Qwen3.5-0.8B LFM2.5-1.2B-ThinkingLFM Open Licensehf.co/LiquidAI/LFM2.5-1.2B-Thinking Ouro-1.4B-ThinkingApache-2.0hf.co/ByteDance/Ouro-1.4B-Thinking R1-Distill-Qwen-1.5BMIThf.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B Qwen3.5-2BApache-2.0hf.co/Qwen/Qwen3.5-2B Ouro-2.6BApache-2.0hf.co/ByteDance/Ouro-2.6B-Thinking Ministral-3-3BMistral Research Licensehf.co/mistralai/Ministral-3-3B-Reasoning-2512 Gemma-3-E4BGemma Licensehf.co/google/gemma-3-4b-it Qwe3-4B baseApache 2.0hf.co/Qwen/Qwen3.5-4B Qwen3.5-4BApache-2.0hf.co/Qwen/Qwen3.5-4B Open-weight models — mid (7–35B) OLMo-3-7B-ThinkApache-2.0hf.co/allenai/Olmo-3-7B-Think R1-0528-Qwen3-8BMIThf.co/deepseek-ai/DeepSeek-R1-0528-Qwen3-8B Ministral-3-8BApache 2.0hf.co/mistralai/Ministral-3-8B-Reasoning-2512 Qwen3.5-9BApache 2.0hf.co/Qwen/Qwen3.5-9B Ministral-3-14BApache 2.0hf.co/mistralai/Ministral-3-14B-Reasoning-2512 gpt-oss-20bApache-2.0hf.co/openai/gpt-oss-20b Nemotron3-Nano-30Bnvidia-open-model-agreementhf.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 Gemma-4-31BGemma Licensehf.co/google/gemma-4-31B Qwen3.5-35B-A3BApache-2.0hf.co/Qwen/Qwen3.5-35B-A3B Open-weight models — large (API-served) R1-Distill-LLama-70BMIThf.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B gpt-oss-120bApache-2.0hf.co/openai/gpt-oss-120b Qwen3.5-122B-A10BApache 2.0hf.co/Qwen/Qwen3.5-122B-A10B Qwen3-235B-A22BApache 2.0hf.co/Qwen/Qwen3-235B-A22B DeepSeek-V3.2MIThf.co/deepseek-ai/DeepSeek-V3.2 Kimi-K2.5Modified MIThf.co/moonshotai/Kimi-K2.5 Closed-source models (API only) Claude-Haiku-4.5Proprietaryanthropic.com/claude/haiku Gemini-2.5-FlashProprietaryai.google.dev/gemini-api/docs/models/gemini-2.5-flash GPT-5.4Proprietaryopenai.com/index/introducing-gpt-5-4 Additional Models Used for Translation tartuNLP/Qwen2.5-3B-Instruct-hsb-dsbCC-BY-4.0hf.co/tartuNLP/Qwen2.5-3B-Instruct-hsb-dsb BSC-LT/salamandraTA-7b-instructApache-2.0hf.co/BSC-LT/salamandraTA-7b-instruct SarvamaiApache-2.0hf.co/sarvamai/collections DeepLProprietaryhttps://w.deepl.com/en/products/api NLLBCC-BY-NC-4.0hf.co/facebook/nllb-200-distilled-600M Table 3: Overview of the licenses and direct links of the resources utilized in this work. B Translations Annotator Instructions Below we reproduce, the written instructions that were shared with every language stakeholder at the start of the project. Instructions for Language Stakeholders — PLURAMATH Project Motivation. We extend the PolyMath mathematical-reasoning benchmark (https:// huggingface.co/datasets/Qwen/PolyMath) beyond resource-rich languages and evaluate the performance of current large language models (LLMs) of various families and sizes on these languages. Language stakeholders contribute to the project on a research basis and are rewarded with co-authorship of the resulting paper. Tasks for Language Stakeholders. 1. Translation preparation. •Recommend the most suitable automatic translation system, so we can generate the first version of the benchmark. We are able to use paid versions of DeepL, Gemini, or Google Translate. •Apart from proprietary models, or in case a very specific translation system is chosen, it is very helpful if the stakeholder can run the translation themselves. 2. Translation check. •Only the question texts need to be checked; answers are typically numeric or LaTeX expres- sions and are universal. • The check should be carried out by native speakers: at least one, preferably two. • Translation guidelines: –Text outside LaTeX code must be fully and correctly translated, using the appropriate target-language terminology. –Text inside LaTeX code must remain in—or be reverted to—the original English. Please use the helper notebook to identify lines with an unbalanced number of$delimiters or otherwise broken LaTeX, then fix them so that the LaTeX is identical to the source. • Please record the number of instances that required correction. 3. Reasoning proof-of-concept study. • Take only the first task of each difficulty level. •Run it through several popular models, e.g., ChatGPT (GPT-5), Gemini, Claude, and Qwen3- 4B / Qwen3-30B-Thinking (the latter via the HuggingFace interface). • Goal: assess whether the models can comprehend the reasoning task at all in the target language. • Compile a short report of your findings. CPLURAMATH Tasks Examples We provide example instances from PLURAMATH for each difficulty level: low 5, medium 6, high 7, and top 8, alongside their corresponding original English tasks from PolyMath. ORIGINAL PROBLEMlow-en-3 James decides to run 3 sprints 3 times a week. He runs 60 meters each sprint. How many total meters does he run a week? hiHINDI जेम्स ने सप्ताह में 3 बार 3 िस्प्र ं ट दौड़ने का फै सला िकया। वह प्रत्येक िस्प्र ं ट में 60 मीटर दौड़ता है। वह एक सप्ताह में कु ल िकतने मीटर दौड़ता है? trTURKISH James haftada 3 kez 3 sprint koşmaya karar verir. Her sprintte 60 metre koşar. Haftada toplam kaç metre koşar? plPOLISH James trenuje bieganie 3 razy w tygodniu, za każdym razem wykonując 3 sprinty po 60 metrów. Ile metrów przebiega tygodniowo? ukUKRAINIAN Джеймс вирішує бігати по 3 спринти 3 рази на тиждень. Він пробігає 60 метрів за кожен спринт. Скільки всього метрів він пробігає на тиждень? uzUZBEK James haftasiga 3 marta 3 tadan sprint yugurishga qaror qildi. U har bir sprintda 60 metr yuguradi. U bir haftada jami necha metr yuguradi? orODIA େଜମ$ ସପ'ାହେର ୩ ଥର ୩ଟି /ି0ଣ2 େଦୗଡ଼ିବାକୁ ସି9ର କରନି'। େସ ପ;େତ=କ /ି0ଣ2େର ୬୦ ମିଟର େଦୗଡ଼ନି'। େସ ସପ'ାହେର େମାଟ େକେତ ମିଟର େଦୗଡ଼ନି'? amAMHARIC ጄምስ በሳምንት 3 ጊዜ 3 የሩጫ ዙሮችን ለመሮጥ ወሰነ። በእያንዳንዱ ዙር 60 ሜትር ይሮጣል። በሳምንት በአጠቃላይ ስንት ሜትር ይሮጣል? elGREEK Ο Τζέιμς αποφασίζει να κάνει 3 σπριντ 3 φορές την εβδομάδα. Τρέχει 60 μέτρα σε κάθε σπριντ. Πόσα μέτρα τρέχει συνολικά την εβδομάδα; kkKAZAKH Джеймс аптасына 3 рет, әр жатығуда 3 спринтен жүгіруді жоспарлады. Әр спринтің ұзындығы 60 метр. Ол бір аптада барлығы қанша метр жүгіреді? csCZECH James se rozhodne absolvovat třikrát týdně běh na krátkou vzdálenost. Na každou krátkou vzdálenost uběhne 60 metrů. Kolik metrů celkem uběhne za týden? heHEBREW הוא ריצה בכל .בשבוע פעמים 3 קצרות ריצות 3 לרוץ מחליט ימס'ג ?בשבוע רץ הוא הכל בסך מטר כמה .מטר 60 רץ srSERBIAN Џејмс одлучује да три пута недељно трчи по три спринта. Сваки спринт трчи 60 метара. Колико укупно метара претрчи недељно? ttTATAR Джеймс атнага өч тапкыр өчәр спринт йөгерергә карар кыла. Ул һәр спринтны 60 метрга йөгерә. Ул атнасына барлыгы ничә метр йөгерә? skSLOVAK James sa rozhodne behať 3 šprinty 3-krát v týždni. Každý šprint pobeží 60 metrov. Koľko metrov celkovo zabehne za týždeň? caCATALAN James decideix córrer 3 curses de velocitat 3 vegades per setmana. Corre 60 metres en cada cursa de velocitat. Quants metres corre en total per setmana? cvCHUVASH Джеймс эрнере куне 3 спринт тума йышăнать. Вăл кирек хăш спринт иртернĕ чухне 60 метр чупать. Вăл эрнере мĕн чухлĕ метр чупса иртет? hsbUPPER SORBIAN James wobzamknje, 3-króć wob tydźeń 3 krótkočarowe běhi absolwować. Na jedyn krótkočarowy běh běži 60 metrow. Kak wjele metrow wón cyłkownje wob tydźeń naběži? dsbLOWER SORBIAN James se rozsuźijo, 3-razow wob tyźeń absolwěrowaś 3- krotkich běgow. Za kuždy krotki běg ženjo 60 metrow. Wjele metrow dogromady wob tyźeń ženjo? Figure 5: Examples of a translated math problem from low level to our PLURAMATH languages. ORIGINAL PROBLEMmedium-en-108 Let $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$). Prove that the sequence $x_n$ converges. Find the limit $ _n → ∞ x_n$. hiHINDI मान लीिजए $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$)। िसद्ध कीिजए िक अनुक्रम $x_n$ अिभसारी है। सीमा $ _n → ∞ x_n$ ज्ञात कीिजए। trTURKISH $x_1=2021$ ve $n 1$ için $x_n^2- 2(x_n+1)x_n+1+2021=0$ olsun. $\x_n\$ dizisinin yakınsadığını kanıtlayınız. $ _n → ∞ x_n$ limitini bulunuz. plPOLISH Niech $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$). Udowodnij, że ciąg $x_n$ jest zbieżny. Wyznacz granicę $ _n → ∞ x_n$. ukUKRAINIAN Нехай $x_1=2021$, $x_n^2 - 2(x_n + 1)x_n+1 + 2021 = 0$ ($n ≥ 1$). Доведіть, що послідовність $x_n$ збігається. Знайдіть границю $ _n → ∞ x_n$. uzUZBEK $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$) bo'lsin. $x_n$ ketma-ketligining konvergensiyasini isbotlang. $ _n → ∞ x_n$ limitini toping. orODIA ଜଧରାଯାଉ $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$)। ପ(ମାଣ କର,- େଯ କ/ ମ $x_n$ ଅଭ2ସାରୀ ଅେଟ। ସୀମା $ _n → ∞ x_n$ ନି89ୟ କର,-। amAMHARIC $x_1=2021$ ይሁን፣ $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$)። ተከታዩ $x_n$ እንደሚሰባሰብ (converge) አረጋግጥ። ወሰኑን $ _n → ∞ x_n$ ፈልግ። elGREEK Έστω $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$). Αποδείξτε ότι η ακολουθία $x_n$ συγκλίνει. Βρείτε το όριο $ _n → ∞ x_n$. kkKAZAKH $x_1=2021$, $x_n^2 - 2(x_n + 1)x_n+1 + 2021 = 0$ ($n ≥ 1$) болсын. $x_n$ тізбегінің жинақталатынын дәлелдеңіз. $ _n → ∞ x_n$ шегін табыңыз. csCZECH Nechť $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$). Dokažte, že posloupnost $x_n$ konverguje. Najděte limitu $ _n → ∞ x_n$. heHEBREW x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$$ יהי הגבול את מצא .מתכנסת $x_n$ שהסדרה הוכח .($n 1$) .$lim_n → ∞ x_n\$ srSERBIAN Нека је $x_1=2021$ и $x_n^2-2(x_n+1)x_n+1+2021=0$ за $n 1$. Докажите да низ $\x_n\$ конвергира. Нађите лимес $ _n→∞x_n$. ttTATAR $x_1=2021$, $x_n^2 - 2(x_n + 1)x_n+1 + 2021 = 0$ ($n ≥ 1$) булсын. $x_n$ эзлеклелегенең җыелучанлыгын исбатлагыз. $ _n → ∞ x_n$ чикләмәсен табыгыз. skSLOVAK Nech $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$). Dokážte, že postupnosť $x_n$ konverguje. Nájdite limitu $ _n → ∞ x_n$. caCATALAN Sigui $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$). Demostreu que la successió $x_n$ convergeix. Trobeu el límit $ _n → ∞ x_n$. cvCHUVASH $x_1=2021$, $x_n^2 - 2(x_n + 1)x_n+1 + 2021 = 0$ ($n ≥ 1$) пултӑр. $x_n$ йӗрке пӗр пӑнчӑ патне туртӑнасине кӑтартса парӑр. $ _n → ∞ x_n$ пределне тупӑр. hsbUPPER SORBIAN Njech je $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n 1$). Dopokazajće, zo slěd $x_n$ konwerguje. Namakajće hraničnu hódnotu $ _n → ∞ x_n$. dsbLOWER SORBIAN Daś jo $x_1=2021$, $x_n^2-2(x_n+1)x_n+1+2021=0$ ($n≥ 1$). Dopokažćo, až slěd $x_n$ konwergěrujo. Namakajśo granicnu gódnotu $ _n → ∞ x_n$. Figure 6: Examples of a translated math problem from medium level to our PLURAMATH languages. ORIGINAL PROBLEMhigh-en-22 Consider the paths of length $16$ that follow the lines from the lower left corner to the upper right corner on an $8× 8$ grid. Find the number of such paths that change direction exactly four times. hiHINDI $16$ लंबाई के उन रास्तों पर िवचार करें जो $8× 8$ िग्रड पर िनचले बाएं कोने से ऊपरी दाएं कोने तक जाने वाली रेखाओं का अनुसरण करते हैं। ऐसे रास्तों की संख्या ज्ञात कीिजए जो ठीक चार बार िदशा बदलते हैं। trTURKISH $8× 8$ bir ızgara üzerinde sol alt köşeden sağ üst köşeye çizgileri takip eden $16$ birim uzunluğundaki yolları düşünün. Bu yollardan kaçının tam dört kez yön değiştirdiğini bulun. plPOLISH Rozważmy ścieżki o długości $16$ które biegną wzdłuż linii siatki od lewego dolnego rogu do prawego górnego rogu na siatce $8× 8$. Wyznacz liczbę takich ścieżek, które zmieniają kierunek dokładnie cztery razy. ukUKRAINIAN Розглянемо шляхи завдовжки $16$, які прямують по лініях із нижнього лівого кута в правий верхній кут на сітці розміром $8 × 8$. Знайдіть кількість таких шляхів, які змінють напрямок рівно чотири рази. uzUZBEK $8× 8$ o'lchamdagi to'rda pastki chap burchakdan yuqori o'ng burchakka chizilgan chiziqlar bo'ylab uzunligi $16$ bo'lgan yo'llarni ko'rib chiqing. Aynan to'rt marta yo'nalishini o'zgartiradigan bunday yo'llarning sonini toping. orODIA $16$ ଲମ#ର ପଥଗୁଡ଼ିକ ବିଷୟେର ବିଚାର କର23 ଯାହାକି $8× 8$ ଗ6ୀଡେର ତଳ ବାମ େକାଣର< ଉପର ଡାହାଣ େକାଣ ପଯ>?ନA େରଖାଗୁଡ଼ିକୁ ଅନୁ ସରଣ କେର | ଏପରି ପଥଗୁଡ଼ିକର ସଂଖ?ା େଖାଜ23 ଯାହା ସଠିI ଭାବେର ଚାରିଥର ଦିଗ ପରିବL>ନ କେର | amAMHARIC በ 8x8 ፍርግርግ ላይ ከታችኛው ግራ ጥግ ወደ ላይኛው ቀኝ ጥግ መስመሮችን የሚከተሉ የ16 ርዝመት ያላቸውን መንገዶች አስቡ። በትክል አራት ጊዜ አቅጣጫ የሚቀይሩ የእንደዚህ አይነት መንገዶችን ብዛት ፈልጉ። elGREEK Εξετάστε τα μονοπάτια μήκους $16$ που ακολουθούν τις γραμές από την κάτω αριστερή γωνία προς την πάνω δεξιά γωνία σε ένα πλέγμα $8\ επί 8$. Βρείτε το πλήθος αυτών των μονοπατιών που αλάζουν κατεύθυνση ακριβώς τέσερις φορές. kkKAZAKH 8×8 торда төменгі сол жақ бұрыштан жоғарғы оң жақ бұрышқа дейінгі сызықтар бойымен өтетін ұзындығы 16 болатын жолдарды қарастырыңыз. Дәл төрт рет бағытын өзгертетін осындай жолдардың санын табыңыз. csCZECH Uvažujte cesty o délce $16$, které vedou po přímkách z levého dolního rohu do pravého horního rohu na mřížce $8× 8$. Najděte počet takových cest, které změní směr přesně čtyřikrát. heHEBREW מהפינה הקוים אחר העוקבים ,16 באורך במסלולים הביטו .8×8 בגודל ברשת העליונה הימנית לפינה התחתונה השמאלית פעמים ארבע בדיוק כיון שמשנים הלו המסלולים מספר את מצאו srSERBIAN Размотрите путање дужине 16 које прате линије мреже од доњег левог угла до горњег десног угла на мрежи $8×8$. Нађите број таквих путања које мењају смер тачно четири пута. ttTATAR $8 × 8$ челтәрдә түбән сул почмактан өске уң почмака сузылган сызыклар буйлап барган озынлыгы $16$ булган юларны карагыз. Төгәл дүрт тапкыр юнәлешне үзгәртүче юлар санын табыгыз. skSLOVAK Uvažujme cesty dĺžky $16$, ktoré vedú z ľavého dolného rohu do pravého horného rohu po mriežke $8 át 8$. Nájdite počet ciest, ktoré zmenia smer práve štyrikrát. caCATALAN Considera els camins de longitud $16$ que segueixen les línies des de la cantonada inferior esquerra fins a la cantonada superior dreta en una quadrícula de $8× 8$. Troba el número d'aquests camins que canvien de direcció exactament quatre vegades. cvCHUVASH $16$ тӑршӗлӗ ҫулсене, $8 × 8$ размерлӑ сеткӑра сулахай аялти кӗтесрен сылтӑм ҫӳлти кӗтесе ҫити линисемпе пыраканисене, пӑхса тухар. Тӑватӑ хут ҫеҫ тӗлне улӑштаракан ҫавӑн пек ҫулсен шутне тупӑр. hsbUPPER SORBIAN Wobhladajće sej šćežki dołhosće $16$, kotrež sćěhuja linijam wot delnjeho lěweho róžka k hornjemu prawemu róžkej na $8× 8$ lěsycy. Namakajće ličbu tajkich šćežkow, kotrež dokładnje štyri króć směr změnja. dsbLOWER SORBIAN Woglědujśo se sćažki z dłujkosću $16$, kótarež slěduju linijam wót dolnego lěwego rožka do górnego pšawego rožka na $8× 8$ pśepletanju. Namakajśo licbu takich sćažkow, kótarež dokradnje styri raze směr změnjaju. Figure 7: Examples of a translated math problem from high level to our PLURAMATH languages. ORIGINAL PROBLEMtop-en-1 Let $Q$ be the set of rational numbers. A function $f: Q → Q$ is called aquaesulian if the following property holds: for every $x,y ∈ Q$,\[ f(x+f(y)) = f(x) + y or f(f(x)+y) = x + f(y). \]Show that there exists an integer $c$ such that for any aquaesulian function $f$ there are at most $c$ different rational numbers of the form $f(r) + f(-r)$ for some rational number $r$, and find the smallest possible value of $c$. hiHINDI मान लीिजए िक $Q$ पिरमेय संख्याओं का समुच्चय है। एक फलन $f: Q → Q$ को एक्वेसुिलयन कहा जाता है यिद िनम्निलिखत गुण सत्य हो: प्रत्येक $x,y ∈ Q$ के िलए, $f(x+f(y)) = f(x) + y या f(f(x)+y) = x + f(y)।$ िसद्ध कीिजए िक एक पूणार्ंक $c$ इस प्रकार मौजूद है िक िकसी भी एक्वेसुिलयन फलन $f$ के िलए, िकसी पिरमेय संख्या $r$ के िलए $f(r) + f(-r)$ के रूप में अिधकतम $c$ िभन्न पिरमेय संख्याएँ हों, और $c$ का सबसे छोटा संभव मान ज्ञात कीिजए। trTURKISH $Q$, rasyonel sayılar kümesi olsun. $f: Q → Q$ fonksiyonuna, aşağıdaki özellik sağlanıyorsa, aquaesulian fonksiyon denir: her $x,y ∈ Q$ için,\[ f(x+f(y)) = f(x) + y veya f(f(x)+y) = x + f(y). \] Herhangi bir aquaesulian fonksiyon $f$ için, $r$ rasyonel sayısı için $f(r) + f(-r)$ biçiminde en fazla $c$ farklı rasyonel sayı olduğunu ve $c$'nin mümkün olan en küçük değerini bulun. plPOLISH Niech $Q$ będzie zbiorem liczb wymiernych. Funkcję $f: Q → Q$ nazywamy aquaesulianem, jeśli zachodzi następująca własność: dla każdego $x,y ∈ Q$,\[ f(x+f(y)) = f(x) + y or f(f(x)+y) = x + f(y). \]Wykaż, że istnieje liczba całkowita $c$ taka, że dla dowolnej funkcji aquaesulianowej $f$ istnieje co najwyżej $c$ różnych liczb wymiernych postaci $f(r) + f(-r)$ dla pewnej liczby wymiernej $r$, oraz znajdź najmniejszą możliwą wartość $c$. ukUKRAINIAN Нехай $Q$ є множиною раціональних чисел. Функція $f: Q → Q$$ називається акуасуліанською, якщо виконується така властивість: для кожного $x, y ∈ Q$,\[ f(x+f(y)) = f(x) + y або f(f(x)+y) = x + f(y). \]Доведіть, що існує ціле число $c$ таке, що для будь-якої акуасуліанської функції $f$ існує не більше ніж $c$ різних раціональних чисел виду $f(r) + f(-r)$ для деяких раціональних чисел $r$, і знайдіть найменше можливе значеня $c$. uzUZBEK $Q$ mantiqiy sonlar to'plami bo'lsin. $f: Q → Q$ funksiyasi aquaesulian deb ataladi, agar quyidagi xususiyat bajarilsa: har qanday $x,y ∈ Q$ uchun, \[ f(x+f(y)) = f(x) + y yoki f(f(x)+y) = x + f(y). \]Har qanday aquaesulian funksiya $f$ uchun, ba’zi bir ratsional $r$ sonlar uchun hosil bo‘ladigan $f(r) + f(-r)$ ko‘rinishdagi turli ratsional sonlar soni $c$ dan oshmasligini ko'rsating va $c$ ning eng kichik mumkin bo'lgan qiymatini toping. orODIA ଧରାଯାଉ $Q$ େହଉଛି ପରିେମୟ ସଂଖ0ାଗୁଡ଼ିକର େସ5। ଏକ ଫ8ସ9 $f: Q → Q$ କୁ ଆେକ;ସୁଲ=ଆ9 କୁହାଯାଏ ଯଦି ନିମ@ଲ=ଖAତ ଗୁଣଟି ପୂରଣ Fଏ: ପGେତ0କ $x,y ∈ Q$ ପାଇଁ, $f(x+f(y)) = f(x) + y ଅଥବା f(f(x)+y) = x + f(y)।$ ଦଶNାOP େଯ ଏପରି ଏକ ପୂQNସଂଖ0ା $c$ ଅଛି େଯଉଁଥAପାଇଁ େଯେକୗଣସି ଆେକ;ସୁଲ=ଆ9 ଫ8ସ9 $f$ ପାଇଁ କିଛି ପରିେମୟ ସଂଖ0ା $r$ ନିମେନS $f(r) + f(-r)$ ରT ପ ର ଅତି େବଶୀ $c$ ଟି ଭ=ନ@ ପରିେମୟ ସଂଖ0ା ରହିେବ, ଏବଂ $c$ ର ସବNନିମ@ ସWାବ0 ମୂଲ0 ନିQNୟ କରOP। amAMHARIC "$Q$ የረሽናል (መጠነኛ) ቁጥሮች ስብስብ ይሁን። አንድ ፈንክሽን $f: Q → Q$ ""aquaesulian"" ይባላል—የሚከተለው ባህሪ ለማንኛውም $x,y ∈ Q$ የሚሰምር ከሆነ፦ $$ f(x+f(y)) = f(x) + y ወይም f(f(x)+y) = x + f(y). $$ ለማንኛውም ""aquaesulian"" ፈንክሽን $f$፣ በ $f(r) + f(-r)$ መልክ ሊገለጹ የሚችሉ (እዚህ ላይ $r$ ረሽናል ቁጥር ነው) የተለያየ ዋጋ ያላቸው ረሽናል ቁጥሮች ብዛት ከ $c$ እንዳይበልጥ የሚያደርግ አንድ ሙሉ ቁጥር $c$ መኖሩን አሳይ፤ እንዲሁም ሊሆን የሚችለውን አነስተኛውን የ $c$ ዋጋ ፈልግ።" elGREEK Έστω $Q$ το σύνολο των ρητών αριθμών. Μια συνάρτηση $f: Q \στο Q$ ονομάζεται ακουαζουλιανή αν ισχύει η ακόλουθη ιδιότητα: για κάθε $x,y \ στο Q$,\[ f(x+f(y)) = f(x) + y or f(f(x)+y) = x + f(y). \]Δείξτε ότι υπάρχει ένας ακέραιος αριθμός $c$ τέτοιος ώστε για κάθε ακουαζουλιανή συνάρτηση $f$ να υπάρχουν το πολύ $c$ διαφορετικοί ρητοί αριθμοί της μορφής $f(r) + f(-r)$ για κάποιο ορθολογικό αριθμό $r$, και βρείτε τη μικρότερη δυνατή τιμή του $c$. kkKAZAKH $Q$ рационал сандар жиыны болсын. Егер кез келген $x, y ∈ Q$ үшін$$ f(x+f(y)) = f(x) + y немесе f(f(x)+y) = x + f(y) $$шарты орындалса, $f: Q → Q$ функциясы акуасулиандық деп аталады. Кез келген акуасулиандық $f$ функциясы үшін қандай да бір рационал $r$ сандары бойынша алынған $f(r) + f(-r)$ түріндегі әр түрлі мәндердің саны $c$-ден аспайтындай қандай да бір $c$ бүтін саны табылатынын дәлелдеңіз және осындай $c$ санының ең кіші мүмкін мәнін табыңыз.. csCZECH Nechť $Q$ je množina racionálních čísel. Funkce $f: Q → Q$ se nazývá aquaesulská, jestliže je splněna následující vlastnost: pro všechny $x,y ∈ Q$ platí\[ f(x+f(y)) = f(x) + y nebo f(f(x)+y) = x + f(y). \]Ukažte, že existuje celé číslo $c$ takové, že pro libovolnou aquaesulovskou funkci $f$ existuje nejvýše $c$ různých racionálních čísel tvaru $f(r) + f(-r)$ pro libovolné racionální číslo $r$, a najděte nejmenší možnou hodnotu $c$. heHEBREW f:$ פונקציה .הרציונלים המספרים קבוצת $mathbbQ\$ תהי מתקימת אם אקוסולית נקראת $Q → Q f(x+f(y)) = ]\,$x,y ∈ Q$ כל עבור :הבאה התכונה כי הוכח[\ .quad f(f(x)+y) = x + f(y)\ אוf(x) + y לכל יש $f$ אקוסולית פונקציה שלכל כך $c$ שלם מספר קים עבור $f(r) + f(-r)$ מהצורה שונים רציונלים מספרים $c$ היותר של האפשרי ביותר הקטן הערך את ומצא ,$r$ כלשהו רציונלי מספר $c$ srSERBIAN Нека $Q$ буде скуп рационалних бројева. Функција $f: Q → Q$ назива се акваезулијанском ако важи следеће својство: за сваки $x,y ∈ Q$,\[ f(x+f(y)) = f(x) + y или f(f(x)+y) = x + f(y). \]. Покажите да постоји цео број $c$ такав да за сваку акваезулијанску функцију $f$ постоји највише $c$ различитих рационалних бројева облика $f(r) + f(-r)$ за неки рационални број $r$, и пронађите најмању могућу вредност за $c$. ttTATAR $Q$ рациональ санар күплеге булсын. $f: Q → Q$ функциясе Аквасулиан дип атала, әгәр түбәндәге үзлек үтәлсә: һәр $x, y ∈ Q$ өчен, \[ f(x+f(y)) = f(x) + y яки f(f(x)+y) = x + f(y). \]Һәр Аквасулиан функциясе $f$ өчен кайбер рациональ санар $r$ өчен $f(r) + f(-r)$ формасындагы рациональ санар саны $c$-дән артмавын дәлиләгез һәм $c$-ның мөмкин булган иң кечкенә кыймәтен табыгыз. skSLOVAK Nech $Q$ je množina racionálnych čísel. Funkcia $f: Q Q$ sa nazýva aquaesulská, ak platí nasledujúca vlastnosť: pre každé $x,y ∈ Q$,\[ f(x+f(y)) = f(x) + y alebo f(f(x)+y) = x + f(y). \]Ukážte, že existuje celé číslo $c$ také, že pre ľubovoľnú aquaesulskú funkciu $f$ existuje najviac $c$ rôznych racionálnych čísel tvaru $f(r) + f(-r)$ pre nejaké racionálne číslo $r$, a nájdite najmenšiu možnú hodnotu $c$. caCATALAN Sigui $Q$ el conjunt dels nombres racionals. Una funció $f: Q → Q$ s'anomena aquaeusuliana si es compleix la següent propietat: per a cada $x,y ∈ Q$,\[ f(x+f(y)) = f(x) + y or f(f(x)+y) = x + f(y). \] Mostra que existeix un nombre enter $c$ tal que per a qualsevol funció aquaesuliana $f$ hi ha com a màxim $c$ nombres racionals diferents de la forma $f(r) + f(-r)$ per a algun nombre racional $r$, i troba el valor més petit possible de $c$. cvCHUVASH "$Q$ рационалӑ хисепсен множестви пултӑр. $f: Q → Q$ функцине акуасулианлӑ теҫӗ, енчен те кашни $x, y ∈ Q$ вали \[ f(x+f(y)) = f(x) + y е f(f(x)+y) = x + f(y) \] свойстви пурнӑҫланать. $c$ тули хисеп пулсан, кашни акуасулианлӑ $f$ функци вали, хӑш-пӗр $r$ рационалӑ хисепсем вали, $f(r) + f(-r)$ евӗрлӗ рационалӑ хисепсен шучӗ $c$-ран нумайрах мар пулсан, $c$ пурине кӑтартса парӑр, тата $c$ чи пӗчӗк пӗлтерӗшне тупӑр." hsbUPPER SORBIAN $Q$ je mnohosć racionalnych ličbow. Funkcija $f: Q → Q$ rěka aquaesulian, hdyž je slědowaca kajkosć spjelnjena: za wšitke $x,y ∈ Q$ płaći\[\ f(x+f(y)) = f(x) + y abo f(f(x)+y) = x + f(y). \]Pokazajće, zo eksistuje cyła ličba $c$, tak zo za kóždu aquaesulianu funkciju $f$ maksimalnje $c$ wšelakich racionalnych ličbow formy $f(r) + f(-r)$ za někajku racionalnu ličbu $r$ eksistuja, a namakajće najmjeńšu hódnotu wot $c$. dsbLOWER SORBIAN Daś jo $Q$ młogosć racionalnych licbow. Funkciji $f:Q $ groni se aquaesuliantna, gaž slědujuca kakosć jo docynjona: za wšykne $x,y $ płaśi \ [f(x+f(y))=f(x)+y f(f(x)+y)=x+f(y).\] Dopokažćo, až wobstoj ceła licba $c$, tak až za kuždu aquaesuliantnu funkciju $f$ w nejwušem paźe eksistěrujo $c$ wšakich racionalnych licbow formy $f(r)+f(-r)$ za někaki racionalny numer $r$, a namakajśo nejmjeńšu gódnotu $c$. Figure 8: Examples of a translated math problem from top level to our PLURAMATH languages. DPLURAMATH and PolyMath Math Tasks Length Comparison We compare math task lengths between high-resource PolyMath languages—English, Spanish, Russian, and German—and the newly introduced languages in PLURAMATH. Tokenization was performed with Qwen3-4Bmodel. Mean and standard deviation statistics are reported in Tables 4,5,6,7, with box plots shown in Figure 9. Task lengths vary substantially across language families: Slavic languages exhibit relatively similar distributions, whereas Indo-Aryan languages such as Hindi and especially Odia, as well as the Semitic language Amharic and Turkic languages, have considerably longer inputs. These length differences may introduce additional challenges for task comprehension and long-context reasoning. LevelEnglishSpanishRussianGerman Low 60±22 76±29 90±34 81±32 Medium 96±53 106±56 115±63 112±61 High 114±78 131±83 144±89 140±86 Top 118±62 137±72 154±80 146±75 Table 4: Length (tokens) per problem (mean±std) for the four high-resource languages from PolyMath, by difficulty level. LevelHindiTurkishPolishUkrainianUzbekOdia Low 219±86 84±32 90±35 126±47 106±41 471±187 Medium 194±119 115±62 117±64 135±76 126±69 346±236 High 279±159 144±88 151±92 179±104 163±99 536±308 Top 317±181 153±78 161±82 195±103 180±91 604±385 Table 5: Length (tokens) per problem (mean±std) for our PLURAMATH target languages (1/3), by difficulty level. LevelAmharicGreekKazakhCzechHebrewSerbian Low 211±76 228±93 149±56 100±39 81±32 118±45 Medium 190±115 190±116 158±88 121±68 110±59 123±67 High 278±161 278±161 217±120 163±97 136±86 157±69 Top 298±169 311±187 241±129 177±91 143±73 181±95 Table 6: Length (tokens) per problem (mean±std) for our PLURAMATH target languages (2/3), by difficulty level. LevelTatarSlovakCatalanChuvashU. SorbianL. Sorbian Low 139±53 102±41 85±33 141±54 112±45 115±47 Medium 154±86 121±68 110±59 148±85 129±71 130±73 High 204±114 161±96 139±88 206±119 171±100 173±101 Top 226±119 173±91 148±77 219±118 182±95 187±100 Table 7: Length (tokens) per problem (mean±std) for our PLURAMATH target languages (3/3), by difficulty level. LowMedHighTop 0 200 400 Length (tokens) English LowMedHighTop Spanish LowMedHighTop Russian LowMedHighTop German LowMedHighTop 0 200 400 Length (tokens) Hindi LowMedHighTop Turkish LowMedHighTop Polish LowMedHighTop 0 200 400 Length (tokens) Ukrainian LowMedHighTop Uzbek LowMedHighTop Odia LowMedHighTop 0 200 400 Length (tokens) Amharic LowMedHighTop Greek LowMedHighTop Kazakh LowMedHighTop 0 200 400 Length (tokens) Czech LowMedHighTop Hebrew LowMedHighTop Serbian LowMedHighTop 0 200 400 Length (tokens) Tatar LowMedHighTop Slovak LowMedHighTop Catalan LowMedHighTop 0 200 400 Length (tokens) Chuvash LowMedHighTop U. Sorbian LowMedHighTop L. Sorbian Figure 9: Fuul box-plot comparison of math task sample lengths between four high-resource languages from PolyMath and the 18 newly introduced languages in PLURAMATH. E Prompt Templates and Examples This section gives the full text of all prompts used in the experiments together with a specific illustrative examples in Ukrainian. Throughout, “HR” denotes the high-resource pivot language (English by default; for non-English HR baselines such as German, Russian, or Spanish we substitute the language name in the directive). E.1 Templates 1. BaseProblem and instruction in the target language. The model is free to reason in any language but is asked to place the final answer in $ $. Base problem tgt closing_instruction tgt 2. Base+EN-CoT Problem in the target language; an additional system-level directive instructs the model to perform its chain-of-thought in English while still placing the final answer in a $ $. Base+EN-CoT [System] You are solving a mathematical problem. Reason step by step in English, then write the final answer in $ $. [User] problem tgt 3. Backtranslated The problem is first machine-translated from the target language back to the high- resource language using X model, then presented to the model entirely in the high-resource language. Backtranslated problem HR closing_instruction HR E.2 Prompts Example in Ukrainian We illustrate every prompt template on a concrete problem from the low-difficulty level in Ukrainian. Native (Ukrainian) Качки Дженет несуть по 16 яєць на день. Вона їсть три на снiданок щоранку i пече кекси для своїх друзiв щодня, використовуючи чотири. Решту вона щодня продає на фермерському ринку по 2 $ за свiже качине яйце. Скiльки в доларах вона заробляє на фермерському ринку щодня? Примiтка: Будь ласка, вставте остаточну вiдповiдь в $ $. English reference for illustration purpose only Janet’s ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers’ market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers’ market? Note: Please put the final answer in $ $. Native+EN-CoT (Ukrainian problem English CoT) [System] You are solving a mathematical problem. Reason step by step in English, then write the final answer in . [User] Качки Дженет несуть по 16 яєць на день. Вона їсть три на снiданок щоранку i пече кекси для своїх друзiв щодня, використовуючи чотири. Решту вона щодня продає на фермерсько- му ринку по 2 $ за свiже качине яйце. Скiльки в доларах вона заробляє на фермерському ринку щодня? Backtranslated (Ukrainian→ English via MT) Janet’s ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers’ market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers’ market? Note: Please put the final answer in . E.3 Per-Language Closing Instructions To keep the input fully monolingual under in Base prompt, we translated the closing instructions recom- mended in the original PolyMat (Wang et al., 2025) to keep the final output structured. Native speakers verified each translation. We list all 18 target-language directives in Table 10. Code Language Translation hi Hindi नोट: क ृ पया अं+तम उ/र को $ $ म 1 रख 1 । tr Turkish Not: Lütfen son cevabı $ $ içine yazın. pl Polish Uwaga: Proszę umieścić końcową odpowiedź w $ $. uk Ukrainian Примітка: Будь ласка, помістіть остаточну відповідь у $ $. uz Uzbek Eslatma: Iltimos, yakuniy javobni $ $ ichiga yozing. or Odia ଟିପ$ଣୀ: ଦୟାକରି େଶଷ ଉ0ରକୁ $ $ ଭ3ତେର ରଖ67। am Amharic ማስታወሻ፡ እባክዎ የመጨረሻውን መልስ በ$ $ ውስጥ ያስገቡ። el Greek Σημείωση: Παρακαλώ βάλτε την τελική απάντηση μέσα στο $ $. k Kazakh Ескерту: Соңғы жауапты $ $ ішіне жазыңыз. cs Czech Poznámka: Prosím vložte konečnou odpověď do $ $. he Hebrew ך ו תב ת י פ ו סה הב ו שתה תא ם י של א נ : הר עה $ $. sr Serbian Напомена: Молимо ставите коначан одговор у $ $. t Tatar Искәрмә: Зинһар, соңгы җавапны $ $ эченә куегыз. sk Slovak Poznámka: Prosím vložte konečnú odpoveď do $ $. ca Catalan Nota: Si us plau, poseu la resposta final dins de $ $. cv Chuvash Асӑрхатару: Юлашки хуравăрăн $ $ ӑшне лартăр. hsb Upper Sorbian Kedźbu: Prošu stajće kónčnu wotmołwu do $ $. dsb Lower Sorbian Glědajśo: Pšosym stajśo kóńcnu wótegrono do $ $. Figure 10: The translated per language final instructions which were added in the end of main tasks. English original: Note: Please put the final answer in the $ $. F Hyperparameter Search We performed a two-stage hyperparameter search before running the main multilingual evaluation. All scores in this appendix are reported with the metrics described in Section 4.2: per-level scores are exact- match accuracies of the content of the$ $expression, and the rightmost column of each table reports PolyMath’s difficulty-weighted accuracy (DW-ACC) aggregated over all four difficulty levels with weights1, 2, 4, 8. Stage 1 – choosing the reasoning effort (English only).We first searched for an appropriate reasoning effort setting using only the English split of PolyMath. For each of the three reasoning models considered, we ran the supported reasoning-effort levels (low, medium, high) at a fixed sampling temperature. Results are reported in Table 8. Stage 2 – temperature sweep on high-resource languages.Using the reasoning-effort settings selected in Stage 1, we then ran a narrower temperature sweep on three high-resource languages (English, German, Russian). The resulting configurations and per-language scores are shown in Table 9. The low+temperature= 0.5configuration was run on English only as a sanity check on temperature sensitivity and was not extended to the other two languages. Reasoning effortModelLowMediumHighTopDW-ACC Low gpt-oss-120b79.239.228.013.625.2 Qwen3-235B93.64.00.00.06.8 Qwen3-30B92.87.23.20.88.4 Medium gpt-oss-120b77.644.831.29.624.6 Qwen3-235B92.83.20.00.06.6 Qwen3-30B93.67.23.20.88.5 High gpt-oss-120b83.232.013.64.015.6 Qwen3-235B93.64.80.00.06.9 Qwen3-30B89.67.22.40.88.0 Table 8: Effect of reasoning effort on English-only PolyMath performance. Per-level scores are exact-match accuracies (%) of the content of ; the DW-ACC column reports the difficulty-weighted accuracy aggregated over all four levels with weights1, 2, 4, 8. EnglishGermanRussian EffortTemp.ModelLowMedHighTopDWLowMedHighTopDWLowMedHighTopDW Low0.1 gpt-oss-120b79.239.228.013.625.261.632.829.611.222.366.436.827.212.823.4 Qwen3-235B93.64.00.00.06.887.29.61.60.07.591.211.24.00.08.6 Qwen3-30B92.87.23.20.88.487.211.24.00.88.889.68.84.80.88.9 Low0.3 gpt-oss-120b79.240.027.28.822.663.236.831.28.822.169.641.628.816.026.4 Qwen3-235B92.83.20.00.06.687.28.81.60.07.492.011.24.80.08.9 Qwen3-30B91.27.21.60.87.987.211.24.00.88.888.012.04.00.89.0 Low0.5 gpt-oss-120b80.038.428.810.423.7– Qwen3-235B93.63.20.00.06.7– Qwen3-30B93.65.62.40.88.1– Medium0.1 gpt-oss-120b77.644.831.29.624.662.443.232.08.022.763.241.628.08.821.9 Qwen3-235B92.83.20.00.06.688.09.61.60.07.692.810.44.00.08.6 Qwen3-30B93.67.23.20.88.587.210.45.60.89.188.810.43.20.88.6 Medium0.3 gpt-oss-120b79.241.628.89.623.665.643.228.89.622.958.442.427.28.821.5 Qwen3-235B93.63.20.00.06.785.67.20.80.06.991.28.84.80.08.5 Qwen3-30B91.26.43.20.88.286.412.04.00.88.989.612.84.00.89.2 Table 9: Hyperparameter search on high-resource languages (English, German, Russian) across all four PolyMath difficulty levels. Each configuration combines a reasoning-effort setting (low or medium) with a sampling temper- ature. Per-level scores are exact-match accuracies (%); the DW column reports the difficulty-weighted accuracy aggregated over all four levels with weights1, 2, 4, 8(denominator15). A dash (–) marks configurations not run for the given language. G Closed Models Inference Budget First, we utilized commercial closed-source APIs for precomputing translations. For 10 PLURAMATH languages, we used the DeepL API at a total cost of ($63.82), while Gemini API usage across three models amounted to approximately ($11.60). We estimate the cost of the benchmark from output-token pricing, which dominates spend for reasoning models. The dataset containsL = 4difficulty levels withT = 125tasks each, givingN = L· T = 500 problems. We cap every reasoning trace atB = 2,000output tokens, so a single model produces N·B = 5×10 2 ·2×10 3 = 10 6 output tokens per language, i.e. exactly one million tokens. The benchmark covers ℓ = 22 languages and p = 3 prompt designs. Letc m be the price (USD) per one million output tokens for modelm. Because each model emits exactly10 6 tokens per language, the cost of evaluatingmunder a single prompt design across all languages reduces to C 1 m = c m · N · B· ℓ 10 6 = c m · ℓ = 22c m , and a full three-design sweep costs C 3 m = p· ℓ· c m = 66c m . Evaluation protocol. To keep the budget tractable we run the full three-design sweep only for the open-weight mid-sized models and the two cheapest large open-weight models, where the marginal cost is negligible. The remaining large models—including all proprietary endpoints—are evaluated under the single base prompt design. Table 10 reports the per-model breakdown. Under this protocol the adopted budget is$236.94(full thee-prompts design group for open-source models)+ $724.90(base-only group for bigger and closed models) = $961.84. ModelAccess$/1M out1 design (×22)3 designs (×66) Full three-prompts design GPT-OSS-20BOpen0.143.089.24 Gemma-4-31B-itOpen0.388.3625.08 Qwen3.5-35B-A3BOpen1.0022.0066.00 DeepSeek-R1-Distill-70BOpen0.8017.6052.80 Nemotron-3-Nano-30B-A3BOpen0.8017.6052.80 GPT-OSS-120BOpen0.194.1812.54 DeepSeek-V3.2Open0.286.1618.48 Subtotal (full sweep)236.94 Base prompt design only Qwen3.5-122B-A10BOpen2.4052.80— Kimi-K2.5Open2.2549.50— Qwen3-235B-A22B-ThinkingOpen2.3050.60— Claude-Haiku-4.5Closed16.00352.00— Gemini-2.5-FlashClosed2.5055.00— GPT-5.4 † Closed15.00165.00— Subtotal (base only)724.90 Adopted budget961.84 Table 10: Estimated inference budget (USD). Costs are derived from output-token pricing: each model emits N·B = 10 6 tokens per language (N = 500problems,B = 2,000tokens per trace), so the cost of one prompt design acrossℓ = 22languages is22c m and a full three-design sweep is66c m . Large and proprietary models are evaluated under a single base prompt design only. † Batched decoding halves the GPT-5.4 line item from $330.00 to $165.00. H Detailed Per-Difficulty-Level and Per-Prompting-Ablations Results Here, we present details results per levels and prompts designs as well as additional results visualizations. Per difficulty levels results: Appendix H.1, models’ reasoning length comparison: Appendix H.2, per prompting types results: Appendix H.3. H.1 Per-Difficulty-Level Results Results for each difficulty level in terms of math answer correctness: low (Table 11), medium (Table 12), high (Table 13), and top (Table 14). Additionally, we put here box-plots results comparison per language (Figure 11) and per model (Figure 12). Overall, most models achieve relatively strong performance on low-level tasks, followed by a substantial decline on higher difficulty levels. Even at the lowest level, Odia, Amharic, and Chuvash remain the most challenging underrepresented languages. For higher difficulty levels, only thegpt-ossvariants and the proprietary GPT-5.4 and Claude-Haiku-4.5 models maintain consistently non-zero performance. High-resource Summary (HR) Our target languages Summary (target) Model en de ru es Avg. Len Lang hi tr pl uk uz or am el k cs he sr t sk ca cv hsb dsb Avg. Len Lang Open-weight — small ( ≤ 4B) Qwen3.5-0.8B 52.8 93 28.8 98 4.8 15 40.0 96 31.6 518 ± 689 TL 5.6 77 5.6 43 8.0 55 15.2 92 0.8 19 4.0 66 0.8 16 17.6 93 5.6 66 19.2 92 16.8 82 8.0 63 7.2 56 14.4 80 4.8 69 0.8 10 0.8 29 2.4 41 7.6 1499 ± 705 TL LFM2.5-1.2B 0.0 0 0.0 0 0.0 0 0.0 0 0.0 3409 ± 1403 EN 0.0 0 10.4 41 43.2 58 0.0 0 0.8 98 0.0 0 0.8 42 0.0 0 0.0 0 0.0 0 0.0 0 2.4 32 0.0 0 0.0 0 20.8 52 0.0 0 0.0 0 0.0 0 4.4 2720 ± 868 EN Ouro-1.4B 90.4 93 73.6 82 67.2 73 82.4 88 78.4 748 ± 481 TL 28.0 50 32.8 54 62.4 75 21.6 36 3.2 37 1.6 18 0.8 57 21.6 41 0.8 10 32.0 51 18.4 51 16.8 33 2.4 13 28.8 50 54.4 71 3.2 18 6.4 35 4.0 25 18.8 1267 ± 471 EN R1-Distill-Qwen-1.5B 48.8 90 0.8 66 0.8 50 5.6 86 14.0 1232 ± 974 TL 0.0 38 0.0 39 0.0 55 0.0 46 4.0 72 0.0 13 0.0 26 0.0 51 0.0 39 0.0 54 0.0 35 11.2 40 0.0 55 0.0 36 4.8 54 0.8 54 0.8 58 0.8 70 1.2 2235 ± 1110 EN Qwen3.5-2B 28.8 55 28.0 84 29.6 86 32.0 77 29.6 1847 ± 282 TL 7.2 57 24.8 75 36.0 94 35.2 82 20.0 74 8.0 56 2.4 15 32.8 75 16.0 57 30.4 89 29.6 79 32.0 87 10.4 58 31.2 82 35.2 83 0.0 37 5.6 49 1.6 77 19.9 1978 ± 225 TL Ouro-2.6B 93.6 94 76.8 88 84.0 90 88.0 93 85.6 751 ± 439 TL 49.6 69 64.0 78 0.0 0 65.6 74 18.4 46 12.0 46 1.6 26 56.8 73 6.4 31 55.2 69 28.0 48 47.2 66 6.4 36 53.6 68 72.0 82 8.0 38 20.0 47 17.6 44 32.4 1160 ± 503 EN Ministral-3-3B 89.6 100 68.8 99 76.8 99 84.0 99 79.8 315 ± 191 TL 59.2 97 37.6 42 73.6 99 72.0 100 17.6 19 0.0 71 0.0 46 64.0 95 48.8 95 60.8 99 63.2 98 33.6 36 33.6 74 62.4 94 76.0 100 4.8 96 20.8 93 19.2 93 41.5 767 ± 434 TL Gemma-3-4B 93.6 100 68.8 89 20.0 27 84.0 94 66.6 522 ± 463 TL 67.2 90 49.6 72 76.0 91 75.2 91 53.6 81 24.0 62 45.6 78 77.6 94 24.0 40 67.2 92 73.6 93 74.4 90 35.2 74 72.0 94 76.8 90 3.2 17 24.0 76 16.8 67 52.0 778 ± 822 TL Qwen3.5-4B 44.0 89 49.6 95 52.8 92 46.4 92 48.2 1763 ± 335 TL 36.0 92 45.6 96 51.2 98 43.2 97 42.4 97 24.8 90 37.6 95 48.8 96 46.4 95 41.6 93 56.8 98 49.6 99 42.4 95 49.6 90 54.4 95 16.0 86 18.4 94 20.8 93 40.3 1852 ± 311 TL Open-weight — mid (7–35B) OLMo-3-7B-Think 72.0 74 64.8 71 38.4 43 79.2 82 63.6 1557 ± 457 TL 27.2 38 49.6 58 67.2 74 53.6 61 0.8 26 21.6 32 0.8 10 40.0 47 17.6 32 36.0 48 36.0 45 3.2 6 4.0 60 35.2 46 58.4 63 0.8 67 5.6 41 2.4 29 25.6 1846 ± 274 EN R1-0528-Qwen3-8B 56.0 70 44.0 66 0.0 2 30.4 63 32.6 2918 ± 749 TL 0.0 87 8.8 37 11.2 39 0.0 95 4.0 56 0.0 91 0.0 88 0.0 98 0.0 92 4.8 86 0.0 94 0.0 83 0.0 89 1.6 76 10.4 62 0.0 80 0.0 44 0.0 62 2.3 3044 ± 612 EN Ministral-3-8B 95.2 100 77.6 100 84.8 99 92.8 100 87.6 310 ± 148 TL 80.8 99 83.2 100 88.8 99 87.2 99 75.2 98 0.0 50 0.0 20 82.4 100 66.4 97 72.0 99 84.8 100 88.8 99 64.0 93 80.0 100 82.4 100 4.8 71 40.0 88 31.2 84 61.8 502 ± 358 EN Qwen3.5-9B 53.6 92 39.2 97 56.8 98 56.0 89 51.4 2746 ± 1167 TL 41.6 98 52.8 98 50.4 98 55.2 98 51.2 94 38.4 93 48.0 94 56.8 99 59.2 98 46.4 96 61.6 98 52.0 99 52.8 100 52.0 98 52.8 98 42.4 97 27.2 98 31.2 96 48.4 2497 ± 846 TL Ministral-3-14B 94.4 99 85.6 100 65.6 68 92.0 99 84.4 633 ± 267 TL 68.8 98 66.4 68 77.6 78 56.8 64 50.4 62 1.6 19 0.0 7 85.6 100 68.8 95 75.2 100 85.6 99 48.8 51 35.2 54 47.2 55 62.4 66 12.0 86 46.4 94 42.4 79 51.7 1138 ± 504 TL gpt-oss-20b 89.6 100 73.6 99 88.0 99 88.0 98 84.8 545 ± 596 EN 83.2 98 89.6 97 84.8 98 87.2 99 84.8 98 77.6 97 48.0 78 82.4 100 83.2 100 75.2 100 84.0 99 83.2 97 80.8 96 79.2 98 82.4 94 11.2 86 56.8 91 54.4 93 73.8 973 ± 836 EN Nemotron3-Nano-30B 74.4 90 60.0 82 63.2 75 60.8 77 64.6 628 ± 1052 TL 55.2 82 54.4 78 72.8 90 36.8 66 22.4 81 1.6 46 0.0 54 56.8 78 12.8 44 60.8 88 74.4 89 64.0 92 5.6 50 28.8 63 64.8 86 0.8 33 12.0 41 12.0 39 35.3 882 ± 1162 EN Gemma-4-31B 83.2 85 77.6 95 88.0 95 88.8 94 84.4 832 ± 405 TL 86.4 98 88.8 98 83.2 94 82.4 96 88.0 97 78.4 93 84.8 95 78.4 91 82.4 97 71.2 92 80.8 96 85.6 96 74.4 89 80.8 94 80.8 93 33.6 73 62.4 84 55.2 76 76.5 1014 ± 429 TL Qwen3.5-35B-A3B 54.4 92 50.4 97 56.0 95 60.8 94 55.4 1636 ± 389 TL 43.2 94 55.2 99 56.8 99 60.8 98 63.2 96 32.8 94 52.0 95 56.8 99 48.0 98 46.4 92 54.4 98 54.4 98 43.2 100 54.4 99 54.4 94 25.6 95 30.4 98 39.2 94 48.4 1771 ± 357 TL Open-weight — large / API R1-Distill-Llama-70B 84.8 98 63.2 100 90.4 98 77.6 100 79.0 569 ± 270 TL 73.6 100 77.6 99 83.2 99 82.4 100 73.6 100 49.6 78 8.0 90 77.6 100 72.0 100 73.6 99 77.6 100 81.6 100 65.6 95 73.6 100 72.0 100 45.6 91 57.6 98 54.4 95 66.6 711 ± 303 TL gpt-oss-120b 77.6 100 62.4 100 64.0 100 74.4 100 69.6 400 ± 238 TL 72.8 100 76.8 100 65.6 100 72.8 100 80.8 100 70.4 100 69.6 98 59.2 100 84.0 100 54.4 100 72.8 100 75.2 100 76.0 100 60.8 100 64.0 100 16.0 100 54.4 100 54.4 100 65.6 594 ± 350 EN Qwen3.5-122B-A10B 69.6 94 51.2 96 60.0 98 56.0 98 59.2 2068 ± 1080 TL 52.0 95 60.0 97 56.8 99 60.8 97 60.8 95 41.6 98 57.6 99 65.6 98 54.4 98 49.6 97 65.6 100 61.6 99 46.4 99 56.0 100 63.2 98 43.2 99 47.2 99 44.0 98 54.8 2144 ± 813 TL Qwen3-235B-A22B 92.8 94 88.0 98 92.8 98 94.4 98 92.0 1262 ± 817 TL 87.2 92 84.8 86 83.2 86 89.6 94 76.8 80 77.6 81 56.8 58 91.2 94 91.2 96 77.6 91 94.4 98 82.4 83 80.0 93 84.8 94 84.8 87 17.6 38 77.6 95 72.8 86 78.4 1700 ± 841 TL DeepSeek-V3.2 96.0 99 85.6 94 83.2 93 89.6 93 88.6 748 ± 566 TL 88.0 94 88.8 95 92.8 95 79.2 94 80.8 88 80.0 86 79.2 84 86.4 94 82.4 93 76.0 89 90.4 94 91.2 94 69.6 86 85.6 94 89.6 92 65.6 88 83.2 88 81.6 90 82.8 972 ± 607 TL Kimi-K2.5 82.4 92 45.6 92 52.8 91 64.0 87 61.2 980 ± 631 TL 46.4 88 68.0 94 62.4 94 52.0 94 52.0 90 53.6 81 51.2 85 56.0 85 55.2 98 45.6 88 54.4 94 52.0 97 44.8 94 60.0 87 58.4 87 30.4 90 57.6 94 51.2 89 52.8 1232 ± 617 TL Closed-source Claude-Haiku-4.5 83.2 100 60.0 100 53.6 100 59.2 100 64.0 222 ± 87 TL 79.2 100 74.4 100 56.8 100 69.6 100 76.8 100 81.6 99 85.6 100 68.0 100 85.6 100 49.6 100 78.4 100 60.0 100 74.4 100 53.6 100 62.4 100 45.6 97 64.8 100 71.2 99 68.8 327 ± 137 TL Gemini-2.5-Flash 82.4 94 76.8 94 71.2 96 84.8 96 78.8 827 ± 349 TL 79.2 97 85.6 98 84.0 97 87.2 95 84.8 93 73.6 90 86.4 96 75.2 96 89.6 98 71.2 94 83.2 96 87.2 97 78.4 92 77.6 92 76.8 97 82.4 94 68.0 93 59.2 91 79.4 982 ± 384 TL GPT-5.4 72.8 100 69.6 100 73.6 100 72.0 100 72.0 128 ± 46 TL 71.2 100 85.6 100 82.4 100 82.4 100 84.0 100 67.2 100 75.2 100 83.2 100 80.8 100 68.8 100 92.0 100 89.6 100 75.2 100 76.8 100 77.6 100 80.8 100 76.0 100 73.6 100 79.0 183 ± 65 TL Table 11: Low-difficulty base-prompting results on our P LURA M ATH languages vs high-resource ones from PolyMath. Each cell shows answer accuracy (%; large) over the answer-format compliance rate (%; small). Cell shading encodes accuracy on a tier-specific scale (pale → deep teal, 0 → 100% ); a numeral is printed in black on light (low-score) cells and in white on dark (high-score) cells solely for legibility—the text colour carries no extra meaning. Best and second-best per column are highlighted. After each language block we report the macro-average accuracy, the mean ± std generation length (tokens, reasoning+answer), and the dominant answer language coded EN (predominantly English) or TL (requested target language). High-resource Summary (HR) Our target languages Summary (target) Model en de ru es Avg. Len Lang hi tr pl uk uz or am el k cs he sr t sk ca cv hsb dsb Avg. Len Lang Open-weight — small ( ≤ 4B) Qwen3.5-0.8B 7.2 27 2.4 42 3.2 18 6.4 40 4.8 2753 ± 1136 TL 0.0 25 0.0 11 0.0 11 3.2 26 0.0 6 0.8 27 0.0 11 2.4 26 0.0 7 2.4 34 1.6 27 0.0 4 1.6 17 2.4 30 0.0 10 0.0 2 2.4 25 1.6 44 1.0 2420 ± 594 OTH LFM2.5-1.2B 0.0 0 0.0 0 0.0 0 0.0 0 0.0 3552 ± 986 EN 0.0 0 2.4 10 4.0 17 0.0 0 1.6 75 0.0 0 3.2 23 0.0 0 0.0 0 0.0 0 0.0 0 3.2 23 0.0 0 0.0 0 2.4 13 0.0 0 0.0 0 0.0 0 0.9 2710 ± 764 EN Ouro-1.4B 4.8 8 10.4 16 4.0 7 6.4 11 6.4 1671 ± 196 TL 4.8 7 4.8 10 4.8 7 0.0 2 2.4 10 2.4 5 2.4 23 3.2 6 2.4 8 3.2 8 3.2 8 2.4 5 3.2 12 3.2 6 5.6 10 0.8 12 2.4 10 4.0 13 3.1 1712 ± 209 EN R1-Distill-Qwen-1.5B 2.4 7 0.8 14 0.8 4 0.8 8 1.2 3134 ± 373 EN 0.8 16 2.4 11 0.8 11 0.0 22 1.6 12 0.8 18 4.0 12 1.6 19 1.6 15 3.2 18 0.0 15 0.8 4 0.8 14 0.8 9 1.6 7 4.0 20 0.8 10 0.0 11 1.4 3280 ± 399 EN Qwen3.5-2B 0.8 3 1.6 5 0.0 6 0.0 4 0.6 1959 ± 31 TL 0.0 14 0.8 3 1.6 8 0.0 2 0.8 9 0.0 20 0.8 8 0.8 4 0.0 6 0.8 8 0.0 9 0.0 11 0.8 11 0.8 5 0.0 4 0.0 42 0.0 18 0.0 32 0.4 1977 ± 50 TL Ouro-2.6B 8.0 14 7.2 12 8.0 11 8.0 13 7.8 1815 ± 200 EN 8.0 12 7.2 10 0.0 0 8.8 14 4.0 10 7.2 14 1.6 14 6.4 10 1.6 9 8.8 15 5.6 12 6.4 10 4.0 13 7.2 12 3.2 8 2.4 17 7.2 15 4.8 13 5.2 1704 ± 188 EN Ministral-3-3B 15.2 34 12.0 34 15.2 53 18.4 62 15.2 1621 ± 590 TL 14.4 32 0.0 2 13.6 30 13.6 55 0.0 2 4.8 33 0.0 5 10.4 24 8.8 42 12.8 29 14.4 34 0.0 0 5.6 22 11.2 25 13.6 62 8.0 42 8.0 32 11.2 28 8.4 1846 ± 446 EN Gemma-3-4B 19.2 82 14.4 84 12.8 74 12.8 77 14.8 1508 ± 446 TL 10.4 78 12.0 77 13.6 75 12.0 84 8.0 78 4.8 43 7.2 69 12.8 80 7.2 48 13.6 82 13.6 78 12.8 75 6.4 43 12.8 73 13.6 85 4.8 53 8.8 71 8.0 65 10.1 1659 ± 633 TL Qwen3.5-4B 0.0 2 1.6 9 0.0 10 0.8 7 0.6 1964 ± 13 TL 1.6 33 0.8 7 1.6 10 1.6 7 1.6 14 0.8 41 0.8 26 2.4 10 0.8 11 2.4 12 1.6 10 0.8 10 1.6 22 3.2 13 0.8 9 2.4 34 0.0 54 0.0 28 1.4 1971 ± 38 TL Open-weight — mid (7–35B) OLMo-3-7B-Think 0.8 2 3.2 4 0.8 2 0.8 4 1.4 2011 ± 106 EN 0.0 2 0.0 2 0.8 2 1.6 5 0.0 8 0.0 2 0.0 10 0.8 1 0.8 6 0.8 5 0.0 2 0.0 1 0.0 24 0.0 2 0.8 2 0.0 28 0.8 13 0.0 18 0.4 2009 ± 92 EN R1-0528-Qwen3-8B 2.4 4 1.6 3 0.0 1 1.6 7 1.4 3159 ± 250 TL 0.0 34 0.0 4 0.0 1 0.8 31 0.0 9 0.0 34 0.0 21 0.0 33 0.0 33 0.8 15 0.0 38 0.0 22 0.0 34 0.8 14 0.8 10 0.0 21 0.8 19 0.8 21 0.3 3174 ± 252 EN Ministral-3-8B 13.6 28 15.2 30 10.4 26 20.8 42 15.0 1788 ± 512 EN 13.6 42 13.6 32 16.8 51 12.8 30 15.2 50 9.6 37 10.4 33 12.8 43 4.8 34 13.6 26 16.8 41 16.8 33 11.2 26 11.2 30 16.8 38 4.8 20 14.4 38 14.4 34 12.8 1778 ± 550 EN Qwen3.5-9B 0.8 5 5.6 22 7.2 20 4.0 16 4.4 3888 ± 195 TL 8.0 43 0.0 8 2.4 8 8.0 22 4.8 16 6.4 48 1.6 22 4.0 23 5.6 27 5.6 23 4.8 25 1.6 7 8.0 42 6.4 22 0.8 14 3.2 59 1.6 63 0.0 66 4.0 3274 ± 166 TL Ministral-3-14B 15.2 30 22.4 59 0.0 2 22.4 58 15.0 1793 ± 418 TL 16.8 31 0.8 3 0.8 5 0.8 1 0.8 5 0.0 2 0.0 4 11.2 30 12.8 26 17.6 36 16.8 47 0.0 2 0.0 11 0.0 3 1.6 2 8.8 30 8.8 18 6.4 19 5.8 1995 ± 320 EN gpt-oss-20b 34.4 61 30.4 60 34.4 57 30.4 59 32.4 3019 ± 1376 EN 28.8 63 22.4 40 22.4 37 32.0 59 20.0 34 30.4 61 20.0 37 31.2 58 27.2 54 29.6 56 28.0 58 18.4 35 30.4 54 28.8 57 20.8 38 16.0 66 27.2 58 25.6 62 25.5 2699 ± 1074 EN Nemotron3-Nano-30B 20.0 45 15.2 35 17.6 35 9.6 35 15.6 2389 ± 1661 EN 16.8 35 11.2 46 10.4 34 9.6 42 4.0 54 2.4 35 8.8 58 17.6 44 8.0 35 13.6 44 17.6 50 12.8 63 8.8 44 12.8 47 10.4 37 5.6 38 13.6 34 8.0 38 10.7 1591 ± 1421 EN Gemma-4-31B 3.2 6 8.8 14 12.0 17 12.0 16 9.0 1924 ± 190 TL 10.4 16 11.2 17 9.6 14 12.8 17 7.2 12 5.6 10 9.6 15 8.0 13 8.0 14 9.6 15 8.8 15 10.4 15 8.8 16 7.2 12 10.4 14 8.0 19 5.6 16 6.4 23 8.8 1944 ± 199 TL Qwen3.5-35B-A3B 0.8 2 1.6 11 1.6 7 0.8 6 1.2 1965 ± 12 TL 0.0 14 0.0 6 0.0 7 2.4 9 2.4 16 0.8 36 3.2 29 2.4 10 4.0 16 0.8 8 1.6 8 0.8 10 1.6 24 1.6 14 1.6 13 1.6 52 4.0 57 1.6 78 1.7 1973 ± 32 TL Open-weight — large / API R1-Distill-Llama-70B 5.6 12 4.8 11 14.4 22 6.4 14 7.8 2022 ± 187 TL 8.8 16 4.8 10 8.0 12 12.0 26 6.4 14 8.0 18 5.6 16 8.8 27 10.4 14 8.0 14 5.6 14 7.2 18 9.6 22 4.8 11 6.4 12 12.0 26 8.0 16 8.8 16 8.0 2003 ± 228 EN gpt-oss-120b 44.8 74 43.2 73 41.6 74 36.0 73 41.4 2712 ± 1302 TL 32.0 65 22.4 45 21.6 45 34.4 66 24.8 42 32.8 65 23.2 41 35.2 72 36.0 72 32.8 71 36.8 73 24.8 45 34.4 66 36.8 67 24.8 43 17.6 78 35.2 78 31.2 71 29.8 2487 ± 1023 EN Qwen3.5-122B-A10B 0.8 10 6.4 22 7.2 19 1.6 14 4.0 3883 ± 207 TL 4.0 29 1.6 8 0.8 7 3.2 22 1.6 11 8.8 43 1.6 16 5.6 29 6.4 24 4.0 22 3.2 20 1.6 9 6.4 38 4.0 26 1.6 11 4.8 54 4.8 71 8.8 78 4.0 3269 ± 191 TL Qwen3-235B-A22B 3.2 9 9.6 14 10.4 18 8.8 17 8.0 3830 ± 539 TL 7.2 13 3.2 4 1.6 4 9.6 18 3.2 6 7.2 16 4.0 5 6.4 13 8.0 14 9.6 17 8.8 15 4.0 6 8.8 18 11.2 17 4.0 5 9.6 19 14.4 19 11.2 18 7.3 3206 ± 399 TL DeepSeek-V3.2 15.2 22 17.6 23 15.2 25 15.2 26 15.8 2009 ± 348 TL 15.2 28 16.0 28 16.8 29 15.2 24 12.0 18 16.0 26 14.4 27 16.8 30 12.8 23 13.6 24 18.4 32 14.4 31 12.8 18 20.0 30 13.6 24 15.2 28 16.0 38 16.8 30 15.3 2014 ± 322 TL Kimi-K2.5 7.2 12 7.2 10 4.0 11 6.4 10 6.2 2083 ± 155 TL 3.2 16 6.4 14 8.0 20 8.0 20 7.2 14 8.8 22 5.6 30 8.0 21 5.6 19 5.6 21 5.6 14 8.0 19 5.6 26 4.8 21 4.0 12 4.8 41 6.4 31 5.6 34 6.2 2093 ± 146 TL Closed-source Claude-Haiku-4.5 49.6 100 44.8 100 43.2 99 45.6 100 45.8 1018 ± 256 TL 43.2 98 49.6 100 39.2 99 48.8 99 44.0 100 42.4 99 45.6 98 45.6 100 44.0 99 41.6 99 42.4 99 42.4 100 44.8 100 40.0 100 46.4 99 31.2 95 46.4 100 39.2 100 43.2 1052 ± 275 TL Gemini-2.5-Flash 2.4 4 4.0 6 3.2 4 3.2 4 3.2 4021 ± 4045 TL 4.0 6 4.0 7 4.8 7 4.8 6 2.4 5 4.0 5 2.4 3 2.4 5 4.8 6 2.4 5 4.0 5 4.0 6 4.0 4 5.6 7 2.4 4 4.8 7 5.6 7 4.0 7 3.9 4292 ± 4199 TL GPT-5.4 32.8 100 34.4 100 32.0 100 32.8 100 33.0 513 ± 218 TL 28.8 100 32.8 100 27.2 100 28.0 100 35.2 100 33.6 100 27.2 100 31.2 100 28.8 100 28.0 100 26.4 100 30.4 100 28.8 100 32.0 100 34.4 100 29.6 100 29.6 100 29.6 100 30.1 566 ± 229 TL Table 12: Medium-difficulty base-prompting results on our P LURA M ATH languages vs high-resource ones from PolyMath. Each cell shows answer accuracy (%; large) over the answer-format compliance rate (%; small). Cell shading encodes accuracy on a tier-specific scale (pale → deep teal, 0 → 50% ); a numeral is printed in black on light (low-score) cells and in white on dark (high-score) cells solely for legibility—the text colour carries no extra meaning. Best and second-best per column are highlighted. After each language block we report the macro-average accuracy, the mean ± std generation length (tokens, reasoning+answer), and the dominant answer language coded EN (predominantly English) or TL (requested target language). High-resource Summary (HR) Our target languages Summary (target) Model en de ru es Avg. Len Lang hi tr pl uk uz or am el k cs he sr t sk ca cv hsb dsb Avg. Len Lang Open-weight — small ( ≤ 4B) Qwen3.5-0.8B 0.8 15 1.6 43 0.0 10 0.8 39 0.8 2885 ± 1030 TL 0.0 28 0.0 2 0.0 5 0.0 19 0.0 2 0.0 37 0.0 6 0.8 21 0.0 12 0.8 30 0.8 25 0.0 3 0.0 6 0.0 24 0.0 3 0.8 9 0.0 14 0.0 48 0.2 2510 ± 602 TL LFM2.5-1.2B 0.0 0 0.0 0 0.0 0 0.0 0 0.0 3510 ± 1229 OTH 0.0 0 0.0 4 0.8 6 0.0 0 0.8 80 0.0 0 1.6 20 0.0 0 0.0 0 0.0 0 0.0 0 0.0 15 0.0 0 0.0 0 1.6 7 0.0 0 0.0 0 0.0 0 0.3 2751 ± 736 OTH Ouro-1.4B 2.4 2 2.4 2 2.4 2 3.2 3 2.6 1743 ± 154 TL 0.0 2 1.6 2 1.6 2 0.0 1 0.0 4 0.0 4 0.8 24 0.8 3 0.0 8 0.0 4 0.0 6 0.0 1 0.0 5 0.0 4 1.6 2 0.0 13 0.0 7 0.8 10 0.4 1736 ± 186 EN R1-Distill-Qwen-1.5B 1.6 2 0.0 7 0.0 6 0.0 14 0.4 3171 ± 485 EN 0.0 17 0.0 7 0.0 11 0.0 20 0.0 8 0.0 20 0.0 5 0.0 15 0.0 16 0.8 7 0.0 21 0.0 2 1.6 18 0.0 16 0.0 10 0.8 22 0.0 10 0.0 19 0.2 3270 ± 442 EN Qwen3.5-2B 0.0 0 0.0 2 0.0 0 0.0 0 0.0 1959 ± 17 – 0.0 7 0.0 0 0.0 3 0.0 1 0.0 3 0.8 14 0.0 3 0.0 0 0.0 2 0.0 5 0.0 3 0.0 1 0.0 1 0.0 4 0.0 1 0.0 26 0.0 10 0.0 23 0.0 1991 ± 63 – Ouro-2.6B 3.2 3 2.4 2 3.2 3 3.2 3 3.0 1896 ± 55 EN 0.0 1 1.6 2 3.2 3 0.8 1 0.0 1 0.0 2 0.0 17 3.2 3 0.0 5 1.6 4 0.0 1 0.8 3 0.8 13 0.8 1 1.6 2 0.8 16 1.6 9 0.8 6 1.0 1806 ± 163 EN Ministral-3-3B 4.0 7 6.4 10 5.6 10 9.6 54 6.4 1912 ± 268 TL 7.2 12 2.4 2 4.8 6 3.2 22 0.0 2 2.4 22 0.0 17 6.4 8 1.6 36 4.0 10 2.4 6 0.0 1 3.2 23 4.0 9 5.6 30 1.6 30 4.0 17 0.0 13 2.9 1974 ± 263 EN Gemma-3-4B 5.6 78 9.6 76 4.0 73 6.4 68 6.4 1537 ± 416 TL 3.2 56 4.0 58 3.2 69 6.4 74 4.8 54 0.8 30 1.6 57 4.8 62 2.4 22 4.0 65 0.8 62 5.6 65 1.6 22 6.4 66 4.8 77 1.6 38 1.6 77 2.4 54 3.3 1932 ± 648 TL Qwen3.5-4B 0.0 0 0.0 0 0.0 0 0.0 0 0.0 1960 ± 13 – 0.0 12 0.0 0 0.8 1 0.0 0 0.0 2 0.0 18 0.8 9 0.0 0 0.0 2 0.0 1 0.0 2 0.8 1 0.0 9 0.0 1 0.0 0 0.0 21 0.0 24 0.0 20 0.1 1970 ± 30 – Open-weight — mid (7–35B) OLMo-3-7B-Think 0.8 1 0.0 0 0.0 0 0.8 1 0.4 2008 ± 115 EN 0.0 3 0.0 0 0.0 1 0.0 0 0.0 22 0.0 2 0.0 10 0.0 1 0.0 2 0.0 1 0.0 2 0.0 3 0.0 33 0.0 1 0.0 4 0.0 39 0.0 21 0.0 24 0.0 2010 ± 116 EN R1-0528-Qwen3-8B 0.0 2 0.0 3 0.0 0 0.0 4 0.0 3299 ± 267 – 0.0 41 0.0 10 0.0 7 0.0 36 0.0 14 0.0 47 0.0 34 0.0 38 0.0 45 0.0 21 0.0 40 0.0 34 0.0 38 0.0 24 0.0 8 0.0 44 0.0 29 0.0 37 0.0 3250 ± 304 EN Ministral-3-8B 6.4 8 5.6 9 6.4 11 5.6 11 6.0 1992 ± 190 EN 8.0 24 7.2 11 5.6 39 5.6 14 4.0 23 0.8 12 1.6 17 6.4 22 1.6 14 4.8 10 6.4 10 4.0 10 3.2 10 4.0 12 5.6 10 1.6 18 2.4 10 4.0 14 4.3 1947 ± 318 EN Qwen3.5-9B 0.0 0 0.8 4 0.8 1 0.8 2 0.6 3915 ± 43 TL 0.8 18 0.0 0 0.0 1 0.8 1 0.0 4 0.0 17 0.0 10 0.8 3 0.0 6 0.8 4 1.6 4 0.0 1 0.0 20 1.6 6 0.0 2 0.0 37 0.0 38 0.0 47 0.4 3282 ± 55 TL Ministral-3-14B 4.0 4 9.6 28 0.0 0 11.2 50 6.2 1957 ± 264 TL 3.2 14 0.8 5 2.4 6 0.0 5 0.0 4 0.0 6 0.0 6 7.2 12 6.4 9 2.4 15 6.4 30 0.0 0 0.0 10 0.0 2 0.0 0 0.8 14 3.2 7 2.4 7 2.0 2050 ± 214 EN gpt-oss-20b 20.0 27 19.2 30 21.6 29 17.6 26 19.6 3831 ± 929 EN 17.6 29 9.6 11 11.2 12 20.8 30 7.2 11 22.4 34 8.0 13 16.8 23 16.0 24 18.4 27 17.6 27 9.6 14 17.6 25 21.6 32 12.0 14 4.0 58 15.2 28 12.8 33 14.4 3246 ± 685 EN Nemotron3-Nano-30B 2.4 20 4.8 16 3.2 15 8.8 29 4.8 2773 ± 1666 EN 4.0 25 4.0 29 7.2 22 8.8 30 7.2 61 1.6 62 1.6 34 5.6 23 5.6 46 3.2 31 7.2 27 2.4 17 7.2 46 6.4 30 3.2 18 0.8 30 1.6 31 2.4 15 4.4 1925 ± 1412 EN Gemma-4-31B 2.4 2 4.0 4 4.8 5 4.0 4 3.8 1945 ± 92 TL 4.8 5 4.8 5 5.6 6 3.2 3 2.4 2 1.6 2 2.4 2 3.2 3 4.0 4 3.2 3 5.6 6 5.6 6 2.4 3 4.0 4 3.2 3 0.0 2 1.6 3 1.6 2 3.3 1963 ± 97 TL Qwen3.5-35B-A3B 0.0 0 0.0 1 0.0 0 0.0 0 0.0 1959 ± 13 – 0.0 6 0.0 1 0.0 2 0.0 0 0.8 6 0.0 20 0.8 15 0.0 2 0.0 3 0.0 2 0.0 2 0.0 1 0.8 14 0.0 2 0.0 2 0.0 48 0.8 43 0.0 69 0.2 1966 ± 21 OTH Open-weight — large / API R1-Distill-Llama-70B 1.6 2 2.4 2 4.0 6 1.6 2 2.4 2078 ± 82 TL 2.4 2 2.4 3 2.4 3 4.0 7 1.6 2 3.2 6 0.0 9 2.4 15 2.4 3 1.6 3 1.6 3 7.2 10 3.2 11 4.0 4 1.6 2 5.6 14 0.8 3 4.8 7 2.8 2071 ± 177 EN gpt-oss-120b 31.2 46 32.0 47 28.0 43 26.4 43 29.4 3554 ± 962 TL 21.6 42 9.6 18 12.0 17 22.4 42 12.0 19 20.8 41 11.2 22 21.6 39 24.8 46 24.0 42 26.4 44 9.6 14 24.8 47 28.8 46 12.0 17 9.6 66 21.6 50 25.6 54 18.8 3084 ± 701 EN Qwen3.5-122B-A10B 0.0 0 2.4 5 1.6 4 0.8 2 1.2 3917 ± 46 TL 0.0 5 0.0 1 0.0 0 2.4 2 0.0 0 4.8 23 0.0 4 0.8 2 1.6 6 1.6 5 0.0 2 0.0 0 2.4 14 3.2 6 0.0 1 1.6 42 1.6 39 0.8 53 1.2 3279 ± 67 TL Qwen3-235B-A22B 0.0 0 1.6 2 4.0 5 0.8 1 1.6 3984 ± 115 TL 0.8 2 0.0 0 0.0 0 3.2 3 0.0 0 0.8 2 0.8 1 2.4 2 3.2 4 2.4 3 5.6 6 0.8 1 4.8 6 3.2 3 0.0 0 3.2 8 3.2 5 3.2 3 2.1 3315 ± 104 TL DeepSeek-V3.2 8.0 10 7.2 9 6.4 10 7.2 10 7.2 2129 ± 140 TL 8.0 14 6.4 10 8.0 15 4.0 7 4.0 8 6.4 10 5.6 20 7.2 18 4.8 6 8.0 13 7.2 17 7.2 16 2.4 6 6.4 10 7.2 13 6.4 10 5.6 15 8.0 10 6.3 2122 ± 140 TL Kimi-K2.5 1.6 2 3.2 4 1.6 5 2.4 2 2.2 2099 ± 92 TL 0.8 20 3.2 14 2.4 14 2.4 14 0.0 7 0.8 18 0.8 18 3.2 8 1.6 13 0.8 9 2.4 10 1.6 15 0.8 28 1.6 10 1.6 7 0.0 19 0.8 26 3.2 28 1.6 2114 ± 100 TL Closed-source Claude-Haiku-4.5 38.4 100 36.0 100 31.2 100 29.6 99 33.8 1103 ± 153 TL 27.2 99 30.4 100 31.2 100 36.0 100 25.6 99 28.0 89 31.2 99 32.0 100 26.4 100 33.6 100 30.4 100 33.6 100 24.0 100 34.4 100 36.0 100 12.8 94 30.4 100 30.4 100 29.6 1185 ± 192 TL Gemini-2.5-Flash 0.0 0 0.8 2 0.8 1 0.0 2 0.4 4693 ± 4631 TL 0.8 1 1.6 2 2.4 3 0.0 1 0.8 2 0.8 4 1.6 4 0.0 0 0.8 3 0.0 1 0.8 2 0.0 1 1.6 4 0.8 1 0.8 2 0.8 2 0.0 2 0.0 2 0.8 5096 ± 4904 TL GPT-5.4 20.8 100 18.4 100 16.8 100 18.4 100 18.6 579 ± 162 TL 20.0 100 19.2 100 18.4 100 17.6 100 20.0 100 16.8 100 14.4 100 16.0 100 18.4 100 17.6 100 20.0 100 16.0 100 17.6 100 19.2 100 18.4 100 15.2 100 16.8 100 18.4 100 17.8 676 ± 191 TL Table 13: High-difficulty base-prompting results on our P LURA M ATH languages vs high-resource ones from PolyMath. Each cell shows answer accuracy (%; large) over the answer-format compliance rate (%; small). Cell shading encodes accuracy on a tier-specific scale (pale → deep teal, 0 → 40% ); a numeral is printed in black on light (low-score) cells and in white on dark (high-score) cells solely for legibility—the text colour carries no extra meaning. Best and second-best per column are highlighted. After each language block we report the macro-average accuracy, the mean ± std generation length (tokens, reasoning+answer), and the dominant answer language coded EN (predominantly English) or TL (requested target language). High-resource Summary (HR) Our target languages Summary (target) Model en de ru es Avg. Len Lang hi tr pl uk uz or am el k cs he sr t sk ca cv hsb dsb Avg. Len Lang Open-weight — small ( ≤ 4B) Qwen3.5-0.8B 0.0 17 0.8 31 0.0 14 0.0 37 0.2 2994 ± 992 TL 0.0 18 0.0 4 0.0 4 0.0 17 0.0 8 1.6 100 0.0 5 0.0 24 0.0 12 0.0 30 1.6 21 0.0 7 0.0 9 0.0 0 0.0 7 0.0 13 0.0 15 0.8 47 0.2 2523 ± 500 TL LFM2.5-1.2B 0.0 0 0.0 0 0.0 0 0.0 0 0.0 3468 ± 1167 TL 0.0 0 0.0 6 0.0 13 0.0 0 0.0 84 0.0 0 0.0 18 0.0 0 0.0 0 0.0 0 0.0 0 0.8 24 0.0 0 0.0 0 0.0 6 0.0 0 0.0 0 0.0 0 0.0 2460 ± 595 EN Ouro-1.4B 0.0 1 0.0 1 0.0 1 0.0 2 0.0 1744 ± 161 TL 0.8 3 0.0 3 0.0 2 0.0 2 0.0 9 0.0 0 0.0 30 0.0 3 0.0 7 0.0 2 0.0 4 0.0 2 1.6 14 0.0 0 0.0 0 0.0 11 0.0 10 0.0 8 0.1 1716 ± 160 EN R1-Distill-Qwen-1.5B 0.0 0 0.0 8 0.0 2 0.0 8 0.0 3228 ± 493 EN 0.0 24 0.0 9 0.0 10 0.0 14 0.0 11 0.0 100 0.8 7 0.0 25 0.0 20 0.8 15 0.0 25 0.0 2 0.0 20 0.0 0 0.0 6 0.0 18 0.0 14 0.0 18 0.1 3345 ± 408 EN Qwen3.5-2B 0.0 2 0.0 1 0.0 1 0.0 1 0.0 1961 ± 20 OTH 0.0 7 0.0 2 0.0 2 0.0 3 0.0 5 0.0 0 0.0 2 0.0 2 0.0 1 0.0 4 0.0 4 0.0 14 0.0 2 0.0 0 0.0 2 0.0 27 0.0 14 0.0 26 0.0 1983 ± 51 OTH Ouro-2.6B 0.0 1 0.0 0 0.0 1 0.0 0 0.0 1885 ± 84 EN 0.0 0 0.0 0 0.0 1 0.0 0 0.0 6 0.0 0 0.0 21 0.0 1 0.0 9 0.0 1 0.8 4 0.0 0 0.0 12 0.0 0 0.0 1 0.0 11 0.0 9 0.0 9 0.0 1730 ± 149 EN Ministral-3-3B 0.0 13 0.0 14 2.4 22 7.2 68 2.4 1838 ± 343 TL 0.0 8 1.6 18 0.8 7 2.4 26 0.0 6 0.0 0 0.0 13 1.6 8 2.4 35 0.8 10 0.8 5 0.0 1 0.0 18 0.0 0 2.4 35 0.8 37 0.8 18 0.0 16 0.8 1948 ± 329 EN Gemma-3-4B 5.6 91 3.2 81 4.8 71 4.8 77 4.6 1482 ± 416 TL 3.2 66 0.8 65 2.4 79 4.0 74 3.2 70 0.0 100 1.6 58 4.8 80 0.8 31 4.0 75 0.8 71 3.2 71 0.8 38 0.0 100 4.0 86 1.6 49 0.0 76 0.8 69 2.0 1680 ± 571 OTH Qwen3.5-4B 0.0 0 0.0 1 0.0 1 0.0 1 0.0 1960 ± 13 OTH 0.0 18 0.0 2 0.0 3 0.0 2 0.0 10 0.0 0 0.0 12 0.0 2 0.0 6 0.0 3 0.0 2 0.0 6 0.0 14 0.0 0 0.0 2 0.0 22 0.0 18 0.0 19 0.0 1977 ± 20 OTH Open-weight — mid (7–35B) OLMo-3-7B-Think 0.0 0 0.0 2 0.0 0 0.0 0 0.0 1990 ± 107 EN 0.0 2 0.0 2 0.0 2 0.0 2 0.0 31 0.0 0 0.0 9 0.0 1 0.0 6 0.0 1 0.0 1 0.0 1 0.0 43 0.0 0 0.0 0 0.0 48 0.0 18 0.0 26 0.0 1974 ± 79 EN R1-0528-Qwen3-8B 0.0 2 0.0 3 0.0 4 0.0 6 0.0 3349 ± 269 EN 0.0 48 0.0 6 0.0 4 0.0 41 0.0 18 0.0 0 0.0 27 0.0 49 0.0 46 0.0 27 0.0 40 0.0 43 0.0 37 0.0 0 0.0 8 0.0 37 0.0 27 0.0 34 0.0 3239 ± 274 EN Ministral-3-8B 1.6 6 0.0 12 1.6 10 0.8 14 1.0 1965 ± 262 EN 0.0 16 0.8 14 4.0 32 0.8 9 0.0 37 0.0 0 1.6 23 0.8 20 0.8 22 1.6 6 1.6 13 0.8 6 0.8 13 0.0 0 0.8 12 0.0 19 0.8 14 1.6 15 0.9 1949 ± 336 EN Qwen3.5-9B 0.0 0 0.0 2 0.0 2 0.0 1 0.0 3919 ± 25 OTH 1.6 15 0.0 3 0.0 3 0.8 5 0.0 5 0.0 0 0.0 11 0.8 4 0.8 9 0.0 3 0.0 2 0.0 2 0.0 21 0.0 0 0.0 3 0.0 38 0.8 24 0.8 27 0.3 3276 ± 27 OTH Ministral-3-14B 0.0 5 0.8 35 0.8 2 6.4 65 2.0 1853 ± 356 TL 1.6 6 0.0 3 0.0 4 0.0 2 0.0 6 0.0 0 0.0 7 0.0 11 0.0 3 2.4 20 1.6 28 0.0 6 0.0 16 0.0 100 0.0 1 0.0 15 0.8 7 0.0 7 0.4 2055 ± 212 EN gpt-oss-20b 3.2 13 3.2 13 5.6 17 6.4 18 4.6 3994 ± 573 EN 3.2 15 1.6 3 0.8 2 5.6 17 0.0 2 0.0 100 2.4 8 4.8 13 4.8 12 6.4 15 4.8 13 2.4 2 4.0 18 0.0 0 2.4 4 0.8 62 4.0 22 4.0 30 2.9 3421 ± 392 EN Nemotron3-Nano-30B 2.4 14 0.8 12 0.8 11 0.0 17 1.0 3039 ± 1589 EN 0.8 10 3.2 33 0.8 23 0.8 36 0.0 49 0.0 100 0.8 45 2.4 18 1.6 44 1.6 25 0.8 25 0.0 32 1.6 51 0.0 0 0.0 13 1.6 58 0.8 30 0.0 30 0.9 2034 ± 1199 EN Gemma-4-31B 0.0 0 0.0 0 0.0 0 0.0 0 0.0 1953 ± 66 TL 0.0 0 0.8 2 0.8 1 0.8 1 0.0 0 0.0 0 0.0 1 0.0 0 0.0 0 0.0 0 0.8 1 0.0 0 0.0 0 0.0 0 0.0 0 0.0 0 0.0 0 0.0 4 0.2 1976 ± 66 TL Qwen3.5-35B-A3B 0.0 0 0.0 3 0.0 2 0.0 2 0.0 1961 ± 13 OTH 0.0 12 0.0 2 0.0 2 0.0 3 0.0 10 0.0 0 0.0 21 0.0 3 0.0 5 0.0 3 0.0 4 0.0 4 0.0 23 0.0 0 0.0 4 0.0 45 0.0 45 0.0 70 0.0 1966 ± 15 OTH Open-weight — large / API R1-Distill-Llama-70B 0.0 0 0.8 1 0.8 5 0.0 0 0.4 2044 ± 53 EN 0.0 1 0.0 1 0.8 2 0.8 6 0.0 0 0.0 0 0.0 9 0.8 10 0.0 1 0.0 0 0.0 0 0.0 7 0.8 8 0.0 0 0.0 0 0.8 11 0.0 2 0.0 4 0.2 2028 ± 139 EN gpt-oss-120b 9.6 36 8.0 37 8.8 38 8.8 39 8.8 3694 ± 821 TL 8.8 42 4.0 14 3.2 10 10.4 38 3.2 14 0.0 100 4.0 14 8.8 42 8.0 47 8.0 51 8.8 38 4.0 11 8.0 40 0.0 100 3.2 14 4.0 78 7.2 53 5.6 53 5.5 3147 ± 523 EN Qwen3.5-122B-A10B 0.0 0 0.0 0 0.0 2 0.0 2 0.0 3920 ± 22 OTH 0.0 13 0.0 3 0.0 1 0.0 3 0.0 4 0.0 0 0.0 14 0.8 6 0.0 4 0.0 3 0.0 4 0.0 2 0.8 12 0.0 0 0.0 3 0.0 48 0.0 43 0.0 40 0.1 3272 ± 22 OTH Qwen3-235B-A22B 0.0 0 0.0 0 0.0 0 0.0 0 0.0 3999 ± 0 OTH 0.0 0 0.0 0 0.0 0 0.0 0 0.0 2 0.0 0 0.0 2 0.0 1 0.0 0 0.0 1 0.0 0 0.0 0 0.8 1 0.0 0 0.0 0 0.0 6 0.0 2 0.0 2 0.0 3331 ± 19 OTH Kimi-K2.5 0.0 0 0.0 2 0.0 7 0.0 2 0.0 2056 ± 46 OTH 0.0 18 0.8 6 0.0 14 0.0 10 0.0 4 0.0 0 0.0 14 0.8 13 0.0 10 0.0 14 0.0 7 0.0 13 0.0 19 0.0 0 0.0 7 0.0 29 0.0 26 0.0 33 0.1 2074 ± 56 OTH DeepSeek-V3.2 0.0 3 0.8 3 0.0 1 0.8 7 0.4 2119 ± 62 OTH 0.0 15 0.0 6 0.8 16 0.8 2 0.8 9 0.0 0 0.0 14 0.8 22 0.8 2 0.8 10 0.0 12 0.8 13 1.6 3 0.0 0 0.0 14 1.6 5 0.0 29 0.0 10 0.5 2113 ± 78 OTH Closed-source Claude-Haiku-4.5 20.8 100 15.2 100 14.4 100 19.2 100 17.4 1047 ± 146 TL 18.4 100 17.6 100 16.8 100 21.6 100 16.0 100 0.0 100 16.0 98 20.0 100 19.2 100 18.4 100 17.6 100 16.0 100 16.8 99 0.0 100 19.2 99 8.8 97 16.8 100 16.8 100 15.3 1173 ± 164 OTH Gemini-2.5-Flash 0.8 2 1.6 3 0.0 3 1.6 2 1.0 5002 ± 4722 TL 0.0 2 0.0 2 0.0 2 0.0 2 1.6 2 0.0 0 0.8 3 0.0 1 0.0 1 0.8 3 0.0 1 0.0 3 0.0 3 0.0 0 0.0 2 0.0 1 0.0 2 0.0 2 0.2 5195 ± 4335 TL GPT-5.4 4.8 100 4.0 100 3.2 100 3.2 100 3.8 574 ± 189 TL 4.0 100 2.4 100 4.0 100 4.8 100 3.2 100 0.0 100 3.2 100 2.4 100 4.0 100 4.8 100 4.8 100 3.2 100 2.4 100 0.0 100 4.0 100 2.4 100 3.2 100 2.4 100 3.1 617 ± 194 TL Table 14: Top-difficulty base-prompting results on our P LURA M ATH languages vs high-resource ones from PolyMath. Each cell shows answer accuracy (%; large) over the answer-format compliance rate (%; small). Cell shading encodes accuracy on a tier-specific scale (pale → deep teal, 0 → 25% ); a numeral is printed in black on light (low-score) cells and in white on dark (high-score) cells solely for legibility—the text colour carries no extra meaning. Best and second-best per column are highlighted. After each language block we report the macro-average accuracy, the mean ± std generation length (tokens, reasoning+answer), and the dominant answer language coded EN (predominantly English), or TL (requested target language), or OTH is a model generated non-fluent answers full of special characters. 0102030 DW-ACC (%) en de es ru hi tr pl uk uz or am el k cs he sr t sk ca chv hsb dsb High-res. Target 0255075100 acc. (%) en de es ru hi tr pl uk uz or am el k cs he sr t sk ca chv hsb dsb low 02040 acc. (%) medium 010203040 acc. (%) en de es ru hi tr pl uk uz or am el k cs he sr t sk ca chv hsb dsb high 01020 acc. (%) top (a) Aggregate (DW-ACC)(b) Per difficulty level Figure 11: AggregatedDW=ACC baseprompting scores distributions per all languages across all models and all levels.We observe a clear trend in mean scores that correlates with language resource rankings, with severely underrepresented languages exhibiting low performance even on low-difficulty tasks. 05101520253035 DW-ACC (%) Qwen3.5-0.8B LFM2.5-1.2B Ouro-1.4B DS-R1-Qwen-1.5B Qwen3.5-2B Ouro-2.6B Ministral-3-3B Gemma-3-4B Qwen3.5-4B Olmo-3-7B DS-R1-Qwen3-8B Ministral-3-8B Qwen3.5-9B Ministral-3-14B gpt-oss-20b Nemotron3-Nano-30B Gemma-4-31B Qwen3.5-35B-A3B DS-R1-Llama-70B gpt-oss-120b Qwen3.5-122B-A10B Qwen3-235B DeepSeek-V3.2 Kimi-K2.5 claude-haiku-4.5 gemini-2.5-flash gpt-5.4 Aggregate (DW-ACC) 020406080100 acc. (%) Qwen3.5-0.8B LFM2.5-1.2B Ouro-1.4B DS-R1-Qwen-1.5B Qwen3.5-2B Ouro-2.6B Ministral-3-3B Gemma-3-4B Qwen3.5-4B Olmo-3-7B DS-R1-Qwen3-8B Ministral-3-8B Qwen3.5-9B Ministral-3-14B gpt-oss-20b Nemotron3-Nano-30B Gemma-4-31B Qwen3.5-35B-A3B DS-R1-Llama-70B gpt-oss-120b Qwen3.5-122B-A10B Qwen3-235B DeepSeek-V3.2 Kimi-K2.5 claude-haiku-4.5 gemini-2.5-flash gpt-5.4 low 01020304050 acc. (%) medium 0510152025303540 acc. (%) Qwen3.5-0.8B LFM2.5-1.2B Ouro-1.4B DS-R1-Qwen-1.5B Qwen3.5-2B Ouro-2.6B Ministral-3-3B Gemma-3-4B Qwen3.5-4B Olmo-3-7B DS-R1-Qwen3-8B Ministral-3-8B Qwen3.5-9B Ministral-3-14B gpt-oss-20b Nemotron3-Nano-30B Gemma-4-31B Qwen3.5-35B-A3B DS-R1-Llama-70B gpt-oss-120b Qwen3.5-122B-A10B Qwen3-235B DeepSeek-V3.2 Kimi-K2.5 claude-haiku-4.5 gemini-2.5-flash gpt-5.4 high 0510152025 acc. (%) top high-resource (en, de, es, ru)target languages Figure 12: Answers correctness scores distributions for base prompting per languages types across all models per each level. Smaller models exhibit substantial performance degradation on underrepresented languages, whereas larger and proprietary models remain comparatively stable across language groups. H.2 Models’ Reasoning Length Comparison Visualization We report here the distributions of reasoning traces and answers length per language (Figure 13) and per model (Figure 14) for aggregated metrics and per each level. We usedQwen3-4Bbase model for tokenization. Overall, reasoning lengths do not differ substantially across languages types. For underrepresented languages, models often generate slightly longer outputs for thelowerlevel, but then for others—even shorter outputs, likely due to failing to reach correct solutions. Notably, the top-performing models exhibit nearly equivalent reasoning lengths and variation across both high-resource and underrepresented languages, while producing substantially more concise outputs overall. 0100020003000400050006000 length (tokens) en de es ru hi tr pl uk uz or am el k cs he sr t sk ca chv hsb dsb Aggregate (all levels) 010002000300040005000 length (tokens) en de es ru hi tr pl uk uz or am el k cs he sr t sk ca chv hsb dsb low 0100020003000400050006000 length (tokens) medium 0100020003000400050006000 length (tokens) en de es ru hi tr pl uk uz or am el k cs he sr t sk ca chv hsb dsb high 01000200030004000500060007000 length (tokens) top Figure 13: Distribution of reasoning and answers length for base prompting per language. 0100020003000400050006000 length (tokens) Qwen3.5-0.8B LFM2.5-1.2B Ouro-1.4B DS-R1-Qwen-1.5B Qwen3.5-2B Ouro-2.6B Ministral-3-3B Gemma-3-4B Qwen3.5-4B Olmo-3-7B DS-R1-Qwen3-8B Ministral-3-8B Qwen3.5-9B Ministral-3-14B gpt-oss-20b Nemotron3-Nano-30B Gemma-4-31B Qwen3.5-35B-A3B DS-R1-Llama-70B gpt-oss-120b Qwen3.5-122B-A10B Qwen3-235B DeepSeek-V3.2 Kimi-K2.5 claude-haiku-4.5 gemini-2.5-flash gpt-5.4 Aggregate (all levels) 010002000300040005000 length (tokens) Qwen3.5-0.8B LFM2.5-1.2B Ouro-1.4B DS-R1-Qwen-1.5B Qwen3.5-2B Ouro-2.6B Ministral-3-3B Gemma-3-4B Qwen3.5-4B Olmo-3-7B DS-R1-Qwen3-8B Ministral-3-8B Qwen3.5-9B Ministral-3-14B gpt-oss-20b Nemotron3-Nano-30B Gemma-4-31B Qwen3.5-35B-A3B DS-R1-Llama-70B gpt-oss-120b Qwen3.5-122B-A10B Qwen3-235B DeepSeek-V3.2 Kimi-K2.5 claude-haiku-4.5 gemini-2.5-flash gpt-5.4 low 0100020003000400050006000 length (tokens) medium 0100020003000400050006000 length (tokens) Qwen3.5-0.8B LFM2.5-1.2B Ouro-1.4B DS-R1-Qwen-1.5B Qwen3.5-2B Ouro-2.6B Ministral-3-3B Gemma-3-4B Qwen3.5-4B Olmo-3-7B DS-R1-Qwen3-8B Ministral-3-8B Qwen3.5-9B Ministral-3-14B gpt-oss-20b Nemotron3-Nano-30B Gemma-4-31B Qwen3.5-35B-A3B DS-R1-Llama-70B gpt-oss-120b Qwen3.5-122B-A10B Qwen3-235B DeepSeek-V3.2 Kimi-K2.5 claude-haiku-4.5 gemini-2.5-flash gpt-5.4 high 01000200030004000500060007000 length (tokens) top high-resource (en, de, es, ru)target languages Figure 14: Distribution of models’ reasoning and answers length for base prompting across languages types. H.3 Per-Prompting-Ablations Results This section reports the full per-language and per-prompt breakdown of the aggregated results in Table 15. For every model we report four prompt designs: • Base — the problem is presented and solved entirely in the target language. •EnCoT — the problem is given in the target language but the system prompt instructs the model to perform its chain-of-thought in English as the best performing high-resource language. •Backtr. — the problem is machine-translated back to the corresponding high-resource language and solved there. We compare three prompting strategies exclusively on open-weight models. Backtranslation into a high-resource language yields only marginal improvements for the LFM model and does not improve performance for the remaining systems. For several mid-sized models, includingGemma-3-4bandNemotron3-Nano-30B, as well as the larger gpt-oss-120b, EN-CoT prompting improves performance for some languages. However, these gains may be partly from the additional instruction to reason step-by-step rather than from the language switch itself. Overall, no substantial stable improvements in average performance are observed with other than base prompting strategies. High-resourceOur target languages ModelPromptenderuesAvg HR hitrplukuzoramelkkcshesrttskcacvhsbdsbAvg TL Open-weight — small (≤4B) Qwen3.5-0.8B Base4.73.10.73.73.10.40.40.51.40.11.20.11.70.41.82.40.50.71.30.3 0.3 0.40.80.8 EnCoT3.72.43.03.63.20.90.20.32.40.20.40.01.90.51.70.90.20.62.00.2 0.4 0.30.50.7 BT –0.10.10.10.20.00.00.00.30.30.30.20.10.30.00.5 0.0 0.00.00.1 LFM2.5-1.2B Base0.00.00.00.00.00.01.03.60.00.50.00.90.00.00.00.01.00.00.02.1 0.0 0.00.00.5 EnCoT 0.00.00.00.00.00.01.44.10.00.50.01.40.00.00.00.01.20.00.02.8 0.0 0.00.00.6 BT–2.21.22.43.40.62.11.43.82.52.33.82.22.91.43.5 0.3 0.50.42.1 R1-Distill-Qwen-1.5B Base4.00.20.20.51.20.10.30.10.00.50.11.00.20.21.10.00.90.50.10.5 0.8 0.20.10.4 EnCoT4.70.50.00.71.50.11.00.50.11.70.20.60.00.10.20.00.80.40.50.9 0.5 0.20.30.4 BT –1.71.72.00.10.91.51.13.30.20.22.72.60.11.50.3 0.3 0.30.01.1 Ouro-1.4B Base7.36.95.77.26.82.93.35.21.40.50.40.62.10.42.61.71.41.42.34.8 0.3 0.71.01.8 EnCoT 5.94.95.05.35.32.61.83.82.00.80.50.01.90.52.91.81.80.52.23.6 0.5 0.60.61.6 BT–3.42.13.64.00.72.71.65.13.02.75.84.03.02.75.1 0.3 0.90.32.8 Qwen3.5-2B Base2.02.12.02.12.10.51.82.62.31.40.70.32.31.12.12.02.10.82.22.3 0.0 0.40.11.4 EnCoT1.91.11.01.41.30.60.71.01.40.70.40.20.70.91.01.31.00.90.81.1 0.0 0.40.00.7 BT –0.50.30.41.60.10.60.51.11.10.61.10.61.10.51.1 0.0 0.10.10.6 Ouro-2.6B Base8.26.77.57.87.54.45.70.95.81.81.80.35.50.65.33.04.21.24.75.7 1.1 2.72.03.1 EnCoT7.97.57.37.07.45.25.06.25.92.21.90.15.21.65.33.35.20.85.16.1 1.2 2.12.13.6 BT –3.22.83.46.11.13.02.26.23.32.95.84.34.53.06.2 0.5 0.40.53.3 Ministral-3-3B Base 9.17.99.9 14.510.37.84.08.48.71.21.30.08.26.17.37.22.23.86.79.7 2.2 3.92.85.1 EnCoT9.27.19.89.89.06.46.47.3 10.4 4.42.12.38.77.57.37.66.75.87.48.1 2.7 6.24.66.2 BT–1.50.61.23.30.51.40.93.02.01.42.71.72.11.01.4 0.2 0.10.01.4 Gemma-3-4B Base 13.3 10.8 6.7 11.610.68.46.49.0 10.5 7.62.55.3 10.7 3.69.57.49.94.18.2 10.3 2.1 3.23.36.8 EnCoT12.7 11.9 12.2 12.312.311.0 9.4 13.4 11.8 7.96.19.0 12.6 8.6 10.4 11.0 10.3 7.48.4 11.7 5.5 7.55.89.3 BT–5.76.69.06.32.84.23.5 10.3 2.75.98.27.13.75.39.0 1.3 2.91.85.4 Qwen3.5-4B Base 2.93.53.53.23.32.63.13.83.13.01.82.83.63.23.14.03.63.03.73.7 1.4 1.21.42.9 EnCoT2.71.82.01.72.01.71.82.11.81.81.61.42.31.91.42.31.81.21.82.1 0.7 1.00.41.6 BT –1.10.60.52.90.11.10.52.21.71.11.80.71.70.52.1 0.2 0.10.11.1 Open-weight — mid (7–35B) OLMo-3-7B-Think Base5.14.72.75.64.51.83.34.63.80.11.40.12.81.32.52.40.20.32.34.0 0.1 0.50.21.8 EnCoT5.95.04.55.45.22.83.94.94.30.11.50.12.91.52.93.30.20.42.74.6 0.1 0.50.52.1 BT–2.21.32.02.40.01.80.13.72.01.23.30.11.61.13.6 0.0 0.20.01.5 Ministral-3-8B Base10.7 8.79.6 10.910.09.39.7 11.8 9.48.11.52.79.35.98.7 10.5 9.77.07.99.7 1.4 5.75.97.5 EnCoT8.88.4 10.1 9.09.19.49.39.18.98.51.92.18.86.97.59.49.87.47.79.3 1.7 4.35.77.1 BT–4.54.06.98.33.14.42.77.96.85.58.46.67.64.99.1 1.1 2.82.15.4 Qwen3.5-9B Base3.73.65.04.54.24.93.53.75.44.13.43.45.05.14.15.23.74.64.73.6 3.3 2.52.54.0 EnCoT 3.51.92.72.02.52.52.62.02.32.31.71.72.32.62.32.62.12.12.72.2 1.9 1.31.22.1 BT–1.40.51.02.80.31.10.81.91.00.71.90.71.40.72.1 0.3 0.20.21.1 gpt-oss-20b Base17.6 15.8 19.2 18.017.715.8 12.4 12.1 18.6 10.2 15.2 9.3 16.7 16.0 17.3 16.6 11.8 16.3 14.9 12.7 4.4 13.6 12.613.7 EnCoT18.8 16.7 17.0 16.017.115.9 11.7 12.2 15.9 11.5 14.2 7.7 15.1 15.1 14.9 16.3 11.5 16.1 15.4 10.2 5.4 11.8 12.413.0 BT –6.56.18.7 11.5 4.74.64.5 10.2 9.46.39.47.49.25.38.4 3.0 4.02.96.8 Nemotron3-Nano-30B Base 9.57.77.87.78.27.47.98.66.53.90.92.08.94.37.69.76.64.35.36.6 1.9 3.52.55.5 EnCoT11.3 11.1 11.6 13.311.811.1 8.3 11.1 11.3 6.71.63.8 11.0 6.1 10.0 10.3 10.0 3.19.89.5 3.9 7.15.37.8 BT–4.15.06.66.44.23.43.77.53.73.66.26.75.14.75.9 1.1 2.11.84.5 Gemma-4-31B Base6.67.48.78.67.88.49.18.78.57.56.47.67.17.66.98.58.66.87.47.6 3.3 5.35.07.2 EnCoT8.17.87.98.58.17.78.17.87.87.46.87.96.77.06.67.57.67.37.37.1 5.4 6.35.77.1 BT–2.82.63.17.32.62.11.74.84.74.24.33.05.92.77.2 1.2 1.81.63.5 Qwen3.5-35B-A3B Base 3.73.63.94.23.92.93.73.84.44.72.34.14.13.73.23.83.73.33.83.8 1.9 2.82.83.5 EnCoT2.82.33.13.02.82.12.62.62.23.02.02.92.42.22.22.12.52.62.92.3 2.2 1.92.42.4 BT–1.30.70.82.60.31.10.72.42.00.92.30.91.90.52.5 0.2 0.20.31.2 Open-weight — large / API R1-Distill-Llama-70B Base6.85.99.46.57.26.76.57.78.66.25.21.37.46.86.46.38.36.96.66.1 6.6 5.16.16.4 EnCoT7.27.68.97.77.97.77.37.5 10.0 7.14.62.79.16.96.37.37.58.17.47.4 7.1 5.75.97.0 BT–3.13.04.47.52.02.71.86.47.13.95.04.37.03.45.5 2.0 1.31.64.0 gpt-oss-120b Base24.6 22.7 22.0 21.522.719.6 12.8 12.2 21.0 13.6 14.6 12.9 19.1 21.3 18.7 21.5 13.0 20.5 16.6 12.5 8.1 17.9 17.616.3 EnCoT 21.1 23.9 20.9 21.821.920.7 14.1 13.9 20.3 13.7 14.7 12.5 20.6 23.4 20.9 21.0 13.7 19.3 16.8 11.7 9.9 19.0 18.716.9 BT–7.26.77.8 11.1 7.54.94.38.69.97.8 10.2 8.38.56.4 11.9 2.6 5.44.27.4 DeepSeek-V3.2 Base 10.6 10.4 9.3 10.310.110.0 9.8 11.0 8.88.59.28.7 10.3 8.99.4 10.4 10.3 7.8 10.1 9.7 9.0 9.29.89.5 EnCoT9.69.07.89.59.09.88.59.88.49.89.28.69.28.18.49.79.48.29.29.6 8.7 8.58.79.0 BT–5.04.96.18.24.24.12.68.46.95.27.56.76.95.48.0 2.3 2.62.55.4 Table 15: DW-Acc metric for prompting techniques comparison. Avg HR and Avg TL (shaded) are the macro-averages over the four high-resource and the 18 PLURAMATH target languages. For each model and each column, the best-performing prompting technique is shown in bold. BT is defined only for the target languages. I Human Assessments Results per All Levels Here, we present the detailed results of an additional human evaluation of model answers and rea- soning traces across languages. The evaluation includes five models of varying sizes:gemma-3-4B, Ministral-3-8B,Nemotron3-30B,gpt-oss-120B, andDeepSeek-V3.2. We evaluate two high-resource languages—English and Russian—as well as11languages from PLURAMATH: Hindi, Turkish, Polish, Ukrainian, Odia, Greek, Kazakh, Czech, Tatar, Chuvash, Lower Sorbian. The annotation was limited to this subset of languages due to annotators time constraints prior to submission; however, it will be expanded in future stages of this research. Still, we are covering the vast majority of the languages from all our experiments. For each language, we selected randomly12(10%) samples per each level per each model. Meaning, each annotator had to go through 48 distinct tasks across 5 models. The annotation interface used for this evaluation is shown in Figure 15. The evaluation criteria are detailed below: Q1 For incorrect predictions: was the answer essentially correct but mismarked due to formatting or added explanation? (Y/N/NA) Q2 Reasoned in target language? (Y/N/Partial) Q3 Reasoning sound and step-by-step? (Y/N/Partial) Q4 Final answer consistent with reasoning? (Y/N/NA) Q5 If no final answer was extracted, does the correct answer appear in the reasoning? (Y/N/NA) Q6 Did the reasoning finish? (Y/N) Figure 15: Interface used for the human assessment of models’ answers and reasoning part. Table 16 presents extended aggregated human evaluation results per all levels. Appendix I.1 illustrates more examples of both successful and problematic models behaviors for one target language. Q1Q2Q3Q4Q5Q6 ModelLevelHUHUHUHUHUHU gemma-3-4B low01100941007310082912310091 medium061008496427857047571 high02968992388854026761 top001008988426267006770 Ministral-3-8B low047551100731008501410095 medium88823100671005508855 high018151006410051051746 top0008100551002400028 Nemotron3_30B low0162206735100572036771 medium02295382010035042549 high0033533221003903853 top0042642271005500849 gpt-oss-120b low50461222100891009810017100100 medium332101010075100760197567 high1460396611007520126747 top981767556100701263340 DeepSeek-V3.2 low041005810088100951006210088 medium1061003210050100500132522 high021002310043100319586 top01100241004310040101 Table 16: Human evaluation per difficulty level. Cells are % “Yes” among applicable annotations (Partial = 0.5 for Q2–Q3; NA/blank excluded). H: high-resource (RU); U: underrepresented (10 PluraMath languages). I.1PLURAMATH Answers and Reasoning Traces Examples We present examples of model outputs, reasoning traces, and corresponding human annotation analyses. For illustration, we use Ukrainian examples covering all models included in the annotation process: gpt-oss-120b,DeepSeek-V3.2,nemotron3_nano_omni_30ba3b,Ministral-3-8B-Reasoning-2512, and gemma-3-4b. Figure 17 presents an example of a correct prediction and reasoning trace fromgpt-oss-120b; however, the model performs its reasoning in English. Figure 16 shows a case where the predicted answer was semantically correct but marked incorrect due to formatting differences preventing an exact match. Overall, gpt-oss-120b predominantly reasoned in English across all evaluated languages. EXAMPLElow-123GOLD50|EXTRACTED50|MATCH QUESTIONРозі може бігти зі швидкістю 10 миль на годину протягом 3 годин. Після цього вона біжить зі швидкістю 5 миль на годину. Скільки миль вона може пробігти за 7 годин? MODEL ANSWERРозі пробіжить10×3 = 30миль за перші три години, а потім 5×(7−3) = 5×4 = 20миль за решту чотирьох годин. Загальна відстань: 30 + 20 = 50миль. INTERNAL REASONING Splits the7h budget at the3h breakpoint: first3h at10mph→3×10 = 30mi; remaining7−3 = 4h at5mph →4×5 = 20mi. Sums the two legs:30 + 20 = 50mi. ANNOTATION Q1 correct but mismarked? Q2 target language? Q3 reasoning sound? Q4 ans.⇔ reasoning? Q5 answer in reasoning? Q6 reasoning finished? –PartialYesYesYesYes Figure 16: Example of a correct answer and reasoning from gpt-oss-120b. EXAMPLElow-51GOLD5|EXTRACTED5годин|FORMAT MISMATCH QUESTIONКорабель Тома може рухатися зі швидкістю 10 миль на годину. Він пливе з 13:00 до 16:00. Потім він рухається назад зі швидкістю 6 миль на годину. Скільки в нього піде на те, щоб повернутися назад? MODEL ANSWERПерший етап: відстань= 10×3 = 30миль. Щоб повернутися, треба пройти ту ж саму відстань при6миль/год:t= 30 6 = 5годин. Отже, на зворотний шлях потрібно 5 годин. INTERNAL REASONING Parses the Ukrainian prompt: forward leg13:00–16:00 = 3h at10mph→30mi. Return covers the same30mi at6mph→t= 30/6 = 5h. Notes the question asks thedurationof the return, not the clock time (arrival would be21:00). Concludes5hours. ANNOTATION Q1 correct but mismarked? Q2 target language? Q3 reasoning sound? Q4 ans.⇔ reasoning? Q5 answer in reasoning? Q6 reasoning finished? YesPartialYesYes–Yes Figure 17: Example of a correct reasoning and semantically correct but mismatched answer from gpt-oss-120b. DeepSeek-V3.2 typically reasoned in the target language, as illustrated by the Ukrainian example in Figure 18. However, the model frequently produced excessively long reasoning traces with repetitive self-verification loops. In the shown example, the model arrived at the correct answer but continued generating additional verification steps and failed to complete the final response properly. EXAMPLElow-2GOLD70 000|EXTRACTED–|NO ANSWER · TRUNCATED QUESTIONДжош вирішує перепродати будинок. Він купує будинок за 80000$ і вкладає 50000$ у ремонт. Це підвищило вартість будинку на 150%. Яку вигоду він отримав? MODEL ANSWERПідсумкову відповідь не виведено — генерація обірвалася, не вийшовши з циклу самоперевірок.⟨no ⟩ INTERNAL REASONING Загальні витрати:80 000 + 50 000 = 130 000.Тлумачить «на 150%» як приріст від ціни купівлі:80 000×1.5 = 120 000⇒нова вартість200 000,прибуток200 000−130 000 =70 000. Далі багаторазово переглядає альтернативні тлумаченя (база50 000;«на» проти «до» 150%; чи додавати ремонт окремо), щоразу перераховуючи ті самі величини й не фіксуючи відповідь. Обривається на200 000−130...⟨truncated⟩ ANNOTATION Q1 correct but mismarked? Q2 target language? Q3 reasoning sound? Q4 ans.⇔ reasoning? Q5 answer in reasoning? Q6 reasoning finished? NoYesYesNoYesNo Figure 18:DeepSeek-V3.2typically reasoned in the target language, but quite often did not finish correct answer generation due to repetitive self-checking loops. Nemotron3_nano_omni_30ba3bfrequently produced non-fluent reasoning traces (Figure 20), often mixing the target language with other languages (Figure 19). The model also commonly failed to complete its reasoning process, although it still generated a final answer in $ $ format, sometimes appearing effectively random. EXAMPLElow-52GOLD15|EXTRACTED1|MISMATCH QUESTIONПортфель стає для Юраї занадто важким. Йому потрібно витягнути звідти 15 фунтів. Його комікси важать 1/4 фунта кожен, а іграшки — 1/2 фунта кожна. Якщо він витягне 30 коміксів, скільки іграшок йому потрібно буде витягнути? MODEL ANSWER<|channel|>assistant<|message|>⟨leaked control tokens⟩a no-markdown warning leaks in mixed script:«Не використовуйте маркdown ... », then<|begin_of_thought|>.Closing line (Czech): “...je potřeba vytáhnout 1 irašku.” 1 INTERNAL REASONING [uk]Починає українською: переводить ваги в однакові одиниці;30коміксів = 30× 1 4 = 7.5фунта.[cs/sk]Mid-sentence the language drifts to Czech with broken tokenisation (“Komíks váží 1 4 funtu...Jura vytáhne 30 komícs”), restating the same7.5. Sets upx· 1 2 but writes the self-contradictory stepx· 1 2 = 7.5−7.5 = 0.5⇒x= 1, never using the15-lb target. Emits1. ANNOTATION Q1 correct but mismarked? Q2 target language? Q3 reasoning sound? Q4 ans.⇔ reasoning? Q5 answer in reasoning? Q6 reasoning finished? NoPartialNoYes–Yes Figure 19: Example of a mixed languages reasoning and not so fluent answer fromnemotron3_nano_omni_30ba3b. In contrast togpt-oss-120b, which consistently reasoned in English, andDeepSeek-V3.2, which predominantly reasoned in the target language,Ministral-3-8B-Reasoning-2512exhibited mixed EXAMPLEmedium-124GOLD18|EXTRACTEDNone|NO ANSWER · DEGENERATE QUESTIONДля додатного цілого числа n і ненульових цифр a , b і c нехай A n позначає n-значне число, кожна з цифр якого дорівнюєa;нехайB n позначаєn-значне число, кожна з цифр якого дорівнюєb,і нехайC n позначає2n-значне (не n-значне) число, кожна з цифр якого дорівнюєc.Яка найбільша можлива величинаa+b+c,за якої існує принаймні два значеняn,для яких C n −B n =A 2 n ? MODEL ANSWERВідповіді не вилучено.⟨no ; output is a non-fluent token loop⟩ INTERNAL REASONING verbatim model output (degenerate repetition, terminating in leaked control tokens): <| 2 3 3 3 \2 2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 \2 3 3 |end_of_text|<|start_header_id|>assistant<|end_header_id> ANNOTATION Q1 correct but mismarked? Q2 target language? Q3 reasoning sound? Q4 ans.⇔ reasoning? Q5 answer in reasoning? Q6 reasoning finished? NoNoNo–NoYes Figure 20:Nemotron3_nano_omni_30ba3bcould also produce absolutely non-fluent reasoning with even leaked control tokens. behavior, alternating between target-language and primarily English reasoning. Most incorrect predictions resulted from incomplete reasoning traces as in Figure 21. Finally,gemma-3-4bconsistently reasoned in the target language and, even when producing incorrect answers, generated coherent reasoning traces containing flawed intermediate steps (Figure 22). EXAMPLEhigh-56GOLDn|EXTRACTEDNA|NO ANSWER · TRUNCATED QUESTIONДано додатне ціле число n . На сітці розміром 3 × n множина квадратів S називається «зв'язаною», якщо для довільних двох різних квадратівAіBуS існує ціле числоl≥2іlквадратівA=C 1 , C 2 ,·, C l =BуSтакі, щоC i іC i+1 мають спільну сторону(i= 1,·, l−1).Знайдіть найбільше ціле числоKз такою властивістю: якщо кожен квадрат сітки випадковим чином пофарбовано в чорний або білий колір, завжди існує зв'язана множинаS, така, що абсолютне значеня різниці між числом чорних і білих квадратів у Sпринаймні дорівнюєK. MODEL ANSWERNo final answer.⟨generation did not terminate; no ⟩ INTERNAL REASONING Understanding the problem.We have a3×ngrid; each square is coloured black or white, independently and at random. A subsetSisconnectedif any two squaresA,B∈Sare joined by a pathA=C 1 ,...,C l =Bof side-adjacent squares inS.Objective:find the largest integerKsuch that, no matter how the grid is coloured, there always exists a connectedSwith|black(S)−white(S)| ≥K. Initial thoughts.This reminds me of the Pigeonhole Principle and Ramsey-type problems, where we must guarantee a structure underanycolouring — so the task is a lower bound on the maximum imbalance over connected subsets. Base case (n= 1).The3×1connected subsets are single squares (imbalance1) or all three squares (imbalance3 if monochromatic, else1). The maximumguaranteedimbalance is1, since in the worst case two squares are one colour and one the other. ⟨steps on induction, a graph-theoretic framing, pigeonhole on the three rows, largest monochromatic paths, and the checkerboard analysis omitted⟩ Step 18.I recall that for any3×ngrid the largest guaranteed imbalance isn: one can always find a path ofn squares with at leastn/2of one colour. But I need to verify this. Verification. Case 1 (all one colour):the entire grid has imbalance3n; any subset has imbalance equal to its size. Case 2 (checkerboard):as discussed, the imbalance is minimal. However, I can take a subset of two rows. ⟨generation cut off here — no further output⟩ ANNOTATION Q1 correct but mismarked? Q2 target language? Q3 reasoning sound? Q4 ans.⇔ reasoning? Q5 answer in reasoning? Q6 reasoning finished? NoNoYes–NoNo Figure 21: Example of the unfinished reasoning from Ministral-3-8B-Reasoning-2512. EXAMPLEtop-44GOLD2− 6 π |EXTRACTED1/2|MISMATCH QUESTIONПрипустимо, що площину вкрито нескінченою шаховою дошкою з одиничних квадратів. Якщо інший одиничний квадрат випадковим чином розміщують на площині так, що його положеня й орієнтація незалежні від розміщеня клітинок на шаховій дошці, якою є ймовірність того, що він не закриє жодного з кутів квадратів шахової дошки? MODEL ANSWERРобить висновок, що шукана ймовірність дорівнює1/2(невірно; правильна відповідь2− 6 π ). INTERNAL REASONING НехайS— розміщений квадрат1×1,аC— квадрат шахової дошки. Хоче знайти ймовірність того, щоSне закриває жодного кутаC;стверджує, що для цього центрS має бути на відстані більше0.5від кожного кутаC. Бере кутиCяк(x,y),(x+1,y),(x,y+1),(x+1,y+1)ітой самийсимвол(x,y)для центраS, тож отримує відстані0,1,1, √ 2(центр співпадає з кутом). Вимагаючи відстань>0.5до кожного кута, виводить суперечливі умови x>0.5,x<0.5,y>0.5,y<0.5,тобтоx,y∈[0.5,0.5]. ⟨the same line “xmust lie in[0.5,0.5]andyin[0.5,0.5]” repeats about twenty times⟩ Орієнтацію (поворот) квадратаSніде не враховано. Без жодного виведеня з наведеного оголошує1/2. ANNOTATION Q1 correct but mismarked? Q2 target language? Q3 reasoning sound? Q4 ans.⇔ reasoning? Q5 answer in reasoning? Q6 reasoning finished? NoYesPartialYes–Yes Figure 22:Gemma-3-4bconsistently reasoned fluently and in the target language, but there could have been wrong reasoning steps. J Detailed Translation Quality and Translations Examples Results with Reasoning Models We evaluate a set of reasoning models on machine translation using the FLORES+ (NLLB Team et al., 2024)devsplit. For the Sorbian languages, we acquire test sets from the authors of Okabe et al. (2025) and sample 300 items from the 4000 segments. For the languages present in FLORES+, we run machine translation both from and into English, whereas for the two Sorbian languages we only run translation from German. We tested a subset of seven opensource tested models:Qwen3.5-0.8B with reasoning and no-think mode,Ministral-3-3BandMinistral-3-14Bwith reasoning and no- think mode,Nemotron-3-Nano-Omni-30B,Qwen3.5-35B-A3B,gpt-oss-120B,Qwen3.5-122B-A10B, and DeepSeek-V3.2. We inference models either locally or via API calls.API models are called with reasoning_effort=“medium”to match the mathematical reasoning experiments, and limited to a maxi- mum output length of2048tokens. With locally inferenced models, we have more fine-grained control: We use the recommended decoding parameters when specified, set the maximum number of output tokens to4096, and use a maximum thinking token budget of2048. When models use up the full thinking token budget, we forcedly stop the reasoning trace by inserting “I’l stop thinking and provide my answer now.”, followed by the appropriate reasoning-end-string. ParameterDefaultQwen3.5Ministral temperature0.61.00.7 top_p0.950.950.95 top_k2020— min_p00.0— repetition_penalty1.051.01.05 presence_penalty—1.5— max_tokens409640964096 thinking_token_budget204820482048 Table 17: Sampling parameters for local models for translation capabilities assessment. We evaluate the models’ performance using chrF++, since this metric supports all our target languages— full results in Table 18. Additionally, we calculate the mean length of the MT and thinking outputs (Table 19), and the percentage of empty outputs (Table 20). Explicit examples of models’ behavior for the translation task are in Figure 23. Our results show rather poor chrF++ scores throughout, though with some variation across language pairs and models. For instance, Olmo consistently adds the most additional information to the machine translation output, leading to long answers and low chrF++ scores. How long the models “thought” was more determined by the model than by the language pair, with a weak negative correlation of reasoning length and chrF++. The smaller, locally-inferenced models very frequently used up the full thinking token budget, while GPT-OSS-120B and DeepSeek-V3.2 were more discerning. Translation into the two Sorbian languages achieved some of the best chrF++ scores, but this must be attributed to these segments being shorter on average. ModelTranslation outputReasoning trace Source (uk)У понеділок, науковці зі Школи медицини Стенфордського університету оголосили про винайденя нового діагностичного інструменту, що може сортувати клітини за їх видами: це малесенький друкований чіп, який можна виготовити за допомогою стандартних променевих принтерів десь по одному центу США за штуку. PromptTranslate the following text from Ukrainian to English: <text> Reference (gold, en)On Monday, scientists from the Stanford University School of Medicine announced the invention of a new diagnostic tool that can sort cells by type: a tiny printable chip that can be manufactured using standard inkjet printers for possibly about one U.S. cent each. Qwen3.5-0.8B 0.8 B · reasoning In the morning, scientists from the Stanford University School of Medicine announced the discovery of a new diagnostic tool designed to sort cells based on their characteristics. This tool consists of a compact printed device that can be produced with ordinary X-ray machines throughout the United States, each producing approximately one dollar per unit. Long, unstable trace (≈1.5k words). Repeatedly second- guesses науковці (wonders whether it is a typo for націоналісти / науковців); cannot fix the weekday, cycling through Monday/Tuesday/Wednesday/Saturday; reads променевих as “X-ray”; oscillates between cent and dollar. Never converges; cut off with “I’l stop thinking...”. Qwen3.5-0.8B 0.8 B · no-think On Monday, researchers from the Stanford University School of Medicine announced the discovery of a new diagnostic tool capable of sorting cells according to their types: this small printed chip can be manufactured using standard X-ray printers available in one to two units within a single United States unit. — reasoning disabled (empty trace) Ministral-3-3B 3 B · no-think “Here is the translation of the provided text from Ukrainian to English:” In Monday, scientists from the School of Medicine of Stanford University announced the invention of a new diagnostic tool that can sort cells by their types: it’s a tiny printed chip that can be made using standard laser printers for about one cent per unit in the US. — no trace emitted Ministral-3-3B 3 B · reasoning On Monday, scientists from the Medicine School of Stanford University announced the invention of a new diagnostic tool that can sort cells by their types: this tiny printed chip that can be made using standard laser printers for about one cent per unit in the USA. Methodical word-by-word gloss. Briefly considers “inkjet” for променевих — even notes inkjet is “more common for printing chips” (the gold reading) — but rejects it and settles on “laser”, reasoning that променевий “can refer to laser”. Converges in a single pass; drops the no-think preamble and fixes the weekday. Ministral-3-14B 14 B · no-think “Here is the translation of your text from Ukrainian to English:” On Monday, scientists from Stanford University School of Medicine announced the development of a new diagnostic tool that can sort cells by their types: it’s a tiny printed chip that can be manufactured using standard ray printers for about one cent in the U.S. per unit. — no trace emitted Ministral-3-14B 14 B · reasoning On Monday, scientists from the Medicine School of Stanford University announced the invention of a new diagnostic tool that can sort cells by their types: this tiny printed chip that can be made using standard laser printers for about one cent per unit in the USA. Trace and output identical to the 3B reasoning variant: weighs “inkjet” (acknowledging it fits “printing chips”) against “laser” and chooses laser. Reasoning here repairs the ray printers error made by the 14B no-think run. Nemotron-3-Nano-Omni 30B-A3B · reasoning A team of scientists from Stanford University’s School of Medicine announced on Monday the invention of a new diagnostic tool that can sort cells by their type: it is a tiny printed chip that can be manufactured using standard laboratory printers at a cost of just one cent per unit in the United States. — empty trace Qwen3.5-35B-A3B 35B-A3B · reasoning On Monday, researchers from Stanford University’s School of Medicine announced the invention of a new diagnostic tool capable of sorting cells by type: a tiny printed chip that can be manufactured using standard laser printers at a cost of approximately one US cent per unit. Coherent, clause-by-clause trace. Flags променевих as ambiguous and resolves it to “laser printers” by recalling the actual 2016 Stanford microfluidics study; renders the cost as ≈ one US cent per unit. Converges cleanly, then cut off with “I’l stop thinking...”. gpt-oss-120b 120 B · reasoning On Monday, scientists from the Stanford University School of Medicine announced the invention of a new diagnostic tool that can sort cells by their type: it is a tiny printed chip that can be produced with standard laser printers for roughly one U.S. cent each. Brief word-by-word gloss. Maps променевих принтерів → “laser printers” and the cost phrase → “roughly one US cent each”. Reaches a clean rendering directly, with no back- tracking. Qwen3.5-122B-A10B 122B-A10B · reasoning — output not provided — Careful sentence-by-sentence breakdown. Notes за їх видами reads naturally as “by cell type”, and that the official name is “Stanford University School of Medicine”. On променевих it observes the literal sense is “beam/ray” but judges that “the technology is likely inkjet printers” — i.e. leaning toward the gold reading. Trace truncated before a final rendering. DeepSeek-V3.2 MoE · reasoning On Monday, scientists from Stanford University School of Medicine announced the invention of a new diagnostic tool that can sort cells by their types: it is a tiny printed chip that can be manufactured using standard inkjet printers for about one U.S. cent each. Short planning trace centred on preserving the quantitative and technical detail. Selects “inkjet printers” for променевих (matching the reference) and “about one U.S. cent each” for the cost; reorders clauses from UK to EN. Closest to gold. Annotation legend. serious mistranslation / hallucination minor drift / awkward phrasing matches reference / acceptable variant extraneous boilerplate (not part of the translation). Figure 23: Translation with reasoning models example for the Ukrainian sample. Some models tend to hallucinate additional information or add in the output unnecessary explanation; bigger models translate better, but tend to paraphrase additionally. DirectionQwen3.5-0.8BQwen3.5-0.8B (NT)Ministral-3BMinistral-3B (NT)OLMo-3-7BMinistral-14BMinistral-14B (NT)Nemotron-30BQwen3.5-35Bgpt-oss-120bQwen3.5-122BDeepSeek-V3.2 Into English (X →en) Amharic→ en 1.38.89.35.21.05.32.12.99.711.60.09.8 Catalan→ en7.59.99.27.72.19.76.48.68.98.91.69.1 Chuvash→ en1.88.79.15.90.98.52.25.99.411.30.09.2 Czech→ en 4.09.810.06.72.29.27.88.78.89.12.39.4 Dutch→ en8.510.610.78.11.610.48.110.79.910.29.610.4 Greek→ en 8.09.79.47.31.49.96.79.48.99.27.29.4 Hebrew→ en5.710.08.67.21.110.06.09.19.39.24.59.0 Hindi→ en 1.19.28.87.11.59.08.04.69.29.50.09.5 Kazakh→ en5.89.49.76.91.49.27.18.79.29.84.110.2 Odia→ en1.69.88.44.81.19.63.45.88.99.24.30.0 Polish→ en6.810.310.17.02.49.47.69.79.99.74.610.5 Romanian→ en6.39.89.37.52.19.37.39.19.39.46.79.7 Serbian→ en2.18.98.65.81.59.54.98.610.08.64.39.5 Slovak→ en6.59.49.16.51.68.77.39.09.39.52.29.0 Tatar→ en3.79.89.36.91.29.37.08.19.29.14.49.3 Ukrainian→ en2.89.89.66.81.810.18.09.810.19.30.09.4 Uzbek→ en1.19.58.95.71.18.44.18.69.10.00.09.1 Out of English (en→ X ) en→ Amharic0.77.812.11.71.45.83.30.65.05.90.06.1 en→ Catalan1.89.08.37.81.09.28.35.49.89.50.09.1 en→ Chuvash 1.111.810.44.71.24.11.81.07.210.10.07.6 en→ Czech1.011.712.48.31.19.010.15.612.412.40.012.2 en→ Dutch1.27.98.16.91.46.27.16.56.47.20.07.0 en→ Greek0.88.410.14.81.410.38.86.59.49.30.09.1 en→ Hebrew0.88.99.05.01.06.86.62.59.310.00.010.4 en→ Hindi1.78.35.85.51.24.26.03.53.87.90.06.9 en→ Kazakh1.29.711.94.31.37.39.52.27.29.10.09.2 en→ Odia1.97.47.51.61.44.84.83.63.58.40.08.2 en→ Polish4.310.811.98.41.012.911.88.111.311.80.011.8 en→ Romanian1.510.711.36.51.211.710.86.610.810.90.011.2 en→ Serbian2.66.710.411.11.312.210.51.811.111.60.011.5 en→ Slovak1.511.612.38.61.212.79.59.610.811.70.012.4 en→ Tatar 0.99.210.44.71.35.42.61.15.99.10.07.9 en→ Ukrainian1.18.412.08.51.410.810.26.410.411.40.08.0 en→ Uzbek1.27.211.23.41.18.38.22.67.19.70.010.2 German pivot de→ L. Sorbian4.316.114.65.41.216.43.32.018.17.013.314.5 de→ U. Sorbian1.215.717.65.71.19.75.51.112.714.10.015.6 Table 18: CHRF++translation quality per model and translation direction. Rows are translation directions grouped intoX →en, en→ X, and German-pivot (Sorbian) directions; columns are models (NT = non-thinking variant). Cell shading encodes CHRF++ (pale→ deep teal); the best model per direction is in bold. Higher is better. DirectionQwen3.5-0.8BQwen3.5-0.8B (NT)Ministral-3BMinistral-3B (NT)OLMo-3-7BMinistral-14BMinistral-14B (NT)Nemotron-30BQwen3.5-35Bgpt-oss-120bQwen3.5-122BDeepSeek-V3.2 Into English (X →en) Amharic→ en6,4641547,9533458,031 10,7477475617,1291,2542,7231,172 Catalan→ en8,2991343,8291897,8938,6611951389,3088673,950931 Chuvash→ en7,3501526,9042839,616 10,4776762497,8212,5402,9271,289 Czech→ en8,0901324,1171898,6848,8951791349,0418223,9091,129 Dutch→ en8,4731323,7311807,6168,5871951339,2527614,035920 Greek→ en7,7791404,7181938,5928,9981991419,2948793,7431,305 Hebrew→ en7,6261364,6251928,8027,8641921338,7748823,6121,303 Hindi→ en7,6661564,5161927,9288,3291951358,5549033,3931,192 Kazakh→ en7,6431445,0732079,6989,2302091478,9481,0683,5961,098 Odia→ en6,9081527,3632786,2687,9704711957,4001,0042,873992 Polish→ en8,2841353,7881898,0128,6221751339,5008213,9991,158 Romanian→ en8,3171353,9701857,7768,5331821369,2828153,972948 Serbian→ en7,8661354,1442088,8447,9322341429,0848933,7421,204 Slovak→ en8,0941354,5151979,3188,7491941369,1918693,8981,081 Tatar→ en7,5331475,3502089,9798,6602241968,7161,2033,4561,166 Ukrainian→ en7,9831354,1531987,7018,3351841379,2068433,8481,065 Uzbek→ en8,1441505,485216 10,6129,0822611599,1971,1193,8141,149 Out of English (en→ X ) en→ Amharic7,2321295,2365015,4326,6012454905,9029502,2361,107 en→ Catalan7,9441464,627157 10,745 10,1131412729,2519663,895915 en→ Chuvash8,3141885,275515 10,219 11,0497481,3247,6651,8402,9071,098 en→ Czech8,2181435,691149 11,523 10,8751402338,9049393,5191,054 en→ Dutch8,2941465,442154 10,666 10,4791451819,4889523,896942 en→ Greek8,0611559,1622668,808 11,7092282838,5659263,647994 en→ Hebrew7,7331088,0052249,337 11,1441289178,1027833,141964 en→ Hindi7,362134 10,0134038,267 11,7643486148,1196783,156989 en→ Kazakh8,2041566,2313659,297 10,5702311,0587,9109933,143960 en→ Odia7,6831387,0844115,0848,8484072716,6689492,465957 en→ Polish8,2271395,965178 11,413 10,8961443059,1689043,693962 en→ Romanian7,9871415,030207 10,7939,9561423629,2598743,874913 en→ Serbian8,0021606,192154 10,906 11,3431391,2588,7371,0283,3191,081 en→ Slovak8,2601315,523163 12,054 10,5781623278,8769663,498923 en→ Tatar7,9781456,3134129,641 11,1553661,4717,8391,2103,016988 en→ Ukrainian7,8291364,9831749,5489,3121353428,7578523,446906 en→ Uzbek8,6192066,594659 10,150 10,2852201,0008,7041,1343,352994 German pivot de→ L. Sorbian9,115674,519211 12,161 12,2882578588,2881,7173,404966 de→ U. Sorbian9,170804,795246 12,349 12,9942459758,3531,8813,3941,002 Table 19: Mean generation length (tokens) per model and translation direction, summing the answer/translation output and the reasoning (“thinking) trace. Rows are translation directions grouped intoX →en, en→ X, and German-pivot (Sorbian) directions; columns are models (NT = non-thinking variant). Larger values indicate longer total generations. DirectionQwen3.5-0.8BQwen3.5-0.8B (NT)Ministral-3BMinistral-3B (NT)OLMo-3-7BMinistral-14BMinistral-14B (NT)Nemotron-30BQwen3.5-35Bgpt-oss-120bQwen3.5-122BDeepSeek-V3.2 Into English (X →en) Amharic→ en0 / 00 / 10021 / 00 / 1000 / 1.561 / 00 / 1000 / 1005.5 / 07.2 / 099 / 01.0 / 0 Catalan→ en0 / 00 / 100 0.1 / 00 / 1000 / 010 / 00 / 1000 / 1001.5 / 00 / 090 / 00 / 0 Chuvash→ en0 / 00 / 100 0.9 / 00 / 1000 / 0.420 / 00 / 1000 / 1005.3 / 053 / 0100 / 01.2 / 0 Czech→ en0 / 00 / 100 0.1 / 00 / 1000 / 010 / 00 / 1000 / 1001.1 / 00.1 / 093 / 00.9 / 0 Dutch→ en0 / 00 / 1000 / 00 / 1000 / 09.0 / 00 / 1000 / 1000.8 / 00 / 092 / 00 / 0 Greek→ en0 / 00 / 100 0.3 / 00 / 1000 / 0.111 / 00 / 1000 / 1002.2 / 00 / 099 / 631.7 / 0 Hebrew→ en0 / 00 / 100 0.5 / 00 / 1000 / 0.47.8 / 00 / 1000 / 1002.1 / 00 / 093 / 01.7 / 0 Hindi→ en0 / 00 / 100 0.3 / 00 / 1000 / 07.9 / 00 / 1000 / 1003.1 / 00 / 099 / 01.1 / 0 Kazakh→ en0 / 00 / 100 0.6 / 00 / 1000 / 0.410 / 00 / 1000 / 1002.7 / 00 / 098 / 00.6 / 0 Odia→ en0.5 / 0.20 / 10039 / 00 / 1000 / 1.969 / 00 / 1000 / 1004.7 / 00 / 099 / 0 3.2 / 3.2 Polish→ en0 / 00 / 100 0.2 / 00 / 1000 / 07.4 / 00 / 1000 / 1001.5 / 00 / 096 / 01.0 / 0 Romanian→ en0 / 00 / 1000 / 00 / 1000 / 09.1 / 00 / 1000 / 1001.6 / 00 / 090 / 00 / 0 Serbian→ en0 / 00 / 1000 / 00 / 1000 / 07.0 / 00 / 1000 / 1002.8 / 00 / 092 / 01.1 / 0 Slovak→ en0 / 00 / 100 0.1 / 00 / 1000 / 09.0 / 00 / 1000 / 1001.5 / 00 / 093 / 00.7 / 0 Tatar→ en0 / 00 / 100 0.2 / 00 / 1000 / 0.89.6 / 00 / 1000 / 1003.1 / 00 / 098 / 00.8 / 0 Ukrainian→ en0 / 00 / 100 0.3 / 00 / 1000 / 08.0 / 00 / 10060 / 1002.0 / 00.3 / 095 / 00.6 / 0 Uzbek→ en0 / 00 / 100 0.4 / 00 / 1000 / 0.19.4 / 00 / 1000 / 1002.2 / 0 76 / 7697 / 00.9 / 0 Out of English (en→ X ) en→ Amharic0.1 / 00 / 10042 / 00 / 100 0.2 / 5.097 / 00 / 100 0.1 / 1004.3 / 022 / 0100 / 01.3 / 0 en→ Catalan0 / 00 / 100 0.7 / 00 / 1000 / 0.312 / 00 / 1000 / 1004.8 / 00.1 / 099 / 00 / 0 en→ Chuvash0.1 / 00 / 100 5.0 / 00 / 1000 / 6.878 / 00 / 1000 / 1003.4 / 040 / 0100 / 00.7 / 0 en→ Czech0.1 / 00 / 100 4.4 / 00 / 100 0.1 / 5.925 / 00 / 1000 / 1004.3 / 00 / 0100 / 00.7 / 0 en→ Dutch0 / 00 / 100 2.2 / 00 / 1000 / 0.414 / 00 / 1000 / 1003.6 / 00 / 099 / 00.2 / 0 en→ Greek0 / 00 / 10029 / 00 / 1000 / 9.046 / 00 / 1000 / 1004.3 / 00.4 / 0100 / 00.3 / 0 en→ Hebrew0 / 00 / 10020 / 00 / 100 0.3 / 8.644 / 00 / 1000 / 1003.7 / 00.1 / 0100 / 00.4 / 0 en→ Hindi0 / 00 / 10032 / 00 / 100 0.1 / 5.348 / 00 / 1000 / 1003.9 / 00 / 0100 / 00.4 / 0 en→ Kazakh0 / 00 / 100 7.9 / 00 / 100 0.2 / 7.233 / 00 / 1000 / 1003.6 / 00 / 0100 / 00.2 / 0 en→ Odia0.2 / 00 / 10017 / 00 / 1000 / 6.481 / 00 / 1000 / 1004.6 / 00.6 / 0100 / 076 / 76 en→ Polish0 / 00 / 100 5.5 / 00 / 1000 / 2.821 / 00 / 1000 / 1003.6 / 00 / 0100 / 00.2 / 0 en→ Romanian0 / 00 / 100 1.2 / 00 / 1000 / 0.614 / 00 / 1000 / 1004.0 / 00 / 099 / 00 / 0 en→ Serbian0 / 00 / 100 4.5 / 00 / 1000 / 6.936 / 00 / 1000 / 1005.6 / 00.1 / 0100 / 00.8 / 0 en→ Slovak0.2 / 00 / 100 3.6 / 00 / 1000 / 7.524 / 00 / 1000 / 1004.6 / 00.3 / 0100 / 00.1 / 0 en→ Tatar0 / 00 / 10012 / 00 / 1000 / 6.448 / 00 / 1000 / 1004.4 / 00.1 / 0100 / 00.3 / 0 en→ Ukrainian0 / 00 / 100 1.4 / 00 / 1000 / 1.712 / 00 / 10059 / 1003.6 / 00.1 / 0100 / 064 / 64 en→ Uzbek0.2 / 00 / 10012 / 00 / 1000 / 6.420 / 00 / 1000 / 1003.9 / 00.2 / 0 100 / 310.2 / 0 German pivot de→ L. Sorbian0 / 00 / 100 1.0 / 00 / 1000 / 6.346 / 00 / 1000 / 1002.0 / 02.0 / 099 / 00 / 0 de→ U. Sorbian0 / 00 / 100 1.0 / 00 / 100 1.0 / 8.061 / 00 / 1000 / 1004.3 / 01.7 / 0100 / 00 / 0 Table 20: Empty-output rate (%) per model and translation direction, reported as answer / thinking: the left value is the share of empty answer/translation outputs, the right value the share of empty reasoning traces. Rows are translation directions grouped intoX →en, en→ X, and German-pivot directions; columns are models (NT = non- thinking variant, which by construction emits no reasoning, hence 100% empty thinking). Lower is better. K Discovered Problems in PolyMath Dataset While annotating PolyMath (Wang et al., 2025) and assessing the associated human reasoning traces, our annotators identified a number of errors in the original release. We group them into (i) incorrect reference answers in the English source and (i) translation and localization errors in the multilingual subsets. We report all instances below and will contribute the corresponding corrections as a pull request to the original dataset. K.1 Incorrect Reference Answers We found three items in the English split whose gold answers are incorrect. The corrected values follow from the problem statements and were verified by our annotators. These errors propagated in other languages of PolyMath benchmark. low-en-94 . Problem. In a neighborhood, the number of rabbits pets is twelve less than the combined number of pet dogs and cats. If there are two cats for every dog, and the number of dogs is 60, how many pets in total are in the neighborhood? Original answer: 348. Corrected answer: 195. high-en-122. Problem. Leta,b, andcbe positive real numbers satisfyingab + bc + ca = abc. Determine the minimum value of a a bc + b b ca + c c ab. Original answer: 729. Corrected answer: 3 28 . top-en-118. Problem. Consider the surfaceSof a cube side lengths. LetPbe one of the vertices of the cube, andD ⊂ Sthe collection of points onSthat are at most √ 2· saway fromP, where distance is measured along the surface. Divide the area of D by the area of S, leaving the answer in its exact form. Original answer: π + 3 √ 3− 3 6 . Corrected answer: 2π− 3 + 3 √ 3 12 . K.2 Translation and Localization Errors Beyond the answer errors, we observed recurring problems in the translated subsets. We label each instance with one or more of the following error types: (U) untranslated source-language content; (L) L A T E X/formatting error; (T) terminology error or inconsistency; (M) mistranslation, including factual distortion; (I) information loss relative to the source; (C) calque or unidiomatic phrasing; (S) symbol/nota- tion placement; and (X) cross-lingual misalignment between subsets. German (de). • high-de-79 (U): the otherwise German statement still contains untranslated English text. • high-de-81(L): the predicate “gut” is typeset in math mode ($gut$) instead of as emphasized text. • medium-de-12(T): inconsistent terminology within a single item—“Kisten” vs. “Boxen” for the containers and “Kugel” vs. “Ball” for the balls. • medium-de-39 (L): missing commas in the conditional clause. • low-de-17 (M): “35 Wochen pro Woche” (35 weeks per week) is nonsensical; the source specifies 35 hours per week, so the item is factually distorted. • low-de-26(C): unidiomatic “3 Paar Shorts”; natural German would not use “Paar” here (cf. “3 Shorts”). • low-de-110(M): the localized problem is incoherent in both the German and the Czech versions and does not yield a solvable task. Russian (ru).low-ru-87(I): relative to the English source, the final question omits information, leaving the requested quantity ambiguous. The English item asks for an annual salary, and the gold answer (9360) is the annual figure; the Russian rendering drops the word “annual” (“годовая”), so a solver cannot tell whether a monthly or an annual amount is wanted. English (source). A company pays each of its employees $600 in a month. The company has a policy of increasing the salaries of each of its employees by 10% of the initial salary every year for those who have stayed in the company for five years. If Sylvie just clocked 5 years in the company last December, what is herannualsalary after three more years of service? (Gold answer: 9360.) Russian (as released). Компания платит каждому из своих сотрудников 600 $ в месяц. В компани действует политика ежегодного повышения заработной платы на 10 % от начальной для каждого сотрудника, проработавшего в компани пять лет. Если Сильви отработала в компани 5 лет в прошлом декабре, сколько составит её заработная плата спустя ещё три года работы? Explanation. “. . . how much will hersalarybe after another three years of work?” — the qualifier “annual” present in the source is absent, and the monthly rate ($600) is also given, so the target unit is underspecified. Spanish (es). The Spanish subset shows several recurring issues: • (M) Mistranslations, e.g. “intersecta” for “interseca”, and “secesión” (secession) for “sucesión” (sequence). • (T) Incorrect terminology, e.g. “casco convexo” instead of “envoltura convexa” for convex hull. • (T) Context-unaware term choices, e.g. “suministros” for supplies, which is wrong in context. •(C) English calques, e.g. “se está poniendo demasiado pesada” for is getting too heavy, a poor choice here. • (S) Currency/unit symbols placed before rather than after the numeric amount. • (M) Nonsensical translations, e.g. low-es-110. •(X)high-es-46: the Spanish and Vietnamese entries are swapped in the release; the content occupying the Spanish slot is in fact Vietnamese, whose L A T E X formatting has additionally been altered. Finally, the Spanish subset predominantly uses Latin American Spanish rather than Peninsular Spanish. L LLMs Usage in This Work Finally, we clarify the role of LLMs in this work—where they were used and where they were not. Neural machine translation systems were used only to generate initial translation drafts for the target languages. All subsequent quality assessment and refinement were performed exclusively through human correction, without automated post-editing systems. Claude was actively used to assist with code development, while ChatGPT was used for language editing and polishing. All research ideas, experimental design, analyses, and findings were developed solely by the authors.