Paper deep dive
LLMs Get Smarter from Targeted Synthetic Multilingual Data
Ishika Agarwal, Arkajyoti Charaborty, Tanner Sorensen, Neha Gupta, Andreas Stolcke
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/22/2026, 3:23:06 AM
Summary
The paper introduces HOTFIXR, a synthetic data generation framework designed to improve multilingual reasoning in Large Language Models (LLMs) by addressing Language-Specific Competency (LSC). HOTFIXR uses a question generation model trained via GRPO to probe student models for linguistic deficits, generating targeted training data that improves in-distribution performance by 6.2% and reduces catastrophic forgetting on out-of-distribution tasks and languages.
Entities (10)
Relation Signals (7)
HOTFIXR â addresses â Language-Specific Competency
confidence 95% ¡ HOTFIXR: a post-training data generation framework that aims to improve an LLMâs multilingual ability... Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt.
HOTFIXR â improves â In-Distribution Performance
confidence 95% ¡ On average, HOTFIXR (1) improves in-distribution performance by 6.2%
Qwen/Qwen2.5-7B-Instruct â evaluatedon â mHotPotQA
confidence 90% ¡ we evaluate Qwen/Qwen2.5-7B-Instruct (25) on multilingual HotPotQA (11)
HOTFIXR â reduces â Catastrophic Forgetting
confidence 90% ¡ reduces catastrophic forgetting (induced by fine-tuning) on OOD tasks by 3.7%
HOTFIXR â uses â GRPO
confidence 90% ¡ The optimization algorithm is GRPO (21).
HOTFIXR â outperforms â DataEnvGym
confidence 85% ¡ Finally, over DataEnvGym, HOTFIXR performs 7.3% better in OOD tasks.
HOTFIXR â generatesdatafor â Nemotron-PTDv2
confidence 80% ¡ We generate đS based on the Nemotron data as well.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt. In other words, a language model outputs different (and potentially incorrect) responses to the same semantic query when prompted in different languages. Prior work attributes this to an internal misalignment of semantic representation across languages. Currently, there are two main approaches to address LSC in the literature: (1) routing all queries through English, improving performance, but limiting language expressivity to English; or (2) training on language-balanced data, equalizing model performance across languages, but reducing overall performance. In this work, we take a data centric perspective and introduce HOTFIXR: Hardness Optimized Training data For Improving X-Lingual Reasoning. It is a data generation framework that uses models to probe and learn a student model's multilingual weaknesses, and generates data to mitigate them. HOTFIXR can generate multilingual synthetic training data that can improve multilingual performance. We evaluate on three in-distribution tasks, three out-of-distribution tasks, and four out-of-distribution languages. On average, HOTFIXR (1) improves in-distribution performance by 6.2%, (2) reduces catastrophic forgetting (induced by fine-tuning) on OOD tasks by 3.7%, and (3) on OOD languages by 7.1%. Overall, as many real-world applications requires multilingual LLMs, our work contributes to the efforts of making LLMs multilingually proficient. We will release code upon acceptance.
Tags
Links
- Source: https://arxiv.org/abs/2608.15964v1
- Canonical: https://arxiv.org/abs/2608.15964v1
Trouble viewing inline? Open PDF directly â
Full Text
60,564 characters extracted from source content.
Expand or collapse full text
LLMs Get Smarter from Targeted Synthetic Multilingual Data Ishika Agarwal Arkajyoti Chakraborty Affiliation: UIUC, Uniphore Correspondence:ishikaa2@illinois.edu Tanner Sorensen Affiliation: UIUC, Uniphore Correspondence:ishikaa2@illinois.edu Neha Gupta Affiliation: UIUC, Uniphore Correspondence:ishikaa2@illinois.edu Andreas Stolcke Affiliation: UIUC, Uniphore Correspondence:ishikaa2@illinois.edu Abstract Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt. In other words, a language model outputs different (and potentially incorrect) responses to the same semantic query when prompted in different languages. Prior work attributes this to an internal misalignment of semantic representation across languages. Currently, there are two main approaches to address LSC in the literature: (1) routing all queries through English, improving performance, but limiting language expressivity to English; or (2) training on language-balanced data, equalizing model performance across languages, but reducing overall performance. In this work, we take a data centric perspective and introduce HOTFIXR: Hardness Optimized Training-data For Improving X-lingual Reasoning. It is a data generation framework that uses models to probe and learn a student modelâs multilingual weaknesses, and generates data to mitigate them. HOTFIXR can generate multilingual synthetic training data that can improve multilingual performance. We evaluate on three in-distribution tasks, three out-of-distribution tasks, and four out-of-distribution languages. On average, HOTFIXR (1) improves in-distribution performance by 6.2%, (2) reduces catastrophic forgetting (induced by fine-tuning) on OOD tasks by 3.7%, and (3) on OOD languages by 7.1%. Overall, as many real-world applications requires multilingual LLMs, our work contributes to the efforts of making LLMs multilingually proficient. We will release code upon acceptance. Performance Language Method ID (â ) OOD (â ) Spread (â ) Base 51.9 65.6 9.6 EngReason 50.8 65.4 8.8 SelectionGT 49.2 60.9 9.8 SelectionGEN 50.2 55.8 9.5 Filtered 48.9 58.2 10.6 Untrained 50.4 63.4 9.7 DataEnvGym 49.0 57.4 10.2 HOTFIXR (ours) 56.2 64.7 9.4 Table 1: HOTFIXR has the best in-distribution performance, the best out-of-distribution performance among the training-based baselines, and remains multilingually consistent. This table shows the cross-lingual consistency with respect to ID and OOD performance. Language spread is the standard deviation of performance across languages. Low spread indicates more consistent performance across languages. Note: the training-based baselines are SelectionGT, SelectionGEN, Filtered, Untrained, and DataEnvGym. 1 Introduction Many LLM-related business use cases require models to reliably converse in languages other than English. Multilingual language models are trained by adding in non-English data (23). However, English data significantly dominates the training data. As a result, a lot of internal LLM processes happen in English (26). Specifically, as an LLM processes Spanish (or any non-English) text, it tries to anchor its understanding in English, before answering in Spanish. This anchoring could upper-bound what an LLM can understand in non-English. This points to the representational gap in LLMs, where they represent the same concept differently, in different languages. We also see empirical evidence of this. In Figure 1, we evaluate Qwen/Qwen2.5-7B-Instruct (25) on multilingual HotPotQA (11) (a multilingual RAG dataset). This plot shows that the model capability varies by language. If the discrepancy was created by imbalanced data, what happens when we balance the pretraining data? 4 trained Cohere/aya-23-8B with a focus on balancing 23 languages in their pretraining data as much as possible. In Figure 1, we also included the performance of this Aya model. With the language-balanced pretraining dataset, Aya is able to perform more consistently across languages. However, it doesnât perform as well as Qwen does on English. Of course, the models differ in model architecture, ingested data, training strategies, and more, making this more illustrative rather than controlled. Still, this pattern matches the lessons from prior work: 5âs âCurse of Multilingualityâ for unsupervised learning states that as more languages are added, the overall performance plateaus. Overall, we are seeing a pattern of degradation when models are trained with multilingual data. To summarize, the multilingual tradeoff is to either optimize for English performance and suffer in non-English tasks, or to optimize for consistency in English/non-English performance and suffer in overall model accuracy. We try to explore a middle ground: if data is truly that important, can we curate post-training data to teach a model perform well both across languages and overall? We introduce HOTFIXR: a post-training data generation framework that aims to improve an LLMâs multilingual ability without degrading its general abilities. HOTFIXR is a synthetic data generation framework that trains a question generation model based on a student modelâs feedback signals. The question generator is trained to probe the studentâs weakness (Section 3). We evaluate HOTFIXR across a variety of baselines and benchmarks to support our claims (Section 4). Figure 1: Motivating results for LSC: the performance of Aya (with a more language-balanced pretraining dataset) versus Qwen (with a more English-dominant pretraining dataset) on multilingual HotPotQA (11). We see that although Aya performs more consistently across languages, Aya degrades in performance (see Qwenâs performance in English). 2 Related Works Multilingual LLMs try to map the same semantic phrases into one unified representation space (28; 7). However, representation misalignments occur within languages themselves. This usually emerges in three behaviors: (1) language-specific knowledge (1) where a model can output a different answer to the same factual query in different languages, (2) non-isomorphism (27) where certain words or phrases donât have direct translations in other languages, (3) non-compositionality (2) where the meaning cannot be derived (and therefore, translated) from individual words, including idioms and metaphors. Previous work has shown this misalignment is mainly due to the English-dominant pretraining data (4). This gears the model towards internally representing language in English (26; 19; 32). Particularly, 26 show that models route all their latent thinking into English spaces in the early-to-middle layers, then route their thinking back to non-English spaces in the later layers. This results in English becoming an upper bound for representation: concepts that can be represented in English are more likely to result in accurate answers. The key is to realign these concepts; Section 3.2 shows evidence that they can be realigned in pretraining, but pretraining is prohibitively expensive in many settings. Mitigating misalignment Misalignment can be mitigated using a few strategies. First there is smart prompting. Researchers craft particular prompts that elicit language-dependent knowledge from a model to increase performance on a multi-lingual task (6). The limitation of this approach is that the prompts are hand-crafted and might not generalize across language models. Models can also be sensitive to the style of the prompt, making this an unreliable method. Next is contrastive learning. Usually, a contrastive learning objective is employed to re-align representations (24; 13; 31) and requires positive-negative paired data in order to improve the representations. This is more reliable than prompting, but still requires high-quality samples with human ground truth. Focusing on multilingual reasoning specifically, one way to mitigate is to make architectural changes. For example, 29 uses a multilingual model to provide embedding inputs to a reasoning model to use the specialized capabilities of each model. Other fine-tuning methods either try to align low-resource language reasoning chains with those from high-resource languages (22) or carefully try to craft correct reasoning traces in other languages to improve the performance (18). Finally, other works also focus on data curation for aligning representations. Multilingual data curation. Most of the data curation works are focused on filtering and cleaning up existing datasets (16; 12; 17; 14). The cleaning strategies involve deduplication and removing multilingual documents. The filtering strategies involve model-based data selection (keeping training samples that are within the distribution of the existing reference set). 12 use a form of self-distillation where they generate cross-lingual samples (instruction in one language, but response in a low-resource language). This helps to generate QA pairs with fewer translation mistakes. 3 HOTFIXR Methodology HOTFIXR is a synthetic data generation framework in which we train a question generation model based on the student modelâs signals. Our framework is inspired by AcquisitionSynthesis (3). First, a question generation model is asked to generate a data sample.11 1 The question generation model is given a prompt that outlines the data sample requirements and we even provide an in-context sample. Prompts are available in Appendix E. Second, the student model is provided with the data sample, where we design an acquisition function22 2 Acquisition functions are primarily known as the selection criteria within data selection and active learning works (20). We use the term âacquisition functionâ because it is a metric that informs us how good our data is for training purposes. that indicates whether the sample uncovers a student modelâs lingual deficit. The acquisition function is described in Section 3.1, but is essentially a difficulty and misalignment score that indicates whether the sample elicits a student model weakness. Finally, this score is given as feedback to the question generation model, where it will be optimized to generate samples that maximize the difficulty score. The optimization algorithm is GRPO (21). Figure 3 contains a visualization of HOTFIXR. Figure 2: Intuition behind HOTFIXR. The âdata spaceâ represents the data samples generated by the question generator. Essentially, we want to hill climb through the data space to find samples that have high lingual deficit scores, as they will be most informative to the student model. By fine-tuning the question generation model to generate samples according to the student modelâs lingual deficit, we can improve the information embedded within data, improving our models on downstream tasks. Figure 3: An illustration of how to train the question generation model in HOTFIXR, along with the reward design. 3.1 Measuring Lingual Deficit Student models have two kinds of weaknesses: language-agnostic incompetencies (cannot answer the question correctly in any language) and language-specific incompetency (cannot answer the some questions correctly in certain languages). We need a metric that can capture both weaknesses. To measure language-agnostic incompetency (LAI), we use the LLMâs uncertainty in answering the question by using English-centric reasoning. Once the student model receives the question generated by the question generation model, we prompt the student model to âanswer the following question, reason using Englishâ. Given a prompt P and response R (where Rt(i)R_t^(i) is the i-th most probable token in the t-th spot), the uncertainty of a model is UâĄ(P,R)=1â(1|R|ââtpθâ(Rt(1)âŁP,R<t))U(P,R)=1- ( 1|R| _tp_θ(R_t^(1) P,R_<t) ) or 1 minus the probability of the most probable token, averaged across the sequence. A higher reward is assigned to the question generator if the student model is uncertain about the question in its strongest language33 3 We assume this to be English, because the majority of pre-training data is in English.. To measure language-specific incompetence (LSI), we prompt the student model to answer the same model-generated question in two ways: one with âanswer the following question, reason using Englishâ (REâNR_EN), and the other with âanswer the following question, reason using Lâ (RLR_L), where L alternates between French, Spanish, Arabic, Portuguese, or Italian. Next, we find the last reasoning token of both REâNR_EN and RLR_L (the token right before the answer delimiter), extract the modelâs hidden state of the last layer for that token, and compute the cosine distance between them. The question generator is assigned a higher reward if the distance is larger between the two. Since the correct answer is invariant to the reasoning language, the two reasoning trajectories should be represented in nearby regions of the representation space. If they differ, this represents a cross-lingual misalignment (26) that the student needs to be trained to reduce. Intuition. A question generation model that is trained on a student modelâs feedback signals learns to generate data that is good for the student to learn from. Figure 2 is a schematic of an arbitrary "data" space (what the question generation model can generate) versus the score from the acquisition function. During GRPO optimization, the question generation model figuratively probes the data space to understand what data elicits the highest peak in the acquisition function. In this case, âgoodâ for the student is where the reasoning diverges: either because of incorrect answers or because of the representational gap between languages. Of course, there will always be some difference in the representation of two languages, but the aim is to minimize it so that models can generalize well to other languages without requiring training data in those languages. 3.2 Training Setup To train our question generator model, we use 500 samples from the training dataset QâGD_QG. These 500 samples are used as in-context samples to inform the data sample generation during rollouts. Our training size is small because we see that with more training samples, models reward hack the student feedback signal, and start to generate the one sample that elicits a local maximum reward. Empirically, 500 samples are a sweet spot between maximizing reward and avoiding reward hacks (more in Appendix A). After training our question generation model, we use it to create a dataset for a student model, SD_S, of 5,000 samples. In both the training dataset QâGD_QG and generated dataset SD_S, we prompt the model to generate a sample in a particular language between English, French, Spanish, Arabic, Portuguese, and Italian. To obtain reliable answers for SD_S, we use Qwen/Qwen2.5-32B-Instruct (25) as our label generation model. We train the student model on SD_S with SFT. Then, we evaluate the student model. The better the student model, the higher quality SD_S is. 4 Experiments In this section, we describe our thorough evaluations. 4.1 Setup For HOTFIXR question generation training set QâTD_QT, we use âź 500 samples of Nemotron-PTDv2âs STEM, MATH, and CHAT subsets (167 samples each). We generate SD_S based on the Nemotron data as well. We use the student models trained on these data to showcase the benefits of our method. Models. We use three models for this task: Qwen/Qwen2.5-7B-Instruct and Qwen/Qwen2.5-14B-Instruct (25), and meta-llama/Llama-3.1-8B-Instruct (8). We chose this set to showcase the robustness of our method across model families and model sizes. All question generation models and student models are different instantiations of the same instruction-tuned model. In other words, a slightly different Qwen model will teach itself. Tasks. We break up our evaluation into four tasks: (1) Nemotron (15) for agentic reasoning, (2) factual parametric knowledge, tested with MMMLU 44 4 https://huggingface.co/datasets/openai/MMMLU, (3) RAG-based reading comprehension, tested with mHotPotQA (11), and (4) translation, tested with OPUS-100 (30). These four tasks are used to determine whether HOTFIXR can teach models that do not forget previously trained information, as well as improve the multilingual performance of LLMs. Languages. As mentioned before, HOTFIXR will generate data in English, French, Spanish, Arabic, Portuguese, and Italian. These are our in-distribution languages â our training datasets are equally balanced among all six languages. In addition, we use German and Japanese as our out-of-distribution languages, to ensure the performance does not degrade across languages. Metrics. Each task requires a different kind of similarity metric. Nemotron STEM and MMMLU are multiple-choice question answering tasks. Hence, these tasks are measured by Accuracy, or the percentage of test samples where the LLM chose the same answer as the ground truth. mHotPotQA (multilingual HotPotQA) has short answers, so we use ROUGE-L to capture the lexical agreement by measuring the n-gram overlap between the few-worded answers. OPUS-100 also has short phrases, but could require a metric to reflect the semantic similarity. Hence, we use an LLM-as-a-Judge (LAJ) to determine on a scale of 1-5 whether an English sentence and the translated version match semantically. We use Prometheus-7b-V2.0 (10) as our LAJ. Finally, we have two other in-distribution tasks that require semantic similarity for evaluation: Nemotron MATH and Nemotron CHAT. We also use LAJ to evaluate these. A prediction with at least a score of 4/5 is considered correct. To make the metrics comparable, we discretize them into binary correct/incorrect labels, and report the % of predictions that are correct. Accuracy is already the binary scale. A prediction with at least an 80% ROUGE score or at least 4/5 LAJ score is considered correct. Baselines. To ensure our method is competitive and comparable, we adopt a variety of baselines to prove the various aspects of our claims. 1. Base: the performance of the base models. 2. EngReason: training-free baseline that measures the performance of base models when the model is prompted to âreason about the question in Englishâ. This baseline represents the strategy of falling back on English as the âlanguage of thoughtâ. 3. SelectionGT: this is a data selection baseline. We use 5,000 samples (with their questions, and ground truth answers) from the Nemetron PTDv2STEM, MATH, and CHAT subsets to train a student model. The data selection baselines help evaluate data synthesis compared to just using the available data. 4. SelectionGEN: this is also a data selection baseline. The only difference between this and SelectionGT is where the labels come from. SelectionGEN generates the labels using Qwen/Qwen2.5-32B-Instruct, instead of obtaining them from the original dataset, and helps isolate the role of ground truth versus generated labels. 5. Filtered: this is another data selection baseline where we use the Lingual Deficit score to rank 10,000 samples. We select 5,000 samples with the highest Lingual Deficit score. 6. Untrained: this is a data synthesis baseline. We use the base, untrained model as question generation models to generate datasets for student models. This clarifies the importance of optimizing the data generation models. 7. DataEnvGym: this is another data synthesis baseline from 9. They use a teacher model to synthesize data based on a studentâs weaknesses, irrespective of language. To adapt it to our setting, we add to their prompts to generate data in a particular language (same as ours), and we keep the teacher and student models the same as in our settings. 4.2 Results All of the results reported in this section are averaged across 3 runs. Weâve only reported the averages in the main plots. To see the error bars, please refer to Appendix B. Figure 4 showcases the performance of all student models from baselines and HOTFIXR for all three datasets, on the in-distribution task of Nemotron STEM, MATH, and CHAT. To clarify, in all the plots, the bars with Ă are the untrained baselines, the bars with circles are data selection baselines, and the bars with diagonal lines are data synthesis baselines. Figure 4: Performance of data curation methods in-distribution. Figure 5: Performance of data curation methods for factual (MMMLU) queries, an out-of-distribution task. Figures 5, 6, and 7 compare the performance of all student models trained (or untrained) by all baselines on out-of-distribution tasks, including factual parametric knowledge, RAG, and translation, respectively. These plots have whiskers on each plot. For clarity, we report the average result across languages for each baseline. The whisker length indicates the minimum and maximum language performance. The shorter the whisker, the more consistent the performance is across languages. In Appendix B, we report the granular per-language results for each task. Figure 6: Performance of data curation methods for RAG (multilingual HotPotQA) queries, an out-of-distribution task. Figure 7: Performance of data curation methods for translation (OPUS-100) queries, an out-of-distribution task. 4.3 Analysis Takeaway 1: HOTFIXR generates more informative in-distribution data. On Nemotron STEM, MATH, and CHAT tasks, HOTFIXR gets an overall improvement of 4.3% over the base model (5.5% on Qwen 7B, 3.1% on Qwen 14B, and 4.4% on Llama 8B), improvement of 6.0% over the best selection method (7.5% on Qwen 7B, 6.2% on Qwen 14B, and 4.4% on Llama 8B), and an improvement of 5.8% over the best synthesis method (6.6% on Qwen 7B, 6.1% on Qwen 14B, and 4.8% on Llama 8B). In-Distribution Out-of-Distribution Baseline Î Wins Î Wins Base +4.3 9/9 -0.9 3/9 EngReason +5.4 9/9 -0.7 4/9 SelectionGT +7.0 9/9 +3.8 9/9 SelectionGEN +6.0 9/9 +8.9 8/9 Filtered +7.3 9/9 +6.6 8/9 Untrained +5.8 9/9 +1.4 7/9 DataEnvGym +7.1 9/9 +7.3 9/9 Average +6.2 9/9 +3.7 7/9 Table 2: Average performance difference of HOTFIXR versus each baseline. The average is over nine settings: three students (Qwen 7B, Qwen 14B, Llama 8B) Ă three ID/OOD tasks. âWinsâ counts the number of settings where HOTFIXR scores higher. Takeaway 2: HOTFIXR minimizes out-of-distribution degradation. Because all baselines train on 5,000 samples, it is expected that the models will lose some OOD performance relative to the base model. There are two ways we evaluate for OOD generalization performance: by task, and by language. By OOD task generalization, we refer to the performance on Factual Knowledge (MMMLU), Translation (OPUS-100), and RAG (mHotPotQA). According to Table 2, HOTFIXR performance degrades by 0.9% compared to the base model (averaged over three OOD tasks and three students). However, HOTFIXR is on average 5.6% better than other training-based baselines. Compared to SelectionGT and SelectionGEN, we achieve a 6.4% improvement. This is because the questions are generic in the selection methods, and not targeted towards a modelâs weaknesses. The effect of Untrained versus HOTFIXR in OOD tasks is small but nontrivial (1.4%), which shows that synthetic data will not harm OOD. Still, HOTFIXR achieves a 5.8% improvement in-distribution, so training data generation models is still important to achieve overall performance. Finally, over DataEnvGym, HOTFIXR performs 7.3% better in OOD tasks. This is because DataEnvGym is designed to generate data that mitigate student mistakes on a particular dataset. Instead, HOTFIXRâs dataset-agnostic reward is able to probe a student modelâs general weaknesses. Method De Ja Ru Zh Avg SelectionGT -3.3 -4.7 -9.5 -10.5 -7.0 SelectionGEN -1.9 -4.0 -22.5 -21.5 -12.5 Filtered -3.8 -4.8 -14.7 -16.3 -9.9 Untrained -2.9 -2.0 -2.3 -5.3 -3.1 DataEnvGym -6.8 -6.0 -12.2 -15.4 -10.1 HOTFIXR -2.6 -0.7 -0.9 -1.3 -1.4 Table 3: OOD language generalization: performance differences relative to the Base model, averaged across all models and corresponding tasks. Averaging the difference in HOTFIXR and other methods, we can compute how much HOTFIXR avoids catastrophic forgetting. For example, averaged across all languages, HOTFIXR avoids catastrophic forgetting by 8.7% compared to DataEnvGym (10.1-1.4). By OOD language generalization, we refer to the performance on languages not included in the training distribution. In Factual Knowledge (MMMLU) and Translation (OPUS-100), these are German and Japanese. In RAG (mHotPotQA), these are Russian and Chinese. Please refer to Appendix B for the per-language results of all OOD tasks. Table 3 summarizes the results. Overall, there is post-training decline in performance. However, HOTFIXR suffers the least: it performs better by 7.1% compared to other training-based baselines. Method Range Trim. Range Std IQR CV Base 20.2 16.0 7.2 9.6 0.12 EngReason 20.7 15.4 7.2 8.8 0.12 SelectionGT 22.3 17.1 7.9 9.8 0.13 SelectionGEN 18.8 14.9 6.9 9.5 0.13 Filtered 22.6 17.1 8.0 10.6 0.14 Untrained 20.1 15.0 7.1 9.7 0.12 DataEnvGym 22.8 17.0 8.0 10.2 0.14 HOTFIXR 19.4 14.6 6.9 9.4 0.11 Table 4: Cross-lingual spread of student performance, averaged over three student models (Qwen2.5-7B, Qwen2.5-14B, Llama-3.1-8B) and three multilingual tasks (OPUS, MMMLU, mHotpotQA) evaluated with each taskâs primary metric. Lower is more consistent for all dispersion columns; Bold marks the best value among data-generation methods. Range is the max â- min performance across all languages. Trimmed Range is the best minus the 2nd worst language accuracy. Std is standard deviation of the performance. IQR is the Inter-Quartile Range (difference in the 25% and 75% quartile). CV is the coeffiecient of variation (the standard devation divided by the mean). Takeaway 3: HOTFIXR modestly improves multilingual consistency from the base model. Table 4 reports various spread measures for determining the multilingual consistency of models. The Range is the max â- min performance across all languages. But, Range is not a reliant metric, as it suffers from outliers of low-resource languages. To avoid this outlier, we also report the Trimmed Range, which is the best minus the 2nd worst language accuracy. We also report the Standard Deviation (Std), the Inter-Quartile Range (IQR), and the Coefficient of Variation (CV). All of these methods show that HOTFIXR is slightly more consistent, but there is no SOTA data curation method for ensuring multilingual consistency. Still, we can say that HOTFIXR, among other data curation methods, can maintain (and sometimes, improve) the multilingual consistency of LLMs. Lingual Deficit score ablations. Table 5 contains an ablation over the lingual deficit score components for training the question generator. We see that the format reward, LAI, and LSI do not have strong effects individually. But combined, they are able to train a strong data generator. Knowledge Distillation. In Appendix C, we test the effects of knowledge distillation (both questions and labels) from Qwen 32B. The results show that although it can improve performance, the majority of the gains come from the learned question generation in HOTFIXR. Continual Learning. In Appendix D, we train a fresh question generator on a continuously trained student, and see significant improvements of 7% on Nemotron tasks, 8% on OOD tasks, and 8% on seen v.s. 7% on unseen languages. ID OOD Model Config Nemo. Trans. Fact. RAG Qwen 7B Format 52.0 57.6 64.0 75.8 + LAI 50.9 57.9 64.3 77.2 + LSI 50.8 57.8 63.6 77.7 + LAI + LSI 57.8 59.1 64.3 76.7 Llama 8B Format 48.2 57.8 49.9 70.2 + LAI 48.5 57.2 49.2 69.2 + LSI 47.5 57.5 49.3 67.8 + LAI + LSI 52.5 58.9 49.3 69.7 Qwen 14B Format 49.3 58.0 71.5 79.8 + LAI 50.9 58.6 66.3 76.0 + LSI 48.8 57.2 62.3 76.8 + LAI + LSI 58.3 58.2 67.9 78.3 Table 5: Reward-signal ablation. All rows include the format reward. âNemo.â is Nemotron, âTrans.â is Translation, âFact.â is Factual. Highlights the impact each reward hasâthere is more effect of the reward functions combined, than individually. LAI is the âLanguage Agnostic Incompetencyâ and LSI is the âLanguage Specific Incompetencyâ, as described in Section 3. 4.4 Discussion In business use cases, where it is important to adapt a model to a particular task while also maintaining the modelâs general reasoning, context understanding, and answer generation abilities, HOTFIXR offers the best of both worlds. Compared to the base model, it achieves a 4.3% improvement, and loses only 0.9% performance on out-of-distribution task. Ultimately, fine-tuned models will lose some generalization performance for a price. Across all the training-based baselines, HOTFIXR achieves the best performance (6.2% average improvement) and generalization (5.6% average improvement). Still, HOTFIXR costs more due to the GRPO optimization of the data generator. With roughly 40 steps of RL optimization and 6 A100 NVIDIA GPUs, HOTFIXR occurs a cost of 0.96min/training data sample (which is roughly 8 hours for |QâG|=500|D_QG|=500) to train a data generator. However, 8 GPU-hours is a one-time cost. Once a model is trained for data generation, it can generate as much data as required. Furthermore, there are no efforts required to obtain high-quality seed data, clean generated data, or remove any PII content. HOTFIXR does increase training time, but it makes up for it with consistent improvements in distribution, and reliable generalization performance on out-of-distribution tasks. 5 Conclusion In this paper, we tackle the problem of adapting models to improve their multilingual abilities without degrading the overall performance. To do so, we present HOTFIXR: a principled data synthesis pipeline that trains data generation models to generate data that targets the improvement of agentic abilities and the consistency in multilingual settings. Our reward for the data generator finds samples that are both difficult (LAI) and have larger multilingual representation gaps (LSI). In our experimentation, we see that HOTFIXR is able to improve performance on in-distribution benchmarks, and remain reliable on OOD tasks and languages. Future work involves scaling experiments in the pretraining regime, where most of the multilingual representations are formed. 6 Limitations Our work primarily focuses on high-resource languages: English, French, Spanish, German, Japanese, Portuguese, Italian. Even though we also test with Arabic, accounting for low-resource languages requires special consideration, that we aim to address in future work. For now, a demonstration of fine-tuning models to generate good data for high-resource languages is a nontrivial task in itself. Furthermore, with LLMs that face heavy training data biases, there is a risk of generating harmful dataâwe restrict our study to verifiable tasks that have clear correct and incorrect answers. We cannot predict whether our findings will extrapolate to unverifiable domains. Finally, our results rely on a strong label generation model that can output reliable labels (Qwen/Qwen2.5-32B-Instruct). In preliminary experiments, we tried generating labels with the respective student models themselves, but their instruction following abilities (in particular, ensuring that they generated labels in the prompted language) are poor. So we opted for generating labels with larger models. In spite of this weakness, we still showcase improved performance with few samplesâin future work, we will explore a synthetic, active learning setup where models can generate questions that they want labeled, but also reduce the amount of data used. This can help ensure the labeling cost (either by LLM or by human) is reduced. References Agarwal et al. (2026a) I. Agarwal, N. B. Bozdag, N. Patel, and D. Hakkani-TĂźr Language specific knowledge: do models know better in X than in English?. External Links: 2505.14990, Link Cited by: §2. Agarwal et al. (2026b) I. Agarwal, Z. He, D. Patil, and D. Hakkani-TĂźr A rising tide lifts all boats: mtqe rewards for idioms improve general translation quality. External Links: 2601.06307, Link Cited by: §2. Agarwal et al. (2026c) I. Agarwal, S. Stoica, E. C. Acikgoz, P. Natarajan, M. Namazifar, J. Ma, and D. Hakkani-TĂźr AcquisitionSynthesis: targeted data generation using acquisition functions. External Links: 2605.13149, Link Cited by: §3. Aryabumi et al. (2024) V. Aryabumi, J. Dang, D. Talupuru, S. Dash, D. Cairuz, H. Lin, B. Venkitesh, M. Smith, K. Marchisio, S. Ruder, A. Locatelli, J. Kreutzer, N. Frosst, P. Blunsom, M. Fadaee, A. ĂstĂźn, and S. Hooker Aya 23: open weight releases to further multilingual progress. External Links: 2405.15032 Cited by: §1, §2. Conneau et al. (2020) A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. GuzmĂĄn, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, p. 8440â8451. External Links: Link, Document Cited by: §1. Donthi et al. (2025) S. Donthi, M. Spencer, O. B. Patel, J. Y. Doh, E. Rodan, K. Zhu, and S. OâBrien Improving LLM abilities in idiomatic translation. In Proceedings of the First Workshop on Language Models for Low-Resource Languages, H. Hettiarachchi, T. Ranasinghe, P. Rayson, R. Mitkov, M. Gaber, D. Premasiri, F. A. Tan, and L. Uyangodage (Eds.), Abu Dhabi, United Arab Emirates, p. 175â181. External Links: Link Cited by: §2. Ghosh et al. (2025) A. Ghosh, D. Datta, S. Saha, and C. Agarwal A survey of multilingual reasoning in language models. External Links: 2502.09457, Link Cited by: §2. Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. GuzmĂĄn, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Ăelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §4.1. Khan et al. (2025) Z. Khan, E. Stengel-Eskin, J. Cho, and M. Bansal DataEnvGym: data generation agents in teacher environments with student feedback. External Links: 2410.06215, Link Cited by: item 7. Kim et al. (2024) S. Kim, J. Suk, S. Longpre, B. Y. Lin, J. Shin, S. Welleck, G. Neubig, M. Lee, K. Lee, and M. Seo Prometheus 2: an open source language model specialized in evaluating other language models. External Links: 2405.01535 Cited by: §4.1. Li et al. (2026) B. Li, Z. Xu, and R. Xie Language drift in multilingual retrieval-augmented generation: characterization and decoding-time mitigation. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), p. 31519â31526. Cited by: Figure 1, §1, §4.1. Li et al. (2024a) C. Li, W. Yang, J. Zhang, J. Lu, S. Wang, and C. Zong X-instruction: aligning language model in low-resource languages with self-curated cross-lingual instructions. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 546â566. External Links: Link, Document Cited by: §2. Li et al. (2024b) G. Li, X. Zhao, A. Jafari, W. Shao, R. Farahbakhsh, and N. Crespi Improving cross-lingual transfer with contrastive negative learning and self-training. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), N. Calzolari, M. Kan, V. Hoste, A. Lenci, S. Sakti, and N. Xue (Eds.), Torino, Italia, p. 8781â8791. External Links: Link Cited by: §2. Messmer et al. (2025) B. Messmer, V. SabolÄec, and M. Jaggi Enhancing multilingual LLM pretraining with model-based data selection. In Proceedings of the 10th edition of the Swiss Text Analytics Conference, J. Gerber, M. Cieliebak, D. Tuggener, and M. HĂźrlimann (Eds.), Winterthur, Switzerland, p. 31â56. External Links: Link Cited by: §2. Nathawani et al. (2025) Nemotron-Post-Training-Dataset-v2 External Links: Link Cited by: §4.1. Nguyen et al. (2024) T. Nguyen, C. V. Nguyen, V. D. Lai, H. Man, N. T. Ngo, F. Dernoncourt, R. A. Rossi, and T. H. Nguyen CulturaX: a cleaned, enormous, and multilingual dataset for large language models in 167 languages. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), N. Calzolari, M. Kan, V. Hoste, A. Lenci, S. Sakti, and N. Xue (Eds.), Torino, Italia, p. 4226â4237. External Links: Link Cited by: §2. Penedo et al. (2025) G. Penedo, H. KydlĂÄek, V. SabolÄec, B. Messmer, N. Foroutan, A. H. Kargaran, C. Raffel, M. Jaggi, L. V. Werra, and T. Wolf FineWeb2: one pipeline to scale them all â adapting pre-training data processing to every language. In Second Conference on Language Modeling, External Links: Link Cited by: §2. Ranaldi and Pucci (2025) L. Ranaldi and G. Pucci Multilingual reasoning via self-training. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 11566â11582. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §2. Schut et al. (2025) L. Schut, Y. Gal, and S. Farquhar Do multilingual llms think in english?. External Links: 2502.15603, Link Cited by: §2. Settles (2012) B. Settles Active learning. Morgan & Claypool Publishers. Cited by: footnote 2. Shao et al. (2024) Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. K. Li, Y. Wu, and D. Guo DeepSeekMath: pushing the limits of mathematical reasoning in open language models. CoRR abs/2402.03300. External Links: Link Cited by: §3. She et al. (2024) S. She, W. Zou, S. Huang, W. Zhu, X. Liu, X. Geng, and J. Chen MAPO: advancing multilingual reasoning through multilingual-alignment-as-preference optimization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 10015â10027. External Links: Link, Document Cited by: §2. Singh et al. (2024) S. Singh, F. Vargus, D. Dsouza, B. F. Karlsson, A. Mahendiran, W. Ko, H. Shandilya, J. Patel, D. Mataciunas, L. OMahony, M. Zhang, R. Hettiarachchi, J. Wilson, M. Machado, L. S. Moura, D. KrzemiĹski, H. Fadaei, I. ErgĂźn, I. Okoh, A. Alaagib, O. Mudannayake, Z. Alyafeai, V. M. Chien, S. Ruder, S. Guthikonda, E. A. Alghamdi, S. Gehrmann, N. Muennighoff, M. Bartolo, J. Kreutzer, A. ĂstĂźn, M. Fadaee, and S. Hooker Aya dataset: an open-access collection for multilingual instruction tuning. External Links: 2402.06619 Cited by: §1. Tan et al. (2023) W. Tan, K. Heffernan, H. Schwenk, and P. Koehn Multilingual representation distillation with contrastive learning. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, A. Vlachos and I. Augenstein (Eds.), Dubrovnik, Croatia, p. 1477â1490. External Links: Link, Document Cited by: §2. Team (2024) Q. Team Qwen2.5: a party of foundation models. External Links: Link Cited by: §1, §3.2, §4.1. Wendler et al. (2024) C. Wendler, V. Veselovsky, G. Monea, and R. West Do llamas work in English? on the latent language of multilingual transformers. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 15366â15394. External Links: Link, Document Cited by: §1, §2, §3.1. Wu et al. (2024) D. Wu, Y. Lei, A. Yates, and C. Monz Representational isomorphism and alignment of multilingual large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 14074â14085. External Links: Link, Document Cited by: §2. Xu et al. (2025) Y. Xu, L. Hu, J. Zhao, Z. Qiu, K. Xu, Y. Ye, and H. Gu A survey on multilingual large language models: corpora, alignment, and bias. Frontiers of Computer Science 19 (11). External Links: ISSN 2095-2236, Link, Document Cited by: §2. Yoon et al. (2024) D. Yoon, J. Jang, S. Kim, S. Kim, S. Shafayat, and M. Seo LangBridge: multilingual reasoning without multilingual supervision. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 7502â7522. External Links: Link, Document Cited by: §2. Zhang et al. (2020) B. Zhang, P. Williams, I. Titov, and R. Sennrich Improving massively multilingual neural machine translation and zero-shot translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, p. 1628â1639. External Links: Link, Document Cited by: §4.1. Zhang et al. (2026) Y. Zhang, H. Mouratidis, and R. Shekhar Speak in context: multilingual asr with speechâcontext alignment via contrastive learning. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), S. Piperidis, N. Bel, H. van den Heuvel, N. Ide, S. Krek, and A. Toral (Eds.), Palma, Mallorca, Spain, p. 5873â5882. External Links: Document Cited by: §2. Zhao et al. (2024) Y. Zhao, W. Zhang, G. Chen, K. Kawaguchi, and L. Bing How do large language models handle multilingualism?. In Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37, p. 15296â15319. External Links: Document, Link Cited by: §2. Appendix A Reward Hacking In Section 3.2, we mention that our training dataset DQâGD_QG is only 500 because the question generator reward hacks. We would like to clarify what that means. Although Figure 2 is an illustrative example, it can still help us understand reward hacking. When a data generation model reward hacks, it finds the sample with the local maximum lingual deficit reward. It will not explore further and gets stuck there. During inference, this results in the data generator generating that one sample (with the local maximum reward) over and over again. Even if we change the in-context sample, the generator will generate the exact same data point repeatedly. This means that the resulting dataset will be 5,000 copies of the exact same sample, which will cause trained student models to fail catastrophically. Hence, to avoid reward hacking, we simply reduce the number of optimization steps in GRPO (to around 41 steps) and reduce the number of training samples. Appendix B Error Bars for Performance on ID and OOD tasks, and Per-Language Performance on OOD tasks. Figure 8 contains the error bars for all Nemotron tasks. Figures 9, 10 and 11 contain the error bars as well as the per-language performance of all methods, for each model on the factual, translation, and RAG tasks, respectively. Because the error bars are quite small on the figures themselves, we also note down the minimum and maximum standard deviations for each subtask for the Nemotron (Table 6), Translation (Table 7), Factual (Table 8) and RAG (Table 9)) tasks. Qwen 7B Qwen 14B Llama 8B Method min max min max min max Base 0.30 1.50 0.10 0.40 0.20 2.00 EngReason 0.25 1.23 0.16 0.75 0.34 1.00 SelectionGT 0.20 3.40 0.10 1.00 0.40 2.20 SelectionGEN 0.30 3.10 0.00 0.90 0.10 2.20 Filtered 0.70 0.90 0.30 0.80 0.30 1.70 Untrained 0.40 1.90 0.00 1.50 0.30 0.70 DataEnvGym 0.40 1.20 0.10 0.70 0.20 1.70 HOTFIXR 0.10 0.90 0.20 0.40 0.05 0.50 Table 6: Minimum and maximum standard deviation (across 3 runs) for the Nemotron task (STEM, MATH, and CHAT), per model and method. This explains the small error bars in Figure 8. Qwen 7B Qwen 14B Llama 8B Method min max min max min max Base 0.43 0.85 0.41 1.55 0.58 1.56 EngReason 0.40 0.83 0.43 1.05 0.55 1.56 SelectionGT 0.10 1.40 0.20 1.00 0.40 1.60 SelectionGEN 0.00 0.80 0.00 1.50 0.10 1.50 Filtered 0.10 1.60 0.10 0.80 0.10 1.60 Untrained 0.20 0.50 0.00 0.70 0.30 1.70 DataEnvGym 0.00 1.10 0.20 0.70 0.20 0.80 HOTFIXR 0.00 0.90 0.00 0.80 0.20 1.10 Table 7: Minimum and maximum standard deviation (across 3 runs) for the Translation task (OPUS-100), per model and method. This explains the small error bars in Figure 10. Qwen 7B Qwen 14B Llama 8B Method min max min max min max Base 0.10 1.50 0.10 1.10 0.00 2.30 EngReason 0.09 1.27 0.28 1.09 0.33 1.73 SelectionGT 0.30 2.30 0.20 1.80 0.10 1.60 SelectionGEN 0.40 2.60 0.10 1.80 0.10 2.30 Filtered 0.10 1.30 0.20 1.20 0.00 1.20 Untrained 0.30 1.00 0.10 1.00 0.20 2.20 DataEnvGym 0.30 2.40 0.10 1.40 0.10 1.20 HOTFIXR 0.30 1.40 0.30 1.30 0.30 2.60 Table 8: Minimum and maximum standard deviation (across 3 runs) for the Factual task (MMMLU), per model and method. This explains the small error bars in Figure 9. Qwen 7B Qwen 14B Llama 8B Method min max min max min max Base 0.10 0.90 0.10 0.80 0.80 2.50 EngReason 0.33 0.99 0.16 1.05 0.53 1.71 SelectionGT 0.30 1.20 0.10 0.80 0.30 1.80 SelectionGEN 0.80 1.40 0.30 1.50 0.00 2.50 Filtered 0.10 1.10 0.40 1.50 0.00 1.30 Untrained 0.20 1.10 0.30 1.10 0.10 1.30 DataEnvGym 0.20 1.40 0.30 0.60 0.30 1.20 HOTFIXR 0.00 0.60 0.20 1.50 0.00 1.40 Table 9: Minimum and maximum standard deviation (across 3 runs) for the RAG task (mHotPotQA), per model and method. This explains the small error bars in Figure 11. Figure 8: Performance of data curation methods for Nemotron queries, per task. Figure 9: Performance of data curation methods for factual (MMMLU) queries, per language. Figure 10: Performance of data curation methods for translation (OPUS-100) queries, per language. Figure 11: Performance of data curation methods for RAG (multilingual HotPotQA) queries, per language. Appendix C Knowledge Distillation In order to separate the gains from the question generation versus Qwen 32Bâs label generation, we use Qwen 32B to generate both the questions and labels. This is similar to the Untrained baseline, but with Qwen 32B. Table 10 contains those results, in which we call the new baseline âDistillation (32b)â. The table shows that while distillation with a larger model does help improve performance, HOTFIXRâs trained question generation model is able to understand model weaknesses much better than simply distilling knowledge from a larger model. Method Nemotron Factual Translation RAG Qwen2.5-7B-Instruct Base 52.3 67.0 59.2 79.0 EngReason 51.2 68.2 59.0 79.3 SelectionGT 48.5 63.9 53.8 61.5 SelectionGEN 50.3 59.9 56.8 41.4 Filtered 48.6 61.7 56.6 66.5 mCOT 47.9 55.3 54.4 53.8 Untrained 51.2 59.2 57.1 72.3 DataEnvGym 49.3 46.2 56.9 47.5 Distillation (32B) 49.6 64.7 58.9 77.3 HOTFIXR 57.8 64.3 59.1 76.7 Qwen2.5-14B-Instruct Base 55.2 73.7 59.4 79.9 EngReason 54.1 73.0 59.5 81.0 SelectionGT 51.4 67.3 54.2 73.8 SelectionGEN 52.1 70.3 57.4 42.8 Filtered 50.9 59.1 59.1 45.5 Untrained 52.1 71.0 57.9 77.5 DataEnvGym 50.3 62.6 57.4 70.4 Distillation (32B) 53.3 69.5 59.0 78.4 HOTFIXR 58.3 67.9 58.2 78.3 Llama-3.1-8B-Instruct Base 48.1 48.5 56.3 67.7 EngReason 47.0 47.3 56.4 65.2 SelectionGT 47.6 48.8 57.6 67.2 SelectionGEN 48.1 48.6 57.6 67.2 Filtered 47.2 48.8 57.7 68.4 Untrained 47.7 49.5 57.1 68.6 DataEnvGym 47.6 48.8 57.5 69.1 Distillation (32B) 47.7 51.1 58.0 70.3 HOTFIXR 52.5 49.3 58.9 69.7 Table 10: Average performance per model, with the added Distillation method. Here, we are able to show that a lot of HOTFIXRâs empirical success comes from the targeted question generation, rather than having correct, distilled labels from a large model. Appendix D An Experiment on Continual Learning In this section, we describe an experiment where we test the effects of a second round of HOTFIXR. The question generator model is trained from scratch (the base model), but the student model is continuously trained (from the checkpoint of Round 0 of HOTFIXR). We call the first iteration Round 0 and the second iteration (of a fresh generator model, but continuously trained student model) Round 1. In Figure 12, we present the aggregate results and the per-task results of the experiment, respectively on the left and right. This experiment is only on the Qwen 7B setting. Figure 12: The delta in performance of student models trained with two rounds of HOTFIXR versus one round. The improvement with two rounds is much more than the improvement with one round. As shown, even though the question generation model is trained from scratch, each iteration of HOTFIXR improves the student model much more significantly, even on unseen, OOD tasks and languages. These improvements show that using HOTFIXR in iterations will yield higher gains not only in cross-linguality but also model response quality. Appendix E Prompts Figures 13-15 contain the three prompts used for HOTFIXRâs question generator training and data generation for each subset of Nemotron (MATH, STEM, and CHAT). Prompt for HOTFIXR training on Nemotron MATH In language, generate a NEW, ORIGINAL math problem that is AS DIFFICULT as the reference below. Do NOT copy, paraphrase, or reuse it in any way. Reference (difficulty calibration only â do not reproduce): <question> question </question> <reasoning> reasoning </reasoning> <answer> answer </answer> Requirements for your generated problem: - The question, reasoning, and answer should all be in language. - Requires non-trivial reasoning steps (no single-step shortcuts) - Draws from: number theory, combinatorics, algebra, geometry, or probability - Is self-contained and precisely stated - Reasoning should include a complete step-by-step derivation - Answer includes just the final result IMPORTANT: generate a (question, reasoning, answer) triplet; wrap your question, reasoning, and answer in the following special tokens: <question> Insert your language question here. </question> <reasoning> Insert the thinking and general reasoning here in language. </reasoning> <answer> Insert your short language answer here. </answer> Figure 13: Prompt used during HOTFIXR training and data generation, for the Nemotron MATH subset. As input, the prompt takes a question, answer, and reasoning from the Nemotron MATH training dataset, and a language that alternates between English, French, Spanish, Arabic, Portuguese, or Italian. Prompt for HOTFIXR training on Nemotron STEM In language, generate a NEW, ORIGINAL STEM multiple-choice question (MCQA) that is AS DIFFICULT as the reference below. Do NOT copy, paraphrase, or reuse it in any way. Reference (difficulty calibration only â do not reproduce): <question> question </question> <reasoning> reasoning </reasoning> <answer> answer </answer> Requirements for your generated question: - The question, reasoning, and answer should all be in language. - Draws from STEM domains: physics, chemistry, biology, computer science, engineering, or mathematics - Requires non-trivial conceptual or quantitative reasoning (no single-step lookups or trivial recall) - Has exactly 4 answer choices labeled (A), (B), (C), (D) â only one is correct - Distractors are plausible and reflect common misconceptions or near-miss reasoning errors - Is self-contained, unambiguous, and precisely stated - Reasoning walks through the correct derivation/justification step by step and explains why each distractor is wrong - Answer is the correct letter only, e.g. "(B)" IMPORTANT: generate a (question, reasoning, answer) triplet; wrap them in the following special tokens: <question> Insert your language question stem followed by the four answer choices (A)â(D). </question> <reasoning> Insert the step-by-step reasoning in language, including why each distractor is incorrect. </reasoning> <answer> Insert only the correct letter, e.g. "(A)". </answer> Figure 14: Prompt used during HOTFIXR training and data generation, for the Nemotron STEM subset. As input, the prompt takes a question, answer, and reasoning from the Nemotron STEM training dataset, and a language that alternates between English, French, Spanish, Arabic, Portuguese, or Italian. Prompt for HOTFIXR training on Nemotron CHAT In language, generate a NEW, ORIGINAL instruction-following task that is AS COMPLEX as the reference below. Do NOT copy, paraphrase, or reuse it in any way. Reference (complexity calibration only â do not reproduce): <question> "question" </question> <reasoning> reasoning </reasoning> <answer> answer </answer> Requirements for your generated task: - The question, reasoning, and answer should all be in language. - Instruction imposes at least as many explicit constraints as the reference (e.g. format, length, style, content restrictions, conditional logic) - Constraints are specific and verifiable â a reader can check whether the response satisfies each one - Draws from: open-ended knowledge tasks (Alpaca-style), conversational requests (LMArena-style), or format-constrained tasks (IFEval-style) - Instruction is self-contained and unambiguous - Reasoning walks through how each constraint is satisfied, step by step - Answer is a complete response that fully obeys every constraint in the instruction IMPORTANT: generate a (question, answer) pair; wrap your question and answer in the following special tokens: <question> Insert your language instruction here. </question> <reasoning> Insert the step-by-step reasoning in language. </reasoning> <answer> Insert the complete response in language that satisfies all constraints. </answer> Figure 15: Prompt used during HOTFIXR training and data generation, for the Nemotron CHAT subset. As input, the prompt takes a question, answer, and reasoning from the Nemotron CHAT training dataset, and a language that alternates between English, French, Spanish, Arabic, Portuguese, or Italian. Appendix F LLM Usage In writing this paper, the only main use of LLMs was for creating figures, specifically Claude Opus 4.8. After providing the experimental results from our results, the authors prompted it to write code using matplotlib to create the figures. Other LLM usages were very minimal: only enhancing the writing and small code suggestions. We did not use LLMs to write code files, perform literature reviews, or describe our methodology/evaluation. All LLM use was reviewed heavily by the authors.