Paper deep dive
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
Yassine Turki, Vinko SabolÄec, Bettina Messmer, Martin Jaggi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/26/2026, 6:06:22 PM
Summary
The paper investigates cross-lingual quality classifiers for multilingual pretraining data selection. It explores whether quality markers in embedding spaces (using XLM-RoBERTa) show cross-lingual consistency, allowing high-resource languages to support low-resource ones. The authors introduce the MKC-e (Extended) dataset and a Q3 (third-quartile) sampling strategy to refine decision boundaries. Results show that massive multilingual pooling (ML) outperforms monolingual baselines (HQ) in rank stability and aggregate accuracy for languages like Chinese and Spanish. Notably, a classifier trained on the Nordic language family outperformed a native French HQ baseline on French data, demonstrating that quality signals can transfer across typologically distant language families.
Entities (7)
Relation Signals (4)
ML Classifier â outperforms â Monolingual HQ Classifier
confidence 100% · The ML strategy achieves the highest aggregate accuracy and best average rank, supporting the hypothesis that quality features are transferable.
Q3 Sampling â refines â Decision Boundary
confidence 100% · We introduce a third-quartile (Q3) sampling strategy that... sharpens decision boundaries
Nordic Family â transfersqualityto â French
confidence 95% · the Nordic classifier is the one with the highest mean normalized accuracy, which even outperformed the HQ baseline trained on French.
XLM-RoBERTa â generatesembeddingsfor â MKC-e
confidence 90% · all samples are processed into 768-dimensional embeddings using the XLM-RoBERTa encoder
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As Large Language Models (LLMs) scale, data curation has shifted from maximizing volume to optimizing the signal-to-noise ratio by performing quality filtering. However, for many languages, native high quality data is insufficient to train robust quality classifiers. This work investigates the idea that quality markers in embedding space may show cross-lingual consistency, which would allow high-resource languages to subsidize the filtering of low-resource ones. We evaluate various filtering strategies, including cross-lingual transfer, third quartile sampling (Q3), and retention rate tuning. Our results demonstrate that massive multilingual pooling frequently outperforms monolingual baselines in both rank stability and aggregate accuracy for a 1B model trained on 103B tokens, delivering gains for high resource languages (1.2% increase in aggregate normalized accuracy for French) and matching or exceeding monolingual baselines for low-resource languages. However, we find that scale alone does not guarantee stability. Furthermore, for high-resource languages like French, we show that refining the decision boundary through third quartile sampling (Q3) or tuning the retention rate is necessary to fully leverage the multilingual signal.
Tags
Links
- Source: https://arxiv.org/abs/2604.20549v1
- Canonical: https://arxiv.org/abs/2604.20549v1
Trouble viewing inline? Open PDF directly â
Full Text
105,095 characters extracted from source content.
Expand or collapse full text
Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. TOWARD CROSS-LINGUAL QUALITY CLASSIFIERS FOR MULTILINGUAL PRETRAINING DATA SELECTION Yassine Turki, Vinko Sabol Ë cec, Bettina Messmer, & Martin Jaggi Machine Learning Optimization Lab Ecole Polytechnique FĂ©dĂ©rale de Lausanne (EPFL) Lausanne, Switzerland yassine.turki,vinko.sabolcec,bettina.messmer,martin.jaggi@epfl.ch ABSTRACT As Large Language Models (LLMs) scale, data curation has shifted from maxi- mizing volume to optimizing the signal-to-noise ratio by performing quality filter- ing. However, for many languages, native high-quality data is insufficient to train robust quality classifiers. This work investigates the idea that quality markers in embedding space may show cross-lingual consistency, which would allow high- resource languages to subsidize the filtering of low-resource ones. We evaluate various filtering strategies, including cross-lingual transfer, third quartile sampling (Q3), and retention rate tuning. Our results demonstrate that massive multilin- gual pooling frequently outperforms monolingual baselines in both rank stability and aggregate accuracy for a 1B model trained on 103B tokens, delivering gains for high resource languages (1.2% increase in aggregate normalized accuracy for French) and matching or exceeding monolingual baselines for low-resource lan- guages. However, we find that scale alone does not guarantee stability. Further- more, for high-resource languages like French, we show that refining the decision boundary through third quartile sampling (Q3) or tuning the retention rate is nec- essary to fully leverage the multilingual signal. 1INTRODUCTION As machine learning architectures and optimization strategies have matured, the focus of the com- munity has shifted toward the foundational element of model performance: data quality. Previous work has demonstrated that âmoreâ is not synonymous with âbetterâ (Raffel et al., 2020; Penedo et al., 2023). Building on this insight, recent efforts have shown that aggressively curating high- quality subsets from web-scale corpora can match or exceed the performance of models trained on far larger datasets (Li et al., 2025; Penedo et al., 2024; Messmer et al., 2025). Historically, data curation relied on heuristic-based filtering (Nguyen et al., 2023; Raffel et al., 2023; Rae et al., 2022; Penedo et al., 2023). However, the emergence of model-based classifiers has enabled a more nuanced selection of semantically dense content. FineWeb-Edu (Penedo et al., 2024) reached the performance of LLMs trained on 350B tokens of raw English data using only 10% of tokens through LLM-based filtering. Building upon this, FineWeb2-HQ (Messmer et al., 2025) extended model-based filtering to multilingual data, training classifiers for 20 languages and demonstrating that 15% of tokens could match the performance of models trained on the full FineWeb2 (Penedo et al., 2025) dataset. While model-based filtering has proven effective for high-resource languages, the multilingual do- main faces a critical challenge: many languages lack sufficient native high-quality data to train effective standalone classifiers. This work investigates whether quality classifiers can generalize across languages by exploiting shared semantic structures in multilingual embedding spaces, which would enable high-resource languages to effectively support filtering for low-resource ones. We hypothesize that quality is a measurable density of information and logical coherence rather than subjective preference. We assume high-quality text may be distinguished in the embedding space by grammatical coherence, where formal structures create distinct activation patterns; lexical density, characterized by specialized vocabulary rather than boilerplate; and information density, derived 1 arXiv:2604.20549v1 [cs.CL] 22 Apr 2026 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. from high logical and factual coherence. If this hypothesis holds, a classifier C trained to recog- nize quality patterns in French should transfer to typologically distant languages like Chinese, either because they share an underlying quality manifold, or because structural and formatting regulari- ties in the positive anchor datasets (such as Wikipedia markup or instruction-tuning templates) are themselves cross-lingual proxies for quality. Our primary contributions are as follows: âą Multilingual Synergy: We demonstrate that massive multilingual pooling frequently out- performs monolingual baselines in both rank stability and aggregate accuracy, delivering gains for high-resource languages and beating state-of-the-art baselines for low-resource ones. âą Empirical Validation of Cross-Family Transfer: We demonstrate that classifiers trained on one language family can effectively curate high-quality tokens in typologically distant languages. âą Q3 Sampling Strategy: We introduce a third-quartile (Q3) sampling strategy that trains against fluent but low-utility text. We find this refinement sharpens decision boundaries and improves performance in high-resource languages such as French and Spanish. 2RELATED WORK Evolution of Web Data Curation. Early large-scale datasets relied primarily on heuristic-based filtering of Common Crawl. Wenzek et al. (2019) introduced CCNet, which utilized FastText (Joulin et al., 2016) for language identification and perplexity-based scoring. This approach was refined by Raffel et al. (2020) with C4 and Penedo et al. (2023) with RefinedWeb, the latter demonstrating that aggressive deduplication and string-matching heuristics could allow web data to match curated dataset performance. However, Penedo et al. (2024) showed that simple heuristics fail to capture the nuanced educational value required for complex reasoning tasks. Cross-lingual Representation Learning Modern NLP has moved beyond language-specific models toward unified multilingual encoders. Models such as XLM-RoBERTa (Conneau et al., 2020) utilize a Transformer architecture trained on a masked language modeling objective across 100+ languages simultaneously. The core power of these models lies in their ability to align seman- tic clusters across languages within a shared embedding space. During pre-training, the model learns that semantically equivalent words occupy similar topological positions regardless of surface form. For example, the English word âScienceâ and the German word âWissenschaftâ are positioned near each other in the 768-dimensional embedding space, as both relate similarly to concepts like âlogicâ and âfact.â This alignment enables classification based on geometric position in latent space rather than surface-level syntax. Model-Based Quality Filtering. A paradigm shift occurred with model-based classifiers. FineWeb-Edu (Penedo et al., 2024) used Llama-3 (Grattafiori et al., 2024) as a judge to score ed- ucational quality and create a knowledge-rich dataset. To scale this approach, Li et al. (2025) and Messmer et al. (2025) used lightweight classifiers based on FastText (Joulin et al., 2016) or MLPs trained on embeddings of an XLM-RoBERTa model. Our work builds directly upon the FineWeb2- HQ pipeline by Messmer et al. (2025). Multilingual Scaling. While English-centric curation is well-established, multilingual curation presents unique challenges. Nguyen et al. (2023), Kudugunta et al. (2023) and Penedo et al. (2025) expanded web-scale cleaning to hundreds of languages using perplexity and basic heuristics. FineWeb2-HQ (Messmer et al., 2025) advanced this by applying model-based filtering to 20+ lan- guages. Despite these advances, two areas remain underexplored. First, while multilingual encoders are widely used, the degree to which one classifier can generalize across language families (e.g., from Nordic to Romance) has not been systematically characterized. Second, prior work primar- ily uses random negative sampling; the potential of smarter sampling techniques to refine decision boundaries remains largely uninvestigated. Our work addresses both gaps. 2 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. 3METHODS 3.1CLASSIFIER DATASETS To train a robust multilingual quality classifier, we curate a diverse set of high-quality âposi- tiveâ samples and contrast them against a baseline of general web data. Our strategy extends the FineWeb2-HQ framework by expanding their high-quality anchors (MKC+) with additional instruction-based and synthetic sources to form the MKC-e (Extended) dataset. Positive Anchors (MKC+). We incorporate the original FineWeb2-HQ anchors, which prioritize structured, knowledge-dense content. These datasets are Multilingual MMLU (OpenAI, 2024), Aya Dataset and Collection (Singh et al., 2024), OpenAssistant-2 (Köpf et al., 2023) and Include-Base- 44 (Romanou et al., 2024). Extended Positives (MKC-e). To generalize the classifierâs ability to recognize natural queries and encyclopedic prose, we expand the MKC+ pool with: âą Tagengo (Devine, 2024): A multilingual chat dataset containing approximately 75,000 conversations in 74 languages between humans and GPT4 (OpenAI et al., 2024). âą MURI-IT (Wikipedia Subset) (Köksal et al., 2024): A dataset containing instruction- output pairs across 200 languages. We specifically extracted the Wikipedia subset to en- sure samples reflect factual, encyclopedic prose and are knowledge-rich. Furthermore, MURI-IT contains samples combining different languages. As our goal is to ablate clas- sifier behaviour for different languages, we decided not to include multiple languages in a given sample. âą EuroBlocks-SFT-Synthetic (Martins et al., 2025): Multilingual synthetic data for Super- vised Fine-Tuning used to train the EuroLLM 9B Instruct model. It spans 35 languages. âą WikiQA (Apertus et al., 2025): A dataset linking real-world user queries to factual Wikipedia answer sentences. With 65 languages, it provides a stronger focus on low- resource languages. Negative Anchors (FineWeb2). We use the raw FineWeb2 (Penedo et al., 2025) corpus as our source of negative samples. We assume a random sample from this web-scale crawl primarily con- tains boilerplate, informal prose, or noise. 3.2SAMPLING AND DATA PREPARATION Balancing and Preprocessing. To ensure consistency, all samples are processed into 768- dimensional embeddings using the XLM-RoBERTa encoder (Conneau et al., 2020) used in the original FineWeb2-HQ release. Consistent with prior work, we perform minimal preprocessing (concatenation of prompt/response) and remove samples with <unk> tokens. For training, we sample 100, 000 positive documents per language, an increase from the 80, 000 used in (Messmer et al., 2025) to improve representation. To counter class imbalance in low-resource lan- guages, we upsample positives by a maximum factor of 3 (analogous to (Chung et al., 2023)), ensur- ing high-resource languages do not dominate the gradient without overfitting to duplicated samples. We sample an equal number of negative documents to maintain a balanced class distribution. For our classifier, we use a simple MLP, with a single hidden layer (256 dim, ReLU, 20% dropout) and sig- moid output, trained to predict positive and negative classes based on XLM-RoBERTa embeddings in the same way as FineWeb2-HQ. Negative Sampling Strategies. We employ two distinct strategies for selecting negative samples from FineWeb2: 1. Random Sampling: The standard approach, selecting documents uniformly at random to distinguish quality content from general web noise. 3 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. 2. Q3 (Hard Negatives): To sharpen the decision boundary, we sample documents that score in the 50thâ75th percentile (the third quartile) of a preliminary classifier. These "Q3 neg- atives" typically represent fluent but low-utility text (e.g., repetitive procedural content), forcing the model to learn subtler distinctions beyond surface-level fluency. 4EVALUATION Qualitative Analysis. To understand how each classifier filters samples, we first verify that score distributions concentrate near zero with a flatter tail at higher scores. We then inspect the top and bottom 25 samples to confirm that highly-ranked samples are knowledge-rich and well-structured, while low-ranked ones are poorly written or uninformative. Using a monolingual FineWeb2-HQ classifier as baseline, we compute rank correlations (Spearman and Kendall) and analyze the top 20 samples with the largest rank increases and decreases across classifiers. This reveals the filter- ing patterns each classifier prioritizes (e.g., grammar, content quality, symbols). Selected ranking changes are shown in Appendix C. Downstream Task Evaluation. To better understand how different filtering methods would in- fluence the performance of an LLM, we conduct experiments by training a small 1B parameter Apertus (Apertus et al., 2025) architecture model on the filtered FineWeb2 by a given classifier. The model is trained on 103B tokens with a sequence length of 4096, and sees each token at most twice, except for Arabic, where for retention rates of 10% and 20%, we have replicated the filtered data 10x and 5x respectively, to account for a lower number of tokens. These tokens are directly from the filtered samples of the classifier, followed by a rehydration step as described in (Penedo et al., 2025). Technical details for the LLM training can be found in Appendix E. To evaluate the models, we use the Language Model Evaluation Harness library (Gao et al., 2024). Our main criterion is normalized accuracy, as recommended by KydlĂ Ë cek et al.. The tasks we use span multiple capabilities, including knowledge retrieval, reasoning and natural language under- standing. Full benchmark list can be found in Appendix G. The average rank across these bench- marks serves as our final measure of a filtering strategyâs robustness. We also report the mean normalized accuracy, as the average rank can be volatile for very small differences in performance. 5EXPERIMENTS AND RESULTS We conduct ablations across four languages representing diverse linguistic profiles: French (Ro- mance, high-resource), Spanish (Romance, high-resource), Arabic (Semitic, morphologically com- plex), and Chinese (Sino-Tibetan, logographic). Our experimental design addresses three core ques- tions: 1. Multilingual Synergy: Does pooling data from multiple languages improve performance compared to monolingual baselines? 2. Cross-lingual Transfer: Can classifiers trained on typologically distant languages identify quality in a target language? 3. Decision Boundary Refinement: Can the Q3 strategy improve classifier precision in high- resource settings? Baseline Configurations. We establish two primary baselines: No filtering, representing random sampling from FineWeb2 (lower bound), and HQ, a monolingual classifier trained following the FineWeb2-HQ pipeline (Messmer et al., 2025) with MKC+ anchors. Multilingual Synergy. We first investigate whether data quality is a general feature shared across samples from different languages, or if it is language-specific. Our hypothesis is that there exist shared regularities in XLM-RoBERTaâs embedding space that correlate with quality. If we train a model to detect these regularities across multiple languages, then we would obtain a general classi- fier that could recognize high-quality samples in unseen languages, and even boost the data quality for languages it was trained on. In order to test this hypothesis, we train a general multilingual classifier (denoted by ML). We use the MKC-e pool to account for the lack of samples for some 4 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 1: Comparison of models trained without filtering, with monolingual high-quality (HQ) fil- tering, and with our multilingual (ML) classifier on Chinese. The ML strategy achieves the highest aggregate accuracy and best average rank, supporting the hypothesis that quality features are trans- ferable. BenchmarkNo filteringHQML Agieval Cn0.36180.36440.3457 ARC0.28550.31450.3171 Belebele_c0.30110.32000.3222 Ceval-valid0.22880.24890.2615 Cmmlu_c0.32060.34710.3608 GMMLU_c0.28000.30750.3200 Include_c0.34680.35230.3541 MMLU_c0.27720.29400.2987 PAWS0.55200.55350.5610 Xcopa0.58600.59200.6080 XNLI0.35460.40720.4189 Xstorycloze0.65590.66250.6625 XWinograd0.68060.68250.6766 Aggregate acc_norm0.40240.41900.4236 Average rank2.851.771.31 Table 2: Comparison of models trained without filtering, with monolingual high-quality (HQ) fil- tering, and with our multilingual (ML) classifier on Spanish. The ML strategy achieves the highest aggregate accuracy and best average rank, supporting the hypothesis that quality features are trans- ferable. BenchmarkNo filteringHQML ARC-Challenge0.29910.30770.3248 Belebele_c0.34220.35330.3456 GMMLU_c0.31000.31500.3250 HellaSwag0.50060.52630.5310 Include_c0.34910.37270.3891 M_MMLU_c0.28140.29850.3050 XNLI0.46180.45380.4783 Aggregate acc_norm0.36350.37530.3855 Average rank2.862.001.14 knowledge, and also to balance the dominance of the Aya Collection dataset in terms of size and representation. The detailed counts for each language can be found in the appendix A. Tables 1 and 2 evaluate the effect of multilingual quality filtering on downstream performance for a 1B model in Chinese and Spanish. The multilingual classifier (ML) achieves the highest aggregate accuracy and lowest average rank in both languages, outperforming both No filtering and monolin- gual high-quality filtering (HQ). Improvements are observed consistently across diverse reasoning and natural language understand- ing benchmarks, such as ARC Clark et al. (2018), GMMLU Singh et al. (2025), XNLI (Conneau et al., 2018), and Include (Romanou et al., 2024). While HQ occasionally yields marginal gains on individual tasks, ML provides more stable and stronger overall performance. These results are consistent with the hypothesis that data quality corresponds to a shared structure in XLM-RoBERTaâs representation space. However, we cannot determine whether this reflects abstract semantic features or cross-lingual formatting regularities in the positive anchor datasets. By jointly modeling quality across languages, the classifier captures features that transfer effectively across linguistic boundaries, leading to improved downstream generalization. 5 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 3: Spearman and Kendall correlations between family-specific classifiers and the French HQ baseline. High correlation in distant families (e.g., Nordic) indicates a shared quality manifold, while the drop in "Romance (no French)" suggests potential syntactic interference. ExperimentSpearmanKendall Romance (spa, fra, por, ita, ron, cat) MKC+0.89280.7173 Nordic (swe, dan, nob, isl) MKC+0.88200.6990 Romance, no french (spa, por, ita, ron, cat) MKC-e0.71390.5228 Cross-lingual Transfer. We have seen that multilingual transfer is possible since a classifier trained on multiple languages performs better than its monolingual counterpart. We note that our ML classifier was trained on these languages; therefore, it already has knowledge in that language. We hypothesize that this knowledge has been augmented by the data from the other languages. Can this ML classifier actually generalize to unseen languages? Would this classifier be able to perform as well (if not better) on a language like French if it has never seen a sample in that language? To answer these questions, we investigate the performance of a classifier on French. First, we train classifiers on different language families (i.e. Romance, Nordic, Germanic, etc.) and apply them on the French split of FineWeb2 to obtain scores. Then, we compute the Spearman and Kendall corre- lations between the scores of these classifiers and the HQ baseline. We would expect to have a very strong correlation score in the Romance family (French, Spanish, Portuguese, Italian, Romanian, Catalan) and a low correlation score in the other language families, which are linguistically distant from French. We display some of these correlation scores in Table 3, and provide the full table for other language families in the Appendix C.1. We can see that intra-family transfer is strong: Romance (MKC+) Ï s = 0.89, indicating near-identical document rankings despite training on mul- tiple Romance languages. This is to be expected, as languages in the same family share the same root and possess a big overlap in vocabulary. What is more interesting is that the classifiers trained on unrelated families, for example Nordic (MKC+): Ï s = 0.88, have nearly matching Romance correlation despite limited lexical and syntactic overlap. Finally, we can see that removing French from Romance hurts more than expected. Romance without French (MKC-e) drops to Ï s = 0.71, correlating worse than some linguistically distant families like Germanic or Uralic. This suggests potential syntactic interference: when the classifier trains on closely related but distinct languages (Spanish, Italian, Portuguese), it may learn Romance-specific syntactic patterns that donât perfectly generalize to French, obscuring the underlying quality signal. To investigate this further, we plot the score distribution of the Nordic classifier in Figure 1 to give us insights on how the classifier behaves. We apply it on a held-out set of French high-quality samples, negative samples, and the general FineWeb2. As we can observe, the classifier recognizes almost perfectly the negative and positive classes. It applies a very strict threshold (90th percentile: 0.027) and assigns generally lower scores, yet still identifies high-quality content effectively. Additional experiments, including the baseline plot, can be found in the Appendix B. Additionally, we train our 1B models with the datasets from our classifiers. Despite the high rank correlation, we would expect the classifier trained on Romance to perform better than Nordic, even though both have never seen French samples in their training data. And we expect the Romance classifier to perform close to the HQ baseline, as these languages should be similar enough to French to give the classifier a good idea of what a high-quality sample should be. We show the results of the experiment in Table 4, which provides a surprising result. The Romance classifier has the same rank as the HQ baseline, however, with slightly less in normalized accuracy. Furthermore, the Nordic classifier is the one with the highest mean normalized accuracy, which even outperformed the HQ baseline trained on French. The success of the Nordic classifier provides evidence for transferable quality features across lan- guage families. Because Nordic languages (Swedish, Danish, Norwegian, Icelandic) share no lexical or syntactic overlap with French, the classifier cannot rely on surface-level patterns. These results suggest that cross-lingual transfer of quality signals is practically effective even across distant language families, though whether this reflects abstract semantic structure or shared format- 6 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. 0.00.20.40.60.81.0 Quality Score 10 2 10 1 10 0 10 1 10 2 Density Test Positive Test Negative FineWeb2 Cutoff (retention 0.10) Figure 1: French quality score distribution from a Nordic classifier. The classifier effectively sep- arates high-quality (positives) from low-quality (negatives) French text despite being trained solely on Nordic languages, suggesting the practical effectiveness of cross-lingual quality transfer. Table 4: Models trained on French data filtered by "Nordic" and "Romance (no French)" classifiers. The Nordic classifier outperforms the native French HQ baseline in aggregate accuracy, providing evidence for transferable quality signals across language families. BenchmarkNo filteringHQRomance (no fra)Nordic ARC-Challenge0.28910.30710.30540.3157 Belebele_c0.34440.35110.36890.3422 GMMLU_c0.26250.29250.27500.3075 HellaSwag0.48830.47480.49080.4673 Include_c0.38660.41530.42720.4535 M_MMLU_c0.28310.29450.29500.2936 XNLI0.47070.48550.48270.4783 XWinograd0.63860.65060.60240.6627 Aggregate acc_norm0.39540.40890.40590.4151 Average rank3.502.122.122.25 ting regularities in the embedding space remains to be characterized. However, we note that these high rank correlations may partially reflect the classifier learning shared dataset artifacts such as Wikipedia formatting or instruction-tuning templates rather than a purely abstract representation of quality. Nonetheless, capturing these cross-lingual artifacts serves as a highly effective and robust proxy for pretraining data selection. Conversely, the underperformance of Romance-without-French suggests that syntactic similarity can be a double-edged sword: the classifier may overfit to Span- ish/Italian grammatical structures that donât perfectly align with French, creating systematic biases that distant language families avoid. Refining the Decision Boundary. We have observed that training on multiple languages provides strong performance boosts, and also can generalize to unseen samples. Nevertheless, upon examin- ing the training pipeline, one could argue that using random samples from FineWeb2 might not be the best strategy for selecting negative samples. Standard negative sampling draws randomly from FineWeb2, teaching the classifier to distinguish quality content from typical web noise (advertise- ments, navigation menus, broken formatting, spam). However, as filtering becomes more precise, a subtler distinction emerges: separating high-quality content from fluent but low-utility text. The Q3 strategy addresses this by sampling negatives from the 50th-75th percentile of the base- line HQ distribution. These samples are grammatically correct and well-formatted, but lacking in information density, logical depth, or educational value. We provide an example in the Appendix D. 7 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 5: Comparison of benchmark results using standard random negatives vs. Q3 negatives (fluent but low-utility text) in French. Q3 sampling consistently yields superior aggregate performance, with the monolingual classifier (Q3) performing slightly better than the ML (Q3) classifier. BenchmarkNo filteringHQQ3MLML (Q3) ARC-Challenge0.28910.30710.31050.31650.3054 Belebele_c0.34440.35110.37110.35110.3633 GMMLU_c0.26250.29250.30250.29000.3075 HellaSwag0.48830.47480.47270.49560.4911 Include_c0.38660.41530.42720.42240.4582 M_MMLU_c0.28310.29450.29920.29270.3007 XNLI0.47070.48550.48760.48070.4767 XWinograd0.63860.65060.66270.60240.6145 Aggregate acc_norm0.39540.40890.41670.40640.4147 Average rank4.503.002.003.002.38 Table 6: Comparison of benchmark results using standard random negatives vs. Q3 negatives (fluent but low-utility text) in Spanish. Q3 sampling consistently yields superior aggregate performance. BenchmarkNo filteringHQMLML (Q3) ARC-Challenge0.29910.30770.32480.3282 Belebele_c0.34220.35330.34560.3533 GMMLU_c0.31000.31500.32500.3275 HellaSwag0.50060.52630.53100.5372 Include_c0.34910.37270.38910.4000 M_MMLU_c0.28140.29850.30500.3083 XNLI0.46180.45380.47830.4667 Aggregate acc_norm0.36350.37530.38550.3887 Average rank3.862.862.001.14 In this approach, once a classifier is obtained, we apply it on FineWeb2 and extract samples from the third quartile (i.e. 50th to 75th percentile) as our negative samples for the training of a new classifier. We conduct these experiments using the HQ baseline (Q3) and the ML classifier (ML Q3). The ML classifier uses negative samples from the third quartile of its scores. We apply this extra step to all languages that have more than 200,000 samples in FineWeb2 to ensure we will not need aggressive token replication. Tables 5 and 6 evaluate the impact of refining the negative sampling strategy using third-quartile bootstrapping (Q3). Compared to random negatives from FineWeb2, harder negatives consistently improve performance, confirming that sharper decision boundaries emerge when the classifier is trained on more ambiguous examples. For French, the monolingual Q3 classifier achieves the strongest aggregate performance, while the multilingual bootstrapped model (ML Q3) remains highly competitive, with only a marginal drop in normalized accuracy. In contrast, for Spanish, ML Q3 yields the best overall performance across benchmarks, surpassing both standard multilingual filtering and monolingual HQ. Taken together, these results indicate that bootstrapped negative sampling systematically strengthens quality discrimination. While language-specific refinement can yield peak in-language performance, the multilingual Q3 model matches or exceeds these gains in some languages and, critically, pre- serves cross-lingual transfer without retraining. This further supports the existence of shared in representation space that can be progressively refined through Q3 sampling. Tuning the Retention Rate. While FineWeb2-HQ was derived using a retention rate of 10% for high-resource languages, we argue that this hyperparameter should be tuned for each language sep- arately. The rate for which we filter samples for a language can have a big impact on the filtered data. 8 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 7: Comparison between 10% and 15% retention using the ML classifier. Increasing the reten- tion rate to 15% improves accuracy across most benchmarks, suggesting the default 10% threshold is overly aggressive when using our classifier on French. BenchmarkMLML (15%) ARC-Challenge0.31650.3139 Belebele_c0.35110.3611 GMMLU_c0.29000.2875 HellaSwag0.49560.5135 Include_c0.42240.4726 M_MMLU_c0.29270.2991 XNLI0.48070.4807 XWinograd0.60240.6386 Aggregate acc_norm0.40640.4209 Average rank1.621.25 Table 8: Comparison between 10% and 15% retention using the ML classifier on Spanish. Similar to French, the 15% retention rate yields a higher aggregate accuracy and better average rank when using our ML classifier. BenchmarkMLML (15%) ARC-Challenge0.32480.3436 Belebele_c0.34560.3500 GMMLU_c0.32500.3225 HellaSwag0.53100.5440 Include_c0.38910.4036 M_MMLU_c0.30500.3080 XNLI0.47830.4562 Aggregate acc_norm0.38550.3897 Average rank1.711.29 Tables 7, 8, and 9 analyze the effect of tuning the retention rate used during multilingual and monolingual filtering. Increasing the retention rate consistently improves aggregate performance in French and Spanish, with ML at 15% outperforming the default 10% across most benchmarks. Gains are particularly pronounced on knowledge-intensive tasks such as Include (Romanou et al., 2024) and HellaSwag (Zellers et al., 2019), indicating that overly aggressive filtering can discard useful high-quality content. In Arabic, instead of using 56% like for FineWeb2-HQ, we try a retention rate of 10% and 20%, resulting in a replication of tokens of 10x and 5x, respectively. We see that the standard retention of 56% yields the best overall performance when combined with multilingual filtering. This suggests that optimal retention rates are language-dependent, and that aggressive repetition of the same high- quality data does not necessarily lead to the best performance. We further note that the ML gain over HQ for Arabic ( 0.5%) is comparable to the seed variance reported in Appendix H ( 0.3%), and should therefore be interpreted with caution pending multi-seed validation. Synthesis: When Does Multilingual Pooling Help? Across our four languages, we observe a pattern: the key insight is not that multilingual pooling always dominates, but that it provides a re- liable baseline across languages, while language-specific optimization (negative sampling strategy, retention rate, anchor curation) can yield comparable or superior results when tuned appropriately. These findings provide evidence that the classifier learns language-agnostic markers in the embed- ding space by training on multiple languages. Note on Stochasticity: While multilingual pooling improves overall rank stability, our ablations show that downstream aggregate accuracy remains sensitive to the classifierâs initial sampling seed (shifting by 0.3% in Arabic and 0.8% in French). We provide a more detailed analysis of these seed variances in Appendix H. 9 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 9: Evaluation of standard (56%) vs. aggressive (10%, 20%) filtering with replication. Unlike high-resource languages, reducing Arabic retention to increase token replication degrades perfor- mance. BenchmarkHQMLHQ (10%)HQ (20%)ML (10%)ML (20%) ARC-Easy0.28980.28550.27200.27710.27410.2762 AlGhafa PIQA-MT0.51450.51120.50190.50900.49480.5128 AlGhafa RACE0.27750.28830.27470.28340.27650.2786 AlGhafa SciQ0.44320.45030.45030.41710.45730.4563 ARC-Challenge0.26600.27970.26690.27720.27120.2626 Belebele_c0.32330.31220.33780.29440.32780.3256 GMMLU_c0.25750.27000.24750.25500.27250.2525 HellaSwag0.38730.39250.35920.37280.37470.3857 Include_c0.26810.30430.27170.27720.28440.2790 M_MMLU_c0.26140.26460.26520.26740.26180.2634 AlGhafa PIQA0.60880.61100.59580.60230.60390.6143 XNLI0.33090.33490.33250.33330.33250.3357 XStoryCloze0.59230.58110.58770.58170.57840.5804 Aggregate acc_norm0.37080.37580.36640.36520.37000.3710 Average rank3.622.314.313.693.693.23 Table 10: Aggregated macro and micro metrics for all strategies using a 10% retention rate. The ML (Q3) strategy achieves the best overall performance in both rank and normalized accuracy, highlighting the robustness of combining both our methods to obtain a robust general multilingual classifier. MethodMacro RankMicro RankMacro AccMicro Acc ML (Q3)1.69831.78570.40820.4113 ML1.98081.92860.40520.4092 HQ2.45562.42860.40110.4052 No filtering3.72483.71430.38710.3907 6CONCLUSION This work investigates whether quality classifiers can generalize across languages by exploiting shared semantic structures in multilingual embedding spaces. Through systematic evaluation across French, Spanish, Arabic, and Chinese, we demonstrate that classifiers trained on typologically dis- tant language families can effectively filter quality content in unrelated languages. Across the four tested languages, multilingual classifiers improved over monolingual baselines, with particularly strong gains in Spanish (3.0 ranks) and Arabic (2.31 ranks). We also introduce the Q3 sampling strategy, which refines decision boundaries by training against fluent but low-utility text rather than random negatives, offering a complementary approach to multilingual pooling for high-resource languages. No single strategy dominates. For French, Q3 negatives, ML (15%), and ML (Q3) all achieve comparable results, suggesting practitioners should select based on computational con- straints and available data. The 10% threshold from prior work may be overly aggressive when using our ML classifier; increasing to 15% improved French performance substantially (40.64% to 42.09%), which indicates that this hyperparameter needs language-specific tuning. However, when comparing all approaches (Table 10), we find that the ML (Q3) strategy is the one dominat- ing in terms of performance, which means combining our methods leads to better results. Overall, our findings suggest that multilingual pooling can help democratize high-quality data curation for underrepresented languages while boosting performance on high-resource languages, though the effectiveness varies by language family and resource level. 10 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. REFERENCES Ebtesam Almazrouei, Ruxandra Cojocaru, Michele Baldo, Quentin Malartic, Hamza Alobeidli, Daniele Mazzotta, Guilherme Penedo, Giulia Campesan, Mugariya Farooq, Maitha Alhammadi, Julien Launay, and Badreddine Noune. AlGhafa evaluation benchmark for Arabic language models. In Hassan Sawaf, Samhaa El-Beltagy, Wajdi Zaghouani, Walid Magdy, Ahmed Ab- delali, Nadi Tomeh, Ibrahim Abu Farha, Nizar Habash, Salam Khalifa, Amr Keleg, Hatem Haddad, Imed Zitouni, Khalil Mrini, and Rawan Almatham (eds.), Proceedings of ArabicNLP 2023, p. 244â275, Singapore (Hybrid), December 2023. Association for Computational Linguis- tics. doi: 10.18653/v1/2023.arabicnlp-1.21. URL https://aclanthology.org/2023. arabicnlp-1.21. Project Apertus, Alejandro HernĂĄndez-Cano, Alexander Hagele, Allen Hao Huang, Angelika Ro- manou, Antoni-Joan Solergibert, Barna Pasztor, Bettina Messmer, Dhia Garbaya, Eduard Frank Ë Durech, Ido Hakimi, Juan GarcĂa Giraldo, Mete Ismayilzada, Negar Foroutan, Skander Moalla, Tiancheng Chen, Vinko Sabol Ë cec, Yixuan Xu, Michael Aerni, Badr AlKhamissi, InĂ©s Altemir Mariñas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clement Charmillot, Jonathan Coles, Jan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzen- haliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, MarĂa Grandury, Diba Hashemi, Alexander Hoyle, Jiaming Jiang, Mark Klein, Andrei Kucharavy, Anastasiia Kucherenko, Frederike LĂŒbeck, Roman Machacek, Theofilos Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montar- iol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, LĂ©o Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaus- tubh Ponkshe, Nathan Ranchin, Javi Rando, Mathieu Sauser, Jakhongir Saydaliev, Muham- mad Ali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush Kumar Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Hao Zhao, Alexander Ilic, Ana Klimovic, Andreas Krause, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian TramĂšr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, and Imanol Schlag. Apertus: Democratizing open and compliant llms for global language environments, 2025. URL https://arxiv.org/abs/2509.14233. Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 749â775. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024. acl-long.44. URL http://dx.doi.org/10.18653/v1/2024.acl-long.44. Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. Piqa: Reasoning about physical commonsense in natural language, 2019. URL https://arxiv.org/abs/1911. 11641. Hyung Won Chung, Noah Constant, Xavier Garcia, Adam Roberts, Yi Tay, Sharan Narang, and Orhan Firat. Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining, 2023. URL https://arxiv.org/abs/2304.09151. Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018. URL https://arxiv.org/abs/1803.05457. Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. Xnli: Evaluating cross-lingual sentence representations, 2018. URL https://arxiv.org/abs/1809.05053. Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco GuzmĂĄn, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Un- supervised cross-lingual representation learning at scale, 2020. URL https://arxiv.org/ abs/1911.02116. 11 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A Rossi, and Thien Huu Nguyen. Okapi: Instruction-tuned large language models in multiple lan- guages with reinforcement learning from human feedback. arXiv e-prints, p. arXivâ2307, 2023. Peter Devine. Tagengo: A multilingual chat dataset, 2024. URL https://arxiv.org/abs/ 2405.12612. Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Fos- ter, Laurence Golding, Jeffrey Hsu, Alain Le Noacâh, Haonan Li, Kyle McDonell, Niklas Muen- nighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. The language model evaluation harness, 07 2024. URL https://zenodo.org/records/12608602. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Ko- renev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, Danny Wyatt, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Francisco GuzmĂĄn, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Govind That- tai, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Kore- vaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jack Zhang, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Ma- hadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jong- soo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Karthik Prasad, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Kushal Lakhotia, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Maria Tsimpoukelli, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Ku- mar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoy- chev, Niladri Chatterji, Ning Zhang, Olivier Duchenne, Onur Ăelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ra- mon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohan Maheswari, Ro- hit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mi- haylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, VĂtor Albiero, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaofang Wang, Xiaoqing Ellen Tan, Xide Xia, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Gold- schlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aayushi Srivastava, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Amos Teo, Anam Yunus, Andrei Lupu, Andres Alvarado, An- drew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, An- nie Dong, Annie Franco, Anuj Goyal, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leon- hardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu 12 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Mon- talvo, Carl Parker, Carly Burton, Catalina Mejia, Ce Liu, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Cynthia Gao, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Eric-Tuan Le, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smoth- ers, Fei Sun, Felix Kreuk, Feng Tian, Filippos Kokkinos, Firat Ozgenel, Francesco Caggioni, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hakan Inan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harri- son Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Hongyuan Zhan, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Ilias Leontiadis, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Janice Lam, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jen- nifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Jun- jie Wang, Kai Wu, Kam Hou U, Karan Saxena, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kiran Jagadeesh, Kun Huang, Kunal Chawla, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Miao Liu, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikhil Mehta, Nikolay Pavlovich Laptev, Ning Dong, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Ro- driguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Rangaprabhu Parthasarathy, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Russ Howes, Ruty Rinott, Sachin Mehta, Sachin Siby, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Mahajan, Saurabh Verma, Seiji Yamamoto, Sharadh Ra- maswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shishir Patil, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satter- field, Sudarshan Govindaprasad, Sumit Gupta, Summer Deng, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Koehler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Ku- mar, Vishal Mangla, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiao- jian Wu, Xiaolan Wang, Xilun Wu, Xinbo Gao, Yaniv Kleinman, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yu Zhao, Yuchen Hao, Yundi Qian, Yunlu Li, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhao- duo Wen, Zhenyu Yang, Zhiwei Zhao, and Zhiyu Ma. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783. Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. 2021. Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. arXiv preprint arXiv:2305.08322, 2023. Alexandra Institute. m_mmlu (revision 18e6c8e), 2025. URL https://huggingface.co/ datasets/alexandrainst/m_mmlu. Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. Bag of tricks for efficient text classification, 2016. URL https://arxiv.org/abs/1607.01759. 13 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A. Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. Madlad-400: A multilingual and document-level large audited dataset, 2023. URL https: //arxiv.org/abs/2309.04662. Hynek KydlĂ Ë cek, Guilherme Penedo, ClĂ©mentine Fourier, Nathan Habib, and Thomas Wolf. Finetasks: Finding signal in a haystack of 200+ multilingual tasks.URL https:// huggingface.co/spaces/HuggingFaceFW/blogpost-fine-tasks. Abdullatif Köksal, Marion Thaler, Ayyoob Imani, Ahmet ĂstĂŒn, Anna Korhonen, and Hinrich SchĂŒtze. Muri: High-quality instruction tuning datasets for low-resource languages via reverse instructions, 2024. URL https://arxiv.org/abs/2409.12958. Andreas Köpf, Yannic Kilcher, Dimitri von RĂŒtte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, RichĂĄrd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. Openassistant conversations â democratizing large language model align- ment, 2023. URL https://arxiv.org/abs/2304.07327. Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. Cmmlu: Measuring massive multitask language understanding in chinese, 2024. URL https://arxiv.org/abs/2306.09212. Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Rein- hard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Al- balak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, Cheng-Yu Hsieh, Dhruba Ghosh, Josh Gardner, Maciej Kilian, Hanlin Zhang, Rulin Shao, Sarah Pratt, Sunny Sanyal, Gabriel Il- harco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang, Khyathi Chandu, Thao Nguyen, Igor Vasiljevic, Sham Kakade, Shuran Song, Sujay Sanghavi, Fartash Faghri, Se- woong Oh, Luke Zettlemoyer, Kyle Lo, Alaaeldin El-Nouby, Hadi Pouransari, Alexander Toshev, Stephanie Wang, Dirk Groeneveld, Luca Soldaini, Pang Wei Koh, Jenia Jitsev, Thomas Kol- lar, Alexandros G. Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar. Datacomp-lm: In search of the next generation of training sets for language models, 2025. URL https://arxiv.org/abs/2406.11794. Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian OâHoro, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, Veselin Stoyanov, and Xian Li. Few-shot learning with multilingual language models. CoRR, abs/2112.10668, 2021. URL https://arxiv.org/abs/2112.10668. Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. Program induction by rationale gen- eration: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 158â167, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1015. URL https://aclanthology.org/P17-1015. Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. Logiqa: A challenge dataset for machine reading comprehension with logical reasoning. In International Joint Conference on Artificial Intelligence, 2020. Pedro Henrique Martins, JoĂŁo Alves, Patrick Fernandes, , Nuno M. Guerreiro, Ricardo Rei, Amin Farajian, Mateusz Klimaszewski, Duarte M. Alves, JosĂ© Pombal, Manuel Faysse, Pierre Colombo, François Yvon, Barry Haddow, JosĂ© G. C. de Souza, Alexandra Birch, and AndrĂ© F. T. Martins. Eurollm-9b: Technical report, 2025. Bettina Messmer, Vinko Sabol Ë cec, and Martin Jaggi. Enhancing multilingual llm pretraining with model-based data selection, 2025. URL https://arxiv.org/abs/2502.10361. Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir 14 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. Crosslingual generalization through multitask finetuning, 2022. Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. Culturax: A cleaned, enormous, and multilingual dataset for large language models in 167 languages, 2023. URL https://arxiv.org/abs/2309. 09400. OALL. Alghafa arabic llm benchmark translated. https://huggingface.co/datasets/ OALL/AlGhafa-Arabic-LLM-Benchmark-Translated, 2023. OpenAI. Mmmlu dataset, 2024. URL https://huggingface.co/datasets/openai/ MMMLU. Hugging Face. OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Moham- mad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brock- man, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, SimĂłn Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gib- son, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hal- lacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Ćukasz Kaiser, Ali Ka- mali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Ćukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David MĂ©ly, Ashvin Nair, Reiichiro Nakano, Ra- jeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen OâKeefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Sel- sam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Pre- ston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe CerĂłn Uribe, Andrea Vallone, Arun Vi- jayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Work- man, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao 15 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. Gpt-4 technical report, 2024. URL https://arxiv.org/abs/2303.08774. Matteo Pagliardini, Pierre Ablin, and David Grangier. The ademamix optimizer: Better, faster, older, 2024. URL https://arxiv.org/abs/2409.03137. Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only, 2023. URL https://arxiv.org/abs/2306.01116. Guilherme Penedo, Hynek KydlĂ Ë cek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. The fineweb datasets: Decanting the web for the finest text data at scale, 2024. URL https://arxiv.org/abs/2406.17557. Guilherme Penedo, Hynek KydlĂ Ë cek, Vinko Sabol Ë cec, Bettina Messmer, Negar Foroutan, Amir Hos- sein Kargaran, Colin Raffel, Martin Jaggi, Leandro Von Werra, and Thomas Wolf. Fineweb2: One pipeline to scale them all â adapting pre-training data processing to every language, 2025. URL https://arxiv.org/abs/2506.20920. Edoardo M. Ponti, Goran GlavaĆĄ, Olga Majewska, Qianchu Liu, Ivan Vuli Ì c, and Anna Korhonen. XCOPA: A multilingual dataset for causal commonsense reasoning. arXiv preprint, 2020. URL https://ducdauge.github.io/files/xcopa.pdf. Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kun- coro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Men- sch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson dâAutume, Yu- jia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Au- relia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. Scaling language models: Methods, analysis & insights from training go- pher, 2022. URL https://arxiv.org/abs/2112.11446. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to- text transformer. Journal of Machine Learning Research, 21(140):1â67, 2020. URL http: //jmlr.org/papers/v21/20-074.html. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer, 2023. URL https://arxiv.org/abs/1910.10683. Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon.Choice of plausible al- ternatives: An evaluation of commonsense causal reasoning. In 2011 AAAI Spring Sympo- sium Series, 2011. URL https://people.ict.usc.edu/~gordon/publications/ AAAI-SPRING11A.PDF. Angelika Romanou, Negar Foroutan, Anna Sotnikova, Zeming Chen, Sree Harsha Nelaturu, Shiv- alika Singh, Rishabh Maheshwary, Micol Altomare, Mohamed A Haggag, Alfonso Amayuelas, et al. Include: Evaluating multilingual language understanding with regional knowledge. arXiv preprint arXiv:2411.19799, 2024. Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model par- allelism, 2020. URL https://arxiv.org/abs/1909.08053. 16 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzemi Ì nski, Hakimeh Fadaei, Irem ErgĂŒn, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Minh Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet ĂstĂŒn, Marzieh Fadaee, and Sara Hooker. Aya dataset: An open-access collection for multilingual instruction tuning, 2024. Shivalika Singh, Angelika Romanou, ClĂ©mentine Fourrier, David I. Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Ray- mond Ng, Shayne Longpre, Wei-Yin Ko, Sebastian Ruder, Madeline Smith, Antoine Bosselut, Al- ice Oh, Andre F. T. Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, and Sara Hooker. Global mmlu: Understanding and addressing cultural and linguis- tic biases in multilingual evaluation, 2025. URL https://arxiv.org/abs/2412.03304. Qwen Team. Qwen3 technical report, 2025. URL https://arxiv.org/abs/2505.09388. Alexey Tikhonov and Max Ryabinin. Itâs all in the heads: Using attention heads as a baseline for cross-lingual transfer in commonsense reasoning, 2021. Siyuan Wang, Zhongkun Liu, Wanjun Zhong, Ming Zhou, Zhongyu Wei, Zhumin Chen, and Nan Duan. From lsat: The progress and challenges of complex reasoning. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:2201â2216, 2021. Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco GuzmĂĄn, Armand Joulin, and Edouard Grave. Ccnet: Extracting high quality monolingual datasets from web crawl data, 2019. URL https://arxiv.org/abs/1911.00359. Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. Paws-x: A cross-lingual adversarial dataset for paraphrase identification, 2019. URL https://arxiv.org/abs/1908.11828. Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a ma- chine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019. Haoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang, Zhiyuan Liu, and Maosong Sun. Jec-qa: A legal-domain question answering dataset. In Proceedings of AAAI, 2020. Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. Agieval: A human-centric benchmark for evaluating foundation models, 2023. ADATASET COUNTS Table 11: Dataset Counts by Language (sorted by total count, descending). Note: OA = openassistant2, MMLU = openai_mmlu, Inc = include, Tag = tagengo, Euro = euroblocks, Wiki = muri_wikipedia, AH = aya_human, AC = aya_collection, WQA = wikiqa. LanguageOAMMLUIncTagEuroWikiAHACWQATotal eng_Latn61,27899,842015,771151,1356,6573,94414,693,823015,032,450 jpn_Jpan75614,0425012,5214,7356,9716,2596,218,45906,254,244 arb_Arab7614,04255278908,5084,9955,857,45817,7065,904,126 tha_Thai1,5030013307,1877245,338,23205,347,779 deu_Latn5,79714,0421395,73914,0817,0092414,689,98904,737,037 fra_Latn3,68614,0424195,36914,8826,9881,4224,285,09404,331,902 tel_Telu00548007,7658,4394,058,53504,075,287 rus_Cyrl13,33605528,0564,7277,0424234,005,16604,039,302 fin_Latn1380551921,0226,9707423,939,94116,3833,965,839 spa_Latn26,81114,0425508,31817,4287,0123,8543,872,86403,950,879 ita_Latn89914,0425487,06315,9637,1007383,890,85203,937,205 urd_Arab00352307,8936543,876,19703,885,099 pol_Latn43105481,0905,3587,0131,4833,841,45116,9643,874,338 por_Latn2,58114,04255112,56413,9667,3678,9973,786,06203,846,130 hin_Deva014,042547207,9827,2171,1533,772,86403,803,825 Continued on next page 17 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 11 â Continued from previous page LanguageOAMMLUIncTagEuroWikiAHACWQATotal fas_Arab0054818407,5041,5783,785,2505,6153,800,679 nld_Latn7205513837,6836,8401,7333,736,93818,7233,772,923 ukr_Cyrl82105503235,19105223,729,74812,9793,750,134 ces_Latn12001794,1056,79303,719,21411,2033,741,506 heb_Hebr24055012006,91603,658,06617,2293,682,905 cmn_Hani014,0425455,33827,5079,3684,9093,606,93503,668,644 hun_Latn11305502144,0107,096983,637,91114,3203,664,312 swe_Latn1002566,4766,5241,3103,632,62214,6503,661,839 tur_Latn37054840607,0844,0463,628,10918,4223,658,652 kor_Hang2014,0425001,6092,9057,3433613,605,89417,6163,650,290 cat_Latn1,19400732237,20903,625,53715,4383,649,674 srp_Cyrl00550607,5291523,636,57303,644,810 ben_Beng114,042548007,6091,5343,601,2878,5323,633,553 vie_Latn203055042907,0408,6763,613,27003,630,168 ind_Latn1214,042550240007863,610,07803,625,708 ron_Latn000713,9937,17003,602,21211,1883,624,634 bul_Cyrl00550562107,22103,602,87811,2073,622,122 hau_Latn000008,0453,5123,608,88303,620,440 tam_Taml00550507,76614,1333,596,70703,619,161 slk_Latn000179907,06103,594,20315,0363,617,307 slv_Latn000102016,87303,593,62616,1253,616,835 dan_Latn40006746,348973,601,9008,2123,616,668 ell_Grek005523085827,4956233,606,2498273,616,636 yor_Latn014,042000011,7583,587,23303,613,033 zsm_Latn00501508,06010,0733,593,31303,611,952 bel_Cyrl00550207,49903,589,91212,8683,610,831 sin_Sinh000507,29014,5243,587,05103,608,870 plt_Latn000006,89514,5973,586,96203,608,454 ibo_Latn000008,7671,5343,597,29203,607,593 swh_Latn014,0420007,5183663,580,06103,601,987 ary_Arab0000008,0903,591,62103,599,711 glg_Latn000006,90603,572,36519,4813,598,752 lit_Latn005341107,2399163,573,28116,3543,598,335 amh_Ethi000307,1321,2073,589,99303,598,335 nob_Latn000265346,87003,572,36517,7423,597,537 eus_Latn2570500607,0759393,573,30415,0693,597,150 ltz_Latn000106,99903,572,36517,6893,597,054 som_Latn000007,0367,7043,582,11103,596,851 ekk_Latn00224181797,02803,572,36516,8243,596,638 isl_Latn0005187,37203,572,36515,4263,595,186 gla_Latn000007,65503,572,36515,0533,595,073 mkd_Cyrl00551507,54803,572,36514,0663,594,535 lvs_Latn000221767,51503,572,36514,3113,594,389 als_Latn00551508,1521203,572,48512,7053,594,018 ydd_Hebr000207,06203,572,36513,4173,592,846 mlt_Latn000004,77103,572,36515,3093,592,445 mar_Deva000107,7433,5453,579,22803,590,517 cym_Latn000006,94403,572,36511,0443,590,353 guj_Gujr000007,4923,9893,578,51103,589,992 mal_Mlym00479207,6601,7493,577,96003,587,850 nno_Latn00000003,572,36514,5183,586,883 npi_Deva00500005,0204,0023,576,36703,585,889 sna_Latn000007,4571,3683,576,30903,585,134 zul_Latn000007,6421,8333,574,43703,583,912 afr_Latn000206,50303,577,28503,583,790 kan_Knda000107,5743343,573,85503,581,764 gle_Latn000006,5491,2453,573,61003,581,404 ceb_Latn000007,1307273,573,09203,580,949 mya_Mymr000207,3674723,572,83703,580,678 hat_Latn000007,9821063,572,47103,580,559 kaz_Cyrl00500207,54703,572,36503,580,414 snd_Arab000007,4702743,572,63903,580,383 azj_Latn00548407,31303,572,36503,580,230 kat_Geor00500107,35103,572,36503,580,217 jav_Latn000006,4212473,573,44103,580,109 khm_Khmr000107,71403,572,36503,580,080 epo_Latn269001707,26603,572,36503,579,917 khk_Cyrl000607,19903,572,36503,579,570 hye_Armn0055030003,576,38203,576,935 xho_Latn000001,3513773,574,80603,576,534 lao_Laoo000103,67203,572,36503,576,038 pbt_Arab0000009893,573,35403,574,343 sun_Latn00000901943,573,76703,574,051 arz_Arab0000005293,572,89403,573,423 sot_Latn0000065803,572,36503,573,023 ars_Arab0000001363,572,50103,572,637 apc_Arab000000813,572,44603,572,527 Continued on next page 18 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 11 â Continued from previous page LanguageOAMMLUIncTagEuroWikiAHACWQATotal ckb_Arab000000793,572,44403,572,523 lat_Latn000407,7560021,83629,596 lij_Latn000006,71505,95516,16228,832 oci_Latn000006,8130017,19424,007 lim_Latn000006,6930016,66823,361 nds_Latn000006,9880016,21623,204 vec_Latn000005,4240017,74923,173 scn_Latn000005,9990016,84222,841 pan_Guru000007,9096,3858,541022,835 bar_Latn000003,0700019,56322,633 hrv_Latn005501130006,91314,96622,470 fao_Latn000006,0650016,10222,167 bre_Latn000106,3810015,53021,912 arg_Latn000005,3040015,43920,743 roh_Latn000002,0620017,82319,885 srd_Latn000005,6060013,85019,456 kmr_Latn000006,9860011,67418,660 ast_Latn0000000018,38318,383 fry_Latn0000000017,45017,450 bos_Latn0000000015,93315,933 sco_Latn0000000015,21215,212 dag_Latn0000000012,84812,848 szl_Latn000004,755007,38812,143 fur_Latn000003,114007,15210,266 lmo_Latn000000009,1539,153 nap_Latn000000007,8797,879 wol_Latn000008572,9143,14606,917 war_Latn000206,4370006,439 pfl_Latn000000006,3216,321 san_Deva000106,0520006,053 tuk_Latn000105,9190005,920 frp_Latn000000005,1235,123 fil_Latn0000001,2411,24102,482 nya_Latn0000085368868802,229 bod_Tibt000101,9980001,999 ltg_Latn000000001,8721,872 rmy_Latn00000000877877 anp_Deva000000005757 BSCORE DISTRIBUTION OF LANGUAGE FAMILIES CLASSIFIER Figures B and 1 show the distribution of quality scores assigned to French documents by different classifiers. Several patterns emerge: âą French and Romance classifiers (Fig B) exhibit similar bimodal distributions: most doc- uments score near 0 (clear negatives) with a long tail toward 1 (clear positives). The 90th percentile cutoffs are comparable (French: 0.045, Romance: 0.048). âą Nordic classifier (Fig 1) produces a markedly different distribution despite successfully separating quality tiers. It applies a much stricter threshold (90th percentile: 0.027) and assigns generally lower scores, yet still identifies high-quality content effectively. âą Romance without French (Fig B) shows intermediate behavior, with more score mass in the middle range, suggesting less confident predictions. CQUALITATIVE ANALYSIS OF CLASSIFIER RANKINGS This section provides a qualitative comparison of how different classifiers rank the same French documents. By observing extreme rank shifts, we can infer the âfeaturesâ prioritized by each model. C.1LANGUAGE FAMILY CLASSIFIERS COMPARISON ON FRENCH FILTERING We provide here the full table for all the language families classifiers filtering on French. 19 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. 0.00.20.40.60.81.0 Quality Score 10 2 10 1 10 0 10 1 10 2 Density Test Positive Test Negative FineWeb2 Cutoff (retention 0.10) 0.00.20.40.60.81.0 Quality Score 10 2 10 1 10 0 10 1 10 2 Density Test Positive Test Negative FineWeb2 Cutoff (retention 0.10) Figure 2: Distribution of quality scores for French baseline (top) and Romance languages classifier including French (bottom) on French FineWeb2 samples 0.00.20.40.60.81.0 Quality Score 10 2 10 1 10 0 10 1 10 2 Density Test Positive Test Negative FineWeb2 Cutoff (retention 0.10) Figure 3: Distribution of quality scores for Romance languages classifier without French C.2MONOLINGUAL (HQ) VS. MULTILINGUAL (ML) IN FRENCH Sample 1: Structured Educational Essay (Poetry) Text: la poĂ©sie- rĂ©alitĂ©/ poĂ©sie: forme dâĂ©vasion du rĂ©el :Introduction La poĂ©sie est un genre littĂ©raire qui permet dâexprimer des sentiments... : ThĂšse La poĂ©sie peut ĂȘtre consid- 20 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 12: Spearman and Kendall correlations between family-specific classifiers and the French HQ baseline. High correlation in distant families (e.g., Uralic, Nordic) suggests cross-lingual transfer of quality signals, while the drop in "Romance (no French)" suggests potential syntactic interference. ExperimentSpearman Kendall Romance (spa, fra, por, ita, ron, cat) MKC+0.89280.7173 Uralic (fin, ekk, hun) MKC+0.88700.7073 Nordic (swe, dan, nob, isl) MKC+0.88200.6990 Nordic (swe, dan, nob, isl) MKC-e0.87900.7014 Germanic (deu, nld, en, afr, ltz) MKC+0.87500.6916 Uralic (fin, ekk, hun) MKC-e0.86510.6823 Germanic (deu, nld, en, afr, ltz) MKC-e0.84650.6585 Slavic (pol, rus, ces, ukr, bul, srp, hrv) MKC+0.83870.6484 Romance (spa, fra, por, ita, ron, cat) MKC-e0.78080.5880 Slavic (pol, rus, ces, ukr, bul, srp, hrv) MKC-e0.74430.5523 Indo-Aryan (hin, urd, ben, pan, mar) MKC+0.71570.5221 Romance, no French (spa, por, ita, ron, cat) MKC-e0.71390.5228 Indo-Aryan (hin, urd, ben, pan, mar) MKC-e0.63480.4565 Ă©rĂ©e comme une forme dâĂ©vasion du rĂ©el... : AntithĂšse Cependant, la poĂ©sie peut Ă©galement ĂȘtre considĂ©rĂ©e comme une reprĂ©sentation de la rĂ©alitĂ©... :Conclusion En fin de compte, la poĂ©sie peut ĂȘtre considĂ©rĂ©e Ă la fois comme une forme dâĂ©vasion du rĂ©el et une reprĂ©senta- tion de la rĂ©alitĂ©... Translation: Poetry-reality/poetry: a form of escape from reality :Introduction Poetry is a literary genre that allows the expression of feelings... :Thesis Poetry can be considered a form of escape from reality because it allows the author and reader to escape... :An- tithesis However, poetry can also be considered a representation of reality... :Conclusion Ultimately, poetry can be considered both a form of escape from reality and a representa- tion of reality... HQ Rank: 18,773,069 (Score: 0.1301)â ML Rank: 10,298 (Score: 0.9988) Shift: +18,762,771 Analysis: The monolingual HQ classifier failed to prioritize this highly structured educational essay. In contrast, the ML classifier correctly identified it as high-quality content. This suggests that Multilingual training sensitizes the model to universal academic markers (like "Introduction," "Thesis," and "Conclusion") which appear across many languages in the MKC+ pool. Sample 2: Grammatically Fluent Nonsense Text: Cela façon dont la personne a une bouteille de vie dans la . Il sâagit dâattraction simplement pas sembler le problĂšme avec vous mĂšnera probablement vous aimez pas ĂȘtre exactement . Et pleine forme auprĂšs de ce que câest dâabord, les cadres supĂ©rieurs et plus faibles qui ne fonctionne de compliments semblerait ĂȘtre. Translation: This way in which the person has a bottle of life in the . It is a matter of attraction simply not to seem the problem with you will probably lead you do not like being exactly . And in great shape with what it is first, the senior and weaker executives who does not work of compliments would seem to be. HQ Rank: 65,072 (Score: 0.9959)â ML Rank: 18,768,564 (Score: 0.0199) Shift: -18,703,492 Analysis: This text uses correct French words and localized syntax, but the meaning is total non- sense (e.g., âa bottle of life in theâ). The monolingual HQ model was fooled by the surface-level fluency, but the ML model correctly identified it as noise. This indicates that multilingual embed- dings help the model verify semantic coherence, as ânonsenseâ rarely aligns well across different languages in latent space. In Figure C.2 we see that this behaviour generalizes to Spanish in much less pronounced way, but does not hold for Chinese. C.3ROMANCE (NO-FRENCH) CLASSIFIER Sample 3: Formal/Liturgical Text (Psalms) 21 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Density T=0.020 Multilingual FW2: 11.7% Syn: 82.0% T=0.130 Monolingual FW2: 13.2% Syn: 95.0% T=0.079 French FW2-HQ FW2: 13.6% Syn: 87.8% Density T=0.007 FW2: 10.6% Syn: 99.9% T=0.012 FW2: 14.7% Syn: 99.9% T=0.003 Chinese FW2: 14.0% Syn: 99.9% 0.00.20.40.60.81.0 Score Density T=0.057 FW2: 13.5% Syn: 98.0% 0.00.20.40.60.81.0 Score T=0.120 FW2: 18.1% Syn: 99.7% 0.00.20.40.60.81.0 Score T=0.116 Spanish FW2: 17.7% Syn: 99.9% Figure 4: Comparison of score distributions for 10K synthetically generated grammatically cor- rect nonsense samples across multilingual (ML), monolingual (HQ), and FineWeb2-HQ (FW2-HQ) classifiers for French, Chinese, and Spanish, compared to 10K random samples from FineWeb2. Synthetic samples were generated using Qwen3-32B (Team, 2025), with grammatically correct nonsense text concatenated to match the document length distribution of FineWeb2. T denotes the threshold used to retain the top 10% of data in our experiments. The colored area represents the proportion of retained documents. Text: |1||2||3|... 1De David. Ăternel, je me tourne vers toi, 2mon Dieu, en toi je me confie. Que je ne sois pas couvert de honte! Que mes ennemis ne se rĂ©jouissent pas Ă mon sujet!... 4Ăternel, fais-moi connaĂźtre tes voies, enseigne-moi tes sentiers! 5Conduis-moi dans ta vĂ©ritĂ©... Translation: |1||2||3|... 1Of David. Eternal, I turn toward you, 2my God, in you I trust. Let me not be covered in shame! Let my enemies not rejoice over me!... 4Eternal, make me know your ways, teach me your paths! 5Lead me in your truth... Baseline Rank: 332,619,231 (Score: 0.0000)â No-Fra Rank: 82,936 (Score: 0.9942) Shift: +332,536,295 Analysis: A classifier trained on other Romance languages (Spanish, Italian, etc.) but not French significantly prioritized this liturgical text. This proves that formal registers and religious signatures are highly conserved across language families. The âqualityâ of this register is recognized zero-shot across languages. Sample 4: Transient News (Sports) Text: Myriam SoumarĂ© sâest qualifiĂ©e pour les demi-finales du 200m des Championnats du monde dâathlĂ©tisme, ce jeudi Ă Moscou, en terminant troisiĂšme de sa sĂ©rie en 22â83. La sprinteuse française, championne dâEurope de la discipline en 2010, tentera de se qualifier pour la finale... Translation: Myriam SoumarĂ© qualified for the semi-finals of the 200m at the World Ath- letics Championships this Thursday in Moscow, finishing third in her heat in 22â83. The French sprinter, European champion in the discipline in 2010, will attempt to qualify for the final... Baseline Rank: 13,957,851 (Score: 0.3891)â No-Fra Rank: 318,322,669 (Score: 0.0000) Shift: -304,364,818 Analysis: While informative, this snippet was heavily penalized by the Romance-transfer model. This confirms our hypothesis that cross-family transfer pushes the model toward âEncyclopedicâ quality anchors (like Wikipedia) and away from the more âcommonâ reporting found in general web crawls. 22 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. C.4NORDIC CLASSIFIER Sample 5: Public Park Information Text: Square de la place de la RĂ©union. Horaires: Ouvert en ce moment |jeudi 21/03||08:00 Ă 18:00|... En 1849, le village dit « Grand Charonne » rejoint le hameau le « Petit Charonne »... Ce jardin contemporain a Ă©tĂ© Ă©co labĂ©lisĂ© en 2012... Le hĂȘtre pourpre au feuillage rouge brun crĂ©e une continuitĂ© avec les coloris pourpres du fossĂ© humide. Translation: Square of the Place de la RĂ©union. Hours: Open now |Thursday 03/21||08:00 to 18:00|... In 1849, the village known as "Grand Charonne" joined the hamlet of "Petit Charonne"... This contemporary garden was eco-labeled in 2012... The purple beech with red-brown foliage creates continuity with the purple colors of the wet ditch. HQ Rank: 327,597,576 (Score: 0.0001)â Nordic Rank: 8,268,964 (Score: 0.3525) Shift: +319,328,612 Analysis: This document is heavily ânoisy,â starting with long tables of opening hours and ending with Twitter handles. The monolingual HQ baseline likely penalized the initial boilerplate so severely that it discarded the entire document. However, the Nordic classifier identifies the knowledge-dense middle section (historical facts and detailed botanical descriptions). This suggests that cross-family transfer models can be more robust to localized boilerplate, focusing instead on the global density of information. Sample 6: E-commerce Boilerplate (Empty Cart) Text: Toutes les catĂ©gories Votre panier est vide! Index des marques: 0 - 9 A B C D E... Ce produit est en rupture de stock. Vous pouvez remplir ce formulaire pour ĂȘtre notifiĂ©... Translation: All categories Your cart is empty! Brand index: 0-9 A B C D E... This product is out of stock. You can fill out this form to be notified... Baseline Rank: 6,322,191 (Score: 0.6707)â Nordic Rank: 320,914,784 (Score: 0.0000) Shift: -314,592,593 Analysis: Even without knowing French, the Nordic model identifies the structural signature of low- utility e-commerce pages (e.g., brand lists, empty cart messages). These patterns are cross-lingual, allowing the model to effectively filter web noise in languages it has never seen, compared to the HQ baseline, which gave it a high score due to its fluent writing. DANALYSIS OF Q3 HARD NEGATIVES The Q3 sampling strategy selects documents that score in the 50th-75th percentile of the baseline distribution. These samples are critical for refining the decision boundary of the classifier. Q3 negative sample: Administrative School Snippets Text: PubliĂ© dans Vie des Ă©coles le 16.09.12. Les Ă©lĂšves des Ă©coles LĂ©onard De Vinci collectent les journaux (quotidiens nationaux et rĂ©gionaux). PubliĂ© dans Vie des Ă©coles le 18.09.10. Consulter les diffĂ©rentes commissions et comitĂ©s de lâannĂ©e scolaire 2012-2013. La grande affaire de la rentrĂ©e 2008/2009 Ă lâĂ©cole Ă©lĂ©mentaire LĂ©onard de Vinci a Ă©tĂ© la mise en place de lâaide personnalisĂ©e pour venir en aide aux enfants en difficultĂ©s. Cette annĂ©e, les effectifs sont en hausse : 118 Ă©lĂšves sont inscrits Ă lâĂ©cole maternelle. Translation: Published in School Life on 16.09.12. Students from the LĂ©onard De Vinci schools collect newspapers (national and regional dailies). Published in School Life on 18.09.10. Consult the various commissions and committees for the 2012-2013 school year. The main development of the 2008/2009 school year at the LĂ©onard de Vinci elementary school was the implementation of personalized support to assist children with difficulties. This year, enrollment is rising: 118 students are enrolled in the nursery school. Label: Negative (Q3 Sample) Analysis: This document is a perfect example of a useful negative. It is flawlessly written, uses proper punctuation, and contains no âweb noiseâ like ads or code. However, its informational content is hyper-local, administrative, and transient (referring to school enrollment numbers and newspaper drives from 2008â2012). While a standard classifier might be tempted to rank this highly because it contains educational keywords like Ă©cole (school) and LĂ©onard De Vinci, the Q3 strategy teaches the 23 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. model that fluency does not equal informational depth. By using such samples as negatives, we force the classifier to look past surface-level grammar and prioritize documents with actual semantic or scientific weight. EMEGATRON CONFIG To train the 1B models, we use the Apertus tokenizer, the Megatron LM library (Shoeybi et al., 2020), a batch size of 2.06M tokens, a learning rate of 0.00015, the AdEMAMix optimizer (Pagliar- dini et al., 2024), 50,000 training steps with 2000 warmup steps, and a WSD learning rate schedule (?). These models were trained using 84 NVIDIA GH200 chips. The full configuration for these models is displayed in Table 13). This training pipeline is very similar to what has been done in (Messmer et al., 2025). Table 13: Training Configuration for Apertus 1B Model ParameterValue Model Architecture Number of Layers16 Hidden Size2048 FFN Hidden Size12288 Number of Attention Heads32 Number of Query Groups8 Maximum Position Embeddings 4096 Position Embedding TypeRoPE RoPE Base500000 RoPE Scaling Factor32 NormalizationRMSNorm Activation FunctionXieLU Training Configuration Micro Batch Size3 Global Batch Size504 Sequence Length4096 Total Training Steps50,000 Total TokensâŒ103B Checkpoint Interval2,000 steps Optimization OptimizerAdEMAMix Learning Rate0.00015 Minimum Learning Rate0.0 LR ScheduleWSD (1-sqrt decay) Warmup Steps2,000 WSD Decay Steps10,000 Weight Decay0.1 Gradient Clipping0.1 Adam ÎČ 1 0.9 Adam ÎČ 2 0.999 AdEMAMix α8 AdEMAMix ÎČ 3 0.9999 AdEMAMix ÎČ 3 Warmup100,000 AdEMAMix α Warmup100,000 Regularization Attention Dropout0.0 Hidden Dropout0.0 Infrastructure Number of Nodes21 GPUs per Node4 Total GPUs84 Continued on next page 24 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 13 â Continued from previous page ParameterValue Tensor Parallelism1 Pipeline Parallelism1 PrecisionBF16 Additional Features Tokenizerswiss-ai/Apertus-70B-2509 Goldfish Loss (k, h)50, 50 Cross-document AttentionEnabled QK LayerNormEnabled Seed28 FRANKING PROCEDURE FOR APPROACHES For each language, we evaluate filtering strategies by training 1B parameter Apertus models on their filtered outputs and benchmarking on language-appropriate tasks (detailed in Appendix G). We rank methods by their performance on each individual benchmark, then compute the average rank across all benchmarks for that language. A lower average rank indicates better overall performance. This ranking approach is robust to scale differences across benchmarks and emphasizes consistency across diverse evaluation tasks. We also include mean normalized accuracy as a metric, as it gives us more quantitative insight for the performance gain of each method. GBENCHMARKS BY LANGUAGE G.1FRENCH The following benchmarks were used to evaluate model performance on French: âą ARC-Challenge (Clark et al., 2018) âą Belebele (Bandarkar et al., 2024) âą Global-MMLU (Singh et al., 2025) âą HellaSwag (Zellers et al., 2019; Dac Lai et al., 2023) âą Include-Base-44 (Romanou et al., 2024) âą Multilingual MMLU (Institute, 2025) âą XNLI (Conneau et al., 2018) âą XWinograd (Muennighoff et al., 2022; Tikhonov & Ryabinin, 2021) G.2ARABIC The following benchmarks were used to evaluate model performance on Arabic: âą ARC-Easy (Clark et al., 2018) âą AlGhafa PIQA-MT (Almazrouei et al., 2023) âą AlGhafa RACE (OALL, 2023) âą AlGhafa SciQ (OALL, 2023) âą ARC-Challenge (Clark et al., 2018) âą Belebele (Bandarkar et al., 2024) âą Global-MMLU (Singh et al., 2025) âą HellaSwag (Zellers et al., 2019; Dac Lai et al., 2023) âą Include-Base-44 (Romanou et al., 2024) 25 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. âą Multilingual MMLU (Institute, 2025) âą AlGhafa PIQA (Bisk et al., 2019; OALL, 2023) âą XNLI (Conneau et al., 2018) âą XStoryCloze (Lin et al., 2021) G.3SPANISH The following benchmarks were used to evaluate model performance on Spanish: âą ARC-Challenge (Clark et al., 2018) âą Belebele (Bandarkar et al., 2024) âą Global-MMLU (Singh et al., 2025) âą HellaSwag (Zellers et al., 2019; Dac Lai et al., 2023) âą Include-Base-44 (Romanou et al., 2024) âą Multilingual MMLU (Institute, 2025) âą XNLI (Conneau et al., 2018) G.4CHINESE The following benchmarks were used to evaluate model performance on Chinese: âą Agieval Cn (Zhong et al., 2023; Ling et al., 2017; Hendrycks et al., 2021; Liu et al., 2020; Zhong et al., 2020; Wang et al., 2021) âą ARC (Clark et al., 2018) âą Belebele (Bandarkar et al., 2024) âą Ceval-valid (Huang et al., 2023) âą Cmmlu (Li et al., 2024) âą Global-MMLU (Singh et al., 2025) âą Include-Base-44 (Romanou et al., 2024) âą Multilingual MMLU (Institute, 2025) âą PAWS-X (Yang et al., 2019) âą Xcopa (Ponti et al., 2020; Roemmele et al., 2011) âą XNLI (Conneau et al., 2018) âą XStoryCloze (Lin et al., 2021) âą XWinograd (Muennighoff et al., 2022; Tikhonov & Ryabinin, 2021) HTHE ROLE OF SCALE AND SEED VARIANCE We have investigated the addition of languages and the curation of negative samples. However, we have to ask ourselves about their stability. If we were to supply different positive/negative samples, we would expect to get very similar results. To test this hypothesis, we vary the sampling seed of the classifier. This would result in selecting different samples from our positive anchors, but also from FineWeb2. This experiment is labeled as HQ seed. Tables 14 and 15 analyze the stability of quality filtering with respect to sampling variance and training scale. While multilingual pooling improves overall rank stability, our ablations show that downstream LLM performance remains sensitive to the classifierâs initial sampling seed. Changing the seed shifts aggregate accuracy by 0.3% in Arabic and 0.8% in French (Tables 14 and 15). This variance suggests that the specific subset of examples used to define the quality boundary heavily influences which knowledge domains are selected. 26 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 14: Comparison of the HQ baseline against a version trained with a different sampling seed (HQ seed) on Arabic. The variation in results across tasks suggests that in lower-resource settings, the specific documents selected for the anchor set can significantly impact the classifierâs decision boundary. BenchmarkNo filteringHQHQ seed ARC-Easy0.27160.28980.2838 AlGhafa PIQA-MT0.51940.51450.5095 AlGhafa RACE0.27150.27750.2879 AlGhafa SciQ0.42610.44320.4492 ARC-Challenge0.26780.26600.2720 Belebele_c0.32220.32330.3289 GMMLU_c0.24750.25750.2450 HellaSwag0.39090.38730.3909 Include_c0.30250.26810.2844 M_MMLU_c0.26690.26140.2653 AlGhafa PIQA0.61320.60880.6126 XNLI0.33490.33090.3317 XStoryCloze0.59030.59230.5930 Aggregate acc_norm0.37110.37080.3734 Average rank1.922.311.69 Table 15: Comparison of standard HQ against variants with different classifier seeds, LLM training seeds, and larger positive anchor sets (HQ all) on French. Results indicate that variance in data selection (classifier seed) has a larger impact on downstream performance than the stochasticity of the LLM training itself. BenchmarkNo filteringHQHQ (seed) HQ seed LLM Training HQ (all) ARC-Challenge0.28910.30710.32160.29850.3080 Belebele_c0.34440.35110.35000.33890.3544 GMMLU_c0.26250.29250.29000.30750.2875 HellaSwag0.48830.47480.47610.47740.4773 Include_c0.38660.41530.40570.41290.4105 M_MMLU_c0.28310.29450.29440.29460.2949 XNLI0.47070.48550.48230.49040.4695 XWinograd0.63860.65060.59040.62650.6386 Aggregate acc_norm0.39540.40890.40130.40580.4051 Average rank3.882.383.382.622.62 To mitigate this variance, we examine two complementary strategies: increasing positive sample coverage by using all available high-quality anchors (HQ all), and varying the LLM training seed independently from the filtering process. While altering the LLM seed introduces minor variability, the dominant source of instability arises from the classifier sampling process itself. Using the full positive set moderately improves robustness but does not consistently outperform standard HQ filtering, suggesting diminishing returns from scale alone. This result is counterintuitive if we assume that âmore data is better.â We hypothesize that this degradation is due to an informational saturation effect: by providing the classifier with the entire, unfiltered anchor pool, we likely introduced a higher ratio of ânon-helpfulâ or marginal samples that are present in the positive datasets but lack strong educational signal. This prevents the classifier from establishing a sharp decision boundary between truly high-quality content and baseline web text. This suggests that representative, curated sampling is a more effective strategy for training quality filters than just increasing the number of training samples. These results motivate multilin- 27 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 16: Evaluation of the monolingual HQ baseline versus a classifier trained on the expanded MKC-e pool on Spanish. The inclusion of instruction and synthetic data improves aggregate nor- malized accuracy by nearly 0.5%, indicating high utility for Spanish domain coverage. BenchmarkNo filteringHQMKC-e ARC-Challenge0.29910.30770.3291 Belebele_c0.34220.35330.3378 GMMLU_c0.31000.31500.3200 HellaSwag0.50060.52630.5075 Include_c0.34910.37270.3964 M_MMLU_c0.28140.29850.3137 XNLI0.46180.45380.4566 Aggregate acc_norm0.36350.37530.3801 Average rank2.571.861.57 Table 17: Comparison of standard HQ versus the MKC-e anchor set. The extended data provides a marginal gain in aggregate accuracy (+0.1%) but significantly improves performance on specific benchmarks like ARC-Challenge and XWinograd. BenchmarkNo filteringHQMKC-e ARC-Challenge0.28910.30710.3259 Belebele_c0.34440.35110.3533 GMMLU_c0.26250.29250.2650 HellaSwag0.48830.47480.4647 Include_c0.38660.41530.4296 M_MMLU_c0.28310.29450.2929 XNLI0.47070.48550.4735 Xwinograd0.63860.65060.6747 Aggregate acc_norm0.39540.40890.4099 Average rank2.751.621.62 gual and bootstrapped approaches introduced in the paper, which effectively average over topic and language variability to yield more stable quality signals. IADDITION OF MKC-E DATASETS. To make our classifier highly multilingual beyond the coverage given by the Aya Collection (Singh et al., 2024), we extend the MKC+ pool with additional datasets (resulting in the MKC-e data). A hypothesis is that, beyond allowing us to train on more languages, this addition provides more topics and diversity, which can improve the performance of our filtering. In order to test this, we train some monolingual classifiers on the MKC-e data and evaluate them using the training of the 1B Apertus. Results are displayed in Tables 16, 17, and 18. While the training on MKC-e improves results for Spanish by almost 0.5% in terms of normalized accuracy, we find that the results are more nuanced for French, where the improvement is of 0.1%. However, the Arabic MKC-e classifier seems to perform worse than the âNo filteringâ baseline, which suggests adding this new data pool provides mixed results depending on the targeted language. JGLOBAL BENCHMARK RESULTS 28 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 18: Arabic performance with extended anchors (MKC-e). Unlike the Romance languages, adding the MKC-e pool to Arabic filtering slightly degrades aggregate performance. This suggests that the instruction-tuning signal in MKC-e may not align as cleanly with "educational quality" for Arabic. BenchmarkNo filteringHQMKC-e ARC-Easy0.27160.28980.2728 AlGhafa PIQA-MT0.51940.51450.4970 AlGhafa RACE0.27150.27750.2753 AlGhafa SciQ0.42610.44320.4422 ARC-Challenge0.26780.26600.2643 Belebele_c0.32220.32330.3244 GMMLU_c0.24750.25750.2550 HellaSwag0.39090.38730.3882 Include_c0.30250.26810.2953 M_MMLU_c0.26690.26140.2628 AlGhafa PIQA0.61320.60880.6121 XNLI0.33490.33090.3349 Xstorycloze0.59030.59230.5890 Aggregate acc_norm0.37110.37080.3703 Average rank1.852.002.08 Table 19: Spanish comprehensive benchmark results. Final comparison of all filtering strategies. ML (Q3) achieves the best average rank, showing that the combination of multilingual signal and sharpened decision boundaries is optimal for Spanish. BenchmarkNo filteringHQHQ seed MKC-eML ML (15%) ML (Q3) ARC-Challenge0.29910.30770.32740.32910.3248 0.3436 0.3282 Belebele_c0.34220.35330.34220.33780.3456 0.3500 0.3533 GMMLU_c0.31000.31500.33750.32000.3250 0.3225 0.3275 HellaSwag0.50060.52630.52190.50750.5310 0.5440 0.5372 Include_c0.34910.37270.37820.39640.3891 0.4036 0.4000 M_MMLU_c0.28140.29850.30730.31370.3050 0.3080 0.3083 XNLI0.46180.45380.47070.45660.4783 0.4562 0.4667 Aggregate acc_norm0.36350.37530.38360.38010.3855 0.3897 0.3887 Average rank6.295.143.714.143.572.712.14 29 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 20: Chinese comprehensive benchmark results. Final comparison of all filtering strategies. The ML strategy secures the best rank and aggregate accuracy, demonstrating that multilingual sig- nal provides the most robust quality filter for Chinese logographic text. BenchmarkNo filteringHQHQ seed MKC-eML ML (15%) ML (Q3) Agieval Cn0.36180.36440.36250.36090.3457 0.3497 0.3658 ARC0.28550.31450.30000.31620.3171 0.3068 0.3085 Belebele_c0.30110.32000.32560.34330.3222 0.3356 0.3478 Ceval-valid0.22880.24890.24440.25560.2615 0.2370 0.2467 Cmmlu_c0.32060.34710.34450.35810.3608 0.3466 0.3688 GMMLU_c0.28000.30750.31750.32000.3200 0.3025 0.3200 Include_c0.34680.35230.35230.35410.3541 0.3780 0.3670 M MMLU_c0.27720.29400.29370.29590.2987 0.2927 0.2978 PAWS0.55200.55350.53600.54350.5610 0.5570 0.5480 Xcopa0.58600.59200.62000.59400.6080 0.6160 0.6020 XNLI0.35460.40720.37790.40560.4189 0.3980 0.3627 Xstorycloze0.65590.66250.66050.66180.6625 0.6678 0.6750 XWinograd0.68060.68250.69640.69250.6766 0.6865 0.6667 Aggregate acc_norm0.40240.41900.41780.42320.4236 0.4211 0.4213 Average rank6.383.854.463.232.693.923.00 Table 21: Arabic comprehensive benchmark results. Final comparison of all filtering strategies. The ML strategy (56% retention) remains the dominant approach, while aggressive filtering (10-20%) consistently underperforms regardless of the classifier used. Benchmark No Filtering HQ HQ seed MKC-eML ML (Q3) HQ (10%) HQ (20%) ML (10%) ML (20%) ARC-Easy0.27160.2898 0.28380.27280.2855 0.2750 0.2720 0.2771 0.2741 0.2762 AlGhafa PIQA-MT0.51940.5145 0.50950.49700.5112 0.5085 0.5019 0.5090 0.4948 0.5128 AlGhafa RACE0.27150.2775 0.28790.27530.2883 0.2792 0.2747 0.2834 0.2765 0.2786 AlGhafa SciQ0.42610.4432 0.44920.44220.4503 0.4322 0.4503 0.4171 0.4573 0.4563 ARC-Challenge0.26780.2660 0.27200.26430.2797 0.2601 0.2669 0.2772 0.2712 0.2626 Belebele_c0.32220.3233 0.32890.32440.3122 0.3444 0.3378 0.2944 0.3278 0.3256 GMMLU_c0.24750.2575 0.24500.25500.2700 0.2575 0.2475 0.2550 0.2725 0.2525 HellaSwag0.39090.3873 0.39090.38820.3925 0.3870 0.3592 0.3728 0.3747 0.3857 Include_c0.30250.2681 0.28440.29530.3043 0.2754 0.2717 0.2772 0.2844 0.2790 M_MMLU_c0.26690.2614 0.26530.26280.2646 0.2615 0.2652 0.2674 0.2618 0.2634 AlGhafa PIQA0.61320.6088 0.61260.61210.6110 0.6072 0.5958 0.6023 0.6039 0.6143 XNLI0.33490.3309 0.33170.33490.3349 0.3373 0.3325 0.3333 0.3325 0.3357 XStoryCloze0.59030.5923 0.59300.58900.5811 0.5831 0.5877 0.5817 0.5784 0.5804 Aggregate acc_norm0.37110.3708 0.37340.37030.3758 0.3699 0.3664 0.3652 0.3700 0.3710 Average rank5.005.774.085.853.465.856.926.086.085.15 KLIMITATIONS We acknowledge several limitations that constrain the generalizability of our findings: Statistical Rigor. Due to the computational cost of training 1B parameter models, all primary results are reported as single runs. Seed variance experiments (Tables 14 and 15) show that changing the classifier sampling seed shifts aggregate accuracy by âŒ0.3% in Arabic and âŒ0.8% in French. For Arabic in particular, the ML gain over HQ (âŒ0.5%) is comparable to this noise range and should therefore be interpreted with caution. Future work should employ multiple seeds per condition to establish confidence intervals. Embedding Space Interpretation. While our cross-lingual transfer results are empirically con- sistent, we cannot determine whether they reflect a genuinely abstract quality structure in the embed- ding space or shared formatting across our positive anchor datasets, such as Wikipedia markup or 30 Published as a paper at 3rd DATA-FM workshop @ ICLR 2026, Brazil. Table 22: French comprehensive benchmark results. Detailed comparison of 12 distinct filtering strategies. While ML (15%) achieves the highest absolute accuracy, the Q3 negatives and Nordic transfer models remain highly competitive, proving that quality can be captured through multiple distinct curation pathways. Benchmark No Filtering HQ HQ seed HQ seed (LLM Training) HQ all Q3 negatives HQ Romance (No Fra) MKC-e NordicML ML (15%) ML (Q3) ARC-Challenge0.28910.3071 0.32160.29850.30800.31050.30540.32590.31570.3165 0.3139 0.3054 Belebele_c0.34440.3511 0.35000.33890.35440.37110.36890.35330.34220.3511 0.3611 0.3633 GMMLU_c0.26250.2925 0.29000.30750.28750.30250.27500.26500.30750.2900 0.2875 0.3075 HellaSwag0.48830.4748 0.47610.47740.47730.47270.49080.46470.46730.4956 0.5135 0.4911 Include_c0.38660.4153 0.40570.41290.41050.42720.42720.42960.45350.4224 0.4726 0.4582 M_MMLU_c0.28310.2945 0.29440.29460.29490.29920.29500.29290.29360.2927 0.2991 0.3007 XNLI0.47070.4855 0.48230.49040.46950.48760.48270.47350.47830.4807 0.4807 0.4767 XWinograd0.63860.6506 0.59040.62650.63860.66270.60240.67470.66270.6024 0.6386 0.6145 Aggregate acc_norm0.39540.4089 0.40130.40580.40510.41670.40590.40990.41510.4064 0.4209 0.4147 Average rank9.886.387.626.757.384.006.006.886.126.504.124.62 instruction-tuning templates. Disentangling these mechanisms, for example through probing classi- fiers or controlled anchor set ablations, remains to be addressed in future work. Limited Language Coverage. Our evaluation focuses on four languages, all of which have sub- stantial representation in XLM-RoBERTaâs pretraining corpus. Results may not generalize to ex- tremely low-resource languages or underrepresented language families (e.g., Niger-Congo, Aus- tronesian) Embedding Model Dependence. All results rely on XLM-RoBERTa embeddings. The existence and accessibility of quality manifolds may vary with a different architecture or embedding model, a different embedding dimension, etc. LFUTURE WORK While our 1B model provides a solid baseline, scaling 3B or 7B parameter models would determine if our ML classifier impact becomes even stronger with more capacity. It would be particularly in- teresting to run cooldown experiments: instead of stopping at the stable phase, we could introduce a short, high-quality annealing phase (last 10-20% of tokens) to see if we can ârecoverâ performance on formal tasks while keeping the benefits of our broader multilingual filtering. There is also a big opportunity to test zero-shot transfer on truly low-resource languages like Swahili or Urdu, where native data is so scarce that cross-lingual âsubsidyâ from high-resource languages is the only vi- able path forward. Another ablation idea would be to also apply a similar technique to the positive samples fed to the classifier. Thanks to our experiments, we argue that selecting better positives and combining that with the Q3 strategy could lead to an even bigger performance boost. Finally, the rigidity of the sample count threshold for the ML (Q3) strategy warrants further investigation. Relying on a static count (e.g., 200, 000 samples) is likely suboptimal. Future work should explore adaptive criteria based on score distribution analysis, variance, or density heuristics to dynamically determine eligibility for Q3 sampling. This would optimize the trade-off between refining the deci- sion boundary and preserving valid high-quality tokens in medium-resource languages. 31