Paper deep dive
F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World
Ziyin Zhang, Zihan Liao, Hang Yu, Peng Di, Rui Wang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 6:14:00 AM
Summary
F2LLM-v2 is a family of multilingual, general-purpose embedding models ranging from 80M to 14B parameters, trained on 60 million samples across 282 natural and 40 programming languages. It utilizes a two-stage training pipeline, Matryoshka Representation Learning (MRL), model pruning, and knowledge distillation to achieve state-of-the-art performance on 11 MTEB benchmarks while maintaining efficiency for resource-constrained environments.
Entities (5)
Relation Signals (3)
F2LLM-v2 â evaluatedon â MTEB
confidence 100% · Extensive evaluations confirm that F2LLM-v2-14B ranks first on 11 MTEB benchmarks
F2LLM-v2 â basedon â Qwen3
confidence 95% · All models adopt a standard dense Transformer decoder architecture based on Qwen3
F2LLM-v2 â usestechnique â Matryoshka Representation Learning
confidence 95% · By integrating Matryoshka Representation Learning (MRL) and a two-stage training pipeline
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present F2LLM-v2, a new family of general-purpose, multilingual embedding models in 8 distinct sizes ranging from 80M to 14B. Trained on a newly curated composite of 60 million publicly available high-quality data samples, F2LLM-v2 supports more than 200 languages, with a particular emphasis on previously underserved mid- and low-resource languages. By integrating a two-stage LLM-based embedding training pipeline with matryoshka learning, model pruning, and knowledge distillation techniques, we present models that are far more efficient than previous LLM-based embedding models while retaining competitive performances. Extensive evaluations confirm that F2LLM-v2-14B ranks first on 11 MTEB benchmarks, while the smaller models in the family also set a new state of the art for resource-constrained applications. To facilitate open-source embedding model research, we release all models, data, code, and intermediate checkpoints.
Tags
Links
- Source: https://arxiv.org/abs/2603.19223v1
- Canonical: https://arxiv.org/abs/2603.19223v1
Trouble viewing inline? Open PDF directly â
Full Text
110,521 characters extracted from source content.
Expand or collapse full text
F2LLM-v2 Technical Report F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World Ziyin Zhang 1,2 Zihan Liao 1 Hang Yu â ,1 Peng Di â,1 Rui Wang â,2 1 Ant Group 2 Shanghai Jiao Tong University § github.com/codefuse-ai/CodeFuse-Embeddings huggingface.co/collections/codefuse-ai/f2llm gte-Qwen2-7B-instruct F2LLM-v2-0.6BF2LLM-v2-1.7B F2LLM-v2-4BF2LLM-v2-8B F2LLM-v2-14B 57.5 60.0 62.5 65.0 67.5 70.0 72.5 European (128) F2LLM-v2-0.6B SFR-Embedding-2_R F2LLM-v2-1.7B F2LLM-v2-4BF2LLM-v2-8B F2LLM-v2-14B 60.0 62.5 65.0 67.5 70.0 72.5 75.0 Scandinavian (49) F2LLM-v2-0.6B multilingual-e5-large-instruct F2LLM-v2-1.7B F2LLM-v2-4BF2LLM-v2-8B F2LLM-v2-14B 65 70 75 80 Indic (118) F2LLM-v2-330M F2LLM-v2-0.6BF2LLM-v2-1.7B F2LLM-v2-4BF2LLM-v2-8B F2LLM-v2-14B 58 60 62 64 66 68 70 German (96) gte-Qwen2-7B-instruct F2LLM-v2-1.7B bge-multilingual-gemma2 F2LLM-v2-4B F2LLM-v2-14B F2LLM-v2-8B 66 68 70 72 74 French (117) F2LLM-v2-330M F2LLM-v2-0.6BF2LLM-v2-1.7B F2LLM-v2-4BF2LLM-v2-8B F2LLM-v2-14B 60 65 70 75 80 Polish (9) F2LLM-v2-0.6B sarashina-embedding-v2-1b F2LLM-v2-1.7B F2LLM-v2-4BF2LLM-v2-8B F2LLM-v2-14B 67.5 70.0 72.5 75.0 77.5 80.0 82.5 Japanese (19) Qwen3-Embedding-4B F2LLM-v2-4B Qwen3-Embedding-8B Octen-Embedding-8B F2LLM-v2-8B F2LLM-v2-14B 61 62 63 64 65 66 67 68 Dutch (100) Hakim-small F2LLM-v2-1.7B Hakim F2LLM-v2-4BF2LLM-v2-8B F2LLM-v2-14B 64 66 68 70 72 74 76 Persian (30) F2LLM-v2-330M F2LLM-v2-0.6BF2LLM-v2-1.7B F2LLM-v2-4BF2LLM-v2-8B F2LLM-v2-14B 52 54 56 58 60 62 64 66 Vietnamese (25) Previous SOTA Figure 1: The top six models on ten language-specific MTEB leaderboards. The previous SOTA performance is given by the horizontal line. In each subplot title, we list the number of submissions with complete results on the corresponding benchmark. For comparison, the English benchmark has 163 complete submissions. Abstract We present F2LLM-v2, a new family of general-purpose, multilingual embedding models in 8 distinct sizes ranging from 80M to 14B. Trained on a newly curated composite of 60 million publicly available high-quality data samples, F2LLM-v2 supports more than 200 languages, with a particular emphasis on previously underserved mid- and low- resource languages. By integrating a two-stage LLM-based embedding training pipeline with matryoshka learning, model pruning, and knowledge distillation techniques, we present models that are far more efficient than previous LLM-based embedding models while retaining competitive performances. Extensive evaluations confirm that F2LLM- v2-14B ranks first on 11 MTEB benchmarks, while the smaller models in the family also set a new state of the art for resource-constrained applications. To facilitate open- source embedding model research, we release all models, data, code, and intermediate checkpoints. â Correspondence to: Hang Yu <hyu.hugo@antgroup.com>, Peng Di <dipeng.dp@antgroup.com>, Rui Wang <wangrui12@sjtu.edu.cn>. 1 arXiv:2603.19223v1 [cs.CL] 19 Mar 2026 F2LLM-v2 Technical Report 1 Introduction Text embedding models serve as the fundamental backbone for a wide array of AI applications, including semantic search, retrieval-augmented generation (RAG), text classification, and clustering. By mapping unstructured text into dense vector spaces, these models allow machines to capture com- plex semantic relationships, enabling efficient and accurate information retrieval and data analysis across massive datasets. This field has recently transitioned from encoder-based architectures (De- vlin et al., 2019; Liu et al., 2019; Conneau et al., 2020) to decoder-based LLM embeddings (Zhang et al., 2025a; Lee et al., 2025a; Zhang et al., 2025b), benefiting from the extensive reasoning and linguistic capabilities acquired during large-scale pre-training and achieving remarkable gains in performance. Despite these advancements, the current state of frontier embedding research is characterized by two significant limitations. First, there is a pervasive English-centric bias in both model training and benchmark evaluation. While benchmarks such as MTEB have been instrumental in standardizing evaluation, the high-resource language subsets therein - such as English and Chinese - receive a disproportionately large share of attention, resulting in an abundance of models that are performant in English but fail to provide global utility. Second, a transparency gap has emerged within the research community. Most top-performing embedding models, such as Gemini-Embedding (Lee et al., 2025b) and Qwen3-Embedding (Zhang et al., 2025a), are released either as closed-source APIs or open-weight models without disclosing the underlying training data or methodologies. This lack of transparency hinders reproducibility and limits our collective understanding of how to build truly inclusive, general-purpose embedding systems. To directly tackle these challenges, we introduce F2LLM-v2, a new family of general-purpose, multilingual embedding models designed to address these critical imbalances. We curate a massive, high-quality training corpus of 60 million samples spanning 282 natural languages and over 40 programming languages solely from publicly available resources. By prioritizing real-world data availability over benchmark-specific optimization, we create a model family that excels across a truly global range of applications, including those involving underserved languages. Besides linguistic inclusivity, we also address computational inclusivity by providing 8 distinct model sizes, ranging from 80M to 14B parameters. By integrating Matryoshka Representation Learning (MRL) and a two-stage training pipeline enhanced by model pruning and novel knowledge distillation, we ensure high performance even in resource-constrained environments. Extensive evaluations confirm that our 14B model achieves state-of-the-art results on 11 MTEB benchmarks, setting a new standard for multilingual embedding capabilities, while the smaller models also outperform previous frontier models with a similar size. To foster an open and equitable research environment, we release the complete training recipe, intermediate checkpoints, and all associated code and data for the F2LLM-v2 family, aiming to drive progress toward a more inclusive future for AI technology. 2 Related Work The previous generation of encoder-based embedding models witnessed a proliferation of massively multilingual embedding models supporting hundreds of languages, represented by XLM-R (Con- neau et al., 2020), mDeBERTaV3 (He et al., 2023), mBART (Liu et al., 2020), and mT5 (Xue et al., 2021). Recently, decoder-based embedding models have become the dominant paradigm, ben- efiting from their extensive capabilities acquired during large-scale pre-training, as verified by state-of-the-art models such as E5-Mistral (Wang et al., 2024), NV-Embed (Lee et al., 2025a), Qwen3- Embedding (Zhang et al., 2025a), and Gemini-Embedding (Lee et al., 2025b). However, this advancement has been accompanied by a shift toward English-centric evaluation. This is evidenced in MTEB (Muennighoff et al., 2023), which has been established as one of the most recognized text embedding benchmarks, covering over 500 evaluation tasks and more than 250 languages (Enevoldsen et al., 2025). Yet, in reality, the MTEB leaderboards exhibit significant linguistic bias. For instance, in the MTEB-Multilingual benchmark, 35 out of the 131 tasks focus exclusively on English, potentially obscuring a modelâs true multilingual efficacy. Furthermore, 2 F2LLM-v2 Technical Report engzho rus spa fra python deu ara nld vie hin kor jpn ita ind por pol tur dan tha swe php fas java ukr ces nor cpp ell cat go ron fin bul tgl glg mya hye khm nep hun eus heb javascript lao swa azj lav sinslk tgk est lit msa hrv isl slv srp urd c# ben aze afr tam kat tel mal ruby mon nno kaz cym mar sqi nob pus mkd hbs ceb jav war kanepo c lat guj uzb amh ocibel azb kir mlg vol ast pan ltz nds hat bre gle sco xho tat bos yor min che arz rust 10 4 10 5 10 6 10 7 Data Size Natural Language Programming Language Figure 2: Top-100 natural languages and top-10 programming languages in our training data. many language-specific benchmarks receive disproportionately less attention compared with the English or Multilingual benchmarks. As an extreme example, the Polish MTEB benchmark had only a single model with complete results before our models were submitted. This disparity is exacerbated by the fact that many top-performing multilingual embedding models - such as Qwen3-Embedding (Zhang et al., 2025a), Gemini-Embedding (Lee et al., 2025b), and EmbeddingGemma (Vera et al., 2025) - are either closed-source APIs or open-weight only without training transparency. KaLM-Embedding (Zhao et al., 2025) represents one of the few exceptions with transparency in training data, but focuses exclusively on the Multilingual leaderboard and is not evaluated on the aforementioned language-specific benchmarks that are critical for truly global applications. 3 F2LLM-v2 3.1 Training Data English (28.7%) Chinese (7.7%) Russian (6.1%) Spanish (5.0%) French (4.3%) German (3.1%) Arabic (2.5%) Dutch (2.5%) Vietnamese (2.1%) Hindi (2.0%) Korean (1.9%) Japanese (1.9%) Italian (1.7%) Indonesian (1.7%) Portuguese (1.6%) Polish (1.6%) English (49.4%) Chinese (44.4%) Multilingual (6.3%) Ours KaLM-Embedding Figure 3: Comparison between the language distribution of our training data (outer circle) and KaLM-Embedding (inner circle).KaLM- Embeddingâs data is only annotated with three labels, while ours are annotated with specific lan- guages. A cornerstone of F2LLM-v2 is the compilation of a vast and diverse training corpus designed to foster both linguistic inclusivity and broad task competency. We aggregate data from 157 publicly available sources, creating a collection of 60 million training samples that span 282 natural languages (as identified by ISO-639-3 codes) and over 40 programming languages. Crucially, our data curation process is driven by real-world data availability rather than op- timizing for specific benchmarks. For instance, our dataset contains substantial data for Span- ish, Arabic, Italian, Indonesian, and Portuguese (Figure 2), despite these languages lacking ded- icated benchmarks in MTEB. This approach, which also includes a long tail of low-resource languages and a significant volume of code, aims to build a model with truly global util- ity and stands in direct contrast to recent open- source datasets such as the one released by KaLM-Embedding (Zhao et al., 2025), which is heavily skewed towards English and Chinese (Figure 3). We provide a more comprehensive linguistic breakdown of our dataset in Appendix A. 3 F2LLM-v2 Technical Report Question Answering 35.5% Bitext Mining 24.8% Instruction Data 11.9% Title Matching 7.4% NLI 2.9% Code-to-Code 2.4% Topic Classification 2.3% Summarization 2.1% Text-to-Code 2.1% Paraphrase Detection 1.8% Sentiment Analysis 1.6% Code-to-Text 1.6% Intent Classification 1.3% Domain Classification 1.3% STS 1.0% Figure 4: Task type distribution in our training data. The functional diversity of our dataset is equally critical for training a general-purpose embedding model. As shown in Figure 4, our collection encompasses a wide spectrum of tasks, ranging from retrieval-focused question answering and bitext mining to classification-oriented sentiment analysis and intent/domain classification. To leverage this heterogeneity within a unified contrastive learning framework, we follow the first generation of F2LLM (Zhang et al., 2025b) and consolidate all data into three canonical formats: retrieval, clustering, and two-way classification. This consolidation allows the model to learn a versatile embedding space by optimizing a single, consistent objective across disparate data sources and task structures. For the retrieval format, data consists of (query, positive document, hard negatives) tuples. We leverage both in-batch negatives, where other documents in a mini-batch serve as negatives, and explicitly provided hard negatives (mined using Qwen3-Embedding-8B) to create a challenging and efficient training signal. For the clustering format, which also ingests multi-class classification tasks, tuples are formed by sampling an anchor, a positive example from the same class, and a hard negative from a different class. Finally, the two-way classification format directly uses class labels, where a given text serves as the anchor, the corresponding label text is the positive, and the opposite label text is the negative. For both clustering and classification, only hard negatives are utilized to avoid introducing false negatives from in-batch samples. 3.2 Model Architecture We train models in 8 distinct sizes: 80M, 160M, 330M, 0.6B, 1.7B, 4B, 8B, and 14B. All models adopt a standard dense Transformer decoder architecture based on Qwen3 (Yang et al., 2025), and utilize the final hidden states of the EOS token as sequence representation. The detailed model configurations are given in Table 1. Models from 0.6B to 14B directly correspond to Qwen3 LLMs, while the 80M, 160M, and 330M models are pruned from the 0.6B model. 3.3 Two-stage Training We adopt a two-stage training strategy following previous works (Lee et al., 2025a; Zhang et al., 2025a). The first stage focuses on building a robust semantic foundation, and 7 retrieval datasets are selected based on their large scale and broad language coverage, totalling 27 million samples: Code- SearchNet, CodeSearchNet-CCR, OpenCodeGeneticInstruct, WebFAQ, MMARCO, CLIRMatrix, and ParaCrawl (refer to Appendix A for details). Five models (0.6B-14B) are trained in this stage, and we employ the raw data without applying any instructional prefix. 4 F2LLM-v2 Technical Report 80M160M330M0.6B1.7B4B8B14B Model Configuration Hidden Size32064089610242048256040965120 MLP Intermediate Size2048153625603072614497281228817408 Transformer Layers89162828363640 Attention Heads1616161616323240 KV Heads88888888 Head Dimension128128128128128128128128 Model Size Embedding Parameters49M97M136M156M311M389M622M778M Non-Embedding Parameters31M62M198M440M1409M3634M6946M13212M Total Parameters80M159M334M596M1721M4022M7568M13990M Training Configuration MRL Support â â â â â â â â Learning Rate4e-53e-52e-51e-59e-67e-66e-65e-6 Epochs43322222 Batch Size512512512512512512512512 Teacher0.6B0.6B0.6B1.7B4B--- Table 1: F2LLM-v2 model and training configurations. The second stage aims to sharpen the modelâs ability to handle the nuances of diverse downstream applications like classification, reranking, and paraphrase detection. For this stage, we sample at most 80 thousand queries from each data source, producing a mixture of 18 million samples. We apply task-specific instructions to the queries, and also randomly apply instructions to 30% of documents and negatives in tasks where queries and documents are symmetric, including clustering, STS, bitext mining, and paraphrase detection. Pruning and Knowledge Distillation After stage 1 training, we prune the 0.6B model to three smaller sizes along three dimensions: hidden size, MLP intermediate size, and number of layers. For hidden size and MLP intermediate size, we prune the rows and columns in associated weight matrices based on activation norms on a small set of calibration data. For the layer dimension, we simply keep the firstnlayers of the model. We also experimented with pruning layers based on the change of activation norms, but found it to underperform this simple method. After pruning, we find that naive training leads to large performance drops (see Table 4). We mitigate this by applying an additional knowledge distillation loss when training the pruned models, computed by the MSE between the studentâs sequence embedding and a teacher âs sequence embedding over input query, document, and negatives. Ablation experiments suggest that this form of knowledge distillation can also benefit larger models, so we apply it to the 0.6B and 1.7B models in the second training stage as well, while the three largest models are trained without distillation due to resource constraints. All models are trained with AdamW optimizer (Loshchilov & Hutter, 2019). Matryoshka Representa- tion Learning (Kusupati et al., 2022) is applied in both training stages, with a minimum matryoshka dimension of 8. The remaining training hyperparameters are given in Table 1 4 Experiments 4.1 Main Results We evaluate F2LLM-v2 on 17 MTEB benchmarks: Multilingual, English, Code, Medical, Euro- pean, Scandinavian, Indic, German, French, Korean, Polish, Chinese, Japanese, Dutch, Russian, Persian, and Vietnamese, totaling 430 tasks across ten types: retrieval, reranking, classification, 5 F2LLM-v2 Technical Report ModelMulti. (131) English (41) Code (12) Medical (12) European (73) Scan. (28) Indic (20) German (19) French (25) 14B68.74 (6)73.08 (10)80.75 (1)65.20 (2)69.89 (1)71.10 (1)78.85 (1)67.02 (1)72.62 (2) 8B68.09 (8)72.86 (11)80.16 (5)64.91 (4)69.22 (2)69.94 (2)77.93 (2)66.81 (2)72.66 (1) 4B67.06 (10)72.41 (12)80.15 (6)64.48 (7)68.63 (3)68.46 (3)76.58 (3)66.06 (3)71.22 (3) 1.7B65.21 (13)71.63 (16)78.76 (8)61.40 (15)66.90 (4)67.21 (4)74.20 (4)65.08 (4)69.85 (5) 0.6B62.74 (17)69.97 (25)77.41 (10)57.95 (25)64.49 (5)64.32 (6)70.11 (6)63.08 (5)68.14 (7) 330M60.84 (26)68.86 (36)75.74 (13)56.44 (31)62.04 (13)61.93 (11)66.92 (11)61.61 (6)66.03 (13) 160M57.98 (38)65.93 (49)70.38 (18)52.39 (40)59.06 (22)57.79 (25)62.09 (20)57.35 (9)61.90 (20) 80M55.23 (50)64.55 (60)67.97 (22)50.74 (42)56.24 (35)55.54 (30)58.39 (34)55.56 (13)60.30 (22) Results continued for remaining languages and average ModelKorean (6) Polish (17) Chinese (32) Japan. (28) Dutch (40) Russian (23) Persian (52) Viet. (50) Avg. 14B74.85 (3)75.13 (1)68.24 (21)79.32 (1)66.39 (1)70.90 (4)73.55 (1)63.56 (1)71.72 8B75.11 (2)74.61 (2)67.73 (24)78.54 (2)65.81 (2)70.57 (5)72.69 (2)63.32 (2)71.23 4B73.63 (5)73.42 (3)67.12 (27)77.43 (3)64.06 (5)69.46 (7)71.66 (3)62.74 (3)70.27 1.7B73.77 (4)72.03 (4)66.41 (31)75.68 (4)62.73 (8)68.52 (9)70.01 (5)61.54 (4)68.88 0.6B70.88 (7)69.63 (5)64.81 (35)73.07 (6)59.54 (10)65.97 (11)67.98 (7)59.56 (5)66.45 330M68.70 (11)66.53 (6)63.02 (37)70.75 (9)57.57 (12)63.89 (15)66.14 (9)57.46 (6)64.38 160M62.55 (18)62.32 (7)59.88 (40)65.74 (17)52.95 (26)59.71 (27)62.16 (17)52.56 (11)60.16 80M59.98 (21)59.82 (8)58.22 (41)58.80 (19)50.54 (30)57.15 (35)59.98 (20)50.52 (15)57.62 Table 2: Performance of F2LLM-v2 on 17 MTEB benchmarks and their rankings on the leaderboard, accessed on March 19th, 2026 and given in (parentheses). The number of tasks in each benchmark is given in (superscript) . clustering, pair classification, multilabel classification, STS, instruction reranking, bitext mining, and summarization. More details on these benchmarks and tasks are given in Appendix B. The main results are presented in Table 2, along with the modelsâ rankings on the leaderboards. To assess the performance of the smaller models more thoroughly, we also compare specifically with individual models with the same sizes from the Qwen3-Embedding (Zhang et al., 2025a) and EmbeddingGemma (Vera et al., 2025) families in Table 3. As the results of these models are not complete on several benchmarks, we use the results from the leaderboards when available, and evaluate them on the remaining tasks using the same prompts as those used to evaluate F2LLM-v2. ModelMulti. (131) English (41) Code (12) Medical (12) European (73) Scan. (28) Indic (20) German (19) French (25) 0.3B EmbedGemma61.1569.6768.7651.2462.5054.3966.1156.2861.90 F2LLM-v260.8468.8675.7456.4462.0461.9366.9261.6166.03 0.6B Qwen3-Embed64.3470.4775.4260.1663.9160.9966.5359.4563.01 F2LLM-v262.7469.9777.4157.9564.4964.3270.1163.0868.14 Results continued for remaining languages and average ModelKorean (6) Polish (17) Chinese (32) Japan. (28) Dutch (40) Russian (23) Persian (52) Viet. (50) Avg. 0.3B EmbedGemma58.2464.7050.4060.8250.9864.5767.1143.4559.55 F2LLM-v268.7066.5363.4570.7557.5763.8966.1457.4664.41 0.6B Qwen3-Embed65.2967.4266.7167.2854.2764.2062.8856.0164.02 F2LLM-v270.8869.6365.2373.0759.5465.9767.9859.5666.47 Table 3: Comparison of our models with EmbeddingGemma and Qwen3-Embedding. The number of tasks in each benchmark is given in (superscript) . These results highlight the scalability of the F2LLM-v2 models. Our 14B model achieves state-of-the- art on 11 of the 17 evaluated benchmarks, capturing deep semantic nuances suitable for enterprise- grade database systems where model inference does not present a bottleneck. In comparison, the smaller variants - particularly the 80M and 160M models - demonstrate remarkable efficiency, which is achieved without a proportional degradation in performance, verifying the effectiveness of our pruning and knowledge distillation pipeline. Notably, the 330M and 0.6B models consistently outperform Qwen3-Embedding and EmbeddingGemma on most language-specific benchmarks 6 F2LLM-v2 Technical Report 80M160M330M0.6B1.7B w. distillation (F2LLM-v2)58.0460.5364.5566.7269.13 w.o. distillation53.3756.2762.7765.8768.58 Table 4: Ablation results on knowledge distillation, averaged over 350 tasks. 8163264128256512102420484096 Dimension 40 45 50 55 60 65 70 80M 160M 330M 0.6B 1.7B 4B 8B 14B Figure 5: Results of evaluating F2LLM-v2 models at different representation sizes. and the code benchmark, providing an ideal tradeoff between performance and efficiency for edge deployment. 4.2 Ablation Studies For ablation studies, we conduct experiments on a subset of 350 tasks, which are selected based solely on evaluation time to speed up model iterations. We first examine the effectiveness of knowledge distillation. Starting from the same stage-1 check- points, we train another series of models in identical settings as F2LLM-v2, but without knowledge distillation. The results are presented in Table 4, which demonstrates a consistent drop in perfor- mance across all five model scales, verifying the effectiveness of our knowledge distillation method in transferring the capabilities of teacher models into significantly more compact students. We also verify the effectiveness of MRL by evaluating the models at different representation sizes. For each model, we truncate the output embeddings to dimensions ranging from 8 to their full size and measure performance on the ablation task subset. The results, plotted in Figure 5, confirm that MRL successfully concentrates the most critical semantic information in the initial representation dimensions. Performance scales gracefully with the embedding dimension, with the steepest perfor- mance gains occurring at the lower dimensions from 8 to 128 and plateauing as the representation approaches its full size. This demonstrates that the leading dimensions effectively capture the most salient semantic features, while subsequent dimensions add progressively finer-grained detail. These results highlight a crucial tradeoff for practitioners. For example, the 330M model using its full 896-dimensional embedding performs comparably to the much larger 8B and 14B models when their embeddings are truncated to 32 dimensions. This flexibility enables users to dynamically select an optimal balance between performance, inference cost, and storage cost, showcasing the practical utility of MRL for deploying high-quality embeddings across a wide spectrum of hardware constraints. 7 F2LLM-v2 Technical Report 5 Conclusion F2LLM-v2 is the latest member of the Codefuse embedding model family (Liao et al., 2024; Zhang et al., 2025b; Qin et al., 2025). By addressing the current gaps of language imbalance and training opacity in embedding model research, F2LLM-v2 represents a significant step forward in democra- tizing high-performance embedding models. With the release of 8 models along with the complete training recipe and intermediate checkpoints, we hope to facilitate transparency in frontier embed- ding research and contribute to a future with truly global equity in AI technology deployment. References Asma Ben Abacha and Dina Demner-Fushman. A question-entailment approach to question answering. BMC Bioinform., 20(1):511:1â511:23, 2019. doi: 10.1186/S12859-019-3119-4. URL https://doi.org/10.1186/s12859-019-3119-4. Negin Abadani, Jamshid Mozafari, Afsaneh Fatemi, Mohammd Ali Nematbakhsh, and Arefeh Kazemi. ParSQuAD: Machine translated SQuAD dataset for persian question answering. In 2021 7th International Conference on Web Research (ICWR). IEEE, may 2021. David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba O. Alabi, Yanke Mao, Haonan Gao, and En-Shiun Annie Lee. SIB-200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects. In Yvette Graham and Matthew Purver (eds.), Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2024 - Volume 1: Long Papers, St. Julianâs, Malta, March 17-22, 2024, p. 226â245. Association for Computational Linguistics, 2024. URL https://aclanthology.org/2024.eacl-long.14. Eneko Agirre, Daniel M. Cer, Mona T. Diab, and Aitor Gonzalez-Agirre. Semeval-2012 task 6: A pilot on semantic textual similarity. In Eneko Agirre, Johan Bos, and Mona T. Diab (eds.), Proceedings of the 6th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2012, MontrĂ©al, Canada, June 7-8, 2012, p. 385â393. The Association for Computer Linguistics, 2012. URL https://aclanthology.org/S12-1051/. Wasi Uddin Ahmad, Somshubra Majumdar, Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi, Jocelyn Huang, Siddhartha Jain, Vahid Noroozi, and Boris Ginsburg. Opencodereasoning-i: A simple test time scaling approach via self-critique. CoRR, abs/2507.09075, 2025. doi: 10.48550/A RXIV.2507.09075. URL https://doi.org/10.48550/arXiv.2507.09075. Mohammed Altaf. Medical instruction 120k, 2023. URLhttps://huggingface.co/datasets/Moha mmed-Altaf/medical-instruction-120k. Mohammad Yasin Ayoubi, Sajjad & Davoodeh. Persianqa: a dataset for persian question answering. https://github.com/SajjjadAyobi/PersianQA, 2021. Yuelin Bai, Xeron Du, Yiming Liang, Leo Jin, Junting Zhou, Ziqiang Liu, Feiteng Fang, Mingshan Chang, Tianyu Zheng, Xincheng Zhang, Nuo Ma, Zekun Moore Wang, Ruibin Yuan, Haihong Wu, Hongquan Lin, Wenhao Huang, Jiajun Zhang, Chenghua Lin, Jie Fu, Min Yang, Shiwen Ni, and Ge Zhang. COIG-CQIA: quality is all you need for chinese instruction fine-tuning. In Luis Chiruzzo, Alan Ritter, and Lu Wang (eds.), Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, p. 8190â8205. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025.FINDINGS-NAACL.457. URL https://doi.org/10.18653/v1/2025.findings-naacl.457. Dmitry Balobin. Syntetic dataset of translated russian instructions, 2024. URLhttps://huggingfac e.co/datasets/d0rj/ru-instruct. Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel EsplĂ -Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz-Rojas, Leopoldo Pla Sempere, Gema RamĂrez-SĂĄnchez, Elsa SarrĂas, Marek Strelec, Brian Thompson, William Waites, 8 F2LLM-v2 Technical Report Dion Wiggins, and Jaume Zaragoza. Paracrawl: Web-scale acquisition of parallel corpora. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, p. 4555â 4567. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.ACL-MAIN.417. URL https://doi.org/10.18653/v1/2020.acl-main.417. Gianluca Barmina, Nathalie Carmen Hau Norman, Peter Schneider-Kamp, and Lukas Galke Poech. Dala: Danish linguistic acceptability evaluation guided by real world errors. CoRR, abs/2512.04799, 2025. doi: 10.48550/ARXIV.2512.04799. URL https://doi.org/10.48550/arXiv.2512.04799. Pavel Blinov. Medical qa ru data, 2021. URLhttps://huggingface.co/datasets/blinoff/medi cal_qa_ru_data. Luiz Henrique Bonifacio, Israel Campiotti, Roberto A. Lotufo, and Rodrigo Nogueira. mmarco: A multilingual version of MS MARCO passage ranking dataset. CoRR, abs/2108.13897, 2021. URL https://arxiv.org/abs/2108.13897. Vera Boteva, Demian Gholipour Ghalandari, Artem Sokolov, and Stefan Riezler. A full-text learning to rank dataset for medical information retrieval. In Nicola Ferro, Fabio Crestani, Marie-Francine Moens, Josiane Mothe, Fabrizio Silvestri, Giorgio Maria Di Nunzio, Claudia Hauff, and Gianmaria Silvello (eds.), Advances in Information Retrieval - 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20-23, 2016. Proceedings, volume 9626 of Lecture Notes in Computer Science, p. 716â722. Springer, 2016. doi: 10.1007/978-3-319-30671-1\_58. URLhttps://doi.org/10.1007/ 978-3-319-30671-1_58. Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. A large annotated corpus for learning natural language inference. In LluĂs MĂ rquez, Chris Callison-Burch, Jian Su, Daniele Pighin, and Yuval Marton (eds.), Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015, Lisbon, Portugal, September 17-21, 2015, p. 632â642. The Association for Computational Linguistics, 2015. doi: 10.18653/V1/D15-1075. URL https://doi.org/10.18653/v1/d15-1075. Iñigo Casanueva, Tadas Temcinas, Daniela Gerz, Matthew Henderson, and Ivan Vulic. Efficient intent detection with dual sentence encoders. CoRR, abs/2003.04807, 2020. URLhttps://arxiv. org/abs/2003.04807. channelcorp. Komagpie-raw, 2024. URLhttps://huggingface.co/datasets/channelcorp/KoMa gpie-raw. Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. BGE m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. CoRR, abs/2402.03216, 2024. doi: 10.48550/ARXIV.2402.03216. URLhttps: //doi.org/10.48550/arXiv.2402.03216. Xi Chen, Ali Zeynali, Chico Q. Camargo, Fabian Flöck, Devin Gaffney, Przemyslaw A. Grabowicz, Scott Hale, David Jurgens, and Mattia Samory. Semeval-2022 task 8: Multilingual news article similarity. In Guy Emerson, Natalie Schluter, Gabriel Stanovsky, Ritesh Kumar, Alexis Palmer, Nathan Schneider, Siddharth Singh, and Shyam Ratan (eds.), Proceedings of the 16th International Workshop on Semantic Evaluation, SemEval@NAACL 2022, Seattle, Washington, United States, July 14-15, 2022, p. 1094â1106. Association for Computational Linguistics, 2022. doi: 10.18653/V1/20 22.SEMEVAL-1.155. URL https://doi.org/10.18653/v1/2022.semeval-1.155. cjadams, Daniel Borkan, inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and nithum. Jigsaw unintended bias in toxicity classification, 2019. URLhttps://kaggle.com/competition s/jigsaw-unintended-bias-in-toxicity-classification. CLUEbenchmark. Qbqtc, 2021. URL https://github.com/CLUEbenchmark/QBQTC. CLUEbenchmark. Simclue, 2022. URL https://github.com/CLUEbenchmark/SimCLUE. 9 F2LLM-v2 Technical Report Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld. SPECTER: document- level representation learning using citation-informed transformers. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, p. 2270â2282. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.ACL- MAIN.207. URLhttps: //doi.org/10.18653/v1/2020.acl-main.207. ChatKoAlpaca Community. Koalpaca-v1.1a, 2023. URLhttps://huggingface.co/datasets/beom i/KoAlpaca-v1.1a. ChatKoAlpaca Community. Koalpaca-realqa: A korean instruction dataset reflecting real user scenarios, 2024. URL https://huggingface.co/datasets/beomi/KoAlpaca-RealQA. Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. XNLI: evaluating cross-lingual sentence representations. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Junâichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, p. 2475â2485. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18-1269. URL https://doi.org/10.18653/v1/d18-1269. Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco GuzmĂĄn, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Un- supervised cross-lingual representation learning at scale. In Dan Jurafsky, Joyce Chai, Na- talie Schluter, and Joel R. Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Associ- ation for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, p. 8440â8451. Associa- tion for Computational Linguistics, 2020. doi: 10.18653/V1/2020.ACL- MAIN.747. URL https://doi.org/10.18653/v1/2020.acl-main.747. Kasra Darvishi, Newsha Shahbodaghkhan, Zahra Abbasiantaeb, and Saeedeh Momtazi. Pquad: A persian question answering dataset. Comput. Speech Lang., 80:101486, 2023. doi: 10.1016/J.CSL.20 23.101486. URL https://doi.org/10.1016/j.csl.2023.101486. Maxime De Bruyn, Ehsan Lotfi, Jeska Buhmann, and Walter Daelemans. MFAQ: a multilingual FAQ dataset. In Proceedings of the 3rd Workshop on Machine Reading for Question Answering, p. 1â13, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.mrqa-1.1. Den4ikAI. mailruqa-big, 2022. URLhttps://huggingface.co/datasets/Den4ikAI/mailruQA-b ig. Ivan Ramovich Denis Petrov. Russian dataset for instruct/chat models, 2023. URLhttps://huggin gface.co/datasets/SiberiaSoft/SiberianDatasetXL. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), p. 4171â4186. Association for Computational Linguistics, 2019. doi: 10.18653/V1/N19-1423. URL https://doi.org/10.18653/v1/n19-1423. Michael Dinzinger, Laura Caspari, Kanishka Ghosh Dastidar, Jelena Mitrovic, and Michael Gran- itzer. Webfaq: A multilingual collection of natural q&a datasets for dense retrieval. In Nicola Ferro, Maria Maistro, Gabriella Pasi, Omar Alonso, Andrew Trotman, and Suzan Verberne (eds.), Proceedings of the 48th International ACM SIGIR Conference on Research and Development in In- formation Retrieval, SIGIR 2025, Padua, Italy, July 13-18, 2025, p. 3802â3811. ACM, 2025. doi: 10.1145/3726302.3731934. URL https://doi.org/10.1145/3726302.3731934. 10 F2LLM-v2 Technical Report Li Du, Hanyu Zhao, Yiming Ju, and Tengfei Pan. Scaling towards the information boundary of instruction set: Infinityinstruct-subject technical report. CoRR, abs/2507.06968, 2025. doi: 10.48550/ARXIV.2507.06968. URL https://doi.org/10.48550/arXiv.2507.06968. Kenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, MĂĄrton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzeminski, Genta Indra Winata, Saba Sturua, Saiteja Utpala, Mathieu Ciancone, Marion Schaeffer, Diganta Misra, Shreeya Dhakal, Jonathan RystrĂžm, Roman Solomatin, Ămer Veysel Ăagatan, Akash Kundu, and et al. MMTEB: massive multilingual text embedding benchmark. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URLhttps://openreview.net/forum ?id=zl3pfz4VCV. Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. ELI5: long form question answering. In Anna Korhonen, David R. Traum, and LluĂs MĂ rquez (eds.), Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, p. 3558â3567. Association for Computational Linguistics, 2019. doi: 10.18653/V1/P19-1346. URL https://doi.org/10.18653/v1/p19-1346. fenffef. Cmnli, 2024. URL https://huggingface.co/datasets/fenffef/cmnli. Katja Filippova and Yasemin Altun. Overcoming the lack of parallel data in sentence compression. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013, 18-21 October 2013, Grand Hyatt Seattle, Seattle, Washington, USA, A meeting of SIGDAT, a Special Interest Group of the ACL, p. 1481â1491. ACL, 2013. doi: 10.18653/V1/D13-1155. URL https://doi.org/10.18653/v1/d13-1155. Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gökhan TĂŒr, and Prem Natarajan. MASSIVE: A 1m-example multilingual natural language understanding dataset with 51 typologically-diverse languages. In Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, p. 4277â4302. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.ACL-LONG.235. URL https://doi.org/10.18653/v1/2023.acl-long.235. Gregor Geigle, Nils Reimers, Andreas RĂŒcklĂ©, and Iryna Gurevych. TWEAC: transformer with extendable QA agent classifiers. CoRR, abs/2104.07081, 2021. URLhttps://arxiv.org/abs/21 04.07081. Mansi Gupta, Nitish Kulkarni, Raghuveer Chanda, Anirudha Rayasam, and Zachary C. Lipton. Amazonqa: A review-based question answering task. In Sarit Kraus (ed.), Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI 2019, Macao, China, August 10-16, 2019, p. 4996â5002. ijcai.org, 2019. doi: 10.24963/IJCAI.2019/694. URLhttps://doi.org/ 10.24963/ijcai.2019/694. Faegheh Hasibi, Fedor Nikolaev, Chenyan Xiong, Krisztian Balog, Svein Erik Bratsberg, Alexander Kotov, and Jamie Callan. Dbpedia-entity v2: A test collection for entity search. In Noriko Kando, Tetsuya Sakai, Hideo Joho, Hang Li, Arjen P. de Vries, and Ryen W. White (eds.), Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, Shinjuku, Tokyo, Japan, August 7-11, 2017, p. 1265â1268. ACM, 2017. doi: 10.1145/3077136.3080751. URL https://doi.org/10.1145/3077136.3080751. Junqing He, Mingming Fu, and Manshu Tu. Applying deep matching networks to chinese medical question answering: a study and a dataset. BMC Medical Informatics Decis. Mak., 19-S(2):91â100, 2019. doi: 10.1186/S12911-019-0761-8. URL https://doi.org/10.1186/s12911-019-0761-8. Pengcheng He, Jianfeng Gao, and Weizhu Chen. Debertav3: Improving deberta using electra- style pre-training with gradient-disentangled embedding sharing. In The Eleventh International 11 F2LLM-v2 Technical Report Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/forum?id=sE7-XhLxHA. Wei He, Kai Liu, Jing Liu, Yajuan Lyu, Shiqi Zhao, Xinyan Xiao, Yuan Liu, Yizhong Wang, Hua Wu, Qiaoqiao She, Xuan Liu, Tian Wu, and Haifeng Wang. Dureader: a chinese machine reading comprehension dataset from real-world applications. In Eunsol Choi, Minjoon Seo, Danqi Chen, Robin Jia, and Jonathan Berant (eds.), Proceedings of the Workshop on Machine Reading for Question Answering@ACL 2018, Melbourne, Australia, July 19, 2018, p. 37â46. Association for Computational Linguistics, 2018. doi: 10.18653/V1/W18-2605. URL https://aclanthology.org/W18-2605/. Karl Moritz Hermann, TomĂĄs KociskĂœ, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. Teaching machines to read and comprehend. In Corinna Cortes, Neil D. Lawrence, Daniel D. Lee, Masashi Sugiyama, and Roman Garnett (eds.), Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, p. 1693â1701, 2015. URLhttps://proceedings. neurips.c/paper/2015/hash/afdec7005c9f14302cd0474fd0f3c96-Abstract.html. Baotian Hu, Qingcai Chen, and Fangze Zhu. LCSTS: A large scale chinese short text summarization dataset. In LluĂs MĂ rquez, Chris Callison-Burch, Jian Su, Daniele Pighin, and Yuval Marton (eds.), Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015, Lisbon, Portugal, September 17-21, 2015, p. 1967â1972. The Association for Computational Linguistics, 2015. doi: 10.18653/V1/D15-1229. URL https://doi.org/10.18653/v1/d15-1229. Hai Hu, Kyle Richardson, Liang Xu, Lu Li, Sandra KĂŒbler, and Lawrence S. Moss. OCNLI: original chinese natural language inference. In Trevor Cohn, Yulan He, and Yang Liu (eds.), Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, p. 3512â3526. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.FINDINGS-EMNLP.314. URLhttps://doi.org/10.18653/v1/2020.f indings-emnlp.314. Junjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, and Nan Duan. Cosqa: 20, 000+ web queries for code search and question answering. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, p. 5690â5700. Association for Computational Linguistics, 2021. doi: 10.18653/V1/2021.ACL-LONG.442. URL https://doi.org/10.18653/v1/2021.acl-long.442. Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. Code- searchnet challenge: Evaluating the state of semantic code search. CoRR, abs/1909.09436, 2019. URL http://arxiv.org/abs/1909.09436. infgrad. Llm retrieval data, 2024. URLhttps://huggingface.co/datasets/infgrad/retrieval_ data_llm. Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. CoRR, abs/2009.13081, 2020. URL https://arxiv.org/abs/2009.13081. Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. Pubmedqa: A dataset for biomedical research question answering. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP- IJCNLP 2019, Hong Kong, China, November 3-7, 2019, p. 2567â2577. Association for Computational Linguistics, 2019. doi: 10.18653/V1/D19-1259. URL https://doi.org/10.18653/v1/D19-1259. Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Regina Barzilay and Min-Yen Kan 12 F2LLM-v2 Technical Report (eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, p. 1601â1611. Association for Computational Linguistics, 2017. doi: 10.18653/V1/P17-1147. URLhttps://doi.org/10.18653 /v1/P17-1147. kardosdrur. synthetic-nordic-classification, 2025a. URLhttps://huggingface.co/datasets/kard osdrur/synthetic-nordic-classification. kardosdrur. synthetic-nordic-retrieval, 2025b. URLhttps://huggingface.co/datasets/kardosdr ur/synthetic-nordic-retrieval. kardosdrur. synthetic-nordic-sts, 2025c. URLhttps://huggingface.co/datasets/kardosdrur/s ynthetic-nordic-sts. kardosdrur. synthetic-nordic-text_matching, 2025d. URLhttps://huggingface.co/datasets/ka rdosdrur/synthetic-nordic-text_matching. Mohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Do Long, Weishi Wang, Md. Rizwan Parvez, and Shafiq Joty. Xcodeeval: An execution-based large scale multilingual multitask benchmark for code understanding, generation, translation and retrieval. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, p. 6766â6805. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.ACL-LONG.367. URL https://doi.org/10.18653/v1/2024.acl-long.367. Daniel Khashabi, Amos Ng, Tushar Khot, Ashish Sabharwal, Hannaneh Hajishirzi, and Chris Callison-Burch. Gooaq: Open question answering with diverse answer types. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, p. 421â433. Association for Computational Linguistics, 2021. doi: 10.18653/V1/ 2021.FINDINGS-EMNLP.38. URL https://doi.org/10.18653/v1/2021.findings-emnlp.38. Mi-Young Kim, Juliano Rabelo, Randy Goebel, Masaharu Yoshioka, Yoshinobu Kano, and Ken Satoh. COLIEE 2022 summary: Methods for legal document retrieval and entailment. In Yasufumi Takama, Katsutoshi Yada, Ken Satoh, and Sachiyo Arai (eds.), New Frontiers in Artificial Intelligence - JSAI-isAI 2022 Workshop, JURISIN 2022, and JSAI 2022 International Session, Kyoto, Japan, June 12-17, 2022, Revised Selected Papers, volume 13859 of Lecture Notes in Computer Science, p. 51â67. Springer, 2022. doi: 10.1007/978-3-031-29168-5\_4. URLhttps://doi.org/10.1007/978-3-031 -29168-5_4. Philipp Koehn. Europarl: A parallel corpus for statistical machine translation. In Proceedings of Machine Translation Summit X: Papers, MTSummit 2005, Phuket, Thailand, September 13-15, 2005, p. 79â86, 2005. URL https://aclanthology.org/2005.mtsummit-papers.11. Abdullatif Köksal, Marion Thaler, Ayyoob Imani, Ahmet ĂstĂŒn, Anna Korhonen, and Hinrich SchĂŒtze. MURI: high-quality instruction tuning datasets for low-resource languages via reverse instructions. Trans. Assoc. Comput. Linguistics, 13:1032â1055, 2025. doi: 10.1162/TACL.A.18. URL https://doi.org/10.1162/tacl.a.18. Andreas Köpf, Yannic Kilcher, Dimitri von RĂŒtte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, RichĂĄrd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexan- der Mattick. Openassistant conversations - democratizing large language model alignment. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. URL http://papers.nips.c/paper_files/paper/2023/hash/949f0f8f32267d297c2d4e3e10a2 e7e-Abstract-Datasets_and_Benchmarks.html. 13 F2LLM-v2 Technical Report Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanu- jan, William Howard-Snyder, Kaifeng Chen, Sham M. Kakade, Prateek Jain, and Ali Farhadi. Matryoshka representation learning. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Bel- grave, K. Cho, and A. Oh (eds.), Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, 2022. URLhttp://papers.nips.c/paper_files/paper/2022/ hash/c32319f4868da7613d78af9993100e42-Abstract-Conference.html. Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. Natural questions: a benchmark for question answering research. Trans. Assoc. Comput. Linguistics, 7:452â466, 2019. doi: 10.1162/TACL\_A\_00276. URLhttps://doi.org/10.1162/ta cl_a_00276. Ken Lang. Newsweeder: Learning to filter netnews. In Armand Prieditis and Stuart Russell (eds.), Machine Learning, Proceedings of the Twelfth International Conference on Machine Learning, Tahoe City, California, USA, July 9-12, 1995, p. 331â339. Morgan Kaufmann, 1995. doi: 10.1016/B978-1-55860 -377-6.50048-7. URL https://doi.org/10.1016/b978-1-55860-377-6.50048-7. Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. Nv-embed: Improved techniques for training llms as generalist embedding models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025a. URL https://openreview.net/forum?id=lgsyLSsDRe. Jinhyuk Lee, Feiyang Chen, Sahil Dua, Daniel Cer, Madhuri Shanbhogue, Iftekhar Naim, Gus- tavo HernĂĄndez Ăbrego, Zhe Li, Kaifeng Chen, Henrique Schechter Vera, Xiaoqi Ren, Shanfeng Zhang, Daniel Salz, Michael Boratko, Jay Han, Blair Chen, Shuo Huang, Vikram Rao, Paul Suganthan, Feng Han, Andreas Doumanoglou, Nithi Gupta, Fedor Moiseev, Cathy Yip, Aashi Jain, Simon Baumgartner, Shahrokh Shahi, Frank Palma Gomez, Sandeep Mariserla, Min Choi, Parashar Shah, Sonam Goenka, Ke Chen, Ye Xia, Koert Chen, Sai Meher Karthik Duddu, Yichang Chen, Trevor Walker, Wenlei Zhou, Rakesh Ghiya, Zach Gleicher, Karan Gill, Zhe Dong, Mojtaba Seyedhosseini, Yun-Hsuan Sung, Raphael Hoffmann, and Tom Duerig. Gemini embedding: Gener- alizable embeddings from gemini. CoRR, abs/2503.07891, 2025b. doi: 10.48550/ARXIV.2503.07891. URL https://doi.org/10.48550/arXiv.2503.07891. Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich KĂŒttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. PAQ: 65 million probably-asked questions and what you can do with them. Trans. Assoc. Comput. Linguistics, 9:1098â1115, 2021. doi: 10.1162/TACL\_A\_0 0415. URL https://doi.org/10.1162/tacl_a_00415. Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji, and Timothy Baldwin. Bactrian-x : A multilin- gual replicable instruction-following model with low-rank adaptation. CoRR, abs/2305.15011, 2023a. doi: 10.48550/ARXIV.2305.15011. URL https://doi.org/10.48550/arXiv.2305.15011. Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad. MTOP: A comprehensive multilingual task-oriented semantic parsing benchmark. In Paola Merlo, Jörg Tiedemann, and Reut Tsarfaty (eds.), Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021, Online, April 19 - 23, 2021, p. 2950â2962. Association for Computational Linguistics, 2021. doi: 10.18653/V1/2021.EACL-MAI N.257. URL https://doi.org/10.18653/v1/2021.eacl-main.257. Xiangyang Li, Kuicai Dong, Yi Quan Lee, Wei Xia, Hao Zhang, Xinyi Dai, Yasheng Wang, and Ruiming Tang. Coir: A comprehensive benchmark for code information retrieval models. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Pro- ceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: 14 F2LLM-v2 Technical Report Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, p. 22074â22091. Associa- tion for Computational Linguistics, 2025. doi: 10.18653/V1/2025.ACL- LONG.1072. URL https://doi.org/10.18653/v1/2025.acl-long.1072. Yudong Li, Yuqing Zhang, Zhe Zhao, Linlin Shen, Weijie Liu, Weiquan Mao, and Hui Zhang. CSL: A large-scale chinese scientific literature dataset. In Nicoletta Calzolari, Chu-Ren Huang, Hansaem Kim, James Pustejovsky, Leo Wanner, Key-Sun Choi, Pum-Mo Ryu, Hsin-Hsi Chen, Lucia Donatelli, Heng Ji, Sadao Kurohashi, Patrizia Paggio, Nianwen Xue, Seokhwan Kim, Younggyun Hahm, Zhong He, Tony Kyungil Lee, Enrico Santus, Francis Bond, and Seung-Hoon Na (eds.), Proceedings of the 29th International Conference on Computational Linguistics, COLING 2022, Gyeongju, Republic of Korea, October 12-17, 2022, p. 3917â3923. International Committee on Computational Linguistics, 2022. URL https://aclanthology.org/2022.coling-1.344. Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, and You Zhang. Chatdoctor: A medical chat model fine-tuned on llama model using medical domain knowledge. CoRR, abs/2303.14070, 2023b. doi: 10.48550/ARXIV.2303.14070. URL https://doi.org/10.48550/arXiv.2303.14070. Zehan Li, Jianfei Zhang, Chuantao Yin, Yuanxin Ouyang, and Wenge Rong. Procqa: A large- scale community-based programming question answering dataset for code search. In Nicoletta Calzolari, Min-Yen Kan, VĂ©ronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue (eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COLING 2024, 20-25 May, 2024, Torino, Italy, p. 13057â13067. ELRA and ICCL, 2024. URL https://aclanthology.org/2024.lrec-main.1143. Wing Lian, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet Vong, and "Teknium". Openorca: An open dataset of gpt augmented flan reasoning traces.https://https://huggingf ace.co/datasets/Open-Orca/OpenOrca, 2023. Zihan Liao, Hang Yu, Jianguo Li, Jun Wang, and Wei Zhang. D2LLM: decomposed and distilled large language models for semantic search. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, p. 14798â14814. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.ACL- LONG.791. URLhttps: //doi.org/10.18653/v1/2024.acl-long.791. Xin Liu, Qingcai Chen, Chong Deng, Huajun Zeng, Jing Chen, Dongfang Li, and Buzhou Tang. LCQMC: A large-scale chinese question matching corpus. In Emily M. Bender, Leon Derczynski, and Pierre Isabelle (eds.), Proceedings of the 27th International Conference on Computational Linguistics, COLING 2018, Santa Fe, New Mexico, USA, August 20-26, 2018, p. 1952â1962. Association for Computational Linguistics, 2018a. URL https://aclanthology.org/C18-1166/. Xueqing Liu, Chi Wang, Yue Leng, and ChengXiang Zhai. Linkso: a dataset for learning to retrieve similar question answer pairs on software development forums. In Yijun Yu, Erik M. Fredericks, and Premkumar T. Devanbu (eds.), Proceedings of the 4th ACM SIGSOFT International Workshop on NLP for Software Engineering, NL4SE@ESEC/SIGSOFT FSE 2018, Lake Buena Vista, FL, USA, November 4, 2018, p. 2â5. ACM, 2018b. doi: 10.1145/3283812.3283815. URLhttps://doi.org/ 10.1145/3283812.3283815. Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692, 2019. URL http://arxiv.org/abs/1907.11692. Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. Multilingual denoising pre-training for neural machine translation. Trans. Assoc. Comput. Linguistics, 8:726â742, 2020. doi: 10.1162/TACL\_A\_00343. URLhttps: //doi.org/10.1162/tacl_a_00343. 15 F2LLM-v2 Technical Report Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel S. Weld. S2ORC: the semantic scholar open research corpus. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, p. 4969â4983. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.ACL-MAIN.447. URL https://doi.org/10.18653/v1/2020.acl-main.447. Dingkun Long, Qiong Gao, Kuan Zou, Guangwei Xu, Pengjun Xie, Ruijie Guo, Jian Xu, Guanjun Jiang, Luxi Xing, and Ping Yang. Multi-cpr: A multi domain chinese dataset for passage retrieval. In Enrique AmigĂł, Pablo Castells, Julio Gonzalo, Ben Carterette, J. Shane Culpepper, and Gabriella Kazai (eds.), SIGIR â22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022, p. 3046â3056. ACM, 2022. doi: 10.1145/34 77495.3531736. URL https://doi.org/10.1145/3477495.3531736. Shayne Longpre, Yi Lu, and Joachim Daiber. MKQA: A linguistically diverse benchmark for multilingual open domain question answering. Trans. Assoc. Comput. Linguistics, 9:1389â1406, 2021. doi: 10.1162/TACL\_A\_00433. URL https://doi.org/10.1162/tacl_a_00433. Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In 7th International Confer- ence on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7. Ehsan Lotfi, Nikolay Banar, and Walter Daelemans. BEIR-NL: zero-shot information retrieval bench- mark for the dutch language. In Proceedings of the 31st International Conference on Computational Linguistics, COLING 2025 - Workshops, Abu Dhabi, UAE, January 19-24, 2025, p. 36â45. Association for Computational Linguistics, 2025. URL https://aclanthology.org/2025.bucc-1.5/. Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Dekang Lin, Yuji Matsumoto, and Rada Mihalcea (eds.), The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference, 19-24 June, 2011, Portland, Oregon, USA, p. 142â150. The Association for Computer Linguistics, 2011. URLhttps://aclanthology.org/P11 -1015/. Wei Chen Maggie, Phil Culliton. Tweet sentiment extraction, 2020. URLhttps://kaggle.com/com petitions/tweet-sentiment-extraction. Rishabh Maheshwary, Vikas Yadav, Hoang Nguyen, Khyati Mahajan, and Sathwik Tejaswi Mad- husudhan. M2lingual: Enhancing multilingual, multi-turn instruction alignment in large language models. In Luis Chiruzzo, Alan Ritter, and Lu Wang (eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, p. 9676â9713. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025.N AACL-LONG.489. URL https://doi.org/10.18653/v1/2025.naacl-long.489. Macedo Maia, Siegfried Handschuh, AndrĂ© Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. Wwwâ18 open challenge: Financial opinion mining and question answer- ing. In Pierre-Antoine Champin, Fabien Gandon, Mounia Lalmas, and Panagiotis G. Ipeirotis (eds.), Companion of the The Web Conference 2018 on The Web Conference 2018, W 2018, Lyon , France, April 23-27, 2018, p. 1941â1942. ACM, 2018. doi: 10.1145/3184558.3192301. URL https://doi.org/10.1145/3184558.3192301. Somshubra Majumdar, Vahid Noroozi, Sean Narenthiran, Aleksander Ficek, Jagadeesh Balam, and Boris Ginsburg. Genetic instruct: Scaling up synthetic generation of coding instructions for large language models. CoRR, abs/2407.21077, 2024. doi: 10.48550/ARXIV.2407.21077. URL https://doi.org/10.48550/arXiv.2407.21077. Philip May. Machine translated multilingual sts benchmark dataset., 2021. URLhttps://github.c om/PhilipMay/stsb-multi-mt. 16 F2LLM-v2 Technical Report Julian J. McAuley and Jure Leskovec. Hidden factors and hidden topics: understanding rating dimensions with review text. In Qiang Yang, Irwin King, Qing Li, Pearl Pu, and George Karypis (eds.), Seventh ACM Conference on Recommender Systems, RecSys â13, Hong Kong, China, October 12-16, 2013, p. 165â172. ACM, 2013. doi: 10.1145/2507157.2507163. URLhttps://doi.org/10.1 145/2507157.2507163. medalpaca. medical_meadow_medical_flashcards, 2023. URLhttps://huggingface.co/dataset s/medalpaca/medical_meadow_medical_flashcards. Yev Meyer, Marjan Emadi, Dhruv Nathawani, Lipika Ramaswamy, Kendrick Boyd, Maarten Van Seg- broeck, Matthew Grossman, Piotr Mlocek, and Drew Newberry. Synthetic-Text-To-SQL: A syn- thetic dataset for training language models to generate sql queries from natural language prompts, April 2024. URL https://huggingface.co/datasets/gretelai/synthetic-text-to-sql. MonoHime. ru_sentiment_dataset, 2021. URLhttps://huggingface.co/datasets/MonoHime/ru_ sentiment_dataset. Niklas Muennighoff, Nouamane Tazi, LoĂŻc Magne, and Nils Reimers. MTEB: massive text embedding benchmark. In Andreas Vlachos and Isabelle Augenstein (eds.), Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May 2-6, 2023, p. 2006â2029. Association for Computational Linguistics, 2023. doi: 10.18653/V1/ 2023.EACL-MAIN.148. URL https://doi.org/10.18653/v1/2023.eacl-main.148. Niklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela. Generative representational instruction tuning. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/forum?id=BC4lIvfSzv. Shashi Narayan, Shay B. Cohen, and Mirella Lapata. Donât give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Junâichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, p. 1797â 1807. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18- 1206. URL https://doi.org/10.18653/v1/d18-1206. Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. Adversarial NLI: A new benchmark for natural language understanding. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, p. 4885â4901. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.ACL- MAIN.441. URLhttps: //doi.org/10.18653/v1/2020.acl-main.441. Dan Saattrup Nielsen. Scandeval: A benchmark for scandinavian natural language processing. In Tanel AlumĂ€e and Mark Fishel (eds.), Proceedings of the 24th Nordic Conference on Computational Linguistics, NoDaLiDa 2023, TĂłrshavn, Faroe Islands, May 22-24, 2023, p. 185â201. University of Tartu Library, 2023. URL https://aclanthology.org/2023.nodalida-1.20. James OâNeill, Polina Rozenshtein, Ryuichi Kiryo, Motoko Kubota, and Danushka Bollegala. I wish I would have loved this one, but I didnât - A multilingual dataset for counterfactual detection in product review. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, p. 7092â7108. Association for Computational Linguistics, 2021. doi: 10.18653/V1/2021.EMNLP-MAIN.568. URL https://doi.org/10.18653/v1/2021.emnlp-main.568. Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. Medmcqa: A large-scale multi- subject multi-choice dataset for medical domain question answering. In Gerardo Flores, George H. Chen, Tom J. Pollard, Joyce C. Ho, and Tristan Naumann (eds.), Conference on Health, Inference, and 17 F2LLM-v2 Technical Report Learning, CHIL 2022, 7-8 April 2022, Virtual Event, volume 174 of Proceedings of Machine Learning Research, p. 248â260. PMLR, 2022. URL https://proceedings.mlr.press/v174/pal22a.html. Dina Pisarevskaya and Tatiana Shavrina. WikiOmnia: filtration and evaluation of the generated QA corpus on the whole Russian Wikipedia. In Proceedings of the 2nd Workshop on Natural Language Generation, Evaluation, and Metrics (GEM), p. 125â135, Abu Dhabi, United Arab Emirates (Hybrid), dec 2022. Association for Computational Linguistics. URLhttps://aclanthology.org/2022.ge m-1.10. Jin Qin, Zihan Liao, Ziyin Zhang, Hang Yu, Peng Di, and Rui Wang. C2LLM technical report: A new frontier in code retrieval via adaptive cross-attention pooling. CoRR, abs/2512.21332, 2025. doi: 10.48550/ARXIV.2512.21332. URL https://doi.org/10.48550/arXiv.2512.21332. Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100, 000+ questions for machine comprehension of text. In Jian Su, Xavier Carreras, and Kevin Duh (eds.), Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, p. 2383â2392. The Association for Computational Linguistics, 2016. doi: 10.18653/V1/D16-1264. URL https://doi.org/10.18653/v1/d16-1264. Chandan K. Reddy, LluĂs MĂ rquez, Fran Valero, Nikhil Rao, Hugo Zaragoza, Sambaran Bandy- opadhyay, Arnab Biswas, Anlu Xing, and Karthik Subbian. Shopping queries dataset: A large-scale ESCI benchmark for improving product search. CoRR, abs/2206.06588, 2022. doi: 10.48550/ARXIV.2206.06588. URL https://doi.org/10.48550/arXiv.2206.06588. Mobashir Sadat and Cornelia Caragea. Mscinli: A diverse benchmark for scientific natural language inference. In Kevin Duh, Helena GĂłmez-Adorno, and Steven Bethard (eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, p. 1610â1629. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.NAACL -LONG.90. URL https://doi.org/10.18653/v1/2024.naacl-long.90. Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. CARER: con- textualized affect representations for emotion recognition. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Junâichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Meth- ods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, p. 3687â 3697. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18- 1404. URL https://doi.org/10.18653/v1/d18-1404. Alexander G. Sboev, Aleksandr Naumov, and Roman B. Rybka. Data-driven model for emotion detection in russian texts. In Alexei V. Samsonovich and Valentin V. Klimov (eds.), Proceedings of the 2020 Annual International Conference on Brain-Inspired Cognitive Architectures for Artificial Intelligence, BICA 2020, Eleventh Annual Meeting of the BICA Society, November 10-15, 2020, Virtual Event / Natal, Rio Grande do Norte, Brazil, volume 190 of Procedia Computer Science, p. 637â642. Elsevier, 2020. doi: 10.1016/J.PROCS.2021.06.075. URL https://doi.org/10.1016/j.procs.2021.06.075. Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Sta- iano. MLSUM: the multilingual summarization corpus. In Bonnie Webber, Trevor Cohn, Yu- lan He, and Yang Liu (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, p. 8051â8067. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.EMNLP- MAIN.647. URL https://doi.org/10.18653/v1/2020.emnlp-main.647. Lakshay Sharma, Laura Graesser, Nikita Nangia, and Utku Evci. Natural language understanding with the quora question pairs dataset. CoRR, abs/1907.01041, 2019. URLhttp://arxiv.org/ab s/1907.01041. Shivalika Singh, Freddie Vargus, Daniel Dâsouza, Börje Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OâMahony, Mike Zhang, Ramith Het- tiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzeminski, Hakimeh 18 F2LLM-v2 Technical Report Fadaei, Irem ErgĂŒn, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Minh Vu Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet ĂstĂŒn, Marzieh Fadaee, and Sara Hooker. Aya dataset: An open-access collection for multilingual instruction tuning. In Lun-Wei Ku, Andre Mar- tins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, p. 11521â 11567. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.ACL-LONG.620. URL https://doi.org/10.18653/v1/2024.acl-long.620. Maosong Sun, Jingyang Li, Zhipeng Guo, Yu Zhao, Yabin Zheng, Xiance Si, and Zhiyuan Liu. Thuctc: An efficient chinese text classifier, 2016. URL http://thuctc.thunlp.org/. Shuo Sun and Kevin Duh. Clirmatrix: A massively large collection of bilingual and multilin- gual datasets for cross-lingual information retrieval. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, p. 4160â4170. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.EMNLP- MAIN.340. URL https://doi.org/10.18653/v1/2020.emnlp-main.340. Flax Sentence Embeddings Team. Stackexchange title-body pairs, 2021a. URLhttps://huggingfac e.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl. Flax Sentence Embeddings Team. Stack exchange question pairs, 2021b. URLhttps://huggingfac e.co/datasets/flax-sentence-embeddings/. MTEB Team. Arxiv raw data, 2022a. URL https://huggingface.co/datasets/mteb/raw_arxiv. MTEB Team. Biorxiv raw data, 2022b. URLhttps://huggingface.co/datasets/mteb/raw_biorx iv. MTEB Team. Medrxiv raw data, 2022c. URLhttps://huggingface.co/datasets/mteb/raw_med rxiv. Sentence Transformers Team. Embedding training data, 2021c. URLhttps://huggingface.co/dat asets/sentence-transformers/. Sentence Transformers Team. Reddit title-body pairs, 2021d. URLhttps://huggingface.co/datas ets/sentence-transformers/reddit-title-body. James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. FEVER: a large- scale dataset for fact extraction and verification. In Marilyn A. Walker, Heng Ji, and Amanda Stent (eds.), Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), p. 809â819. Association for Computational Linguistics, 2018. doi: 10.18653/V1/N18-1074. URL https://doi.org/10.18653/v1/n18-1074. George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R. Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Poly- chronopoulos, Yannis Almirantis, John Pavlopoulos, Nicolas Baskiotis, Patrick Gallinari, Thierry ArtiĂšres, Axel-Cyrille Ngonga Ngomo, Norman Heino, Ăric Gaussier, Liliana Barrio-Alvers, Michael Schroeder, Ion Androutsopoulos, and Georgios Paliouras. An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition. BMC Bioinform., 16:138:1â138:28, 2015. doi: 10.1186/S12859-015-0564-6. URLhttps://doi.org/10.1186/s12859 -015-0564-6. Ustinian. Lawzhidao, 2020. URLhttps://w.heywhale.com/mw/dataset/5e953ca8e7ec38002d 02fca7. 19 F2LLM-v2 Technical Report Henrique Schechter Vera, Sahil Dua, Biao Zhang, Daniel Salz, Ryan Mullins, Sindhu Raghuram Panyam, Sara Smoot, Iftekhar Naim, Joe Zou, Feiyang Chen, Daniel Cer, Alice Lisak, Min Choi, Lucas Gonzalez, Omar Sanseviero, Glenn Cameron, Ian Ballantyne, Kat Black, Kaifeng Chen, Weiyi Wang, Zhe Li, Gus Martins, Jinhyuk Lee, Mark Sherwood, Ju-yeong Ji, Renjie Wu, Jingxiao Zheng, Jyotinder Singh, Abheesht Sharma, Divyashree Sreepathihalli, Aashi Jain, Adham Elarabawy, AJ Co, Andreas Doumanoglou, Babak Samari, Ben Hora, Brian Potetz, Dahun Kim, Enrique Alfonseca, Fedor Moiseev, Feng Han, Frank Palma Gomez, Gustavo HernĂĄndez Ăbrego, Hesen Zhang, Hui Hui, Jay Han, Karan Gill, Ke Chen, Koert Chen, Madhuri Shanbhogue, Michael Boratko, Paul Suganthan, Sai Meher Karthik Duddu, Sandeep Mariserla, Setareh Ariafar, Shanfeng Zhang, Shijie Zhang, Simon Baumgartner, Sonam Goenka, Steve Qiu, Tanmaya Dabral, Trevor Walker, Vikram Rao, Waleed Khawaja, Wenlei Zhou, Xiaoqi Ren, Ye Xia, Yichang Chen, Yi-Ting Chen, Zhe Dong, Zhongli Ding, Francesco Visin, GaĂ«l Liu, Jiageng Zhang, Kathleen Kenealy, Michelle Casbon, Ravin Kumar, Thomas Mesnard, Zach Gleicher, Cormac Brick, Olivier Lacombe, Adam Roberts, Qin Yin, Yun-Hsuan Sung, Raphael Hoffmann, Tris Warkentin, Armand Joulin, Tom Duerig, and Mojtaba Seyedhosseini. Embeddinggemma: Powerful and lightweight text representations. CoRR, abs/2509.20354, 2025. doi: 10.48550/ARXIV.2509.20354. URLhttps: //doi.org/10.48550/arXiv.2509.20354. Henning Wachsmuth, Shahbaz Syed, and Benno Stein. Retrieval of the best counterargument without prior topic knowledge. In Iryna Gurevych and Yusuke Miyao (eds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, p. 241â251. Association for Computational Linguistics, 2018. doi: 10.18653/V1/P18-1023. URL https://aclanthology.org/P18-1023/. David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. Fact or fiction: Verifying scientific claims. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, p. 7534â7550. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.EMNLP- MAIN.609. URLhttps: //doi.org/10.18653/v1/2020.emnlp-main.609. Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. Improving text embeddings with large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, p. 11897â11916. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.ACL- LONG.642. URLhttps: //doi.org/10.18653/v1/2024.acl-long.642. Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Darrin Eide, Kathryn Funk, Rodney Kinney, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuansan Wang, Chris Wilhelm, Boya Xie, Douglas Raymond, Daniel S. Weld, Oren Etzioni, and Sebastian Kohlmeier. CORD-19: the covid-19 open research dataset. CoRR, abs/2004.10706, 2020. URLhttps://arxiv. org/abs/2004.10706. Xidong Wang, Jianquan Li, Shunian Chen, Yuxuan Zhu, Xiangbo Wu, Zhiyi Zhang, Xiaolong Xu, Junying Chen, Jie Fu, Xiang Wan, Anningzhe Gao, and Benyou Wang. Huatuo-26m, a large-scale chinese medical QA dataset. In Luis Chiruzzo, Alan Ritter, and Lu Wang (eds.), Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, p. 3828â3848. Association for Computational Linguistics, 2025. doi: 10.18653/V1/20 25.FINDINGS-NAACL.211. URL https://doi.org/10.18653/v1/2025.findings-naacl.211. Xiangpeng Wei, Haoran Wei, Huan Lin, Tianhao Li, Pei Zhang, Xingzhang Ren, Mei Li, Yu Wan, Zhiwei Cao, Binbin Xie, Tianxiang Hu, Shangjie Li, Binyuan Hui, Bowen Yu, Dayiheng Liu, Baosong Yang, Fei Huang, and Jun Xie. Polylm: An open source polyglot large language model. CoRR, abs/2307.06018, 2023. doi: 10.48550/ARXIV.2307.06018. URLhttps://doi.org/10.48550 /arXiv.2307.06018. 20 F2LLM-v2 Technical Report Adina Williams, Nikita Nangia, and Samuel R. Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In Marilyn A. Walker, Heng Ji, and Amanda Stent (eds.), Proceedings of the 2018 Conference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), p. 1112â1122. Association for Computational Linguistics, 2018. doi: 10.18653/V1/N18-1101. URL https://doi.org/10.18653/v1/n18-1101. Fei Xia, Bin Li, Yixuan Weng, Shizhu He, Kang Liu, Bin Sun, Shutao Li, and Jun Zhao. Medconqa: Medical conversational question answering system based on knowledge graphs. In Wanxiang Che and Ekaterina Shutova (eds.), Proceedings of the The 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022 - System Demonstrations, Abu Dhabi, UAE, December 7-11, 2022, p. 148â158. Association for Computational Linguistics, 2022. doi: 10.18653/V1/2022.EMN LP-DEMOS.15. URL https://doi.org/10.18653/v1/2022.emnlp-demos.15. Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. C-pack: Packaged resources to advance general chinese embedding. CoRR, abs/2309.07597, 2023. doi: 10.48550/ARXIV.2309.07 597. URL https://doi.org/10.48550/arXiv.2309.07597. Xiaohui Xie, Qian Dong, Bingning Wang, Feiyang Lv, Ting Yao, Weinan Gan, Zhijing Wu, Xiangsheng Li, Haitao Li, Yiqun Liu, and Jin Ma. T2ranking: A large-scale chinese benchmark for passage ranking. In Hsin-Hsi Chen, Wei-Jou (Edward) Duh, Hen-Hsen Huang, Makoto P. Kato, Josiane Mothe, and Barbara Poblete (eds.), Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2023, Taipei, Taiwan, July 23-27, 2023, p. 2681â2690. ACM, 2023. doi: 10.1145/3539618.3591874. URLhttps://doi.org/10.1145/353961 8.3591874. Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, and Zhenzhong Lan. CLUE: A chinese language understanding evaluation benchmark. In Donia Scott, NĂșria Bel, and Chengqing Zong (eds.), Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, p. 4762â 4772. International Committee on Computational Linguistics, 2020. doi: 10.18653/V1/2020.COL ING-MAIN.419. URL https://doi.org/10.18653/v1/2020.coling-main.419. Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-TĂŒr, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, p. 483â498. Association for Computational Linguistics, 2021. doi: 10.18653/V1/2021.NAACL- MAIN.41. URLhttps: //doi.org/10.18653/v1/2021.naacl-main.41. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jian Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu. Qwen3 technical report. CoRR, abs/2505.09388, 2025. doi: 10.48550/ARXIV.2505.09388. URL https://doi.org/10.48550/arXiv.2505.09388. Dongjie Yang, Ruifeng Yuan, Yuantao Fan, , Yifei Yang, Zili Wang, and Shusen Wang. Refgpt: Reference-to-dialogue by gpt and for gpt, 2023. URLhttps://github.com/ziliwangnlp/RefGPT. 21 F2LLM-v2 Technical Report Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, p. 3685â3690. Association for Computational Linguistics, 2019. doi: 10.18653/V1/D19-1382. URL https://doi.org/10.18653/v1/D19-1382. Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Junâichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, p. 2369â2380. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18-1259. URL https://doi.org/10.18653/v1/d18-1259. Weizhe Yuan, Jane Yu, Song Jiang, Karthik Padthe, Yang Li, Dong Wang, Ilia Kulikov, Kyunghyun Cho, Yuandong Tian, Jason E. Weston, and Xian Li. Naturalreasoning: Reasoning in the wild with 2.8m challenging questions. CoRR, abs/2502.13124, 2025. doi: 10.48550/ARXIV.2502.13124. URL https://doi.org/10.48550/arXiv.2502.13124. Sheng Zhang, Xin Zhang, Hui Wang, Lixiang Guo, and Shanshan Liu. Multi-scale attentive interac- tion networks for chinese medical question answer selection. IEEE Access, 6:74061â74071, 2018. doi: 10.1109/ACCESS.2018.2883637. URL https://doi.org/10.1109/ACCESS.2018.2883637. Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. Character-level convolutional networks for text classification. In Corinna Cortes, Neil D. Lawrence, Daniel D. Lee, Masashi Sugiyama, and Roman Garnett (eds.), Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, p. 649â657, 2015. URLhttps://proceedings.neurips.c/paper/2015/hash/250cf8b51c773f3f8dc8b4b e867a9a02-Abstract.html. Xinlu Zhang, Chenxin Tian, Xianjun Yang, Lichang Chen, Zekun Li, and Linda Ruth Petzold. Alpacare: Instruction-tuned large language models for medical application. CoRR, abs/2310.14558, 2023a. doi: 10.48550/ARXIV.2310.14558. URL https://doi.org/10.48550/arXiv.2310.14558. Xinyu Zhang, Xueguang Ma, Peng Shi, and Jimmy Lin. Mr. tydi: A multi-lingual benchmark for dense retrieval. CoRR, abs/2108.08787, 2021. URL https://arxiv.org/abs/2108.08787. Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin. MIRACL: A multilingual retrieval dataset covering 18 diverse languages. Trans. Assoc. Comput. Linguistics, 11:1114â1131, 2023b. doi: 10.1162/TACL\_A\_00595. URL https://doi.org/10.1162/tacl_a_00595. Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou. Qwen3 embedding: Advanc- ing text embedding and reranking through foundation models. CoRR, abs/2506.05176, 2025a. doi: 10.48550/ARXIV.2506.05176. URL https://doi.org/10.48550/arXiv.2506.05176. Ziyin Zhang, Yikang Liu, Weifang Huang, Junyu Mao, Rui Wang, and Hai Hu. MELA: multilingual evaluation of linguistic acceptability. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, p. 2658â2674. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.ACL- LONG.146. URLhttps: //doi.org/10.18653/v1/2024.acl-long.146. Ziyin Zhang, Zihan Liao, Hang Yu, Peng Di, and Rui Wang. F2LLM technical report: Matching SOTA embedding performance with 6 million open-source data. CoRR, abs/2510.02294, 2025b. doi: 10.48550/ARXIV.2510.02294. URL https://doi.org/10.48550/arXiv.2510.02294. 22 F2LLM-v2 Technical Report Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. Wild- chat: 1m chatgpt interaction logs in the wild. In The Twelfth International Conference on Learn- ing Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=Bl8u7ZRlbM. Xinping Zhao, Xinshuo Hu, Zifei Shan, Shouzheng Huang, Yao Zhou, Zetian Sun, Zhenyu Liu, Dongfang Li, Xinyuan Wei, Qian Chen, Youcheng Pan, Yang Xiang, Meishan Zhang, Haofen Wang, Jun Yu, Baotian Hu, and Min Zhang. Kalm-embedding-v2: Superior training techniques and data inspire A versatile embedding model. CoRR, abs/2506.20923, 2025. doi: 10.48550/ARX IV.2506.20923. URL https://doi.org/10.48550/arXiv.2506.20923. Michal Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. The united nations parallel corpus v1.0. In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, HĂ©lĂšne Mazo, AsunciĂłn Moreno, Jan Odijk, and Stelios Piperidis (eds.), Proceedings of the Tenth International Conference on Language Resources and Evaluation LREC 2016, PortoroĆŸ, Slovenia, May 23-28, 2016. European Language Resources Association (ELRA), 2016. URL http://w.lrec-conf.org/proceedings/lrec2016/summaries/1195.html. 23 F2LLM-v2 Technical Report A Training Data Details ISO Code LanguageSamples ISO Code LanguageSamples ISO Code LanguageSamples engEnglish16,059,324msaMalay175,111tatTatar22,327 zhoChinese4,280,372hrvCroatian146,132bosBosnian21,175 rusRussian3,426,943islIcelandic124,495yorYoruba20,139 spaSpanish2,771,907slvSlovenian122,065minMinangkabau19,868 fraFrench2,426,075srpSerbian113,663cheChechen19,518 deuGerman1,749,922urdUrdu113,258arzEgyptian Arabic16,783 araArabic1,402,943benBengali87,787lmoLombard16,575 nldDutch1,395,135azeAzerbaijani82,209argAragonese16,500 vieVietnamese1,159,472afrAfrikaans81,233bakBashkir16,451 hinHindi1,106,611tamTamil78,384somSomali16,369 korKorean1,083,205katGeorgian77,567alsTosk Albanian15,655 jpnJapanese1,082,466telTelugu77,362idoIdo15,613 itaItalian960,595malMalayalam76,518szlSilesian14,845 indIndonesian952,218monMongolian59,851wuuWu Chinese14,762 porPortuguese919,257nnoNorwegian Nynorsk58,298newNepal Bhasa14,714 polPolish885,709kazKazakh55,317chvChuvash13,759 turTurkish694,051 cymWelsh53,951pnbWestern Panjabi13,657 danDanish675,313marMarathi53,803fryWestern Frisian13,649 thaThai654,534sqiAlbanian53,602sndSindhi13,210 sweSwedish577,848nobNorwegian BokmĂ„l52,905oriOriya12,791 fasPersian535,107pusPushto52,753pltPlateau Malagasy12,717 ukrUkrainian471,079mkdMacedonian52,504scnSicilian12,694 cesCzech467,569hbsSerbo-Croatian48,523kurKurdish11,872 norNorwegian424,995 cebCebuano47,408sunSundanese11,793 ellModern Greek378,202javJavanese47,283barBavarian11,193 catCatalan370,156 warWaray (Philippines)45,348yidYiddish10,785 ronRomanian344,845kanKannada44,534ckbCentral Kurdish9,829 finFinnish332,394epoEsperanto44,266faoFaroese9,825 bulBulgarian321,379latLatin43,335inaInterlingua9,782 tglTagalog313,985gujGujarati40,184glaScottish Gaelic9,769 glgGalician301,386uzbUzbek39,951bugBuginese9,662 myaBurmese294,167amhAmharic38,763queQuechua9,406 hyeArmenian288,622ociOccitan37,413bpyBishnupriya9,400 khmKhmer287,530belBelarusian33,330sanSanskrit8,730 nepNepali276,057azbSouth Azerbaijani31,815limLimburgan8,573 hunHungarian271,802 kirKirghiz29,319hauHausa8,435 eusBasque270,551 mlgMalagasy28,661maiMaithili8,180 hebHebrew263,869volVolapĂŒk27,187zsmStandard Malay8,179 laoLao244,750astAsturian26,004iboIgbo8,132 swaSwahili241,497panPanjabi25,096vecVenetian8,121 azjNorth Azerbaijani213,046ltzLuxembourgish25,092iloIloko7,968 lavLatvian212,058ndsLow German24,713asmAssamese7,042 sinSinhala207,903hatHaitian23,940sahYakut7,011 slkSlovak202,049breBreton23,931arbStandard Arabic6,945 tgkTajik200,631 gleIrish23,148snaShona6,933 estEstonian191,063scoScots23,032mltMaltese6,911 litLithuanian184,391 xhoXhosa22,799zulZulu6,669 Table 5: Natural language distribution in the training data of F2LLM-v2 (part1). 24 F2LLM-v2 Technical Report ISO Code LanguageSamples ISO Code LanguageSamples ISO Code LanguageSamples mznMazanderani6,352tsnTswana1,539lvsStandard Latvian800 uigUighur6,190mwlMirandese1,491magMagahi800 ossIron Ossetic5,893divDhivehi1,387mniManipuri800 tukTurkmen5,854 kbpKabiyĂš1,349mosMossi800 aryMoroccan Arabic5,703chmMari1,238nqoNâKo800 wlnWalloon5,408eweEwe1,220nusNuer800 cdoMin Dong Chinese5,175smoSamoan1,175oryOdia800 npiNepali5,156tsoTsonga1,174prsDari800 napNeapolitan4,778fijFijian1,122quyAyacucho Quechua800 aceAchinese4,758 bamBambara1,061satSantali800 mrjWestern Mari4,728linLingala1,046tpiTok Pisin800 xmfMingrelian4,712navNavajo1,028tumTumbuka800 pesIranian Persian4,414 rohRomansh999tzmCentral Atlas Tamazight800 diqDimli4,140sswSwati982umbUmbundu800 apcLevantine Arabic4,079 awaAwadhi954uznNorthern Uzbek800 wolWolof4,068pagPangasinan954yddEastern Yiddish800 pbtSouthern Pashto3,796corCornish938yueYue Chinese800 nsoPedi3,676dzoDzongkha928dyuDyula799 srdSardinian3,614udmUdmurt890luaLuba-Lulua799 banBalinese3,581fonFon883twiTwi798 lijLigurian3,487konKongo873aebTunisian Arabic793 hsbUpper Sorbian3,440 glvManx864kikKikuyu761 acqTaâizzi-Adeni Arabic3,188tirTigrinya851tyvTuvinian626 crhCrimean Tatar2,985 pmsPiemontese842avaAvaric588 mriMaori2,922myvErzya840aymAymara587 eglEmilian2,859sagSango827krcKarachay-Balkar587 arsNajdi Arabic2,712runRundi825fulFulah581 grnGuarani2,705acmMesopotamian Arabic800ormOromo548 nyaChichewa2,607akaAkan800stqSaterfriesisch461 hifFiji Hindi2,566ayrCentral Aymara800lahLahnda450 kasKashmiri2,363 bemBemba800tonTonga391 furFriulian2,284bhoBhojpuri800mdfMoksha314 swhSwahili2,253cjkChokwe800hawHawaiian299 filFilipino2,156 dikSouthwestern Dinka800niaNias297 smeNorthern Sami2,156fuvNigerian Fulfulde800bisBislama272 shnShan2,098gazWest Central Oromo800altSouthern Altai250 sotSouthern Sotho2,089hneChhattisgarhi800srnSranan Tongo204 kinKinyarwanda2,031kabKabyle800venVenda194 lugGanda1,994kacKachin800kbdKabardian172 papPapiamento1,974kamKamba (Kenya)800xalKalmyk122 cosCorsican1,949keaKabuverdianu800dinDinka104 mhrEastern Mari1,633 khkHalh Mongolian800jamJamaican Creole English100 bjnBanjar1,600kmbKimbundu800kalKalaallisut92 kncCentral Kanuri1,600kmrNorthern Kurdish800ikuInuktitut84 taqTamasheq1,600 ltgLatgalian800gucWayuu52 komKomi1,583luoLuo800chrCherokee51 bodTibetan1,563lusLushai800adyAdyghe33 Table 6: Natural language distribution in the training data of F2LLM-v2 (part2). 25 F2LLM-v2 Technical Report LanguageSamplesLanguageSamples python1,972,390css1,003 php553,651typescript888 java483,469 r636 cpp393,514 lisp467 go351,586 jsx436 javascript245,632objective-c327 c#92,008json264 ruby68,317xml258 c43,487yaml180 rust12,924assembly171 kotlin11,284powershell162 sql6,826 vba157 pascal6,299lua126 d5,278 matlab114 haskell4,967 dart107 scala4,120bash105 html2,777http99 shell2,095graphql89 perl2,009svg82 swift1,907vb.net75 ocaml1,894groovy63 csharp1,891 Misc.2,301 Table 7: Programming language distribution in the training data of F2LLM-v2. 26 F2LLM-v2 Technical Report NameLanguageFormatSizeURL Bitext Mining UNPC (Ziemski et al., 2016)6Retrieval2,922,245 huggingface.co/datasets/Helsinki-NLP/un_pc ParaCrawl (Bañón et al., 2020)30Retrieval10,684,184 paracrawl.eu/index.php BactrianX Translation (Li et al., 2023a)52Clustering491,282 huggingface.co/datasets/MBZUAI/Bactrian-X Europarl (Koehn, 2005)21Clustering477,566 huggingface.co/datasets/Helsinki-NLP/europarl Question Answering WebFAQ (Dinzinger et al., 2025)49Retrieval4,368,504 huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval mMARCO (Bonifacio et al., 2021)14Retrieval5,470,174 huggingface.co/datasets/unicamp-dl/mmarco PAQ (Lewis et al., 2021)enRetrieval938,771 huggingface.co/datasets/sentence-transformers/paq SQuAD (Rajpurkar et al., 2016)enRetrieval89,509 huggingface.co/datasets/rajpurkar/squad Stack Exchange (Team, 2021b)enRetrieval754,705 huggingface.co/datasets/flax-sentence-embeddings/stackexchange_titlebody_best_v oted_answer_jsonl Arguana (Wachsmuth et al., 2018)enRetrieval22,848 huggingface.co/datasets/BeIR/arguana-generated-queries Natural Questions (Kwiatkowski et al., 2019)enRetrieval97,209 huggingface.co/datasets/sentence-transformers/natural-questions HotpotQA (Yang et al., 2018)enRetrieval120,528 huggingface.co/datasets/mteb/hotpotqa ELI5 (Fan et al., 2019)enRetrieval161,345 huggingface.co/datasets/Pavithree/eli5 FiQA2018 (Maia et al., 2018)enRetrieval7,452 huggingface.co/datasets/mteb/fiqa BioASQ (Tsatsaronis et al., 2015)enRetrieval125,248 huggingface.co/datasets/BeIR/bioasq-generated-queries NFCorpus (Boteva et al., 2016)enRetrieval1,283 huggingface.co/datasets/mteb/nfcorpus TriviaQA (Joshi et al., 2017)enRetrieval60,025 huggingface.co/datasets/sentence-transformers/trivia-qa-triplet PubMedQA (Jin et al., 2019)enRetrieval60,227 huggingface.co/datasets/qiaojin/PubMedQA Amazon QA (Gupta et al., 2019)enRetrieval59,340 github.com/amazonqa/amazonqa MIRACL (Zhang et al., 2023b)16Retrieval26,740 huggingface.co/datasets/miracl/miracl Mr.TyDi (Zhang et al., 2021)11Retrieval48,619 huggingface.co/datasets/mteb/mrtidy MLDR (Chen et al., 2024)13Retrieval40,264 huggingface.co/datasets/Shitao/MLDR MKQA (Longpre et al., 2021)26Retrieval69,287 huggingface.co/datasets/mteb/MKQARetrieval StackOverflowQA (Li et al., 2025)enRetrieval13,820 huggingface.co/datasets/mteb/stackoverflow-qa ProCQA (Li et al., 2024)11Retrieval485,780 github.com/jordane95/procqa Yahoo_Answers (Zhang et al., 2015)enRetrieval196,645 huggingface.co/datasets/sentence-transformers/yahoo-answers GooAQ (Khashabi et al., 2021)enRetrieval473,876 github.com/allenai/gooaq T2Ranking (Xie et al., 2023)zhRetrieval85,521 huggingface.co/datasets/sentence-transformers/t2ranking DuReader (He et al., 2018)zhRetrieval78,023 huggingface.co/datasets/sentence-transformers/dureader cMedQAv2 (Zhang et al., 2018)zhRetrieval23,105 huggingface.co/datasets/sentence-transformers/cmedqa-v2 Huatuo_kgqa (Wang et al., 2025)zhRetrieval53,835 huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa Huatuo_encqa (Wang et al., 2025)zhRetrieval253,523 huggingface.co/datasets/FreedomIntelligence/huatuo_encyclopedia_qa Multi CPR Medical (Long et al., 2022)zhRetrieval62,085 github.com/Alibaba-NLP/Multi-CPR HealthCareMagic (Li et al., 2023b)enRetrieval78,626 github.com/Kent0n-Li/ChatDoctor MedicalQA_ru (Blinov, 2021)ruRetrieval71,932 huggingface.co/datasets/blinoff/medical_qa_ru_data LLM Retrieval Data (infgrad, 2024)zhRetrieval177,850 huggingface.co/datasets/infgrad/retrieval_data_llm RefGPT (Yang et al., 2023)zhRetrieval184,332 github.com/sufengniu/RefGPT Lawzhidao (Ustinian, 2020)zhRetrieval11,899 heywhale.com/mw/dataset/5e953ca8e7ec38002d02fca7/content MedMCQA (Pal et al., 2022)enRetrieval16,526 huggingface.co/datasets/openlifescienceai/medmcqa CMCQA (Xia et al., 2022)zhRetrieval108,529 github.com/WENGSYX/CMCQA MedQA (Jin et al., 2020)en, zhRetrieval13,458 github.com/jind11/MedQA webMedQA (He et al., 2019)zhRetrieval27,122 github.com/hejunqing/webMedQA MedQuAD (Abacha & Demner-Fushman, 2019)enRetrieval14,268 github.com/abachaa/MedQuAD Medical Flashcards (medalpaca, 2023)enRetrieval33,183 huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards MailruQA (Den4ikAI, 2022)ruRetrieval150,777 huggingface.co/datasets/Den4ikAI/mailruQA-big WikiOmnia (Pisarevskaya & Shavrina, 2022)ruRetrieval462,693 huggingface.co/datasets/RussianNLP/wikiomnia PersianQA (Ayoubi, 2021)faRetrieval6,277 github.com/SajjjadAyobi/PersianQA PQuAD (Darvishi et al., 2023)faRetrieval46,699 github.com/AUT-NLP/PQuAD ParSQuAD (Abadani et al., 2021)faRetrieval41,879 github.com/BigData-IsfahanUni/ParSQuAD MQA (De Bruyn et al., 2021)38Retrieval5,131,895 huggingface.co/datasets/clips/mqa HotpotQA-NL (Lotfi et al., 2025)nlRetrieval81,192 huggingface.co/datasets/clips/beir-nl-hotpotqa Instruction Data Aya (Singh et al., 2024)65Retrieval126,965 huggingface.co/datasets/CohereLabs/aya_dataset MURI (Köksal et al., 2025)194Retrieval720,782 huggingface.co/datasets/akoksal/muri-it OASST2 (Köpf et al., 2023)26Retrieval12,449 huggingface.co/datasets/OpenAssistant/oasst2 MultiAlpaca (Wei et al., 2023)11Retrieval125,447 huggingface.co/datasets/DAMO-NLP-MT/multialpaca WildChat Zhao et al. (2024)76Retrieval638,781 huggingface.co/datasets/allenai/WildChat-4.8M M2Lingual (Maheshwary et al., 2025)75Retrieval158,251 huggingface.co/datasets/ServiceNow-AI/M2Lingual Natural Reasoning (Yuan et al., 2025)enRetrieval845,682 huggingface.co/datasets/facebook/natural_reasoning Infinity Instruct (Du et al., 2025)en, zhRetrieval757,439 huggingface.co/datasets/BAAI/Infinity-Instruct COIG (Bai et al., 2025)zhRetrieval42,415 huggingface.co/datasets/m-a-p/COIG-CQIA Medinstruct (Zhang et al., 2023a)enRetrieval51,539 github.com/XZhang97666/AlpaCare CodeFeedbackST (Li et al., 2025)137Retrieval115,971 huggingface.co/datasets/mteb/codefeedback-st CodeFeedbackMT (Li et al., 2025)pythonRetrieval52,221 huggingface.co/datasets/mteb/codefeedback-mt OpenOrca (Lian et al., 2023)enRetrieval896,450 huggingface.co/datasets/Open-Orca/OpenOrca MEDI2 (Muennighoff et al., 2025)enRetrieval668,036 huggingface.co/datasets/GritLM/MEDI2 MedicalInstruction (Altaf, 2023)enRetrieval75,268 huggingface.co/datasets/Mohammed-Altaf/medical-instruction-120k SiberianDataset (Denis Petrov, 2023)ruRetrieval255,663 huggingface.co/datasets/SiberiaSoft/SiberianDatasetXL Ru Instruct (Balobin, 2024)ruRetrieval452,574 huggingface.co/datasets/d0rj/ru-instruct KoAlpaca-RealQA (Community, 2024)koRetrieval17,599 huggingface.co/datasets/beomi/KoAlpaca-RealQA KoAlpaca (Community, 2023)koRetrieval21,126 huggingface.co/datasets/beomi/KoAlpaca-v1.1a KoMagpie (channelcorp, 2024)koRetrieval428,780 huggingface.co/datasets/channelcorp/KoMagpie-raw Nordic Text Matching (kardosdrur, 2025d)da, sv, noRetrieval182,485 huggingface.co/datasets/kardosdrur/synthetic-nordic-text_matching Nordic Retrieval (kardosdrur, 2025b)da, sv, noRetrieval172,437 huggingface.co/datasets/kardosdrur/synthetic-nordic-retrieval Nordic Classification (kardosdrur, 2025a)da, sv, noClassification199,280 huggingface.co/datasets/kardosdrur/synthetic-nordic-classification Title Matching S2ORC-Title-Abstract (Lo et al., 2020)enRetrieval250,000 huggingface.co/datasets/sentence-transformers/s2orc CORD 19 (Wang et al., 2020)enRetrieval373,674 huggingface.co/datasets/medalpaca/medical_meadow_cord19 Multi CPR ECom (Long et al., 2022)zhRetrieval90,850 github.com/Alibaba-NLP/Multi-CPR ESCI (Reddy et al., 2022)en, ja, esRetrieval80,468 huggingface.co/datasets/tasksource/esci CLIRMatrix (Sun & Duh, 2020)137Retrieval3,275,561 github.com/ssun32/CLIRMatrix DBPedia (Hasibi et al., 2017)enRetrieval288,736 huggingface.co/datasets/BeIR/dbpedia-entity-generated-queries NLI SNLI (Bowman et al., 2015)enRetrieval54,585 huggingface.co/datasets/stanfordnlp/snli MNLI (Williams et al., 2018)enRetrieval112,075 huggingface.co/datasets/nyu-mll/multi_nli ANLI (Nie et al., 2020)enRetrieval18,801 huggingface.co/datasets/facebook/anli XNLI (Conneau et al., 2018)14Retrieval1,400,600 huggingface.co/datasets/mteb/xnli OCNLI (Hu et al., 2020)zhRetrieval6,616 huggingface.co/datasets/dirtycomputer/OCNLI CMNLI (fenffef, 2024)zhRetrieval113,914 huggingface.co/datasets/fenffef/cmnli MSciNLI (Sadat & Caragea, 2024)enRetrieval19,185 huggingface.co/datasets/sadat2307/MSciNLI Table 8: Number of samples in our collected training dataset (part 1). 27 F2LLM-v2 Technical Report NameLanguageFormatSizeURL Topic Classification Arxiv Clustering P2P (Team, 2022a)enClustering83,476 huggingface.co/datasets/mteb/raw_arxiv Arxiv Clustering S2S (Team, 2022a)enClustering83,486 huggingface.co/datasets/mteb/raw_arxiv Biorxiv Clustering P2P (Team, 2022b)enClustering57,296 huggingface.co/datasets/mteb/raw_biorxiv Biorxiv Clustering S2S (Team, 2022b)enClustering57,296 huggingface.co/datasets/mteb/raw_biorxiv Medrxiv Clustering P2P (Team, 2022c)enClustering18,659 huggingface.co/datasets/mteb/raw_medrxiv Medrxiv Clustering S2S (Team, 2022c)enClustering18,659 huggingface.co/datasets/mteb/raw_medrxiv MLSUM Clustering (Scialom et al., 2020)de, es, fr, ruClustering325,739 huggingface.co/datasets/mteb/mlsum TwentyNewsgroups (Lang, 1995)enClustering11,060 huggingface.co/datasets/SetFit/20_newsgroups SIB200ClusteringS2S (Adelani et al., 2024)205Clustering163,302 huggingface.co/datasets/mteb/sib200 Reddit Clustering P2P (Team, 2021d)enClustering80,000 huggingface.co/datasets/sentence-transformers/reddit-title-body Reddit Clustering S2S (Geigle et al., 2021)enClustering58,141 github.com/UKPLab/TWEAC-qa-agent-selection/tree/master/data/reddit/train Stack Exchange Clustering P2P (Team, 2021a)enClustering80,000 huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl Stack Exchange Clustering S2S (Geigle et al., 2021)enClustering56,731 github.com/UKPLab/TWEAC-qa-agent-selection/tree/master/data/stackexchange/train THUCNews (Sun et al., 2016)zhClustering100,000 huggingface.co/datasets/SirlyDreamer/THUCNews TNews (Xu et al., 2020)zhClustering49,726 huggingface.co/datasets/C-MTEB/TNews-classification CSL (Li et al., 2022)zhClustering100,000 huggingface.co/datasets/neuclir/csl Summarization XSum (Narayan et al., 2018)enRetrieval184,383 huggingface.co/datasets/EdinburghNLP/xsum CNN_DM (Hermann et al., 2015)enRetrieval100,000 huggingface.co/datasets/abisee/cnn_dailymail MLSUM Retreival (Scialom et al., 2020)5Retrieval801,159 huggingface.co/datasets/mteb/mlsum Sentence Compression (Filippova & Altun, 2013)enRetrieval175,477 huggingface.co/datasets/sentence-transformers/sentence-compression Text-to-Code OCGI (Majumdar et al., 2024)pythonRetrieval1,052,849 huggingface.co/datasets/nvidia/OpenCodeGeneticInstruct OpenCodeReasoning-2 (Ahmad et al., 2025)python, cppRetrieval16,632 huggingface.co/datasets/nvidia/OpenCodeReasoning-2 xCodeEval NL2Code (Khan et al., 2024)17Retrieval51,072 huggingface.co/datasets/NTU-NLP-sg/xCodeEval CosQA (Huang et al., 2021)pythonRetrieval9,409 huggingface.co/datasets/mteb/cosqa SyntheticText2SQL (Meyer et al., 2024)sqlRetrieval99,617 huggingface.co/datasets/mteb/synthetic-text2sql Code-to-Code xCodeEval Code2Code (Khan et al., 2024)17Retrieval37,056 huggingface.co/datasets/NTU-NLP-sg/xCodeEval xCodeEval Translation (Khan et al., 2024)11Clustering500,000 huggingface.co/datasets/NTU-NLP-sg/xCodeEval CodeSearchNet-ccr (Li et al., 2025)6Retrieval905,195 huggingface.co/datasets/CoIR-Retrieval/CodeSearchNet-ccr Code-to-Text CodeSearchNet (Husain et al., 2019)6Retrieval936,813 huggingface.co/datasets/CoIR-Retrieval/CodeSearchNet Paraphrase Detection StackExchangeDupQuestions-S2S (Team, 2021c)enRetrieval183,559 huggingface.co/datasets/sentence-transformers/stackexchange-duplicates StackExchangeDupQuestions-P2P (Team, 2021c)enRetrieval203,060 huggingface.co/datasets/sentence-transformers/stackexchange-duplicates QQP (Sharma et al., 2019)enRetrieval243,598 gluebenchmark.com/tasks StackOverflowDupQuestions (Liu et al., 2018b)enRetrieval19,847 huggingface.co/datasets/mteb/stackoverflowdupquestions-reranking PawsX (Yang et al., 2019)7Retrieval216,219 huggingface.co/datasets/google-research-datasets/paws-x LCQMC (Liu et al., 2018a)zhRetrieval167,213 huggingface.co/datasets/C-MTEB/LCQMC Sentiment Analysis Amazon Polarity (McAuley & Leskovec, 2013)enClassification100,000 huggingface.co/datasets/mteb/amazon_polarity IMDb (Maas et al., 2011)enClassification24,904 huggingface.co/datasets/mteb/imdb Toxic Conversations (cjadams et al., 2019)enClassification49,900 huggingface.co/datasets/mteb/toxic_conversations_50k Amazon Counterfactual (OâNeill et al., 2021)en, de, jaClassification14,870 huggingface.co/datasets/mteb/amazon_counterfactual Waimai (Xiao et al., 2023)zhClassification7,999 huggingface.co/datasets/C-MTEB/waimai-classification Amazon Reviews (McAuley & Leskovec, 2013)6Clustering600,000 huggingface.co/datasets/mteb/amazon_reviews_multi Emotion (Saravia et al., 2018)enClustering17,944 huggingface.co/datasets/mteb/emotion Tweet Sentiment Extraction (Maggie, 2020)enClustering26,732 huggingface.co/datasets/mteb/tweet_sentiment_extraction RuSentiment (MonoHime, 2021)ruClustering100,000 huggingface.co/datasets/MonoHime/ru_sentiment_dataset CEDR (Sboev et al., 2020)ruClustering4,376 huggingface.co/datasets/sagteam/cedr_v1 Intent Classification Massive Intent (FitzGerald et al., 2023)51Clustering661,923 huggingface.co/datasets/mteb/amazon_massive_intent MTOP Intent (Li et al., 2021)6Clustering83,922 huggingface.co/datasets/mteb/mtop_intent Banking77 (Casanueva et al., 2020)enClustering9,993 huggingface.co/datasets/mteb/banking77 Domain Classification Massive Scenario (FitzGerald et al., 2023)51Clustering661,923 huggingface.co/datasets/mteb/amazon_massive_scenario MTOP Domain (Li et al., 2021)6Clustering83,922 huggingface.co/datasets/mteb/mtop_domain Language Classification BactrianX Language Classification (Li et al., 2023a)52Clustering491,405 huggingface.co/datasets/MBZUAI/Bactrian-X Citation Prediction S2ORC-TItle-Citation (Lo et al., 2020)enRetrieval132,879 huggingface.co/datasets/sentence-transformers/s2orc S2ORC-Abstract-Citation (Lo et al., 2020)enRetrieval231,587 huggingface.co/datasets/sentence-transformers/s2orc SPECTER (Cohan et al., 2020)enRetrieval24,717 huggingface.co/datasets/sentence-transformers/specter Linguistic Acceptability MELA (Zhang et al., 2024)10Classification40,267 huggingface.co/datasets/Geralt-Targaryen/MELA ScaLA (Nielsen, 2023)9Classification128,471 huggingface.co/datasets/alexandrainst/scala DaLA (Barmina et al., 2025)daClassification6,508 huggingface.co/datasets/giannor/dala_large Claim Verification FEVER (Thorne et al., 2018)enRetrieval106,605 huggingface.co/datasets/mteb/fever SciFact (Wadden et al., 2020)enRetrieval859 huggingface.co/datasets/mteb/scifact COLIEE (Kim et al., 2022)enRetrieval454 w.modelscope.cn/datasets/sentence-transformers/coliee FEVER-NL (Lotfi et al., 2025)nlRetrieval94,987 huggingface.co/datasets/clips/beir-nl-fever STS STS12 (Agirre et al., 2012)enRetrieval1,858 huggingface.co/datasets/mteb/sts12-sts STS22 (Chen et al., 2022)enRetrieval389 huggingface.co/datasets/mteb/sts22-crosslingual-sts STSBenchmark (May, 2021)enRetrieval3,297 huggingface.co/datasets/mteb/stsbenchmark-sts STS22-Crosslingual (Chen et al., 2022)7Retrieval1,469 huggingface.co/datasets/mteb/sts22-crosslingual-sts BQ (Xiao et al., 2023)zhRetrieval2,436 huggingface.co/datasets/C-MTEB/BQ QBQTC (CLUEbenchmark, 2021)zhRetrieval37,139 github.com/CLUEbenchmark/QBQTC SimCLUE (CLUEbenchmark, 2022)zhRetrieval213,301 github.com/CLUEbenchmark/SimCLUE LCSTS (Hu et al., 2015)zhRetrieval278,146 huggingface.co/datasets/hugcyp/LCSTS Nordic STS (kardosdrur, 2025c)da, sv, noRetrieval70,617 huggingface.co/datasets/kardosdrur/synthetic-nordic-sts Table 9: Number of samples in our collected training dataset (part 2). 28 F2LLM-v2 Technical Report B Details on MTEB Evaluation The Massive Text Embedding Benchmark (MTEB) is widely recognized as the de facto standard for the comprehensive evaluation of text embedding models. Originally introduced by Muen- nighoff et al. (2023), it was vastly expanded into the Massive Multilingual Text Embedding Bench- mark (MMTEB) through a large-scale, open-science collaboration (Enevoldsen et al., 2025). This community-driven effort has established a rigorous and diverse evaluation framework, encom- passing over 500 quality-controlled tasks that span more than 250 languages and a wide array of domains. The significance of MTEB lies in its unprecedented scale and diversity, which addresses the critical limitations of previous benchmarks that were often constrained to a few languages (mostly English), specific domains (e.g., news), or a single task type (e.g., retrieval). To provide a holistic assessment of a modelâs capabilities, MTEB organizes its evaluation tasks into ten distinct categories: âąRetrieval: Assesses a modelâs ability to find relevant documents from a large corpus for a given query. âąReranking: Measures the ability to reorder a given list of candidate documents by their relevance to a query. âą Classification: Evaluates performance on standard text classification tasks (e.g., sentiment analysis, topic classification). âą Clustering: Tests how well embeddings group semantically similar documents together. âąPair Classification: Involves predicting the relationship between a pair of texts (e.g., para- phrase detection, natural language inference). âąSemantic Textual Similarity (STS): Measures the ability to predict the degree of semantic similarity between two sentences on a continuous scale. âąBitext Mining: Assesses the ability to identify translated sentence pairs from a collection of sentences in two languages. âą Summarization: Evaluates the semantic similarity between a model-generated summary and a reference summary. âą Instruction Reranking: A more challenging reranking variant where the model must follow a detailed natural language instruction to determine relevance. âąMultilabel Classification: A classification variant where each document can be assigned multiple labels. The hundreds of tasks are further organized into benchmarks, which are curated subsets of tasks grouped by language, domain, or a combination of both. This includes language-specific benchmarks such as English, Chinese, and Russian; domain-specific benchmarks such as Code and Medical; and aggregated benchmarks like Multilingual, European, and Scandinavian, which test performance across a broad and diverse set of languages. This hierarchical structure allows for both a fine-grained analysis of a modelâs performance on a specific language or domain and a high-level view of its overall multilingual and multi-domain capabilities. In this work, we leverage the breadth of MTEB to provide a robust and thorough evaluation of our models. We evaluate on 17 benchmarks, totaling 430 unique tasks: Multilingual, Code, Medical, English, Russian, French, German, Polish, Dutch, Indic, Persian, Chinese, Japanese, Korean, Vietnamese, European, and Scandinavian. This extensive evaluation allows for a robust and fine- grained assessment of our modelsâ capabilities, directly supporting our claims of multilingual inclusivity and broad domain competence. 29