Paper deep dive
PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering
Srikar Kashyap Pulipaka
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/24/2026, 5:13:52 AM
Summary
This paper describes the PSK submission to the WMT 2026 Multilingual Instruction Shared Task (MIST). The system utilizes the 3.35B-parameter Tiny Aya Global model equipped with three task-specific QLoRA adapters for summarization, context-based question answering, and open-ended question answering. The authors demonstrate that specialized adapters outperform a multitask adapter, particularly for summarization and context QA, while open QA results vary based on answer length and evaluation metrics. Three system variants were submitted, differing primarily in the open-QA adapter configuration.
Entities (16)
Relation Signals (15)
PSK → submittedto → WMT 2026 MIST
confidence 99% · We describe the PSK submission to the WMT 2026 Multilingual Instruction Shared Task.
PSK → usesmodel → Tiny-aya-global
confidence 98% · Our system uses the 3.35B-parameter Tiny Aya Global model with three QLoRA adapters
PSK → usestechnique → QLoRA
confidence 97% · Our system uses the 3.35B-parameter Tiny Aya Global model with three QLoRA adapters
Tiny-aya-global → hasadapterfortask → summarization
confidence 95% · The adapters are trained on multilingual document-summary pairs
Tiny-aya-global → hasadapterfortask → Context Question Answering
confidence 95% · The adapters are trained on ... passage-based question answering
Tiny-aya-global → hasadapterfortask → Open Question Answering
confidence 95% · The adapters are trained on ... filtered standalone question answering.
Context Question Answering → usesdataset → Belebele
confidence 93% · The context-QA training set combines Belebele
Context Question Answering → usesdataset → TyDi QA
confidence 93% · The context-QA training set combines ... answerable TyDi QA
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We describe the PSK submission to the WMT 2026 Multilingual Instruction Shared Task. Our system uses the 3.35B-parameter Tiny Aya Global model with three QLoRA adapters, one for each task. The adapters are trained on multilingual document-summary pairs, passage-based question answering, and filtered standalone question answering. The summarization data also includes scientific papers with their author-written abstracts. On our held-out split, the context and summarization adapters perform better than our multitask adapter, which was trained only on data supplied by the organizers. Results for open QA are mixed and vary with answer length and evaluation method. We therefore submit three systems with the same context and summarization adapters but different open-QA adapters.
Tags
Links
- Source: https://arxiv.org/abs/2608.20757v1
- Canonical: https://arxiv.org/abs/2608.20757v1
Trouble viewing inline? Open PDF directly →
Full Text
17,820 characters extracted from source content.
Expand or collapse full text
PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering Srikar Kashyap Pulipaka Independent Researcher srikar.kashyap@gmail.com Abstract We describe the PSK submission to the WMT 2026 Multilingual Instruction Shared Task. Our system uses the 3.35B-parameter Tiny Aya Global model with three QLoRA adapters, one for each task. The adapters are trained on mul- tilingual document–summary pairs, passage- based question answering, and filtered stan- dalone question answering. The summarization data also includes scientific papers with their author-written abstracts. On our held-out split, the context and summarization adapters per- form better than our multitask adapter, which was trained only on data supplied by the orga- nizers. Results for open QA are mixed and vary with answer length and evaluation method. We therefore submit three systems with the same context and summarization adapters but differ- ent open-QA adapters. 1 Introduction The WMT 2026 Multilingual Instruction Shared Task (MIST) evaluates models under a 10B- parameter limit on three tasks: context-based ques- tion answering, summarization from a document in languageXto languageY, and open-ended generation (WMT 2026 MIST Organizers, 2026). The test set covers 24 languages and contains both same-language and cross-lingual generation. MIST follows the broader multilingual evaluation effort introduced in WMT 2025, which found substan- tial variation across tasks and languages and also showed that automatic metrics are imperfect for open generation (Kocmi et al., 2025). We use one multilingual backbone with three task-specific adapters. Tiny Aya Global (Sala- manca et al., 2026) is first adapted to the provided data and then continued separately for each task. This design preserves a single 3.35B backbone while allowing the supervision and decoding policy to follow the output structure of each task. 2 Task and System Overview MIST provides a task label at inference time. We use it to select one of three LoRA adapters and the corresponding decoding configuration (Figure 1); there is no language-specific routing. 2.1 Base Model and Initial Adaptation We useCohereLabs/tiny-aya-global, a 3.35B- parameter model covering 70 languages (Sala- manca et al., 2026). Its 8K context window sup- ports the longer task inputs. An initial multitask adapter is trained on 20,293 provided examples us- ing a deterministic 85/15 split; the resulting 3,626 held-out examples are reused for model selection. The task-specific adapters are initialized from this model. 3 Training Data Table 1 summarizes the task-specific training data. Validation examples and exact test prompts are excluded from training. 3.1 Summarization The 12,000-example mixture is divided evenly be- tween general and scientific text. General examples combine provided data with CrossSum (Bhattachar- jee et al., 2023), WikiLingua (Ladhak et al., 2020), and UPDESH (Chitale et al., 2026). The scien- tific portion uses full ACL Anthology papers (Bird et al., 2008) with author-written abstracts. No Lan- guage Left Behind (NLLB) (NLLB Team, 2024) supplies non-English targets, which are filtered for language, length, repetition, and semantic consis- tency. We remove test overlaps and placeholder- heavy papers. We also screen the provided data for likely misalignment by comparing each summary with chunks of its source document and discarding a pair only when both its best LaBSE similarity and best chrF score fall below fixed thresholds. arXiv:2608.20757v1 [cs.CL] 21 Aug 2026 Prompt + task label Task router Summary adapter Context-QA adapter Open-QA adapter Task-specific greedy decoding and repetition control Output Figure 1: The three routes share one frozen Tiny Aya Global backbone. The known task label selects an adapter and its decoding policy; language does not change the route. AdapterRowsContextMain sources and construction Multitask adapter20,2938,192Provided MIST sample training split across all three tasks Summarization12,0008,1926,000 general examples from the provided data, CrossSum, Wik- iLingua, UPDESH, and targeted language pools; 6,000 ACL papers paired with author abstracts Context QA8,5124,096Belebele, answerable TyDi QA, MLQA, MCIF, UPDESH cultural multihop data, and small Czech/Yoruba Aya pools Open QA6,4002,048Provided open-QA examples plus filtered Aya and WMT25 MIST examples; includes Hindi questions and answers translated into Bhojpuri Table 1: Supervised data used by the selected task adapters. Each is initialized from the multitask adapter. 3.2 Context Question Answering The context-QA training set combines Belebele (Bandarkar et al., 2024), answerable TyDi QA (Clark et al., 2020), MLQA (Lewis et al., 2020), MCIF (Papi et al., 2026), UPDESH, and small Czech and Yoruba Aya subsets. MLQA supplies parallel and cross-lingual examples. 3.3 Open-Ended Question Answering Open QA combines provided examples, WMT25 MIST questions, and the Aya Dataset and Col- lection (Singh et al., 2024). We retain short fac- tual and explanatory QA while excluding passage- dependent, multiple-choice, creative-writing, code, translation, and continuation tasks. We also trans- late a small set of Hindi questions and answers into Bhojpuri with NLLB. 4 Training and Inference We use QLoRA (Dettmers et al., 2023) with 4-bit NF4 quantization and BF16 computation. LoRA (Hu et al., 2022) is applied to the attention and feed-forward projections with rank 16, alpha 32, and dropout 0.05. Each adapter is trained for one epoch with effective batch size 16; learning rates are1× 10 −4 for summarization,8× 10 −5 for open QA, and 2× 10 −4 for context QA. Inference uses greedy decoding without sam- pling. After observing repetitive outputs on the de- velopment set, we add repetition penalties of 1.05 for open QA and 1.03 for summarization. We also block repeated 4-grams in summarization outputs. 5 Development Results We report exact match (EM), chrF (Popovi ́ c, 2015), ROUGE-L (Lin, 2004), and LaBSE cosine simi- larity (Feng et al., 2022). For open QA, we use chrF, ROUGE-L, and LaBSE together with manual inspection of validation outputs. 5.1 Task-Level Results Table 2 reports development scores for the evalu- ated adapters. The 8.5k context mixture and the 12k summary mixture are selected for submission. The targeted context continuation is marginally higher on ROUGE-L and LaBSE, whereas the 8.5k context mixture is higher on EM and chrF. We therefore select the latter. 5.2 Scientific Summarization Table 3 reports results on a 336-example scientific validation set. Gold targets are author abstracts; non-English targets are translated and quality fil- tered. The 12k mixture is strongest across all three metrics. Blocking repeated 4-grams removes the ob- served summary loops. 5.3 Open QA and Submitted Systems Best-score QA has the strongest aggregate auto- matic scores. Long-form QA is stronger on a manu- TaskTraining variantnEMchrFROUGE-LLaBSE Summarization Multitask adapter1,6510.0022.970.15710.6681 12k general/scientific mix1,6510.0026.810.17800.6919 Context QA Multitask adapter1,51766.2577.070.68340.8813 8.5k context mixture1,51768.8978.580.69590.8897 +4.8k targeted continuation1,51768.3678.360.69660.8906 Open QA Multitask adapter7943.9028.340.22870.6524 Long-form QA7944.0327.960.23070.6600 Best-score QA7944.1628.410.23290.6638 Table 2: Internal model-selection results on identical held-out rows. EM is a percentage and LaBSE is cosine similarity. ModelchrFROUGE-LLaBSE Multitask adapter13.270.07420.6556 8k summary mix25.280.18120.6959 12k summary mix29.760.20820.7535 Table 3: Results on the scientific summarization valida- tion set (n = 336). The summary mixtures are evenly divided between general and scientific data. SystemSummaryContextOpen PSK-primary12k mix8.5k mixLong-form PSK-variant_A12k mix8.5k mixBest-score PSK-variant_B12k mix8.5k mixMultitask Table 4: Submitted routes. The open-QA variants are named for their selection criterion. ally identified set of long-form prompts and reaches the generation limit less often. Our main submis- sion uses Long-form QA. A second submission uses Best-score QA. Table 4 summarizes the three submitted routes. 6 Analysis Separate adapters improve context QA and sum- marization on our development set. Scientific sum- marization also improves after adding papers and their abstracts to the training data. Open QA has no clear winner. Best-score QA performs better on automatic metrics, while Long-form QA works bet- ter on the longer questions we checked manually and is less likely to hit the output limit. The shared summarization route improves chrF over the multitask adapter in all 28 represented languages. The shared context-QA route improves exact match in 19 of 25 languages, ties in four, and declines in two. Open QA remains mixed: the primary route improves chrF in 9 of 19 languages, and variant A improves it in 10 of 19. Appendix A gives the complete language-level changes and cell sizes. 7 Conclusion We present a routed Tiny Aya Global system with one QLoRA adapter per MIST task. Development results select the 8.5k context mixture and the 12k summary mixture, while the three submissions vary the less stable open-QA route. Official evaluation will provide the appropriate shared-task compari- son. Limitations Reference metrics do not directly measure factual- ity and are particularly limited for open QA. Trans- lated scientific targets may retain artifacts. Results use one backbone and one development split. The router also assumes that the task label is available. Ethics Statement Public and challenge-provided datasets are used under their respective terms. Machine translation may reproduce source and model biases, and open- ended outputs may contain incorrect or harmful claims. The system should not be treated as a fac- tual authority without verification. References Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. The Belebele benchmark: a parallel reading comprehension dataset in 122 lan- guage variants. In Proceedings of the 62nd Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), pages 749–775, Bangkok, Thailand. Association for Computational Linguistics. Abhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ah- mad, Yuan-Fang Li, Yong-Bin Kang, and Rifat Shahriyar. 2023. CrossSum: Beyond English-centric cross-lingual summarization for 1,500+ language pairs. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 2541–2564, Toronto, Canada. Association for Computational Linguistics. Steven Bird, Robert Dale, Bonnie Dorr, Bryan Gibson, Mark Joseph, Min-Yen Kan, Dongwon Lee, Brett Powley, Dragomir Radev, and Yee Fan Tan. 2008. The ACL Anthology reference corpus: A reference dataset for bibliographic research in computational linguistics. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08), Marrakech, Morocco. European Lan- guage Resources Association (ELRA). Pranjal A. Chitale, Varun Gumma, Sanchit Ahuja, Prashant Kodali, Manan Uppadhyay, Deepthi Sud- harsan, and Sunayana Sitaram. 2026. UPDESH: Synthesizing grounded instruction tuning data for 13 Indic languages. In Proceedings of the 64th An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 37997– 38041, San Diego, California, United States. Associ- ation for Computational Linguistics. Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. TyDi QA: A benchmark for information-seeking question answering in typo- logically diverse languages. Transactions of the As- sociation for Computational Linguistics, 8:454–470. Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient fine- tuning of quantized LLMs. In Advances in Neural Information Processing Systems, volume 36, pages 10088–10115, New Orleans, Louisiana, USA. Neural Information Processing Systems Foundation. Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Ari- vazhagan, and Wei Wang. 2022. Language-agnostic BERT sentence embedding. In Proceedings of the 60th Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 878–891, Dublin, Ireland. Association for Computa- tional Linguistics. Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations. Tom Kocmi, Sweta Agrawal, Ekaterina Artemova, Eleft- herios Avramidis, Eleftheria Briakou, Pinzhen Chen, Marzieh Fadaee, Markus Freitag, Roman Grund- kiewicz, Yupeng Hou, Philipp Koehn, Julia Kreutzer, Saab Mansour, Stefano Perrella, Lorenzo Proietti, Parker Riley, Eduardo Sánchez, Patricia Schmidtova, Mariya Shmatova, and Vilém Zouhar. 2025. Findings of the WMT25 multilingual instruction shared task: Persistent hurdles in reasoning, generation, and eval- uation. In Proceedings of the Tenth Conference on Machine Translation, pages 414–435, Suzhou, China. Association for Computational Linguistics. Faisal Ladhak, Esin Durmus, Claire Cardie, and Kath- leen McKeown. 2020. WikiLingua: A new bench- mark dataset for cross-lingual abstractive summariza- tion. In Findings of the Association for Computa- tional Linguistics: EMNLP 2020, pages 4034–4048, Online. Association for Computational Linguistics. Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. MLQA: Evalu- ating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 7315– 7330, Online. Association for Computational Lin- guistics. Chin-Yew Lin. 2004. ROUGE: A package for auto- matic evaluation of summaries. In Text Summariza- tion Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics. NLLB Team. 2024. Scaling neural machine translation to 200 languages. Nature, 630(8018):841–846. Sara Papi, Maike Züfle, Marco Gaido, Beatrice Savoldi, Danni Liu, Ioannis Douros, Luisa Bentivogli, and Jan Niehues. 2026. MCIF: Multimodal crosslin- gual instruction-following benchmark from scientific talks. In The Fourteenth International Conference on Learning Representations. Maja Popovi ́ c. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal. Association for Computational Linguistics. Alejandro R. Salamanca, Diana Abagyan, Daniel D’souza, Ammar Khairi, David Mora, Saurabh Dash, Viraat Aryabumi, Sara Rajaee, Mehrnaz Mofakhami, Ananya Sahu, Thomas Euyang, Brittawnya Prince, Madeline Smith, Hangyu Lin, Acyr Locatelli, Sara Hooker, Tom Kocmi, Aidan Gomez, Ivan Zhang, and 7 others. 2026. Tiny Aya: Bridging scale and multi- lingual depth. Preprint, arXiv:2603.11510. Shivalika Singh, Freddie Vargus, Daniel D’souza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataci- unas, Laura O’Mahony, Mike Zhang, Ramith Het- tiarachchi, Joseph Wilson, Marina Machado, Luisa Moura, Dominik Krzemi ́ nski, Hakimeh Fadaei, Irem Ergun, Ifeoma Okoh, and 14 others. 2024. Aya dataset: An open-access collection for multilingual instruction tuning. In Proceedings of the 62nd An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11521– 11567, Bangkok, Thailand. Association for Compu- tational Linguistics. WMT 2026 MIST Organizers. 2026. Multilingual in- struction shared task.https://w2.statmt.org/ wmt26/multilingual-instruction.html.Ac- cessed 7 August 2026. A Language-Level Changes Tables 5 and 6 report changes from the multitask adapter on identical held-out examples. All submit- ted systems share the summary and context routes. The primary system uses Long-form QA, variant A uses Best-score QA, and variant B retains the mul- titask open-QA route, for which every open-QA change is zero. Context-QA changes are exact- match percentage points; summary and open-QA changes are chrF points. These are descriptive re- sults, and some language cells are small. Languagen S ∆ S n C ∆ C Arabic85+4.39 90+4.44 Bengali66+8.18 20 +10.00 Bhojpuri14 +26.89– Czech14 +12.80– Central Kurdish 14+9.23 45-2.22 German34+5.97 93+3.23 English70+0.87 93+3.23 Finnish14+6.59 45+2.22 French68+1.11 45+0.00 Haitian Creole34+1.86 45+2.22 Hindi85+5.40 45+2.22 Indonesian85+3.26 90+4.44 Italian34+3.83 93+2.15 Japanese82+3.02 90+4.44 Korean79+2.03 90-1.11 Marathi79+5.53 45+0.00 Persian69+5.59 45+6.67 Portuguese84+0.09 45+2.22 Russian85+5.10 90+2.22 Slovak34+8.72 45+2.22 Spanish85+4.17 45+4.44 Swahili53+1.88 45+4.44 Telugu53+0.82 45+0.00 Thai48+2.30 45+2.22 Turkish84+2.45 45+8.89 Vietnamese84+4.07 45+2.22 Yoruba47+2.60– Chinese68+1.93 93+0.00 Table 5: Changes for the shared routes.Sdenotes sum- marization (∆chrF), andCdenotes context QA (∆EM in percentage points). Languagen Primary∆chrF Var. A∆chrF Arabic52+0.22-0.16 Bengali11+0.15-0.72 Czech11+1.53+1.62 German43-2.02-1.34 English11-0.18+1.82 French45-1.89-0.79 Hindi52-1.44+1.75 Indonesian53+0.86+1.50 Italian45+2.94+2.16 Japanese52-1.25-0.84 Korean45-0.43+0.02 Marathi45-4.24-2.62 Portuguese45-0.70-0.68 Russian52-1.65-1.11 Spanish45-2.42-1.57 Turkish45+0.21+1.33 Vietnamese 45+1.62+1.00 Yoruba45+1.87+1.57 Chinese52+1.38+0.03 Table 6: Open-QA chrF changes. The primary system uses Long-form QA; variant A uses Best-score QA. Variant B is the multitask reference and has zero change.