Paper deep dive
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
Zhuolin Yang, Zihan Liu, Yang Chen, Wenliang Dai, Boxin Wang, Sheng-Chieh Lin, Chankyu Lee, Yangyi Chen, Dongfu Jiang, Jiafan He, Renjie Pi, Grace Lam, Nayeon Lee, Alexander Bukharin, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 99%
Last extracted: 3/22/2026, 6:13:56 AM
Summary
Nemotron-Cascade 2 is a 30B MoE LLM with 3B activated parameters, utilizing a Cascade RL framework and multi-domain on-policy distillation to achieve state-of-the-art reasoning and agentic performance, including Gold Medal-level results in IMO 2025, IOI 2025, and ICPC World Finals 2025.
Entities (6)
Relation Signals (3)
Nemotron-Cascade 2 โ achievedperformance โ IMO 2025
confidence 100% ยท Nemotron-Cascade-2-30B-A3B achieves breakthrough performance in mathematical and coding reasoning, securing gold-medal results in both the 2025 International Mathematical Olympiad
Multi-domain On-Policy Distillation โ integratedinto โ Cascade RL
confidence 100% ยท we introduce multi-domain on-policy distillation from the strongest intermediate teacher models for each domain throughout the Cascade RL process
Nemotron-Cascade 2 โ trainedusing โ Cascade RL
confidence 100% ยท we substantially expand Cascade RL to cover a much broader spectrum of reasoning and agentic domains
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We introduce Nemotron-Cascade 2, an open 30B MoE model with 3B activated parameters that delivers best-in-class reasoning and strong agentic capabilities. Despite its compact size, its mathematical and coding reasoning performance approaches that of frontier open models. It is the second open-weight LLM, after DeepSeekV3.2-Speciale-671B-A37B, to achieve Gold Medal-level performance in the 2025 International Mathematical Olympiad (IMO), the International Olympiad in Informatics (IOI), and the ICPC World Finals, demonstrating remarkably high intelligence density with 20x fewer parameters. In contrast to Nemotron-Cascade 1, the key technical advancements are as follows. After SFT on a meticulously curated dataset, we substantially expand Cascade RL to cover a much broader spectrum of reasoning and agentic domains. Furthermore, we introduce multi-domain on-policy distillation from the strongest intermediate teacher models for each domain throughout the Cascade RL process, allowing us to efficiently recover benchmark regressions and sustain strong performance gains along the way. We release the collection of model checkpoint and training data.
Tags
Links
- Source: https://arxiv.org/abs/2603.19220v1
- Canonical: https://arxiv.org/abs/2603.19220v1
Trouble viewing inline? Open PDF directly โ
Full Text
174,587 characters extracted from source content.
Expand or collapse full text
2026-03-16 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Zhuolin Yang โ , Zihan Liu * , Yang Chen * , Wenliang Dai * , Boxin Wang * , Sheng-Chieh Lin, Chankyu Lee, Yangyi Chen, Dongfu Jiang, Jiafan He โก , Renjie Pi, Grace Lam, Nayeon Lee, Alexander Bukharin, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping *โ Abstract We introduce Nemotron-Cascade 2, an open 30B MoE model with 3B activated parameters that delivers best-in- class reasoning and strong agentic capabilities. Despite its compact size, its mathematical and coding reasoning performance approaches that of frontier open models. It is the second open-weight LLM, after DeepSeek- V3.2-Speciale-671B-A37B, to achieve Gold Medal-level performance in the 2025 International Mathematical Olympiad (IMO), the International Olympiad in Informatics (IOI), and the ICPC World Finals, demonstrating remarkably high intelligence density with 20ร fewer parameters. In contrast to Nemotron-Cascade 1, the key technical advancements are as follows. After SFT on a meticulously curated dataset, we substantially expand Cascade RL to cover a much broader spectrum of reasoning and agentic domains. Furthermore, we introduce multi-domain on-policy distillation from the strongest intermediate teacher models for each domain throughout the Cascade RL process, allowing us to efficiently recover benchmark regressions and sustain strong performance gains along the way. We release the collection of model checkpoint and training data. โข Nemotron-Cascade-2-30B-A3B: the post-trained model based on Nemotron-3-Nano-30B-A3B-Base. โข Nemotron-Cascade-2-SFT-Data: collection of SFT datasets for Nemotron-Cascade-2. โข Nemotron-Cascade-2-RL-Data: collection of RL datasets for Nemotron-Cascade-2. LiveCodeBench V6 20 40 60 80 Nemotron-Cascade-2 30B-A3B Nemotron- Cascade-2 30B-A3B 100 88.4 87.2 74.6 Qwen3.5- 35B-A3B 83.6 Qwen3.5- 397B-A17B 85.0 Kimi-K2.5-1T Thinking LiveCodeBench Pro 25Q1 Medium 10 20 30 40 Nemotron- Cascade-2 30B-A3B 50 45.2 39.2 25.6 Qwen3.5- 35B-A3B 44.4 Qwen3.5- 397B-A17B 45.6 Kimi-K2.5-1T Thinking HMMT February 2025 60 70 80 90 Nemotron- Cascade-2 30B-A3B 100 94.694.8 89.0 Qwen3.5- 35B-A3B 95.4 IMO ProofBench 20 40 60 80 Nemotron- Cascade-2 30B-A3B 100 72.9 80.2 DeepSeek Math-V2 671B-A37B 76.7 Gemini Deep Think (IMO Gold) ArenaHard v2 20 40 60 80 Nemotron- Cascade-2 30B-A3B 100 83.5 TODO 65.4 Qwen3.5- 35B-A3B IFBench prompt 20 40 60 80 Nemotron- Cascade-2 30B-A3B 100 82.9 76.5 70.2 Qwen3.5- 35B-A3B SWE Verified OpenHands 20 40 60 80 Nemotron- Cascade-2 30B-A3B 50.2 38.8 Nemotron-3 Nano 30B-A3B 69.2 Qwen3.5- 35B-A3B 60.5 Nemotron-3 Super 120B-A12B Humanity's Last Exam 5 10 15 20 Nemotron- Cascade-2 30B-A3B 25 17.7 10.6 Nemotron-3 Nano 30B-A3B 22.4 Qwen3.5- 35B-A3B 18.3 Nemotron-3 Super 120B-A12B 0 Thinking Thinking + Tool Calling (Python) Qwen3.5- 397B-A17B Kimi-K2.5-1T Thinking 70.2 Qwen3.5- 397B-A17B Qwen3.5- 397B-A17B Kimi-K2.5-1T Thinking LiveCodeBench V6 20 40 60 80 Nemotron- Cascade-2 30B-A3B 100 88.4 74.6 Qwen3.5- 35B-A3B 83.6 Qwen3.5- 397B-A17B 85.0 Kimi-K2.5-1T Thinking LiveCodeBench Pro 25Q1 Medium 10 20 30 40 Nemotron- Cascade-2 30B-A3B 50 45.2 25.6 Qwen3.5- 35B-A3B 44.4 Qwen3.5- 397B-A17B 45.6 Kimi-K2.5-1T Thinking HMMT February 2025 60 70 80 90 Nemotron- Cascade-2 30B-A3B 100 94.694.8 89.0 Qwen3.5- 35B-A3B 95.4 IMO ProofBench 20 40 60 80 Nemotron- Cascade-2 30B-A3B 100 72.9 80.2 DeepSeek Math-V2 671B-A37B 76.7 Gemini Deep Think (IMO Gold) ArenaHard v2 20 40 60 80 Nemotron- Cascade-2 30B-A3B 100 83.5 83.9 65.4 Qwen3.5- 35B-A3B IFBench prompt 20 40 60 80 Nemotron- Cascade-2 30B-A3B 100 82.9 76.5 70.2 Qwen3.5- 35B-A3B SWE Verified OpenHands 20 40 60 80 Nemotron- Cascade-2 30B-A3B 50.2 38.8 Nemotron-3 Nano 30B-A3B 69.2 Qwen3.5- 35B-A3B 60.5 Nemotron-3 Super 120B-A12B Humanity's Last Exam 5 10 15 20 Nemotron- Cascade-2 30B-A3B 25 17.7 10.6 Nemotron-3 Nano 30B-A3B 22.4 Qwen3.5- 35B-A3B 18.3 Nemotron-3 Super 120B-A12B 0 Qwen3.5- 397B-A17B Kimi-K2.5-1T Thinking 70.2 Qwen3.5- 397B-A17B Qwen3.5- 397B-A17B Kimi-K2.5-1T Thinking โ Equal contribution, with authors listed in reverse alphabetical order by first name. โก Reviewed and scored our model-generated solutions for IMO 2025 as a gold medalist at the IMO 2015. โ Leads the effort. ยฉ 2026 NVIDIA. All rights reserved. arXiv:2603.19220v1 [cs.CL] 19 Mar 2026 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Contents 1 Introduction4 2 Main Results4 3 Supervised Fine-Tuning6 3.1 Training Framework . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 3.1.1 Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 3.1.2 Chat Template . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 3.2 SFT Data Curation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 3.2.1 Math . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 3.2.2 Code Reasoning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 3.2.3 Science . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.2.4 Long Context . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.2.5 General Chat . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.2.6 Instruction Following . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.2.7 Safety . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.2.8 Conversational Agent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 3.2.9 Software Engineering Agent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 3.2.10 Terminal Agent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 4 Cascade RL and Multi-Domain On-Policy Distillation9 4.1 Training Framework . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 4.1.1 What determines the ordering of Cascade RL . . . . . . . . . . . . . . . . . . . . . . . . 10 4.1.2 RL Training Configuration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 4.2 Instruction-Following Reinforcement Learning (IF-RL) . . . . . . . . . . . . . . . . . . . . . . 11 4.2.1 Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 4.2.2 Training recipe . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 4.3 Multi-domain RL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 4.4 Multi-domain On-Policy Distillation (MOPD) . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 4.5 Reinforcement Learning from Human Feedback (RLHF) . . . . . . . . . . . . . . . . . . . . . 14 4.5.1 Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 4.5.2 Training recipe . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 4.5.3 Hyper-parameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 4.6 Long-context RL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 4.7 Code RL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 4.7.1 Data Curation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 4.7.2 Training Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 4.8 Software Engineering Reinforcement Learning (SWE RL) . . . . . . . . . . . . . . . . . . . . . 15 4.8.1 Agentless RL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 4.8.2 Execution-based RL for Agentic SWE Scaffold . . . . . . . . . . . . . . . . . . . . . . . 16 5 International Mathematical Olympiad (IMO)16 5.1 IMO 2025 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 5.2 IMO-ProofBench . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 6 Competitive Coding18 6.1 IOI 2025 and ICPC World Finals 2025 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 6.2 Competitive Coding Benchmark Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 2 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 7 Acknowledgments19 A Benchmarks and Evaluation Setups20 A.1 Math . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 A.1.1 Non-proof Math . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 A.1.2 Math Proof . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 A.2 Code Reasoning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 A.3 Knowledge and STEM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 A.4 Alignment and Instruction-Following . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 A.5 Long Context and Context Learning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 A.6 Agentic Tasks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 A.7 Multilingual . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 B Training Hyperparameters24 C Prompt Templates26 C.1 Prompt Templates for Test-Time Scaling on IOI 2025 . . . . . . . . . . . . . . . . . . . . . . . 26 C.2 HLE Judge Prompt . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 D ELO Rating Analysis27 E IMO 2025 Model Solutions30 3 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 1. Introduction Reinforcement Learning (RL) (Guo et al., 2025; Ouyang et al., 2022) has emerged as the cornerstone of LLM post-training, driving advances in reasoning, agentic capabilities, and real-world problem-solving. As models are tasked with increasingly sophisticated requirements, the primary challenge lies in successfully incorporating a broader array of RL environments and very diverse reasoning and agentic tasks. Scaling RL to encompass multifaceted, real-world applications necessitates robust frameworks capable of handling varied reward signals and complex environmental feedback without destabilizing the training process. Our previous work, Nemotron-Cascade 1 (Wang et al., 2025), introduced Cascade RL, a framework that orches- trates sequential, domain-wise RL training across specialized task domains. Cascade RL significantly simplifies the engineering complexity associated with multi-domain RL while achieving state-of-the-art performance across a wide range of benchmarks. The advantages of Cascade RL are threefold. First, domain-specific RL stages are remarkably resistant to catastrophic forgetting. They rarely degrade benchmark performance attained in earlier domains and may even improve it. Second, it allows RL hyperparameters and the training curriculum to be carefully tailored to each specific domain, enabling optimized learning dynamics and improved final performance. Third, task homogeneity within each RL stage also yields substantial compute savings, as response lengths and verification wall-clock times are more uniform within a domain than across multiple domains trained jointly. In this work, we introduce Nemotron-Cascade 2, an open 30B Mixture-of-Experts (MoE) model with 3B activated parameters. Similar to its predecessor, Nemotron-Cascade 2 further scales Cascade RL on high-priority domains to preserve the benefits of domain-wise training, enabling us to push the limits of reasoning performance in key domains to state-of-the-art levels. Furthermore, we incorporate on-policy distillation (Xiao et al., 2026; Zeng et al., 2026) into Cascade RL training stages. By distilling knowledge from the best-performing intermediate teacher models within each specific domain during Cascade RL, this mechanism effectively recovers any benchmark regressions that can occur when training in increasingly complex RL environments. In addition, we integrate multi-domain RL into Cascade RL for groups of tasks with similar response formats and comparable verification costs, allowing them to be trained jointly to scale up for more RL environments and improve training efficiency when cross-task interference is minimal. Our Nemotron-Cascade-2-30B-A3B achieves breakthrough performance in mathematical and coding reasoning, securing gold-medal results in both the 2025 International Mathematical Olympiad (IMO) and the International Olympiad in Informatics (IOI) despite being only a 30B MoE model, 1 while also delivering best-in-class performance across a broad range of benchmarks, including alignment, instruction-following, long context (e.g., 1M context window), and agentic tasks. See Table 1 for the full results. We fully open source the model weights, training data, and methodological details, enabling the research community to reproduce, analyze, and extend the proposed Cascade RL training paradigm. We organize the remainder of this report as follows. Section ยง2 summarizes the main results. Section ยง3 describes the supervised fine-tuning (SFT) with details on data curation. Section ยง4 presents Cascade RL framework intergrated with the multi-domain on-policy distillation. Section ยง5 details the evaluation setup and results on IMO, while Section ยง6 presents the evaluation setup and results on IOI and the ICPC World Finals. 2. Main Results We evaluate Nemotron-Cascade 2 on a comprehensive suite of benchmarks covering mathematical and coding reasoning, knowledge and STEM, alignment and instruction following, long-context understanding and in- context learning, multilingual capabilities, and agentic tasks. The main results are shown in Table 1, and the benchmarks and detailed evaluation setups are described in Appendix A. 1 Our model is the second open-weight LLM, after DeepSeek-V3.2-Speciale-671B-A37B (Liu et al., 2025), to achieve gold-medal performance in both the IMO and IOI. 4 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Table 1: Main results. Nemotron-Cascade-2-30B-A3B achieves gold-medal performance in both the IMO 2025 and IOI 2025, which demonstrate remarkably high intelligence density. โ Numbers in brackets refers to Tool-Integrated Reasoning (TIR) results. โก For the baseline models, we use official numbers when available, otherwise evaluate them using the recommended settings. Benchmark Metric: pass@1 Nemotron-3-Nano 30B-A3B Nemotron-3-Super 120B-A12B Qwen3.5 35B-A3B Nemotron-Cascade-2 30B-A3B Math IMO 2025โ 35 pts IMO AnswerBench70.4 โก 77.2 โก 74.8 โก 79.3 IMO ProofBenchโ72.9 AIME 202589.190.291.9 โก 92.4 (98.6)โ AIME 202689.9 โก 89.8 โก 91.1 โก 90.9 (95.0)โ HMMT Feb2584.6 โก 93.789.0 94.6 Code Reasoning IOI 2025โ348.6 โก 439.28 ICPC World Finals 2025โ10/12 LiveCodeBench v6 (2408-2505)68.378.774.6 87.2 (88.4)โ LiveCodeBenchPro 25Q2 (Easy)54.5 โก 81.7 โก 81.1 โก 87.0 (89.3)โ LiveCodeBenchPro 25Q2 (Med)3.50 โก 23.2 โก 17.8 โก 27.6 (36.8)โ SciCode33.342.138.036.4 Knowledge & STEM MMLU-Reduxโ93.386.3 MMLU-Pro78.383.785.379.8 GPQA-Diamond73.079.284.2 76.1 HLE (no tool)10.618.322.417.7 Alignment & Instruction Following ArenaHard v2 (Avg.)67.7โ65.4 โก 83.5 โ Hard Prompt72.173.964.5 โก 88.2 โ Creative Writing63.2โ66.3 โก 78.7 IFBench (prompt)71.572.670.2 82.9 Scale AI Multi-Challenge38.555.260.0 45.3 Long Context & Context Learning A-LCR35.958.358.5 39.1 LongBench v239.6โ59.040.3 NIAH@1M (RULER Subset)94.898.394.3 โก 99.0 CL-Bench12.0 โก โ15.5 โก 12.2 Agentic BFCL v453.8โ67.3 52.9 ํ 2 -Bench49.061.281.258.9 Terminal Bench 2.08.531.040.521.1 SWE Verified (OpenHands)38.860.569.250.2 Multilingual MMLU-ProX59.579.481.0 72.5 WMT24++ (en -> x)86.286.787.6 โก 84.1 From Table 1, Nemotron-Cascade-2-30B-A3B outperforms both the latest released Qwen3.5-35B-A3B (2026-02- 24) (Qwen Team, 2026) and the larger Nemotron-3-Super-120B-A12B (2026-03-11) (Blakeman et al., 2025), and achieves best-in-class performance across benchmarks in mathematics, code reasoning, alignment, and instruction following. Notably, despite being only a 30B MoE model, Nemotron-Cascade 2 achieves gold-medal 5 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Table 2: Performance of Nemotron-Cascade-2-30B-A3B model on IMO 2025, IOI 2025, and ICPC World Finals 2025 competitions. Nemotron-Cascade-2 model achieved solid gold medal on all these top-tier competitions. Our IMO 2025 solutions are evaluated by human expert (IMO 2015 Gold medalist) while IOI 2025 and ICPCWF 2025 solutions are verified through OnlineJudge with official testcases. Competition P1P2P3 P4P5P6OverallMedal IMO 202577 โ 777035/42Gold IOI 202539 88.53 100 100 28.75 83 439.28/600 Gold CompetitionA B C D E F G H I J K L Overall Medal ICPC World Finals 2025 + - + + + + - + + + + + 10/12 Gold โ For IMO 2025 P2, we use LLM grader with reference solution and marking schema from ProofBench (Ma et al., 2025) due to the extensive analytic geometry approach of the model, which human expert could be hard to verify all the intermediate derivation steps. performance on IMO 2025, IOI 2025 and ICPC World Finals 2025, results previously thought to be attainable only by frontier proprietary models (Gemini Team, 2025) (i.e., Gemini Deep Think) and frontier-sized open models (Liu et al., 2025) (i.e., DeepSeek-V3.2-Speciale-671B-A37B). The detailed performance of our model is reported in Table 2. It underperforms Qwen3.5-35B-A3B primarily on knowledge-intensive and agentic tasks, highlighting the importance of stronger knowledge-intensive pretraining and agentic RL in future work. Nemotron-Cascade-2-30B-A3B also outperforms Nemotron-3-Nano-30B-A3B on nearly all benchmarks, even though both models are post-trained from the same pretrained model, Nemotron-3-Nano-30B-A3B-Base (NVI- IDA, 2025). This result further demonstrates the effectiveness of our Cascade RL plus MOPD training pipeline. 3. Supervised Fine-Tuning In this section, we describe the training framework and data curation process for supervised fine-tuning (SFT), the first stage of our post-training pipeline. This stage equips the model with foundational capabilities, including reasoning, conversational ability, instruction following, and agentic and software engineering skills. 3.1. Training Framework 3.1.1. Overview Our SFT data spans a broad range of domains, including mathematics, coding, science, tool use, agentic tasks, and software engineering, as well as more general domains such as multi-turn dialogue, knowledge-intensive question answering, creative writing, role-playing, safety, and instruction following. We pack all SFT samples into sequences of up to 256K tokens and train the model in a single stage. Empirically, we find that the SFT model reaches optimal performance after approximately 1.5 epochs. The SFT training hyperparameters can be found in Appendix B. 3.1.2. Chat Template Our chat template is depicted in Figure 1. There are two changes to the chat template compared with Nemotron- Cascade (Wang et al., 2025). First, we remove the/thinkand/no_thinktags for simplicity. Second, we prepend an empty <think></think> block to activate the non-thinking mode. For tool calling task, we specify all available tools in the system prompt within the<tools>and</tools> tags, and instruct the model to perform tool calls wrapped within the<tool_call>and</tool_call> tags. 6 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Chat Template <|im_start|>system You are a helpful and harmless assistant. <|im_start|>user Write a paragraph about San Francisco. <|im_end|> <|im_start|>assistant <think></think>San Francisco, located on the northern coast of California, is a vibrant and iconic city known for its stunning Golden Gate Bridge, steep hills, and eclectic neighborhoods like ...... <|im_end|> <|im_start|>user What should I prepare for a summer trip to San Francisco? <|im_end|> <|im_start|>assistant <think> Okay, the user is asking about preparing for a summer trip to San Francisco. They've already seen my description of the city, so they're likely planning their visit. Hmm, summer in SF is famously unpredictable with microclimates and fog, so the key advice should emphasize layers and rain gear. The user might not realize how drastically temperatures can change between neighborhoods or elevations. I should also highlight practical tips they might overlook: comfy shoes for hills, hydration due to dry air, and booking reservations since summer is peak season. ...... </think> Here's what to prepare for a summer trip to San Francisco, focusing on essentials and maximizing your experience: ...... <|im_end|> System Prompt for Tool Calling <|im_start|>system You are a customer service agent that helps the user. The policy that determines how you should respond to requests from users is described below between <policy> and </policy> tags. <policy> ...... </policy> Try to be helpful and always follow the policy. <|im_start|>user Book a hotel in San Francisco. <|im_end|> ...... <|im_start|>assistant You are not allowed to use any tools.<|im_end|> <tools> <function> <name>function-name</name> <description>function-description</description> <parameters> <parameter> <name>parameter-name</name> <type>parameter-type</type> <description>parameter-description</description> <title>parameter-title</title> </parameter> ...... (more parameters to add for this function) </parameters> </function> ...... (more functions to add in the tool list) </tools> # Tools You have access to the following functions: If you choose to call a function, ONLY reply in the following format with NO suffix: <tool_call> <function=example_function_name> <parameter=example_parameter_1> value_1 </parameter> </function> </tool_call><|im_end|> <|im_start|>user Tell me more about it. <|im_end|> <|im_start|>assistant <think></think>A summer trip to San Francisco can be amazing, but the experience is a bit different from typical summer destinations. Here are some additional things that will help you plan better: ...... <|im_end|> Figure 1: (Left) The chat template uses adjacent <think></think> tokens to indicate non-thinking mode, and a single <think> followed by to indicate thinking mode. (Right) For tool calling, the available tools are listed in the system prompt. The model is instructed to call tools within the <tool_call> and </tool_call> tags. 3.2. SFT Data Curation 3.2.1. Math Our non-proof math prompts are primarily sourced from Nemotron-Cascade (Wang et al., 2025) and Nemotron- Math-v2 (Du et al., 2025), from which we collect 1.8M tool calling (i.e., python) samples and 1.9M non-tool samples, with responses generated by DeepSeek-V3.2 and DeepSeek-V3.2-Speciale (Liu et al., 2025), respectively. In addition, we collect 676K samples from the generation-selection category (without tool calling) of Nemotron- 3-Nano (Blakeman et al., 2025), with responses generated by GPT-OSS-120B (Agarwal et al., 2025). In total, the competition math SFT comprises 1.8M tool-calling samples and 2.6M samples without tool use. For mathematical natural language proof, we collect 98K mathematical proof problems from the AOPS split of Nemotron-Math-Proofs-v1 (Du et al., 2025). We generate multiple samples per problem to cover two capabilities including proof generation (410K) and proof verification (400K) using DeepSeek-V3.2-Speciale (Liu et al., 2025), resulting in a total of 816K samples. 3.2.2. Code Reasoning Built on Nemotron-Cascade 1 (Wang et al., 2025), we curate approximately 165K unique coding prompts from several open-source datasets, including OpenCode-Stage2 (Huang et al., 2024), OpenCodeReasoning (Ahmad et al., 2025), and HardTests (He et al., 2025). These prompts are originally sourced from competitive programming platforms such as Codeforces, AtCoder, AIZU, and CodeChef. To encourage prompt diversity and reduce redundancy in our SFT training set, we apply strict deduplication using two methods: (1) sample I/O fingerprinting and (2) n-gram-based text analysis. This process removes approximately 24.2% of self-duplicated coding prompts. 7 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation We choose GPT-OSS-120B (Agarwal et al., 2025) as our SFT teacher model due to its strong code reasoning capabilities. For each coding prompt with verifiable test cases, we apply correctness filtering to the teacherโs reasoning traces, retaining only those that generate correct code. For prompts without verifiable test cases, we generally select longer reasoning traces under the assumption that they reflect more thorough problem analysis. This pipeline yields a final dataset comprising 1.9M Python reasoning traces, 1.0M C++14 reasoning traces, and 1.3M Python tool-calling reasoning traces for competitive coding. Scientific Coding: We further collect scientific research coding prompts spanning the domains of biology, material science, physics, chemistry, and mathematics. The responses to these prompts are generated by GPT-OSS-120B (Agarwal et al., 2025), resulting in a total of 1.1M SFT samples. 3.2.3. Science The science prompts we collect span physics, chemistry, and biology. We use 1.4M science SFT samples from Nemotron-Cascade (Wang et al., 2025) and an additional 1.3M samples from Nemotron-3-Nano (Blakeman et al., 2025). Responses in both datasets are generated by GPT-OSS-120B (Agarwal et al., 2025). 3.2.4. Long Context We adopt the 160K long context SFT data from Nemotron-3-Nano (Blakeman et al., 2025), which has an average sequence length of 128K tokens. In addition, we collect another 74K long context SFT from ChatQA-2 (Xu et al., 2024), which has an average length of 29K tokens. 3.2.5. General Chat We source prompts from Nemotron-Cascade 1 (Wang et al., 2025) and construct 4.9M reasoning-on and 372K reasoning-off samples. Responses for reasoning-on samples are generated by GPT-OSS-120B (Agarwal et al., 2025). For reasoning-off samples, 300K responses are drawn from high-quality annotated short answers within the dataset itself, while an additional 330K are generated by DeepSeek-V3-0324 (Liu et al., 2024) to improve response quality. To enhance multi-turn dialogue capabilities, we synthesize approximately 700K multi-turn conversation samples using two GPT-OSS-120B (Agarwal et al., 2025) instances in a role-playing setup, where one instance plays the user and the other the assistant. The user-side model may terminate the conversation at any point to prevent repetitive exchanges. We additionally incorporate 4.6M reasoning-on chat samples from Nemotron-3-Nano (Blakeman et al., 2025), with prompts drawn from LMSYS (Zheng et al., 2023) and WildChat (Zhao et al., 2024). Responses are generated by GPT-OSS-120B (Agarwal et al., 2025), Qwen3-235B-A22B-Thinking-2507, and Qwen3-235B- A22B-Instruct-2507 (Yang et al., 2025). 3.2.6. Instruction Following We source prompts from Nemotron-Cascade 1 (Wang et al., 2025) and generate approximately 230K reasoning- on responses using GPT-OSS-120B (Agarwal et al., 2025) and 64K reasoning-off responses using DeepSeek- V3-0324 (Liu et al., 2024). In addition, we incorporate 497K instruction-following samples from Nemotron-3- Nano (Blakeman et al., 2025), including 457K reasoning-on and 40K reasoning-off responses. These responses are generated by GPT-OSS-120B (Agarwal et al., 2025), Qwen3-235B-A22B-Thinking-2507, and Qwen3-235B- A22B-Instruct-2507 (Yang et al., 2025). 3.2.7. Safety We collect 4K safety SFT samples from Nemotron-3-Nano (Blakeman et al., 2025) to enable models to exhibit appropriate refusal behavior when encountering unsafe inputs. The SFT prompts are originally sourced from Nemotron Content Safety v2 (Ghosh et al., 2025), Gretel Safety Alignment v1 (gre, 2024), Harmful Tasks (Hasan et al., 2024), and Red-Team-2K (Luo et al., 2024). 8 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 3.2.8. Conversational Agent Aside from the Python tool-use data for math and code reasoning, we further gather tool-use samples in multi-turn conversational settings, where multiple tools are available and the assistant must determine which tools to invoke and how to use them effectively. We collect 822K conversational tool-use samples from Nemotron- 3-Nano (Blakeman et al., 2025), with responses generated by Qwen3-235B-A22B-Thinking-2507, Qwen3-32B, Qwen3-235B-A22B-Instruct-2507 (Yang et al., 2025), and GPT-OSS-120B (Agarwal et al., 2025). 3.2.9. Software Engineering Agent We curate the software engineering (SWE) data using various agentic scaffolds, including OpenHands (Wang et al., 2025), SWE-Agent (Yang et al., 2024), Mini-SWE-Agent, and the agentless scaffold proposed by Wei et al. (2025), to enhance the modelsโ agentic software engineering capabilities. First, we utilize the data from Nemotron 3 Nano (Blakeman et al., 2025) and Super (Blakeman et al., 2025), which includes SWE agentic trajectories generated using Qwen3-Coder-480B-A35B-Instruct (Yang et al., 2025). The problem instances are drawn from SWE-Gym (Pan* et al., 2025), SWE-rebench (Badertdinov et al., 2025), and R2E-Subset (Jain et al., 2025). Second, we employ SWE agentless data from Nemotron-Cascade 1 (Wang et al., 2025), which includes three main tasks: (1) buggy code localization, (2) code repair, and (3) test case generation. Following the established procedure in Wang et al. (2025), we reconstruct the code repair data using DeepSeek-V3.2 (Liu et al., 2025). Our preliminary study shows that incorporating SWE agentless data improves modelsโ effectiveness on SWE agentic tasks. For example, fine-tuning solely on agentic data achieves Pass@1 of 48.9 and Pass@4 of 62.8, whereas fine-tuning on a combination of agentic and agentless data improves performance to Pass@1 of 49.9 and Pass@4 of 65.2 on SWE-bench Verified using OpenHands. Based on this observation, we combine 125K agentic samples and 389K agentless samples as the supervised fine-tuning (SFT) data for SWE tasks. Our models are trained in non-thinking mode on SWE agentic data and in thinking mode on SWE agentless data. 3.2.10. Terminal Agent To enhance agentic capabilities for terminal use, we adopt the Terminal-Task-Gen methodology (Pi et al., 2026) to curate our training tasks. This framework consists of (1) dataset adapters that transform static data into interactive terminal formats, and (2) synthetic tasks generated from both diverse seed prompts and a structured terminal skill taxonomy. Using this framework, we curate 490K samples in total. Specifically, we first adapt 162K math, 32K code, and 32K SWE-specific samples from existing high-quality sources (Wang et al., 2025), which establishes broad foundational coverage. To further improve targeted skill refinement, we synthesize 120K seed-based and 140K skill-based tasks. For trajectory construction, we leverage the tasks curated from above, and employ DeepSeek-V3.2 (Liu et al., 2025) as the core engine to generate step-by-step solution traces via an execution-feedback loop within isolated Docker environments. The Terminus 2 agent framework (Merrill et al., 2026) serves as the underlying scaffolding and tool-use protocol, enabling the model to interact with the terminal and complete complex tasks. 4. Cascade RL and Multi-Domain On-Policy Distillation Following a similar approach to Nemotron-Cascade 1 (Wang et al., 2025), we apply Cascaded Reinforcement Learning (Cascade RL) as our post training pipeline. 4.1. Training Framework We illustrate our training process in Figure 2. In this work, we start the Cascade RL process with IF-RL (ยง4.2) to establish foundational instruction adherence, followed by multi-domain RL (ยง4.3) to enhance the modelโs tool-calling capabilities, STEM reasoning, and response format adherence. We then transition to Multi-domain On-policy Distillation (ยง4.4) to unify specialized expertise into a single, cohesive policy to mitigate performance degradation. We continue with RLHF (ยง4.5) for human alignment, Long-context RL (ยง4.6) to enhance reasoning 9 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Base Model Training Pipeline SFT Instruction-Following RL Multi-domain RL Multi-domain On-policy Distiilation RLHFLong-context RL Code RLSWE RL Nemotron-Cascade 2 3/19/26, 2:01 AMcascade2 file:///C:/Users/boxinw/Downloads/cascade2.drawio.html1/1 Figure 2: Nemotron-Cascade 2 applies Cascade RL with the sequential, domain-wise ordering after SFT, leading to substantial improvements across the corresponding domains. over massive input sequences, Code RL (ยง4.7) for competitive coding problems, and finally SWE RL (ยง4.8) for mastering agentic software interactions. 4.1.1. What determines the ordering of Cascade RL The optimal ordering of stages within a Cascade RL pipeline is not a universal constant; rather, it is a dynamic function of the modelโs underlying behaviors and learning trajectories. In contrast to the original Nemotron Cascade (Wang et al., 2025), our current work Nemotron-Cascade 2 introduces significant improvements in SFT data quality and substantially scales the complexity of the RL environments and tasks. These advancements have fundamentally altered the modelโs behavioral dynamics, which require us to adopt a different order to better accommodate the evolving capabilities of LLMs. Rule of thumb: Mitigating Inter-Domain Interference. Specifically, the rationale for this ordering is primarily driven by the need to mitigate catastrophic forgetting as the model interacts with increasingly diverse envi- ronments. Cascade RL provides a granular lens through which we can observe how specific domains compete or conflict, such as strict instruction adherence in IF-RL versus human preference alignment in RLHF. Our core design principle is to identify an ordering that minimizes negative interference across domains while thoroughly optimizing the highest-priority domains. By identifying which tasks serve as foundational priors and which act as specialized refinements, we can mitigate inter-domain interference. Scaling via Multi-Domain Integration. Following this principle, the Cascade RL pipeline can incorporate multi-domain RL stages when specific domains are found to be non-conflicting or beneficial to the overall performance. This integrated approach is particularly effective as RL environments and datasets grow in complexity, while ensuring that the model maintains a broad performance profile across various benchmarks, as detailed in ยง4.3. Stabilization through On-policy Distillation. Furthermore, We find that Multi-domain On-policy Distillation (ยง4.4) serves as a critical stabilization point in this ordering. It is effective at recovering benchmark performance that may have regressed during earlier, more specialized stages of the cascade RL, leading to a more balanced and robust final policy model. 4.1.2. RL Training Configuration Throughout the entire Cascade RL process, we use Group Relative Policy Optimization (GRPO) algorithm (Shao et al., 2024) with strict on-policy training following Nemotron Cascade (Wang et al., 2025). We adopt on-policy training for improved stability and higher accuracy. We conduct our training using the Nemo-RL repository (NVIDIA, 2025). At each iteration, we generate a group ofํบrollouts from the current policyํ ํ and then perform a single gradient 10 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation update. This ensures that the policy used for data collection always matches the one being updated, making the importance sampling ratio exactly 1. This on-policy setup contributes to stable RL training and mitigates entropy collapse. In addition, we remove KL divergence term entirely, which simplifies the GRPO objective to the standard REINFORCE objective (Williams, 1992) with group-normalized rewards and token-level loss (Yu et al., 2025): ํฅ GRPO (ํ) = E (ํ,ํ)โผํ, ํ ํ ํบ ํ=1 โผํ ํ (ยท|ํ) [๏ธ 1 โ๏ธ ํบ ํ=1 |ํ ํ | ํบ โ๏ธ ํ=1 |ํ ํ | โ๏ธ ํก=1 ห ํด ํ,ํก ]๏ธ , where ห ํด ํ,ํก = ํ ํ โ mean(ํ ํ ํบ ํ=1 ) std(ํ ํ ํบ ํ=1 ) for all ํก, (1) andํ ํ ํบ ํ=1 denotes the group of G rewards assigned to the sampled responsesํ ํบ ํ=1 for a given questionํ drawn from the datasetํ, verified against the ground-truth answerํin RLVR. For RLHF,ํ ํ is the aggregated reward score from the generative reward model for responseํ ํ and questionํ. Details of the reward functions for different domains will be provided in the corresponding subsections. 4.2. Instruction-Following Reinforcement Learning (IF-RL) In this subsection, we describe our instruction-following RL recipe, which serves as the first stage of our Cascade RL. We demonstrate that applying verifiable IF-RL significantly improves instruction adherence, achieving a state-of-the-art accuracy of 83.13% on IFBench (Pyatkin et al., 2025). 4.2.1. Dataset We use the same instruction-following training data used for NVIDIA Nano-v3 post-training (Blakeman et al., 2025). The instructions in this dataset are designed for objective verifiability, for instance, requiring a response to be under 200 words. This making the dataset well-suited for training and evaluating models on strict adherence. Given the high baseline quality of the data, our curation process mainly resolves formatting inconsistencies within the keyword arguments for certain instruction types (e.g.,count_increment_word). 4.2.2. Training recipe Following (Wang et al., 2025), we also apply dynamic filtering (Yu et al., 2025). This technique filters out samples where all rollouts are either entirely correct or entirely incorrect. By ensuring that every prompt in a batch provides effective gradients, dynamic filtering stabilizes IF-RL training and pushes the upper bound of model performance. Furthermore, we observed that extended IF-RL training can lead to excessive token usage, which is often unnecessary for fulfilling specific constraints in general chat domains. To mitigate this, we apply overlong penalty, which penalizes samples that fail to complete generation within the maximum sequence length with a zero reward. Unlike Nemotron Cascade (Wang et al., 2025), we position IF-RL as the first stage of our Cascade RL training for two primary reasons: (i) IF-RL can negatively impact human alignment capabilities (e.g., ArenaHard), while our subsequent generative-reward-model-based RLHF has a negligible impact on instruction following scores. By prioritizing instruction adherence first, we can focus on maximizing instruction following performance and then utilize the later stages to recover and refine human preference alignment. (i) An early IF-RL stage produces a model with superior instruction-following capabilities, which serves as a strong teacher for subsequent multi-domain on-policy distillation. Another difference from Nemotron Cascade (Wang et al., 2025) is that our IF-RL is trained exclusively in โthinking modeโ without incorporating a reward model. We found that the โthinking modeโ yields higher accuracy on instruction-following benchmarks (e.g., IFBench (Pyatkin et al., 2025)). Because subsequent RL stages recover any regressions in human preference alignment introduced during IF-RL, we can focus entirely on maximizing instruction adherence without incurring the computational overhead of an auxiliary reward model. We use a batch size of 128, sampling 16 responses per prompt with temperature 1.0 and top-p 1.0. We adopt a learning rate of 2e-6 with AdamW (Kingma, 2014), and set both the entropy loss coefficient and KL loss 11 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation coefficient to 0. Our IF-RL with dynamic filtering takes around 180 steps. The full set of hyperparameters is provided in Appendix B. 4.3. Multi-domain RL Following IF-RL, we conduct an additional stage of multi-domain RL that covers three capabilities: multi-choice question answering (MCQA) in the STEM domain, agentic tool calling, and structured output for instruction following. The datasets are drawn from the NVIDIA Nano-v3 RL training blend (Blakeman et al., 2025). The data mixture consists of approximately 55% MCQA, 30% agentic tool calling using the Workplace Assistant setup (Blakeman et al., 2025), and 15% structured output. We group these domains into a single multi-domain RL stage for two main reasons. First, we do not observe performance degradation across evaluation benchmarks when training on the blended domains. Instead, the model exhibits consistent improvements on benchmarks including MMLU-Pro,ํ 2 -Bench, and IF-Bench. Second, the response lengths and verification times of these datasets are similar, which minimizes training inefficiencies caused by waiting for longer generations or slower environment verification. During training, we use a batch size of 128 and sample 16 responses per prompt with temperature 1.0 and top-p 1.0 (see Appendix B). We adopt a learning rate of3ร 10 โ6 with AdamW (Kingma, 2014), and set both the entropy loss coefficient and KL loss coefficient to zero. This multi-domain RL stage runs for approximately 70 training steps. 4.4. Multi-domain On-Policy Distillation (MOPD) While well-designed Cascade RL substantially reduces catastrophic forgetting compared with vanilla sequential RL in an arbitrary order, it does not fully eliminate capability drift as the number of training environments increases. In practice, we observe noticeable fluctuations across different benchmark categories tracked throughout training, and the dominant trade-offs differ by stage. For example, certain RLVR training often reduces model entropy and shortens reasoning traces, thus can negatively impact mathematical reasoning performance, while RLHF-oriented optimization can partially trade off against instruction-following behavior. These observations motivate an additional training stage for re-balancing capabilities within the Cascade RL process. We therefore adopt multi-domain on-policy distillation (MOPD) (Agarwal et al., 2024; Lu and Lab, 2025; Xiao et al., 2026; Yang et al., 2025; Zeng et al., 2026) as a complementary post-training stage. In our setting, MOPD is particularly attractive for three reasons. First, teacher checkpoints can be selected directly from the Cascade RL pipeline by choosing the strongest validation checkpoint for each benchmark category, which makes it easy to assemble a capability-diverse teacher pool without introducing external model families. Second, because these teachers are derived from the same SFT initialization, they share the same tokenizer and vocabulary as the student, reducing distribution shift and avoiding additional alignment issues. Third, MOPD provides a dense token-level training advantage, which is especially useful compared with sparse outcome rewards, and in Figure 3(c) we show its training-efficiency benefits compared with GRPO. MOPD objective. Letํ inf denote the student policy used for response generation in the inference engine, and letํ train denote the student policy optimized by the training engine. For each promptํฅ, we sample a responseํฆ = (ํฆ 1 ,...,ํฆ ํ )โผ ํ inf (ยท | ํฅ). We then select a domain teacherํ domain ํ for that training example, whereํํํํํํ ํ indicates the capability domain associated with the chosen teacher. Writingํ ํก = (ํฅ,ํฆ <ํก )for the decoding state at stepํก, we define the token-level distillation advantage using reverse-KL as ํ MOPD ํก = logํ domain ํ (ํฆ ํก | ํ ํก )โ logํ train (ํฆ ํก | ํ ํก ).(2) 12 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 204060 0.03 0.02 0.01 (a) Reverse KL 204060 0.2 0.4 0.6 0.8 (b) grad_norm 0 102030405060 89 90 91 92 GRPO MOPD Teacher (c) AIME25 (avg@64) Figure 3: Training dynamics and downstream evaluation. Intuitively, this term is positive when the domain teacher assigns a higher probability to the sampled token than the current training policy, and therefore serves as a dense token-level distillation advantage that converges toward 0 during training. The log-probability difference is computed only on the student-sampled token rather than over the full vocabulary. Because responses are sampled underํ inf but optimized underํ train , we apply truncated importance weighting to account for trainโinfer mismatch: ํ ํก = ํ train (ํฆ ํก | ํ ํก ) ํ inf (ํฆ ํก | ํ ํก ) , ํค ํก = sg[ํ ํก ]1[ํ low โค ํ ํก โค ํ high ],(3) where sg[ยท] denotes stop-gradient. We then optimize the surrogate objective โ MOPD =โE ํฅโผํ,ํฆโผํ inf (ยท|ํฅ) โก โฃ 1 |ํฑ(ํฆ)| โ๏ธ ํกโํฑ(ํฆ) ํค ํก sg [๏ธ ํ MOPD ํก ]๏ธ logํ train (ํฆ ํก | ํ ํก ) โค โฆ ,(4) whereํฑ(ํฆ) is the set of valid response tokens retained by the token mask. Hyperparameters. Unless otherwise specified, we use a rollout size of 4 and 128 prompts per update, giving an effective batch size of 512 responses. In later experiments, we find that using 512 prompts with rollout size 1 yields slightly more stable optimization while producing similar final results. We use a learning rate of 2ร 10 โ6 with linear warm-up over the first 30 optimization steps, starting from2ร 10 โ7 . Training typically converges within 40-50 optimization steps (Fig. 3(a)). We find the warm-up stage important for stability: gradient norms are substantially larger at the beginning of training and decrease rapidly after the warm-up phase (Fig. 3(b)). For truncated importance weighting, we setํ low = 0.5andํ high = 2.0. In the main experiments, we use three domain teachers corresponding to math, RLVR, and RLHF. The math teacher is the initial SFT checkpoint, the RLHF teacher is a checkpoint optimized for RLHF, and the RLVR teacher is selected from the early IF-RL + Multi- domain RL stage. We sample prompts accordingly from the RL training data pool and AceReason-Math (Chen et al., 2025). Training efficiency advantage. MOPD provides a dense token-level distillation advantage, whereas GRPO relies on a sparse sequence-level outcome reward that is shared across all generated tokens. This makes MOPD substantially more sample- and step-efficient in practice. Starting from the same initial checkpoint, MOPD consistently reaches stronger perfor- mance in fewer optimization steps. On AIME25 (Figure 3(c)), under math-only training, GRPO improves from 89.9 to 91.0 after 25 steps, while MOPD reaches 92.0 within 30 steps and recovers teacher-level performance. A similar trend appears on ArenaHard v2.0 (Table 3). After 52 steps, MOPD improves Hard Prompt from 71.5 13 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Table 3: Comparison of MOPD and RLHF at matched evaluation checkpoints on ArenaHard V2.0. ArenaHard V2.0 Method Steps Hard Prompt Creative Writing Initial071.540.6 RLHF 10081.768.6 16080.771.2 MOPD5285.571.0 to 85.5 and Creative Writing from 40.6 to 71.0. In contrast, RLHF training requires 160 steps to reach 80.7 on Hard Prompt and 71.2 on Creative Writing. These results show that the dense token-level advantage in on-policy distillation lead to much faster training convergence. 4.5. Reinforcement Learning from Human Feedback (RLHF) Building on multi-domain on-policy distillation, our RLHF recipe focuses on human preferece learning. This process further enhances creative writing and non-verifiable problem-solving in coding and mathematics, as measured by ArenaHard v2, while maintaining performance across other domains without degradation. 4.5.1. Dataset We adopt the RLHF training dataset from NVIDIA Nano-v3 (Blakeman et al., 2025), which comprises HelpSteer3 (Wang et al., 2025), a commercially-friendly subset of the arena-human-preference-140k dataset (Chiang et al., 2024), and a synthetic safety blend (Blakeman et al., 2025). Following the NVIDIA Nano-v3 (Blakeman et al., 2025), we utilize Qwen3-235B-A22B-Thinking-2507 as our generative reward model (GenRM), trained via the HelpSteer3 framework (Wang et al., 2025). Given a conversation history, a user request, and two candidate responses, the GenRM first reasons through the strengths and weaknesses of each response before producing individual helpfulness scores and a final comparative ranking. 4.5.2. Training recipe Following a training recipe similar to NVIDIA Nano-v3 (Blakeman et al., 2025), we conduct RLHF using the GenRM. To ensure the training signals are of high quality, we adopt pair-wise comparisons for all pairs of rollouts per prompt. We aggreagte the reward scores in the same way as NVIDIA Nano-v3 RLHF training, and apply the same length-normalized reward adjustment and quality-gated conciseness bonus (Blakeman et al., 2025). These mechanisms encourage shorter responses without sacrificing quality, effectively mitigating the rapid growth of inference token usage. Different from Nemotron Cascade (Wang et al., 2025), we train RLHF exclusively in the thinking mode. While incorporating both thinking and non-thinking modes can improve training convergence and yield slight gains on evaluation benchmarks, we observe a significant degradation in instruction-following performance. The resulting drop is substantial enough that the gains obtained in the earlier RLVR stage cannot be fully recovered. 4.5.3. Hyper-parameters We use a batch size of 128, generating 16 rollout per prompt with a temperature of 1.0 and a top-p value of 1.0. We use a maximum response length of 16K during RLHF without applying overlong filtering. We adopt a learning rate of 3e-6 with AdamW (Kingma, 2014). We set the entropy loss coefficient to 0 and the KL loss coefficient to 0.03 to keep the model capabilities on other domains. The training takes around 30 steps. 4.6. Long-context RL Following RLHF, we conduct a stage of long-context RL to further enhance the modelโs long-context understand- ing and reasoning capabilities. We use the NVIDIA Nano-v3 RL data blend (Blakeman et al., 2025), but restrict 14 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation this phase to long-context datasets only. In our experiments, incorporating other domains during long-context RL negatively affects performance on unrelated benchmarks, motivating this domain-specific training setup. We adopt the Nemo-Gym RL environment (NVIDIA, 2025) and use Qwen3-235B-A22B-Instruct-2507 as an LLM judge to evaluate model rollouts for question answering tasks. During training, input sequences are limited to 32K tokens, and the maximum sequence length is set to 49K tokens without applying overlength filtering. We train with a batch size of 128, generating 16 rollouts per prompt with temperature 1.0 and top-p 1.0. Optimization is performed using AdamW (Kingma, 2014) with a learning rate of3ร 10 โ6 , while both the entropy and KL loss coefficients are set to zero. Training runs for approximately 30 steps, as we observe a rapid increase in generated inference tokens beyond this point. 4.7. Code RL 4.7.1. Data Curation We construct our Code RL training set from the Nemotron-Cascade coding corpus (Wang et al., 2025), which contains coding prompts sourced from modern competitive programming platforms such as AtCoder, Codeforces, and AIZU with robust test cases for reward verification. To improve training efficiency and strengthen deep reasoning, we aggressively filter out prompts that GPT-OSS-120B solves correctly in all 8 of 8 rollouts, yielding a compact final set of only 3.5K samples. We find that high-difficulty prompts paired with strong test cases are critical for further boosting model performance. 4.7.2. Training Details We conduct Code RL using a batch size of 128 and a learning rate of3ร 10 โ6 with the AdamW optimizer. Compared to Nemotron-Cascade, we increase the maximum response length during RL to 118K tokens and the number of rollouts per sample to 16, enabling the policy to better capture sparse reward signals on extremely difficult problems that require long reasoning traces. We adopt the strict binary reward function to avoid potential reward hacking and keep the whole training to be fully on-policy for stability. To support the resulting verification throughput of128ร 16 = 2, 048code executions per RL step, we deploy an asynchronous reward verification server that completes each batch in 427.2 seconds across 384 CPU cores. 4.8. Software Engineering Reinforcement Learning (SWE RL) 4.8.1. Agentless RL Training Details and Hyperparameters. To enhance the modelsโ code repair capability, we adopt the same data source as Wang et al. (2025) for agentless code repair reinforcement learning (RL) training. Since most instances do not provide executable Docker environments, we employ GPTOOS-120B as a reward model to evaluate the quality of code repairs generated by our models. Following Wang et al. (2025), for each instance we construct prompts using both the golden localization and the top-5retrieved localizations, and filter out relatively easy samples. We perform agentless SWE RL with a batch size of128ร 16 = 2,048(128 prompts with 16 rollouts per prompt), a maximum sequence length of 98,304, and a learning rate of3ร 10 โ6 using the AdamW optimizer. We sample responses with temperature 1.0 and top-p 1.0. During training, we mask the loss for prompts for which none of the rollouts receives a reward greater than 0.5. We observe that these difficult prompts degrade the stability and effectiveness of agentless SWE RL training. Our agentless RL training typically converges within 40โ50 steps. Can Agentless RL Training Helps Agentic Tasks? Table 4 shows that agentless RL training not only improves model performance within the agentless framework but also enhances the modelsโ ability to solve SWE tasks in agentic settings. Note that for Agentless Mini evaluation, we employ a code embedding model, NV-Embed-Code (Sohrabizadeh et al., 2025), to retrieve 5 candidate files whose code contents are semantically similar to the problem context. This result suggests that 15 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Table 4: Effectiveness of Agentless RL on SWE-bench Verified. Scaffold Agentless MiniOpenHands avg@4 pass@4 avg@4 pass@4 Init.41.9% 55.2% 49.8% 64.2% after Agentless RL 44.3% 57.4% 50.8% 65.0% improving modelsโ code repair capability alone can generalize across different scaffolds, consistent with the observations from Yang et al. (2026). 4.8.2. Execution-based RL for Agentic SWE Scaffold Modern software engineering agents rely on scaffolding frameworks that coordinate repository interaction, tool calling, code editing, and test execution. Training agents to operate effectively within these environments requires optimizing not only individual model outputs but the entire problem-solving trajectory. To address this, we apply Reinforcement Learning from Verifiable Rewards (RLVR) directly within agentic SWE scaffolds, enabling end-to-end optimization of the full agent workflow. Our training environments integrate established OpenHands frameworks (Wang et al., 2025), which provide structured tool usage, repository interaction, and iterative patch generation. We train agents using execution-based reinforcement learning in fully executable software environments, where each episode corresponds to resolving a software issue instance from benchmarks such as SWE-bench. The agent operates inside an instrumented repository that exposes tools for file inspection, search, code editing, and test execution. Candidate patches generated by the agent are executed within the environment, which returns verifiable signals from compilation results and unit test outcomes, enabling automatic reward computation without human annotation. Through the OpenHands scaffolding framework, the agent iteratively localizes defects, proposes patches, and validates them through test execution. Environment feedbackโincluding compilation errors, failing tests, or successful test passesโprovides deterministic rewards that directly reflect functional correctness. Specifically, we conduct execution-based agentic reinforcement learning with a batch size of 1024, corresponding to 16 prompts with 64 rollouts per prompt. The maximum context length is set to 256k tokens, and the agent is allowed up to 200 interaction turns, providing a larger reasoning token budget during agentic coding problem solving. Training data is drawn from SWE-Gym (Pan* et al., 2025) and R2E-Subset (Jain et al., 2025). We generate 16 rollouts per instance using our intermediate model and evaluate them using the verification pipeline. Instances for which all rollouts pass verification (100% accuracy), indicating overly simple problems, are removed from the dataset. For instances where none of the rollouts pass verification (0% accuracy), indicating extremely difficult problems, we randomly discard 90% of such cases to reduce their proportion in the training data. 5. International Mathematical Olympiad (IMO) 5.1. IMO 2025 In Table 2, we evaluate Nemotron-Cascade-2-30B-A3B on the IMO 2025 problem set using a self-improving test-time scaling framework (Shao et al., 2025), in which the model iteratively generates candidate solutions, verifies them, and refines them based on its own feedback. Remarkably, despite its relatively modest 30B-A3B scale, the model successfully solves the first five problems. We provide the full model solutions in Appendix E, together with comments from the human expert. These results are particularly encouraging, as they suggest that strong olympiad-level mathematical reasoning can emerge from a comparatively compact model when paired with effective inference-time scaling. There remain several promising directions for improvement: 16 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation expert review indicates that some proofs are longer than necessary, include superfluous intermediate steps or definitions, occasionally expose traces of intermediate reasoning, and sometimes contain minor typographical issues. For Problem 2, the model adopts an analytic solution strategy, similar to OpenAIโs approach, rather than a more geometric approach such as that used by Gemini Deep Think (IMO Gold). 5.2. IMO-ProofBench Table 5: IMO-ProofBench (Luong et al., 2025) reports scores split into the Basic (30 problems) and Advanced (30 problems) subtasks, as well as Overall (60 problems). Expert-evaluated results are taken from the IMO- ProofBench leaderboard (accessed on 2026/3/9). ModelIMO-ProofBench Basic (30) Advanced (30) Overall (60) Aletheia (Feng et al., 2026)-91.9- Gemini 3 Deep Think (Gemini Team, 2026)-76.7- Gemini Deep Think (IMO Gold) (Gemini Team, 2025)89.065.776.7 DeepSeek-Math-V2-671B-A37B (Shao et al., 2025)99.061.980.2 DeepSeek-Math-V2-671B-A37B (our reproduced score) โ 99.557.778.6 Nemotron-Cascade-2-30B-A3B โ 92.553.472.9 GPT-5.2-Thinking (high) (OpenAI, 2025)-35.7- Gemini 3 Pro (Gemini Team, 2025)-30.0- GPT-5 Pro (OpenAI, 2025)-28.6- โ Use DeepSeek-V3.2-Speciale as the judge model with LLM ProofAutoGrader prompt (Luong et al., 2025). As shown in Table 5, Nemotron-Cascade-2-30B-A3B achieves72.9on IMO-ProofBench with generate-verify- refine test-time scaling, placing it within 8 points of DeepSeek-Math-V2-671B-A37B despite using 10รfewer active parameters. It reaches 90+ on Basic split and surpass the QED-Nano-4B (54.0) (LM-Provers et al., 2026) by 18 points, though the latter is not directly comparable due to judge model. Re-evaluating the provided DeepSeek-Math-V2 proofs under our LLM-judge setup yields a score within 4 points of the reported human rating, suggesting that our protocol does not substantially overestimate performance (more details in Appendix A.1.2). In Figure 4, we show that increasing test-time compute improves Nemotron-Cascade-2-30B- A3B on IMO-ProofBench (Advanced), raising the score from 40.7 at round 1 to 53.4 at round 5 and narrowing the gap to DeepSeek-Math-V2 under the same grader. 12345 40 50 60 40.7 45 50.5 50.1 53.4 generate-verify-refine rounds score (%) IMO-ProofBench (Advanced) Nemotron-Cascade-2-30B-A3B DeepSeek-Math-V2 DeepSeek-Math-V2 (our reproduced score) Figure 4: IMO-ProofBench (Advanced) score graded by LLM ProofAutoGrader (DeepSeek-V3.2-Speciale). 17 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 6. Competitive Coding Table 6: Competitive programming results on comprehensive benchmarks, evaluated against a significantly expanded set of proprietary and open-source baseline models. Models LiveCodeBenchLiveCodeBench ProCodeforces v625Q125Q22501 - 2507 2408 - 2505EasyMedHard EasyMedHardELOPercentile GPT-5.2 (high)-96.675.05.991.859.623.1259099.9 Gemini-3 Pro90.794.470.05.994.845.67.7244099.8 GPT-o4-mini (high)80.285.451.70.084.529.80.0226699.5 DeepSeek-v3.2-Speciale88.789.748.10.088.543.10.0235399.7 GPT-OSS-120B (high)87.088.841.90.788.531.10.0232099.6 Kimi-K2.5-1T-thinking85.088.545.60.090.237.90.0233399.7 Qwen-3.5-397B-A17B83.689.344.40.088.131.40.0235099.7 Qwen-3.5-122B-A10B78.987.635.60.084.324.20.0223399.4 Qwen-3.5-35B-A3B74.684.625.60.081.117.80.0218199.1 Nemotron-3-Super-120B-A12B78.783.031.00.081.723.20.0221299.4 Qwen3-235B-A22B-Thinking-250778.775.818.80.077.617.50.0211998.6 Nemotron-Cascade-14B74.671.616.30.068.910.50.0200497.9 Qwen3-Next-80B-A3B-Thinking73.268.516.30.069.17.50.0189496.8 Nemotron-3-Nano-30B-A3B68.360.36.00.054.53.50.0168193.1 Nemotron-Cascade-2-30B-A3B87.288.139.20.787.027.60.0232099.6 Nemotron-Cascade-2-30B-A3B (TIR)88.491.045.22.289.336.80.0234599.7 6.1. IOI 2025 and ICPC World Finals 2025 For IOI 2025, we adapt the IOI Test-Time Scaling pipeline from Nemotron-Cascade (Wang et al., 2025), which can be viewed as a multi-round generate-select-submit framework that exploits the modelโs reasoning ability under IOIโs official rules. Each subtask is allotted at most 50 rounds. Within each round, we prompt our model to generate 40 candidate solutions, aggregated with (1) submission history with official judge verdicts from previous rounds, and (2) shared insights from high scored or fully solved subtasks within the same main task. The complete chat template is provided in Appendix C.1. Using this approach, we achieved full score on Problem 3 and 4, achieving a gold-medal score of 439.28 within at most40ร 50 = 2000model generations, while the score of 507.66 is achievable within 5000 generations. Notably, on Problem 2 which requires designing and optimizing a heuristic algorithm, our pipeline reached over 86 points in just 5 rounds (at most 200 model generations), demonstrating the effectiveness of self-refinement and cross-subtask insights. For ICPC World Finals 2025, we generate up to 1000 solutions per problem and submit them for official evaluation after initial filtering. We successfully solved 10 out of 12 problems, achieving the #4 Gold medal placement, with 8 problems (except Problems A and I) solved within only 100 submissions. 6.2. Competitive Coding Benchmark Results We evaluate our Nemotron-Cascade-2-30B-A3B model on various competitive coding benchmarks, including LiveCodeBench v6 (Jain et al., 2024), and LiveCodeBench Pro (Zheng et al., 2025)โs 25Q1 and 25Q2 splits. We also estimate Codeforces ELO score through simulated participation on 40 Div.1/Div.2 Codeforces Rounds held from 2501 to 2507. We report our avg@8 results under 128K-token thinking budget, the sampling temperature of 1.0 and thetop_pof 0.95. For Tool-Integrated Reasoning (TIR) results, we allow our model to call a stateful Python executor for up to 100 calls. For baseline model evaluation, we follow their recommended inference configurations, ensuring a thinking budget of at least 128K tokens to at most 256K tokens. More evaluation details can be found in Appendix A and Appendix D. As shown in Table 6, Nemotron-Cascade-2-30B-A3B achieves magnificent Pass@1 accuracy and ELO rating, even compared with frontier open-source models with over 100B total params, such as Nemotron-3-Super- 18 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation 120B-A12B, GPT-OSS-120B, and Qwen-3.5-122B-A10B. With Tool-Integrated Reasoning (TIR), our modelโs performance can be further boosted especially on hard problems, and match the strongest open-source models with more than 300B total parameters, such as Kimi-K2.5-1T-Thinking, Qwen-3.5-397B-A17B, and DeepSeek- v3.2-Speciale, which either lack TIR support for deep reasoning or perform poorly with Python TIR. Notably, Nemotron-Cascade-2-30B-A3B achieves above 0% on the LiveCodeBench Pro hard split within 8 attempts, demonstrating strong reasoning ability on problems that are extremely difficult even for humans. 7. Acknowledgments We would like to extend our gratitude to the NVIDIA Nemo team for the valuable discussion and collaboration on building reasoning models. We especially wish to thank Boris Ginsburg, Oleksii Kuchaiev, Igor Gitman, Olivier Delalleau, Zhilin Wang, Olivier Delalleau, Banghua Zhu, Tugrul Konuk, Wei Du, Somshubra Majumdar, Wasi Uddin Ahmad, Siddhartha Jain, Jiaqi Zeng, Yi Dong, Alexander Bukharin, Vahid Noroozi, Khushi Bhardwaj, Sugam Dipak Devare, Jian Zhang, and Jonathan Cohen. We thank Ying Lin for helpful discussions and useful input in building the knowledge-intensive SFT dataset. We also thank Atefeh Sohrabizadeh, Jialin Song, and Jonathan Raiman for valuable discussions on SWE-bench. 19 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Appendix A. Benchmarks and Evaluation Setups A.1. Math A.1.1. Non-proof Math For non-proof math reasoning tasks, we include โข AIME 2025 (MAA, 2025) consists of 30 problems from American Invitational Mathematics Examination at 2025. โขAIME 2026 (MAA, 2026) consists of 30 problems from American Invitational Mathematics Examination at 2026. โขHMMT Feb 2025 (HMMT, 2025) consists of 30 problems from Harvard-MIT Mathematics Tournament 2025 February math competition. โข IMO-AnswerBench (Luong et al., 2025) consists of 400 problems with verifiable answers carefully chosen from past Olympiad competitions and then altered by experts to avoid memorization. For Nemotron-Cascade-2-30B-A3 evaluated on AIME 2025, AIME 2026 and HMMT 2025 Feb, we set the thinking budget (maximum response length) to 131K tokens, the sampling temperature to 1.0, the top-p value to 1.0. For the with-tool setting, we enable tool use by appending a system-prompt postfix, allowing the model to call a stateful Python executor for up to 100 tool calls with a maximum response length of 131K tokens. For IMO-AnswerBench, we set to 256K tokens because we found the questions are significantly more difficult. We use and report the LLM-Judge score using GPT-OSS-120B (Agarwal et al., 2025) as the judge and the AnswerAutoGrader prompt (Luong et al., 2025) for answer correctness on IMO-AnswerBench as the short answers are complicated for rule-based verifier to compute. Following Liu et al. (2024, 2026), we report avg@64 for AIME/HMMT and avg@16 for IMO-AnswerBench. For baseline models, we use official numbers from their reports or evaluate them with the recommended settings if the official numbers are unavailable. A.1.2. Math Proof For math proof tasks, we include โข IMO 2025 (IMO, 2025) consists of 6 problems from IMO 2025. โข IMO-ProofBench (Luong et al., 2025) is designed to evaluate the ability of AI models to construct comprehensive and valid mathematical arguments. This benchmark consists of 60 proof-based problems, curated to mirror the kinds of problems found in the IMO. For Nemotron-Cascade-2-30B-A3, we apply test-time scaling following the DeepSeek-Math-V2 generate-verify- refine pipeline, using the same instructions. We implement this pipeline with NeMo-Skills (NVIDIA, 2025). We use the default hyperparameters from DeepSeek-Math-V2: 128 proof generations, 64 verifications per proof, selection of the top 32 proofs for refinement, and 8 verification analyses paired with each proof, prioritizing the lowest-rated analyses. We then generate 4 refined proofs and continue for up to 8 rounds, or until the average proof score reaches the threshold of 0.99999. We set the maximum generation length to 256K tokens, with temperature 1.0 and top-p 0.95. For IMO-ProofBench Basic and 11 problems from the Advanced split (i.e., Problems 1, 4, 7, 13, 14, 17, 19, 22, 25, 26, and 28), we reduce the compute budget to 32 proof generations, 16 verifications, top 8 proofs, and 2 rounds to save compute. For IMO-ProofBench evaluation, we use DeepSeek-V3.2-Speciale to make sure the results are reproducible later and run 64 grading attempts with the ProofAutoGrader prompt (Luong et al., 2025). We found that reporting mean score yields 73.8 for DeepSeek-Math-V2 on the Advanced 20 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation split, which is substantially more generous than the human rating of 61.9. We therefore adopt a simple aggregation rule based on analysis:if any judge assigns a score of 0, the final score is set to 0; otherwise, return the mean score. Under this rule, DeepSeek-Math-V2 obtains 57.7, which is much closer to the human rating and reduces the discrepancy from 11.9 points to 4.2 points. A.2. Code Reasoning For code generation tasks, we include โขLiveCodeBench (Jain et al., 2024) contains diverse algorithm coding problems with unit tests, collected from AtCoder, LeetCode platforms. We evaluate models competitive coding capability on LiveCodeBench v6 (2024/08-2025/05, 454 problems in total). We report pass@1 accuracy in thinking mode, averaged over 8 generations (avg@8). โขLiveCodeBench Pro (Zheng et al., 2025) contains daily-updated challenging competitive coding problems with strong unit tests, collected mainly from top-tier coding contests. We report pass@1 accuracy on Easy/Med difficulty splits in thinking mode, averaged over 8 generations (avg@8) on two recently released subsets: 2025Q1 (2025/01-2025/04, 166 problems in total) and 2025Q2 (2025/04-2025/07, 167 problems in total). โขIOI and ICPC World Finals represent the most challenging and prestigious annual algorithmic coding competitions, gathering the worldโs top human contestants. The IOI awards gold medals to approximately the top 8.3% (one-twelfth) of participants, while the ICPC World Finals (ICPCWF) limits gold medals to only the top 4 teams globally. โขSciCode (Tian et al., 2024) serves as a challenging benchmark to evaluate modelโs ability on solving realistic scientific research tasks from STEM domains. It contains 338 subproblems from 80 main tasks. For Nemotron-Cascade-2-30B-A3B evaluated on LiveCodeBench v6 and LiveCodeBench Pro, we use a 128K- token thinking budget, a sampling temperature of 1.0, a top-p of 0.95. For the with-tool setting, we enable tool use by appending a system-prompt postfix, allowing the model to call a stateful Python executor for up to 100 tool calls with a maximum response length of 131K tokens. We evaluate baseline models with their recommended inference configurations, ensuring a thinking budget of at least 128K tokens. A.3. Knowledge and STEM For knowledge reasoning tasks, we include: โขMMLU-Redux (Gema et al., 2024) is a benchmark consisting of a subset of 3,000 manually re-annotated questions across 30 MMLU subjects (Hendrycks et al., 2020), which eliminates the original annotation errors. We evaluate the models in thinking mode and, due to the large test set size, report exact match (EM) accuracy based on a single generation per question. โข MMLU-Pro (Wang et al., 2024) is an enhanced version of the original MMLU benchmark that mitigates model saturation by expanding to over 12,000 graduate-level questions and increasing answer choices from four to ten. We report EM accuracy in thinking mode using one generation per question. โขGPQA-Diamond (Rein et al., 2024) is a benchmark for assessing an LLMโs scientific reasoning capability. It consists of the highest quality 198 GPQA questions covering graduate-level physics, biology, and chemistry. We report pass@1 accuracy in thinking mode, averaged over 8 generations per question (avg@8) to reduce variance. โข HLE (Phan et al., 2025) is a frontier academic reasoning benchmark spanning a broad range of expert-level subjects. We evaluate on its text-only split, which contains 2,158 examples. For Nemotron-Cascade-2-30B-A3B evaluated on MMLU-Redux, MMLU-Pro, GPQA-Diamond and HLE in thinking mode, we use a temperature of 1.0, a top-p value of 0.95, and a 128K-token thinking budget (maximum response length). For HLE, we use the default system prompt and append โPlease place your final answer inside 21 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation โ to each question, and use GPT-OSS-120B as the LLM judge for answer extraction and correctness verification with the prompt in Appendix C.2. Compared with the official HLE response format, which requests an explanation, an answer, and a confidence score, this boxed-answer prompt improves the accuracy by 6โ7 points, primarily on the math subset, by better aligning with the answer format used in our math SFT data. A.4. Alignment and Instruction-Following For alignment tasks, we include: โขArenaHard 2.0 (Li et al., 2024) is a human-preference alignment benchmark featuring 750 diverse and rigorous real-user prompts. The dataset is specifically structured with 500 prompts targeting open-ended software engineering problems and complex mathematical questions, while the remaining 250 focus on creative writing. It uses an automatic LLM-as-Judge approach to estimate human preferences relative to a baseline model, enabling fully automated, low-cost, and fast evaluation without human intervention. In our experiments, we report results without style control to allow for straightforward comparison with the officially reported numbers of other models. We evaluate the models in thinking mode, and use GPT-4.1 as the automated judge. โข IFBench (Pyatkin et al., 2025) extends IFEval (Zhou et al., 2023) by introducing 58 new, diverse, and challenging verifiable out-of-domain instruction constraints. It provides a separate constraint list to ensure no overlap between training and test constraints, enabling evaluation of an LLMโs generalization ability. The test set contains 294 prompts. We report pass@1 accuracy in thinking mode, averaged over 8 generations (avg@8). โขScale AI Multi-Challenge (Deshpande et al., 2025) is a benchmark designed to evaluate LLMs in multi-turn conversations with human users. It consists of four challenge categories: Instruction Retention, Inference Memory, Reliable Versioned Editing, and Self-Coherence. These tasks require models to simultaneously perform accurate instruction following, effective context management, and in-context reasoning. The test set contains 273 conversations in total. We report pass@1 accuracy in thinking mode, averaged across 10 generations (avg@10). For Nemotron-Cascade models evaluated on IFEval in non-thinking mode, on IFBench and ArenaHard in thinking mode, we use a temperature of 0.6, a top-p value of 0.95, and a maximum response length of 32K tokens. For baseline models, we use officially reported results whenever available; if such results are absent, we evaluate them using their recommended inference configuration or the same settings as ours. A.5. Long Context and Context Learning For long context and context learning tasks, we include: โขA-LCR (Team, 2025) consists of 100 challenging text-based questions that require reasoning over multiple long, real-world documents, including company reports, government consultations, legal documents, and academic papers. Each sample contains a document set averaging approximately 100k tokens. The questions are designed such that answers cannot be directly retrieved from the documents and instead require reasoning across multiple sources of information. We report pass@1 accuracy in thinking mode, averaged over 16 generations (avg@16). โขLongBench v2 (Bai et al., 2025) contains 503 challenging multiple-choice questions with context lengths ranging from 8k to 2M words. The benchmark spans six task categories: single-document QA, multi- document QA, long in-context learning, long dialogue history understanding, code repository understand- ing, and long structured data understanding. The questions are designed to be difficult; even human experts equipped with document search tools may require substantial time to answer them correctly. We evaluate models in thinking mode and report pass@1 accuracy averaged over four generations (avg@4). โขNIAH@1M (Ruler Subset) refers to the needle-in-a-haystack (NIAH) tasks from the RULER bench- mark (Hsieh et al., 2024). The NIAH test (Kamradt, 2023) assesses an LLMโs long-context ability to retrieve 22 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation a specific piece of information (the โneedleโ) embedded within long distractor text (the โhaystackโ). The RULER benchmark defines four variants of this task: Single NIAH, Multi-keys NIAH, Multi-values NIAH, and Multi-queries NIAH. Following Blakeman et al. (2025), we evaluate 100 instances from each category using a 1M-token context setting. Models are evaluated in reasoning-off mode, and we report pass@1 accuracy from a single generation (avg@1). โขCL-Bench (Dou et al., 2026) evaluates an LLMโs ability to learn from provided context and apply the acquired knowledge to solve tasks, a process referred to as context learning. The benchmark contains 1,899 test samples spanning 500 complex contexts and 31,607 verification rubrics, all developed by experienced domain experts. The knowledge required to complete these tasks largely falls outside what existing models typically learn during pre-training, requiring models to learn directly from the provided context. Models are evaluated in thinking mode, and we report pass@1 accuracy from a single generation (avg@1). A.6. Agentic Tasks For agentic tasks, we include: โข BFCL v4 (Patil et al., 2025) offers a comprehensive agentic evaluation framework for LLMs, covering tasks such as web search, memory reading and writing, and function invocation across multiple programming languages. We follow the official BFCL V4 evaluation protocol and report scores across a combination of Agentic, multi-turn, live, and non-live categories. Models are evaluated in thinking mode, and we report pass@1 accuracy based on a single generation (avg@1). โขSWE-bench Verified (OpenAI, 2024) is a subset of the original test set from SWE-bench (Jimenez et al., 2023), consisting of 500 samples verified to be non-problematic by human annotators. We evaluate models in non-thinking mode and report pass@1 accuracy, averaged over 4 generations per prompt (avg@4). โข ํ 2 -Bench (Barres et al., 2025) evaluates multi-turn customer-service agents in environments with explicit policies, tool use, and shared world-state updates. We evaluate on the three official subsets: airline (50 examples), retail (114 examples), and telecom (114 examples). To keep the standard error within 1.5, we report avg@16 on airline and avg@8 on both retail and telecom. โข Terminal Bench 2.0(Merrill et al., 2026) is adopted for evaluating agents in terminal-based environments, which comprises of 89 human-validated tasks across specialized fields such as scientific computing, machine learning, and system administration. Moving beyond simple code generation, this benchmark focuses on end-to-end workflows, requiring agents to demonstrate proficiency in holistic operations like model training, system configuration, and software debugging rather than just producing isolated functions. We evaluate the model using the default Terminus-2 scaffolding. We report avg@5 task success rate. For SWE-bench Verified, we use the OpenHands scaffold (Wang et al., 2025) as the agentic coding evaluation framework. We adopt a full interaction retention policy for agent trajectories, preserving the complete history of tool calls, observations, and model outputs across turns. This includes prior file views, search results, executed commands, and intermediate patches, enabling the model to maintain state and reason effectively over long-horizon debugging processes. We set the maximum context length to 256K tokens and allow up to 200 turns, consistent with our execution-based agentic SWE-RL training configuration. Notably, this evaluation setup closely mirrors our training environment, as both rely on execution-based feedback and multi-turn interaction within the same tool-augmented scaffold. This alignment reduces trainโtest mismatch and enables the model to more effectively transfer learned behaviors, such as iterative debugging, hypothesis refinement, and tool-driven reasoning, to the evaluation setting. Forํ 2 -Bench evaluation, we adopt a latest-turn thought retention policy for managing reasoning traces in multi- turn interactions: we retain the modelโs reasoning content after the most recent user turn, while discarding reasoning content from earlier turns. The officialํ 2 -Bench evaluation code follows a no thought carry-over policy, which removes all prior reasoning content; in our experiments, this evaluation setup consistently reduces scores by 3โ5 points relative to latest-turn thought retention. We attribute this gap to trainโtest mismatch, 23 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation since our SFT data forํ 2 -style interactions is constructed with the same latest-turn thought retention policy, which is also the thought-state management strategy used in Nemotron-3-Nano-v3 and DeepSeek-V3.2. For the telecom subset, we additionally modify the system prompt to emphasize the dual-control setting by repeating the instruction โMake sure you guide the user through the steps, do not perform user-side actions yourself.โ three times. We also tested a full thought retention policy, which preserves reasoning content from all previous turns and more closely matches RL training, but found it gives similar accuracy to latest-turn thought retention while incurring substantially longer contexts. We therefore report our final ํ 2 -Bench results using latest-turn thought retention. A.7. Multilingual For multilingual tasks, we include: โข MMLU-ProX (Xuan et al., 2025) expands the challenging MMLU-Pro benchmark to include 29 languages. Following Blakeman et al. (2025), six languages are selected for evaluation: English (en), German (de), Spanish (es), French (fr), Italian (it), and Japanese (ja). The model is evaluated in thinking mode, and we report pass@1 accuracy from a single generation (avg@1). โข WMT24++ (Deutsch et al., 2025) extends the WMT24 machine translation benchmark to cover 55 languages. Following Blakeman et al. (2025), we evaluate on five translation pairs: English to German (en โ de), English to Spanish (en โ es), English to French (en โ fr), English to Italian (en โ it), and English to Japanese (en โ ja). We use XCOMET-XXL (Guerreiro et al., 2024) as the evaluation metric to assess the translation quality. Our model is evaluated in thinking mode, and we report pass@1 accuracy based on a single generation (avg@1). B. Training Hyperparameters We list the training hyperparameters for the Nemotron-Cascade-2-30B-A3B during all stages in Table 7, 9, 10. Table 7: Training hyperparameters for Nemotron-Cascade-2-30B-A3B in SFT. Hyperparameters Global batch size64 Packed sequence length256K Max learning rate5ร 10 โ5 Min learning rate5ร 10 โ6 Learning rate warmup steps200 Schedulercosine Max Steps40,000 OptimizerAdamW Optimizer configํฝ 1 = 0.9, ํฝ 2 = 0.98 Weight decay0.1 # of training steps33,000 24 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Table 8: Training hyperparameters of Nemotron-Cascade-2-30B-A3B in Cascade RL (IF-RL, Multi-domain RL, MOPD). Hyper-parametersIF-RLMulti-domain RLMOPD Max response length49K49K98K Batch size128128128 # Rollout size16164 Learning rate2ร 10 โ6 3ร 10 โ6 3ร 10 โ6 Steps 1807052 AdamWAdamAdamW Optimizerํฝ 1 = 0.9ํฝ 1 = 0.9ํฝ 1 = 0.9 ํฝ 2 = 0.95ํฝ 2 = 0.95ํฝ 2 = 0.95 Temperature 1.01.01.0 Top-p1.01.01.0 Overlong filteringFalseTrueFalse Table 9: Training hyperparameters of Nemotron-Cascade-2-30B-A3B in Cascade RL (RLHF, Long-context RL, Code RL). Hyper-parametersRLHFLong-context RLCode RL Max response length16K49K118K Batch size128128128 # Rollout size161616 Learning rate3ร 10 โ6 3ร 10 โ6 3ร 10 โ6 Steps253022 AdamWAdamAdamW Optimizerํฝ 1 = 0.9ํฝ 1 = 0.9ํฝ 1 = 0.9 ํฝ 2 = 0.95ํฝ 2 = 0.95ํฝ 2 = 0.95 Temperature1.01.01.0 Top-p1.01.00.95 Overlong filtering TrueTrueTrue 25 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Table 10: Training hyperparameters of Nemotron-Cascade-2-30B-A3B model in execution-based agentic SWE- RL. Hyperparameters # prompts per step16 # rollout64 Temperature0.8 Max sequence length256ํ Max turn200 Max learning rate3ร 10 โ6 Min learning rate0 Learning rate warmup steps10 C. Prompt Templates C.1. Prompt Templates for Test-Time Scaling on IOI 2025 Write Python code to solve the problem. Please place the solution code in the following format: โโpython # Your solution code here โโ problem_statement Below you are provided the accepted correct solutions but with different input constraints. You may use them as a reference for your insights. ======================= ## Different Constraints (for reference only): subtask_constraints ### Accepted Code: [CODE] ======================= ## Different Constraints (for reference only): ... ======================= From here, you are also given your submission history containing **incorrect** code and their corre- sponding official judgement verdicts as reference โ Official judgement verdicts and problem statement/- conditions are 100% reliable. You should make improvements from them if they could help: ======================= ### Incorrect Code [CODE] Judgement Verdict: [VERDICT], Score: [SCORE] ======================= ### Incorrect Code ... ======================= 26 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation C.2. HLE Judge Prompt Judge whether the following [response] to [question] is correct or not based on the precise and unambiguous [correct_answer] below. [question]: question [response]: response Your judgement must be in the format and criteria specified below: extracted_final_answer: The final exact answer extracted from the [response]. Put the extracted answer as โNoneโ if there is no exact, final answer to extract from the response. [correct_answer]: correct_answer reasoning: Explain why the extracted_final_answer is correct or incorrect based on [correct_answer], fo- cusing only on if there are meaningful differences between [correct_answer] and the extracted_final_answer. Do not comment on any background to the problem, do not attempt to solve the problem, do not argue for any answer different than [correct_answer], focus only on whether the answers match. correct: Answer โyesโ if extracted_final_answer matches the [correct_answer] given above, or is within a small margin of error for numerical problems. Answer โnoโ otherwise, i.e. if there if there is any inconsistency, ambiguity, non-equivalency, or if the extracted answer is incorrect. confidence: The extracted confidence score between 0|%| and 100|%| from [response]. Put 100 if there is no confidence score available. D. ELO Rating Analysis We perform ELO rating analysis on our Nemotron-Cascade-2-30B-A3B model based on 40 recent Div.1 and Div.2 Codeforces contests held between 2501โ2507. Problems and evaluations are provided by LiveCodeBench Pro (Zheng et al., 2025). We adopt similar rating estimation approach as in Wang et al. (2025), by allowing model with up toํ = 8submissions to each contest problems, estimating model performance and relative ranking to human contestants with expected penalty consideration. We generate the modelโs responses using a temperature of 1.0,top-pof 0.95, and a maximum token budget of 128K. The performance details of our Nemotron-Cascade-2-30B-A3B model (with and without python-tool use) can be found in Table 11 and Table 12, respectively. We observed our modelโs strong code reasoning ability on solving really tough problems and achieving high ranking even on some Div. 1 rounds (Round 999, 1012, 1015, 1021 etc.), while maintaining stable performance on solving easy-medium level problems. However, the models still has weakness on dealing with problems that requiring constructive algorithms, interactive manner, and hypothesis-driven ideas. 27 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Table 11: Nemotron-Cascade-2-30B-A3B performance details on 40 Div.1 and Div.2 Codeforces Rounds ranging from 2501 to 2507 without python-tool use. We attempt each problem withํ = 8times in total. For regular codeforces rounds, we present the score after considering expected penalties for each problem. For ICPC style rounds, we mark passed/failed problems as + and - correspondingly. We compute the estimated rank to human contestants and the corresponding Elo score as shown in rightmost two columns. Contest NameContest ProblemsScore Penalty Est. Rank ELO Hello 2025 ABCDE1E2FGH 10779.46-13/16703 3449 500.001000.00 1493.75 2235.710.01900.000.03650.000.0 Codeforces Round 996 (Div. 2) ABCDEF 5793.75-2/212322198 500.00 993.751475.000.02825.000.0 Codeforces Round 997 (Div. 2) ABCDEF1F2 9378.75-1/188232198 493.751250.00 1475.000.02225.00 2710.00 1225.00 IAEPC Preliminary Contest (Codeforces Round 999, Div. 1 + Div. 2) ABCDEF1F2GH1H2I 9278.75-43/12647 3076 500.001000.00 1500.00 1493.75 1960.000.00.00.02825.000.00.0 Codeforces Round 1000 (Div. 2) ABCDEF1F2 10935.71-1/171692200 500.00 985.711500.00 2250.00 2687.50 1687.50 1325.00 Ethflow Round 1 (Codeforces Round 1001, Div. 1 + Div. 2) ABCDE1E2FGH 2493.75-1727/16234 1898 500.00 993.751000.000.00.00.00.00.00.0 Codeforces Round 1002 (Div. 2) ABCDE1E2 3300.00-1102/19443 1882 500.00 975.000.01825.000.00.0 Codeforces Round 1004 (Div. 1) ABCD1D2EF 2681.25-145/1030 2666 0.0687.501250.00743.750.00.00.0 Codeforces Round 1004 (Div. 2) ABCDEFG 5397.50-8/167492098 500.00 960.000.00.01687.50 2250.000.0 Codeforces Round 1005 (Div. 2) ABCDEF 9198.21-1/176212260 493.751000.00 1243.75 1735.71 2075.00 2650.00 Educational Codeforces Round 174 (Rated for Div. 2) ABCDEF 42.86156/16701 2242 ++++-- Educational Codeforces Round 175 (Rated for Div. 2) ABCDEF 40.00234/16060 2195 ++++-- Codeforces Round 1007 (Div. 2) ABCD1D2EF 8429.46-1/162542198 500.001000.00 1485.71 1743.75 1225.00 2475.000.0 Codeforces Round 1008 (Div. 1) ABCDEFG 2000.00-355/9092312 500.000.01500.000.00.00.00.0 Codeforces Round 1008 (Div. 2) ABCDEFG 6825.00-9/146412008 500.00 750.001250.00 1575.000.02750.000.0 Educational Codeforces Round 176 (Rated for Div. 2) ABCDEF 510.862/181592198 +++++- Codeforces Round 1011 (Div. 2) ABCDEF1F2 10137.50-1/159062200 500.001250.00 1250.00 1743.75 2500.00 1993.75900.00 Codeforces Round 1012 (Div. 1) AB1B2C1C2DE 2985.00-24/6533057 710.00 975.00 325.00 975.000.00.00.0 Codeforces Round 1012 (Div. 2) ABCDE1E2F1F2 9945.00-1/85362007 500.00 960.001750.00 1960.00 1975.00825.001975.000.0 Codeforces Round 1014 (Div. 2) ABCDEF 6500.00-2/158422213 500.00 750.001250.00 1750.00 2250.000.0 Teza Round 1 (Codeforces Round 1015, Div. 1 + Div. 2) ABCDEFG1G2H 12521.43-4/112063830 750.001000.00 1500.00 1735.71 2235.71 2825.00 2475.000.00.0 Neowise Labs Contest 1 (Codeforces Round 1018, Div. 1 + Div. 2) ABCDEFGH 4400.00-493/12771 2312 500.00 750.001500.00 1650.000.00.00.00.0 Codeforces Round 1019 (Div. 2) ABCDEF 4825.00-47/14465 2202 500.001000.00 1500.00 1825.000.00.0 Codeforces Round 1021 (Div. 1) ABCDEF 3218.75-75/6512760 493.75 900.000.01825.000.00.0 Codeforces Round 1021 (Div. 2) ABCDEF 8468.75-1/58242019 500.001250.00 1493.75 2150.000.03075.00 Educational Codeforces Round 178 (Rated for Div. 2) ABCDEFG 612.504/117062215 ++++++- Codeforces Round 1022 (Div. 2) ABCDEF 3087.50-308/11127 2132 500.001187.50 1400.000.00.00.0 Codeforces Round 1023 (Div. 2) ABCDEF1F2 6506.25-6/116362209 250.00 750.001493.75 1937.500.02075.000.0 Codeforces Round 1024 (Div. 1) ABCDEF 1729.46-477/8572149 485.711243.750.00.00.00.0 Codeforces Round 1024 (Div. 2) ABCDEF 3479.46-34/11201 1998 250.00 500.00 985.711743.750.00.0 Codeforces Round 1025 (Div. 2) ABC1C2C3DEF 7985.71-1/159452197 500.00 985.711243.75575.00 500.001687.50 2493.750.0 Codeforces Round 1026 (Div. 2) ABCDEF 9897.50-1/176682198 500.00 750.001500.00 1960.00 2250.00 2937.50 Codeforces Round 1028 (Div. 1) ABCDEF1F2 2710.00-75/9562865 500.000.00.02210.000.00.00.0 Codeforces Round 1028 (Div. 2) ABCDEF 5453.75-4/183142018 493.75 750.001250.000.00.02960.00 Educational Codeforces Round 179 (Rated for Div. 2) ABCDEFG 560.0094/12301 2231 +-++++- Codeforces Round 1030 (Div. 2) ABCD1D2EF 7003.75-2/183352205 500.00 975.001000.00 1243.75960.002325.000.0 Codeforces Round 1031 (Div. 2) ABCDEF 4060.71-20/11032 2216 500.00 735.710.00.00.02825.00 Codeforces Round 1033 (Div. 2) and CodeNite 2025 ABCDEFG 9623.21-1/129482216 493.75 750.001250.00 1735.71 2493.75 2900.000.0 Educational Codeforces Round 180 (Rated for Div. 2) ABCDEF 533.758/171282253 +++++- Codeforces Round 1035 (Div. 2) ABCDEF 2985.71-587/15624 2008 500.001000.00 1485.710.00.00.0 28 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation Table 12: Nemotron-Cascade-2-30B-A3B performance details on 40 Div.1 and Div.2 Codeforces Rounds ranging from 2501 to 2507 with python-tool use. We attempt each problem withํ = 8times in total. For regular codeforces rounds, we present the score after considering expected penalties for each problem. For ICPC style rounds, we mark passed/failed problems as + and - correspondingly. We compute the estimated rank to human contestants and the corresponding Elo score as shown in rightmost two columns. Contest NameContest ProblemsScore Penalty Est. Rank ELO Hello 2025 ABCDE1E2FGH 11712.50-11/16703 3497 500.001000.00 1500.00 2225.00937.501900.000.03650.000.0 Codeforces Round 996 (Div. 2) ABCDEF 5025.00-2/212322198 500.00 975.001475.00 2075.000.00.0 Codeforces Round 997 (Div. 2) ABCDEF1F2 11187.50-1/188232198 493.751250.00 1500.00 1900.00 2075.00 2743.75 1225.00 IAEPC Preliminary Contest (Codeforces Round 999, Div. 1 + Div. 2) ABCDEF1F2GH1H2I 9416.96-40/12647 3097 500.001000.00 1493.75 1485.71 2000.000.00.00.02937.500.00.0 Codeforces Round 1000 (Div. 2) ABCDEF1F2 10981.25-1/171692200 500.001000.00 1500.00 2243.75 2725.00 1687.50 1325.00 Ethflow Round 1 (Codeforces Round 1001, Div. 1 + Div. 2) ABCDE1E2FGH 2493.75-1727/16234 1898 500.00 993.751000.000.00.00.00.00.00.0 Codeforces Round 1002 (Div. 2) ABCDE1E2 3300.00-1102/19443 1882 500.00 975.000.01825.000.00.0 Codeforces Round 1004 (Div. 1) ABCD1D2EF 2743.75-122/1030 2721 0.0743.751250.00750.000.00.00.0 Codeforces Round 1004 (Div. 2) ABCDEFG 5487.50-6/167492098 500.00 993.750.00.01743.75 2250.000.0 Codeforces Round 1005 (Div. 2) ABCDEF 9230.71-1/176212260 500.00 985.711250.00 1710.00 2210.00 2575.00 Educational Codeforces Round 174 (Rated for Div. 2) ABCDEF 42.50156/16701 2242 ++++-- Educational Codeforces Round 175 (Rated for Div. 2) ABCDEF 55.003/160602198 +++++- Codeforces Round 1007 (Div. 2) ABCD1D2EF 8443.75-1/162542198 500.001000.00 1500.00 1725.00 1225.00 2493.750.0 Codeforces Round 1008 (Div. 1) ABCDEFG 2000.00-355/9092312 500.000.01500.000.00.00.00.0 Codeforces Round 1008 (Div. 2) ABCDEFG 6975.00-5/146412008 500.00 750.001250.00 1725.000.02750.000.0 Educational Codeforces Round 176 (Rated for Div. 2) ABCDEF 52.502/181592198 +++++- Codeforces Round 1011 (Div. 2) ABCDEF1F2 10137.50-1/159062200 500.001250.00 1250.00 1743.75 2493.75 2000.00900.00 Codeforces Round 1012 (Div. 1) AB1B2C1C2DE 2693.75-66/6532745 725.00 975.000.0993.750.00.00.0 Codeforces Round 1012 (Div. 2) ABCDE1E2F1F2 9193.75-1/85362007 500.001000.00 1750.00 1975.00 1975.000.01993.750.0 Codeforces Round 1014 (Div. 2) ABCDEF 6500.00-2/158422213 500.00 750.001250.00 1750.00 2250.000.0 Teza Round 1 (Codeforces Round 1015, Div. 1 + Div. 2) ABCDEFG1G2H 9723.21-55/11206 3008 750.001000.00 1500.00 1743.75 2243.750.02485.710.00.0 Neowise Labs Contest 1 (Codeforces Round 1018, Div. 1 + Div. 2) ABCDEFGH 6397.50-70/12771 2933 500.00 750.001500.00 1687.50 1960.000.00.00.0 Codeforces Round 1019 (Div. 2) ABCDEF 7725.00-2/144652202 500.001000.00 1500.00 1825.000.02900.00 Codeforces Round 1021 (Div. 1) ABCDEF 4899.46-21/6513143 493.75 985.711460.00 1960.000.00.0 Codeforces Round 1021 (Div. 2) ABCDEF 11399.46-1/58242019 500.001250.00 1493.75 2235.71 2710.00 3210.00 Educational Codeforces Round 178 (Rated for Div. 2) ABCDEFG 612.114/117062215 ++++++- Codeforces Round 1022 (Div. 2) ABCDEF 3235.71-300/11127 2137 500.001250.00 1485.710.00.00.0 Codeforces Round 1023 (Div. 2) ABCDEF1F2 6537.50-6/116362209 250.00 743.751493.75 1900.000.02150.000.0 Codeforces Round 1024 (Div. 1) ABCDEF 4075.00-156/8572590 500.001250.000.02325.000.00.0 Codeforces Round 1024 (Div. 2) ABCDEF 6325.00-2/112011998 250.00 500.001000.00 1750.000.02825.00 Codeforces Round 1025 (Div. 2) ABC1C2C3DEF 8087.14-1/159452197 500.00 985.711235.71710.00 460.001710.00 2485.710.0 Codeforces Round 1026 (Div. 2) ABCDEF 7931.25-8/176682198 500.00 750.001500.000.02243.75 2937.50 Codeforces Round 1028 (Div. 1) ABCDEF1F2 3960.71-36/9563093 500.001225.000.02235.710.00.00.0 Codeforces Round 1028 (Div. 2) ABCDEF 7454.46-1/183142018 493.75 750.001250.00 1975.000.02985.71 Educational Codeforces Round 179 (Rated for Div. 2) ABCDEFG 640.113/123012231 ++++++- Codeforces Round 1030 (Div. 2) ABCD1D2EF 7160.71-2/183352205 500.00 993.751000.00 1243.75937.502485.710.0 Codeforces Round 1031 (Div. 2) ABCDEF 4143.75-20/11032 2216 500.00 743.750.00.00.02900.00 Codeforces Round 1033 (Div. 2) and CodeNite 2025 ABCDEFG 9650.00-1/129482216 500.00 750.001250.00 1750.00 2500.00 2900.000.0 Educational Codeforces Round 180 (Rated for Div. 2) ABCDEF 510.868/171282253 +++++- 29 IMO2025IMO 2025 Model Proofs E. IMO 2025 Model Solutions Problem 1 A line in the plane is called sunny if it is not parallel to any of theํฅ-axis, the ํฆ-axis, or the line ํฅ + ํฆ = 0. Letํโฅ 3be a given integer. Determine all nonnegative integersํsuch that there exist ํ distinct lines in the plane satisfying both of the following: โขfor all positive integersํandํwithํ +ํโค ํ + 1, the point(ํ,ํ)lies on at least one of the lines; and โข exactly ํ of the ํ lines are sunny. Short Answer k = 0, 1, 3 for all n Model Proof Notation. For an integer ํโฅ 3 define ํ ํ =(ํ,ํ)โ Z 2 | ํโฅ 1, ํโฅ 1, ํ + ํโค ํ + 1. A line in the plane is called sunny if its slope is neither0(horizontal),โ(vertical), norโ1 (the slope of the line ํฅ + ํฆ = 0). We have to determine all non-negative integersํfor which there existํdistinct lines covering all points of ํ ํ with exactly ํ of them sunny. 1. Constructions showing that ํ = 0, 1, 3 are attainable ํ = 0. Take theํvertical linesํฅ = 1,ํฅ = 2,...,ํฅ = ํ. Every point(ํ,ํ)โ ํ ํ hasํโค ํ, hence lies on the lineํฅ = ํ. These lines are not sunny. Clearly there areํdistinct lines, soํ = 0is possible for every ํโฅ 3. ํ = 1. For ํโฅ 3 take the ํโ 1 vertical lines ํฅ = 1,...,ํฅ = ํโ 1 and the line ํฟ : through (ํ, 1) and (ํโ 2, 2). The slope of ํฟ is 2โ1 (ํโ2)โํ =โ 1 2 ; it is not 0,โ,โ1, so ํฟ is sunny. - Points with ํโค ํโ 1 are covered by the vertical line ํฅ = ํ. - The only point with ํ = ํ is (ํ, 1), which lies on ํฟ. Thus all points of ํ ํ are covered. Hence ํ = 1 is attainable for all ํโฅ 3. ํ = 3. Case ํ = 3. ํ 3 consists of the six points (1, 1), (1, 2), (1, 3), (2, 1), (2, 2), (3, 1). Take the three sunny lines โ 1 : ํฆ = ํฅ, โ 2 : ํฆ =โ2ํฅ + 5, โ 3 : ํฆ =โ 1 2 ํฅ + 5 2 . Their slopes are1,โ2,โ 1 2 - none is0,โ,โ1, so they are sunny. One checks directly that each of the six points belongs to at least one of them. Therefore ํ = 3 is possible for ํ = 3. Case ํโฅ 4. 30 of 63 IMO2025IMO 2025 Model Proofs Take the ํโ 3 vertical lines ํฅ = 1,ํฅ = 2,...,ํฅ = ํโ 3. The remaining points of ํ ํ are those with ํโฅ ํโ 2: ํ =(ํโ 2, 1), (ํโ 2, 2), (ํโ 2, 3), (ํโ 1, 1), (ํโ 1, 2), (ํ, 1). Cover ํ by the following three sunny lines: ํฟ ํ : through (ํโ 2, 1) and (ํโ 1, 2) (slope 1), ํฟ ํ : through (ํโ 2, 2) and (ํ, 1)(slope โ 1 2 ), ํฟ ํ : through (ํโ 2, 3) and (ํโ 1, 1) (slope โ 2). All slopes are different from 0,โ,โ1, hence these lines are sunny. Together with theํโ 3vertical lines we have exactlyํdistinct lines. Every point with ํโค ํโ 3lies on a vertical line; the six points ofํlie onํฟ ํ ,ํฟ ํ ,ํฟ ํ . Thus the wholeํ ํ is covered. Consequently ํ = 3 is attainable for all ํโฅ 3. 2. Impossibility: no other value of ํ can occur We prove by induction on ํ that any covering of ํ ํ by ํ lines must have ํ โ0, 1, 3. 2.1. The boundary set ํ and a lemma Define the three special lines ํ 1 : ํฅ = 1, ํ 2 : ํฆ = 1, ํ 3 : ํฅ + ํฆ = ํ + 1. These are the sides of the right triangle that containsํ ํ . Letํbe the set of points ofํ ํ lying on these three sides; i.e. ํ =(1,ํ)| 1โค ํโค ํโช(ํ, 1)| 1โค ํโค ํโช(ํ,ํ)| ํ + ํ = ํ + 1. A1 A1 | Similar to definition of ํ ํ , we need to require ํ,ํ are integers. There are|ํ| = 3ํโ 3 distinct points (the three corners are counted twice) A2 A2 | What means of counted twice? Two options: 1. Each edge have ํ point, so totally 3ํ in 3 edges, but three corners counted twice so 3ํโ3 distinct. 2. Directly say 3ํโ3 distinct points without mentioning counted twice is also okay. Lemma. A line that is not one of ํ 1 ,ํ 2 ,ํ 3 contains at most two points of ํ. Proof. If such a line contained two distinct points from the same side, it would coincide with that side, contradicting the assumption. Hence it can contain at most one point from each side. A3 Moreover, a straight line can intersect the boundary of a triangle in at most two points A3 | This sentence is not related to the proof. (unless it coincides with a side). Therefore it contains at most two points of ํ.โก 2.2. A counting argument for ํโฅ 4 Consider any covering ofํ ํ byํlines, and letํbe the number of sunny lines A4 . For each A4 | The definition is unnecessary since ํ is not used in 2.2. lineโdenote byํ (โ)the number of points ofํthat lie onโ. Because every point ofํis covered at least once, โ๏ธ โ ํ (โ)โฅ|ํ| = 3ํโ 3.(1) If a line is one of the special sidesํ 1 ,ํ 2 ,ํ 3 , then it contains allํpoints of that side, so ํ (โ) = ํ. If it is not a special side (whether sunny or not), the lemma gives ํ (โ)โค 2. Letํ be the number of the three special sides that actually appear among theํlines. Then โ๏ธ โ ํ (โ)โค ํ ยท ํ + (ํโ ํ )ยท 2 = 2ํ + ํ (ํโ 2).(2) Combining (1) and (2) we obtain 3ํโ 3โค 2ํ + ํ (ํโ 2)=โ ํโ 3โค ํ (ํโ 2).(3) Forํโฅ 4we haveํโ 2 > 0 A5 , so (3) forcesํ โฅ 1. Hence every covering ofํ ํ withํโฅ 4 A5 | ํโ 3 > 0 lines must contain at least one of the sides ํฅ = 1, ํฆ = 1 or ํฅ + ํฆ = ํ + 1. 31 of 63 IMO2025IMO 2025 Model Proofs 2.3. Reducing the problem to ํโ 1 Assumeํโฅ 4and that a covering ofํ ํ withํlines andํsunny lines exists. By the previous paragraph, one of the special sides is present; we show how to produce a covering ofํ ํโ1 with ํโ 1 lines and the same ํ. Case A: ํฅ = 1 is present. Remove the line ํฅ = 1. The remaining ํโ 1 lines cover all points of ํ ํ with ํโฅ 2. Let ํ โฒ =(ํ,ํ)โ ํ ํ | ํโฅ 2, and define the translationํ : ํ โฒ โ ํ ํโ1 byํ(ํ,ํ) = (ํโ 1,ํ) A6 . This is a bijection. A6 | ํ((ํ,ํ)) since input is a pair (ํ,ํ) not two parameter ํ and ํ. Translate the remainingํโ 1lines by(โ1, 0): for each lineโ(withโ ฬธ= (ํฅ = 1)), set โ โฒ = (ํฅโ 1,ํฆ) | (ํฅ,ํฆ) โ โ. Translation preserves slopes, soโ โฒ is sunny exactly when โis sunny, and distinct lines remain distinct. Because the original lines coveredํ โฒ , their translates coverํ(ํ โฒ ) = ํ ํโ1 . Thus we obtain a covering ofํ ํโ1 byํโ 1lines with exactly ํ sunny lines. Case B: ํฆ = 1 is present. Analogously, removeํฆ = 1, translate the remaining lines by(0,โ1), and obtain a covering of ํ ํโ1 with ํโ 1 lines and ํ sunny lines. Case C: ํฅ + ํฆ = ํ + 1 is present. Remove this line. The set(ํ,ํ) โ ํ ํ | ํ + ํ โค ํis exactlyํ ํโ1 . The remainingํโ 1 lines coverํ ํโ1 directly (they cover all ofํ ํ , and the removed line contains no point of ํ ํโ1 ). No translation is needed, and the number of sunny lines stays ํ. In all three cases we have produced a covering of ํ ํโ1 by ํโ 1 lines with the same ํ. 2.4. Base case ํ = 3 and induction Base case ํ = 3. We have already exhibited coverings withํ = 0, 1, 3. It remains to show thatํ = 2is impossible. The setํ 3 is exactly the boundaryํforํ = 3. Observe that the only lines containing two points of ํ 3 and being sunny are โ 1 : through (1, 1), (2, 2) (ํ ํํํํ1), โ 2 : through (1, 2), (3, 1) (ํ ํํํํโ 1/2), โ 3 : through (1, 3), (2, 1) (ํ ํํํํโ 2). Any other line containing two points has slope 0,โ, orโ1, hence is non-sunny. Suppose a covering ofํ 3 by three lines has exactlyํ = 2sunny lines. Then there is exactly one non-sunny line ํ. Ifํis one of the special sides (sayํฅ = 1), thenํcovers three points. The remaining three points must be covered by the two sunny lines. But any line through two of those three points is either vertical, horizontal, or of slopeโ1, all non-sunny. Hence each sunny line can contain at most one of these points, so they cover at most two points - contradiction. Ifํis not a special side, thenํcontains at most two points. Letํ (โ)be the number of points ofํ 3 on lineโ. To cover all six points we need โ๏ธ ํ (โ)โฅ 6 . For a sunny line that is not special,ํ (โ)โค 2(by the lemma). If any line hadํ (โ)โค 1, then the total sum would be at most1 + 2 + 2 = 5 < 6. Therefore each line must satisfyํ (โ) = 2. In particular, the two sunny lines each contain exactly two points, andํalso contains exactly two points. Thus the three lines partition ํ 3 into three disjoint collinear pairs. The only collinear pairs that can be covered by a sunny line areโ 1 ,โ 2 ,โ 3 . Hence the two sunny lines must be two ofโ 1 ,โ 2 ,โ 3 . The uncovered pair is exactly the pair of the remainingโ ํ . Forํto cover that pair,ํwould have to be the line through those two points, which isโ ํ , a sunny line - contradicting thatํis non-sunny. Thereforeํ = 2is impossible. Consequently, for ํ = 3 the only possible ํ are 0, 1, 3. Induction step for ํโฅ 4. Assume the statement holds forํโ 1. Given any covering ofํ ํ byํlines withํsunny lines, we have shown how to produce a covering ofํ ํโ1 byํโ 1lines with the sameํ. By the induction hypothesis, ํ โ0, 1, 3. Hence the statement holds for ํ as well. By induction, for everyํโฅ 3any covering ofํ ํ byํlines must haveํ โ0, 1, 3. Together with the constructions in ยง1, we conclude that the set of all non-negative integersํfor which such a configuration exists is exactly 32 of 63 IMO2025IMO 2025 Model Proofs 0, 1, 3. Human Expert Comment 7/7 33 of 63 IMO2025IMO 2025 Model Proofs Problem 2 Letโฆandฮbe circles with centresํandํ, respectively, such that the radius ofโฆis less than the radius ofฮ. Supposeโฆandฮintersect at two distinct pointsํดandํต. Lineํํintersectsโฆatํถandฮ atํท, so thatํถ,ํ,ํ,ํทlie onํํin that order. Letํbe the circumcentre of triangleํดํถํท. Lineํดํ meets โฆ again at ํธ ฬธ= ํด and meets ฮ again at ํน ฬธ= ํด. Let ํป be the orthocentre of triangle ํํํ. Prove that the line through ํป parallel to ํดํ is tangent to the circumcircle of triangle ํตํธํน . Short Answer No short answer provided. Model Proof We set up a coordinate system with ํ = (0, 0) and ํ = (ํ, 0), ํ > 0. Circles: ฮฉ : ํฅ 2 + ํฆ 2 = ํ 2 (radius ํ), ฮ : (ํฅโ ํ) 2 + ํฆ 2 = ํ 2 (radius ํ ), 0 < ํ < ํ . They intersect in two points; let ํด = (ํฅ ํด ,ํฆ ํด ) with ํฆ ํด > 0, ํต = (ํฅ ํด ,โํฆ ํด ). Solving the circle equations gives ํฅ ํด = ํ 2 + ํ 2 โ ํ 2 2ํ , ํฆ 2 ํด = (๏ธ (ํ + ํ) 2 โ ํ 2 )๏ธ(๏ธ ํ 2 โ (ํ โ ํ) 2 )๏ธ 4ํ 2 . Define ํ = ํ + ํ, ํ = ํ โ ํ, ํพ = ํ 2 โ ํ 2 , ํฟ = ํ 2 โ ํ 2 . Then 4ํฆ 2 ํด = ํฟํพ ํ 2 . (Note that ํ satisfies ํ < ํ < ํ because the circles intersect in two distinct points.) The line ํํ (the ํฅ-axis) meets ฮฉ at ํถ = (โํ, 0) and ฮ at ํท = (ํ + ํ , 0). The order on ํํ is ํถ,ํ,ํ,ํท. 1. Circumcenter ํ ofโณํดํถํท Since ํถ and ํท lie on the ํฅ-axis, the perpendicular bisector of ํถํท is the vertical line ํฅ = โ where โ = โํ + (ํ + ํ ) 2 = ํ + ํ โ ํ 2 = ํ + ํ 2 . Thus ํ = (โ,ํ) for some ํ. From ํํด = ํํถ we obtain (โโ ํฅ ํด ) 2 + (ํโ ํฆ ํด ) 2 = (โ + ํ) 2 + ํ 2 . Solving for ํ yields ํ =โ ํพ ํ 4ํํฆ ํด ,where ํ = ํ + ํ + ํ = ํ + ํ. The vector v = ํ โ ํด is v = (ํฃ ํฅ ,ํฃ ํฆ ) = (๏ธ ํ ํ 2ํ , โ ํพํํ 4ํ 2 ํฆ ํด )๏ธ . 2. Second intersections ํธ,ํน and midpoint ํ of ํธํน The line ํดํ consists of points ํด + ํกv. Substituting into ฮฉ (using ํดยท ํด = ํ 2 ) gives 2ํกํดยท v + ํก 2 |v| 2 = 0, so the non-zero root is ํก ํธ =โ2ํดยท v/|v| 2 . Compute ํดยท v =โ ํํ 2 =โ ํก ํธ = ํํ |v| 2 . Similarly, for ฮ we use (ํดโ ํ )ยท v =โํ ํ/2 and obtain 34 of 63 IMO2025IMO 2025 Model Proofs ํก ํน = ํ ํ |v| 2 . Hence ํธ = ํด + ํก ํธ v, ํน = ํด + ํก ํน v. The midpoint ํ of ํธํน is ํ = ํด + ํก ํ v with ํก ํ = ํก ํธ + ํก ํน 2 = (ํ + ํ )ํ 2|v| 2 = ํํ 2|v| 2 . 3. Unit vectors along and perpendicular to ํดํ Lete = v/|v|andn = (โํฃ ํฆ ,ํฃ ํฅ )/|v|. Theneis the direction ofํดํ,nis perpendicular toํดํ, andํธํนis parallel toe. The half-length of ํธํน is โ = |ํก ํน โ ํก ํธ ||v| 2 = ํ ํ 2|v| , โ 2 = ํ 2 ํ 2 4|v| 2 . Thus ํธ = ํโ โe, ํน = ํ + โe. 4. Orthocenter ํป ofโณํํํ Points: ํ = (0, 0), ํ = (ํ, 0), ํ = (โ,ํ). The altitude from ํ is the vertical line ํฅ = โ. The altitude fromํis perpendicular toํํ; slope ofํํisโํ/(ํโ โ), so the altitude fromํhas equationํฆ = ํโโ ํ ํฅ . Intersection gives ํป = (๏ธ โ, โ(ํโ โ) ํ )๏ธ . Now โ(ํโ โ) = (ํ+ํ )(ํโํ ) 4 = ํ 2 โํ 2 4 = ํพ 4 . Using ํ =โ ํพํ 4ํํฆ ํด , we obtain ํป = (๏ธ โ, ํพ/4 โํพํ/(4ํํฆ ํด ) )๏ธ = (๏ธ โ, โ ํํฆ ํด ํ )๏ธ . 5. Computation of ํฝ = (ํโ ํต)ยท n and ํพ = (ํโ ํป)ยท n Because ํ = ํด + ํก ํ v and nโฅ v, we have (ํโ ํต)ยท n = (ํดโ ํต)ยท n,(ํโ ํป)ยท n = (ํดโ ํป)ยท n. Thus ํฝ = (ํดโ ํต)ยท n, ํพ = (ํดโ ํป)ยท n. Now ํดโ ํต = (0, 2ํฆ ํด ) and n = (โํฃ ํฆ ,ํฃ ํฅ )/|v|, so ํฝ = 2ํฆ ํด ํฃ ํฅ |v| = 2ํฆ ํด |v| ยท ํ ํ 2ํ = ํฆ ํด ํ ํ ํ|v| . To compute ํพ, we use wยท n =โ wรv |v| (for any w). Hence ํพ =โ (ํดโ ํป)ร v |v| . We compute the cross product. Coordinates: ํฅ ํด โ โ =โ ํ ํ 2ํ , ํฆ ํด โ ํป ํฆ = ํฆ ํด + ํํฆ ํด ํ = ํฆ ํด ํ + ํ ํ , ํฃ ํฅ = ํ ํ 2ํ , ํฃ ํฆ =โ ํพํํ 4ํ 2 ํฆ ํด . Then (ํดโ ํป)ร v = (ํฅ ํด โ โ)ํฃ ํฆ โ (ํฆ ํด โ ํป ํฆ )ํฃ ํฅ = (๏ธ โ ํ ํ 2ํ )๏ธ(๏ธ โ ํพํํ 4ํ 2 ํฆ ํด )๏ธ โ (๏ธ ํฆ ํด ํ + ํ ํ )๏ธ(๏ธ ํ ํ 2ํ )๏ธ = ํ ํพํํ 2 8ํ 3 ํฆ ํด โ ํ ํฆ ํด (ํ + ํ) 2ํ . 35 of 63 IMO2025IMO 2025 Model Proofs Factor ํ 2ํ : (ํดโ ํป)ร v = ํ 2ํ (๏ธ ํพํํ 2 4ํ 2 ํฆ ํด โ ํฆ ํด (ํ + ํ) )๏ธ . Let ํต = ํพํํ 2 4ํ 2 ํฆ ํด โ ํฆ ํด (ํ + ํ). Multiply by 2ํฆ ํด : 2ํฆ ํด ํต = ํพํํ 2 2ํ 2 โ 2ํฆ 2 ํด (ํ + ํ). Using 4ํฆ 2 ํด = ํฟํพ ํ 2 we have ํฆ 2 ํด = ํฟํพ 4ํ 2 , so 2ํฆ 2 ํด (ํ + ํ) = ํฟํพ(ํ + ํ) 2ํ 2 . Thus 2ํฆ ํด ํต = ํพ 2ํ 2 (๏ธ ํํ 2 โ ํฟ(ํ + ํ) )๏ธ . Now ํ = ํ + ํ and ํฟ = ํ 2 โ ํ 2 . Compute ํํ 2 โ ํฟ(ํ + ํ) = ํ (ํ + ํ ) 2 โ (ํ 2 โ ํ 2 )(2ํ + ํ ) = 2ํ 2 ํ. Hence 2ํฆ ํด ํต = ํพ 2ํ 2 ยท 2ํ 2 ํ = ํพํ , so ํต = ํพํ 2ํฆ ํด . Consequently, (ํดโ ํป)ร v = ํ 2ํ ยท ํพํ 2ํฆ ํด = ํ ํพํ 4ํํฆ ํด . Therefore ํพ =โ ํ ํพํ 4ํํฆ ํด |v| . Comparing with ํฝ = ํฆ ํด ํ ํ ํ|v| , we obtain ํพ =โ ํพ 4ํฆ 2 ํด ํฝ. Moreover, using 4ํฆ 2 ํด = ํฟํพ ํ 2 we can write ํพ 4ํฆ 2 ํด = ํ 2 ํฟ =โ ํพ =โ ํ 2 ํฟ ํฝ. 6. Expression for|ํโ ํต| 2 Write ํดโ ํต = ํผe + ํฝn, where ํผ = (ํดโ ํต)ยท e = (ํดโ ํต)ยท v |v| . Compute (ํดโ ํต)ยท v = 2ํฆ ํด ํฃ ํฆ = 2ํฆ ํด (๏ธ โ ํพํํ 4ํ 2 ํฆ ํด )๏ธ =โ ํพํํ 2ํ 2 , so ํผ =โ ํพํํ 2ํ 2 |v| . Since ํ = ํด + ํก ํ v and v =|v|e, ํโ ํต = (ํดโ ํต) + ํก ํ v = (ํผe + ํฝn) + ํก ํ |v|e = (ํผ + ํก ํ |v|)e + ํฝn. Now ํก ํ |v| = ํํ 2|v| . Thus ํผ + ํก ํ |v| =โ ํพํํ 2ํ 2 |v| + ํํ 2|v| = ํํ 2|v| (๏ธ 1โ ํพ ํ 2 )๏ธ = ํํ 2|v| ยท ํ 2 ํ 2 = ํํ 2 ํ 2ํ 2 |v| . 36 of 63 IMO2025IMO 2025 Model Proofs Hence |ํโ ํต| 2 = (ํผ + ํก ํ |v|) 2 + ํฝ 2 = ํ 2 ํ 4 ํ 2 4ํ 4 |v| 2 + ํฆ 2 ํด ํ 2 ํ 2 ํ 2 |v| 2 . Substitute ํฆ 2 ํด = ํฟํพ 4ํ 2 to obtain |ํโ ํต| 2 = ํ 2 ํ 2 |v| 2 (๏ธ ํ 2 ํ 2 4ํ 4 + ํฟํพ 4ํ 4 )๏ธ = ํ 2 ํ 2 4ํ 4 |v| 2 (๏ธ ํ 2 ํ 2 + ํฟํพ )๏ธ . 7. Circumcenter ํ ofโณํตํธํน and the tangency condition The perpendicular bisector ofํธํนis the line throughํparallel ton, so we can writeํ = ํ + ํnfor some realํ. The circumradius ํ satisfies ํ 2 =|ํโ ํธ| 2 = โ 2 + ํ 2 . The condition|ํโ ํต| 2 = ํ 2 gives |(ํโ ํต) + ํn| 2 = โ 2 + ํ 2 . Expanding, and using (ํโ ํต) = ํผ โฒ e + ํฝn with ํผ โฒ = ํผ + ํก ํ |v| (so that|ํโ ํต| 2 = ํผ โฒ2 + ํฝ 2 ), we get ํผ โฒ2 + (ํฝ + ํ) 2 = โ 2 + ํ 2 =โ |ํโ ํต| 2 + 2ํํฝ = โ 2 .(1) Thus 2ํํฝ = โ 2 โ|ํโ ํต| 2 . Now consider the line through ํป parallel to ํดํ . Its distance to ํ equals the circumradius ofโณํตํธํน iff โ โ (ํโ ํป)ยท n โ โ = ํ. Since ํโ ํป = (ํโ ํป) + ํn and (ํโ ํป)ยท n = ํพ, we have (ํโ ํป)ยท n = ํพ + ํ. Therefore tangency is equivalent to (ํพ + ํ) 2 = ํ 2 + โ 2 โโ ํพ 2 + 2ํพํ = โ 2 .(2) Substituting ํ = โ 2 โ|ํโ ํต| 2 2ํฝ from (1) into (2) yields ํพ 2 + ํพ ํฝ (๏ธ โ 2 โ|ํโ ํต| 2 )๏ธ = โ 2 . Multiplying by ํฝ and rearranging gives ํพ (๏ธ โ 2 โ|ํโ ํต| 2 )๏ธ = ํฝ (๏ธ โ 2 โ ํพ 2 )๏ธ .(3) 8. Verification of (3) We now use the explicit expressions: ํฝ = ํฆ ํด ํ ํ ํ|v| , ํพ =โ ํ 2 ํฟ ํฝ, โ 2 = ํ 2 ํ 2 4|v| 2 , |ํโ ํต| 2 = ํ 2 ํ 2 4ํ 4 |v| 2 (๏ธ ํ 2 ํ 2 + ํฟํพ )๏ธ . Set ํ =โ ํ 2 ํฟ , so that ํพ = ํํฝ. Then (3) becomes (after dividing by the non-zero factor ํฝ) ํ (๏ธ โ 2 โ|ํโ ํต| 2 )๏ธ = โ 2 โ ํ 2 ํฝ 2 .(3โ) Compute each side: ํ 2 ํฝ 2 = ํ 4 ํฟ 2 ยท ํฆ 2 ํด ํ 2 ํ 2 ํ 2 |v| 2 = ํ 4 ํฟ 2 ยท ํฟํพํ 2 ํ 2 4ํ 4 |v| 2 = ํพํ 2 ํ 2 4ํฟ|v| 2 . Hence โ 2 โ ํ 2 ํฝ 2 = ํ 2 ํ 2 4|v| 2 โ ํพํ 2 ํ 2 4ํฟ|v| 2 = ํ 2 ํ 2 4|v| 2 (๏ธ 1โ ํพ ํฟ )๏ธ = ํ 2 ํ 2 4|v| 2 ยท ํฟโ ํพ ํฟ . Next, 37 of 63 IMO2025IMO 2025 Model Proofs โ 2 โ|ํโ ํต| 2 = ํ 2 ํ 2 4|v| 2 โ ํ 2 ํ 2 4ํ 4 |v| 2 (๏ธ ํ 2 ํ 2 + ํฟํพ )๏ธ = ํ 2 ํ 2 4|v| 2 [๏ธ 1โ ํ 2 ํ 2 + ํฟํพ ํ 4 ]๏ธ . Thus the left-hand side of (3โ) is ํ (๏ธ โ 2 โ|ํโ ํต| 2 )๏ธ =โ ํ 2 ํฟ ยท ํ 2 ํ 2 4|v| 2 [๏ธ 1โ ํ 2 ํ 2 + ํฟํพ ํ 4 ]๏ธ =โ ํ 2 ํ 2 4ํฟ|v| 2 (๏ธ ํ 2 โ ํ 2 ํ 2 + ํฟํพ ํ 2 )๏ธ . Equation (3โ) is therefore equivalent (after multiplying both sides by 4ํฟ|v| 2 /ํ 2 ํ 2 ) to โํ 2 + ํ 2 ํ 2 + ํฟํพ ํ 2 = ํฟโ ํพ. Multiply by ํ 2 : โํ 4 + ํ 2 ํ 2 + ํฟํพ = (ํฟโ ํพ)ํ 2 . Now substitute ํฟ = ํ 2 โ ํ 2 , ํพ = ํ 2 โ ํ 2 . The left side expands to โํ 4 + ํ 2 ํ 2 + (ํ 2 โ ํ 2 )(ํ 2 โ ํ 2 ) =โํ 4 + ํ 2 ํ 2 + ํ 2 ํ 2 โ ํ 2 ํ 2 โ ํ 4 + ํ 2 ํ 2 = ํ 2 (ํ 2 + ํ 2 โ 2ํ 2 ). The right side is (ํฟโ ํพ)ํ 2 = (๏ธ (ํ 2 โ ํ 2 )โ (ํ 2 โ ํ 2 ) )๏ธ ํ 2 = (ํ 2 + ํ 2 โ 2ํ 2 )ํ 2 , which is identical. Hence (3โ) holds, and consequently (3) and (2) are satisfied. Therefore the line through ํป parallel to ํดํ is tangent to the circumcircle ofโณํตํธํน .โก LLM Judge Comment Instruction: adopt from ProofBench(Ma et al., 2025) With Reference Solution and Marking Scheme (Strict) You are an **expert math proof grader**. You are judging the correctness of an LLM-generated proof for a math problem. ### Input Your input will consist of: * **Problem Statement**: A mathematical problem that the proof is attempting to solve. * **Reference Solution**: A correct solution or proof provided for reference. This is **not necessarily the only valid solution**. If the problem requires a final numeric or algebraic answer, this section contains the correct answer, which should be the only accepted final answer (though alternative reasoning paths are valid). * **Marking Scheme**: A problem-specific grading rubric (0-7 scale) with checkpoints, zero-credit items, and deductions. You must follow this scheme when assigning points. * **Proof Solution**: The proof that you need to evaluate. This proof may contain errors, omissions, or unclear steps. The proof was generated by another language model. ### Task Analyze the proof carefully. * Compare the proof against the reference solution and the marking scheme. * Award points according to the marking schemeโs checkpoints, zero-credit items, and deductions. * Identify logical errors, incorrect steps, or unclear reasoning. * Give a score between 0 and 7 with a brief overall assessment. * Show clearly how the score was derived: * Which checkpoints were earned (with awarded points). Grade this as an IMO marker using ultra-strict standards. Do not fill in missing steps, do not repair the proof, and do not infer claims that are not explicitly proved. Treat any later use of an unstated or unreproved bound as a gap. First list every logical dependency line by line, then check whether each dependency has been established earlier in the written solution. Distinguish clearly between: (1) correct idea, (2) complete written proof, and (3) patchable gap. Score according to the written proof, not the intended argument problem: problem Reference Solution: reference solution 38 of 63 IMO2025IMO 2025 Model Proofs Marking Scheme: marking scheme proof: proof โโ GPT-5.4-Thinking(Extensive) Thought for 2m 35s> Here is the grading, done as an ultra-strict IMO marker on the written proof only. Overall judgment: the submitted proof is mathematically sound and, despite being very calculation-heavy, it is a complete coordinate proof. It does not follow the reference solutionโs synthetic route, but it replaces the official constructions by explicit coordinate/algebraic identities and closes the tangency criterion directly. I do not see a fatal logical gap. Dependency audit 1. The proof fixes coordinates (M=(0,0)), (N=(d,0)), with circles ฮฉ : ํฅ 2 + ํฆ 2 = ํ 2 ,ฮ : (ํฅโ ํ) 2 + ํฆ 2 = ํ 2 ,0 < ํ < ํ . This is a valid coordinate setup. No gap. 2. It writes ํด = (ํฅ ํด ,ํฆ ํด ), ํต = (ํฅ ํด ,โํฆ ํด ), and derives ํฅ ํด = ํ 2 + ํ 2 โ ํ 2 2ํ , ํฆ 2 ํด = ((ํ + ํ) 2 โ ํ 2 )(ํ 2 โ (ํ โ ํ) 2 ) 4ํ 2 . This follows from solving the two circle equations and using symmetry about the (x)-axis. Established. 3. It defines ํ = ํ + ํ, ํ = ํ โ ํ, ํพ = ํ 2 โ ํ 2 , ํฟ = ํ 2 โ ํ 2 , and notes (V<d<U) because the circles intersect in two distinct points. Hence(ํพ > 0), (ํฟ > 0), (ํฆ ํด > 0). Established. 4. It identifies ํถ = (โํ, 0), ํท = (ํ + ํ , 0), with order (C,M,N,D). Correct and established. 5. It finds the circumcenter (P) of (โณ ACD) as (P=(h,p)), where โ = โํ + (ํ + ํ ) 2 = ํ + ํ 2 . This uses that the perpendicular bisector of (CD) is vertical. Established. 6. From (PA=PC), it states ํ =โ ํพํ 4ํํฆ ํด , ํ = ํ + ํ. This is a compressed algebra step, but it is a direct consequence of the displayed equation and the earlier definitions. I checked it; it is correct. No logical gap. 7. It defines (v = ํ โ ํด) and gives v = (ํฃ ํฅ ,ํฃ ํฆ ) = (๏ธ ํ ํ 2ํ ,โ ํพํํ 4ํ 2 ํฆ ํด )๏ธ . Again, compressed algebra, but correct and dependent only on already established formulas. 8. It parametrizes line (AP) as (ํด + ํกv). Substituting into (ฮฉ), it derives the second intersection parameter ํก ํธ = ํํ |v| 2 . This follows from 2ํก,ํดยท v + ํก 2 |v| 2 = 0 39 of 63 IMO2025IMO 2025 Model Proofs and the computed identity (ํดยท v =โํํ/2). Established. 9. Similarly, substituting into (ฮ), it derives ํก ํน = ํ ํ |v| 2 . This depends on ((ํดโ ํ )ยท v =โํ ํ/2). Established. 10. Hence ํธ = ํด + ํก ํธ v, ํน = ํด + ํก ํน v, and the midpoint (S) of (EF) is ํ = ํด + ํก ํ v, ํก ํ = ํํ 2|v| 2 . Established. 11. It defines e = v |v| ,n = (โํฃ ํฆ ,ํฃ ํฅ ) |v| , so (e) is the direction of (AP), (nโฅ ํดํ ), and (ํธํน โ ํดํ ). Established. 12. It computes the half-length of (EF): โ = |ํก ํน โ ํก ํธ |,|v| 2 = ํ ํ 2|v| . Correct and established. 13. It computes the orthocenter (H) of (โณํํํ). From the altitude equations it gets ํป = (๏ธ โ, โ(ํโ โ) ํ )๏ธ , and then ํป = (๏ธ โ,โ ํํฆ ํด ํ )๏ธ . This is correct and fully supported by earlier formulas. 14. It defines ํฝ = (ํโ ํต)ยท n, ํพ = (ํโ ํป)ยท n. Using (ํ = ํด + ํก ํ v) and (nโฅ v), it reduces these to ํฝ = (ํดโ ํต)ยท n, ํพ = (ํดโ ํป)ยท n. Established. 15. It computes ํฝ = ํฆ ํด ํ ํ ํ|v| . Correct and established. 16. It computes ํพ =โ (ํดโ ํป)ร v |v| , then works through the cross product and arrives at ํพ =โ ํ ํพํ 4ํํฆ ํด |v| =โ ํพ 4ํฆ 2 ํด ํฝ =โ ํ 2 ํฟ ํฝ. This is the densest calculation in the proof, but every ingredient used has already been introduced, and the final identity checks out. No missing geometric lemma is being smuggled in. 17. It decomposes (A-B) into the ((e, n))-basis, then obtains |ํโ ํต| 2 = ํ 2 ํ 2 4ํ 4 |v| 2 (๏ธ ํ 2 ํ 2 + ํฟํพ )๏ธ . This is a long but direct algebraic consequence of earlier formulas. Established. 40 of 63 IMO2025IMO 2025 Model Proofs 18. It introduces (O), the circumcenter of (โณํตํธํน). Since (ํธํน โ e), the perpendicular bisector of (EF) is the line through (S) parallel to (n), so ํ = ํ + ํn. Correct. 19. Using (|ํํต| =|ํํธ|), it derives |ํโ ํต| 2 + 2ํํฝ = โ 2 .(1) Correct. 20. It then expresses the line through (H) parallel to (AP) as a line with direction (e), so its distance from (O) is (|(ํโ ํป)ยท n|). Tangency is therefore equivalent to (ํพ + ํ) 2 = ํ 2 + โ 2 โโ ํพ 2 + 2ํพํ = โ 2 .(2) This is a valid tangent-to-circle criterion. 21. Eliminating (ํ) between (1) and (2), it obtains ํพ(โ 2 โ|ํโ ํต| 2 ) = ํฝ(โ 2 โ ํพ 2 ).(3) Correct algebra. 22. It substitutes the explicit expressions ํฝ, ํพ, โ 2 , |ํโ ํต| 2 into (3), rewrites with (ํ =โํ 2 /ํฟ), and reduces the whole claim to โํ 4 + ํ 2 ํ 2 + ํฟํพ = (ํฟโ ํพ)ํ 2 . Correct reduction. 23. Finally, substituting (ํฟ = ํ 2 โ ํ 2 ) and (ํพ = ํ 2 โ ํ 2 ), both sides simplify to ํ 2 (ํ 2 + ํ 2 โ 2ํ 2 ), so (3) holds, hence (2) holds, hence the line through (H) parallel to (AP) is tangent to the circumcircle of (โณํตํธํน). Complete. Correct idea vs complete proof vs patchable gaps Correct idea: yes. The idea is to replace the synthetic construction by a coordinate model, compute (P,E,F,H), write the circumcenter (O) of (โณํตํธํน ) in the ((e, n))-frame, and reduce tangency to an explicit scalar identity. Complete written proof: yes. Every major object used later is explicitly defined, and the final tangency condition is actually proved, not merely asserted. Patchable gaps: there are a few compressed algebra jumps (โsolving for (p)โ, the dot-product simplifications, the cross-product simplification, the final reduction in Section 8), but these are routine algebraic consequences of already displayed formulas, not new unproved geometric claims. Under IMO marking, I would not count these as logical gaps. Checkpoint scoring Because this is a full alternative coordinate proof, I map the official checkpoints by equivalent accomplished work rather than by the exact synthetic route. 1. Foundational Angle Chase โ 1/1 Earned by equivalent foundational setup: the proof correctly derives the full coordinate description of (A,B,C,D,P), including the intersection formulas and all sign/positivity conditions needed later. 2. Parallelogram Construction and Alignment โ 2/2 The proof does not use the official (Aโ)-parallelogram route, but it establishes the equivalent control of the (E,F) configuration directly: it computes the line (AP), the second intersections (E,F), the midpoint (S), and the direction/normal frame ((e, n)). This is enough to replace the official CP2 machinery. 3. Properties of Point (T) โ 2/2 The proof bypasses the auxiliary point (T) entirely and instead analyzes the actual circumcenter (O) of (โณํตํธํน ) directly. This is a valid stronger substitute for the official CP3 structure. 41 of 63 IMO2025IMO 2025 Model Proofs 4. Properties of (H) and Final Conclusion โ 2/2 Earned. The proof computes (H) explicitly, computes its signed normal offset (ํพ), and then verifies the tangency criterion exactly. Deductions No cap applies, because the proof reaches the final tangency statement. No (-1) deduction applies. I do not see a nontrivial unproved leap in the main line. Final score 7/7 Brief assessment This is a valid full-score coordinate solution. It is not elegant, and it is denser than the reference synthetic proof, but as written it does close all the required dependencies. Under ultra-strict marking, I would still award full marks because the proof is complete rather than merely suggestive. 42 of 63 IMO2025IMO 2025 Model Proofs Problem 3 LetNdenote the set of positive integers. A functionํ : Nโ Nis said to be bonza if ํ (ํ) divides ํ ํ โ ํ (ํ) ํ(ํ) for all positive integers ํ and ํ. Determine the smallest real constantํsuch thatํ (ํ) โค ํํfor all bonza functions ํ and all positive integers ํ. Short Answer 4 Model Proof We will prove that the smallest constantํsuch thatํ (ํ)โค ํํfor every bonza functionํis ํ = 4. 1. Basic properties From the definition, for all ํ,ํโ N, ํ (ํ)| ํ ํ โ ํ (ํ) ํ(ํ) .(*) โขTakingํ = ํgivesํ (ํ)| ํ ํ โ ํ (ํ) ํ(ํ) . Sinceํ (ํ) ํ(ํ) is a multiple ofํ (ํ), we obtain ํ (ํ)| ํ ํ . Hence every prime divisor of ํ (ํ) divides ํ. โข Takingํ = ํ = 1yieldsํ (1)| 1โํ (1) ํ(1) . The right-hand side is congruent to1modulo ํ (1), so ํ (1)| 1. Thus ํ (1) = 1. 2. Behaviour on primes Letํbe a prime. Because every prime divisor ofํ (ํ)must divideํ, we haveํ (ํ) = ํ ํก for some integer ํกโฅ 0 (where ํก = 0 means ํ (ํ) = 1). Define ํ =ํ prime| ํ (ํ)ฬธ= 1 =ํ| ํกโฅ 1. B1 B1 | ํก should be a function w.r.t ํ, like ํก(ํ) or ํก ํ , instead of a constant If ํโ ํ, then ํ (ํ) = ํ ํก with ํกโฅ 1. Using (*) with ํ = ํ and an arbitrary ํ we get ํ ํก | ํ ํ โ ํ (ํ) ํ ํก . In particular, reducing modulo ํ, we have ํ ํ โก ํ (ํ) ํ ํก (mod ํ). We now show thatํฅ ํ ํก โก ํฅ (mod ํ) for every integerํฅ. This follows by induction onํกusing Fermatโs little theorem: for ํก = 0 it is trivial; assuming ํฅ ํ ํก โก ํฅ (mod ํ), then ํฅ ํ ํก+1 = (ํฅ ํ ํก ) ํ โก ํฅ ํ (mod ํ)โก ํฅ (mod ํ), where the last congruence is Fermatโs little theorem. Henceํ (ํ) ํ ํก โก ํ (ํ) (mod ํ) . Also ํ ํ โก ํ (mod ํ) by Fermat. Therefore ํโ ํ (ํ)โก ํ ํ โ ํ (ํ) ํ ํก โก 0 (mod ํ), so ํ (ํ)โก ํ (mod ํ)for all ํโ N.(1) 43 of 63 IMO2025IMO 2025 Model Proofs 3. When ํ is infinite Assume thatํis infinite. Then for any fixedํ, the integerํ (ํ)โ ํis divisible by every prime inํ. An integer that is divisible by infinitely many distinct primes must be zero. Hence ํ (ํ) = ํfor allํ. Consequently, the only bonza function withํinfinite is the identity ํ (ํ) = ํ. 4. Odd primes in ํ force ํ to be infinite Suppose that an odd primeํbelongs toํ. Letํbe any prime withํ ฬธโก 1 (mod ํ). From (1) we haveํ (ํ) โก ํ (mod ํ). Becauseํ (ํ) | ํ ํ , the only prime divisors ofํ (ํ)areํ; thusํ (ํ)is a power ofํ, sayํ (ํ) = ํ ํ . Ifํ = 0thenํ (ํ) = 1, which would implyํ โก 1 (mod ํ), contradicting the choice of ํ. Hence ํ โฅ 1, and therefore ํ | ํ (ํ), so ํ โ ํ. We now prove that there are infinitely many primesํwithํ ฬธโก 1 (mod ํ). Assume, to the contrary, that only finitely many such primes exist; list them asํ 1 ,ํ 2 ,...,ํ ํ (this list contains ํ and 2 because both are notโก 1 (mod ํ)). Consider ํ = ํยท (ํ 1 ํ 2 ยทํ ํ )โ 1. B2 Clearlyํ > 1. Any prime divisorํofํis different from everyํ ํ , for otherwiseํ B2 | correct, but no need to have a multiply as ํ here, since you already saidํand2in the list. would divideํยท โ๏ธ ํ ํ and hence divideํ + 1, forcingํ | 1, impossible. By our assumption, every prime divisor ofํsatisfiesํ โก 1 (mod ํ). Consequently,ํ โก 1 (mod ํ). But ํ = ํยท โ๏ธ ํ ํ โ 1โกโ1 (mod ํ), a contradiction. Hence there are infinitely many primesํ withํ ฬธโก 1 (mod ํ). By the previous paragraph, each such prime lies inํ, soํis infinite. Then by ยง3, ํ is the identity. Thus, ifํis not the identity, the setํcannot contain any odd prime. Therefore the only possibilities for a non-identity bonza function are (i) ํ = โ or(i) ํ =2. 5. Case ํ = โ This means ํ (ํ) = 1 for every prime ํ. We show that then ํ is the constant function 1. Suppose, for contradiction, that there existsํwithํ (ํ) > 1. Letํbe a prime divisor ofํ (ํ). From ยง1 we haveํ | ํ. Sinceํis prime andํ = โ , we haveํ (ํ) = 1. Apply (*) withํ(the original number) and ํ = ํ: ํ (ํ)| ํ ํ โ ํ (ํ) ํ(ํ) = ํ ํ โ 1. Becauseํ | ํ (ํ), we obtainํ | ํ ํ โ 1. Butํ ํ โก 0 (mod ํ), soํ ํ โ 1 โก โ1 (mod ํ), contradiction. Hence no suchํexists, andํ (ํ) = 1for allํ. The constant function1is clearly bonza because 1 divides any integer. 6. Case ํ =2 Hereํ (2)ฬธ= 1(andํ (2)| 2 2 = 4, soํ (2) = 2or4), whileํ (ํ) = 1for every odd primeํ. We will prove that ํ (ํ)โค 4ํ for all ํ. 6.1 Odd arguments Letํ > 1be odd. Assumeํ (ํ) > 1and letํbe a prime divisor ofํ (ํ). Thenํ | ํ(since ํ (ํ)| ํ ํ ), soํis an odd prime. Becauseํis odd,ํ (ํ) = 1. Using (*) withํ = ํandํ = ํ gives ํ (ํ)| ํ ํ โ 1. Butํ | ํ (ํ)impliesํ | ํ ํ โ 1, which is impossible becauseํ ํ โก 0 (mod ํ). Therefore ํ (ํ) = 1 for all odd ํ. Also ํ (1) = 1. B3 B3 | Logical issue. It should be โTherefore, ํ(ํ) = 1 for odd ํ > 1. Also ํ(1) = 1, so ํ(odd) = 1โ. 6.2 Even arguments Now let ํ be even and write ํ = 2 ํ ยท ํ with ํ = ํฃ 2 (ํ)โฅ 1 and ํ odd. First we show thatํ (ํ)is a power of2. Suppose an odd primeํdividesํ (ํ). Then, as in ยง6.1, ํ | ํ and ํ (ํ) = 1 (since ํ is odd), and (*) with ํ = ํ B4 yields ํ (ํ)| ํ ํ โ 1, leading B4 | missing ํ = ํ to the same contradictionํ | ํ ํ โ 1. Hence all prime factors ofํ (ํ)are2, soํ (ํ) = 2 ํ for some ํ โฅ 0. Now take ํ to be any odd prime ํ B5 . Because ํ (ํ) = 1, condition (*) gives B5 | missing ํ = ํ 44 of 63 IMO2025IMO 2025 Model Proofs 2 ํ | ํ ํ โ 1.(2) Thus ํ โค ํฃ 2 (ํ ํ โ 1) for every odd prime ํ. Consequently, ํ โคmin odd prime ํ ํฃ 2 (ํ ํ โ 1). We compute this minimum. Sinceํis even, writeํ = 2 ํ ยท ํwithํodd B6 . For an odd B6 | already defined at the beginning of 6.2 prime ํ, we have (using the lifting-the-exponent lemma or a direct argument) ํฃ 2 (ํ ํ โ 1) = ํฃ 2 (ํ 2 ํ โ 1) = ํฃ 2 (ํโ 1) + ํฃ 2 (ํ + 1) + ํโ 1.(3) Proof of (3): Becauseํis odd,ํฃ 2 (ํ ํ โ 1) = ํฃ 2 (ํ 2 ํ โ 1). Forํ โฅ 1, one proves by induction thatํฃ 2 (ํ 2 ํ โ 1) = ํฃ 2 (ํโ 1) +ํฃ 2 (ํ + 1) +ํโ 1. The baseํ = 1isํ 2 โ 1 = (ํโ 1)(ํ + 1). For the inductive step,ํ 2 ํ โ 1 = (ํ 2 ํโ1 โ 1)(ํ 2 ํโ1 + 1). The first factor contributesํฃ 2 (ํ 2 ํโ1 โ 1), and forํ โฅ 2one hasํฃ 2 (ํ 2 ํโ1 + 1) = 1becauseํ 2 ํโ1 โก 1 (mod 8)(asํodd impliesํ 2 โก 1 (mod 8), and by induction all higher powers areโก 1 (mod 8)). This gives the formula. For an odd primeํ, the two consecutive even numbersํโ 1andํ + 1have the following property: one of them is divisible by2exactly once, the other is divisible by4(at least). Hence ํฃ 2 (ํโ 1) + ํฃ 2 (ํ + 1)โฅ 1 + 2 = 3, and the minimum3is attained for infinitely many primes, e.g.ํ = 3, 5, 11, 13,.... Therefore min odd prime ํ ํฃ 2 (ํ ํ โ 1) = 3 + ํโ 1 = ํ + 2. Thus ํ โค ํ + 2, and ํ (ํ) = 2 ํ โค 2 ํ+2 = 4ยท 2 ํ . Since ํ = 2 ํ ยท ํ with ํโฅ 1, we have 2 ํ โค ํ. Hence ํ (ํ)โค 4ยท 2 ํ โค 4ํ. For oddํwe already haveํ (ํ) = 1โค 4ํ. The identity functionํ (ํ) = ํalso satisfies this bound. Consequently, every bonza function satisfies ํ (ํ)โค 4ํ for all ํ. B7 B7 | You divided into 3 cases, and here ignored the situation with ํ(ํ) = 1 always. 7. Sharpness We construct a bonza functionํนfor whichํน (4) = 16, showing thatํcannot be smaller than 4. Define ํน (1) = 1, ํน (2) = 2, ํน (ํ) = 1 for odd ํ > 1, and for even ํโฅ 4, let ํ = ํฃ 2 (ํ) and set ํน (ํ) = 4ยท 2 ํ = 2 ํ+2 . (One could also take ํน (2) = 4; both work, but we choose ํน (2) = 2 for definiteness.) We verify that ํน satisfies the bonza condition. โข Case ํ odd: Then ํน (ํ) = 1 and 1| ํ ํ โ ํน (ํ) 1 trivially. โข Case ํ = 2: ํ odd: ํน (2) = 2 divides ํ 2 โ 1 because ํ 2 is odd, so ํ 2 โ 1 is even. ํ even: Both ํ 2 and ํน (ํ) 2 are even, so their difference is even, hence divisible by 2. โขCaseํ โฅ 4even: Writeํ = 2 ํ ยท ํwithํ = ํฃ 2 (ํ) โฅ 1andํodd, soํน (ํ) = 2 ํ+2 . Consider ํ. ํ odd: ํน (ํ) = 1. For any odd ํ and even ํ, we have ํฃ 2 (ํ ํ โ 1) = ํฃ 2 (ํโ 1) + ํฃ 2 (ํ + 1) + ํโ 1โฅ 3 + ํโ 1 = ํ + 2, so 2 ํ+2 | ํ ํ โ 1. ํ even: Writeํน (ํ) = 2 ํก withํกโฅ 1(ifํ = 2thenํก = 1; ifํโฅ 4thenํก = ํฃ 2 (ํ) + 2โฅ 3). 45 of 63 IMO2025IMO 2025 Model Proofs Then ํฃ 2 (ํ ํ ) = ํยท ํฃ 2 (ํ)โฅ ํโฅ ํ + 2 (since ํโฅ 4 and even), and ํฃ 2 (๏ธ ํน (ํ) ํน(ํ) )๏ธ = ํกยท ํน (ํ) = ํกยท 2 ํ+2 โฅ 2 ํ+2 . B8 Hence bothํ ํ andํน (ํ) ํน(ํ) are multiples of2 ํ+2 , and therefore their difference is B8 | ํ + 2 should be enough also a multiple of 2 ํ+2 . Thus in all situationsํน (ํ) | ํ ํ โ ํน (ํ) ํน(ํ) . Thereforeํนis bonza. Forํ = 4we have ํฃ 2 (4) = 2, soํน (4) = 2 4 = 16, givingํน (4)/4 = 4. Hence any constantํwithํ (ํ)โค ํํfor all bonza functions must satisfy ํโฅ 4. 8. Conclusion We have shown that every bonza functionํsatisfiesํ (ํ)โค 4ํfor allํโ N, and we have exhibited a bonza function for which equalityํ (4) = 4ยท 4is attained (i.e.,ํ (4)/4 = 4). Consequently, the smallest possible constant is 4. Human Expert Comment 7/7 46 of 63 IMO2025IMO 2025 Model Proofs Problem 4 A proper divisor of a positive integerํis a positive divisor ofํother than ํ itself. The infinite sequenceํ 1 ,ํ 2 ,ยทconsists of positive integers, each of which has at least three proper divisors. For eachํโฅ 1, the integerํ ํ+1 is the sum of the three largest proper divisors of ํ ํ . Determine all possible values of ํ 1 . Short Answer The set of all possible values forํ 1 are integersํof the formํ = 2 ํฅ 3 ํฆ ํ, where: 1.ํฅis a positive odd integer. 2.ํฆis a positive integer satisfying ํฆ โฅ ํฅ+1 2 . 3.ํis a positive integer whose prime factors are all greater than or equal to 7. Model Proof 1. Preliminaries For a positive integerํwith at least three proper divisors, list all its divisors in increasing order: 1 = ํ 1 < ํ 2 < ํ 3 < ํ 4 <ยท < ํ ํ = ํ. Because of the pairing ํ ํ ยท ํ ํ+1โํ = ํ, the three largest proper divisors are ํ ํ 2 , ํ ํ 3 , ํ ํ 4 . Hence we can write ํ (ํ ) := sum of the three largest proper divisors of ํ = ํ (๏ธ 1 ํ 2 + 1 ํ 3 + 1 ํ 4 )๏ธ . (1) C1 C1 | In the correct formulation, the ํ 2 ,ํ 3 ,ํ 4 should be function of ํ. The sequence is defined byํ ํ+1 = ํ (ํ ํ ), and we require that every term has at least three proper divisors. 2. Two key lemmas Lemma 1 (odd numbers). If ํ is odd and has at least three proper divisors, then ํ (ํ ) < ํ and ํ (ํ ) is odd. Proof. Because all divisors of an odd number are odd, we have ํ 2 โฅ 3, ํ 3 โฅ 5, ํ 4 โฅ 7. Hence 1 ํ 2 + 1 ํ 3 + 1 ํ 4 โค 1 3 + 1 5 + 1 7 = 71 105 < 1, soํ (ํ ) < ํ. Moreover, each ofํ/ํ 2 , ํ/ํ 3 , ํ/ํ 4 is odd (odd divided by odd), so their sum is odd.โก Lemma 2 (even numbers not divisible by 3). If ํ is even, 3 โค ํ, and has at least three proper divisors, then (i) ํ (ํ ) < ํ, and (i) ํ (ํ ) is not a multiple of 6. Proof. Sinceํis even,ํ 2 = 2. The smallest possible values forํ 3 ,ํ 4 that maximize the sum in (1) areํ 3 = 4andํ 4 = 5(e.g.,ํ = 20). In any caseํ 3 โฅ 4andํ 4 โฅ 5(because4is the only 47 of 63 IMO2025IMO 2025 Model Proofs even number between2and5, and5is the smallest integer >4 that does not force a factor 3). Consequently 1 2 + 1 ํ 3 + 1 ํ 4 โค 1 2 + 1 4 + 1 5 = 19 20 < 1, so ํ (ํ ) < ํ. We now show that ํ (ํ ) cannot be divisible by 6. Write ํ = 2 ํ ยท ํ with ํ odd and 3 โค ํ (ํโฅ 1). Consider two cases. Case ํโฅ 2. Then 4| ํ, so ํ 3 = 4. Hence ํ (ํ ) = ํ 2 + ํ 4 + ํ ํ 4 = 3ํ 4 + ํ ํ 4 . Becauseํ/4is an integer,3ํ/4is a multiple of3. Sinceํis not divisible by3,ํ/ํ 4 is also not divisible by3. Thusํ (ํ )โก ํ/ํ 4 ฬธโก 0 (mod 3), and thereforeํ (ํ )is not a multiple of 3; in particular it cannot be a multiple of 6. Case ํ = 1. Thenํ = 2ํwithํodd,3 โค ํ. Letํbe the smallest prime divisor ofํ; thenํโฅ 5and ํ 3 = ํ. The fourth divisor ํ 4 is either โข an odd divisor ํ with ํ < ํ < 2ํ (if such a divisor exists), or โข 2ํ (if no odd divisor lies between ํ and 2ํ). We examine the two possibilities. โข Subcase ํ 4 odd. Thenํ/ํ 4 = 2(ํ/ํ)is even,ํ/2 = ํis odd, andํ/ํ = 2(ํ/ํ)is even. Hence ํ (ํ ) = odd + even + even = odd, soํ (ํ )is not even and certainly not a multiple of6. โข Subcase ํ 4 = 2ํ (even). Here ํ (ํ ) = ํ 2 + ํ ํ + ํ 2ํ = ํ + 2 ํ ํ + ํ ํ = ํ + 3 ํ ํ . Sinceํis not divisible by3, we haveํ (ํ ) โก ํ (mod 3) ฬธ= 0. Thusํ (ํ )is not a multiple of 3 and therefore not a multiple of 6. In every subcase ํ (ํ ) is not a multiple of 6, completing the proof.โก 3. All terms must be multiples of 6 Assume, for contradiction, that some term ํ ํ is not a multiple of 6. Ifํ ํ is odd, Lemma 1 gives thatํ ํ+1 = ํ (ํ ํ )is odd andํ ํ+1 < ํ ํ . Repeating, we obtain an infinite strictly decreasing sequence of positive odd integers - impossible by the well-ordering principle. Ifํ ํ is even but3 โค ํ ํ , Lemma 2 tells us thatํ ํ+1 = ํ (ํ ํ ) < ํ ํ andํ ํ+1 is not a multiple of 6. There are two possibilities for ํ ํ+1 : โข ํ ํ+1 is odd. Then by Lemma 1 all subsequent terms are odd and strictly decreasing - again impossible. โข ํ ํ+1 is even and3 โค ํ ํ+1 . Then we can apply the same reasoning toํ ํ+1 : we obtain ํ ํ+2 < ํ ํ+1 andํ ํ+2 not a multiple of6, and so on. This would produce an infinite strictly decreasing sequenceํ ํ > ํ ํ+1 > ํ ํ+2 > ...of positive integers, which cannot exist. Thus, in every case, the assumption leads to a contradiction. Consequently, every term of the infinite sequence is a multiple of 6. 4. Reduction to the form 6ํ Writeํ ํ = 6ํ ํ withํ ํ โ N. Because6ํ ํ always has at least the divisors1, 2, 3, 6, we have ํ (6ํ ํ )โฅ 4 (i.e., at least three proper divisors). 48 of 63 IMO2025IMO 2025 Model Proofs Forํ = 6ํ, the three smallest divisors greater than1are alwaysํ 2 = 2andํ 3 = 3. The fourth divisor ํ 4 depends on ํ: โข Ifํis evenโ 4| 6ํ, and since4is the smallest integer larger than3, we haveํ 4 = 4. โข If ํ is odd and 5| ํ โ 5| 6ํ and 5 < 6, so ํ 4 = 5. โขIfํis odd and5 โค ํ โthe next divisor is6(because6 | 6ํand neither4nor5 divides 6ํ), so ํ 4 = 6. Using (1) we compute ํ (6ํ ) = 6ํ 2 + 6ํ 3 + 6ํ ํ 4 = 3ํ + 2ํ + 6ํ ํ 4 = 5ํ + 6ํ ํ 4 .(2) Now analyse the three cases. 1. ํ odd, 5 โค ํ โ ํ 4 = 6 ํ (6ํ ) = 5ํ + ํ = 6ํ. Hence 6ํ is a fixed point. 2. ํ odd, 5| ํ โ ํ 4 = 5 ํ (6ํ ) = 5ํ + 6ํ 5 = 31 5 ํ. Sinceํis odd and divisible by5,ํ/5is odd, so the result is odd. Therefore it is not a multiple of 6. By the result of ยง3, such a number cannot appear in an infinite sequence. 3. ํ evenโ ํ 4 = 4 ํ (6ํ ) = 5ํ + 6ํ 4 = 13 2 ํ. For this to be a multiple of 6 (necessary for the next term to be admissible) we need 13 2 ํ โก 0 (mod 6) โโ 13ํ โก 0 (mod 12) โโ ํ โก 0 (mod 12), C2 C2 | Notation inconsistency. Just useโค and|. because13 โก 1 (mod 12). Hence, ifํis even but not divisible by12, thenํ (6ํ ) would not be a multiple of6, contradicting ยง3. Therefore, in an infinite sequence, every even ํ must be divisible by 12, and then ํ ํ+1 = ํ (6ํ ) = 6ยท (๏ธ 13ยท ํ 12 )๏ธ , so the new parameter is ํ โฒ = 13ยท (ํ/12). 5. Characterising admissible ํ 1 = ํ 1 /6 Letํ 1 = ํ 1 /6. We have shown that all terms must be multiples of6, and the recurrence for the corresponding ํ ํ is: โข If ํ is odd and 5 โค ํ โ fixed point. โข If ํ is even and divisible by 12โ ํ โฆโ 13ยท (ํ/12). โข Any other situation leads to a contradiction. For the sequence to be infinite, we must start fromํ 1 and, after finitely many applications of the even step C3 , reach an oddํthat is not divisible by5. This forcesํ 1 to have a very C3 | even step? 1. Second cases ํ โ 13ํ/12. 2. even ํ specific structure. Let ํ be the largest integer such that 12 ํ | ํ 1 (the exponent of 12 in ํ 1 ). Write ํ 1 = 12 ํ ยท ํ, 49 of 63 IMO2025IMO 2025 Model Proofs whereํis a positive integer not divisible by12(i.e.,ํis the โremainderโ after removing all factors of 12). We claim that the sequence is infinite iff ํ is odd and 5 โค ํ. Sufficiency Assume ํ 1 = 12 ํ ยท ํ with ํ โฅ 0, ํ odd, and 5 โค ํ. We prove by induction that ํ ํ+1 = 13 ํ ยท ํ 1 12 ํ = 12 ํโํ ยท 13 ํ ยท ํ(0โค ํโค ํ). For ํ = 0 this is the definition of ํ 1 . Inductive step: Forํ < ํ, we haveํ ํ = 12 ํโํ ยท 13 ํ ยทํ. Sinceํโํโฅ 1,ํ ํ is even and divisible by12(because it contains the factor12 ํโํ ). Hence we may apply the even-step rule, obtaining ํ ํ+1 = 13ยท ํ ํ 12 = 13ยท 12 ํโํโ1 ยท 13 ํ ยท ํ = 12 ํโํโ1 ยท 13 ํ+1 ยท ํ, which matches the formula. Thus forํ = ํโ 1we getํ ํ = 12 1 ยท 13 ํโ1 ยท ํ, which is even and divisible by12. Applying the rule once more gives ํ ํ+1 = 13ยท ํ ํ 12 = 13 ํ ยท ํ, which is odd (since13 ํ is odd andํis odd) and not divisible by5. By case 1 of ยง4,6ํ ํ+1 is a fixed point. Consequently, ํ ํ+1 = 6ยท 13 ํ ยท ํ, k+2and the sequence becomes constant thereafter. All terms have at least three proper divisors, so the sequence is infinite. Necessity Suppose, for contradiction, thatํ 1 does not have the form12 ํ ยท ํwithํodd and5 โค ํ. Writeํ 1 = 12 ํ ยท ํas above (nowํmay be even, or odd but divisible by5). C4 We consider C4 | This formulation should mention 12โค ํ two cases. โข ํ is even. C5 C5 | No need to consider for even or odd, we can directly say ํ ํ+1 = 13 ํ ยท ํ and then contradictory. We already have the dynamic of ํ ํ at the beginning. Thenํ 1 is even. Ifํ = 0, thenํ 1 is even but not divisible by12(becauseํeven and not a multiple of12). By ยง4, case 3, this would giveํ 2 not a multiple of6- contradiction. If ํ โฅ 1, then ํ 1 is even and divisible by 12. Apply the even step once to obtain ํ 2 = 13ยท ํ 1 12 = 13ยท 12 ํโ1 ยท ํ. Since ํ is even, ํ 2 is even. We now check whether ํ 2 is divisible by 12. โ Ifํโ 1โฅ 1(i.e.,ํ โฅ 2), then12 ํโ1 already contains a factor12, soํ 2 is divisible by 12. But then we can continue the process. However, after exactlyํsteps we will obtain ํ ํ+1 = 13 ํ ยท ํ, which is even (becauseํis even). Sinceํis not divisible by12(by the definition of ํ),ํ ํ+1 is not divisible by12. Hence at that moment we would have an evenํthat is not divisible by12, which forcesํ (6ํ ํ+1 )not to be a multiple of6- contradiction. โIfํ = 1, thenํ 1 = 12ยทํwithํeven. Thenํ 2 = 13ํis even but not divisible by12 (because13ํlacks the necessary factor3and C6 the factor4ifํis not a multiple of C6 | or 12). Again, ํ 2 fails the requirement that every even ํ must be divisible by 12. In all subcases, contradiction arises. โข ํ is odd but 5| ํ. Then after ํ applications of the even step we obtain ํ ํ+1 = 13 ํ ยท ํ, 50 of 63 IMO2025IMO 2025 Model Proofs which is odd and divisible by5. Consequently,ํ ํ+1 = 6ํ ํ+1 falls into case 2 of ยง4, yieldingํ (ํ ํ+1 )odd. As in ยง3, an odd term would lead to an infinite strictly decreasing sequence of odd numbers - contradiction. Thus, the only way to avoid contradictions is that ํ is odd and 5 โค ํ. 6. Final answer Combining everything, we obtain that the initial valueํ 1 can be exactly those numbers of the form ํ 1 = 6ยท 12 ํ ยท ํ(ํ โ N 0 , ํ odd, 5 โค ํ). Equivalently, ํ 1 = 2 2ํ+1 ยท 3 ํ+1 ยท ํ with ํ odd and not divisible by 5. Human Expert Comment 7/7 51 of 63 IMO2025IMO 2025 Model Proofs Problem 5 Alice and Bazza are playing the inekoalaty game, a two-player game whose rules depend on a positive real numberํwhich is known to both players. On the ํth turn of the game (starting with ํ = 1) the following happens: โข If ํ is odd, Alice chooses a nonnegative real number ํฅ ํ such that ํฅ 1 + ํฅ 2 +ยท + ํฅ ํ โค ํํ. โข If ํ is even, Bazza chooses a nonnegative real number ํฅ ํ such that ํฅ 2 1 + ํฅ 2 2 +ยท + ํฅ 2 ํ โค ํ. If a player cannot choose a suitableํฅ ํ , the game ends and the other player wins. If the game goes on forever, neither player wins. All chosen numbers are known to both players. Determine all values ofํfor which Alice has a winning strategy and all those for which Bazza has a winning strategy. Short Answer The outcome of the inekoalty game depends on the value of ํ as follows: โข If ํ > โ 2 2 , Alice has a winning strategy. โข If ํ < โ 2 2 , Bazza has a winning strategy. โขIfํ = โ 2 2 , neither player has a winning strategy, and the game can go on forever. Model Proof Notation and basic facts Define ํ ํ = โ๏ธ ํ ํ=1 ํฅ ํ and ํ ํ = โ๏ธ ํ ํ=1 ํฅ 2 ํ . The rules: on odd ํ, ํ ํ โค ํํ; on even ํ, ํ ํ โค ํ. From Cauchy-Schwarz, after any even turn 2ํ we have ํ 2 2ํ โค (2ํ)ํ 2ํ โค (2ํ)ยท 2ํ = (2ํ) 2 =โ ํ 2ํ โค 2ํ.(1) For ํ โฅ 1 define the slack after the even turn 2ํ: ํ ํ = ํ(2ํ + 1)โ ํ 2ํ . Alice can move on turn 2ํ + 1 iff ํ ํ โฅ 0 (she may choose ํฅ 2ํ+1 = 0). If at some odd turn Alice can legally choose a number ํข > โ 2 while ํ 2ํ = 2ํ, then ํ 2ํ+1 = 2ํ +ํข 2 > 2ํ + 2 and Bazza will have no legal move on the next turn, so Alice wins immediately. In the general case, we will use the slack to analyse the game. Key strategy for Bazza We consider the following natural strategy for Bazza on even turns: after Aliceโs moveํขon turn 2ํ + 1, if the game has not ended, Bazza chooses ํฃ = โ๏ธ 2ํ + 2โ ํ 2ํ+1 , i.e. the largest possible number that still satisfies ํ 2ํ+2 โค 2ํ + 2. This choice is always legal as long asํ 2ํ+1 โค 2ํ + 2and makesํ 2ํ+2 = 2ํ + 2(the 52 of 63 IMO2025IMO 2025 Model Proofs maximal possible sum of squares). We call this the maximal strategy for Bazza. Claim 1. As long as Alice never picks a numberํข > โ 2(which would make her win immediately), the maximal strategy forces ํ 2ํ = 2ํ for every ํ โฅ 1. Proof. By induction. Forํ = 1, after Aliceโs first moveํฅ 1 = ํ(with0โค ํโค ํ), Bazza can chooseํฃ = โ 2โ ํ 2 becauseํ 2 โค 2(this holds for all relevantํ; in the cases where we apply it, we haveํโค โ 2/2 < โ 2, and for largerํAlice would win earlier). Thenํ 2 = ํ 2 +ํฃ 2 = 2. Inductive step: assumeํ 2ํ = 2ํ. Aliceโs moveํขsatisfiesํ 2ํ+1 = 2ํ +ํข 2 โค 2ํ + 2because otherwiseํข 2 > 2and Alice would have won (this can only happen ifํ 2ํ = 2ํand ํข > โ 2). D1 If the game continues, we must haveํขโค โ 2. Thenํ 2ํ+1 = 2ํ + ํข 2 โค 2ํ + 2. Do not understand. Reads like a thinking draft process. Bazza chooses ํฃ = โ๏ธ 2ํ + 2โ (2ํ + ํข 2 ) = โ 2โ ํข 2 , yielding ํ 2ํ+2 = 2ํ + 2.โก Analysis of the slack under the maximal strategy Suppose the maximal strategy is followed and that the game has not ended (so far Alice has never picked ํข > โ 2). Let ํข = ํฅ 2ํ+1 and ํฃ = ํฅ 2ํ+2 . Then ํ 2ํ+2 = ํ 2ํ + ํข + ํฃ, ํ 2ํ+2 = 2ํ + 2, and therefore ํ ํ+1 = ํ(2ํ + 3)โ ํ 2ํ+2 = ํ ํ + 2ํโ (ํข + ํฃ).(2) For0 โค ํข โค โ 2defineโ(ํข) = ํข + โ 2โ ํข 2 . One checks thatโ(ํข) โฅ โ 2, with equality exactly at ํข = 0 and ํข = โ 2, and maximum 2 at ํข = 1. Consequently ํ ํ+1 โค ํ ํ + 2ํโ โ 2,(3) and equality can be achieved by choosing ํข = 0 (or ํข = โ 2, if allowed). Thus the maximal possible increase ofํ ํ (when Alice tries to keep her slack large) is2ํโ โ 2, attained by taking ํข = 0. Three regimes 1. ํ < โ 2 2 Then 2ํโ โ 2 < 0. Bazzaโs winning strategy. Bazza adopts the maximal strategy described above. First, compute an upper bound for ํ 1 . After the first two moves, ํ 2 = ํ + โ 2โ ํ 2 with0โค ํโค ํ D2 . The functionํ (ํ) = ํ + โ 2โ ํ 2 satisfiesํ (ํ)โฅ โ 2 D2 | Better to use ํฅ 1 not ํ (minimum at ํ = 0), so ํ 1 = 3ํโ ํ 2 โค 3ํโ โ 2.(4) Because ํ < โ 2/2, we have 3ํโ โ 2 < โ 2/2 < โ 2. Now, from (3) we obtain for any possible play ํ ํ โค ํ 1 + (ํโ 1)(2ํโ โ 2)โค (3ํโ โ 2) + (ํโ 1)(2ํโ โ 2).(5) (The first inequality follows by induction usingํ ํ+1 โค ํ ํ + 2ํโ โ 2, the second uses the bound on ํ 1 .) Since2ํโ โ 2 < 0, the right-hand side of (5) tends toโโasํ โโ. Hence there exists a finite ํพ with ํ ํพ < 0. Whenํ ํพ < 0, at the beginning of turn2ํพ + 1Alice cannot move (evenํฅ = 0would violate the sum condition), so Bazza wins. Moreover, we claim that Alice never gets the chance to win earlier by pickingํข > โ 2. Indeed, from (5) and2ํโ โ 2 < 0, we haveํ ํ โค 3ํโ โ 2 < โ 2for allํ. Henceํ ํ never exceeds โ 2, so she cannot legally pick a number larger than โ 2(which would be required to win immediately). Thus Bazzaโs strategy is winning. Conclusion for ํ < โ 2/2: Bazza has a winning strategy. 53 of 63 IMO2025IMO 2025 Model Proofs 2. ํ = โ 2 2 Here 2ํโ โ 2 = 0. Bazza does not have a winning strategy. We exhibit a strategy for Alice that prevents Bazza from ever winning. Let Alice always choose ํฅ 1 = 0 and thereafter ํฅ 2ํ+1 = 0 on every odd turn. Consider any Bazza moves. On even turns we have only Bazzaโs numbers, sayํฆ 1 ,...,ํฆ ํ at turn2ํ. The constraints are โ๏ธ ํ ํ=1 ํฆ 2 ํ = ํ 2ํ โค 2ํ (because odd terms are zero). By Cauchy-Schwarz, ํ 2ํ = ํ โ๏ธ ํ=1 ํฆ ํ โค โฏ โธ โธ โท ํ ํ โ๏ธ ํ=1 ํฆ 2 ํ โค โ ํยท 2ํ = โ 2ํ. Therefore ํ ํ = ํ(2ํ + 1)โ ํ 2ํ โฅ โ 2 2 (2ํ + 1)โ โ 2ํ = โ 2 2 > 0. Thus Alice can always move (she may choose0). Also, she never wins because to win she would need to makeํ 2ํ+1 > 2ํ + 2, which would requireํฅ 2 2ํ+1 > 2ํ + 2โ ํ 2ํ โฅ 2, i.e.ํฅ 2ํ+1 > โ 2. Butํฅ 2ํ+1 โค ํ ํ and we have not shown an upper bound onํ ํ ; D4 however, D4 | This conclusion only focuses on a specific ํตโs policy, which already discussed at the beginning. This looks like thinking draft process. the crucial point is that with this Alice strategy, Bazza never wins because Alice never loses. So this suffices to prove that Bazza does not have a winning strategy (he cannot force a win against every Alice strategy). Alice does not have a winning strategy. We now show that Bazza has a strategy to prevent Alice from winning. Let Bazza adopt the maximal strategy (as in Case 1). We prove that under this strategy, Alice can never win. First, by Claim 1, as long as the game continues,ํ 2ํ = 2ํ. Using (3) withํ = โ 2/2we get ํ ํ+1 = ํ ํ + โ 2โ (ํข + ํฃ)โค ํ ํ ,(6) because ํข + ํฃ โฅ โ 2. Hence ํ ํ is non-increasing. Moreover, from (4) we have ํ 1 โค 3ํโ โ 2 = โ 2/2. Thusํ ํ โค โ 2/2 < โ 2for allํ. In particular,ํ ํ never exceeds โ 2, so Alice can never pick a number ํข > โ 2 (which would be necessary to win immediately). Could Alice win by picking someํขโค โ 2and later build up? To win, she must at some odd turn makeํ 2ํ+1 > 2ํ + 2. Sinceํ 2ํ = 2ํ, this would requireํข 2 > 2, i.e.ํข > โ 2. Thus any winning move would necessarily beํข > โ 2. Becauseํ ํ โค โ 2/2 < โ 2, such a move is impossible. Therefore Alice never wins. D5 D5 | ??? Reads like a thinking draft process. Additionally, we check thatํ ํ never becomes negative, so Alice always can move. Indeed, using (2) and the fact thatํขโค ํ ํ (soํขโค โ 2/2 < 1), the smallest possibleํ ํ+1 for a given ํ ํ is attained when ํข = ํ ํ , giving ํ ํ+1 = โ 2โ โ๏ธ 2โ ํ 2 ํ โฅ 0. Hence ํ ํ โฅ 0 for all ํ. D6 D6 | No need. The proof is finished. Thus the maximal strategy by Bazza prevents Alice from winning, and also avoids his own loss (the game can continue indefinitely, e.g., if Alice always chooses0, thenํ ํ = โ 2/2 constant andํ 2ํ = 2ํ, so the play never ends). Consequently, Alice cannot force a win; she has no winning strategy. Since neither player has a strategy that forces a win, the game is a draw for ํ = โ 2/2. 3. ํ > โ 2 2 Now 2ํโ โ 2 > 0. Aliceโs winning strategy. She will always choose 0 on her odd turns, until she can force a win. Specifically, on turn1she setsํฅ 1 = 0. On each subsequent odd turn2ํ + 1(ํ โฅ 1), she examines the current slackํ ํ = ํ(2ํ + 1)โ ํ 2ํ . Ifํ ํ > โ 2ํ + 2โ ํ 2ํ , she chooses a 54 of 63 IMO2025IMO 2025 Model Proofs number ํข with โ๏ธ 2ํ + 2โ ํ 2ํ < ํขโค ํ ํ , such thatํข 2 > 2ํ + 2โ ํ 2ํ (this is possible becauseํ ํ is strictly larger). She then wins on the next turn becauseํ 2ํ+1 = ํ 2ํ + ํข 2 > 2ํ + 2. Ifํ ํ โค โ 2ํ + 2โ ํ 2ํ , she simply chooses ํฅ 2ํ+1 = 0. We must verify that this strategy is legal and that it indeed leads to a win. First, note that as long as Alice always picks0, all her moves contribute nothing toํandํ 2ํ is the sum of Bazzaโs even moves. For any Bazza strategy we have ํ 2ํ = ํ โ๏ธ ํ=1 ํฅ 2 2ํ โค 2ํ, ํ 2ํ = ํ โ๏ธ ํ=1 ํฅ 2ํ . By Cauchy-Schwarz, ํ 2 2ํ โค ํํ 2ํ โค 2ํ 2 =โ ํ 2ํ โค โ 2ํ.(7) Consequently, ํ ํ = ํ(2ํ + 1)โ ํ 2ํ โฅ ํ(2ํ + 1)โ โ 2ํ = (2ํโ โ 2)ํ + ํ.(8) Because2ํโ โ 2 > 0, the right-hand side of (8) tends to+โasํ โโ. In particular, there exists an integer ํพ such that ํ ํพ > โ๏ธ 2ํพ + 2โ ํ 2ํพ . (Since the right side is at most โ 2ํพ + 2, it suffices that(2ํโ โ 2)ํพ + ํ > โ 2ํพ + 2; this holds for all sufficiently large ํพ.) At the start of turn 2ํพ + 1, we have ํ ํพ > โ 2ํพ + 2โ ํ 2ํพ . Choose ํข = โ๏ธ 2ํพ + 2โ ํ 2ํพ + ํ, with ํ > 0 small enough that ํขโค ํ ํพ (possible because ํ ํพ is strictly larger). Then ํ 2ํพ+1 = ํ 2ํพ + ํขโค ํ 2ํพ + ํ ํพ = ํ(2ํพ + 1), so the sum condition is satisfied, and ํ 2ํพ+1 = ํ 2ํพ + ํข 2 > ํ 2ํพ + (2ํพ + 2โ ํ 2ํพ ) = 2ํพ + 2. Therefore on turn2ํพ + 2Bazza has no legal move (evenํฅ = 0would giveํ > 2ํพ + 2). Alice wins. It remains to check that Alice can always make the choiceํฅ 2ํ+1 = 0whenํ ํ โค โ 2ํ + 2โ ํ 2ํ . The condition for choosing0is that0 โค ํ ํ , i.e.ํ ํ โฅ 0. From (8), ํ ํ โฅ (2ํ โ โ 2)ํ + ํ > 0for allํ(since the right side is positive forํ = 1because 3ํโ โ 2 > 0forํ > โ 2/2). Hence Alice can always pick0when she does not want to win yet. Thus her strategy is well defined and guarantees a win in finite time. Conclusion for ํ > โ 2/2: Alice has a winning strategy. Final answer โข Alice has a winning strategy iff ํ > โ 2 2 . โข Bazza has a winning strategy iff ํ < โ 2 2 . 55 of 63 IMO2025IMO 2025 Model Proofs โข For ํ = โ 2 2 neither player can force a win; the game can be drawn. Human Expert Comment Some parts of the proof read like a thinking process. It feels like it realizes that c has not been proved, then decides that proving c is not actually necessary. The proof seems finished, but then it starts thinking again. 7/7 56 of 63 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation References [1]Gretel synthetic safety alignment dataset, 12 2024. URLhttps://huggingface.co/datasets/ gretelai/gretel-safety-alignment-en-v1. 8 [2] Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from self-generated mistakes. In The twelfth international conference on learning representations, 2024. 12 [3]Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K Arora, Yu Bai, Bowen Baker, Haiming Bao, et al. gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925, 2025. 7, 8, 9, 20 [4] Wasi Uddin Ahmad, Sean Narenthiran, Somshubra Majumdar, Aleksander Ficek, Siddhartha Jain, Jo- celyn Huang, Vahid Noroozi, and Boris Ginsburg. Opencodereasoning: Advancing data distillation for competitive coding. arXiv preprint arXiv:2504.01943, 2025. 7 [5]Ibragim Badertdinov, Alexander Golubev, Maksim Nekrashevich, Anton Shevtsov, Simon Karasik, Andrei Andriushchenko, Maria Trofimova, Daria Litvintseva, and Boris Yangel. Swe-rebench: An automated pipeline for task collection and decontaminated evaluation of software engineering agents. arXiv preprint arXiv:2505.20411, 2025. 9 [6]Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, et al. Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3639โ3664, 2025. 22 [7]Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik Narasimhan.ํ 2 -bench: Evaluating conversational agents in a dual-control environment, 2025. URLhttps://arxiv.org/abs/2506. 07982. 23 [8]Aaron Blakeman, Aaron Grattafiori, Aarti Basant, Abhibha Gupta, Abhinav Khattar, Adi Renduchintala, Aditya Vavre, Akanksha Shukla, Akhiad Bercovich, Aleksander Ficek, et al. Nemotron 3 nano: Open, efficient mixture-of-experts hybrid mamba-transformer model for agentic reasoning. arXiv preprint arXiv:2512.20848, 2025. 7, 8, 9, 11, 12, 14, 23, 24 [9]Aaron Blakeman, Aaron Grattafiori, Aarti Basant, Abhibha Gupta, Abhinav Khattar, Adi Renduchintala, Aditya Vavre, Akanksha Shukla, Akhiad Bercovich, Aleksander Ficek, et al. Nvidia nemotron 3: Efficient and open intelligence. arXiv preprint arXiv:2512.20856, 2025. 5, 9 [10]Yang Chen, Zhuolin Yang, Zihan Liu, Chankyu Lee, Peng Xu, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. Acereason-nemotron: Advancing math and code reasoning through reinforcement learning. Advances in neural information processing systems, 2025. 13 [11]Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. Chatbot arena: An open platform for evaluating llms by human preference, 2024. 14 [12] Kaustubh Deshpande, Ved Sirdeshmukh, Johannes Baptist Mols, Lifeng Jin, Ed-Yeremai Hernandez- Cardona, Dean Lee, Jeremy Kritz, Willow E Primack, Summer Yue, and Chen Xing. Multichallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier llms. In Findings of the Association for Computational Linguistics: ACL 2025, pages 18632โ18702, 2025. 22 57 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation [13]Daniel Deutsch, Eleftheria Briakou, Isaac Rayburn Caswell, Mara Finkelstein, Rebecca Galor, Juraj Juraska, Geza Kovacs, Alison Lui, Ricardo Rei, Jason Riesa, et al. Wmt24++: Expanding the language coverage of wmt24 to 55 languages & dialects. In Findings of the Association for Computational Linguistics: ACL 2025, pages 12257โ12284, 2025. 24 [14]Shihan Dou, Ming Zhang, Zhangyue Yin, Chenhao Huang, Yujiong Shen, Junzhe Wang, Jiayi Chen, Yuchen Ni, Junjie Ye, Cheng Zhang, et al. Cl-bench: A benchmark for context learning. arXiv preprint arXiv:2602.03587, 2026. 23 [15] Wei Du, Shubham Toshniwal, Branislav Kisacanin, Sadegh Mahdavi, Ivan Moshkov, George Armstrong, Stephen Ge, Edgar Minasyan, Feng Chen, and Igor Gitman. Nemotron-math: Efficient long-context distillation of mathematical reasoning from multi-mode supervision. arXiv preprint arXiv:2512.15489, 2025. 7 [16]Tony Feng, Trieu H. Trinh, Garrett Bingham, Dawsen Hwang, Yuri Chervonyi, Junehyuk Jung, Joonkyung Lee, Carlo Pagano, Sang-hyun Kim, Federico Pasqualotto, Sergei Gukov, Jonathan N. Lee, Junsu Kim, Kaiying Hou, Golnaz Ghiasi, Yi Tay, YaGuang Li, Chenkai Kuang, Yuan Liu, Hanzhao Lin, Evan Zheran Liu, Nigamaa Nayakanti, Xiaomeng Yang, Heng-Tze Cheng, Demis Hassabis, Koray Kavukcuoglu, Quoc V. Le, and Thang Luong. Towards autonomous mathematics research. arXiv preprint arXiv:2602.10177, 2026. doi: 10.48550/arXiv.2602.10177. 17 [17] Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, and Pasquale Minervini. Are we done with mmlu?, 2024. 21 [18]Gemini Team.A new era of intelligence with gemini 3.https://blog.google/ products-and-platforms/products/gemini/gemini-3/ , 2025. Google Blog, November 18, 2025. 17 [19]Gemini Team. Advanced version of gemini with deep think officially achieves gold-medal standard at the international mathematical olympiad. https://deepmind.google/blog/advanced-version-of-gemini-with- deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/, 2025. Google DeepMind Blog, July 21, 2025. 6, 17 [20]Gemini Team.Gemini 3 deep think: Advancing science, research and engineering. https://blog.google/innovation-and-ai/models-and-research/gemini-models/ gemini-3-deep-think/, 2026. Google Blog, February 12, 2026. 17 [21]Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, and Christopher Parisien. Aegis2. 0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 5992โ6026, 2025. 8 [22]Nuno M Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and Andrรฉ FT Martins. xcomet: Transparent machine translation evaluation through fine-grained error detection. Transactions of the Association for Computational Linguistics, 12:979โ995, 2024. 24 [23]Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. 4 58 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation [24]Adib Hasan, Ileana Rugina, and Alex Wang. Pruning for protection: Increasing jailbreak resistance in aligned llms without fine-tuning. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 417โ430, 2024. 8 [25]Zhongmou He, Yee Man Choi, Kexun Zhang, Jiabao Ji, Junting Zhou, Dejia Xu, Ivan Bercovich, Aidan Zhang, and Lei Li. Hardtests: Synthesizing high-quality test cases for llm coding. arXiv preprint arXiv:2505.24098, 2025. 7 [26]Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Stein- hardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020. 21 [27] HMMT. Harvard-mit mathematics tournament february 2025, 2025. 20 [28] Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. Ruler: Whatโs the real context size of your long-context language models? arXiv preprint arXiv:2404.06654, 2024. 22 [29] Siming Huang, Tianhao Cheng, Jason Klein Liu, Jiaran Hao, Liuyihan Song, Yang Xu, J Yang, Jiaheng Liu, Chenchen Zhang, Linzheng Chai, et al. Opencoder: The open cookbook for top-tier code large language models. arXiv preprint arXiv:2411.04905, 2024. 7 [30] IMO. International Mathematical Olympiad, 2025. 20 [31]Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar- Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. arXiv preprint arXiv:2403.07974, 2024. 18, 21 [32]Naman Jain, Jaskirat Singh, Manish Shetty, Liang Zheng, Koushik Sen, and Ion Stoica. R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents, 2025. URLhttps: //arxiv.org/abs/2504.07164. 9, 16 [33]Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770, 2023. 23 [34] Gregory Kamradt. Needle in a haystack - pressure testing llms. Github, 2023. URLhttps://github. com/gkamradt/LLMTest_NeedleInAHaystack/tree/main. 22 [35] Diederik P Kingma. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 11, 12, 14, 15 [36]Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. arXiv preprint arXiv:2406.11939, 2024. 22 [37]Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-V3 technical report. arXiv preprint arXiv:2412.19437, 2024. 8 [38]Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, et al. Deepseek-v3. 2: Pushing the frontier of open large language models. arXiv preprint arXiv:2512.02556, 2025. 4, 6, 7, 9 [39]Zihan Liu, Yang Chen, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. AceMath: Advancing frontier math reasoning with post-training and reward modeling. ACL, 2024. 20 59 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation [40]Zihan Liu, Zhuolin Yang, Yang Chen, Chankyu Lee, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. Acereason-nemotron 1.1: Advancing math and code reasoning through sft and rl synergy. ICLR, 2026. 20 [41]LM-Provers, Yuxiao Qu, Amrith Setlur, Jasper Dekoninck, Edward Beeching, Jia Li, Ian Wu, Lewis Tunstall, and Aviral Kumar. Qed-nano: Teaching a tiny model to prove hard theorems. https://huggingface.co/spaces/lm-provers/qed-nano-blogpost, 2026. Blog post. 17 [42]Kevin Lu and Thinking Machines Lab. On-policy distillation. Thinking Machines Lab: Connectionism, 2025. doi: 10.64434/tml.20251026. https://thinkingmachines.ai/blog/on-policy-distillation. 12 [43]Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks. arXiv preprint arXiv:2404.03027, 2024. 8 [44]Minh-Thang Luong, Dawsen Hwang, Hoang H Nguyen, Golnaz Ghiasi, Yuri Chervonyi, Insuk Seo, Junsu Kim, Garrett Bingham, Jonathan Lee, Swaroop Mishra, et al. Towards robust mathematical reasoning. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 35406โ 35430, 2025. 17, 20 [45] Wenjie Ma, Andrei Cojocaru, Neel Kolhe, Bradley Louie, Robin Said Sharif, Haihan Zhang, Vincent Zhuang, Matei Zaharia, and Sewon Min. Reliable fine-grained evaluation of natural language math proofs. arXiv preprint arXiv:2510.13888, 2025. 6, 38 [46] MAA. American Invitational Mathematics Examination - AIME 2025, 2025. 20 [47] MAA. American Invitational Mathematics Examination - AIME 2026, 2026. 20 [48] Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, E. Kelly Buchanan, Junhong Shen, Guanghao Ye, Haowei Lin, Jason Poulos, Maoyu Wang, Marianna Nezhurina, Jenia Jitsev, Di Lu, Orfeas Menis Mastromichalakis, Zhiwei Xu, Zizhao Chen, Yue Liu, Robert Zhang, Leon Liangyu Chen, Anurag Kashyap, Jan-Lucas Uslu, Jeffrey Li, Jianbo Wu, Minghao Yan, Song Bian, Vedang Sharma, Ke Sun, Steven Dillmann, Akshay Anand, Andrew Lanpouthakoun, Bardia Koopah, Changran Hu, Etash Guha, Gabriel H. S. Dreiman, Jiacheng Zhu, Karl Krauth, Li Zhong, Niklas Muennighoff, Robert Amanfu, Shangyin Tan, Shreyas Pimpalgaonkar, Tushar Aggarwal, Xiangning Lin, Xin Lan, Xuandong Zhao, Yiqing Liang, Yuanli Wang, Zilong Wang, Changzhi Zhou, David Heineman, Hange Liu, Harsh Trivedi, John Yang, Junhong Lin, Manish Shetty, Michael Yang, Nabil Omi, Negin Raoof, Shanda Li, Terry Yue Zhuo, Wuwei Lin, Yiwei Dai, Yuxin Wang, Wenhao Chai, Shang Zhou, Dariush Wahdany, Ziyu She, Jiaming Hu, Zhikang Dong, Yuxuan Zhu, Sasha Cui, Ahson Saiyed, Arinbjรถrn Kolbeinsson, Jesse Hu, Christopher Michael Rytting, Ryan Marten, Yixin Wang, Alex Dimakis, Andy Konwinski, and Ludwig Schmidt. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces, 2026. URL https://arxiv.org/abs/2601.11868. 9, 23 [49]NVIDIA. Nemo gym: An open source library for scaling reinforcement learning environments for llm. https://github.com/NVIDIA-NeMo/Gym, 2025. GitHub repository. 15 [50]NVIDIA. NeMo RL: A Scalable and Efficient Post-Training Library.https://github.com/ NVIDIA-NeMo/RL, 2025. GitHub repository. 10 [51] NVIDIA. NeMo-Skills.https://github.com/NVIDIA-NeMo/Skills, 2025. GitHub repository. 20 [52]NVIIDA. Nemotron-3-Nano-30B-A3B-Base-BF16, 2025. URLhttps://huggingface.co/nvidia/ NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16. 6 [53] OpenAI. Introducing SWE-bench Verified, 2024. 23 60 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation [54] OpenAI. Introducing GPT-5, 2025. 17 [55]Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. NeurIPS, 35, 2022. 4 [56]Jiayi Pan*, Xingyao Wang*, Graham Neubig, Navdeep Jaitly, Heng Ji, Alane Suhr, and Yizhe Zhang. Training software engineering agents and verifiers with swe-gym. In ICML, 2025. URLhttps://arxiv. org/abs/2412.21139. 9, 16 [57]Shishir G. Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E. Gonzalez. The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models. In ICML, 2025. 23 [58]Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, and et al. Humanityโs last exam, 2025. URL https://arxiv.org/abs/2501.14249. 21 [59]Renjie Pi, Grace Lam, Mohammad Shoeybi, Pooya Jannaty, Bryan Catanzaro, and Wei Ping. On data engineering for scaling llm terminal capabilities, 2026. URLhttps://arxiv.org/abs/2602.21193. 9 [60] Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert, and Hannaneh Hajishirzi. Generalizing verifiable instruction following, 2025. URLhttps: //huggingface.co/datasets/allenai/IF_multi_constraints_upto5. 11, 22 [61] Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. URLhttps://qwen.ai/ blog?id=qwen3.5. 5 [62]David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling, 2024. 21 [63]Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. DeepseekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. 10 [64]Zhihong Shao, Yuxiang Luo, Chengda Lu, Z Ren, Jiewen Hu, Tian Ye, Zhibin Gou, Shirong Ma, and Xiaokang Zhang. Deepseekmath-v2: Towards self-verifiable mathematical reasoning. arXiv preprint arXiv:2511.22570, 2025. 16, 17 [65]Atefeh Sohrabizadeh, Jialin Song, Mingjie Liu, Rajarshi Roy, Chankyu Lee, Jonathan Raiman, and Bryan Catanzaro. Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity. In Forty-second International Conference on Machine Learning, 2025. 15 [66] Artificial Analysis Team. Artificial analysis long context reasoning benchmark(lcr), 2025. 22 [67] Minyang Tian, Luyu Gao, Shizhuo D Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, et al. Scicode: A research coding benchmark curated by scientists. Advances in Neural Information Processing Systems, 37:30624โ30650, 2024. 21 [68] Boxin Wang, Chankyu Lee, Nayeon Lee, Sheng-Chieh Lin, Wenliang Dai, Yang Chen, Yangyi Chen, Zhuolin Yang, Zihan Liu, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. Nemotron-Cascade: Scaling cascaded reinforcement learning for general-purpose reasoning models. arXiv preprint arXiv:2512.13607, 2025. 4, 6, 7, 8, 9, 10, 11, 14, 15, 18, 27 61 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation [69]Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brennan, Hao Peng, Heng Ji, and Graham Neubig. Openhands: An open platform for AI software developers as generalist agents. In The Thirteenth International Conference on Learning Representations, 2025. URLhttps: //openreview.net/forum?id=OJd3ayDDoF. 9, 16, 23 [70]Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. arXiv preprint arXiv:2406.01574, 2024. 21 [71]Zhilin Wang, Jiaqi Zeng, Olivier Delalleau, Hoo-Chang Shin, Felipe Soares, Alexander Bukharin, Ellie Evans, Yi Dong, and Oleksii Kuchaiev. Helpsteer3-preference: Open human-annotated preference data across diverse tasks and languages, 2025. URL https://arxiv.org/abs/2505.11475. 14 [72]Zhilin Wang, Jiaqi Zeng, Olivier Delalleau, Hoo-Chang Shin, Felipe Soares, Alexander Bukharin, Ellie Evans, Yi Dong, and Oleksii Kuchaiev. Helpsteer3-preference: Open human-annotated preference data across diverse tasks and languages, 2025. URL https://arxiv.org/abs/2505.11475. 14 [73] Yuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux, Lingming Zhang, Daniel Fried, Gabriel Synnaeve, Rishabh Singh, and Sida I Wang. Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution. arXiv preprint arXiv:2502.18449, 2025. 9 [74]Ronald J Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8:229โ256, 1992. 11 [75]Bangjun Xiao, Bingquan Xia, Bo Yang, Bofei Gao, Bowen Shen, Chen Zhang, Chenhong He, Chiheng Lou, Fuli Luo, Gang Wang, et al. Mimo-v2-flash technical report. arXiv preprint arXiv:2601.02780, 2026. 4, 12 [76]Peng Xu, Wei Ping, Xianchao Wu, Chejian Xu, Zihan Liu, Mohammad Shoeybi, and Bryan Catanzaro. Chatqa 2: Bridging the gap to proprietary llms in long context and rag capabilities. arXiv preprint arXiv:2407.14482, 2024. 8 [77]Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, et al. Mmlu-prox: A multilingual benchmark for advanced large language model evaluation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 1513โ1532, 2025. 24 [78] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. 8, 9, 12 [79]John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 50528โ50652. Curran Associates, Inc., 2024. doi: 10. 52202/079017-1601. URLhttps://proceedings.neurips.c/paper_files/paper/2024/ file/5a7c947568c1b1328c5230172e1e7c-Paper-Conference.pdf. 9 [80] Zonghan Yang, Shengjie Wang, Kelin Fu, Wenyang He, Weimin Xiong, Yibo Liu, Yibo Miao, Bofei Gao, Yejie Wang, YINGWEI MA, Yanhao Li, Yue Liu, Zhenxing Hu, kaitai zhang, Shuyi Wang, Huarong Chen, Flood Sung, Yang Liu, Yang Gao, Zhilin Yang, and Tianyu Liu. Kimi-dev: Agentless training as skill prior for SWE-agents. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=tYppHuGhxJ. 16 62 Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation [81]Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al. DAPO: An open-source LLM reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025. 11 [82]Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chengx- ing Xie, Cunxiang Wang, et al. Glm-5: from vibe coding to agentic engineering. arXiv preprint arXiv:2602.15763, 2026. 4, 12 [83]Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. Wildchat: 1m chatgpt interaction logs in the wild. arXiv preprint arXiv:2405.01470, 2024. 8 [84] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric. P Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang. Lmsys-chat-1m: A large-scale real-world llm conversation dataset, 2023. 8 [85]Zihan Zheng, Zerui Cheng, Zeyu Shen, Shang Zhou, Kaiyuan Liu, Hansen He, Dongruixuan Li, Stanley Wei, Hangyi Hao, Jianzhu Yao, et al. Livecodebench pro: How do olympiad medalists judge llms in competitive programming? arXiv preprint arXiv:2506.11928, 2025. 18, 21, 27 [86] Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models, 2023. URLhttps://arxiv.org/ abs/2311.07911. 22 63