Paper deep dive
MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis
Zixiong Yu, Jun Rao, Guhan Chen, Songtao Tian, Bohan Li, Jiansheng Wei, Min Zhang, Xiaojun Meng
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/14/2026, 2:31:41 AM
Summary
MathAgent is a hierarchical data synthesis framework that uses a Legislator-Executor paradigm to generate high-quality mathematical reasoning data. The Legislator adversarially evolves constraint graphs (blueprints) to ensure structural complexity and diversity, while the Executor instantiates these blueprints into natural language problems. Experiments show that models fine-tuned on 1K MathAgent samples outperform existing datasets like LIMO and s1K across eight mathematical benchmarks.
Entities (6)
Relation Signals (4)
Executor → instantiates → Constraint Graph
confidence 100% · the Executor instantiates these specifications into diverse natural language scenarios
Legislator → optimizes → Constraint Graph
confidence 100% · the Legislator (meta-level) adversarially optimizes the combination of problem elements over a constraint graph
MathAgent → outperforms → LIMO
confidence 100% · models fine-tuned on 1K synthesized samples outperform widely-used datasets of comparable scale (LIMO, s1K)
MathAgent → utilizes → Legislator-Executor paradigm
confidence 100% · We introduce a Legislator-Executor paradigm: The Legislator adversarially evolves structured generation blueprints
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Synthesizing high-quality mathematical reasoning data without human priors remains a significant challenge. Current approaches typically rely on seed data mutation or simple prompt engineering, often suffering from mode collapse and limited logical complexity. This paper proposes a hierarchical synthesis framework that formulates data synthesis as an unsupervised optimization problem over a constraint graph followed by semantic instantiation, rather than treating it as a direct text generation task. We introduce a Legislator-Executor paradigm: The Legislator adversarially evolves structured generation blueprints encoding the constraints of the problem, while the Executor instantiates these specifications into diverse natural language scenarios. This decoupling of skeleton design from linguistic realization enables a prioritized focus on constructing complex and diverse logical structures, thereby guiding high-quality data synthesis. Experiments conducted on a total of 10 models across the Qwen, Llama, Mistral, and Gemma series demonstrate that our method achieves notable results: models fine-tuned on 1K synthesized samples outperform widely-used datasets of comparable scale (LIMO, s1K) across eight mathematical benchmarks, exhibiting superior out-of-distribution generalization.
Tags
Links
- Source: https://arxiv.org/abs/2604.11188v1
- Canonical: https://arxiv.org/abs/2604.11188v1
Trouble viewing inline? Open PDF directly →
Full Text
104,558 characters extracted from source content.
Expand or collapse full text
MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis Zixiong Yu1,2, Jun Rao3, Guhan Chen2, Songtao Tian2, Bohan Li2,4, Jiansheng Wei1, Min Zhang3, and Xiaojun Meng1 1Huawei Large Model Data Technology Lab 2Tsinghua University 3Harbin Institute of Technology, Shenzhen 4Kyoto University yuzx19,tiansongtao.2020,libh19@tsinghua.org.cn rao7jun@gmail.com chen-gh23@mails.tsinghua.edu.cn zhangmin2021@hit.edu.cn weijiansheng,xiaojun.meng@huawei.com Co-first Author.Corresponding Author. Abstract Synthesizing high-quality mathematical reasoning data without human priors remains a significant challenge. Current approaches typically rely on seed data mutation or simple prompt engineering, often suffering from mode collapse and limited logical complexity. This paper proposes a hierarchical synthesis framework that formulates data synthesis as an unsupervised optimization problem over a constraint graph followed by semantic instantiation, rather than treating it as a direct text generation task. We introduce a Legislator-Executor paradigm: The Legislator adversarially evolves structured generation blueprints encoding the constraints of the problem, while the Executor instantiates these specifications into diverse natural language scenarios. This decoupling of skeleton design from linguistic realization enables a prioritized focus on constructing complex and diverse logical structures, thereby guiding high-quality data synthesis. Experiments conducted on a total of 10 models across the Qwen, Llama, Mistral, and Gemma series demonstrate that our method achieves notable results: models fine-tuned on 1K synthesized samples outperform widely-used datasets of comparable scale (LIMO, s1K) across eight mathematical benchmarks, exhibiting superior out-of-distribution generalization. MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis Zixiong Yu1,2, Jun Rao3†thanks: Co-first Author., Guhan Chen2, Songtao Tian2, Bohan Li2,4, Jiansheng Wei1, Min Zhang3, and Xiaojun Meng1†thanks: Corresponding Author. 1Huawei Large Model Data Technology Lab 2Tsinghua University 3Harbin Institute of Technology, Shenzhen 4Kyoto University yuzx19,tiansongtao.2020,libh19@tsinghua.org.cn rao7jun@gmail.com chen-gh23@mails.tsinghua.edu.cn zhangmin2021@hit.edu.cn weijiansheng,xiaojun.meng@huawei.com 1 Introduction In recent years, Large Language Models (LLMs; Vaswani et al., 2017; Brown et al., 2020; Zhao et al., 2023) have become a central pillar of modern artificial intelligence (AI). Although theoretical understanding of their underlying mechanisms remains relatively limited (Jacot et al., 2018; Li et al., 2024b; Yu et al., 2025), LLMs have demonstrated strong reasoning abilities in practice and achieved remarkable success on complex tasks (Wei et al., 2022a, 2026). These capabilities have in turn enabled rapid expansion into a wide range of domains, such as embodied AI (Driess et al., 2023; Zeng et al., 2026) and LLM-driven scientific discovery (Boiko et al., 2023; Lu et al., 2024a). This progress has been driven by multiple factors, including the scaling of model parameters and training data (Kaplan et al., 2020; Hoffmann et al., 2022), reasoning-oriented techniques such as chain-of-thought (CoT) prompting (Wei et al., 2022b; Zeng et al., 2025a), and, equally importantly, the quality of training data (Zhou et al., 2023; Ye et al., 2025; Zhao et al., 2026). However, as high-quality human-generated corpora become increasingly difficult to scale, the field faces a growing data bottleneck (Villalobos et al., 2024). Consequently, synthetic data generation, which uses generative models to produce training samples, has emerged as a major research direction (Honovich et al., 2023; Rao et al., 2025a; Ke et al., 2025). Current synthesis paradigms primarily fall into two categories: (i) Seed-based methods, such as Self-Instruct (Wang et al., 2023), expand upon human-curated seeds. While effective, their diversity is inherently upper-bounded by the semantic span of the initial seeds. (i) Zero-shot methods like Magpie (Xu et al., 2025) probe model distributions directly but often suffer from mode collapse and logical hallucinations due to the lack of structural guidance (Shumailov et al., 2023). We argue that by framing data synthesis as a mere text generation task rather than a structured optimization problem, current methods often confine models to superficial narrative imitation without mastering core reasoning capabilities (Gudibande et al., 2023). To address this, we propose a hierarchical synthesis framework anchored by a bi-level Legislator-Executor paradigm, which effectively decouples structural specifications from their textual instantiation. By pre-establishing high-level task blueprints that incorporate logical relations and constraints, the framework more effectively guides the generation of high-quality mathematical problems. We instantiate this architecture as MathAgent, where the Legislator (meta-level) adversarially optimizes the combination of problem elements over a constraint graph, while the Executor (base-level) transforms these abstract blueprints into natural language. This decoupling mechanism enables a prioritized focus on orchestrating structural diversity and complexity. Through iterative adversarial evolution, MathAgent continuously explores the underlying structural space, thereby progressively pushing the frontiers of model generation capabilities. Compared to direct probing methods confined by high-frequency patterns and seed dataset augmentation methods limited by initial semantic ranges, our framework excels at capturing scarce data characterized by high difficulty and quality. Consequently, relying solely on basic conceptual primitives rather than seed data, MathAgent synthesizes corpora with high structural complexity and rich diversity, while flexibly regulating the complexity of data distributions through an adaptive early-stop iteration mechanism. Our contributions are summarized as follows: • We propose the Legislator-Executor paradigm, a hierarchical synthesis framework that decouples task specification from textual realization to facilitate the guided synthesis of reasoning data. • We introduce a constraint-graph-based adversarial evolutionary mechanism to explore structural spaces, generating high-difficulty, high-quality problems often absent in standard datasets. • Extensive experiments demonstrate that models fine-tuned on 1K MathAgent samples outperform mainstream datasets of comparable scale (LIMO & s1K) across eight benchmarks, exhibiting superior out-of-distribution generalization. 2 Related Work Data Synthesis A prevalent paradigm in data synthesis involves the iterative expansion of seed examples. Methods like Self-Instruct (Wang et al., 2023) and WizardMath (Luo et al., 2025) utilize evolution strategies to amplify task complexity, while MathGenie (Lu et al., 2024c) employs a backward mechanism, augmenting seed solutions to back-translate new questions. Nevertheless, these approaches are constrained by the "semantic radius" of their initial seeds, often failing to explore the unknown regions of the problem space. Similarly, for preference-pair data, methods such as SeaPO (Rao et al., 2025a) construct preference pairs by generating contrastive responses based on existing answers; however, such approaches typically place higher demands on the fine-grained and controllable editing capabilities of LLMs (Zeng et al., 2025c). To avoid relying on seed datasets, zero-shot methods such as Magpie (Xu et al., 2025) and schema-driven frameworks like Condor (Maosongcao et al., 2025) attempt to synthesize data from scratch. While successful in high-resource domains, they lack the structural incentives to discover samples in the long-tail distribution, where complex reasoning capabilities are often forged. Multi-Agent and Adversarial Generation The deployment of LLMs has progressed from single-turn prompting to complex Multi-Agent Systems. Frameworks such as CAMEL (Li et al., 2023) and MetaGPT (Hong et al., 2024) demonstrate that role-playing agents can effectively decompose tasks through cooperation, while AgentDropout (Wang et al., 2025) further improves efficiency and coordination via dynamic agent elimination. In the realm of data synthesis, MATRIX (Tang et al., 2025) utilizes multi-agent simulation to construct virtual societies, generating instruction data grounded in realistic social scenarios. Concurrently, multi-agent debate (Du et al., 2024) has emerged as a pivotal mechanism for enhancing reasoning reliability. For instance, Debate4MATH (Zhang and Xiong, 2025) employs fine-grained step verification to rectify logical errors, while Liang et al. (2024) leverage debate to stimulate divergent thinking for higher-quality problem solving. Our work adapts adversarial dynamics for data synthesis, aiming to drive continuous evolution of the training data distribution and explore the generation of complex samples. 3 Method We propose a hierarchical synthesis framework to synthesize mathematical data by optimizing the constraint graph of problem structures. Specifically instantiated as MathAgent, our approach diverges from standard methods that operate directly in the token space by decoupling the synthesis process into two distinct phases: (1) Structural Evolution (Meta-level), governed by a Legislator agent that optimizes a constraint graph (acting as the synthesis blueprint); and (2) Semantic Instantiation (Base-level), conducted by an Executor that grounds the graph structure into natural language scenarios. 3.1 Problem Formulation We formulate the skeleton of a mathematical problem as a structure comprising a Constraint Graph =(,ℰ)G=(V,E) and Style Tokens S. Specifically: • Nodes (V) represent mathematical concepts. • Edges (ℰE) represent logical relations. • Style Tokens (S) control global attributes (e.g., problem category or difficulty level). Serving as a blueprint for problem synthesis, this graph facilitates automated structural evolution, enabling the targeted generation of high-complexity and richly diverse reasoning data. Formally, our objective is to explore the space of graph topologies, optimizing the complexity while ensuring strict solvability (conditioned on S): ∗=arg max∈ℋ()s.t.valid(∣)=1, ^*=arg\,max_G \;H(G) .t. _valid(G )=1, where G is the search space, ℋ(⋅)H(·) estimates the complexity, and valid(⋅)I_valid(·) is a binary validity indicator. Note that while the optimization seeks to push the reasoning frontier (finding ∗G^*), the evolutionary trajectory yields a diverse curriculum of graphs. The resulting tuple (∗,)(G^*,S) is then passed to the Executor for textual realization. Figure 1: The MathAgent Framework. The framework consists of two decoupled phases: (1) Meta-Level Structural Evolution, where a tri-agent Legislator system (Proposer, Critic, and Moderator) iteratively optimizes a Constraint Graph G based on Style Tokens S; and (2) Base-Level Semantic Instantiation, where the Executor grounds the optimized structural blueprint into natural language problems Q and reasoning chains A. 3.2 Phase 1: The Legislator (Meta-Level) To address the optimization objective, we design the Legislator as a tri-agent evolutionary system. Instead of directly manipulating text, the system iteratively optimizes the dynamic constraint graph tG_t through inter-agent collaboration under given style token S conditions. The evolutionary process is jointly driven by three distinct roles: Proposer (PA_P) As the driving engine of structural evolution, the proposer PA_P optimizes tG_t to t+1G_t+1 guided by feedback from previous iterations. It resolves logical contradictions while ensuring the graph’s characteristics align with the structural specifications defined in style tokens S. In particular, if the current structure has not reached the target complexity required by style tokens S, the proposer PA_P proactively expands knowledge nodes or strengthens constraints to enhance structural depth. Critic (CA_C) As a key component of adversarial evolution, the critic CA_C scrutinizes t+1G_t+1 across three dimensions based on the style tokens S: (1) Internal Consistency: it verifies whether logical contradictions exist in the current graph; (2) Specification Alignment: it checks if the graph complies with the constraints specified in the style tokens S; (3) Optimization Potential: it proactively probes for superior configurations that transcend the current design. The results are synthesized into a comprehensive refinement report. Moderator (MA_M) The moderator MA_M serves as the strategic decision-maker, adjudicating the evolution of t+1G_t+1 by weighing the refinement report against the global objective. Each cycle yields one of two outcomes: • Adaptive Truncation: If t+1G_t+1 satisfies S and the potential for further gain is marginal, MA_M terminates the process and outputs ∗G^*. • Iterative Guidance: Otherwise, MA_M directs PA_P to implement the critic’s suggestions to resolve inconsistencies or enhance structural depth. Initialization To ensure high initial diversity and eliminate human intervention, we deploy a similar adversarial mechanism before the evolutionary loop to construct an initial pool: the proposer PA_P activates latent information to propose candidate attributes, while the critic CA_C filters these attributes based on requirements such as orthogonality, validity, and diversity. This adversarial process builds a self-organized initial pool, containing: • Style Tokens (S): A rich and diverse set of stylistic constraint dimensions, each offering as comprehensive a range of options as possible. • Concept Taxonomy (C): A comprehensive atlas of mathematical domains. At the onset (t=0t=0), the system randomly samples from this initial pool to generate the initial graph 0G_0, ensuring that the entire dataset is driven solely by the model’s intrinsic representational diversity. In practice, the concept taxonomy can be derived from a variety of sources, such as existing knowledge bases, interactions between humans and LLMs, or weaknesses identified from evaluations of the target LLM (Rao et al., 2025b). 3.3 Phase 2: The Executor (Base-Level) The Executor is a conditional generative model that performs semantic instantiation. It receives the linearized textual representation of the constraint graph ∗G^* alongside the set of Style Tokens S: (Q,A)∼Pexecutor(⋅∣∗,) (Q,A) P_executor(\,·\, ^*,S) where Q denotes the natural language problem statement and A is the step-by-step reasoning chain. By conditioning the generation on ∗G^*, the executor is freed from the burden of exploring and constructing complexity and diversity, allowing it to focus solely on language itself, thereby enabling the generation of diverse textual scenarios. To further ensure the reliability of the synthesized question-answer pairs, we adopt a general model-based verification scheme (Zheng et al., 2023), which involves employing an external model as a judge to evaluate the logical correctness of the generated questions and answers, as well as the consistency between their descriptions. Only samples that pass this verification are retained. Model Dataset Elementary Middle Competition Avg. GSM8K MATH 500 Minerva Math Gaokao 2023en Olympiad Bench AIME24 (Avg@8) AIME25 (Avg@8) AMC23 (Avg@8) Qwen3 Series Models Qwen3-14B-Base Base 94.7 80.6 37.5 67.5 46.2 15.0 15.8 66.6 53.0 LIMO 91.8 86.2 39.0 76.9 50.8 33.8 27.5 70.0 59.5 S1K 87.5 86.4 40.8 76.1 52.6 37.9 25.0 75.9 60.3 Ours 95.4 91.8 39.0 79.0 56.3 38.8 30.0 80.6 63.9 Qwen3-8B-Base Base 92.0 76.8 32.7 64.9 41.8 17.5 13.8 58.4 49.7 LIMO 88.3 80.4 35.7 68.6 44.9 19.6 24.2 60.3 52.8 S1K 87.6 81.4 37.5 71.7 44.4 19.2 23.3 60.3 53.2 Ours 93.3 87.2 39.7 75.8 50.5 27.1 25.4 74.7 59.2 Qwen3-4B-Base Base 82.8 72.4 20.2 60.8 37.9 10.4 7.5 50.3 42.8 LIMO 82.0 74.0 30.9 68.1 40.0 13.8 19.6 56.9 48.2 S1K 84.8 76.8 33.5 68.6 40.6 16.7 20.0 53.1 49.3 Ours 92.0 81.8 35.3 69.6 44.6 19.2 20.4 65.0 53.5 Qwen2.5 Series Models Qwen2.5-7B Base 87.4 63.0 24.3 56.1 29.0 5.4 3.8 36.9 38.2 LIMO 88.6 71.4 29.4 61.8 34.7 10.0 14.6 45.9 44.6 S1K 89.4 68.8 30.5 62.9 35.9 11.7 9.6 43.8 44.1 Ours 90.1 74.6 30.9 68.3 38.5 14.2 15.0 55.9 48.4 Qwen2.5-Math-7B Base 66.7 64.0 12.1 56.1 28.3 11.7 4.2 38.8 35.2 LIMO 87.4 72.2 31.6 63.1 37.9 10.8 14.6 47.2 45.6 S1K 87.9 73.2 33.5 64.2 35.6 11.7 11.2 49.7 45.9 Ours 91.6 82.2 34.2 70.4 47.0 18.8 18.3 65.3 53.5 Qwen2.5-7B-QwQ Base 90.4 75.4 32.7 66.8 39.6 17.9 19.6 55.0 49.7 LIMO 90.7 78.4 32.7 70.4 44.1 16.2 20.8 57.5 51.4 S1K 90.7 78.2 30.5 64.9 44.7 17.9 20.0 55.0 50.2 Ours 91.7 83.8 33.8 73.5 50.7 27.5 21.7 70.6 56.7 Llama, Mistral and Gemma Models Llama-3.1-8B Base 40.9 13.4 4.8 15.1 3.1 0.0 0.0 4.7 10.3 LIMO 66.0 24.6 6.2 24.9 6.4 0.0 0.4 8.8 17.2 S1K 66.9 24.0 8.5 28.8 5.9 0.0 0.8 7.8 17.8 Ours 67.8 27.6 9.6 30.9 6.8 0.8 0.8 10.9 19.4 Llama-3.2-3B Base 25.8 7.4 2.6 9.1 2.5 0.0 0.0 3.8 6.4 LIMO 22.8 8.6 3.3 11.9 2.7 0.0 0.0 2.5 6.5 S1K 20.4 7.8 2.2 11.2 1.8 0.0 0.0 3.4 5.9 Ours 37.4 10.8 4.0 17.1 3.3 0.0 0.4 3.8 9.6 Mistral-7B-v0.3 Base 18.0 7.4 2.6 9.4 2.2 0.0 0.0 2.5 5.3 LIMO 33.1 12.6 4.0 14.5 3.0 0.0 0.4 2.2 8.7 S1K 33.4 7.8 5.9 11.9 1.8 0.0 0.0 2.8 8.0 Ours 43.5 13.2 6.6 15.6 3.4 0.4 0.4 3.1 10.8 Gemma-2-9B Base 54.8 23.4 8.5 24.7 6.5 0.4 0.0 7.5 15.7 LIMO 76.6 40.6 14.7 38.4 15.7 2.1 0.8 20.6 26.2 S1K 76.4 41.0 17.3 43.1 13.5 1.2 0.4 20.3 26.7 Ours 75.9 43.8 16.2 43.4 16.0 2.9 1.2 23.8 27.9 Table 1: Main Results. Model performance comparison on eight mathematical benchmarks (scores in %). Our synthesized dataset outperforms existing open-source baselines (LIMO and s1K) for SFT across multiple models. 4 Experiment 4.1 Setup Data Synthesis Implementation We implement MathAgent through a multi-model pipeline. First, DeepSeek-V3 (DeepSeek-AI et al., 2025b) is employed to construct candidate pools for style tokens S and concept taxonomy C, establishing a diverse initial state for meta-level evolution, with this model also responsible for generating the final solutions. Next, the intermediate cyclic adversarial evolution led by the Legislator, the semantic instantiation performed by the Executor, and the subsequent verification process are all driven by Qwen2.5-32B-Instruct (Yang et al., 2025b). All generation stages maintain a uniform temperature of 0.3. We synthesize a final corpus of 1K instances, aligning the data scale with the baseline datasets introduced in the following section to ensure a fair comparison of data efficiency. Baseline We benchmark the proposed synthetic data against datasets curated through rigorous human-designed filtering pipelines in supervised fine-tuning (SFT). To facilitate extensive experimentation, we select two well-known small-scale open-source datasets, LIMO (Ye et al., 2025) and s1K (Muennighoff et al., 2025) (containing approximately 0.8K and 1K samples, respectively), as representatives of such high-quality filtered data. Previous studies have indicated that the quality rather than the quantity of instruction-tuning datasets is critical (Zhou et al., 2023), making SFT experiments at this scale sufficiently informative. Details of these datasets are provided in Appendix D. Models Our fine-tuning experiments are primarily conducted on the Qwen series of models, specifically including the Qwen3-14B/8B/4B-Base (Yang et al., 2025a) models, Qwen2.5-7B (Yang et al., 2025b), and Qwen2.5-Math-7B (Yang et al., 2024). We also include Qwen2.5-7B-QwQ, which is fine-tuned from the Math-Base model on 15K QwQ samples (Qwen, 2024). To ensure cross-architecture generalization, we extend our evaluation to Llama-3.1-8B (Grattafiori et al., 2024), Llama-3.2-3B (Meta, 2024), Mistral-7B-v0.3 (Jiang et al., 2023), and Gemma-2-9B (Riviere et al., 2024). Comprehensive training details are provided in Appendix A and B. Evaluation We evaluate the model’s mathematical capability after SFT using eight mathematical test sets across the following three difficulty levels: • Elementary (Elem.) GSM8K (Cobbe et al., 2021) & MATH500 (Hendrycks et al., 2021b); • Middle (Mid.) Minerva Math (Lewkowycz et al., 2022), Gaokao 2023en (Liao et al., 2024) & Olympiad Bench (He et al., 2024); • Competition (Comp.) AIME 2024 & 2025 (AoPS, 2025) and AMC23 (AoPS, 2023). For Elementary and Middle-level sets, we report greedy decoding accuracy in a zero-shot setting. For Competition-level sets, we perform 8 sampling iterations per problem and report the average accuracy to mitigate variance. During answer generation for these datasets, we set the temperature to 0.1 and top_p to 0.95. A maximum generation length of 8192 tokens is applied across all test sets. 4.2 Main Results Superior Performance As presented in Table˜1, our synthesized dataset yields significant mathematical reasoning improvements across a comprehensive range of model architectures (Qwen, Llama, Mistral, Gemma), scales (3B–14B), and initialization stages (Base, Math, and SFT). Notably, our method consistently outperforms competitive open-source baselines (LIMO and s1K) across all eight benchmarks, highlighting the efficacy of the Legislator-Executor paradigm. Crucially, these gains are not a byproduct of data contamination; the synthesis process remains entirely independent of the evaluation benchmarks, as substantiated by the similarity analysis in Appendix C. Cross-Difficulty Performance As shown in Table˜1, the performance gains of our method are most prominent on high-difficulty benchmarks. For a clearer comparison, we selected representative models (The Mistral and Llama series were excluded due to their consistently poor performance on challenging test sets, while the Qwen3 series post-training versions were omitted as they incorporate thinking modes that preclude direct comparability) and compared them with their official instruction-tuned (Instruct) versions, with detailed results presented in Table˜2. While Instruct models maintain a slight edge on elementary benchmarks (potentially by capitalizing on high-frequency linguistic patterns within their large-scale training sets), our method significantly outperforms them on intermediate and competition-level tasks. This underscores the efficacy of adversarial evolution in capturing complex skeletons and logical dependencies, fostering robust reasoning capabilities that transcend simple pattern matching. Model Method Elem. Mid. Comp. Avg. Qwen2.5-7B (Yang et al., 2025b) Instruct 84.4 45.8 24.4 47.4 Ours 83.3 45.9 28.4 48.7 Qwen2.5-Math-7B (Yang et al., 2024) Instruct 89.5 48.5 28.6 51.3 Ours 86.9 50.5 34.1 53.5 Gemma-2-9B (Riviere et al., 2024) Instruct 61.1 21.8 6.0 25.7 Ours 59.9 25.2 9.3 27.9 Table 2: Performance Comparison of Our Method versus the Official instruction-tuned Version. Superior performance of our method on more challenging test sets. 4.3 Ablation Study To verify the necessity of the core components, we conduct an ablation study using Qwen2.5-7B as the base model and Qwen2.5-Math-72B-Instruct (Yang et al., 2024) as the resource-efficient solution annotator. The experimental configurations and results are summarized in Table˜3. Method / Variant Avg. Δ Full MathAgent 45.4 - Impact of Legislator-Executor paradigm w/o Constraint Graph (Direct Gen.) 42.4 -3.0 Impact of Adversarial Evolution w/o Roundtable (One-pass) 43.1 -2.3 Table 3: Ablation Study. Δ indicates the performance drop relative to the full framework. The results underscore the necessity of both structural decoupling and adversarial evolution for high-quality synthesis. The performance degradation observed in both variants highlights the synergy between our structural and evolutionary components. Eliminating the constraint graph results in a substantial drop (-3.0%), as the model reverts to superficial narrative imitation and high-frequency patterns in the absence of an explicit structural blueprint. Similarly, bypassing adversarial evolution (-2.3%) forces a reliance on initial semantic intuition, confirming that continuous structural refinement and difficulty-stretching via adversarial mechanisms are essential for pushing the boundaries of the model’s generative capacity for high-quality problem synthesis. 4.4 Isolating Problem Quality In this section, we disentangle the impact of problem quality from answer generation to verify that the performance gains in Table˜1 are primarily driven by our synthesis approach. SFT with Consistent Response Generation To eliminate the impact of differing response generation methodologies across datasets, we adopt a uniform response generation protocol. Specifically, all responses are generated in a single pass using the Qwen2.5-Math-72B-Instruct model without post-generation filtering. For the baseline problem datasets, in addition to LIMO and s1K datasets already used in Table˜1, we included two additional datasets: NuminaMath (Li et al., 2024a) and Magpie (Xu et al., 2025). Detailed descriptions of these datasets are provided in Appendix D. We conducted SFT on the Qwen2.5-7B model using response-standardized versions of all datasets to ensure a rigorous and controlled comparison. The results in Table˜4 reveal that datasets relying on rigorous heuristic curation (LIMO and s1K) outperform both the unfiltered NuminaMath and the unconstrained, single-prompt synthesis of Magpie. In addition, even after standardizing response generation, our proposed method maintains its superiority across all comparisons. Method Type Size Avg. Qwen2.5-7B (Yang et al., 2025b) - - 38.2 + NuminaMath (Li et al., 2024a) SFT 1K 41.1 + Magpie (Xu et al., 2025) SFT 1K 41.5 + LIMO (Ye et al., 2025) SFT 0.8K 43.2 + S1K (Muennighoff et al., 2025) SFT 1K 43.0 + MathAgent (Ours) SFT 1K 45.4 Qwen2.5-7B-Instruct (Yang et al., 2025b) - 47.4 + NuminaMath (Li et al., 2024a) IDPO 32K 47.9 + Magpie (Xu et al., 2025) IDPO 32K 48.0 + MathAgent (Ours) IDPO 32K 49.1 Table 4: Isolating Problem Quality. By standardizing response generation for all datasets, we observe consistent gains under both SFT and IDPO. This confirms that the superiority of MathAgent stems from the inherent logical quality of the synthesized problems, independent of response-level variance. IDPO We also adopt the Iterative Direct Preference Optimization111https://github.com/RLHFlow/Online-DPO-R1 (IDPO; Zhang et al., 2025; Rao et al., 2026) method for evaluation. This approach requires the model to autonomously explore the solution space during training, thereby effectively distinguishing the intrinsic quality differences among various problem sets. The experiments are conducted using the Qwen2.5-7B-Instruct as the base model, and the LIMO and s1K datasets are excluded from this comparison due to their limited scale. Results in Table˜4 show that despite the limited headroom for improvement in an already instruction-tuned base model, our method still attains the superior performance. 4.5 Cross-Task Generalization Fine-tuning on domain-specific data often incurs a trade-off in general capability degradation. To verify that our method maintains robust cross-task generalization while enhancing mathematical reasoning, we extend our evaluation to the following benchmarks: BBH (Suzgun et al., 2023), HumanEval (Chen et al., 2021), MMLU (Hendrycks et al., 2021a), and TruthfulQA (Lin et al., 2022). Adopting the same SFT training setup as in Table˜4, the results are summarized in Table˜5. Method BBH Human Eval MMLU TruthfulQA MC1 MC2 Qwen2.5-7B 69.7 60.4 74.3 38.9 56.3 + NuminaMath 67.3 58.5 73.8 38.8 55.9 + Magpie 68.2 60.4 73.6 37.7 54.9 + LIMO 67.8 62.2↑\, 73.8 37.9 56.2 + S1K 67.4 63.4↑\, 73.2 37.2 55.1 + MathAgent 68.4 64.0↑\, 74.1 38.9 56.0 Table 5: Cross-Task Generalization Benchmarks. Bold and underlined values denote the best and second-best results among the fine-tuned models, respectively. The superscript ↑ highlights performance improvements over the base model. TruthfulQA is evaluated using the standard MC1 (single-choice) and MC2 (multiple-choice) metrics. Our method effectively mitigates catastrophic forgetting while demonstrating positive transfer to code generation (HumanEval). Overall, performance variations across datasets are marginal, highlighting two primary trends. First, a slight decline on non-coding benchmarks suggests minor forgetting, yet this effect is constrained within a narrow range, likely due to the limited scale of our training samples. Notably, our method exhibits the most robust capability retention among all fine-tuned models. Second, on the HumanEval benchmark, which shares high logical synergy with mathematical reasoning, reasoning-enhanced models including LIMO, s1K, and our approach consistently show performance gains over the base model. Among these, our method achieves the most significant improvement, demonstrating effective positive transfer to programming tasks. 5 Analysis This section analyzes the characteristics of the data generated by our approach. To this end, we compare our method against the aforementioned Magpie Xu et al. (2025) and OpenR1-Math OpenR1 (2025). Specifically, OpenR1-Math is derived from NuminaMath but serves as a rigorously filtered, high-quality subset, thus offering higher analytical value than random sampling from the raw source. 5.1 Data Quality and Difficulty Following the protocol in Chen et al. (2024), we employ Qwen2.5-32B-Instruct to evaluate both the quality and difficulty of the datasets on a five-level scale (The relevant prompts are detailed in Appendix G.2). The resulting distributions are visualized in Figure˜2(a) and (b), respectively. For the manual evaluation of overall quality through sampling, please refer to Section˜5.4. Data Quality As shown in Figure˜2(a), both synthetic datasets yield a higher proportion of "excellent" samples compared to OpenR1-Math, with our method further outperforming the Magpie baseline. While the synthetic datasets exhibit a slightly higher share of "very poor" instances than OpenR1-Math, the absolute proportion remains negligible. This is an expected consequence of the synthetic data being evaluated in its raw state, whereas OpenR1-Math has undergone post-filtering. Figure 2: Quality and Difficulty Distributions. Quality and difficulty increase from left to right. Our method shows a significant advantage in generating high-quality, high-difficulty mathematical problems. Data Difficulty Figure˜2(b) demonstrates that our method holds a significant advantage in synthesizing high-difficulty problems, whereas Magpie is limited by its tendency to generate low-complexity, high-frequency data. In fact, our approach allows for flexible control over the adversarial iteration process via prompt engineering (e.g., style tokens). This capability enables the generation of problems at targeted difficulty levels, facilitating the construction of model-adaptive synthetic datasets. 5.2 Dataset Diversity Dataset diversity is widely acknowledged as a pivotal determinant of dataset quality. Despite the lack of a universal standard for its quantification, we employ two primary metrics in this study: intra-dataset similarity and dataset coverage. Intuitively, lower intra-similarity coupled with higher coverage indicates superior diversity. For experimental details of this section, please refer to Appendix E. Dataset Intra-Similarity To quantify dataset diversity, we measure intra-dataset similarity by calculating the average similarity between data instances within each dataset. We utilize the complete 220K-version OpenR1-Math dataset and randomly sample an equivalent number of instances (220K) from both our dataset and Magpie. The distribution of average similarity scores is shown in Figure˜3. The results demonstrate that our method achieves the lowest intra-dataset similarity. Figure 3: Intra-dataset Similarity Distribution. The bars represent frequency counts, while the curves denote Kernel Density Estimation (KDE). The results demonstrate that our method exhibits lower intra-similarity, indicating superior dataset diversity. Dataset Coverage The coverage of mathematical problems is primarily characterized by the diversity of the underlying knowledge points. To evaluate this, we adopt an approach integrating methods from InsTag Lu et al. (2024b) and Zhao et al. (2024). As illustrated in Figure˜4, we employ t-SNE (Maaten and Hinton, 2008) to project the semantic embeddings of knowledge point tags derived from 10K randomly sampled instances into a two-dimensional space. The results show that the distribution of the data generated by our method substantially encompasses the coverage areas of both Magpie and OpenR1-Math, which further validates the superior diversity of our approach in generating mathematical problems. Figure 4: t-SNE Visualization of Knowledge Points. The extensive coverage of the dark blue points (representing our method) demonstrates the significant diversity of the generated mathematical problems. 5.3 Data Scaling In this section, we investigate the scaling laws governing our approach by analyzing the correlation between model performance and dataset size. We employ Qwen2.5-Math-7B as the base model and adjust the training schedule to 3 epochs to facilitate efficient experimentation. As illustrated in Figure˜5, performance across all three datasets exhibits a trend of rapid initial growth followed by stabilization (or slight saturation) as data scale increases, with peak performance achieved at approximately 100K samples. Crucially, our method consistently yields superior performance gains compared to other methods, demonstrating its robustness and data efficiency. Figure 5: Performance Scaling Analysis. The x-axis is plotted on a logarithmic scale for clarity. While performance generally improves with increased data scale, our method maintains a consistent and significant performance advantage over the baselines. 5.4 Human Evaluation Given the potential systematic bias of relying solely on LLM-as-a-judge, we also conduct a human evaluation. Specifically, we randomly sample 100 instances from each of Magpie, OpenR1, and MathAgent, and perform a blinded evaluation with source anonymization on a 5-point scale, where higher scores indicate better quality. The results show that MathAgent achieves the highest average score (4.2/5.0), outperforming Magpie (3.5/5.0) and OpenR1 (3.9/5.0). Expert feedback further suggests that, although Magpie-generated samples are often grammatically fluent, their logical structures tend to be simplistic or repetitive, directly resulting in the lowest scores for Magpie in this evaluation. In contrast, MathAgent demonstrates stronger structural complexity and logical depth, consistent with the high-difficulty distribution observed in the automated analysis. We also observe cases where the LLM judge fails to identify subtle logical flaws recognized by human experts, highlighting the limitations of automated evaluation alone. Detailed evaluation settings and discussions of limitations are provided in Appendix F. 6 Conclusion By decoupling blueprint design from linguistic realization, this paper proposes the MathAgent framework, which maximizes structural diversity through adversarial optimization on constraint graphs to guide the generation of high-quality, long-tail distributed data. Experimental results demonstrate that models fine-tuned on samples synthesized via this approach outperform existing human-filtered datasets. This provides a scalable pathway for constructing high-quality mathematical reasoning data, breaking through the limitations of manual annotation and avoiding the mode collapse issues inherent in previous synthesis methods. Limitations Despite the performance gains and robust scaling characteristics demonstrated by MathAgent, our study has several limitations: • Domain Specificity. Our framework currently focuses on mathematical reasoning where logic is highly structured and objective. Its applicability to more open-ended or less formal domains, such as legal reasoning or creative writing, requires further exploration into how to define effective constraint graphs for non-mathematical tasks. • Computational Overhead. The adversarial evolution process involves multiple iterations between the legislator and executor models. This iterative cycle inevitably leads to higher computational costs and longer synthesis times compared to single-pass methods that do not require multi-round refinement. • Ground-Truth Verification. As the adversarial process pushes problem complexity to extreme levels, ensuring the absolute correctness of the generated ground-truth solutions becomes increasingly difficult. Future work could benefit from integrating external formal verifiers or symbolic solvers to guarantee the accuracy of synthesized reasoning chains. These limitations also point to potential research directions for the future. Ethics Statement This work complies with the ACL Ethics Policy. We utilize publicly available datasets and open-source models, ensuring that all resources are properly cited and the experiments are reproducible. We acknowledge the inherent risks associated with LLMs, including the potential for hallucinations and the generation of non-factual content. While our method aims to enhance reasoning reliability in the mathematical domain, users should exercise caution and verify model outputs when deploying such systems in critical or real-world applications. Acknowledgments The authors would like to express their sincere gratitude to the PhD students in the Department of Mathematical Sciences and the Department of Statistics and Data Science at Tsinghua University for their dedicated efforts in the manual verification process. We also thank the Huawei Large Model Data Technology Lab for its invaluable guidance and suggestions. Furthermore, we are grateful to the anonymous reviewers and the meta-reviewer for their insightful and constructive comments, which have greatly improved the quality of this work. References A. o. P. S. AoPS (2023) AMC Problems and Solutions. Note: https://artofproblemsolving.com/wiki/index.php/AMC_12_Problems_and_SolutionsAccessed: 2025-09-30 Cited by: 3rd item. A. o. P. S. AoPS (2025) AIME Problems and Solutions. Note: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_SolutionsAccessed: 2025-09-30 Cited by: 3rd item. D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes (2023) Autonomous chemical research with large language models. Nature 624 (7992), p. 570–578. External Links: Link Cited by: §1. T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei (2020) Language models are few-shot learners. In Advances in Neural Information Processing Systems, Vol. 33, p. 1877–1901. External Links: Link Cited by: §1. L. Chen, S. Li, J. Yan, H. Wang, K. Gunaratna, V. Yadav, Z. Tang, V. Srinivasan, T. Zhou, H. Huang, and H. Jin (2024) AlpaGasus: training a better alpaca with fewer data. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §G.2, §5.1. M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba (2021) Evaluating large language models trained on code. External Links: 2107.03374, Link Cited by: §4.5. S. Chen, C. Tian, B. Hu, K. Chen, Z. Liu, Z. Zhang, and J. Zhou (2025) Arrows of math reasoning data synthesis for large language models: diversity, complexity and correctness. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management, CIKM ’25, New York, NY, USA, p. 4665–4669. External Links: ISBN 9798400720406, Link, Document Cited by: Appendix D. K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman (2021) Training verifiers to solve math word problems. External Links: 2110.14168, Link Cited by: 1st item. DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang (2025a) DeepSeek-r1: incentivizing reasoning capability in llms via reinforcement learning. External Links: 2501.12948, Link Cited by: Appendix D. DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Wang, J. Chen, J. Chen, J. Yuan, J. Qiu, J. Li, J. Song, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Xu, L. Xia, L. Zhao, L. Wang, L. Zhang, M. Li, M. Wang, M. Zhang, M. Zhang, M. Tang, M. Li, N. Tian, P. Huang, P. Wang, P. Zhang, Q. Wang, Q. Zhu, Q. Chen, Q. Du, R. J. Chen, R. L. Jin, R. Ge, R. Zhang, R. Pan, R. Wang, R. Xu, R. Zhang, R. Chen, S. S. Li, S. Lu, S. Zhou, S. Chen, S. Wu, S. Ye, S. Ye, S. Ma, S. Wang, S. Zhou, S. Yu, S. Zhou, S. Pan, T. Wang, T. Yun, T. Pei, T. Sun, W. L. Xiao, W. Zeng, W. Zhao, W. An, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, X. Q. Li, X. Jin, X. Wang, X. Bi, X. Liu, X. Wang, X. Shen, X. Chen, X. Zhang, X. Chen, X. Nie, X. Sun, X. Wang, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yu, X. Song, X. Shan, X. Zhou, X. Yang, X. Li, X. Su, X. Lin, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. X. Zhu, Y. Zhang, Y. Xu, Y. Xu, Y. Huang, Y. Li, Y. Zhao, Y. Sun, Y. Li, Y. Wang, Y. Yu, Y. Zheng, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Tang, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Wu, Y. Ou, Y. Zhu, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Zha, Y. Xiong, Y. Ma, Y. Yan, Y. Luo, Y. You, Y. Liu, Y. Zhou, Z. F. Wu, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Huang, Z. Zhang, Z. Xie, Z. Zhang, Z. Hao, Z. Gou, Z. Ma, Z. Yan, Z. Shao, Z. Xu, Z. Wu, Z. Zhang, Z. Li, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Gao, and Z. Pan (2025b) DeepSeek-v3 technical report. External Links: 2412.19437, Link Cited by: §4.1. D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence (2023) PaLM-e: an embodied multimodal language model. In Proceedings of the 40th International Conference on Machine Learning, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (Eds.), Proceedings of Machine Learning Research, Vol. 202, p. 8469–8488. External Links: Link Cited by: §1. Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch (2024) Improving factuality and reasoning in language models through multiagent debate. In Forty-first International Conference on Machine Learning, External Links: Link Cited by: §2. Google (2024) Gemini 2.0 flash thinking mode (gemini-2.0-flash-thinking-exp-1219). Note: Accessed via Google Cloud Vertex AI External Links: Link Cited by: Appendix D. A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Guzmán, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Çelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma (2024) The llama 3 herd of models. External Links: 2407.21783, Link Cited by: Table 6, Table 6, §4.1. A. Gudibande, E. Wallace, C. Snell, X. Geng, H. Liu, P. Abbeel, S. Levine, and D. Song (2023) The false promise of imitating proprietary llms. External Links: 2305.15717, Link Cited by: §1. E. K. Guha, R. Marten, S. Keh, N. Raoof, G. Smyrnis, H. Bansal, M. Nezhurina, J. Mercat, T. Vu, Z. R. Sprague, A. Suvarna, B. Feuer, L. L. Chen, Z. Khan, E. Frankel, S. Grover, C. Choi, N. Muennighoff, S. Su, W. Zhao, J. Yang, S. Pimpalgaonkar, K. sharma, C. C. Ji, Y. Deng, S. M. Pratt, V. Ramanujan, J. Saad-Falcon, S. Acharya, J. Li, A. Dave, A. Albalak, K. Arora, B. Wulfe, C. Hegde, G. Durrett, S. Oh, M. Bansal, S. Gabriel, A. Grover, K. Chang, V. Shankar, A. Gokaslan, M. A. Merrill, T. Hashimoto, Y. Choi, J. Jitsev, R. Heckel, M. Sathiamoorthy, A. Dimakis, and L. Schmidt (2026) OpenThoughts: data recipes for reasoning models. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: Appendix D. C. He, R. Luo, Y. Bai, S. Hu, Z. Thai, J. Shen, J. Hu, X. Han, Y. Huang, Y. Zhang, J. Liu, L. Qi, Z. Liu, and M. Sun (2024) OlympiadBench: a challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 3828–3850. External Links: Link, Document Cited by: 2nd item. D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt (2021a) Measuring massive multitask language understanding. In International Conference on Learning Representations, External Links: Link Cited by: §4.5. D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt (2021b) Measuring mathematical problem solving with the math dataset. External Links: 2103.03874, Link Cited by: 1st item. J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre (2022) Training compute-optimal large language models. External Links: 2203.15556, Link Cited by: §1. S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu, and J. Schmidhuber (2024) MetaGPT: meta programming for a multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §2. O. Honovich, T. Scialom, O. Levy, and T. Schick (2023) Unnatural instructions: tuning language models with (almost) no human labor. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, p. 14409–14428. External Links: Link, Document Cited by: §1. A. Jacot, F. Gabriel, and C. Hongler (2018) Neural tangent kernel: convergence and generalization in neural networks. In Advances in Neural Information Processing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31, p. . External Links: Link Cited by: §1. A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed (2023) Mistral 7b. External Links: 2310.06825, Link Cited by: §4.1. J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei (2020) Scaling laws for neural language models. External Links: 2001.08361, Link Cited by: §1. X. Ke, H. Deng, X. Liu, J. Rao, Z. Song, J. Yu, and M. Zhang (2025) AQuilt: weaving logic and self-inspection into low-cost, high-relevance data synthesis for specialist LLMs. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 5752–5785. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1. A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra (2022) Solving quantitative reasoning problems with language models. External Links: 2206.14858, Link Cited by: 2nd item. G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem (2023) CAMEL: communicative agents for "mind" exploration of large language model society. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, p. 51991–52008. External Links: Link Cited by: §2. J. Li, E. Beeching, L. Tunstall, B. Lipkin, R. Soletskyi, S. Huang, K. Rasul, L. Yu, A. Q. Jiang, Z. Shen, et al. (2024a) Numinamath: the largest public dataset in ai4maths with 860k pairs of competition math problems and solutions. Hugging Face repository. External Links: Link Cited by: Appendix D, Appendix D, §4.4, Table 4, Table 4. Y. Li, Z. Yu, G. Chen, and Q. Lin (2024b) On the eigenvalue decay rates of a class of neural-network related kernel functions defined on general domains. Journal of Machine Learning Research 25 (82), p. 1–47. External Links: Link Cited by: §1. T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, S. Shi, and Z. Tu (2024) Encouraging divergent thinking in large language models through multi-agent debate. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 17889–17904. External Links: Link, Document Cited by: §2. M. Liao, C. Li, W. Luo, W. Jing, and K. Fan (2024) MARIO: MAth reasoning with code interpreter output - a reproducible pipeline. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 905–924. External Links: Link, Document Cited by: 2nd item. S. Lin, J. Hilton, and O. Evans (2022) TruthfulQA: measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, p. 3214–3252. External Links: Link, Document Cited by: §4.5. I. Loshchilov and F. Hutter (2017) SGDR: stochastic gradient descent with warm restarts. In International Conference on Learning Representations, External Links: Link Cited by: Appendix B. I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. In International Conference on Learning Representations, External Links: Link Cited by: Appendix B. C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha (2024a) The ai scientist: towards fully automated open-ended scientific discovery. External Links: 2408.06292, Link Cited by: §1. K. Lu, H. Yuan, Z. Yuan, R. Lin, J. Lin, C. Tan, C. Zhou, and J. Zhou (2024b) #InsTag: instruction tagging for analyzing supervised fine-tuning of large language models. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §5.2. Z. Lu, A. Zhou, H. Ren, K. Wang, W. Shi, J. Pan, M. Zhan, and H. Li (2024c) MathGenie: generating synthetic data with question back-translation for enhancing mathematical reasoning of LLMs. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 2732–2747. External Links: Link, Document Cited by: §2. H. Luo, Q. Sun, C. Xu, P. Zhao, J. Lou, C. Tao, X. Geng, Q. Lin, S. Chen, Y. Tang, and D. Zhang (2025) WizardMath: empowering mathematical reasoning for large language models via reinforced evol-instruct. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §2. L. v. d. Maaten and G. Hinton (2008) Visualizing data using t-sne. Journal of machine learning research 9 (Nov), p. 2579–2605. External Links: Link Cited by: §5.2. M. Maosongcao, T. Zhang, M. Li, C. Zhang, Y. Liu, C. He, H. Duan, S. Zhang, and K. Chen (2025) Condor: enhance LLM alignment with knowledge-driven data synthesis and refinement. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 22392–22412. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §2. T. Meta (2024) Llama 3.2: revolutionizing edge ai and vision with open, customizable models. Note: Blog Post External Links: Link Cited by: §4.1. N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei-Fei, H. Hajishirzi, L. Zettlemoyer, P. Liang, E. Candès, and T. Hashimoto (2025) S1: simple test-time scaling. External Links: 2501.19393, Link Cited by: Appendix D, Appendix D, §4.1, Table 4. T. OpenR1 (2025) open-r1/OpenR1-Math-220k. Note: https://huggingface.co/datasets/open-r1/OpenR1-Math-220kAccessed: 2025-09-30 Cited by: Appendix D, Appendix D, §5. T. Qwen (2024) QwQ: reflect deeply on the boundaries of the unknown. Note: Blog Post External Links: Link Cited by: §4.1. S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He (2020) ZeRO: memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, Vol. , p. 1–16. External Links: Document, Link Cited by: Appendix B. J. Rao, Y. Liao, X. Liu, Z. Lin, L. Lian, D. Jin, S. Cheng, J. Yu, and M. Zhang (2025a) SeaPO: strategic error amplification for robust preference optimization of large language models. In Findings of the Association for Computational Linguistics: EMNLP 2025, External Links: Link Cited by: §1, §2. J. Rao, Z. Lin, X. Liu, X. Ke, L. Lian, D. Jin, S. Cheng, J. Yu, and M. Zhang (2025b) APT: improving specialist LLM performance with weakness case acquisition and iterative preference training. In Findings of the Association for Computational Linguistics: ACL 2025, Vienna, Austria, p. 20958–20980. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §3.2. J. Rao, X. Liu, H. Deng, Z. Lin, Z. Yu, J. Wei, X. Meng, and M. Zhang (2026) Dynamic sampling that adapts: iterative dpo for self-aware mathematical reasoning. In Findings of the Association for Computational Linguistics: ACL 2026, External Links: Link Cited by: §4.4. J. Rao, X. Liu, H. Yan, J. Shen, H. Mo, Y. Dong, Z. Yan, Z. Wang, Z. Lin, X. Meng, Z. Yu, L. Deng, J. Wei, Y. Wang, and M. Zhang (2025c) A data-centric perspective on the lifecycle of large language models. TechRxiv 2025 (1220), p. . External Links: Document, Link, https://w.techrxiv.org/doi/pdf/10.36227/techrxiv.176620610.03288677/v1 Cited by: Appendix D. N. Reimers and I. Gurevych (2019) Sentence-BERT: sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), K. Inui, J. Jiang, V. Ng, and X. Wan (Eds.), Hong Kong, China, p. 3982–3992. External Links: Link, Document Cited by: Appendix C. T. G. M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, J. Ferret, P. Liu, P. Tafti, A. Friesen, M. Casbon, S. Ramos, R. Kumar, C. L. Lan, S. Jerome, A. Tsitsulin, N. Vieillard, P. Stanczyk, S. Girgin, N. Momchev, M. Hoffman, S. Thakoor, J. Grill, B. Neyshabur, O. Bachem, A. Walton, A. Severyn, A. Parrish, A. Ahmad, A. Hutchison, A. Abdagic, A. Carl, A. Shen, A. Brock, A. Coenen, A. Laforge, A. Paterson, B. Bastian, B. Piot, B. Wu, B. Royal, C. Chen, C. Kumar, C. Perry, C. Welty, C. A. Choquette-Choo, D. Sinopalnikov, D. Weinberger, D. Vijaykumar, D. Rogozińska, D. Herbison, E. Bandy, E. Wang, E. Noland, E. Moreira, E. Senter, E. Eltyshev, F. Visin, G. Rasskin, G. Wei, G. Cameron, G. Martins, H. Hashemi, H. Klimczak-Plucińska, H. Batra, H. Dhand, I. Nardini, J. Mein, J. Zhou, J. Svensson, J. Stanway, J. Chan, J. P. Zhou, J. Carrasqueira, J. Iljazi, J. Becker, J. Fernandez, J. van Amersfoort, J. Gordon, J. Lipschultz, J. Newlan, J. Ji, K. Mohamed, K. Badola, K. Black, K. Millican, K. McDonell, K. Nguyen, K. Sodhia, K. Greene, L. L. Sjoesund, L. Usui, L. Sifre, L. Heuermann, L. Lago, L. McNealus, L. B. Soares, L. Kilpatrick, L. Dixon, L. Martins, M. Reid, M. Singh, M. Iverson, M. Görner, M. Velloso, M. Wirth, M. Davidow, M. Miller, M. Rahtz, M. Watson, M. Risdal, M. Kazemi, M. Moynihan, M. Zhang, M. Kahng, M. Park, M. Rahman, M. Khatwani, N. Dao, N. Bardoliwalla, N. Devanathan, N. Dumai, N. Chauhan, O. Wahltinez, P. Botarda, P. Barnes, P. Barham, P. Michel, P. Jin, P. Georgiev, P. Culliton, P. Kuppala, R. Comanescu, R. Merhej, R. Jana, R. A. Rokni, R. Agarwal, R. Mullins, S. Saadat, S. M. Carthy, S. Cogan, S. Perrin, S. M. R. Arnold, S. Krause, S. Dai, S. Garg, S. Sheth, S. Ronstrom, S. Chan, T. Jordan, T. Yu, T. Eccles, T. Hennigan, T. Kocisky, T. Doshi, V. Jain, V. Yadav, V. Meshram, V. Dharmadhikari, W. Barkley, W. Wei, W. Ye, W. Han, W. Kwon, X. Xu, Z. Shen, Z. Gong, Z. Wei, V. Cotruta, P. Kirk, A. Rao, M. Giang, L. Peran, T. Warkentin, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, D. Sculley, J. Banks, A. Dragan, S. Petrov, O. Vinyals, J. Dean, D. Hassabis, K. Kavukcuoglu, C. Farabet, E. Buchatskaya, S. Borgeaud, N. Fiedel, A. Joulin, K. Kenealy, R. Dadashi, and A. Andreev (2024) Gemma 2: improving open language models at a practical size. External Links: 2408.00118, Link Cited by: §4.1, Table 2. I. Shumailov, Z. Shumaylov, Y. Zhao, Y. Gal, N. Papernot, and R. Anderson (2023) The curse of recursion: training on generated data makes models forget. External Links: 2305.17493, Link Cited by: §1. M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. Le, E. Chi, D. Zhou, and J. Wei (2023) Challenging BIG-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, p. 13003–13051. External Links: Link, Document Cited by: §4.5. S. Tang, X. Pang, Z. Liu, B. Tang, R. Ye, T. Jin, X. Dong, Y. Wang, and S. Chen (2025) Synthesizing post-training data for LLMs through multi-agent simulation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 23306–23335. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §2. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. In Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30, p. . External Links: Link Cited by: §1. P. Villalobos, A. Ho, J. Sevilla, T. Besiroglu, L. Heim, and M. Hobbhahn (2024) Position: will we run out of data? limits of LLM scaling based on human-generated data. In Forty-first International Conference on Machine Learning, External Links: Link Cited by: §1. Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi (2023) Self-instruct: aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, p. 13484–13508. External Links: Link, Document Cited by: Appendix D, §1, §2. Z. Wang, Y. Wang, X. Liu, L. Ding, M. Zhang, J. Liu, and M. Zhang (2025) AgentDropout: dynamic agent elimination for token-efficient and high-performance LLM-based multi-agent collaboration. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 24013–24035. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §2. J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus (2022a) Emergent abilities of large language models. Transactions on Machine Learning Research. Note: Survey Certification External Links: ISSN 2835-8856, Link Cited by: §1. J. Wei, X. Wang, D. Schuurmans, M. Bosma, b. ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou (2022b) Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35, p. 24824–24837. External Links: Link Cited by: §1. K. Wei, R. Shan, D. Zou, J. Yang, B. Zhao, J. Zhu, and J. Zhong (2026) Mirage: scaling test-time inference with parallel graph-retrieval-augmented reasoning chains. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 33818–33826. External Links: Link Cited by: §1. Z. Xu, F. Jiang, L. Niu, Y. Deng, R. Poovendran, Y. Choi, and B. Y. Lin (2025) Magpie: alignment data synthesis from scratch by prompting aligned LLMs with nothing. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: Appendix D, Appendix D, §G.2, §1, §2, §4.4, Table 4, Table 4, §5. A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu (2025a) Qwen3 technical report. External Links: 2505.09388, Link Cited by: Table 6, §4.1. A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu (2025b) Qwen2.5 technical report. External Links: 2412.15115, Link Cited by: §4.1, §4.1, Table 2, Table 4, Table 4. A. Yang, B. Zhang, B. Hui, B. Gao, B. Yu, C. Li, D. Liu, J. Tu, J. Zhou, J. Lin, K. Lu, M. Xue, R. Lin, T. Liu, X. Ren, and Z. Zhang (2024) Qwen2.5-math technical report: toward mathematical expert model via self-improvement. External Links: 2409.12122, Link Cited by: §4.1, §4.3, Table 2. Y. Ye, Z. Huang, Y. Xiao, E. Chern, S. Xia, and P. Liu (2025) LIMO: less is more for reasoning. In Second Conference on Language Modeling, External Links: Link Cited by: Appendix D, Appendix D, §1, §4.1, Table 4. Z. Yu, S. Tian, and G. Chen (2025) Divergence of empirical neural tangent kernel in classification problems. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1. S. Zeng, X. Chang, M. Xie, X. Liu, Y. Bai, Z. Pan, M. Xu, and X. Wei (2025a) FutureSightDrive: thinking visually with spatio-temporal cot for autonomous driving. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §1. S. Zeng, D. Qi, X. Chang, F. Xiong, X. Shichao, X. Wu, S. Liang, M. Xu, and X. Wei (2026) JanusVLN: decoupling semantics and spatiality with dual implicit memory for vision-language navigation. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1. W. Zeng, Y. Huang, Q. Liu, W. Liu, K. He, Z. MA, and J. He (2025b) SimpleRL-zoo: investigating and taming zero reinforcement learning for open base models in the wild. In Second Conference on Language Modeling, External Links: Link Cited by: Appendix A. Y. Zeng, W. Yu, Z. Li, T. Ren, Y. Ma, J. Cao, X. Chen, and T. Yu (2025c) Bridging the editing gap in LLMs: FineEdit for precise and targeted text modifications. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 2193–2206. External Links: Link, ISBN 979-8-89176-335-7 Cited by: §2. H. Zhang, J. Yao, C. Ye, W. Xiong, and T. Zhang (2025) Online-dpo-r1: unlocking effective reasoning without the ppo overhead, 2025. Notion Blog. External Links: Link Cited by: §4.4. S. Zhang and D. Xiong (2025) Debate4MATH: multi-agent debate for fine-grained reasoning in math. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 16810–16824. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §2. W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, Y. Du, C. Yang, Y. Chen, Z. Chen, J. Jiang, R. Ren, Y. Li, X. Tang, Z. Liu, P. Liu, J. Nie, and J. Wen (2023) A survey of large language models. External Links: 2303.18223, Link Cited by: §1. W. Zhao, X. Ren, J. Hessel, C. Cardie, Y. Choi, and Y. Deng (2024) WildChat: 1m chatGPT interaction logs in the wild. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §5.2. X. Zhao, X. Hu, Z. Shan, S. Huang, Y. Zhou, X. Zhang, Z. Sun, zhenyu liu, D. Li, X. Wei, Y. Pan, Y. Xiang, M. Zhang, H. Wang, J. Yu, B. Hu, and M. Zhang (2026) KaLM-embedding-v2: superior training techniques and data inspire a versatile embedding model. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1. L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica (2023) Judging LLM-as-a-judge with MT-bench and chatbot arena. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §3.3. Y. Zheng, R. Zhang, J. Zhang, Y. Ye, and Z. Luo (2024) LlamaFactory: unified efficient fine-tuning of 100+ language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), Y. Cao, Y. Feng, and D. Xiong (Eds.), Bangkok, Thailand, p. 400–410. External Links: Link, Document Cited by: Appendix B. C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y. Mao, X. Ma, A. Efrat, P. Yu, L. YU, S. Zhang, G. Ghosh, M. Lewis, L. Zettlemoyer, and O. Levy (2023) LIMA: less is more for alignment. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, p. 55006–55021. External Links: Link Cited by: §1, §4.1. Appendix A Prompts for model training and testing Standard prompt designs in current mathematical reasoning tasks typically combine zero-shot Chain-of-Thought (CoT) guidance with structured output constraints. This approach strikes a balance between maintaining model accuracy and facilitating automated answer extraction and evaluation. Thus, we adopt this strategy in our study (refer to the Complex Prompt in Figure˜6). However, existing research indicates that such complex prompts may impose a burden on models with limited instruction-following capabilities (Zeng et al., 2025b), potentially leading to performance degradation. Consequently, a simplified prompt design is adopted for such models (refer to the Simple Prompt in Figure˜6). Simple Prompt Question: input Answer: Let’s think step by step. Complex Prompt <|im_start|>system You are a helpful assistant.<|im_end|> <|im_start|>user input Please reason step by step, and put your final answer within \ .<|im_end|> <|im_start|>assistant output Figure 6: Comparison between simple prompts and more complex prompts (using the Qwen series prompt templates as an example). We empirically verify this phenomenon in Table˜6, using Qwen3-8B-Base and Llama-3.1-8B as representative models. Experimental results reveal that the non-instruction-tuned Llama-3.1-8B suffers significant performance degradation under complex prompts (adapted to the corresponding Llama template), whereas Qwen3-8B-Base demonstrates improved performance in similar settings. Based on further observations, we adopt the simple prompt for the Llama, Mistral, and Gemma series, while retaining the standard complex prompt format for the Qwen series. Model Prompt Template Avg. Qwen3-8B-Base (Yang et al., 2025a) simple 32.5 Qwen-complex 49.7 Llama-3.1-8B (Grattafiori et al., 2024) simple 10.3 Llama-complex 1.4 Llama-3.1-8B-Instruct (Grattafiori et al., 2024) simple 23.5 Llama-complex 31.0 Table 6: Performance comparison using Simple vs. Complex prompts. Complex prompts degrade the performance of Llama-Base but benefit Qwen-Base and Llama-Instruct. Notably, although our experiments primarily applied complex prompt strategies to Qwen models, the applicability of this design is not limited to this series. Instead, as previously discussed, it is closely tied to the model’s instruction-following capability. As shown in Table˜6, the instruction-tuned Llama-3.1-8B-Instruct also demonstrates significant performance improvement when using complex prompts (adapted to the corresponding Llama template). However, considering that instruction-tuned models are generally less suitable for few-shot supervised fine-tuning, our main experiments do not focus on them. Appendix B Training Details for SFT We perform full-parameter fine-tuning of the model using LLaMA-Factory222https://github.com/hiyouga/LLaMA-Factory (Zheng et al., 2024). Adhering to the strategies discussed in Appendix˜A, we apply the corresponding prompt templates (Simple or Complex) for each model during the data formatting stage. To optimize memory usage and training efficiency, we employ the DeepSpeed ZeRO Stage 3 (Rajbhandari et al., 2020) strategy. Training utilizes the AdamW (Loshchilov and Hutter, 2019) optimizer (β1=0.9,β2=0.95 _1=0.9, _2=0.95, weight decay=1×10−4weight decay=1× 10^-4) combined with a cosine learning rate decay schedule (Loshchilov and Hutter, 2017) and a warmup ratio of 0.050.05. The maximum sequence length is limited to 40964096 tokens, and the maximum gradient norm is clipped at 1.01.0. The training process uses bfloat16 precision throughout to ensure numerical stability and computational efficiency, and operates in a distributed environment across 88 devices, with the random seed fixed at 4242 to ensure reproducibility. Certain hyperparameters, such as the initial learning rate, batch size, and number of training epochs, vary across models; these specific configurations are detailed in Table˜7. Model Peak Learning Rate Batch Size Training Epochs Qwen-7B/8B Series 1e-5 2×82× 8 5 & Llama-3.1-8B Qwen3-14B-Base 5e-6 1×81× 8 5 & Gemma-2-9B Qwen3-4B-Base 2e-5 4×84× 8 5 & Llama-3.2-4B Mistral-7B-v0.3 1e-5 2×82× 8 3 Table 7: Training Hyperparameters for Different Models. Note that "×8× 8" in the batch size column denotes the total batch size calculated as per-device batch size × number of devices (8). Appendix C Training-Test Set Similarity Figure 7: Average Maximum Similarity Between Training and Test Datasets: A lower similarity score indicates a reduced likelihood of data leakage. This figure demonstrates that our synthetic data does not carry a higher risk of data leakage compared to LIMO and s1K. To prevent potentially biased test results caused by data leakage during synthetic data generation, we calculated the similarity between each of the three training sets and the test set, respectively. We adopt the average maximum similarity (AMS) in the embedding space as our metric. Specifically, we first map all mathematical problems into the embedding space using the all-mpnet-base-v2333https://huggingface.co/sentence-transformers/all-mpnet-base-v2 model (Reimers and Gurevych, 2019), which enables the computation of pairwise similarity between any two problems. For each instance in the test set, we retrieve the sample in the training set with the highest similarity to it and record this maximum similarity value. The final metric is then calculated as the average of these maximum similarity values across all test instances. The final results are presented in Figure˜7. It is evident that the similarity between our synthetic dataset and the test sets is not significantly higher than that of the other training sets. Conversely, for most test sets, the similarity metric obtained by our method is noticeably lower. This indicates that no additional information leakage occurs during our synthetic data generation process, and thus it does not introduce bias into the experimental results. Appendix D Details for Baseline Data Comparisons In this section, we describe the selection criteria and implementation details of the baseline data sources used in our performance evaluation and data characteristic analysis. Overall Criteria Open-source datasets and synthesis methods are highly diverse (Rao et al., 2025c). Because our method targets seed-free synthesis of mathematical problems, we mainly compare against methods in the same paradigm, such as Magpie (Xu et al., 2025), rather than seed-based approaches (e.g., Self-Instruct; Wang et al. 2023). Methods such as OpenThought (Guha et al., 2026), which focus on constructing high-quality reasoning traces from existing public datasets, are complementary to our work and therefore excluded from direct comparison. In addition, program-assisted generation frameworks such as ARROWS (Chen et al., 2025) are effective when problems can be formulated as executable programs, but are limited to such settings and are therefore excluded from direct comparison. Since many open-source math problems are derived from NuminaMath (Li et al., 2024a), we include samples from it as a raw-data baseline, together with several representative curated subsets built upon it (Ye et al., 2025; Muennighoff et al., 2025; OpenR1, 2025). LIMO & s1K We employ the open-source LIMO (Ye et al., 2025) and s1K (Muennighoff et al., 2025) datasets as high-quality baselines. These datasets are constructed through expert-designed pipelines and rigorous screening to ensure reasoning depth, representing the state-of-the-art in small-scale, curated reasoning data. For s1K, which offers two reasoning model versions based on Gemini (Google, 2024) and DeepSeek-R1 (DeepSeek-AI et al., 2025a) respectively, we specifically select the latter for the experiments in Table˜1 to establish a more competitive and rigorous baseline. NuminaMath NuminaMath (Li et al., 2024a) serves as a large-scale open-source collection of mathematical problems, providing a rich resource of raw mathematical tasks for research. In our experiments, we randomly sample the required number of problems from the complete dataset to evaluate the model’s performance across a broad distribution of raw mathematical tasks. Magpie To construct the Magpie (Xu et al., 2025) baseline, we employ Qwen2.5-32B-Instruct as the question generator. Following the original protocol, we adopt the math-specific prompt from Figure 2 of Xu et al. (2025) as the system prompt. Critically, Magpie’s generation quality is highly sensitive to the sampling temperature T. While higher temperatures can enhance dataset diversity, they often lead to a non-negligible decline in synthesis quality; thus, we adopt a balanced setting of T=1T=1 to ensure a fair and effective comparison. Unlike our iterative adversarial approach, Magpie relies on a static, uniform system prompt for direct generation. For the comprehensive data characteristic analysis, we scale this process to generate a candidate pool of 300,000 samples. OpenR1-Math OpenR1-Math (OpenR1, 2025) is an open-source dataset derived from NuminaMath. As a rigorously filtered and high-quality subset, it offers significantly higher analytical value than the original raw data. The dataset consists of two versions, including a further curated set of 94k problems and a complete version containing 220k instances. For experiments that prioritize data quality over quantity, such as assessments of data quality and difficulty as well as SFT experiments for data scaling performance, we utilize OpenR1-Math-94k as the comparison baseline. In contrast, for diversity analysis, we select the larger OpenR1-Math-220k version to provide a more representative and comprehensive benchmark. Appendix E Details for Diversity Analysis In the internal similarity and coverage analysis, we uniformly use the all-mpnet-base-v2 model to generate semantic embedding vectors, mapping problem descriptions and knowledge tags into a 768-dimensional vector space. The internal similarity is calculated as the average cosine similarity between the embedding vector of each sample and all other samples in the dataset pool. In the coverage analysis, the knowledge point extraction stage employs Qwen2.5-32B-Instruct to perform annotation on 10,000 randomly sampled mathematical problems, followed by dimensionality reduction visualization using t-SNE. Notably, the similarity metric used here differs from that in Appendix˜C. While the latter employs maximum similarity to detect potential data contamination from near-identical samples, the average similarity used in this section is designed to evaluate global semantic density. Furthermore, due to the high baseline similarity inherent in mathematical reasoning data, maximum similarity tends to saturate (approach 1.0) at large scales, thereby reducing its discriminative power. Consequently, we adopt average similarity to provide a more robust and distinguishable measure of dataset diversity. Appendix F Human Evaluation Details Evaluation Setup To complement the automated evaluation, we conduct an expert-based human evaluation on samples drawn from Magpie, OpenR1, and MathAgent. We invite four Ph.D. researchers with expertise in different areas of mathematics, including Algebra, Geometry & Analysis, Statistics, and Mathematical Physics, to ensure broad domain coverage. We randomly sample 100 instances from each dataset, resulting in 300 evaluated instances in total. The evaluation is conducted under a blinded protocol with source anonymization, so evaluators do not know which method produced each sample. Each instance is rated on a 5-point scale, where higher scores indicate better overall quality: • 1 (Serious Flaws): The sample contains major logical or conceptual errors. • 2 (Minor Defects): The sample has small errors that do not completely invalidate it. • 3 (Mediocre): The sample is largely correct but trivial, repetitive, or lacking in depth. • 4 (Good Quality): The sample is correct, clear, and reasonably well-structured. • 5 (Insightful): The sample demonstrates strong logical depth, originality, or pedagogical value. Additional Qualitative Findings Beyond the quantitative scores, the evaluators provide several consistent qualitative observations. First, Magpie-generated samples are often grammatically fluent and superficially coherent, but they frequently rely on repetitive reasoning patterns or exhibit limited structural complexity. As a result, their overall quality is often judged as moderate rather than strong. Second, samples generated by MathAgent are more likely to exhibit multi-step reasoning, richer structural organization, and greater logical depth. These observations are consistent with the high-difficulty tendencies identified in the automated evaluation. Third, we observe noticeable discrepancies between human judgments and LLM-based evaluation in some cases. In particular, the LLM judge occasionally assigns favorable scores to samples that contain subtle logical gaps or weak inferential steps that are readily identified by human experts. This suggests that automated evaluation alone may fail to fully capture the quality of mathematical reasoning data. Evaluation Limitations Despite its usefulness, the human evaluation has several limitations. First, due to the high cost and cognitive demand of expert review, the evaluation is conducted on a relatively small sample of 300 instances, which is much smaller than the full scale of the synthesized dataset. Second, although we intentionally recruit evaluators from diverse mathematical areas, evaluator coverage remains limited relative to the full breadth of mathematical reasoning tasks. Some problems may fall closer to the expertise of certain evaluators than others. As a result, especially under limited evaluation time, some particularly difficult samples may be marked as uncertain or out-of-scope. Third, human evaluation inevitably contains some degree of subjectivity, especially for high-level criteria such as insightfulness, elegance, or pedagogical value. While the scoring criteria help standardize judgments, they cannot completely eliminate individual variation in scoring. Appendix G Prompts for Synthetic Data Generation and Analysis G.1 Prompts for Data Synthesis In this section, we delineate the specific prompts employed throughout the synthetic mathematical problem generation pipeline. The hierarchical structure is organized as follows: within the Legislator module, the prompt for the Proposer is detailed in Figure˜9, the Critic’s evaluation protocol is provided in Figure˜10, and the Moderator’s decision-making logic is given in Figure˜11. Regarding the Executor module, we focus on its primary function, which is the Semantic Instantiation process. Figure˜12 illustrates the prompt design for transforming abstract constraint graphs (∗G^*) and style tokens (S) into fluent, natural language mathematical problems. To maintain focus on our structural innovations, the subsequent solution generation and model-based verification processes are not further elaborated herein, as they follow established heuristic workflows. G.2 Prompts for Analysis The prompts used for evaluating the quality and difficulty of the dataset (as discussed in Section˜5.1) are designed with reference to the evaluation protocol proposed by Chen et al. (2024) and are adapted from the methodology of Xu et al. (2025). The detailed prompt templates are illustrated in Figure˜13. Appendix H Case Study We illustrate the core iterative process of the MathAgent framework with a representative example. For demonstration, the following style tokens are selected: Difficulty: Medium, Question Type: Calculation, Context: Real-world Application, and Knowledge Level: Undergraduate. It should be noted that although this case is very similar to the one presented in Figure˜1, Figure˜1 primarily focuses on providing an intuitive understanding of the process. For ease of layout and process presentation, the case has been adapted accordingly. As shown in Figure˜8 (illustrating a two-round iteration; the original JSON output has been reformatted for clarity), the i-th node is denoted as viv_i, and the directed edge from node i to node j is denoted as eije_ij. Nodes follow the “Concept: Description” format, conforming to the structural specifications defined in Figure˜9. Iteration 1 The initial topic is set to the Saddle Surface, which serves as a rudimentary conceptual primitive. Driven by the Proposer (PA_P), Graph 1 is generated. The Critic (CA_C) then provides a key assessment regarding specification alignment: if the parameters a and b in the saddle surface equation z=x2/a2−y2/b2z=x^2/a^2-y^2/b^2 are not constrained by a=ba=b, the resulting surface integral may involve elliptic integrals, exceeding the intended Medium difficulty. Consequently, the Critic recommends adding this constraint and further suggests including explicit nodes for partial derivatives required for surface area calculation, while leveraging symmetry to simplify the computation. Iteration 2 The Moderator (MA_M) adopts these suggestions and provides iterative guidance to the Proposer. Based on this input, the Proposer updates the graph topology to produce Graph 2. Upon reviewing Graph 2, the Critic finds no logical contradictions but offers supplementary recommendations for optimization potential: • Introduce a node dedicated to symmetry arguments, potentially referencing group theory for theoretical depth; • Incorporate dimensionless scaling analysis; • Strengthen the association between v2v_2 (Context) and v4v_4 (Domain Constraint). Final Decision The Moderator deems these additional suggestions non-essential for the current task specification, noting they might divert the problem from its core objective. Therefore, the Moderator executes adaptive truncation to terminate the iteration. Graph 2 is output as the final Constraint Graph (∗G^*), which is then passed to the Executor for semantic instantiation into the finalized mathematical problem. Graph 0 v1v_1: Saddle Surface Graph 1 v1v_1. Saddle Surface: The abstract geometric primitive characterized by opposing curvatures. v2v_2. Potato Chip Context: Real-world physical object (e.g., Pringles) used to frame the problem. v3v_3. Hyperbolic Paraboloid Equation: The specific algebraic model, e.g., z=x2/a2−y2/b2z=x^2/a^2-y^2/b^2. v4v_4. Cylindrical Domain Constraint: The boundary condition restricting the surface to a finite disk, x2+y2≤R2x^2+y^2≤ R^2. v5v_5. Surface Area Integral: The double integral calculation required to find the total area of the curved surface. e21e_21. instantiates e13e_13. formalized_by e35e_35. integrand_source e45e_45. defines_limits Graph 2 v1v_1: Saddle Surface: The abstract geometric primitive characterized by opposing curvatures. v2v_2: Potato Chip Context: Real-world physical object (e.g., Pringles) used to frame the problem. v3v_3: Hyperbolic Paraboloid Equation: Simplified symmetric model: z=c(x2−y2)z=c(x^2-y^2) with c>0c>0 constant, or equivalently a=b=1a=b=1 in the form z=(x2−y2)/a2z=(x^2-y^2)/a^2. v4v_4: Cylindrical Domain Constraint: The boundary condition restricting the surface to a finite disk, x2+y2≤R2x^2+y^2≤ R^2. v5v_5: Surface Area Integral: The double integral calculation for the total area of the curved surface. Now includes symmetry exploitation: integrate over first quadrant and multiply by 4. v6v_6: Partial Derivative Calculation: Compute ∂z/∂x∂ z/∂ x and ∂z/∂y∂ z/∂ y for the surface area element dS=1+(∂z/∂x)2+(∂z/∂y)2dxdydS= 1+(∂ z/∂ x)^2+(∂ z/∂ y)^2\,dxdy. e21e_21: instantiates e13e_13: formalized_by e36e_36: differentiated_to e65e_65: integrand_source e45e_45: defines_limits Final Synthesized Problem A gourmet potato chip manufacturer designs its product to follow the precise geometry of a saddle surface, known mathematically as a hyperbolic paraboloid. When centered at the origin, the surface of a single chip is modeled by the equation z=c(x2−y2)z=c(x^2-y^2), where c is a positive constant. To ensure uniformity, each chip is trimmed so that its vertical projection onto the xyxy-plane is bounded by the circle x2+y2≤R2x^2+y^2≤ R^2. By first calculating the partial derivatives ∂z/∂x∂ z/∂ x and ∂z/∂y∂ z/∂ y to determine the surface area element dSdS, set up and evaluate a double integral to find the total surface area of the chip. In your calculation, exploit the symmetry of the surface by integrating over the first quadrant of the cylindrical domain and multiplying the result by four. Express your final answer in terms of c and R. Figure 8: Case Study. Prompt for Legislator - Proposer (PA_P) You are the Proposer (PA_P) in the Legislator-Executor framework. Your objective is to drive the meta-level structural evolution of a mathematical problem by optimizing a Constraint Graph =(,ℰ)G=(V,E). ### Input Data: - Style Tokens (S): STYLE_TOKENS_INPUT - Current Graph (tG_t): CURRENT_GRAPH_DATA - Feedback: ITERATIVE_GUIDANCE_FROM_MODERATOR ### Operational Directives: - Evolution & Revision: Perform topological mutations to transition tG_t to t+1G_t+1. - Graph-Style Alignment: Expand nodes V and logical edges ℰE to achieve the graph-related stylistic goals (particularly complexity) specified in S. - Consistency Maintenance: Rectify any logical contradictions identified in previous feedback. ### Task Workflow: Step 1: Internal Analysis & Planning Analyze the gap between the current graph tG_t and the target specifications in S. Plan specific mutations (e.g., adding concepts, nesting operators, or refining constraints) to bridge this gap while resolving any reported flaws. Step 2: Structured Output (JSON) Generate the updated graph t+1G_t+1 following the strict JSON schema below. ### Final Output Format: Analysis and Planning: [Your detailed step-by-step thinking process here] Final Optimized Graph (JSON): "graph_id": "G_t+1", "nodes": ["id": "v_n", "concept": "string", "description": "string"], "edges": ["source": "v_i", "target": "v_j", "relation": "string"], "mutation_log": "Summary of changes made in this iteration." (Constraint: Ensure all referenced nodes in ‘edges‘ exist in the ‘nodes‘ list) Figure 9: Prompt for the Proposer. Prompt for Legislator - Critic (CA_C) You are the Critic (CA_C). Your goal is not merely to check for correctness, but to identify the "evolutionary headroom" of the graph t+1G_t+1 to push it from functional to exceptional. ### Input for Review: - Style Tokens (S): STYLE_TOKENS_INPUT - Proposed Graph (t+1G_t+1): PROPOSED_GRAPH_DATA ### Evaluation Dimensions: - Internal Consistency: Scrutinize for logical contradictions or ill-defined constraints. - Specification Alignment: Verify if the graph strictly complies with the complexity and category requirements in S. - Optimization Potential: Even if requirements are met, provide several actionable suggestions for potential optimization. ### Final Output Format: - Analysis: [Your detailed step-by-step thinking process here] - Critical Flaws: [List any issues that violate consistency or S (Output "None" if perfect)] - Refinement Suggestions: [Propose at least 2-3 specific actions to further optimize the graph’s complexity or elegance] - Expected Utility: [Estimate the marginal gain of these optimizations (High / Medium / Low) to assist the Moderator’s decision] Figure 10: Prompt for the Critic. Prompt for Legislator - Moderator (MA_M) You are the Moderator (MA_M). You adjudicate the state of graph t+1G_t+1 based on the Critic’s report. ### Data for Decision: - Critic’s Report: CRITIC_REPORT - Style Tokens (S): STYLE_TOKENS_INPUT - Proposed Graph (t+1G_t+1): PROPOSED_GRAPH_DATA ### Decision Logic: - Adaptive Truncation: If t+1G_t+1 satisfies S and the potential for further gain is marginal, MA_M terminates the process and outputs the graph. - Iterative Guidance: Otherwise, direct specific modifications to the Proposer to extend structure or rectify flaws. ### Final Output Format: - Analysis: [Your detailed step-by-step thinking process here] - Decision: [Suspend/Continue Iteration] - Guidance for the Proposer: [If ITERATE: Provide a concise instruction list for the Proposer. If TERMINATE: Output "None"] - Final Graph: [If TERMINATE: Output the full JSON of t+1G_t+1. If ITERATE: Output "N/A"] Figure 11: Prompt for the Moderator Prompt for Executor - Question Synthesizer Your task is to perform Semantic Instantiation: converting an abstract Constraint Graph ∗G^* into a high-quality, natural language mathematical problem. ### Input Data: - Style Tokens (S): STYLE_TOKENS_INPUT - Final Constraint Graph (∗G^*): FINAL_GRAPH_DATA ### Operational Directives: - Structural Fidelity: Every node v∈v and edge e∈ℰe must be reflected in the problem. Do not omit constraints. - Style Alignment: The generated mathematical problem should conform to the constraints specified in the style tokens. - Semantic Fluency: The problem must be linguistically fluid, not a robotic list of conditions. Ensure logical transitions between the situational narrative and the technical specifications. - Output Constraint: Generate ONLY the natural language question (Q). Do not provide solutions, explanations, or meta-comments. ### Final Output Format: - Analysis: [Step-by-step plan: How to map ∗G^* nodes to S context while maintaining fluency] - Question: [The finalized natural language problem statement] Figure 12: Prompt for the Question Synthesizer. Prompt for Generating Quality of Problems # Instruction You need to rate the quality of the math problem based on its clarity, accuracy, and logical coherence. The rating scale is as follows: – Very poor: The problem description is ambiguous, conditions are incomplete, or contains logical contradictions. It lacks essential information and context required for solving, or the given instruction is not a mathematical problem. – Poor: The problem is somewhat unclear or lacks important details. It requires significant clarification to define the solving requirements. – Average: The problem is moderately clear and accurate but may contain imprecise expressions. Additional information might be needed for a complete solution. – Good: The problem is clearly structured, with well-defined conditions and logical coherence. It provides sufficient information to support the solving process. – Excellent: The problem is precisely formulated, with complete conditions and rigorous logic. It contains all necessary elements for solving without redundant information. ## Math Problem to Evaluate math_problem ## Output Format First, provide an assessment highlighting the strengths and/or weaknesses of the math problem. Then, output a rating by filling in the placeholders: "explanation": "[Your assessment analysis]", "quality": "[very poor / poor / average / good / excellent]". Prompt for Generating Difficulty of Problems # Instruction You are an expert in mathematics education and cognitive task analysis. Your responsibility is to evaluate the complexity of mathematical problems presented by users. For each mathematical problem, you must first identify the required knowledge points, and then assess the difficulty level based on the mathematical concepts involved, problem-solving steps, and cognitive demands. ## Math Problem to Evaluate math_problem ## Output Format Given the provided mathematical problem, in your output you must first determine the knowledge points required to solve it. Then, rate the difficulty level of the mathematical problem as ’very easy’, ’easy’, ’medium’, ’hard’, or ’very hard’. Please output the difficulty level below in the following format by filling in the placeholders in […]: "explanation": "[Your detailed explanation and reasoning]", "knowledge": "[list specific mathematical concepts, procedures, or knowledge domains]", "difficulty": "[very easy / easy / medium / hard / very hard]". Figure 13: Prompts for Evaluating Mathematical Problem Quality and Difficulty.