Paper deep dive
Grounded Chess Reasoning in Language Models via Master Distillation
Zhenwei Tang, Qianfeng Wen, Seth Grief-Albert, Yahya Elgabra, Blair Yang, Honghua Dong, Ashton Anderson
Intelligence
Status: succeeded | Model: anthropic/claude-sonnet-4.6 | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/24/2026, 3:24:40 AM
Summary
This paper introduces Master Distillation, a framework for distilling expert system reasoning (e.g., Stockfish chess engine) into natural language chain-of-thought explanations for compact language models. The approach combines Feigned Discovery Prompting, supervised fine-tuning (SFT), and reinforcement learning with verifiable rewards (RLVR) with theme-balanced data sampling. The resulting 4B parameter model C1 achieves 48.1% accuracy on chess puzzle solving, outperforming all open-source models and most frontier proprietary systems, and surpassing its distillation teacher Gemini-3-Flash (40.8%), while generating solutions in two orders of magnitude fewer tokens than baselines.
Entities (41)
Relation Signals (22)
C1 → achievesaccuracy → 48.1% on chess puzzle solving
confidence 99% · Our 4B parameter model, C1, advances from a near-zero baseline to 48.1% accuracy
Master Distillation → appliedto → chess puzzle solving
confidence 99% · We demonstrate this approach in chess, a canonical reasoning domain where language models continue to underperform
C1 → surpasses → Gemini 3 Flash
confidence 99% · C1-4B surpasses its distillation teacher Gemini-3-Flash (40.8%)
C1 → trainedwith → Master Distillation
confidence 99% · We demonstrate this approach in chess... Our pipeline combines supervised fine-tuning and reinforcement learning with theme-balanced data sampling
Zhenwei Tang → affiliatedwith → University of Toronto
confidence 98% · University of Toronto... Correspondence to: Zhenwei Tang <joseph-tang@cs.toronto.edu>
C1 → finetunedfrom → Qwen3-4B-Instruct-2507
confidence 98% · We perform full fine-tuning on Qwen3-4B-Instruct-2507
C1 → outperforms → all open-source models
confidence 98% · outperforming all open-source models and most frontier proprietary systems
Master Distillation → uses → RLVR
confidence 98% · Master Distillation demonstrates how to inject expert-level knowledge into compact models... offering a recipe for unlocking RLVR
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Language models often lack grounded reasoning capabilities in specialized domains where training data is scarce but bespoke systems excel. We introduce a general framework for distilling expert system reasoning into natural language chain-of-thought explanations, enabling compact models to acquire domain expertise and the ability to generate faithful, grounded explanations. Rather than distilling only final outputs, we capture the full reasoning process, transforming opaque expert computations into transparent, step-by-step explanations. We demonstrate this approach in chess, a canonical reasoning domain where language models continue to underperform. Our 4B parameter model, C1, advances from a near-zero baseline to 48.1% accuracy, outperforming all open-source models and most frontier proprietary systems. Notably, C1 surpasses its distillation teacher and generates solutions in two orders of magnitude fewer tokens than baselines. Unlike prior neural chess approaches that predict only best moves, C1 generates explainable solutions revealing strategic reasoning. Our pipeline combines supervised fine-tuning and reinforcement learning with theme-balanced data sampling for comprehensive tactical coverage. Master Distillation demonstrates how to inject expert-level knowledge into compact models for under-optimized domains, offering a recipe for unlocking RLVR where LLMs lack sufficient base capabilities.
Tags
Links
- Source: https://arxiv.org/abs/2603.20510v1
- Canonical: https://arxiv.org/abs/2603.20510v1
Trouble viewing inline? Open PDF directly →
Full Text
57,291 characters extracted from source content.
Expand or collapse full text
Grounded Chess Reasoning in Language Models via Master Distillation Zhenwei Tang 1 Qianfeng Wen 1 Seth Grief-Albert 2 Yahya Elgabra 1 Blair Yang 3 Honghua Dong 1 Ashton Anderson 1 Abstract Language models often lack grounded reasoning capabilities in specialized domains where training data is scarce but bespoke systems excel. We in- troduce a general framework for distilling expert system reasoning into natural language chain-of- thought explanations, enabling compact models to acquire domain expertise and the ability to gen- erate faithful, grounded explanations. Rather than distilling only final outputs, we capture the full reasoning process, transforming opaque expert computations into transparent, step-by-step expla- nations. We demonstrate this approach in chess, a canonical reasoning domain where language models continue to underperform. Our 4B param- eter model, C1, advances from a near-zero base- line to 48.1% accuracy, outperforming all open- source models and most frontier proprietary sys- tems. Notably, C1 surpasses its distillation teacher and generates solutions in two orders of magni- tude fewer tokens than baselines. Unlike prior neural chess approaches that predict only best moves, C1 generates explainable solutions reveal- ing strategic reasoning. Our pipeline combines supervised fine-tuning and reinforcement learn- ing with theme-balanced data sampling for com- prehensive tactical coverage. Master Distillation demonstrates how to inject expert-level knowl- edge into compact models for under-optimized domains, offering a recipe for unlocking RLVR where LLMs lack sufficient base capabilities. 1. Introduction Large language models exhibit jagged intelligence: super- human performance in some domains while remaining sur- prisingly clumsy in others. While the reasons are multi- faceted, weak performance often emerges in niche domains 1 University of Toronto 2 Queen’s University 3 Coolwei AI Lab.Correspondence to:Zhenwei Tang<joseph- tang@cs.toronto.edu>. Preprint. March 24, 2026. where high-quality reasoning data is sparse in pretraining corpora, even when strong specialized systems exist. Chess exemplifies this phenomenon: despite the existence of su- perhuman engines like Stockfish (Stockfish Team, 2023), state-of-the-art LLMs struggle with basic tactical puzzles, and recent attempts to apply reinforcement learning with verifiable rewards (RLVR) have failed to yield meaningful improvements (Hwang et al., 2025a). Prior work has explored distilling chess engines into neural networks, but these approaches transfer only move predic- tions without the underlying reasoning process (Ruoss et al., 2024). We present Master Distillation, a data synthesis approach that distills both outcome and process from spe- cialized systems into language models. The core insight is that expert systems know the correct answer but cannot articulate their reasoning, while LLMs can generate fluent explanations but lack domain knowledge. Master Distilla- tion combines these complementary strengths: we leverage a master system to provide ground-truth solutions and prompt a frontier LLM to verbalize the reasoning process through Feigned Discovery Prompting, generating chain-of-thought traces that are both verifiably correct and pedagogically natural. We focus on chess puzzle solving as our testbed. Chess offers perfectly verifiable rewards through engine evalua- tion, and puzzles have an unambiguous ground truth with a single best move, making them ideal for both evaluation and reward design. The domain is rich with tactical concepts spanning multiple difficulty levels, enabling systematic anal- ysis of model capabilities. Most importantly, the failure of prior RLVR attempts on chess (Hwang et al., 2025a) makes it a compelling challenge: if Master Distillation can un- lock RLVR here, the approach likely generalizes to other domains where LLMs currently fall short while master sys- tems exist. Our experiments demonstrate that Master Distillation en- ables effective SFT+RLVR training for chess reasoning. Our C1-4B model achieves 48.1% accuracy, outperform- ing all open-source models and most proprietary models. Notably, C1-4B surpasses its distillation teacher Gemini-3- Flash (40.8%), demonstrating that the student can exceed teacher performance. Beyond accuracy, our models exhibit 1 arXiv:2603.20510v1 [cs.AI] 20 Mar 2026 Grounded Chess Reasoning in Language Models via Master Distillation Figure 1. C1 achieves competitive performance against frontier models orders of magnitude larger on chess puzzle-solving. remarkable token efficiency, generating solutions in roughly two orders of magnitude fewer tokens than both proprietary and open-source baselines, reflecting how humans actually solve chess puzzles through focused pattern recognition rather than extensive deliberation. In summary, we show that Master Distillation unlocks RLVR for domains where LLMs lack sufficient base capabilities, offering a recipe for improving language model performance in specialized domains with strong bespoke systems. We open-source our code and data 1 to foster further research. 2. Related Work Reinforcement Learning with Verifiable Rewards. RLVR enhances language model reasoning by replacing learned reward models with deterministic verification (Guo et al., 2025; Hwang et al., 2025b; Shao et al., 2024; Yu et al., 2025), providing objective signals that scale efficiently un- like preference-based RLHF (Ouyang et al., 2022). The success of recent reasoning models, including OpenAI o1 (OpenAI, 2024), DeepSeek-R1 (Guo et al., 2025), and Qwen3 (Yang et al., 2025a), demonstrates the effective- ness of RLVR. These methods build on policy gradient foundations like PPO (Schulman et al., 2017), with GRPO (Shao et al., 2024) marking a critical pivot by eliminat- ing critic networks via grouped rollout baselines. Subse- quent work addresses GRPO’s limitations: DAPO (Yu et al., 1 https://github.com/CSSLab/C1 2025) introduces clip-higher bounds and dynamic sampling, GSPO (Zheng et al., 2025) stabilizes long sequences via sequence-level ratios, Dr. GRPO (Liu et al., 2025) corrects length biases, PRIME (Cui et al., 2025) enables implicit pro- cess rewards, and RLOO (Ahmadian et al., 2024) revisits REINFORCE-style optimization. RLVR has not yet been shown to be effective in chess (Hwang et al., 2025b). Chess Move Modeling. Traditional chess AI focuses on finding optimal moves through search and evaluation. Al- phaZero (Silver et al., 2016; 2017a;b) and Leela Chess Zero (Pascutto & Linscott, 2018; Pascutto et al., 2018) com- bine Monte Carlo tree search with neural network evaluation, achieving superhuman performance. Stockfish (Stockfish Team, 2023) integrates Efficiently Updateable Neural Net- works (NNUE) (Nasu, 2018) for fast position evaluation. Recent work explores replacing search with pure neural pre- diction: Ruoss et al. (2024) distill search-based policies into transformers, and Monroe & Chalmers (2024) train trans- formers directly on grandmaster games. These approaches focus on move prediction but do not produce human- readable explanations of their reasoning. A parallel line of research models human chess behavior rather than optimal play. Maia (McIlroy-Young et al., 2020) predicts population- level human moves, extended by Maia-2 (Tang et al., 2024) and ALLIE (Zhang et al., 2025b). Individual-level modeling has been explored in Maia-Individual (McIlroy-Young et al., 2022) and Maia4All (Tang et al., 2026). While these models capture human decision-making patterns, they similarly lack 2 Grounded Chess Reasoning in Language Models via Master Distillation explicit reasoning traces. Chess LLMs. Several benchmarks evaluate LLM chess capabilities across different dimensions: MATE (Wang et al., 2025), PGN2FEN (Cooper, 2025), LLM Chess Puz- zles (Kagi Search, 2025), LLM Chess Leaderboard (Saplin, 2025), Chess LLM Arena (Jannala, 2025), Kaggle Game Arena (Lee et al., 2025), and ChessQA (Wen et al., 2025). Kim et al. (2025) explores concept-guided chess commen- tary generation. These benchmarks reveal that current LLMs struggle with chess reasoning despite strong per- formance in other domains. Early efforts in training LLMs for chess include commentary generation (Jhamtani et al., 2018; Zang et al., 2019), ChessGPT (Feng et al., 2023), and self-supervised approaches using GPT-2 (Noever et al., 2020). St ̈ ockl (2021) analyze language model learning dy- namics on chess, while ChePT (Swingle et al., 2021) com- bines move prediction with commentary generation. Zhang et al. (2025a) train on complete game transcripts to improve move prediction. However, none of these approaches explic- itly address distilling reasoning from master systems like chess engines into language models for explainable puzzle solving, which is the focus of our work. Most recently, Chess-R1 (Hwang et al., 2025a) attempted to apply RLVR directly to chess but found that neither direct RLVR nor SFT followed by RLVR yields meaningful improvements. Our work addresses this challenge through Master Distil- lation, demonstrating that the SFT+RLVR pipeline can be made effective for chess reasoning with a properly designed distillation method. 3. Methodology 3.1. Chess Puzzle Solving with RLVR Chess Puzzles. Chess puzzles are carefully curated tactical positions with a single objectively best move sequence, de- signed to test concrete calculation and pattern recognition across various tactical themes such as forks, pins, and mat- ing attacks. Compared to arbitrary game positions where multiple reasonable moves may exist, puzzles offer unam- biguous ground truths, making them particularly suitable for evaluation and reward design in our setting. Chess Puzzle Solving as a Distillation Task. We formu- late chess puzzle solving as predicting the correct first move. While this may appear simple, finding the optimal first move requires the model to reason about future positions, antici- pate opponent responses, and evaluate resulting outcomes before committing to a decision. This lookahead reasoning is precisely what distinguishes tactical calculation from su- perficial pattern matching. Prior work on distilling chess engines into neural networks (Ruoss et al., 2024; Nasu, 2018) focuses on transferring the engine’s move predictions or position evaluations, but not the underlying reasoning process. In contrast, we focus on distilling both the outcome and the process: the model generates a chain-of-thought that explores candidate moves and evaluates their consequences before outputting a final answer. The model is evaluated on first-move accuracy, which serves as a proxy for the quality of its internal reasoning learned from distillation. ¬ Example Puzzle Prompt 8 0Zks0Z0s 7 opo0ZNop 6 0ZnapZ0Z 5 Z0Z0Z0M0 4 0ZPO0ZnZ 3 Z0Z0A0Zq 2 PO0ZBO0Z 1 S0ZQZRJ0 a b c d e f g h Black To Move. Board shown for reference only, not included in prompt. You are given a chess position in FEN: 2kr3r/p2Npp/2nbp3/6N1/2P2n1/4B2q/ P2BP2/R2Q1RK1 b - - 2 15. Find the best move for the side to play. PIECEARRANGEMENT. LEGALMOVES. Analyze step by step and explain your reasoning. Finish with a single line formatted EXACTLY as: FINAL ANSWER: <answer>. Use UCI notation (e.g., e2e4, c2b1q) for the final answer. Ground Truth h3h2 Established Post Training Pipeline. Reinforcement learn- ing with verifiable rewards (RLVR), with optional super- vised fine-tuning (SFT) as a cold-start phase, has emerged as the dominant paradigm for enhancing LLM reasoning ca- pabilities. This pipeline has proven effective across diverse domains: in mathematics, coding, deep research, and tool use (Guo et al., 2025; Yang et al., 2025b; Jin et al., 2025; Li et al., 2025). The success of RLVR relies on a critical prerequisite: the base model must already possess sufficient capability to occasionally produce correct solutions, which RL can then amplify through reward-driven optimization. Chess, however, presents a unique challenge. Despite being a domain with perfectly verifiable rewards, neither direct RLVR nor SFT followed by RLVR yields meaningful im- 3 Grounded Chess Reasoning in Language Models via Master Distillation provements (Hwang et al., 2025a). This is because LLMs’ base chess capabilities remain too weak to bootstrap from, stemming from the relative scarcity of high-quality chess reasoning data in pretraining corpora. Unlike mathematics, where step-by-step solutions are abundant online, chess anal- ysis rarely includes the detailed chain-of-thought reasoning needed for effective SFT. 3.2. Master Distillation We present a data synthesis approach that combines the domain expertise of specialized systems, e.g., the master, with the linguistic capabilities of frontier LLMs. The core insight is that bespoke systems like Stockfish (Stockfish Team, 2023) know the correct answer but cannot articulate their symbolic reasoning, while LLMs can generate fluent explanations but lack the domain knowledge to reliably solve hard problems. This motivates our Master Distilla- tion approach: by leveraging a specialized master system to provide ground-truth solutions and a frontier LLM to verbalize the symbolic reasoning process, we can generate high-quality training data that bridges this gap and enables effective RLVR for chess reasoning. 3.2.1. CONTEXT ENGINEERING The first dimension that requires careful design towards realizing is context engineering: we must provide the LLM with sufficient information to understand the position and generate grounded analysis. Master Solution. The principal variation (PV) is the se- quence of optimal moves for both sides leading to the puz- zle’s objective, which may be checkmate, material gain, or achieving a decisive advantage. The PV represents the sym- bolic reasoning process of the master system, and is the core knowledge we aim to distill from the symbolic master into the language model student. We leverage Stockfish (Stock- fish Team, 2023) at depth 24, which is more than capable of solving most puzzles accurately, as the master system to compute the PV for each puzzle position. We choose not to use transformer-based chess models (Pascutto & Linscott, 2018; Ruoss et al., 2024) because they do not explicitly provide PVs or allow trivial extraction of move sequences. We consider two settings: multi-PV, which generates the top-kcandidate move sequences and provides additional context about alternative move options, and best-move PV, which provides only the single optimal line. While multi- PV offers richer information, best-move PV introduces less noise, allowing the model to focus on what is correct. This is particularly suitable for puzzles where there is one ab- solute best move and alternatives are significantly worse. We provide empirical comparison between these settings in Table 2. Board and Move Representation. Chess positions are en- coded in Forsyth-Edwards Notation (FEN), a compact string representation that captures piece placement, active color, castling rights, en passant targets, and move clocks (Ed- wards, 1994). Moves are represented in Universal Chess Interface (UCI) notation, which specifies moves by their source and destination squares (e.g., e2e4). While Stan- dard Algebraic Notation (SAN) is more human-readable and prevalent in online chess discussions and literature, it introduces ambiguities since the same symbol can refer to different pieces (e.g., B for bishop versus the b-file) and includes special characters like + and # that complicate eval- uation and reward computation through inconsistent string matching. Prior work (Tang et al.) has shown that LLMs struggle to recognize FEN strings due to tokenization issues, conflating chess reasoning ability with position recognition. To separate these concerns and provide complete positional context, we augment the FEN with an explicit piece list describing the location of each piece and a list of all legal moves in the current position, for both the puzzle-solving task and distillation. An example chess puzzle is shown in the blue box titled “Example Puzzle”. Additional Hints. Puzzles may be associated with ad- ditional metadata, as provided in public databases like Lichess 2 . We provide the teacher LLM with additional context to guide its analysis. The opponent’s last move pro- vides crucial context for understanding the tactical situation, as many puzzles arise from the opponent’s mistake or create threats that must be addressed. The tactical themes (e.g., fork, pin, back-rank mate) hint at the underlying pattern the LLM should explain, helping it focus on the relevant tactical motifs rather than generic analysis. The puzzle rat- ing signals the expected complexity, allowing the LLM to calibrate the depth and length of its explanation accordingly. Crucially, while all this information guides the teacher’s generation, the student model never sees these hints during training or evaluation. 3.2.2. FEIGNED DISCOVERY PROMPTING Our goal is to produce chain-of-thought traces that are both verifiably correct, grounded in the master’s solution, and pedagogically natural, reading like genuine reasoning rather than post-hoc rationalization. Naively conditioning on the answer risks generating justifications for a known conclu- sion rather than authentic problem-solving processes. We want the student model to learn how to arrive at solutions, not merely how to justify them. This creates a fundamental tension between faithfulness to the master trace and authen- tic reasoning. If the generated reasoning is too faithful, the LLM simply parrots the moves without producing transfer- able reasoning patterns. If it is too exploratory, the LLM 2 https://database.lichess.org/#puzzles 4 Grounded Chess Reasoning in Language Models via Master Distillation may hallucinate and lose the grounding benefit of the master solution. Our key insight is to instruct the teacher LLM to reason as if the solution is unknown, while secretly steering toward the master trace. The teacher must pretend not to know the answer, simulating the discovery process a human solver would experience. This feigned ignorance is criti- cal: it forces the LLM to generate reasoning that explains why a move is correct through genuine analysis, rather than simply asserting correctness because the answer was pro- vided. The master’s PV and additional context acts as a soft constraint, ensuring the reasoning naturally converges to the correct solution while maintaining the appearance of authentic exploration. We impose several additional constraints in the prompt to en- sure the generated reasoning traces are concise and pedagog- ically useful. First, we enforce length constraints of 4–10 sentences scaled to puzzle complexity, as chess learners ben- efit from focused analysis rather than verbose explanations. Second, we require the teacher to bridge chess concepts with notation by annotating moves with natural language clarifications (e.g., “Qxh7+ (queen takes h7 with check)”) and grounding analysis in explicit board coordinates (e.g., “the queen on h5”, “king on g1”) while explaining tacti- cal relationships such as attacks, pins, and defender counts. This bridging is important because student LLMs have lim- ited exposure to chess-specific notation during pretraining. Third, we enforce stylistic constraints: an objective voice without phrases like “I see” or “I notice”, and strict prohibi- tion against mentioning engine scores, ratings, themes, or any indication that the solution was provided. In essence, we are role-playing the LLM as a chess grandmaster who an- alyzes positions naturally and arrives at the solution through genuine reasoning. These constraints collectively ensure the generated traces read as natural expert analysis rather than machine-generated explanations. An example of the distillation prompt is shown in the blue box titled “Master Distillation Prompt” in the Appendix. 3.3. C1 Training Data Balancing. We apply theme-balanced sampling to construct both SFT and RLVR training data, as described in Algorithm 1. Since each puzzle is annotated with mul- tiple tactical themes, we prioritize balancing by the rarest themes, ensuring that infrequent concepts such as underPro- motion and balestraMate receive sufficient representation. By sampling puzzles that contain these rare themes, more frequent themes like fork and pin are naturally covered as co-occurring labels. We targetM=800 samples for each of theK=50 rarest themes, resulting in approximately 40,000 samples for both SFT and RLVR with non-overlapping puz- zle sets. Detailed statistics are provided in Table 9 and per-theme distributions in Table 10. Algorithm 1 Theme-Balanced Data Sampling Require: DatasetD where each puzzle p has theme setT (p) Require: Number of rare themes to balance K Require: Maximum samples per theme M Ensure: Balanced subsetD bal Compute theme frequencies:f(t) ← |p ∈ D : t ∈ T (p)| for all themes t Select rare themes: T rare ← arg min K f(t) Initialize selected IDs: S ←∅ Initialize output: D bal ←∅ for each theme t∈T rare do C t ←p∈D : t∈T (p)∧ id(p) /∈S Sample min(M,|C t |) puzzles fromC t without replacement D bal ←D bal ∪ sampled puzzles S ←S∪id(p) : p∈ sampled puzzles end for returnD bal Supervised Fine-Tuning. We perform full fine-tuning on Qwen3-4B-Instruct-2507, which achieves near-zero accu- racy on our evaluation out of the box. SFT is a necessary prerequisite for RLVR: the base model must solve puzzles correctly some of the time, otherwise all rollouts receive zero reward and no learning signal is available. We opt for full fine-tuning over parameter-efficient methods like LoRA (Hu et al., 2022) because our task requires knowl- edge acquisition in a novel domain rather than adaptation of existing capabilities, which benefits from updating all model parameters. Reinforcement Learning with Verifiable Rewards. We build upon DAPO (Yu et al., 2025), which introduces im- provements over GRPO (Shao et al., 2024): Overlong Reward Shaping, Token-Level Policy Gradient Loss, Dy- namic Sampling, Clip-Higher, and Removing KL Diver- gence. However, we exclude two of these modifications: Overlong Reward Shaping and Removing KL Divergence. The original DAPO removes KL divergence under the as- sumption that long-CoT reasoning benefits from allowing the model distribution to diverge significantly from the ini- tial policy. In our setting, the SFT data already consists of concise, to-the-point reasoning traces, and we aim to preserve this brevity rather than encourage free exploration. Similarly, Overlong Reward Shaping is unnecessary since our responses are inherently short. We denote this modi- fied algorithm as DAPO-C1. We use a binary reward based solely on whether the final move matches the Stockfish- verified optimal move. Following the Bitter Lesson (Sutton, 2019), we avoid complex reward shaping that may introduce biases or enable reward hacking, instead relying on a clean signal that scales with data and compute. 5 Grounded Chess Reasoning in Language Models via Master Distillation Table 1. Performance comparison across difficulty levels and models. Avg Acc represents average accuracy (%), and Avg Tokens indicates average generation length. BeginnerIntermediateAdvancedExpertTheme-SplitAvg AccAvg Tokens Proprietary models gpt-595.084.054.031.085.276.712,193 gemini-3-pro88.086.070.044.078.275.43,182 gemini-3-flash65.059.034.019.038.040.86,418 gpt-5-chat 52.039.027.018.041.838.3925 gemini-2.5-pro37.031.029.019.031.030.19,668 claude-sonnet-4.532.029.015.011.028.625.63,227 claude-sonnet-4 35.019.016.010.026.823.88,028 claude-haiku-4.533.024.014.011.025.623.38,111 gemini-2.5-flash9.04.06.05.08.27.29,991 Open-source models deepseek-chat-v3.127.021.06.016.022.020.011,249 qwen3-next-80b-a3b 24.014.014.08.017.616.413,938 deepseek-r1-0528 11.010.014.016.016.014.614,442 qwen3-max 22.015.03.016.013.813.93,393 llama-4-maverick 12.08.05.010.08.68.71,092 mistral-medium-3.19.06.07.04.08.07.32,818 llama-4-scout0.00.00.01.00.40.3806 gemma-3-27b0.00.00.00.00.00.0705 Ours C1-SFT-4B51.030.030.026.046.240.9188 C1-SFT-8B57.036.027.027.046.642.2189 C1-4B65.039.039.022.053.648.1178 4. Experiments 4.1. Experimental Settings We implement SFT using LLaMA-Factory (Zheng et al., 2024) and RLVR using VERL (Sheng et al., 2024). All ex- periments are conducted on 4×H100 GPUs. For both SFT and RLVR, we use a held-out validation set of chess puzzles for early stopping to prevent overfitting. Detailed hyperpa- rameters for SFT and RLVR are provided in Tables 7 and 8 in the Appendix, respectively. We follow the evaluation protocol of Wen et al. (2025). The test set contains 900 puz- zles: 500 puzzles sampled across 20 tactical themes (25 per theme) and 400 puzzles sampled across 4 difficulty levels (100 per level: Beginner, Intermediate, Advanced, Expert). We report pass@1 accuracy as the primary metric. For all baseline models, we enable reasoning mode where appli- cable. We use OpenRouter for API access to proprietary models during evaluation. We use Gemini-3-Flash-Preview via OpenRouter as the teacher model for Master Distillation. 4.2. Performance Comparison Our C1-4B model achieves the highest accuracy among all open-source models by a substantial margin and outper- forms most frontier proprietary models, demonstrating the effectiveness of Master Distillation for domain-specific rea- soning. Notably, all C1 models surpass Gemini-3-Flash, the model used to generate these reasoning traces, indicating that combining symbolic expertise with language model reasoning enables the student to exceed its teacher. This im- provement is consistent across difficulty levels, with partic- ularly strong gains on Intermediate puzzles. On the Theme- Split evaluation, C1-4B shows even more pronounced advan- tages, with detailed per-theme results provided in Appendix Table 11. Beyond accuracy, our models exhibit remarkable token effi- ciency, generating solutions in roughly two orders of mag- nitude fewer tokens than both proprietary and open-source baselines. This reflects how humans actually solve chess puzzles: through focused pattern recognition and precise calculation rather than extensive deliberation. It also aligns with the practical needs of chess learners, who benefit from concise, to-the-point explanations rather than verbose rea- 6 Grounded Chess Reasoning in Language Models via Master Distillation soning chains. Comparing C1-4B (SFT+RLVR) with C1-SFT-4B reveals the effect of reinforcement learning. RLVR provides consis- tent improvements on Beginner and Intermediate puzzles, but shows diminished or negative gains on Expert-level problems. This pattern suggests that the chain-of-thought reasoning traces may not faithfully capture the underlying logic for harder puzzles, limiting what RLVR can learn from outcome-based rewards. Unlike mathematical reasoning, where RLVR often yields dramatic improvements, the less reliable reasoning chains in complex chess puzzles constrain the potential gains from reinforcement learning. 4.3. Data Ablations Table 2. Ablation study on SFT data configurations. ScaleDistributionQualityContextSFT 8krandomflashfull19.3 8khardflashfull16.2 8kbalancedprofull22.8 8kbalancedflashfull20.1 16kbalancedflashfull29.7 8kbalancedflashMulti PVs17.6 8kbalancedflashw/o Theme 17.3 8kbalancedflashw/o Feigned 16.3 39kbalancedflashfull40.9 SFT. Table 2 presents ablation studies on SFT data config- urations, revealing several key insights. First, distillation quality matters: using Gemini-3-Pro as the teacher out- performs Gemini-3-Flash, suggesting that higher-quality chain-of-thought traces lead to better student performance. However, due to budget constraints, we use Flash for scaling experiments. Second, data distribution significantly impacts results: hard sampling yields the worst performance because chain-of-thought quality degrades on difficult puzzles, while random sampling underperforms balanced sampling due to insufficient coverage of rare tactical concepts. Third, our context engineering and prompting design choices are vali- dated: Multi-PV introduces noise and performs worse than best-move PV, removing theme hints causes the reasoning to lose focus, and removing our Feigned Discovery Prompt (w/o Feigned) leads to the largest degradation, demonstrat- ing that our prompting strategy is critical for generating effective training data. Fourth, data scale matters substan- tially, with performance improving from 19.3% at 8k to 29.7% at 16k and 40.9% at 39k samples. These results col- lectively justify our final configuration (highlighted in gray): 39k balanced samples distilled from Gemini-3-Flash with full context and Feigned Discovery Prompting. Table 3. Ablation study on RLVR Data. DistributionReuseSFTRLVRImpr randomfalse40.947.1+6.2 hardfalse40.943.0+1.1 balancedtrue 40.946.7+5.8 balancedfalse40.948.1+7.2 RLVR. Table 3 presents ablations on RLVR data selection strategies. Random sampling draws problems i.i.d. from the training distribution, while hard sampling prioritizes the most difficult puzzles. Balanced sampling ensures equal representation across tactical themes for broader concept coverage. The results show that balanced sampling achieves the best performance, suggesting that diverse tactical expo- sure benefits reinforcement learning more than focusing on difficulty alone. Notably, hard sampling yields the small- est improvement, consistent with our earlier observation that unfaithful reasoning traces on difficult puzzles limit RLVR’s effectiveness. We also examine whether to reuse SFT samples during RLVR. As shown in 3, reusing SFT samples slightly hurts performance compared to using fresh problems. This is likely because training on previously seen examples encourages the model to reinforce existing behav- ior rather than acquire new knowledge, shifting the learning objective toward consistency with prior outputs rather than exploration of improved reasoning strategies. 4.4. Algorithm Ablations Table 4. Ablation study on SFT configurations. ConfigurationSFT LoRA (r = 16,α = 32)32.1 LoRA (r = 64,α = 128)38.8 Full Fine-Tuning40.9 SFT. Table 5 validates our choice of full fine-tuning over parameter-efficient methods. While increasing LoRA rank from 16 to 64 improves performance from 32.1% to 38.8%, full fine-tuning achieves the best result at 40.9%. This confirms our hypothesis that chess puzzle solving requires knowledge acquisition in a novel domain rather than adap- tation of existing capabilities, benefiting from updating all model parameters rather than low-rank approximations. RLVR. Table 5 presents ablation studies on RLVR configu- rations. DAPO outperforms GRPO, validating the benefits of Token-Level Policy Gradient Loss, Dynamic Sampling, and Clip-Higher for our setting. Our modified DAPO-C1, which retains KL divergence and removes Overlong Reward Shaping to preserve concise reasoning, achieves the best 7 Grounded Chess Reasoning in Language Models via Master Distillation Table 5. Ablation study on RLVR configurations. AlgorithmRewardSFTRLVRImpr GRPOCorrectness40.945.5+4.6 DAPOCorrectness40.947.3+6.4 DAPO-C1Partial (η=0.5)40.939.3-1.6 DAPO-C1Partial (η=0.2) 40.941.6+0.7 DAPO-C1Correctness40.948.1+7.2 Table 6. Ablation study on base models. The 8B model was not trained with RLVR due to computational constraints. ModelSFTRLVR Qwen3-8B42.2- Qwen3-4B40.547.6 Qwen3-4B-Instruct-250740.948.1 Gemma3-4B36.241.0 performance with a 7.2% improvement over SFT. We also experiment with partial credit rewards, where moves with correct source or destination squares receive a fractional rewardη. However, partial rewards withη=0.5 hurts perfor- mance, andη=0.2 yields only marginal gains. This supports our design choice of using binary correctness rewards: fol- lowing the Bitter Lesson (Sutton, 2019), simple and clean reward signals scale better than hand-crafted shaping, which may introduce biases or enable reward hacking. Our final configuration (highlighted in gray) uses DAPO-C1 with binary correctness rewards. 4.5. Model Ablations Table 6 compares different base models after SFT and RLVR training. Qwen3-4B-Instruct-2507 achieves the best overall performance, slightly outperforming the reasoning variant Qwen3-4B, possibly because reasoning models are more rigid in their established reasoning patterns and thus harder to adapt to domain-specific chain-of-thought styles. Gemma3-4B lags behind the Qwen models, reflecting differ- ences in pretraining data and architecture. Notably, the 8B model shows only marginal gains over the 4B model after SFT (42.2% vs 40.9%), indicating that data quality is a more significant bottleneck than model capacity at this scale. This suggests that simply scaling parameters without improving training data yields diminishing returns for domain-specific reasoning tasks. While scaling to larger models would likely yield further improvements, we focus on the 4B model due to computational constraints and leave scaling experiments to future work. 5. Discussion Towards Better Performance. Our ablation studies re- veal several key factors for effective chess reasoning: high- quality chain-of-thought traces from capable teacher models, balanced data distribution for broad tactical coverage, care- ful context engineering and Feigned Discovery Prompting, and sufficient data scale. Each of these dimensions offers opportunities for improvement. The most straightforward path to better performance is scaling. Our experiments show consistent gains from 8k to 16k to 39k training samples, with no sign of saturation. Similarly, using Gemini-3-Pro in- stead of Flash as the teacher improves performance, suggest- ing that even stronger teachers would yield higher-quality distillation data. Notably, our distilled C1 models surpass their teacher Gemini-3-Flash, demonstrating that Master Distillation combined with RLVR can produce students that exceed teacher performance. Due to resource constraints, we focused on 4B parameter models, but our methodology naturally extends to larger architectures where we would expect further gains. Chess Foundational Model. While C1 demonstrates strong performance on puzzle solving, a true chess foundation model would generalize across diverse tasks and position types. Our current focus on puzzles, which have unambigu- ous ground truth, represents one end of a spectrum. Strategic positions, where multiple reasonable continuations exist and success depends on positional judgment rather than forced calculation, present a complementary challenge. Extending Master Distillation to such positions would require rethink- ing reward design, as evaluation becomes less binary and more nuanced. Beyond move prediction, a chess foundation model could support multiple tasks: game annotation, open- ing preparation, endgame analysis, mistake identification, and personalized training recommendations. Mid-training on large-scale chess corpora, combined with multimodal ca- pabilities to process board images directly, could further en- hance accessibility and practical utility. The ultimate goal is to serve as an accessible chess education tool. Current chess engines provide optimal moves but offer little pedagogical value, as their evaluations are opaque and their suggestions often beyond human comprehension. A foundation model trained with Master Distillation could explain concepts at appropriate skill levels, identify common mistake patterns, and guide learners through tactical and strategic themes, democratizing access to high-quality chess instruction. Master Distillation Generalization. Master Distillation is not specific to chess. The framework applies to any domain where strong bespoke systems exist but language models lack sufficient reasoning capabilities due to sparse training data. The key requirements are a master system that can provide verifiable solutions and problems with clear correct- 8 Grounded Chess Reasoning in Language Models via Master Distillation ness criteria for reward design. Several niche domains fit this profile. Circuit design has simulation tools that validate functionality, but generating correct designs with explana- tions remains challenging. Music composition has harmonic analyzers and counterpoint checkers that verify rule adher- ence, yet LLMs lack grounding in music theory. Molecular docking has scoring functions that evaluate binding affin- ity, while the reasoning behind favorable configurations is difficult for LLMs to articulate. While the master system component transfers directly across domains, Feigned Dis- covery Prompting requires domain-specific redesign. The core principle remains constant: the LLM must pretend to discover the solution through genuine reasoning rather than reveal that it was given the answer. However, what constitutes authentic reasoning varies by domain. Impact Statement This paper presents Master Distillation, a framework for distilling expert system reasoning into language models. While we demonstrate the approach on chess, a recreational domain with limited direct societal risks, we discuss broader implications. On the positive side, our work contributes to making AI reasoning more transparent and explainable. Unlike opaque neural systems that output only predictions, models trained with Master Distillation generate human- readable explanations grounded in expert knowledge. This transparency could benefit educational applications, helping learners understand not just what the correct answer is, but why. The general framework of distilling specialized systems into language models could extend to other domains. While this offers opportunities for democratizing expert knowl- edge, it also raises considerations about the reliability of generated explanations and the potential for overreliance on AI-generated reasoning in high-stakes domains. We empha- size that our current work focuses on chess, where errors have minimal consequences, and caution that applications to domains with greater societal impact would require careful validation. References Ahmadian, A., Cremer, C., Gall ́ e, M., Fadaee, M., Kreutzer, J., Pietquin, O., ̈ Ust ̈ un, A., and Hooker, S. Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024. URLhttps://arxiv. org/abs/2402.14740. Cooper, A.Pgn2fen:A benchmark for eval- uating llm chess reasoning.Blog post, May 2025. URLhttps://w.aidancooper.co.uk/ pgn2fen-benchmark/. Accessed 2025-09-16. Cui, G., Yuan, L., Wang, Z., Wang, H., Zhang, Y., Chen, J., Li, W., He, B., Fan, Y., Yu, T., Xu, Q., Chen, W., Yuan, J., Chen, H., Zhang, K., Lv, X., Wang, S., Yao, Y., Han, X., Peng, H., Cheng, Y., Liu, Z., Sun, M., Zhou, B., and Ding, N. Process reinforcement through implicit rewards, 2025. URL https://arxiv.org/abs/2502.01456. Edwards, S. J. Portable game notation specification and im- plementation guide, 1994. Chess programming standard. Feng, X., Luo, Y., Wang, Z., Tang, H., Yang, M., Shao, K., Mguni, D., Du, Y., and Wang, J. Chessgpt: Bridging policy learning and language modeling. Advances in Neural Information Processing Systems, 36:7216–7262, 2023. Guo, D., Yang, D., and Zhang, H. e. a.Deepseek- r1 incentivizes reasoning in llms through reinforce- ment learning. Nature, 645(8081):633–638, Septem- ber 2025.ISSN 1476-4687.doi:10.1038/ s41586-025-09422-z. URLhttp://dx.doi.org/ 10.1038/s41586-025-09422-z. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al. Lora: Low-rank adaptation of large language models. ICLR, 1(2):3, 2022. Hwang, D., Lee, H., Choo, J., Park, D., and Park, J. Can large language models develop strategic reasoning? post- training insights from learning chess. arXiv preprint arXiv:2507.00726, 2025a. Hwang, D., Lee, H., Choo, J., Park, D., and Park, J. Can large language models develop strategic reasoning? post-training insights from learning chess, 2025b. URL https://arxiv.org/abs/2507.00726. Jannala, S. D. Chess LLM Arena: A Framework for Eval- uating Strategic Decision-Making in Large Language Models. PhD thesis, Dublin, National College of Ireland, 2025. Jhamtani, H., Gangal, V., Hovy, E., Neubig, G., and Berg- Kirkpatrick, T. Learning to generate move-by-move commentary for chess games from large-scale social fo- rum data. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), p. 1661–1671, Melbourne, Aus- tralia, July 2018. Association for Computational Lin- guistics. doi: 10.18653/v1/P18-1154. URLhttps: //aclanthology.org/P18-1154/. Jin, B., Zeng, H., Yue, Z., Yoon, J., Arik, S., Wang, D., Zamani, H., and Han, J. Search-r1: Training llms to reason and leverage search engines with reinforcement learning. arXiv preprint arXiv:2503.09516, 2025. 9 Grounded Chess Reasoning in Language Models via Master Distillation Kagi Search. llm-chess-puzzles: Benchmark llm reason- ing capability by solving chess puzzles. GitHub repos- itory, April 2025.URLhttps://github.com/ kagisearch/llm-chess-puzzles. 1000 FEN puzzles; updated 2025-04-26. Accessed 2025-09-16. Kim, J., Goh, J., Hwang, I., Cho, J., and Ok, J. Bridging the gap between expert and language models: Concept- guided chess commentary generation and evaluation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 9497–9516, 2025. Langley, P. Crafting papers on machine learning. In Langley, P. (ed.), Proceedings of the 17th International Conference on Machine Learning (ICML 2000), p. 1207–1216, Stan- ford, CA, 2000. Morgan Kaufmann. Lee, A., Gulli, A., Chang, B., Fraser, B., Doerschuk-Tiberi, B., Woodford, C., Prichard, C., Sugnet, C., Hennes, D., Yeroshenko, D., Sterling, D., Dong, E., Wang, H., Jobe, H., Gemp, I., Hwang, J., Ren, J., Schultz, J., Lipovetz, J., Peng, J., Chiu, J., Hakimzadeh, K., Ol- szewska, K., Prince, L., Hightower, L., Lanctot, M., Plomecka, M., Risdal, M., O’Connell, M., Chen, M., Keating, N., Tomasev, N., Kelly, O., Firat, O., Kirk, P., Jones, R., Daniel, R., Trostle, R., Sharifzadeh, S., Liu, S., Chung, T., Mason, T., Cukierski, W., Xu, Y., Yan, Y., Su, Y., Zhuang, Y., Han, Y., and Zhai, Y. Chess Text In- put.https://w.kaggle.com/benchmarks/ kaggle/chess-text, 2025.Google DeepMind, Google Cloud, Kaggle. Li, K., Zhang, Z., Yin, H., Zhang, L., Ou, L., Wu, J., Yin, W., Li, B., Tao, Z., Wang, X., et al. Websailor: Navigating super-human reasoning for web agent. arXiv preprint arXiv:2507.02592, 2025. Liu, Z., Chen, C., Li, W., Qi, P., Pang, T., Du, C., Lee, W. S., and Lin, M. Understanding r1-zero-like training: A critical perspective, 2025. URLhttps://arxiv. org/abs/2503.20783. McIlroy-Young, R., Sen, S., Kleinberg, J., and Anderson, A. Aligning superhuman ai with human behavior: Chess as a model system. In Proceedings of the 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, p. 1677–1687, Virtual Event, CA, USA, 2020. ACM. doi: 10.1145/3394486.3403219. URLhttps://dl. acm.org/doi/10.1145/3394486.3403219. McIlroy-Young, R., Wang, R., Sen, S., Kleinberg, J., and Anderson, A. Learning models of individual behavior in chess. In Proceedings of the 28th ACM SIGKDD Con- ference on Knowledge Discovery and Data Mining, p. 1253–1263, 2022. Monroe, D. and Chalmers, P. A. Mastering chess with a transformer model. arXiv preprint arXiv:2409.12272, 2024. Nasu, Y. Efficiently updatable neural-network-based eval- uation functions for computer shogi. The 28th World Computer Shogi Championship Appeal Document, 2018. Original NNUE paper (in Japanese). Noever, D., Ciolino, M., and Kalin, J. The chess transformer: Mastering play using generative language models. arXiv preprint arXiv:2008.04057, 2020. OpenAI. Openai o1 system card, 2024. URLhttps: //arxiv.org/abs/2412.16720. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback, 2022. URLhttps: //arxiv.org/abs/2203.02155. Pascutto, G.-C. and Linscott, G. Leela chess zero, 2018. URL http://lczero.org/. Pascutto, G.-C., Linscott, G., et al. Leela zero: Go en- gine with no human provided knowledge.https:// github.com/leela-zero/leela-zero, 2018. Ruoss, A., Del ́ etang, G., Medapati, S., Grau-Moya, J., Wen- liang, L. K., Catt, E., Reid, J., Lewis, C. A., Veness, J., and Genewein, T. Amortized planning with large-scale transformers: A case study on chess. Advances in Neural Information Processing Systems, 37:65765–65790, 2024. Saplin, M.Llm chess leaderboard.Project website, 2025.URLhttps://maxim-saplin.github. io/llm_chess/. Elo vs. calibrated engine and ran- dom baselines. Accessed 2025-09-16. Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O.Proximal policy optimization algo- rithms, 2017.URLhttps://arxiv.org/abs/ 1707.06347. Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y. K., Wu, Y., and Guo, D. Deepseekmath: Pushing the limits of mathemat- ical reasoning in open language models, 2024. URL https://arxiv.org/abs/2402.03300. Sheng, G., Zhang, C., Ye, Z., Wu, X., Zhang, W., Zhang, R., Peng, Y., Lin, H., and Wu, C. Hybridflow: A flexi- ble and efficient rlhf framework. arXiv preprint arXiv: 2409.19256, 2024. 10 Grounded Chess Reasoning in Language Models via Master Distillation Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. Mastering the game of go with deep neural networks and tree search. nature, 529(7587):484–489, 2016. Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Grae- pel, T., et al. Mastering chess and shogi by self-play with a general reinforcement learning algorithm. arXiv preprint arXiv:1712.01815, 2017a. Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. Mastering the game of go without human knowledge. nature, 550(7676):354–359, 2017b. Stockfish Team. Stockfish: A strong open-source chess engine, 2023. URLhttps://stockfishchess. org/. NNUE integration since version 12 (2020). St ̈ ockl, A. Watching a language model learning chess. In Mitkov, R. and Angelova, G. (eds.), Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), p. 1369–1379, Held Online, September 2021. INCOMA Ltd. URLhttps://aclanthology.org/2021. ranlp-1.153/. Sutton, R. The bitter lesson. Incomplete Ideas (blog), 13(1): 38, 2019. Swingle, C., Mellsop, H., and Langshur, A. ChePT – applying deep neural transformer models to chess move prediction and self-commentary. CS224N course project report, Stanford University, 2021.URLhttps: //web.stanford.edu/class/cs224n/ reports/final_reports/report087.pdf. Stanford CS224N: Natural Language Processing with Deep Learning. Tang, Z., Jiao, D., Yang, B., and Anderson, A. Seam: Se- mantically equivalent across modalities benchmark for vision-language models. In Second Conference on Lan- guage Modeling. Tang, Z., Jiao, D., McIlroy-Young, R., Kleinberg, J., Sen, S., and Anderson, A. Maia-2: A unified model for human- ai alignment in chess. Advances in Neural Information Processing Systems, 37:20919–20944, 2024. Tang, Z., Jiao, D., Xue, E., McIlroy-Young, R., Kleinberg, J., Sen, S., and Anderson, A. Learning to imitate with less: Efficient individual behavior modeling in chess. Transac- tions on Machine Learning Research, 2026. ISSN 2835- 8856. URLhttps://openreview.net/forum? id=iw4kjcw319. Wang, S., Ji, L., Wang, R., Zhao, W., Liu, H., Hou, Y., and Wu, Y. N. Explore the reasoning capability of llms in the chess testbed. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associa- tion for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), p. 611–622, 2025. Wen, Q., Tang, Z., and Anderson, A. Chessqa: Evaluating large language models for chess understanding. arXiv preprint arXiv:2510.23948, 2025. Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. Qwen3 technical report, 2025a. URLhttps: //arxiv.org/abs/2505.09388. Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025b. Yu, Q., Zhang, Z., Zhu, R., Yuan, Y., Zuo, X., Yue, Y., Dai, W., Fan, T., Liu, G., Liu, L., Liu, X., Lin, H., Lin, Z., Ma, B., Sheng, G., Tong, Y., Zhang, C., Zhang, M., Zhang, W., Zhu, H., Zhu, J., Chen, J., Chen, J., Wang, C., Yu, H., Song, Y., Wei, X., Zhou, H., Liu, J., Ma, W.- Y., Zhang, Y.-Q., Yan, L., Qiao, M., Wu, Y., and Wang, M. Dapo: An open-source llm reinforcement learning system at scale, 2025. URLhttps://arxiv.org/ abs/2503.14476. Zang, H., Yu, Z., and Wan, X. Automated chess com- mentator powered by neural chess engine. In Proceed- ings of the 57th Annual Meeting of the Association for Computational Linguistics, p. 5952–5961, Florence, Italy, July 2019. Association for Computational Lin- guistics. doi: 10.18653/v1/P19-1597. URLhttps: //aclanthology.org/P19-1597/. Zhang, Y., Han, X., Li, H., Chen, K., and Lin, S. Complete chess games enable llm become a chess master. In Pro- ceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), p. 1–7, 2025a. Zhang, Y., Jacob, A. P., Lai, V., Fried, D., and Ippolito, D. Human-aligned chess with a bit of search. In The 11 Grounded Chess Reasoning in Language Models via Master Distillation Thirteenth International Conference on Learning Rep- resentations, 2025b. URLhttps://openreview. net/forum?id=bc2H72hGxB. Zheng, C., Liu, S., Li, M., Chen, X.-H., Yu, B., Gao, C., Dang, K., Liu, Y., Men, R., Yang, A., Zhou, J., and Lin, J. Group sequence policy optimization, 2025. URL https://arxiv.org/abs/2507.18071. Zheng, Y., Zhang, R., Zhang, J., Ye, Y., Luo, Z., Feng, Z., and Ma, Y. Llamafactory: Unified efficient fine- tuning of 100+ language models. In Proceedings of the 62nd Annual Meeting of the Association for Computa- tional Linguistics (Volume 3: System Demonstrations), Bangkok, Thailand, 2024. Association for Computational Linguistics. URL http://arxiv.org/abs/2403. 13372. 12 Grounded Chess Reasoning in Language Models via Master Distillation Figure 2. Chess puzzle-solving accuracy of C1 compared to frontier LLMs. ¬ Master Distillation Prompt General Instruction You are a chess grandmaster generating training data for teaching small LLMs to solve chess puzzles. CRITICAL: The small model only sees the FEN string. Your analysis must show how to go from FEN→ piece positions→ tactical relationships→ solution. Ground everything in explicit square names. Context (provided to teacher, hidden from student) FEN: fen Pieces: piecearrangement Side to Move: side tomove Opponent’s Last Move: last move Solution: move Rating: rating Themes: themes PVs: principal variation (omitted for mateIn1) Legal Moves: legalmoves 13 Grounded Chess Reasoning in Language Models via Master Distillation Task Description TASK: Write natural chain-of-thought analysis arriving at move. Your analysis should cover (in natural prose, vary the structure): • Where key pieces are (use square names: ‘the queen on h5’, ‘king on g1’) • What tactical relationship exists (attacks, pins, weak squares, defender counts) • Why the move works (what it threatens, why opponent can’t respond adequately) Style: • Objective voice, no ‘I see/notice’ • Standard notation with brief clarification when helpful: ‘Qxh7+ (queen takes h7 with check)’ • Never mention engine scores, ratings, themes, or that you were given the solution • 4--10 sentences, scaled to complexity End with exactly: FINAL ANSWER: move Table 7. RLVR Training Hyperparameters. HyperparameterValue Model Configuration Base ModelC1-SFT-4B AlgorithmDAPO-C1 Gradient CheckpointingEnabled Data Configuration Max Prompt Length1024 Max Response Length512 Generation Batch Size96 Training Batch Size32 Rollout Samples32 Sampling Parameters Temperature1.0 Top-p1.0 Top-k-1 GPU Memory Utilization0.8 Optimizer Learning Rate1× 10 −6 PPO Mini Batch Size8 PPO Micro Batch Size (per GPU)32 Log Prob Micro Batch Size256 RL Specific KL Loss Coefficient0.001 Entropy Coefficient0 Clip Ratio (low / high)0.2 / 0.28 Loss AggregationToken-mean KL in RewardDisabled Filter GroupsEnabled (max batches: 5) Training Total Epochs1 Save Frequency500 Test Frequency10 Num GPUs4 Early-Stop Tolerance5 14 Grounded Chess Reasoning in Language Models via Master Distillation Table 8. Supervised Fine-Tuning Hyperparameters ParameterSetting Base ModelQwen3-4B-Instruct Training MethodFull-parameter fine-tuning OptimizerAdamW (DeepSpeed ZeRO-3) Learning Rate1e−5 Batch Size16× 4 (grad accum)× 4 GPUs = 256 Epochs10 SchedulerCosine with 10% warmup Max Length1024 tokens PrecisionBF16 Validation Interval50 Steps Early-Stop Tolerance5 Table 9. Data statistics for SFT and RLVR training. SFTRLVR Total samples39,60939,585 Unique themes6666 Balanced themes (K)5050 Target per theme (M )800800 Rating range399–3,154399–3,196 Average rating1,5401,539 15 Grounded Chess Reasoning in Language Models via Master Distillation Table 10. Theme distribution in SFT and RLVR training data. We apply balanced sampling withK = 50rare themes and targetM = 800 samples per theme. endgamemateshortmiddlegamecrushinglongadvantagemateIn2oneMovemateIn1masterveryLong SFT20,817 17,971 17,205 16,210 12,957 10,236 7,858 6,930 6,856 6,824 6,133 5,294 RLVR20,752 17,872 17,056 16,210 13,054 10,292 7,834 6,767 6,860 6,833 6,056 5,359 sacrifice advancedPawn defensiveMove kingsideAttackforkattractionopeningpinmateIn3quietMoverookEndgamediscoveredAttackdeflectionexposedKingpawnEndgamepromotion SFT4,932 3,288 3,277 3,001 2,948 2,672 2,572 2,370 2,335 2,307 2,195 2,164 2,136 2,019 1,829 1,820 RLVR5,037 3,316 3,178 2,957 2,939 2,718 2,614 2,364 2,395 2,278 2,159 2,245 2,279 2,005 1,850 1,901 backRankMatemasterVsMasterhangingPiecequeensideAttackskewerclearancedoubleCheckmateIn4intermezzoqueenEndgamebishopEndgamequeenRookEndgamezugzwangattackingF2F7trappedPieceknightEndgamemateIn5capturingDefender SFT1,596 1,557 1,439 1,393 1,165 1,083 1,077 1,055 1,040 966 964 891 874 872 867 866 827 827 RLVR1,594 1,587 1,515 1,425 1,122 1,081 1,061 1,047 1,053 967 973 898 892 871 874 850 830 819 xRayAttackcornerMateinterferencearabianMatehookMateanastasiaMateequalityvukovicMateblindSwineMatetriangleMatesuperGMenPassantkillBoxMatecastlingbodenMatedoubleBishopMatedovetailMatesmotheredMate balestraMate underPromotion SFT825 818 817 810 808 806 805 805 805 805 803 803 802 800 800 800 800 800 697 512 RLVR830 822 814 807 810 808 807 808 805 808 802 805 800 800 800 801 801 801 668 517 16 Grounded Chess Reasoning in Language Models via Master Distillation Table 11. Performance comparison across chess tactic themes. All values represent accuracy (%). advancedPawnattractionbackRankMatecapturingDefenderdefensiveMovedeflectiondiscoveredAttackdoubleCheckforkhangingPiecemateIn1mateIn2pinpromotionqueensideAttacksacrificeskewertrappedPiecexRayAttackzugzwang Proprietary models gpt-584921008468928080889696968476929284688072 gemini-3-pro8068928068688472888084847668968884568068 gemini-3-flash28 323244 40 44 36 20 40 44 56 52 60 24 44 52 36 24 28 24 gpt-5-chat28 406836 44 48 32 36 48 32 44 64 32 40 44 52 48 28 40 32 gemini-2.5-pro28 2052836 20 20 52 32 36 40 28 28 44 36 28 32 28 20 32 claude-sonnet-4.5 44 125216820 48 40 48816 36 20 56 12 40 32040 24 claude-sonnet-436 284416 16 28 20 32 36 28 24 24 28 32 20 40 24444 12 claude-haiku-4.540 284812 16 40 32 32 48 20 32 20 20 36 24 24 120208 gemini-2.5-flash16 12000120124412 1242448160168 claude-haiku-3.588848041681681288048084 Open-source models deepseek-chat-v3.140 244428 36 12420 24 20 16420 24 32 16 208840 qwen3-next-80b-a3b16 28824 12 24 24 12 36 24 12 12 12 16 12820824 20 deepseek-r1-052816 20424 24 24 12 16 32 36 12412 2088168420 qwen3-max16 20412816812 24 12 16 40820 128200812 llama-4-maverick2404884416 12 12 1281644012012 12 mistral-medium-3.1 128048812816081288124120128 llama-4-scout00004000000000000004 gemma-3-27b 00000000000000000000 Ours C1-SFT-4B48 527228 48 36 28 52 48 52 56 40 24 48 60 44 44 12 60 72 C1-SFT-8B 52 448420 60 40 56 44 40 52 64 40 36 48 60 36 44848 56 C1-4B646484566036445236646456285268605245276 17