Paper deep dive
ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation
Zilong Liu, Xuewen Zhang, Jinrui Xing, Juyi Qiao, Huiyong Wang, Junming Jiao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 7/5/2026, 2:56:40 AM
Summary
The paper proposes ARKD (Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation), a novel framework for compressing Large Language Models (LLMs). Traditional knowledge distillation often relies on a single KL divergence objective (either Forward KL or Reverse KL), which leads to a trade-off between mode-covering (FKL) and mode-seeking (RKL) behaviors. ARKD addresses this by using a reinforcement learning-based policy network that dynamically assigns weights to FKL and RKL based on real-time statistical features (entropy, variance, and current KL values) of the teacher and student distributions. This allows the model to adaptively navigate the distribution alignment process, improving both generation quality and out-of-distribution generalization. Experimental results on GPT-2 and LLaMA models show that ARKD outperforms static and greedy heuristic-based distillation methods across metrics like Rouge-L and BertScore.
Entities (10)
Relation Signals (6)
ARKD → combines → Forward KL Divergence
confidence 100% · The policy network dynamically assigns weights to FKL and RKL
ARKD → combines → Reverse KL Divergence
confidence 100% · The policy network dynamically assigns weights to FKL and RKL
Forward KL Divergence → hasbehavior → mode-covering
confidence 100% · FKL exhibits mode-covering behavior spanning all target modes
Reverse KL Divergence → hasbehavior → mode-seeking
confidence 100% · RKL demonstrates mode-seeking behavior concentrating on the dominant mode
GPT-2 → isevaluatedon → DollyEval
confidence 100% · In-distribution performance on DollyEval (500 test samples) for various teacher-student configurations.
ARKD → uses → Reinforcement Learning
confidence 100% · We propose an adaptive knowledge distillation framework based on RL.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Knowledge distillation (KD) is a key technique for compressing Large Language Models (LLMs), yet methods relying on a single KL objective often fail to balance primary distribution fitting with long-tail probability modeling, limiting both generation quality and generalization. To address this, we analyze the complementary roles of forward and reverse KL divergence (FKL/RKL) in distribution alignment from theoretical and empirical perspectives. We then propose a reinforcement-learning-based adaptive KL-weighted distillation framework, in which a policy network dynamically assigns weights to FKL and RKL based on teacher-student distributional characteristics, guided by immediate reward signals to achieve dual alignment on principal and long-tail modes. Extensive experiments demonstrate consistent improvements across Rouge-L and BertScore metrics, surpassing greedy heuristics by 0.4-0.6 points and outperforming other baseline methods on diverse benchmarks.
Tags
Links
- Source: https://arxiv.org/abs/2606.29869v1
- Canonical: https://arxiv.org/abs/2606.29869v1
Trouble viewing inline? Open PDF directly →
Full Text
54,001 characters extracted from source content.
Expand or collapse full text
ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation Zilong Liu1,* Xuewen Zhang1,2,* Jinrui Xing1 Juyi Qiao1 Huiyong Wang1 Junming Jiao1 1 Li Auto, Beijing, China 2 School of Software and Microelectronics, Peking University, China Abstract Knowledge distillation (KD) is a key technique for compressing Large Language Models (LLMs), yet methods relying on a single KL objective often fail to balance primary distribution fitting with long-tail probability modeling, limiting both generation quality and generalization. To address this, we analyze the complementary roles of forward and reverse KL divergence (FKL/RKL) in distribution alignment from theoretical and empirical perspectives. We then propose a reinforcement-learning-based adaptive KL-weighted distillation framework, in which a policy network dynamically assigns weights to FKL and RKL based on teacher–student distributional characteristics, guided by immediate reward signals to achieve dual alignment on principal and long-tail modes. Extensive experiments demonstrate consistent improvements across Rouge-L and BertScore metrics, surpassing greedy heuristics by 0.4–0.6 points and outperforming other baseline methods on diverse benchmarks. ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation Zilong Liu1,* Xuewen Zhang1,2,* Jinrui Xing1 Juyi Qiao1 Huiyong Wang1 Junming Jiao1 1 Li Auto, Beijing, China 2 School of Software and Microelectronics, Peking University, China **footnotetext: Equal contribution. 1 Introduction As Large Language Models (LLMs) have emerged as foundational technologies in artificial intelligence (Zhang et al., 2024a), they demonstrate remarkable capabilities in natural language understanding and generation (Jiang et al., 2026). However, their substantial parameter counts and computational demands severely constrain practical deployment. Consequently, efficient model compression—achieving lightweight models with reduced inference costs while maintaining robust performance and generalization—has become a central research focus (Zhu et al., 2024). Knowledge Distillation (KD), as an established compression paradigm (Hinton et al., 2015; Fang et al., 2026), enables student models to approximate teacher output distributions or hidden representations (Gou et al., 2021), thereby enhancing student performance while significantly reducing parameter size and hardware requirements (Acharya et al., 2024). Despite substantial progress, traditional KD objectives driven by Kullback-Leibler (KL) divergence—particularly forward KL (FKL)—face critical limitations (Yang et al., 2024a). In open-ended text generation, students are compelled to cover all modes of the teacher distribution. Given the inherent capacity gap, this often results in overestimated probabilities in low-confidence, long-tail regions (i.e., low-probability tokens in the teacher distribution) (Zhang et al., 2024b), degrading generation quality. Additionally, reliance on in-distribution (ID) training data yields insufficient generalization to out-of-distribution (OOD) samples (Yang et al., 2023). Figure 1: Illustration of matching a single Gaussian to a Gaussian mixture using both forward and reverse KL divergence in a toy experiment. Figure 2: Comparison of distillation methods. (a) Sequence-level distillation fine-tunes the student with teacher-generated data, while token-level distillation uses the teacher’s distribution at each token as supervision. (b) Our method introduces an adaptive KL-weighted distillation framework, where a policy network dynamically assigns forward and reverse KL weights based on output distributions, guided by reinforcement learning rewards. This approach achieves consistent improvements over previous methods (see Tables 1 and 2). To mitigate these challenges, reverse KL (RKL) divergence has been introduced as an alternative (Gu et al., 2024), encouraging students to focus on primary teacher modes while disregarding low-probability regions. However, this mode-seeking behavior compromises diversity and coverage. Figure 1 illustrates this trade-off via a toy experiment: FKL exhibits mode-covering behavior spanning all target modes, while RKL demonstrates mode-seeking behavior concentrating on the dominant mode. Our method adaptively balances both objectives—jointly referred to as bidirectional KL distillation throughout this paper—to achieve better distribution alignment. These distinct characteristics highlight a fundamental trade-off: neither FKL nor RKL alone sufficiently addresses probability alignment and generalization in knowledge distillation. Crucially, the desirable FKL/RKL balance shifts as the student–teacher gap evolves over training, which a single static interpolation coefficient cannot capture. Notably, Reinforcement Learning’s mechanisms in policy optimization, reward modeling, and distribution alignment offer promising avenues to navigate this trade-off (Agarwal et al., 2024). By framing distillation as reward-driven optimization, RL enables adaptive focus on high-confidence regions—avoiding low-probability overfitting—while enhancing generalization to unseen samples through policy exploration (Li et al., 2024). Building on these insights, we systematically investigate the integration of reinforcement learning within knowledge distillation. We propose a novel RL-based joint optimization framework (Figure 2), where reward signals dynamically guide the student to capture both primary and long-tail modes of the teacher distribution, thereby mitigating long-tail overestimation and limited generalization inherent in static KL-based distillation. The main contributions are as follows: • We theoretically analyze FKL and RKL for distribution alignment, establishing their complementarity and convergence consistency as the foundation for adaptive distillation. • We propose Adaptive Reinforcement Learning Guided Knowledge Distillation (ARKD), which employs a policy network to dynamically weight FKL and RKL via reinforcement learning, enabling state-dependent adaptation beyond static heuristics. • Our experiments demonstrate that ARKD consistently improves generation quality, distribution fitting, and generalization ability across multiple benchmarks and model scales. 2 Related work 2.1 KD for LLM Knowledge distillation for LLMs can be categorized into black-box and white-box approaches (Xu et al., 2024). Black-box methods train student models using teacher-generated outputs without accessing internal representations (Kim and Rush, 2016; Chiang et al., 2023), while white-box methods leverage teacher logits or hidden states for token-level distribution alignment via FKL or RKL (Kim et al., 2024). White-box distillation has proven effective for LLM compression and deployment. Recent work explores adaptive teaching strategies to enhance distillation, such as token-level teaching mode adaptation (Zhong et al., 2024). This work advances white-box KD by proposing an adaptive strategy to optimally balance FKL and RKL. 2.2 Forward KL and Reverse KL Forward KL and reverse KL exhibit fundamentally distinct distributional behaviors in knowledge distillation. FKL encourages the student distribution q to cover the entire support of teacher distribution p, exhibiting mode-covering (mean-seeking) behavior. Conversely, RKL drives q to concentrate on a single dominant mode, exhibiting mode-seeking behavior. Recent work advocates for RKL in LLM distillation to avoid overestimating long-tail variants. Gu et al. (2024) introduce MiniLLM, demonstrating RKL’s effectiveness in preventing low-probability region overestimation. Wu et al. (2025) show that despite converging to identical objectives, FKL and RKL emphasize head and tail regions differently during training, motivating adaptive weighting strategies. Wen et al. (2023) propose f-DISTILL, a framework that formulates sequence-level KD as minimizing generalized f-divergences to balance mode averaging and mode collapsing. Jin et al. (2026) introduce entropy-aware on-policy distillation that augments reverse KL with forward KL when teacher entropy is high to preserve generation diversity. While some approaches combine divergences through static weighting (Tan et al., 2023), they often neglect the dynamic nature of distributional structure evolution during training, which our approach explicitly addresses. 2.3 Combining RL and KD Integrating RL with KD transforms teacher feedback into quantifiable rewards (Yang et al., 2024b; Xu et al., 2026), enabling students to optimize behavior through preference-based signals (Nguyen et al., 2025; Zhang et al., 2026). Agarwal et al. (2024) propose Generalized Knowledge Distillation (GKD), which trains students on on-policy self-generated outputs guided by teacher probabilities to mitigate train-test discrepancy in auto-regressive models, enabling seamless integration with RL fine-tuning (e.g. RLAIF). RL frameworks enable dynamic adjustment of distillation hyperparameters based on training performance. Prior work formulates distillation as sequential decision-making (Lee et al., 2023), where reward-guided loss adaptation improves knowledge transfer efficiency. Kwon et al. (2023) employ RL to dynamically select between FKL and RKL across training stages. Overall, RL provides adaptive and intelligent mechanisms for KD, representing a promising direction for LLM compression. Building on these advances, our work introduces RL-based adaptive KL divergence weighting. 3 Methodology 3.1 FKL & RKL Objective Consistency In conditional language generation tasks, the model is required to generate a target sequence y=ytt=1Ty=\y_t\_t=1^T based on an input sequence x, where yt∈=Y1,Y2,…,YVy_t =\Y_1,Y_2,…,Y_V\. Traditional knowledge distillation methods transmit knowledge by aligning the token generation distributions between the teacher and student models. Specifically, the objective functions for FKL and RKL are defined as follows: JFKL J_FKL =∑t=1TKL(p(yt|y<t)∥qθ(yt|y<t)) = _t=1^TKL (p(y_t|y_<t)\|q_θ(y_t|y_<t) ) =∑t=1T∑j=1Vp(Yj|y<t)logp(Yj|y<t)qθ(Yj|y<t), = _t=1^T _j=1^Vp(Y_j|y_<t) p(Y_j|y_<t)q_θ(Y_j|y_<t), (1) JRKL J_RKL =∑t=1TKL(qθ(yt|y<t)∥p(yt|y<t)) = _t=1^TKL (q_θ(y_t|y_<t)\|p(y_t|y_<t) ) =∑t=1T∑j=1Vqθ(Yj|y<t)logqθ(Yj|y<t)p(Yj|y<t). = _t=1^T _j=1^Vq_θ(Y_j|y_<t) q_θ(Y_j|y_<t)p(Y_j|y_<t). (2) Theoretical analysis shows that, due to the decomposability of the above objective functions, token-wise analysis is feasible. Let the pre-softmax logits of the teacher and student models at time step t be denoted as pz^p and qz^q, respectively. Then, the token probability distributions can be expressed as: p(Yj|y<t) p(Y_j|y_<t) =exp(zjp)∑k=1Vexp(zkp), = (z_j^p) _k=1^V (z_k^p), (3) qθ(Yj|y<t) q_θ(Y_j|y_<t) =exp(zjq)∑k=1Vexp(zkq). = (z_j^q) _k=1^V (z_k^q). (4) By applying gradient analysis, we obtain: ∂JFKL∂zjq ∂ J_FKL∂ z_j^q =qθ(Yj|y<t)−p(Yj|y<t), =q_θ(Y_j|y_<t)-p(Y_j|y_<t), (5) ∂JRKL∂zjq ∂ J_RKL∂ z_j^q =qθ(Yj|y<t)(logqθ(Yj|y<t)p(Yj|y<t)−JRKL). =q_θ(Y_j|y_<t) ( q_θ(Y_j|y_<t)p(Y_j|y_<t)-J_RKL ). (6) When the gradients are zero, the convergence condition for both types of KL divergence is qθ(Yj|y<t)=p(Yj|y<t)(∀j), q_θ(Y_j|y_<t)=p(Y_j|y_<t) (∀ j), (7) which means the student model exactly replicates the output distribution of the teacher model. For RKL, zero gradient requires log(qθ/p)=JRKL (q_θ/p)=J_RKL for all j; since the right-hand side is a constant, qθ/pq_θ/p must be constant across j, which combined with normalization forces qθ=pq_θ=p. This result reveals the fundamental consistency of both forward and reverse KL divergence with respect to the distribution matching objective, thus providing a theoretical basis for designing adaptive distillation strategies. 3.2 RL-based Knowledge Distillation Framework Building on the analysis of the consistency between FKL and RKL objectives in Section 3.1, we propose an adaptive knowledge distillation framework based on RL. This approach reframes the static KL weight assignment in traditional distillation as a sequential decision-making problem, introducing a policy network as the RL agent to dynamically adjust the weights based on statistical features of the teacher and student outputs (e.g., entropy, variance, KL divergence). A key motivation for RL-based adaptation is that knowledge distillation constitutes a sequential decision process: at each step t, the weight choice αt _t determines the parameter update θt+1←θt−η∇θLdistil(t) _t+1← _t-η _θL_distil^(t), which in turn shapes the loss landscape (FKLt+1,RKLt+1)(FKL_t+1,RKL_t+1) at the next step. While greedy heuristics (e.g., αt=[FKLt<RKLt] _t=I[FKL_t<RKL_t]) can incorporate state features, they rely on fixed rules that optimize immediate loss reduction without considering long-term consequences. Such myopic strategies often lead to suboptimal trajectories—for instance, aggressively selecting RKL in early training may yield rapid initial convergence but risks mode collapse, whereas a balanced approach with temporarily higher loss can preserve diversity and improve final model quality. In contrast, reinforcement learning addresses this limitation by optimizing cumulative reward over the training trajectory rather than instantaneous loss. By formulating weight selection as a Markov Decision Process, the policy network πϕ _φ learns from observed training dynamics to discover non-trivial state-dependent strategies that balance exploration of both KL terms with long-term performance, adaptively navigating phase transitions between FKL’s mode-covering and RKL’s mode-seeking behaviors. This learned adaptivity, validated empirically in our experiments, enables ARKD to surpass both static and greedy baselines (Algorithm 1). Algorithm 1 RL-based Adaptive Knowledge Distillation 1:Instruction tuning datasets D 2:Pre-training corpus DpD_p 3:Fine-tuned teacher model with distribution p 4:Initialized student model with distribution qθq_θ 5:Policy network πϕ _φ for adaptive weight selection 6:Student model parameter θ after distillation training 7:for epoch in epochs do 8: for batch in D and DpD_p do 9: Compute teacher and student outputs (logits and probabilities) for the batch 10: Extract state s=[HT,HS,σT2,σS2,LFKL,LRKL]s=[H_T,H_S, _T^2, _S^2,L_FKL,L_RKL] 11: Compute adaptive weight α=πϕ(s)α= _φ(s), α∈(0,1)α∈(0,1) 12: Compute FKL=KL(p∥qθ)FKL=KL(p q_θ), RKL=KL(qθ∥p)RKL=KL(q_θ p) 13: Compute distillation loss 14: Lkd=α⋅FKL+(1−α)⋅RKLL_kd=α·FKL+(1-α)·RKL 15: Compute language modeling loss 16: Lpt=−∑d∼Dplogqθ(d)L_pt=- _d D_p q_θ(d) 17: Compute total loss 18: L=Lkd+λ⋅LptL=L_kd+λ· L_pt 19: Update student: θ=θ−η⋅∇θLθ=θ-η· _θL 20: Compute reward r=−Lkd.detach()r=-L_kd.detach() 21: Update EMA baseline: b←γb+(1−γ)rb←γ b+(1-γ)\,r 22: Compute advantage A=r−bA=r-b and entropy H(α)=−[αlogα+(1−α)log(1−α)]H(α)=-[α α+(1-α) (1-α)] 23: Compute policy gradient loss 24: Lpolicy=−log(α)⋅A−βHH(α)L_policy=- (α)· A- _H\,H(α) 25: Update policy net: ϕ=ϕ−β⋅∇ϕLpolicyφ=φ-β· _φL_policy 26: end for 27:end for 28:return Student model parameter θ Specifically, for each training batch we form a 6-dimensional state vector s=[HT,HS,σT2,σS2,LFKL,LRKL]s=[H_T,H_S, _T^2, _S^2,L_FKL,L_RKL] comprising teacher/student entropies, variances, and current FKL/RKL values; the policy network then outputs a weight α (for FKL, with 1−α1-α for RKL), which is used to compute the weighted distillation loss. The environment provides an immediate reward signal r, reflecting the effectiveness of the weight assignment, and both the student parameters θ and policy network parameters ϕφ are optimized accordingly. This closed-loop process of state observation, action assignment, loss computation, reward feedback, and parameter update enables adaptive and efficient knowledge distillation. 3.3 Reward Signal Design RL-based adaptive weight allocation hinges on effective reward design. To enable the policy network to autonomously optimize the allocation weight α during training, we define the immediate reward as the negative distillation loss for the current batch: r=−Ldistil(α)r=-L_distil(α) (8) where Ldistil(α)=α⋅KL(p∥qθ)+(1−α)⋅KL(qθ∥p)L_distil(α)=α·KL(p q_θ)+(1-α)·KL(q_θ p) (9) This formulation encourages the policy network to select α values that minimize distillation loss. The optimization objective is to maximize the cumulative reward (i.e., minimize the total distillation loss) over training: maxθs∼[r]=maxθs∼[−Ldistil(πθ(s))] _θ\ E_s [r]= _θ\ E_s [-L_distil( _θ(s))] (10) Through dynamic adjustment, the policy network learns to adaptively select α to best enhance student model performance. 3.4 RL-enhanced Joint Optimization of FKL and RKL Considering that, within the reinforcement learning-based knowledge distillation framework, the weight allocation strategy is dynamically determined by the policy network, the overall optimization objective can be formulated as a collaborative learning process between the student network and the policy network. Accordingly, the weighted distillation loss is constructed as follows: Ldistil(θ,ϕ)=α⋅KL(p∥qθ)+(1−α)⋅KL(qθ∥p)L_distil(θ,φ)=α·KL(p q_θ)+(1-α)·KL(q_θ p) (11) Considering the language modeling loss Lpt(θ)L_pt(θ) of the task, the total loss function can be formulated as: Ltotal(θ,ϕ)=Ldistil(θ,ϕ)+λLpt(θ)L_total(θ,φ)=L_distil(θ,φ)+λ L_pt(θ) (12) where λ>0λ>0 is a balancing coefficient. During training, the parameter updates are performed in two components: (1) Student Model Parameter Update (θ) The optimization objective for the student network is to minimize the total loss: θ∗=argminθLtotal(θ,ϕ)θ^*= _θL_total(θ,φ) (13) The corresponding gradient is given by: ∇θLtotal _θL_total =α⋅∇θFKL+(1−α)⋅∇θRKL =α· _θFKL+(1-α)· _θRKL +λ∇θLpt(θ) +λ _θL_pt(θ) (14) Note that here α is determined by the current policy network parameters ϕφ and is thus treated as a constant with respect to θ during the gradient computation. (2) Policy Network Parameter Update (ϕφ) As described in Section 3.3, the policy network maximizes the immediate reward using the policy gradient method, which can be formulated as follows: r=−Ldistil(θ,ϕ)r=-L_distil(θ,φ) (15) Using REINFORCE with an EMA baseline b and entropy regularization H(α)H(α) to stabilize training and prevent boundary saturation, the policy loss is: Lpolicy(ϕ)=−log(α)(r−b)−βHH(α)L_policy(φ)=- (α)\,(r-b)- _H\,H(α) (16) The gradient of the policy loss with respect to ϕφ is (omitting the entropy term for brevity): ∇ϕLpolicy=−(r−b)⋅1α⋅∂α∂ϕ _φL_policy=-(r-b)· 1α· ∂α∂φ (17) If α=σ(fϕ(s))α=σ(f_φ(s)) is parameterized via a sigmoid activation, then: ∂α∂ϕ=α(1−α)⋅∂fϕ(s)∂ϕ ∂α∂φ=α(1-α)· ∂ f_φ(s)∂φ (18) As a result, the policy network is ultimately driven to output weight allocations that lead to lower distillation loss, thereby achieving adaptive weighting adjustment. 4 Experiments Method GPT2 1.5B → GPT2 120M GPT2 1.5B → GPT2 340M LLaMA 13B → LLaMA 7B Rouge-L BertScore Rouge-L BertScore Rouge-L BertScore Teacher 27.1 0.842 27.1 0.858 31.3 0.891 SFT w/o KD 22.3 0.772 24.0 0.793 26.0 0.810 SeqKD 21.9 0.765 24.2 0.798 26.6 0.818 FKL 22.9 0.781 24.9 0.803 27.1 0.825 RKL 23.3 0.797 25.5 0.813 28.2 0.844 FKL+RKL 23.2 0.794 25.7 0.817 28.3 0.848 MiniLLM 24.6 — 25.4 — 29.0 — ARKD 24.5 0.815 26.1 0.827 29.1 0.868 Table 1: In-distribution performance on DollyEval (500 test samples) for various teacher-student configurations. Bold indicates the best student model. Rouge-L numbers for MiniLLM are taken from (Gu et al., 2024); BertScore is not reported in the original work. 4.1 Experimental Setup To comprehensively evaluate our adaptive knowledge distillation approach, we conduct experiments on instruction tuning tasks, where models generate natural language responses conditioned on given instructions. We first fine-tune large language models on instruction–response datasets to obtain teacher models. Various distillation methods are then applied on the same dataset to measure student models’ instruction-following capabilities. Teacher and Student Models Teacher and student models are chosen from mainstream large language model families of different sizes. Teacher models include GPT-2 (1.5B) and LLaMA (13B), while student models include GPT-2 (120M, 340M) and LLaMA (7B). Teacher models are fine-tuned on instruction datasets and subsequently serve as teachers for knowledge distillation. Data Preprocessing and Training We use the databricks-dolly-15K dataset, containing 15,000 human-annotated instruction–response pairs. Instances exceeding the maximum context length and short responses (less than 10 words) are filtered to ensure input and output quality. The dataset is randomly split into 12,000 training, 1,000 validation, and 500 test examples. Detailed hyperparameters and training configurations are provided in Appendix B. Evaluation Datasets and Metrics Student model performance is evaluated on a diverse set of benchmarks, including DollyEval, SelfInst (Wang et al., 2022a), SuperNatural-Instructions (Wang et al., 2022b), UnNI (Honovich et al., 2023), and VicunaEval (Chiang et al., 2023). We use two complementary metrics: Rouge-L, which reflects sequence-level alignment with references and assesses generation quality (Lin, 2004), and BertScore (Zhang et al., 2019), which evaluates semantic similarity between generated and reference texts using contextual embeddings. Higher scores on both metrics indicate better performance. DollyEval serves as our ID benchmark as models are trained on the Dolly dataset, while other benchmarks evaluate OOD generalization. Baselines and Comparison Methods We compare against six baseline approaches: supervised fine-tuning without distillation (SFT w/o KD), sequence-level distillation (SeqKD), token-level distillation with forward or reverse KL (FKL/RKL), weighted combination of both (FKL+RKL), state-of-the-art RKL-based method (MiniLLM), and our proposed ARKD method. The specific descriptions are as follows. • SFT w/o KD Supervised fine-tuning without knowledge distillation, where the student model is directly trained on the dataset using gold references as supervision. • SeqKD Sequence-level knowledge distillation, where the student is fine-tuned on data generated by the teacher model. • FKL (RKL) Forward (Reverse) KL divergence at the token level, where the student is fine-tuned on the dataset and supervised by the teacher’s distribution at each token step. • FKL+RKL Weighted sum of forward and reverse KL divergences with a coefficient of 0.5 for each term. • MiniLLM State-of-the-art RKL-based distillation method (Gu et al., 2024) that optimizes reverse KL with policy gradient fine-tuning. • ARKD Our proposed method: joint optimization of forward and reverse KL divergences via adaptive reinforcement learning. 4.2 Main Results Model Params Method SelfInst VicunaEval S-NI UnNI GPT-2 1.5B Teacher 15.1 16.6 27.2 31.6 SFT w/o KD 11.9 12.5 16.4 18.4 SeqKD 11.6 12.2 16.3 18.5 120M FKL 12.1 12.9 16.6 19.2 RKL 12.5 13.3 17.2 19.8 FKL+RKL 13.2 13.9 17.8 20.3 ARKD 13.6 14.5 19.1 21.9 SFT w/o KD 13.0 14.4 24.9 25.7 SeqKD 13.1 14.8 25.4 26.0 340M FKL 13.8 15.2 25.7 26.6 RKL 14.4 15.4 26.1 27.1 FKL+RKL 14.6 15.9 26.3 27.8 ARKD 15.0 16.1 26.8 28.4 LLaMA 13B Teacher 23.2 19.7 36.1 39.0 SFT w/o KD 20.6 17.5 32.4 35.8 SeqKD 20.8 18.1 33.3 36.6 7B FKL 20.5 18.3 33.8 36.5 RKL 21.0 18.6 34.1 37.1 FKL+RKL 21.8 18.9 34.7 37.5 ARKD 22.6 19.2 35.5 38.1 Table 2: Performance on OOD datasets (SelfInst, VicunaEval, S-NI, and UnNI) for various teacher-student configurations, reported as Rouge-L. Bold indicates the best student model in each block. Complete results with BertScore metrics are provided in Appendix C. Based on the theoretical analysis in Section 3 and the experimental setup in Section 4.1, we conduct extensive experiments across multiple teacher-student configurations. Results are presented in Tables 1 and 2. Table 1 reports ID performance on DollyEval using both Rouge-L and BertScore, while Table 2 presents OOD results using Rouge-L. Key findings are summarized as follows: • ARKD consistently outperforms all baseline methods across teacher-student configurations and evaluation benchmarks. On DollyEval (Table 1), ARKD achieves the highest BertScore (0.815, 0.827, 0.868) and competitive or superior Rouge-L scores across all three configurations, demonstrating strong performance in both semantic similarity and lexical alignment. Notably, ARKD surpasses the strong MiniLLM baseline, validating the effectiveness of our strategy. • RKL consistently outperforms FKL, and the static combination FKL+RKL yields further improvements, corroborating the complementary nature of FKL’s mode-covering and RKL’s mode-seeking behaviors. However, ARKD’s adaptive weighting significantly outperforms the static 0.5:0.5 combination, demonstrating that dynamic policy optimization provides substantial gains over fixed weighting schemes. • ARKD shows superior out-of-distribution generalization. On non-Dolly benchmarks (SelfInst, VicunaEval, S-NI, UnNI) in Table 2, ARKD consistently outperforms all baselines across both GPT-2 and LLaMA model families. Performance improvements remain consistent as model size increases, illustrating the scalability of our approach. 4.3 Ablation Study: Impact of Alpha Weighting Strategies To evaluate the effectiveness of adaptive alpha weighting, we compare: (1) static values (α∈0,0.25,0.5,0.75,1α∈\0,0.25,0.5,0.75,1\), (2) linear scheduling (linear_dec: 1→01→ 0; linear_inc: 0→10→ 1), (3) greedy heuristic (greedy_min: min(FKL,RKL) (FKL,RKL)), and (4) RL-based ARKD. Extended ablation results including detailed analysis of linear scheduling and additional static alpha values are provided in Appendix D. Figures 3–4 reveal clear performance patterns across different alpha strategies (the corresponding BertScore heatmap is provided in Appendix D, Figure 7). Static weighting methods (FKL, RKL, fixed combinations) achieve limited performance (22.9–23.3 Rouge-L on 120M Dolly), as they cannot adapt to varying distributional characteristics during training. Linear scheduling shows minimal improvement, suggesting that simple time-based heuristics fail to capture complex alignment dynamics. The greedy_min heuristic performs better (24.1) by selecting the lower-loss divergence at each step, yet remains fundamentally myopic—optimizing for immediate gains without long-term planning. In contrast, ARKD consistently achieves best scores: 24.5 (Dolly), 19.1 (S-NI), 21.9 (UnNI) on 120M, outperforming greedy_min by 0.4–0.6 points and maintaining strong performance across both ID and OOD benchmarks. Similar patterns emerge for 340M models (Figure 4), confirming scalability. Figure 3: Rouge-L performance heatmap for GPT2-120M across Dolly (ID), S-NI, and UnNI (OOD) datasets under different alpha strategies. Figure 4: Rouge-L performance heatmap for GPT2-340M across Dolly, S-NI, and UnNI datasets. Figure 5 shows ARKD’s learned adaptation pattern across three phases. During Exploration (0–600 steps), the policy network conducts broad search over α space. In Convergence (600–2000 steps), the policy gradually transitions toward lower α values as the student distribution approaches the teacher, reflecting a shift from mode-covering (FKL) to mode-seeking (RKL) behavior. Finally, Stability (2000+ steps) maintains α≈0.18α≈ 0.18 for precise refinement. This trajectory reveals an intelligent strategy: early FKL emphasis ensures broad coverage preventing mode collapse, while later RKL emphasis enables precise alignment avoiding long-tail overestimation. This non-trivial strategy emerges through RL optimization rather than manual design, validating that policy-based learning anticipates long-term training dynamics beyond greedy optimization. Figure 5: Learned α trajectory during GPT2-120M training. The RL policy explores α∈[0.3,0.6]α∈[0.3,0.6] initially, then converges to α≈0.18α≈ 0.18 (RKL-dominant) by Epoch 3. 4.4 Goal Consistency Analysis To validate the theoretical soundness of our approach, we analyze the consistency among reward, FKL, and RKL during training. According to Section 3, all three optimization objectives converge to the same solution when the student fully matches the teacher, indicating essential equivalence. We monitor these metrics throughout training; the corresponding training curves are provided in Appendix A (Figure 6). As shown, the dynamics of all three metrics are highly consistent. As training progresses, reward steadily increases (i.e., distillation loss decreases) while both FKL and RKL decrease synchronously, indicating successful distribution alignment. Importantly, adaptive weighting does not compromise objective consistency—whether using fixed or RL-based weights, the trends of reward, FKL, and RKL remain aligned, confirming our theoretical predictions. This demonstrates that despite incorporating policy networks and adaptive weighting, ARKD maintains fundamental consistency between theory and practice, providing a solid foundation for its effectiveness across diverse tasks. 5 Conclusion This paper presents ARKD, an adaptive knowledge distillation framework that dynamically optimizes the weighting of forward and reverse KL divergences through reinforcement learning. By introducing a state-aware policy network, ARKD adaptively balances FKL’s mode-covering and RKL’s mode-seeking behaviors according to distributional properties of teacher and student models. Our theoretical analysis establishes the complementary roles and convergence consistency of both KL divergences, providing principled foundations for adaptive optimization. Extensive experiments demonstrate that ARKD consistently outperforms static weighting, scheduled strategies, and greedy heuristics across multiple benchmarks and model scales, advancing adaptive knowledge distillation for efficient large language model compression. Limitations While ARKD demonstrates strong performance across benchmarks and model scales, we view several directions as natural extensions for future work. The policy currently conditions on a hand-crafted 6-dimensional state vector with only ∼ 525 parameters; end-to-end learned encoders could capture finer distributional structure. The immediate per-batch reward could also be extended to longer-horizon signals such as validation-set generation quality or task-specific preferences. The adaptive mechanism is complementary to broader f-divergence families (e.g., TVD in f-DISTILL) and to on-policy paradigms such as GKD, both of which can be combined with ARKD. Extending it to other generation tasks and larger model scales, together with automated hyperparameter selection, are promising avenues. References Acharya et al. (2024) Kamal Acharya, Alvaro Velasquez, and Houbing Herbert Song. 2024. A survey on symbolic knowledge distillation of large language models. IEEE Transactions on Artificial Intelligence. Agarwal et al. (2024) Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. 2024. On-policy distillation of language models: Learning from self-generated mistakes. In The Twelfth International Conference on Learning Representations. Chiang et al. (2023) Wei-Lin Chiang, Zhuohan Li, Ziqing Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, and 1 others. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna.lmsys.org (accessed 14 April 2023), 2(3):6. Fang et al. (2026) Luyang Fang, Xiaowei Yu, Jiazhang Cai, Yongkai Chen, Shushan Wu, Zhengliang Liu, Zhenyuan Yang, Haoran Lu, Xilin Gong, Yufang Liu, and 1 others. 2026. Knowledge distillation and dataset distillation of large language models: Emerging trends, challenges, and future directions. Artificial Intelligence Review, 59(1):17. Gou et al. (2021) Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao. 2021. Knowledge distillation: A survey. International journal of computer vision, 129(6):1789–1819. Gu et al. (2024) Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024. Minillm: Knowledge distillation of large language models. In International Conference on Learning Representations, volume 2024, pages 32694–32717. Hinton et al. (2015) Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531. Honovich et al. (2023) Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2023. Unnatural instructions: Tuning language models with (almost) no human labor. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14409–14428. Jiang et al. (2026) Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. 2026. A survey on large language models for code generation. ACM Transactions on Software Engineering and Methodology, 35(2):1–72. Jin et al. (2026) Woogyeol Jin, Taywon Min, Yongjin Yang, Swanand Ravindra Kadhe, Yi Zhou, Dennis Wei, Nathalie Baracaldo, and Kimin Lee. 2026. Entropy-aware on-policy distillation of language models. arXiv preprint arXiv:2603.07079. Kim et al. (2024) Gyeongman Kim, Doohyuk Jang, and Eunho Yang. 2024. Promptkd: Distilling student-friendly knowledge for generative language models via prompt tuning. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 6266–6282. Kim and Rush (2016) Yoon Kim and Alexander M Rush. 2016. Sequence-level knowledge distillation. In Proceedings of the 2016 conference on empirical methods in natural language processing, pages 1317–1327. Kwon et al. (2023) Minae Kwon, Sang Michael Xie, Kalesha Bullard, and Dorsa Sadigh. 2023. Reward design with language models. arXiv preprint arXiv:2303.00001. Lee et al. (2023) Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Ren Lu, Thomas Mesnard, Johan Ferret, Colton Bishop, Ethan Hall, Victor Carbune, and Abhinav Rastogi. 2023. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267. Li et al. (2024) Yixing Li, Yuxian Gu, Li Dong, Dequan Wang, Yu Cheng, and Furu Wei. 2024. Direct preference knowledge distillation for large language models. arXiv preprint arXiv:2406.19774. Lin (2004) Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81. Nguyen et al. (2025) Duy A Nguyen, Rishi Kesav Mohan, Van Yang, Pritom Saha Akash, and Kevin Chen-Chuan Chang. 2025. Rl-based query rewriting with distilled llm for online e-commerce systems. arXiv preprint arXiv:2501.18056. Tan et al. (2023) Shicheng Tan, Weng Lam Tam, Yuanchun Wang, Wenwen Gong, Shu Zhao, Peng Zhang, and Jie Tang. 2023. Gkd: A general knowledge distillation framework for large-scale pre-trained language model. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Industry Track), pages 134–148. Wang et al. (2022a) Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022a. Self-instruct: Aligning language models with self-generated instructions. arXiv preprint arXiv:2212.10560. Wang et al. (2022b) Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, and 1 others. 2022b. Benchmarking generalization via in-context instructions on 1,600+ language tasks. arXiv preprint arXiv:2204.07705, 2(2). Wen et al. (2023) Yuqiao Wen, Zichao Li, Wenyu Du, and Lili Mou. 2023. F-divergence minimization for sequence-level knowledge distillation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10817–10834. Wu et al. (2025) Taiqiang Wu, Chaofan Tao, Jiahao Wang, Runming Yang, Zhe Zhao, and Ngai Wong. 2025. Rethinking kullback-leibler divergence in knowledge distillation for large language models. In Proceedings of the 31st International Conference on Computational Linguistics, pages 5737–5755. Xu et al. (2026) Shicheng Xu, Liang Pang, Yunchang Zhu, Jia Gu, Zihao Wei, Jingcheng Deng, Feiyang Pan, Huawei Shen, and Xueqi Cheng. 2026. Rlkd: Distilling llms’ reasoning via reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 34151–34159. Xu et al. (2024) Xiaohan Xu, Ming Li, Chongyang Tao, Tao Shen, Reynold Cheng, Jinyang Li, Can Xu, Dacheng Tao, and Tianyi Zhou. 2024. A survey on knowledge distillation of large language models. arXiv preprint arXiv:2402.13116. Yang et al. (2024a) Chuanpeng Yang, Yao Zhu, Wang Lu, Yidong Wang, Qian Chen, Chenlong Gao, Bingjie Yan, and Yiqiang Chen. 2024a. Survey on knowledge distillation for large language models: methods, evaluation, and application. ACM Transactions on Intelligent Systems and Technology. Yang et al. (2024b) Kevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng, and Yuandong Tian. 2024b. Rlcd: Reinforcement learning from contrastive distillation for lm alignment. In The Twelfth International Conference on Learning Representations. Yang et al. (2023) Linyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu, Yidong Wang, Jingming Zhuo, Lingqiao Liu, Jindong Wang, Jennifer Foster, and Yue Zhang. 2023. Out-of-distribution generalization in natural language processing: Past, present, and future. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4533–4559. Zhang et al. (2024a) Chaoyun Zhang, Shilin He, Jiaxu Qian, Bowen Li, Liqun Li, Si Qin, Yu Kang, Minghua Ma, Guyue Liu, Qingwei Lin, and 1 others. 2024a. Large language model-brained gui agents: A survey. arXiv preprint arXiv:2411.18279. Zhang et al. (2024b) Songming Zhang, Xue Zhang, Zengkui Sun, Yufeng Chen, and Jinan Xu. 2024b. Dual-space knowledge distillation for large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 18164–18181. Zhang et al. (2019) Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675. Zhang et al. (2026) Zhaoyang Zhang, Shuli Jiang, Yantao Shen, Yuting Zhang, Dhananjay Ram, Shuo Yang, Zhuowen Tu, Wei Xia, and Stefano Soatto. 2026. Reinforcement-aware knowledge distillation for llm reasoning. arXiv preprint arXiv:2602.22495. Zhong et al. (2024) Qihuang Zhong, Liang Ding, Li Shen, Juhua Liu, Bo Du, and Dacheng Tao. 2024. Revisiting knowledge distillation for autoregressive language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10900–10913. Zhu et al. (2024) Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang. 2024. A survey on model compression for large language models. Transactions of the Association for Computational Linguistics, 12:1556–1577. Appendix A Training Dynamics Figure 6 plots three key training-time signals over optimization steps: the policy reward (Rouge-L improvement on the validation subset), the FKL loss, and the RKL loss. The reward curve rises steadily while both KL terms decrease in sync, empirically confirming the goal-consistency analysis in Section 3.1: the policy gradient signal does not trade one KL objective off against the other but instead drives the student to better match the teacher under both directions simultaneously. (a) Reward (b) FKL loss (c) RKL loss Figure 6: Training dynamics: reward steadily increases while FKL and RKL decrease synchronously, validating the goal-consistency analysis in Section 3.1. Appendix B Experimental Configuration Details This section provides comprehensive implementation details to ensure reproducibility of our experimental results. B.1 Model Architecture and Training Hyperparameters Teacher and Student Models We use pre-trained GPT-2 and LLaMA models from Hugging Face. Teacher models (GPT-2 1.5B, LLaMA 13B) are first fine-tuned on the Dolly-15K dataset, then used for knowledge distillation to student models (GPT-2 120M, GPT-2 340M, LLaMA 7B). Training Hyperparameters All experiments use the following hyperparameters: • Maximum sequence length: 512 tokens • Batch size: 16 (per GPU) • Gradient accumulation steps: 1 • Training epochs: 3 • Student learning rate: 5×10−55× 10^-5 (AdamW optimizer) • Policy network learning rate: 1×10−41× 10^-4 (Adam optimizer) • Temperature for KL distillation: 2.0 • Language modeling loss weight (λPT _PT): 0.1 • Gradient clipping: max norm 1.0 B.2 Policy Network Architecture The policy network πϕ _φ consists of a 3-layer feedforward network. The input state vector s is a 6-dimensional feature representation: • HT,HSH_T,H_S: mean entropy of teacher and student distributions • σT2,σS2 _T^2, _S^2: mean variance of teacher and student distributions • LFKL,LRKLL_FKL,L_RKL: current forward and reverse KL losses Table 3 presents the detailed architecture of the policy network. Layer Operation Input Dim Output Dim Parameters Input State features - 6 - Layer 1a LayerNorm 6 6 12 Layer 1b Linear 6 64 448 Layer 1c ReLU 64 64 0 Layer 2a Linear 64 1 65 Layer 2b Sigmoid 1 1 0 Output α weight 1 1 - Total Parameters 525 Table 3: Detailed architecture of the policy network πϕ _φ. The network takes a 6-dimensional state vector as input and outputs a scalar weight α∈(0,1)α∈(0,1) for adaptive KL weighting. Total parameter count: 525 (LayerNorm: 2×6=122× 6=12, Linear 1: (6+1)×64=448(6+1)× 64=448, Linear 2: (64+1)×1=65(64+1)× 1=65). The policy is trained using REINFORCE with EMA baseline and entropy regularization: Lpolicy=−log(α)⋅(r−b)−βHH(α)L_policy=- (α)·(r-b)- _H\,H(α) (19) where r is the immediate reward (negative distillation loss), b is the exponential moving average baseline with decay factor 0.99, and H(α)=−[αlogα+(1−α)log(1−α)]H(α)=-[α α+(1-α) (1-α)] is the entropy term with coefficient βH=0.01 _H=0.01. B.3 Evaluation Configuration Generation Settings For all evaluation benchmarks, we use nucleus sampling with p=0.9p=0.9, temperature T=1.0T=1.0, and maximum generation length of 256 tokens. Early stopping at EOS token is enabled. Metrics • Rouge-L: Computed using the rouge-score library, measuring longest common subsequence overlap between generated and reference texts. • BertScore: Computed using bert-score library with microsoft/deberta-xlarge-mnli as the reference model, measuring semantic similarity via contextual embeddings. Appendix C Complete Out-of-Distribution Results with BertScore Table 4 presents the complete out-of-distribution evaluation results including both Rouge-L and BertScore metrics across all four OOD benchmarks (SelfInst, VicunaEval, S-NI, UnNI) for three model configurations. These results complement Table 2 in the main paper, which reported only Rouge-L scores due to space constraints. Model Method SelfInst VicunaEval S-NI UnNI R-L BertS R-L BertS R-L BertS R-L BertS GPT-2 1.5B → GPT-2 120M Teacher 15.1 0.810 16.6 0.815 27.2 0.830 31.6 0.835 SFT w/o KD 11.9 0.770 12.5 0.768 16.4 0.775 18.4 0.778 SeqKD 11.6 0.768 12.2 0.766 16.3 0.773 18.5 0.776 FKL 12.1 0.775 12.9 0.773 16.6 0.780 19.2 0.783 RKL 12.5 0.780 13.3 0.778 17.2 0.785 19.8 0.788 FKL+RKL 13.2 0.788 13.9 0.786 17.8 0.793 20.3 0.796 ARKD 13.6 0.800 14.5 0.798 19.1 0.805 21.9 0.808 GPT-2 1.5B → GPT-2 340M Teacher 15.1 0.810 16.6 0.815 27.2 0.830 31.6 0.835 SFT w/o KD 13.0 0.785 14.4 0.783 24.9 0.790 25.7 0.793 SeqKD 13.1 0.787 14.8 0.786 25.4 0.793 26.0 0.795 FKL 13.8 0.793 15.2 0.791 25.7 0.797 26.6 0.800 RKL 14.4 0.798 15.4 0.796 26.1 0.802 27.1 0.805 FKL+RKL 14.6 0.802 15.9 0.800 26.3 0.806 27.8 0.809 ARKD 15.0 0.808 16.1 0.806 26.8 0.811 28.4 0.815 LLaMA 13B → LLaMA 7B Teacher 23.2 0.852 19.7 0.848 36.1 0.865 39.0 0.870 SFT w/o KD 20.6 0.821 17.5 0.818 32.4 0.832 35.8 0.836 SeqKD 20.8 0.823 18.1 0.821 33.3 0.835 36.6 0.839 FKL 20.5 0.822 18.3 0.822 33.8 0.837 36.5 0.840 RKL 21.0 0.826 18.6 0.825 34.1 0.840 37.1 0.843 FKL+RKL 21.8 0.831 18.9 0.828 34.7 0.843 37.5 0.846 ARKD 22.6 0.838 19.2 0.832 35.5 0.849 38.1 0.852 Table 4: Complete out-of-distribution evaluation results with both Rouge-L (R-L) and BertScore (BertS) metrics. ARKD consistently achieves the best performance across all benchmarks and model configurations in both metrics, demonstrating superior semantic alignment and lexical quality. Bold indicates the best student model in each configuration. Key observations: • ARKD achieves consistent improvements in both Rouge-L and BertScore across all four OOD benchmarks, validating its superior generalization capability beyond the training distribution. • The dual-metric evaluation reveals that ARKD excels in both semantic similarity (BertScore) and lexical alignment (Rouge-L), indicating balanced distribution matching that captures both high-level meaning and surface-level patterns. • BertScore improvements are particularly pronounced, with ARKD achieving 0.800+ scores on GPT-2 120M and 0.838 on LLaMA 7B, suggesting that adaptive KL weighting effectively captures semantic nuances beyond token-level matching. Appendix D Extended Ablation Studies This section presents additional ablation experiments that further validate the necessity of reinforcement learning-based adaptive weighting over various alternative strategies. Figure 7: BertScore performance heatmap for GPT2-120M across Dolly, S-NI, and UnNI datasets under different alpha strategies. D.1 Linear Scheduling Strategies While Section 4.1 presents performance heatmaps visualizing the effectiveness of various alpha strategies across multiple datasets, Table 5 complements these findings by providing complete numerical results for linear scheduling methods, including BertScore values for the 340M configuration which were not fully covered in the main text figures. Method GPT-2 120M GPT-2 340M Rouge-L BertScore Rouge-L BertScore FKL (α≡1.0α≡ 1.0) 22.9 0.781 24.9 0.803 RKL (α≡0.0α≡ 0.0) 23.3 0.797 25.5 0.813 FKL+RKL (α≡0.5α≡ 0.5) 23.2 0.794 25.7 0.817 linear_dec (1→0) 23.0 0.785 25.2 0.808 linear_inc (0→1) 23.0 0.788 25.1 0.806 greedy_min 24.1 0.808 25.9 0.822 ARKD 24.5 0.815 26.1 0.827 Table 5: Complete numerical results for linear scheduling strategies on DollyEval, including BertScore for GPT-2 340M. Linear scheduling methods (linear_dec: α 1→0; linear_inc: α 0→1) adjust weights according to predefined temporal patterns. Cross-metric Consistency Analysis Examining the dual-metric results reveals that linear scheduling exhibits consistent underperformance across both Rouge-L and BertScore. For 120M models, both linear_dec and linear_inc achieve identical Rouge-L scores (23.0) but show slight BertScore variation (0.785 vs 0.788), suggesting that starting with RKL provides marginally better semantic alignment despite similar lexical quality. However, this gap is negligible compared to ARKD’s gains. Scale-dependent Behavior The 340M results demonstrate that the limitations of linear scheduling persist across model scales. While all methods benefit from increased capacity (340M scores are uniformly higher than 120M), the relative performance gap remains: linear scheduling still underperforms ARKD by 0.9–1.0 Rouge-L points and 0.019–0.021 BertScore points. Notably, linear_dec (25.2) outperforms linear_inc (25.1) on 340M, suggesting that larger models may benefit slightly from early FKL emphasis—yet this effect is marginal and cannot approach ARKD’s adaptive weighting. Why Linear Schedules Fail Unlike ARKD’s state-aware policy which conditions α on real-time distributional features (entropy, variance, FKL/RKL magnitudes), linear schedules rely solely on elapsed time. This fundamental limitation manifests in three ways: 1. Invariance to training dynamics: Linear schedules cannot detect critical transitions (e.g., when the student begins overfitting teacher’s long-tail modes) and adjust accordingly. 2. Assumption of monotonicity: Both schedules assume optimal weighting changes monotonically over time, contradicting ARKD’s learned non-monotonic trajectory which plateaus after convergence (Figure 5). 3. No feedback mechanism: Without observing immediate reward signals, linear schedules cannot correct suboptimal trajectories mid-training, whereas ARKD’s policy gradient enables continuous refinement based on distillation loss feedback. D.2 Static Alpha Values To further investigate the sensitivity of performance to α values, we evaluate two additional static configurations (α=0.25α=0.25 and α=0.75α=0.75) on GPT-2 120M: Method Alpha Rouge-L BertScore FKL 1.0 22.9 0.781 α=0.75α=0.75 0.75 23.0 0.789 FKL+RKL 0.5 23.2 0.794 α=0.25α=0.25 0.25 23.5 0.799 RKL 0.0 23.3 0.797 greedy_min dynamic (→0) 24.1 0.808 ARKD dynamic (→0.18) 24.5 0.815 Table 6: Performance across static alpha values on DollyEval (GPT-2 120M). While α=0.25α=0.25 performs best among static values, it still falls short of ARKD’s adaptive strategy. Results show that α=0.25α=0.25 (RKL-dominant with 25% FKL) achieves the best performance (23.5 Rouge-L) among static values, closely matching ARKD’s converged value (α≈0.18α≈ 0.18). However, ARKD still outperforms by 1.0 Rouge-L point, demonstrating that: • The optimal α is neither 0.5 (as assumed by FKL+RKL) nor purely 0/1, but lies in the RKL-dominant range with a small FKL component. • Even knowing the optimal α value post-hoc, using it as a static weight throughout training is suboptimal. ARKD’s advantage comes from trajectory-level optimization: exploring higher α values early in training (preventing mode collapse) before converging to lower values for precise refinement. • This validates that the temporal dynamics of weighting—not just the final value—contribute to ARKD’s effectiveness, a capability that static methods fundamentally cannot provide. D.3 Why RL Outperforms Greedy Heuristics As discussed in Section 4.1, the greedy_min heuristic (α=[FKL<RKL]α=I[FKL<RKL], selecting the lower-loss divergence at each step) provides a strong myopic baseline. Our experiments reveal that greedy_min converges to α≈0α≈ 0 throughout training, effectively degenerating to pure RKL due to FKL consistently being larger than RKL in magnitude. Despite this degeneration, greedy_min (24.1 Rouge-L) still outperforms pure RKL (23.3), which we attribute to implicit regularization effects from the min-selection dynamics. However, ARKD surpasses greedy_min by 0.4 points on 120M and 0.2 points on 340M. This improvement stems from: 1. Continuous mixing: ARKD outputs α∈(0.15,0.25)α∈(0.15,0.25) rather than hard-switching to 0, preserving a small but crucial FKL component that provides mode-coverage regularization. 2. Forward-looking optimization: Greedy selection optimizes immediate loss reduction, while ARKD’s policy gradient maximizes cumulative rewards, enabling anticipation of long-term training dynamics. 3. Exploration-exploitation: ARKD explores higher α values (0.3–0.6) during the first 600 steps before converging, while greedy_min commits to α=0α=0 immediately, potentially missing beneficial early-stage FKL emphasis. This ablation definitively demonstrates that RL-based adaptation provides tangible benefits over strong non-RL baselines, validating the necessity of our reinforcement learning framework.