Paper deep dive
EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning
Can Xie, Yuyi Zhou, Wen Yang, Ziyi zhang, Siyao Song, Yingzhuo Deng, Shuo Ren, Jiajun Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/25/2026, 7:45:14 AM
Summary
The paper introduces EDGE (Experience-Distillation for Guided Exploration), a framework for agentic reinforcement learning that treats external experiences as temporary training-time scaffolds rather than persistent inference-time memory. EDGE partitions rollouts into experience-conditioned and experience-free trajectories to estimate marginal gains, using a gain-gated reverse-KL distillation to internalize beneficial behaviors into the policy. A co-evolutionary experience bank manages the lifecycle of experiences by synthesizing new ones from failure modes and pruning obsolete entries. Experiments on ALFWorld and WebShop show EDGE outperforms GRPO and other experience-augmented methods, retaining 96% of performance without external retrieval at inference.
Entities (9)
Relation Signals (6)
EDGE → evaluatedon → ALFWorld
confidence 95% · On ALFWorld and WebShop, EDGE improves over GRPO
EDGE → evaluatedon → WebShop
confidence 95% · On ALFWorld and WebShop, EDGE improves over GRPO
EDGE → improvesover → GRPO
confidence 95% · EDGE improves over GRPO by 8.3 and 12.5 success-rate points at the 7B scale
EDGE → uses → Experience Bank
confidence 95% · A co-evolutionary experience bank further synthesizes guidance from emerging failure modes
EDGE → outperforms → SkillRL
confidence 90% · EDGE retains 96.0% of its scaffolded performance when external experiences are removed... compared with 82.9% for SkillRL
EDGE → uses → Reverse-KL Divergence
confidence 90% · distills the induced behavior into the base policy via a reverse-KL objective
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy update. Existing experience-augmented approaches retrieve historical guidance at inference time, but they apply experiences without accounting for the policy's evolving capability and create persistent dependencies on external retrieval. We propose EDGE (Experience-Distillation for Guided Exploration), a framework that treats retrieved experiences as temporary training-time scaffolds and progressively internalizes their benefits into the parametric policy. Concretely, EDGE partitions each rollout group into experience-conditioned and experience-free trajectories to estimate and admit only positive marginal gains without extra sampling, then distills the induced behavior into the base policy via a reverse-KL objective on its own empirical support. A co-evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries as the policy evolves. On ALFWorld and WebShop, EDGE improves over GRPO by 8.3 and 12.5 success-rate points at the 7B scale and retains 96.0% of its scaffolded performance when external experiences are removed at inference time. The code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.21946v1
- Canonical: https://arxiv.org/abs/2608.21946v1
Trouble viewing inline? Open PDF directly →
Full Text
121,041 characters extracted from source content.
Expand or collapse full text
EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning Can Xie Affiliation: [5pt] School of Artificial Intelligence, University of Chinese Academy of Sciences Affiliation: Institute of Automation, Chinese Academy of Sciences Yuyi Zhou Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences Wen Yang Affiliation: [5pt] School of Artificial Intelligence, University of Chinese Academy of Sciences Affiliation: Institute of Automation, Chinese Academy of Sciences Ziyi Zhang Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences Siyao Song, Yingzhuo Deng, Shuo Ren, Jiajun Zhang11footnotemark: 1 Thanks: Corresponding authors Affiliation: [5pt] School of Artificial Intelligence, University of Chinese Academy of Sciences Affiliation: [5pt] School of Artificial Intelligence, University of Chinese Academy of Sciences Affiliation: [5pt] School of Artificial Intelligence, University of Chinese Academy of Sciences Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Wuhan AI Researchxiecan2024,shuo.ren@ia.ac.cn, jjzhang@nlpr.ia.ac.cn Abstract Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy update. Existing experience-augmented approaches retrieve historical guidance at inference time, but they apply experiences without accounting for the policy’s evolving capability and create persistent dependencies on external retrieval. We propose EDGE (Experience-Distillation for Guided Exploration), a framework that treats retrieved experiences as temporary training-time scaffolds and progressively internalizes their benefits into the parametric policy. Concretely, EDGE partitions each rollout group into experience-conditioned and experience-free trajectories to estimate and admit only positive marginal gains without extra sampling, then distills the induced behavior into the base policy via a reverse-KL objective on its own empirical support. A co-evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries as the policy evolves. On ALFWorld and WebShop, EDGE improves over GRPO by 8.3 and 12.5 success-rate points at the 7B scale and retains 96.0% of its scaffolded performance when external experiences are removed at inference time. The code is available at https://github.com/xvolcano02/EDGE. Figure 1: Overview of the EDGE framework. (a) Experience-guided exploration scaffolding uses contrastive rollouts to estimate the marginal gain Δe _e of a retrieved experience. (b) Gain-gated distillation internalizes beneficial scaffold behavior into the base policy only when Δe>0 _e>0. (c) The experience bank co-evolves with the policy via failure-driven expansion and utility-based pruning, yielding a retrieval-free policy at deployment. 1 Introduction Reinforcement learning (RL) with outcome-based objectives such as GRPO (24) has become a standard post-training paradigm for LLM-based agents that reason, plan, and act through multi-turn interactions (38; 25; 37; 26; 9; 31; 11). Yet current agentic RL uses its own experience poorly. A single rollout may contain reusable patterns—effective task decompositions, recovery strategies, dead-end avoidance, but once a trajectory contributes a scalar advantage to one policy update, these patterns are largely discarded. The agent must therefore rediscover them from scratch, a particularly costly failure mode in sparse-reward, long-horizon environments where agentic RL is most needed. A growing body of work addresses this inefficiency by augmenting agents with external experience at inference time, through episodic reflections (25; 45), persistent memory (4; 8), or skill libraries (34; 13). While effective, these approaches share two structural limitations that become apparent as training progresses. Experience utility is inherently policy-dependent: guidance that accelerates an undertrained policy can become redundant---or actively harmful---once the corresponding behavior has been internalized, yet most methods apply experience unconditionally or filter it by static heuristics. Moreover, if the agent must retrieve experience at deployment, part of its competence resides in the context window rather than in the model parameters, incurring persistent token overhead and sensitivity to retrieval noise. 11 1 These are not hypothetical concerns: in our experiments, naive memory augmentation (e.g., EvolveR, GRPO+Mem0) degrades performance well below vanilla GRPO, confirming that unfiltered experience injection is unreliable. These observations suggest that experience reuse in agentic RL should be treated as a dynamic lifecycle rather than a static retrieval mechanism. External experience could first guide exploration, as a temporary training-time scaffold, then be exploited by consolidating its useful behavioral effect into the experience-free policy, and finally be retired once its utility vanishes. Under this view, experience reuse becomes a scaffold-to-parameter learning problem rather than a retrieve-and-prompt augmentation problem. In this paper, we instantiate this idea in EDGE (Experience-Distillation for Guided Exploration), a framework that turns retrieved experience from persistent inference-time memory into dynamically validated training-time scaffolding. EDGE uses online marginal-gain estimation to decide when an experience should guide exploration, be distilled into the policy, or be retired as the policy evolves, via three core mechanisms. First, experience-guided exploration scaffolding partitions each GRPO rollout group into experience-conditioned and experience-free trajectories, estimating the marginal gain of a retrieved experience without additional environment sampling and retaining only positive-gain instances. Then, gain-gated privileged distillation serves as the exploitation step, transferring the behavior induced by the privileged context into the standard policy via reverse-KL divergence computed on the student’s own empirical support, activating only when the estimated gain is positive to avoid negative transfer. Finally, experience bank evolution and management enables the experience bank to co-evolve with the policy throughout training: new experiences are synthesized from failure-mode analysis, while obsolete ones are pruned based on tracked utility, keeping the scaffold aligned with the agent’s evolving capability. On ALFWorld (26) and WebShop (37), EDGE consistently outperforms skill-free RL and prior experience-augmented methods, with the largest gains on exploration-intensive subtasks (e.g., Heat, Cool, Pick2). When external experiences are withheld at inference time, EDGE retains 96.0% of its scaffolded performance—compared with 82.9% for SkillRL and 92.3% for EMPO2—confirming that useful exploration priors have been absorbed into the parametric policy. Our contributions are as follows: • We identify two failure modes of experience-augmented agentic RL—policy-dependent experience utility and persistent inference-time retrieval dependence—and reframe experience reuse as a dynamic scaffold-to-parameter transition. • We introduce experience-guided exploration scaffolding, which partitions rollout groups to estimate the marginal value of each retrieved experience under the current policy without requiring additional environment sampling. • We propose gain-gated privileged distillation, which internalizes scaffold-induced behavior via reverse-KL on the student’s own empirical support, gated by the estimated marginal gain to prevent negative transfer, together with a co-evolutionary experience bank that expands and prunes in response to the policy’s evolving needs. • Experiments on ALFWorld and WebShop demonstrate that EDGE improves task performance and training efficiency while enabling robust deployment without external experience retrieval. 2 Preliminaries We formalize the multi-turn decision-making process of an LLM-based agent as a Partially Observable Markov Decision Process (POMDP) (12) ⟨,,,,ℛ⟩ ,A,O,T,R . The observation space O consists of natural-language strings emitted by the environment, and the action space A comprises token sequences generated autoregressively by the policy πθ _θ. Given a task instruction x, the agent interacts with the environment over a sequence of turns: at step t it receives observation ot∈o_t and produces action at∈a_t according to at∼πθ(⋅∣x,ht,ot),a_t _θ(· x,h_t,o_t), (1) where ht=(o1,a1,…,ot−1,at−1)h_t=(o_1,a_1,…,o_t-1,a_t-1) is the interaction history. An episode terminates upon task completion or at a maximum step limit, yielding a sparse binary reward R(τ)∈0,1R(τ)∈\0,1\ and a full trajectory τ=(x,o1,a1,…,oT,aT,R(τ)).τ=(x,\;o_1,a_1,\;…,\;o_T,a_T,\;R(τ)). (2) We build on Group Relative Policy Optimization (GRPO) (24), which samples a group of G parallel trajectories τii=1G\ _i\_i=1^G per task and computes group-normalized advantages: A^i=R(τi)−mean(R(τj)j=1G)std(R(τj)j=1G). A_i= R( _i)-mean(\R( _j)\_j=1^G)std(\R( _j)\_j=1^G). (3) The policy is updated by maximizing a clipped surrogate objective with KL regularization: ℒRL(θ)=−[1G∑i=1G1|τi|∑t=1|τi|(min(ρi,tA^i, _RL(θ)=-E\! [ 1G _i=1^G 1| _i| _t=1^| _i| ( \! ( _i,t\, A_i,\; clip(ρi,t,−ϵ,+ϵ)A^i)−βDKL(πθ∥πref))], ( _i,t,1\!-\!ε,1\!+\!ε)\, A_i )-β\,D_KL ( _θ\| _ref ) ) ], (4) where ρi,t=πθ(at∣x,ht,ot)/πθold(at∣x,ht,ot) _i,t= _θ(a_t x,h_t,o_t)\,/\, _ _old(a_t x,h_t,o_t) is the importance sampling ratio, πref _ref is the reference policy, and β controls KL penalty strength. By contrasting outcomes within each group, GRPO steers θ toward successful action sequences without a learned value function. However, in sparse-reward, partially observable environments, unguided exploration often yields groups in which few or no trajectories succeed, rendering the advantage estimate uninformative—a limitation we address in the following section. 3 Method: EDGE We present EDGE (Experience-Distillation for Guided Exploration), a framework that iteratively strengthens the multi-turn reasoning capability of LLM-based agents (Figure 1). Central to our approach is treating retrieved external experience not as a static inference-time prompt—which inflates the context window and creates persistent retrieval dependence—but as a training-time scaffold that guides exploration and is discarded at deployment. The agent first explores with privileged access to experience, then distills only the empirically beneficial behavior into its own parameters, so the deployed policy requires no external scaffold. Section 3.1 introduces the experience-guided exploration scaffolding, Section 3.2 details the gain-gated privileged distillation mechanism, and Section 3.3 describes the experience bank evolution and management strategy. 3.1 Experience-Guided Exploration Scaffolding To provide exploratory guidance in the sparse-reward, partially observable environments typical of agentic tasks, EDGE introduces a controlled information asymmetry within the standard GRPO rollout group. Given a task instruction x and an experience bank ℰE, we retrieve the top-m most relevant experiences by embedding similarity, using the task instruction and initial environmental observations as the query, and select the highest-scoring entry e∈ℰe . Recall that GRPO samples a group of G parallel trajectories per task to estimate relative advantages (Eq. (3)). We partition this group into two equal subsets without adding extra rollouts: • Teacher rollouts (T T, G/2G/2 trajectories): conditioned on the privileged context c=x⊕ec T=x e, where ⊕ denotes concatenation under a unified chat template. • Student rollouts (T S, G/2G/2 trajectories): conditioned on the standard context c=xc S=x alone. Because both subsets share the same policy πθ _θ and differ only in whether the retrieved experience is visible, any performance gap can be attributed to the informational advantage provided by e. We quantify this gap via the instantaneous marginal gain: Δe=1||∑τ∈R(τ)−1||∑τ∈R(τ), _e\;=\; 1|T T| _τ TR(τ)\;-\; 1|T S| _τ SR(τ), (5) where R(τ)R(τ) is the binary outcome reward defined in §2. A positive Δe _e signals that the experience provides useful guidance beyond the agent’s current capability; a non-positive value indicates it is redundant or harmful. This estimate gates the distillation objective (§3.2) and drives experience bank updates (§3.3). Beyond gating distillation, Δe _e controls which rollouts enter the RL loss itself. When Δe>0 _e>0, advantages (Eq. (3)) are computed over all G trajectories. The pooled baseline therefore reflects the performance level achievable under the scaffold, giving student rollouts a stronger calibrated reference than a within-subset baseline would provide. When Δe≤0 _e≤ 0, teacher-conditioned trajectories are excluded from the RL loss entirely: advantages are computed over T S alone, reducing the update to a standard GRPO step. Without this masking, non-beneficial teacher rollouts would still shift the group baseline and distort student advantages—contaminating the policy gradient even though distillation is gated off. 3.2 Gain-Gated Privileged Distillation The scaffolding in §3.1 exposes the agent to successful reasoning patterns it may not discover on its own, but because the retrieved experience is unavailable at deployment, these gains remain ephemeral unless consolidated into the policy itself. This motivates a complementary mechanism: distilling the scaffold-induced improvements into the base policy parameters so that the agent can reproduce them from the standard context alone. Our approach realizes this through asymmetric self-distillation. Rather than relying on a separate, unconditionally superior teacher, we use the same policy πθ _θ under two informational conditions: the teacher is πθ _θ evaluated with the privileged context c T (gradients stopped), while the student operates under c S. The asymmetry therefore lies in the information available to each role, not in model capacity. The gain gate (Δe>0)I( _e>0) further ensures that distillation is triggered only when the scaffold yields a verified performance advantage, allowing the policy to selectively internalize reusable reasoning patterns while avoiding negative transfer. To guard against covariate shift, we construct the distillation target on the student’s own empirical support. For a gain-gated task instance (Δe>0 _e>0) and a student trajectory τ∈τ S with token sequence (y1,…,ym)(y_1,…,y_m) generated under c S, we perform a no-gradient forward pass of πθ _θ over the same tokens conditioned on c T. Because both passes share identical actions, the comparison isolates the informational advantage of e without introducing out-of-distribution transitions. We adopt the reverse KL divergence DKL(πstudent∥πteacher)D_KL( _student\| _teacher) as the distillation objective. Unlike the forward KL, which compels the student to cover the full teacher support and can induce mode-covering artifacts, the reverse KL is mode-seeking: it encourages the policy to concentrate on the most effective reasoning mode under the scaffold. The token-level scaffold internalization loss is: ℒdistill(θ) _distill(θ) =τ∼[(Δe>0) =E_τ S\! [I( _e\!>\!0) ∑t=1m∑y∈πθ(y)δt(y)], _t=1^m _y _θ S(y)\, _t(y) ], δt(y) _t(y) =logπθ(y∣c<t)πsg(θ)(y∣c<t), = _θ(y c S_<t) _sg(θ)(y c T_<t), (6) where πθ(y)≜πθ(y∣c<t) _θ S(y) _θ(y c S_<t) for brevity, δt(y) _t(y) is the per-token log-ratio between the student and the privileged teacher, and sg(⋅)sg(·) denotes stop-gradient. This loss is combined with the GRPO objective (Eq. (2)) to form the joint actor loss: ℒactor(θ)=ℒRL(θ)+λℒdistill(θ),L_actor(θ)=L_RL(θ)+λ\,L_distill(θ), (7) where λ controls the relative weight of scaffold internalization versus the RL signal. Through this joint optimization, the deployed agent reproduces privileged reasoning without any external scaffold at test time. 3.3 Experience Bank Evolution and Management A static experience library cannot adequately support the agent throughout training: as πθ _θ improves, it encounters situations where existing experiences provide insufficient guidance; conversely, previously beneficial experiences may become redundant once the policy has internalized the corresponding behavior, or harmful if they conflict with newly discovered strategies. EDGE therefore treats ℰE as a living repository that co-evolves with the policy through both expansion and pruning. Following 34, we generate new experiences from the agent’s own rollout trajectories. After each training step, we identify task categories whose success rate falls below a threshold ξ and collect representative failed and successful trajectories from these categories. An external reflector LLM then analyzes the contrast between them to synthesize new experiential guidance: enew=freflect(τ+,τ−),e_new=f_reflect\! (τ^+,\,τ^- ), (8) where τ+τ^+ and τ−τ^- denote a successful and a failed trajectory from the same task category, respectively. The generated experiences are deduplicated against existing entries and inserted into ℰE for retrieval in subsequent training iterations. Meanwhile, the utility of existing experiences is continuously tracked via an Exponential Moving Average (EMA) score for each experience e: Ue(t)=(1−μ)Ue(t−1)+μΔe,U_e^(t)=(1-μ)\,U_e^(t-1)+μ\, _e, (9) where μ∈(0,1)μ∈(0,1) is the momentum coefficient and Δe _e is the instantaneous marginal gain from Eq. (5). The EMA smooths over stochastic fluctuations while remaining responsive to genuine shifts in experience utility. Experiences whose tracked utility Ue(t)U_e^(t) falls below a threshold η are removed from ℰE, retiring scaffolds that have been absorbed into the policy and reducing overhead during retrieval and rollout. This yields a co-evolutionary dynamic: the policy improves by internalizing useful experiences, exposing new failure modes that drive bank expansion, while obsolete experiences are simultaneously pruned away. Type Method ALFWorld WebShop Pick Look Clean Heat Cool Pick2 All Score Succ. Closed-Source Model Prompting GPT-4o 75.3 60.8 31.2 56.7 21.6 49.8 48.0 31.8 23.7 Prompting Gemini-2.5-Pro 92.8 63.3 62.1 69.0 26.6 58.7 60.3 42.5 35.9 Qwen2.5-1.5B-Instruct Prompting Base Model 5.9 5.5 3.3 9.7 4.2 0.0 4.1 23.1 5.2 Prompting ReAct 17.4 20.5 15.7 6.2 7.7 2.0 12.8 40.1 11.3 Prompting Reflexion 35.3 22.2 21.7 13.6 19.4 3.7 21.8 55.8 21.9 Post-Training GRPO 85.3 53.7 84.5 78.2 59.7 53.5 72.8 75.8 56.8 Post-Training SkillRL 91.2±4.3 64.3±4.6 78.1±5.4 76.9±6.3 70.8±6.5 54.6±6.1 74.2±4.7 77.3±3.5 60.9±3.7 Post-Training EMPO2 86.9±3.3 66.2±5.2 79.3±4.9 79.1±4.5 75.3±5.7 64.8±6.6 76.8±3.7 78.2±3.7 63.2±3.3 Post-Training EDGE 80.6±2.2 73.7±5.5 76.0±4.3 87.3±4.0 85.6±6.1 68.4±5.6 79.7±2.3 78.8±2.1 65.6±3.2 Qwen2.5-7B-Instruct Prompting Base Model 33.4 21.6 19.3 6.9 2.8 3.2 14.8 26.4 7.8 Prompting ReAct 48.5 35.4 34.3 13.2 18.2 17.6 31.2 46.2 19.5 Prompting Reflexion 62.0 41.6 44.9 30.9 36.3 23.8 42.7 58.1 28.8 Post-Training EvolveR† 64.9 33.3 46.4 13.3 33.3 33.3 43.8 42.5 17.6 Post-Training GRPO+Mem0† 78.1 54.8 56.1 31.0 65.0 26.9 54.7 58.1 37.5 Post-Training OPSD† 50.0 60.0 22.7 21.4 17.6 9.5 32.8 4.5 2.3 Post-Training GRPO 92.8 85.7 89.3 75.7 74.5 67.7 82.1 80.3 70.1 Post-Training GRPO+OPSD† 91.4 61.5 100 87.5 76.5 52.2 80.4 86.8 76.5 Post-Training SkillRL 94.1±2.4 83.3±4.3 88.4±3.6 85.6±5.1 90.2±5.6 78.6±4.9 88.2±3.7 86.2±3.1 76.7±3.3 Post-Training EMPO2 93.3±3.7 88.9±3.9 91.2±4.4 88.5±5.3 89.4±4.6 79.8±5.1 89.1±2.8 88.3±2.6 77.1±4.0 Post-Training EDGE 96.0±1.9 85.1±4.8 93.4±3.2 90.0±2.6 92.9±4.5 81.6±6.0 90.4±2.2 89.6±2.8 82.6±3.8 Table 1: Performance on ALFWorld and WebShop. We report the average success rate (%) per subtask and overall for ALFWorld, and both the average score and success rate (%) for WebShop. Results of EDGE and other experience-augmented training methods are averaged over 3 random seeds (mean± ). † denotes results replicated from 34 and 14. 4 Experiments We evaluate EDGE to examine whether experience scaffolding can improve agentic RL while being progressively internalized by the policy. Our experiments address four questions: (1) Does EDGE improve over prompting, vanilla RL, and prior experience-augmented methods, particularly on exploration-intensive tasks? (2) Does the trained policy retain its performance when external experiences are removed at inference time? (3) How much does each component—gain gating, distillation, and pruning—contribute, and how do they interact? (4) Is experience utility truly non-stationary, and does the co-evolutionary bank adapt accordingly during training? 4.1 Experiment Setup Environments. We evaluate on two interactive environments with sparse outcome-level feedback. ALFWorld (26) is a text-based household environment with 3,827 task instances spanning six types: Pick & Place (Pick), Examine in Light (Look), Clean & Place (Clean), Heat & Place (Heat), Cool & Place (Cool), and Pick Two & Place (Pick2). WebShop (37) simulates an e-commerce website with over 1.1M products and 12k human-written shopping instructions, requiring agents to search, browse, and purchase products matching detailed user specifications. Baselines. We compare against three families of methods: closed-source LLM agents (GPT-4o (17) and Gemini-2.5-Pro (6)), prompting agents (ReAct (38) and Reflexion (25)), and training-based agents. The post-training baselines include GRPO (24), OPSD (46), EvolveR (33), GRPO augmented with Mem0 (4), SkillRL (34), and EMPO2 (13). Detailed descriptions are provided in Appendix B.1. Training details. We use Qwen2.5-1.5B/7B-Instruct (21) as our base models. For ALFWorld and WebShop, all Post-Training methods use exactly the same hyperparameter configurations. The rollout group size G for group-based RL methods is set to 8. For experience retrieval, we encode experiences with Qwen3-Embedding-0.6B (44) and rank candidates by cosine similarity against the task instruction and initial observations. For experience bank expansion, we use GPT-4o (17) as the reflector LLM to contrast successful and failed trajectories and synthesize new experiences. Full training settings and hyperparameter details are provided in Appendix B.2. 4.2 Main Results Table 1 presents results on ALFWorld and WebShop across two model scales. At 7B, EDGE achieves 90.4% success rate on ALFWorld and 82.6% on WebShop, improving over GRPO by 8.3 and 12.5 points and over the strongest experience-augmented baseline EMPO2 by 1.3 and 5.5 points, with consistent gains at 1.5B. The subtask breakdown reveals that EDGE’s advantage concentrates on exploration-heavy tasks—Heat, Cool, and Pick2—which demand longer action sequences with fewer intermediate rewards. On WebShop, the success-rate improvement over GRPO (+12.5) exceeds the score improvement (+9.3), indicating that EDGE helps agents complete full decision chains rather than merely accumulate partial credit. These results highlight that experience is most effective as selective exploration guidance rather than unconditional augmentation. Naive memory approaches (EvolveR, GRPO+Mem0) degrade well below vanilla GRPO, and standalone self-distillation (OPSD) provides insufficient signal without effective exploration. Combining the two (GRPO+OPSD) recovers competitive aggregate performance but remains brittle across subtasks, whereas EDGE’s marginal-gain gating ensures only beneficial experiences contribute, yielding both higher overall success and more uniform subtask coverage. Figure 2: Performance retention after scaffold removal at inference time. EDGE preserves 96.0% of its scaffolded performance without external experiences, compared with 82.9% for SkillRL and 92.3% for EMPO2, indicating more effective internalization into the parametric policy at the 7B scale. Figure 3: Training dynamics of EDGE vs. GRPO on ALFWorld with Qwen2.5-7B-Instruct. EDGE achieves higher validation success while reducing both environment steps and trajectory length more rapidly, indicating that scaffolded exploration is progressively internalized into more efficient experience-free behavior. Inference without External Scaffolds. Figure 2 poses a stricter test: how much performance survives when all external experiences are withheld at inference time? EDGE preserves 96.0% of its scaffolded performance, compared with 92.3% for EMPO2 and 82.9% for SkillRL, confirming that the reverse-KL distillation stage successfully transfers scaffold-induced behavior into the parametric policy. The residual 3.6-point gap suggests that a small fraction of experience-conditioned exploration strategies resist distillation, a direction we leave to future work. 4.3 Analysis Ablation Studies. Method Pick Look Clean Heat Cool Pick2 All GRPO 92.8 85.7 89.3 75.7 74.5 67.7 82.1 EDGE 96.0 85.1 93.4 90.0 92.9 81.6 90.4 w/o ExpPruning 90.5 82.2 90.1 84.3 87.6 76.8 86.7 w/o Distillation 91.2 81.5 91.8 78.7 77.2 71.1 83.6 w/o Gain-Gating 83.1 72.5 79.7 65.3 61.8 61.2 72.3 Table 2: Ablation results on ALFWorld with Qwen2.5-7B-Instruct. We report success rate (%) per subtask and overall. Table 2 isolates each component on ALFWorld with Qwen2.5-7B-Instruct. The most striking finding is that removing the gain gate drops overall success to 72.3%—9.8 points below vanilla GRPO—revealing that unfiltered experience injection does not merely fail to help but actively harms the policy, particularly on exploration-heavy subtasks where misleading guidance compounds over long horizons. This result also addresses a natural concern about the G/2+G/2G/2+G/2 rollout partition: since all other ablated variants retain the same split yet outperform GRPO, the partition itself does not dilute the RL signal; the degradation is attributable entirely to distilling low-quality experiences. The remaining two components contribute complementary benefits. Without distillation, performance falls to 83.6%, only 1.5 points above GRPO, with losses concentrated on the same exploration-heavy subtasks—confirming that scaffolded exploration provides transient guidance but does not durably reshape the policy without explicit behavioral transfer. Without experience pruning, performance declines more modestly to 86.7% with losses spread evenly across subtasks, indicating that pruning acts as a maintenance mechanism that keeps the experience bank aligned with the policy’s evolving capability rather than targeting any specific failure mode. Training Dynamics. Figure 3 reveals a clear divergence between EDGE and GRPO after approximately step 100: GRPO plateaus and begins to regress, whereas EDGE continues to improve steadily toward 90% validation success. We attribute GRPO’s decline to an exploration–exploitation collapse: once the policy commits to locally successful strategies, it loses the diversity needed to solve the remaining hard tasks, and further optimization erodes earlier gains. EDGE’s experience scaffolds counteract this by continually injecting structured exploration guidance, while gain gating ensures that this guidance remains beneficial as the policy strengthens. The efficiency panels corroborate this interpretation—EDGE reduces both environment steps and trajectory length more rapidly than GRPO, indicating that the policy internalizes increasingly direct action sequences rather than relying on extended trial-and-error, consistent with the scaffold-removal results in Figure 2. Figure 4: Experience gain dynamics during training. Experience Gain Tracking. Figure 4 tracks the mean marginal gain Δe _e (Eq. (5)) across training to test a key premise of §3.3: that experience utility is non-stationary. With a static bank, the initially positive gain decays and turns negative after roughly step 100—confirming that once-useful experiences become actively harmful as the policy outgrows them, and explaining why removing gain gating in Table 2 degrades performance below vanilla GRPO. Experience evolution delays this decay, but unchecked bank growth (reaching over 650 entries) introduces retrieval noise that keeps the gain signal volatile. The full co-evolutionary configuration sustains the most stable positive gain: pruning not only curbs bank growth but causes the bank to shrink after step 100, indicating that the policy absorbs existing experiences faster than new failure modes generate replacements—a direct signature of successful internalization. Further details on bank evolution dynamics appear in Appendix C.2. 5 Related Work Reinforcement Learning for LLM Agents. RL has become a standard post-training paradigm for LLM agents in multi-turn environments (37; 26; 9; 31; 11), progressing from PPO-based methods (23) to critic-free objectives such as GRPO (24) and RLOO (2), with further improvements in multi-turn credit assignment through turn-level or stepwise signals (9; 32; 29; 43). However, the reusable exploration patterns within trajectories are still consumed once and discarded. EDGE retains the group-sampling backbone of GRPO but repurposes it to estimate which retrieved experiences currently improve exploration and to transfer their effects into the policy. Experience-Augmented LLM Agents. External memory and experience reuse have been explored through prompting-based reflections (25; 45; 36; 8; 18) and persistent retrieval during interaction (4; 34; 33; 16; 30; 41; 15). Recent methods selectively replay past reasoning traces (35; 42; 20; 40), but primarily target single-turn tasks without the multi-turn, partially observable interaction loops of agentic settings. EMPO2 (13), the closest prior work, encourages memory-free behavior but selects experiences by static heuristics and distills off-policy, risking distribution mismatch with the student’s own visitation. EDGE measures experience utility online via instantaneous marginal gain and transfers scaffold-induced behavior through on-policy self-distillation on the student’s empirical support. Privileged Information and Self-Distillation. Leveraging training-time information unavailable at deployment spans LUPI (28), asymmetric actor-critic methods (19), and context distillation in LLMs (27; 5; 13; 3), though these approaches are largely off-policy and suffer from distribution mismatch with the student’s own visitation. On-policy distillation (1; 39) addresses this by supervising the student on its own sequences, and on-policy self-distillation (OPSD) (46; 10; 14) further removes the need for a separate teacher. EDGE shares this structure but adds gain-gating to activate distillation only under verified positive gain and co-evolves the experience bank through utility-driven expansion and pruning. 6 Conclusion We presented EDGE, a framework for using retrieved experience as a temporary training-time scaffold rather than a persistent inference-time dependency. EDGE estimates the marginal utility of retrieved experiences under the current policy, admits only positive-gain scaffolds, and distills their behavioral effect into the standard experience-free policy. Across ALFWorld and WebShop, EDGE improves over GRPO and prior experience-augmented agents, with especially large gains on exploration-intensive subtasks and strong retention after scaffold removal. These results suggest a practical principle for agentic RL: external experience is most useful when it is dynamically validated, selectively applied, and ultimately internalized. Limitations While EDGE requires no extra environment rollouts, it introduces training-time overhead for maintaining the experience bank, computing teacher–student comparisons, and invoking the reflector LLM. The effectiveness of gain-gating also depends on reward quality; as with all outcome-based RL methods, noisy or misspecified rewards would reduce the reliability of the marginal-gain estimates. In terms of empirical scope, our evaluation covers two text-based benchmarks with Qwen2.5 models at the 1.5B and 7B scales; broader validation across environments, model families, and larger scales remains future work. More broadly, EDGE assumes experiences are discrete textual artifacts; extending the scaffold-to-parameter principle to latent memory representations is an open problem. References Agarwal et al. (2024) R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. R. Garea, M. Geist, and O. Bachem On-policy distillation of language models: learning from self-generated mistakes. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, External Links: Link Cited by: §5. Ahmadian et al. (2024) A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, O. Pietquin, A. Üstün, and S. Hooker Back to basics: revisiting reinforce-style optimization for learning from human feedback in llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), p. 12248–12267. External Links: Link, Document Cited by: §5. Cheng et al. (2026) Z. Cheng, Z. Liu, Y. Shan, X. Wang, X. Zhu, Y. Ma, H. Wang, Y. Guo, W. Lin, and Y. Wang Mem2^2evolve: towards self-evolving agents via co-evolutionary capability expansion and experience distillation. External Links: 2604.10923, Link Cited by: §5. Chhikara et al. (2025) P. Chhikara, D. Khant, S. Aryan, T. Singh, and D. Yadav Mem0: building production-ready ai agents with scalable long-term memory. External Links: 2504.19413, Link Cited by: item •, §1, §4.1, §5. Choudhury and Sodhi (2025) S. Choudhury and P. Sodhi Better than your teacher: LLM agents that learn from privileged AI feedback. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, External Links: Link Cited by: §5. Comanici et al. (2025) G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, L. Marris, S. Petulla, C. Gaffney, A. Aharoni, N. Lintz, T. C. Pais, H. Jacobsson, I. Szpektor, N. Jiang, K. Haridasan, A. Omran, N. Saunshi, D. Bahri, G. Mishra, E. Chu, T. Boyd, B. Hekman, A. Parisi, C. Zhang, K. Kawintiranon, T. Bedrax-Weiss, O. Wang, Y. Xu, O. Purkiss, U. Mendlovic, I. Deutel, N. Nguyen, A. Langley, F. Korn, L. Rossazza, A. Ramé, S. Waghmare, H. Miller, N. Byrd, A. Sheshan, R. Hadsell, S. Bhardwaj, P. Janus, T. Rissa, D. Horgan, A. Abdagic, L. Belenki, J. Allingham, A. Singh, T. Guidroz, S. Srinivasan, H. Schmit, K. Chiafullo, A. Elisseeff, N. Jha, P. Kolhar, L. Berrada, F. Ding, X. Si, S. B. Mallick, F. Och, S. Erell, E. Ni, T. Latkar, S. Yang, P. Sirkovic, Z. Feng, R. Leland, R. Hornung, G. Wu, C. Blundell, H. Alvari, P. Huang, C. Yip, S. Deur, L. Liu, G. Surita, P. Duque, D. Damen, J. Jia, A. Guez, M. Mircea, A. Sinha, A. Magni, P. Stradomski, T. Marian, V. Galić, W. Chen, H. Husain, A. Singhal, D. Grewe, F. Aubet, S. Song, L. Blanco, L. Rechis, L. Ho, R. Munoz, K. Zheng, J. Hamrick, K. Mather, H. Taitelbaum, E. Rutherford, Y. Lei, K. Chen, A. Shukla, E. Moreira, E. Doi, B. Isik, N. Shabat, D. Rogozińska, K. Kolipaka, J. Chang, E. Vušak, S. Venkatachary, S. Noghabi, T. Bharti, Y. Jun, A. Zaks, S. Green, J. Challagundla, W. Wong, M. Mohammad, D. Hirsch, Y. Cheng, I. Naim, L. Proleev, D. Vincent, A. Singh, M. Krikun, D. Krishnan, Z. Ghahramani, A. Atias, R. Aggarwal, C. Kirov, D. Vytiniotis, C. Koh, A. Chronopoulou, P. Dogra, V. Ion, G. Tyen, J. Lee, F. Weissenberger, T. Strohman, A. Balakrishna, J. Rae, M. Velic, R. de Liedekerke, O. Elyada, W. Yuan, C. Liu, L. Shani, S. Kishchenko, B. Alessio, Y. Li, R. Song, S. Kwei, O. Jankowski, A. Pappu, Y. Namiki, Y. Ma, N. Tripuraneni, C. Cherry, M. Ikonomidis, Y. Ling, C. Ji, B. Westberg, A. Wright, D. Yu, D. Parkinson, S. Ramaswamy, J. Connor, S. H. Yeganeh, S. Grover, G. Kenwright, L. Litchev, C. Apps, A. Tomala, F. Halim, A. Castro-Ros, Z. Li, A. Boral, P. Sho, M. Yarom, E. Malmi, D. Klinghoffer, R. Lin, A. Ansell, P. K. S, S. Zhao, S. Zuo, A. Santoro, H. Cheng, S. Demmessie, Y. Liu, N. Brichtova, A. Culp, N. Braun, D. Graur, W. Ng, N. Mehta, A. Phillips, P. Sundberg, V. Godbole, F. Liu, Y. Katariya, D. Rim, M. Seyedhosseini, S. Ammirati, J. Valfridsson, M. Malihi, T. Knight, A. Toor, T. Lampe, A. Ittycheriah, L. Chiang, C. Yeung, A. Fréchette, J. Rao, H. Wang, H. Srivastava, R. Zhang, R. Rhodes, A. Brand, D. Weesner, I. Figotin, F. Gimeno, R. Fellinger, P. Marcenac, J. Leal, E. Marcus, V. Cotruta, R. Cabrera, S. Luo, D. Garrette, V. Axelrod, S. Baltateanu, D. Barker, D. Chen, H. Toma, B. Ingram, J. Riesa, C. Kulkarni, Y. Zhang, H. Liu, C. Wang, M. Polacek, W. Wu, K. Hui, A. N. Reyes, Y. Su, M. Barnes, I. Malhi, A. Siddiqui, Q. Feng, M. Damaschin, D. Pighin, A. Steiner, S. Yang, R. S. Boppana, S. Ivanov, A. Kandoor, A. Shah, A. Mujika, D. Huang, C. A. Choquette-Choo, M. Patel, T. Yu, T. Creswell, Jerry, Liu, C. Barros, Y. Razeghi, A. Roy, P. Culliton, B. Xiong, J. Pan, T. Strohmann, T. Powell, B. Seal, D. DeCarlo, P. Shyam, K. Katircioglu, X. Wang, C. Hardin, I. Odisho, J. Broder, O. Chang, A. Nair, A. Shtefan, M. O’Brien, M. Agarwal, S. Potluri, S. Goyal, A. Jhindal, S. Thakur, Y. Stuken, J. Lyon, K. Toutanova, F. Feng, A. Wu, B. Horn, A. Wang, A. Cullum, G. Taubman, D. Shrivastava, C. Shi, H. Tomlinson, R. Patel, T. Tu, A. M. Oflazer, F. Pongetti, M. Yang, A. A. Taïga, V. Perot, N. W. Pierse, F. Han, Y. Drori, I. Iturrate, A. Chakrabarti, L. Yeung, D. Dopson, Y. Chen, A. Kulshreshtha, T. Guo, P. Pham, T. Schuster, J. Chen, A. Polozov, J. Xing, H. Zhou, P. Kacham, D. Kukliansky, A. Miech, S. Yaroshenko, E. Chi, S. Douglas, H. Fei, M. Blondel, P. Myla, L. Madmoni, X. Wu, D. Keysers, K. Kjems, I. Albuquerque, L. Yu, J. D’sa, M. Plantan, V. Ionescu, J. S. Elias, A. Gupta, M. R. Vuyyuru, F. Alcober, T. Zhou, K. Ji, F. Hartmann, S. Puttagunta, H. Song, E. Amid, A. Stefanoiu, A. Lee, P. Pucciarelli, E. Wang, A. Raul, S. Petrov, I. Tian, V. Anklin, N. Nti, V. Gomes, M. Schumacher, G. Vesom, A. Panagopoulos, K. Bousmalis, D. Andor, J. Jacob, Y. Zhang, B. Rosgen, M. Kecman, M. Tung, A. Belias, N. Goodman, P. Covington, B. Wieder, N. Saxena, E. Davoodi, M. Huang, S. Maddineni, V. Roulet, F. Campbell-Ajala, P. G. Sessa, Xintian, Wu, G. Lai, P. Collins, A. Haig, V. Sakenas, X. Xu, M. Giustina, L. E. Shafey, P. Charoenpanit, S. Garg, J. Ainslie, B. Severson, M. G. Arenas, S. Pathak, S. Rajayogam, J. Feng, M. Bakker, S. Li, N. Wichers, J. Rogers, X. Geng, Y. Li, R. Jagerman, C. Jia, N. Olmert, D. Sharon, M. Mauger, S. Mariserla, H. Ma, M. Mohabey, K. Kim, A. Andreev, S. Pollom, J. Love, V. Jain, P. Agrawal, Y. Schroecker, A. Fortin, M. Warmuth, J. Liu, A. Leach, I. Blok, G. P. Girirajan, R. Aharoni, B. Uria, A. Sozanschi, D. Goldberg, L. Ionita, M. T. Ribeiro, M. Zlocha, V. Birodkar, S. Lachgar, L. Yuan, H. Choudhury, M. Ginsberg, F. Zheng, G. Dibb, E. Graves, S. Lokhande, G. Rasskin, G. Muraru, C. Quick, S. Tata, P. Sermanet, A. Chawla, I. Karo, Y. Wang, S. Zhang, O. Keller, A. Dragan, G. Su, I. Chou, X. Liu, Y. Tao, S. Prabhakara, M. Wilson, R. Liu, S. Wang, G. Evans, D. Du, A. Castaño, G. Prasad, M. E. Mahdy, S. Gerlach, M. Reid, J. Kahn, A. Zait, T. S. Pillai, T. Ulrich, G. Wang, J. Wassenberg, E. Farkash, K. Yalasangi, C. Wang, M. Bauza, S. Bucher, T. Liu, J. Yan, G. Leung, V. Sindhwani, P. Barnes, A. Singh, I. Jurin, J. Chang, N. K. Bhumihar, S. Eiger, G. Citovsky, B. Withbroe, Z. Li, S. Xue, N. D. Santo, G. Stoyanov, Y. Raimond, S. Zheng, Y. Gao, V. Listík, S. Kwasiborski, R. Saputro, A. Ozturel, G. Mallya, K. Majmundar, R. West, P. Caron, J. Wei, L. Castrejon, S. Vikram, D. Ramachandran, N. Dhawan, J. Park, S. Smoot, G. van den Driessche, Y. Blau, C. Malik, W. Liang, R. Hirsch, C. N. dos Santos, E. Weinstein, A. van den Oord, S. Lall, N. FitzGerald, Z. Jiang, X. Yang, D. Webster, A. Elqursh, A. Pope, G. Rotival, D. Raposo, W. Zhu, J. Dean, S. Alabed, D. Tran, A. Gupta, Z. Gleicher, J. Austin, E. Rosseel, M. Umekar, D. Das, Y. Sun, K. Chen, K. Misiunas, X. Zhou, Y. Di, A. Loo, J. Newlan, B. Li, V. Ramasesh, Y. Xu, A. Chen, S. Gandhe, R. Soricut, N. Gupta, S. Hu, S. El-Sayed, X. Garcia, I. Brusilovsky, P. Chen, A. Bolt, L. Huang, A. Gurney, Z. Zhang, A. Pritzel, J. Wilkiewicz, B. Seybold, B. K. Shamanna, F. Fischer, J. Dean, K. Gill, R. Mcilroy, A. Bhowmick, J. Selier, A. Yang, D. Cheng, V. Magay, J. Tan, D. Varma, C. Walder, T. Kocisky, R. Nakashima, P. Natsev, M. Kwong, I. Gog, C. Zhang, S. Dieleman, T. Jimma, A. Ryabtsev, S. Brahma, D. Steiner, D. Du, A. Žužul, M. Žanić, M. Raghavachari, W. Gierke, Z. Zheng, D. Petrova, Y. Dauphin, Y. Liu, I. Kessler, S. Hand, C. Duvarney, S. Kim, H. Lee, L. Hussenot, J. Hui, J. Smith, D. Jain, J. Xia, G. S. Tomar, K. Amiri, D. Phan, F. Fuchs, T. Weyand, N. Tomasev, A. Cordell, X. Liu, J. Mallinson, P. Joshi, A. Crawford, A. Suggala, S. Chien, N. Fernando, M. Sanchez-Vargas, D. Williams, P. Crone, X. Luo, I. Karpov, J. Shan, T. Thurk, R. Strudel, P. Voigtlaender, P. Patil, T. Dozat, A. Khodaei, S. Singla, P. Ambroszczyk, Q. Wu, Y. Chang, B. Roark, C. Hegde, T. Ding, A. Filos, Z. Wu, A. S. Pinto, S. Liu, S. Khanna, A. Pandey, S. Mcloughlin, Q. Li, S. Haves, A. Zhou, E. Buchatskaya, I. Leal, P. de Boursac, N. Akazawa, N. Anderson, T. Chen, K. Somandepalli, C. Liang, S. Goenka, S. Winkler, A. Grushetsky, Y. Ding, J. Smith, F. Ye, J. Pont-Tuset, E. Li, R. Li, T. Golany, D. Wegner, T. Jiang, O. Barak, Y. Shangguan, E. Vértes, R. Wong, J. Bornschein, A. Tudor, M. Bevilacqua, T. Schaul, A. S. Rawat, Y. Zhao, K. Axiotis, L. Meng, C. McLean, J. Lai, J. Beattie, N. Kushman, Y. Liu, B. Kutzman, F. Lang, J. Ye, P. Netrapalli, P. Mishra, M. Khan, M. Goel, R. Willoughby, D. Tian, H. Zhuang, J. Chen, Z. Tsai, T. Kementsietsidis, A. Khare, J. Keeling, K. Xu, N. Waters, F. Altché, A. Popat, B. Mittal, D. Saxton, D. E. Badawy, M. Mathieu, Z. Zheng, H. Zhou, N. Ranka, R. Shin, Q. Duan, T. Salimans, I. Mihailescu, U. Shaham, M. Chang, Y. Assael, N. Dikkala, M. Izzard, V. Cohen-Addad, C. Graves, V. Feinberg, G. Chung, D. Strouse, D. Karmon, S. Sharifzadeh, Z. Ashwood, K. Pham, J. Blanton, A. Vasiloff, J. Barber, M. Geller, A. Zhou, F. Zubach, T. Huang, L. Zhang, H. Gupta, M. Young, J. Proskurnia, R. Votel, V. Gabeur, G. Barcik, A. Tripathi, H. Yu, G. Yan, B. Changpinyo, F. Pavetić, A. Coyle, Y. Fujii, J. G. Mendez, T. Zhou, H. Rajamani, B. Hechtman, E. Cao, D. Juan, Y. Tan, V. Dalibard, Y. Du, N. Clay, K. Yao, W. Jia, D. Vijaykumar, Y. Zhou, X. Bai, W. Hung, S. Pecht, G. Todorov, N. Khadke, P. Gupta, P. Lahoti, A. Autef, K. Duddu, J. Lee-Thorp, A. Bykovsky, T. Misiunas, S. Flennerhag, S. Thangaraj, J. McGiffin, Z. Nado, M. Kunesch, A. Noever, A. Hertz, M. Liang, V. Stone, E. Palmer, S. Daruki, A. Pramanik, S. Põder, A. Kyker, M. Khan, E. Sluzhaev, M. Ritter, A. Ruderman, W. Zhou, C. Nagpal, K. Vodrahalli, G. Necula, P. Barham, E. Pavlick, J. Hartford, I. Shafran, L. Zhao, M. Mikuła, T. Eccles, H. Shimokawa, K. Garg, L. Vilnis, H. Chen, I. Shumailov, K. Lee, A. Abdelhamed, M. Xie, V. Cohen, E. Hlavnova, D. Malkin, C. Sitawarin, J. Lottes, P. Coquinot, T. Yu, S. Kumar, J. Zhang, A. Mahendru, Z. Ahmed, J. Martens, T. Chen, A. Boag, D. Peng, C. Devin, A. Klimovskiy, M. Phuong, D. Vainstein, J. Xie, B. Ramabhadran, N. Howard, X. Yu, G. Goswami, J. Cui, S. Shleifer, M. Pinto, C. Yeh, M. Yang, S. Javanmardi, D. Ethier, C. Lee, J. Orbay, S. Kotecha, C. Bromberg, P. Shaw, J. Thornton, A. G. Rosenthal, S. Gu, M. Thomas, I. Gemp, A. Ayyar, A. Ushio, A. Selvan, J. Wee, C. Liu, M. Majzoubi, W. Yu, J. Abernethy, T. Liechty, R. Pan, H. Nguyen, Qiong, Hu, S. Perrin, A. Arora, E. Pitler, W. Wang, K. Shivakumar, F. Prost, B. Limonchik, J. Wang, Y. Gao, T. Cour, S. Buch, H. Gui, M. Ivanova, P. Neubeck, K. Chan, L. Kim, H. Chen, N. Goyal, D. Chung, L. Liu, Y. Su, A. Petrushkina, J. Shen, A. Joulin, Y. Xu, S. X. Lin, Y. Kulizhskaya, C. Chelba, S. Vasudevan, E. Collins, V. Bashlovkina, T. Lu, D. Fritz, J. Park, Y. Zhou, C. Su, R. Tanburn, M. Sushkov, M. Rasquinha, J. Li, J. Prendki, Y. Li, P. LV, S. Sharma, H. Fitoussi, H. Huang, A. Dai, P. Dao, M. Burrows, H. Prior, D. Qin, G. Pundak, L. L. Sjoesund, A. Khurshudov, Z. Zhu, A. Webson, E. Kemp, T. Tan, S. Agrawal, S. Sargsyan, L. Cheng, J. Stephan, T. Kwiatkowski, D. Reid, A. Byravan, A. H. Michaely, N. Heess, L. Zhou, S. Goenka, V. Carpenter, A. Levskaya, B. Wang, R. Roberts, R. Leblond, S. Chikkerur, S. Ginzburg, M. Chang, R. Riachi, Chuqiao, Xu, Z. Borsos, M. Pliskin, J. Pawar, M. Lustman, H. Kirkwood, A. Anand, A. Chaudhary, N. Kalb, K. Milan, S. Augenstein, A. Goldie, L. Prince, K. Raman, Y. Sun, V. Xia, A. Cohen, Z. Huo, J. Camp, S. Ellis, L. Zilka, D. V. Torres, L. Patel, S. Arora, B. Chan, J. Adler, K. Ayoub, J. Liang, F. Jamil, J. Jiang, S. Baumgartner, H. Sun, Y. Karov, Y. Akulov, H. Zheng, I. Cai, C. Fantacci, J. Rubin, A. R. Acha, M. Wang, N. D’Souza, R. Sathyanarayana, S. Dai, S. Rowe, A. Simanovsky, O. Goldman, Y. Kuang, X. Pan, A. Rosenberg, T. Rojas-Esponda, P. Dutta, A. Zeng, I. Jurenka, G. Farquhar, Y. Bansal, S. Iqbal, B. Roelofs, G. Joung, P. Beak, C. Ryu, R. Poplin, Y. Wu, J. Alayrac, S. Buthpitiya, O. Ronneberger, C. Habtegebriel, W. Li, P. Cavallaro, A. Wei, G. Bensky, T. Denk, H. Ganapathy, J. Stanway, P. Joshi, F. Bertolini, J. Lo, O. Ma, Z. Charles, G. Sampemane, H. Sahni, X. Chen, H. Askham, D. Gaddy, P. Young, J. Tan, M. Eyal, A. Bražinskas, L. Zhong, Z. Wu, M. Epstein, K. Bailey, A. Hard, K. Lee, S. Goldshtein, A. Ruiz, M. Badawi, M. Lochbrunner, J. Kearns, A. Brown, F. Pardo, T. Weber, H. Yang, P. Jiang, B. Akin, Z. Fu, M. Wainwright, C. Zou, M. Gaba, P. Manzagol, W. Kan, Y. Song, K. Zainullina, R. Lin, J. Ko, S. Deshmukh, A. Jindal, J. Svensson, D. Tyam, H. Zhao, C. Kaeser-Chen, S. Baird, P. Moradi, J. Hall, Q. Guo, V. Tsang, B. Liang, F. Pereira, S. Ganesh, I. Korotkov, J. Adamek, S. Thiagarajan, V. Tran, C. Chen, C. Tar, S. Jain, I. Dasgupta, T. Bilal, D. Reitter, K. Zhao, G. Vezzani, Y. Gehman, P. Mehta, L. Beltrone, X. Dotiwalla, S. Guadarrama, Z. Abbas, S. Karp, P. Georgiev, C. Ferng, M. Brockschmidt, L. Peng, C. Hirnschall, V. Verma, Y. Bi, Y. Xiao, A. Dabush, K. Xu, P. Wallis, R. Parker, Q. Wang, Y. Xu, I. Safarli, D. Tewari, Y. Zhang, S. Kim, A. Gesmundo, M. Thomas, S. Levi, A. Chowdhury, K. Rao, P. Garst, S. Conway-Rahman, H. Ran, K. McKinney, Z. Xiao, W. Yu, R. Agrawal, A. Stjerngren, C. Ionescu, J. Chen, V. Sharma, J. Chiu, F. Liu, K. Franko, C. Sanford, X. Cai, P. Michel, S. Ganapathy, J. Labanowski, Z. Garrett, B. Vargas, S. Sun, B. Gale, T. Buschmann, G. Desjardins, N. Ghelani, P. Jain, M. Verma, C. Asawaroengchai, J. Eisenschlos, J. Harlalka, H. Kazawa, D. Metzler, J. Howland, Y. Jian, J. Ades, V. Shah, T. Gangwani, S. Lee, R. Ring, S. M. Hernandez, D. Reich, A. Sinha, A. Sathe, J. Kovac, A. Gill, A. Kannan, A. D’olimpio, M. Sevenich, J. Whang, B. Kim, K. C. Sim, J. Chen, J. Zhang, S. Lall, Y. Matias, B. Jia, A. Friesen, S. Nasso, A. Thapliyal, B. Perozzi, T. Yu, A. Shekhawat, S. Huda, P. Grabowski, E. Wang, A. Sreevatsa, H. Dib, M. Hassen, P. Schuh, V. Milutinovic, C. Welty, M. Quinn, A. Shah, B. Wang, G. Barth-Maron, J. Frye, N. Axelsson, T. Zhu, Y. Ma, I. Giannoumis, H. Sedghi, C. Ye, Y. Luan, K. Aydin, B. Chandra, V. Sampathkumar, R. Huang, V. Lavrenko, A. Eleryan, Z. Hong, S. Hansen, S. M. Carthy, B. Samanta, D. Ćevid, X. Wang, F. Li, M. Voznesensky, M. Hoffman, A. Terzis, V. Sehwag, G. Fidel, L. He, M. Cai, Y. He, A. Feng, M. Nikoltchev, S. Phatale, J. Chase, R. Lawton, M. Zhang, T. Ouyang, M. Tragut, M. H. Manshadi, A. Narayanan, J. Shen, X. Gao, T. Bolukbasi, N. Roy, X. Li, D. Golovin, L. Panait, Z. Qin, G. Han, T. Anthony, S. Kudugunta, V. Patraucean, A. Ray, X. Chen, X. Yang, T. Bhatia, P. Talluri, A. Morris, A. Ražnatović, B. Brownfield, J. An, S. Peng, P. Kane, C. Zheng, N. Duduta, J. Kessinger, J. Noraky, S. Liu, K. Rong, P. Veličković, K. Rush, A. Goldin, F. Wei, S. M. R. Garlapati, C. Pantofaru, O. Kwon, J. Ni, E. Noland, J. D. Trapani, F. Beaufays, A. G. Roy, Y. Chow, A. Turker, G. Cideron, L. Mei, J. Clark, Q. Dou, M. Bošnjak, R. Leith, Y. Du, A. Yazdanbakhsh, M. Nasr, C. Kwak, S. S. Sheth, A. Kaskasoli, A. Anand, B. Lakshminarayanan, S. Jerome, D. Bieber, C. Chu, A. Senges, T. Shen, M. Sridhar, N. Ndebele, B. Beyret, S. Mohamed, M. Chen, M. Freitag, J. Guo, L. Liu, P. Roit, H. Chen, S. Yan, T. Stone, J. Co-Reyes, J. Cole, S. Scellato, S. Azizi, H. Hashemi, A. Jin, A. Iyer, M. Valentine, A. György, A. Ahuja, D. H. Diaz, C. Lee, N. Clement, W. Kong, D. Garmon, I. Watts, K. Bhatia, K. Gupta, M. Miecnikowski, H. Vallet, A. Taly, E. Loper, S. Joshi, J. Atwood, J. Chick, M. Collier, F. Iliopoulos, R. Trostle, B. Gunel, R. Leal-Cavazos, A. M. Hrafnkelsson, M. Guzman, X. Ju, A. Forbes, J. Emond, K. Chauhan, B. Caine, L. Xiao, W. Zeng, A. Moufarek, D. Murphy, M. Meng, N. Gupta, F. Riedel, A. Das, E. Lawal, S. Narayan, T. Sosea, J. Swirhun, L. Friso, B. Neyshabur, J. Lu, S. Girgin, M. Wunder, E. Yvinec, A. Pyne, V. Carbune, S. Rijhwani, Y. Guo, T. Doshi, A. Briukhov, M. Bain, A. Hitron, X. Wang, A. Gupta, K. Chen, C. Du, W. Zhang, D. Shah, A. Akula, M. Dylla, A. Kachra, W. Kuo, T. Zou, L. Wang, L. Xu, J. Zhu, J. Snyder, S. Menon, O. Firat, I. Mordatch, Y. Yuan, N. Ponomareva, R. Blevins, L. Moore, W. Wang, P. Chen, M. Scholz, A. Dwornik, J. Lin, S. Li, D. Antognini, T. I, X. Song, M. Miller, U. Kalra, A. Raveret, O. Akerlund, F. Wu, A. Nystrom, N. Godbole, T. Liu, H. DeBalsi, J. Zhao, B. Liu, A. Caciularu, L. Lax, U. Khandelwal, V. Langston, E. Bailey, S. Lattanzi, Y. Wang, N. Kovelamudi, S. Mondal, G. Guruganesh, N. Hua, O. Roval, P. Wesołowski, R. Ingale, J. Halcrow, T. Sohn, C. Angermueller, B. Raad, E. Stickgold, E. Lu, A. Kosik, J. Xie, T. Lillicrap, A. Huang, L. L. Zhang, D. Paulus, C. Farabet, A. Wertheim, B. Wang, R. Joshi, C. Ko, Y. Wu, S. Agrawal, L. Lin, X. Sheng, P. Sung, T. Breland-King, C. Butterfield, S. Gawde, S. Singh, Q. Zhang, R. Apte, S. Shetty, A. Hutter, T. Li, E. Salesky, F. Lebron, J. Kanerva, M. Paganini, A. Nguyen, R. Vallu, J. Peter, S. Velury, D. Kao, J. Hoover, A. Bortsova, C. Bishop, S. Jakobovits, A. Agostini, A. Agarwal, C. Liu, C. Kwong, S. Tavakkol, I. Bica, A. Greve, A. GP, J. Marcus, L. Hou, T. Duerig, R. Moroshko, D. Lacey, A. Davis, J. Amelot, G. Wang, F. Kim, T. Strinopoulos, H. Wan, C. L. Lan, S. Krishnan, H. Tang, P. Humphreys, J. Bai, I. H. Shtacher, D. Machado, C. Pang, K. Burke, D. Liu, R. Aravamudhan, Y. Song, E. Hirst, A. Singh, B. Jou, L. Bai, F. Piccinno, C. K. Fu, R. Alazard, B. Meiri, D. Winter, C. Chen, M. Zhang, J. Heitkaemper, J. Lambert, J. Lee, A. Frömmgen, S. Rogulenko, P. Nair, P. Niemczyk, A. Bulyenov, B. Xu, H. Shemtov, M. Zadimoghaddam, S. Toropov, M. Wirth, H. Dai, S. Gollapudi, D. Zheng, A. Kurakin, C. Lee, K. Bullard, N. Serrano, I. Balazevic, Y. Li, J. Schalkwyk, M. Murphy, M. Zhang, K. Sequeira, R. Datta, N. Agrawal, C. Sutton, N. Attaluri, M. Chiang, W. Farhan, G. Thornton, K. Lin, T. Choma, H. Nguyen, K. Dasgupta, D. Robinson, I. Comşa, M. Riley, A. Pillai, B. Mustafa, B. Golan, A. Zandieh, J. Lespiau, B. Porter, D. Ross, S. Rajayogam, M. Agarwal, S. Venugopalan, B. Shahriari, Q. Yan, H. Xu, T. Tobin, P. Dubov, H. Shi, A. Recasens, A. Kovsharov, S. Borgeaud, L. Dery, S. Vasanth, E. Gribovskaya, L. Qiu, M. Mahdieh, W. Skut, E. Nielsen, C. Zheng, A. Yu, C. G. Bostock, S. Gupta, A. Archer, C. Rawles, E. Davies, A. Svyatkovskiy, T. Tsai, Y. Halpern, C. Reisswig, B. Wydrowski, B. Chang, J. Puigcerver, M. H. Taege, J. Li, E. Schnider, X. Li, D. Dena, Y. Xu, U. Telang, T. Shi, H. Zen, K. Kastner, Y. Ko, N. Subramaniam, A. Kumar, P. Blois, Z. Dai, J. Wieting, Y. Lu, Y. Zeldes, T. Xie, A. Hauth, A. Ţifrea, Y. Li, S. El-Husseini, D. Abolafia, H. Zhou, W. Ding, S. Ghalebikesabi, C. Guía, A. Maksai, Á. Weisz, S. Arik, N. Sukhanov, A. Świetlik, X. Jia, L. Yu, W. Wang, M. Brand, D. Bloxwich, S. Kirmani, Z. Chen, A. Go, P. Sprechmann, N. Kannen, A. Carin, P. Sandhu, I. Edkins, L. Nooteboom, J. Gupta, L. Maggiore, J. Azizi, Y. Pritch, P. Yin, M. Gupta, D. Tarlow, D. Smith, D. Ivanov, M. Babaeizadeh, A. Goel, S. Kambala, G. Chu, M. Kastelic, M. Liu, H. Soltau, A. Stone, S. Agrawal, M. Kim, K. Soparkar, S. Tadepalli, O. Bunyan, R. Soh, A. Kannan, D. Kim, B. J. Chen, A. Halumi, S. Roy, Y. Wang, O. Sercinoglu, G. Gibson, S. Bhatnagar, M. Sano, D. von Dincklage, Q. Ren, B. Mitrevski, M. Olšák, J. She, C. Doersch, Jilei, Wang, B. Liu, Q. Tan, T. Yakar, T. Warkentin, A. Ramirez, C. Lebsack, J. Dillon, R. Mathews, T. Cobley, Z. Wu, Z. Chen, J. Simon, S. Nath, T. Sainath, A. Bendebury, R. Julian, B. Mankalale, D. Ćurko, P. Zacchello, A. R. Brown, K. Sodhia, H. Howard, S. Caelles, A. Gupta, G. Evans, A. Bulanova, L. Katzen, R. Goldenberg, A. Tsitsulin, J. Stanton, B. Schillings, V. Kovalev, C. Fry, R. Shah, K. Lin, S. Upadhyay, C. Li, S. Radpour, M. Maggioni, J. Xiong, L. Haas, J. Brennan, A. Kamath, N. Savinov, A. Nagrani, T. Yacovone, R. Kappedal, K. Andriopoulos, L. Lao, Y. Li, G. Rozhdestvenskiy, K. Hashimoto, A. Audibert, S. Austin, D. Rodriguez, A. Ruoss, G. Honke, D. Karkhanis, X. Xiong, Q. Wei, J. Huang, Z. Leng, V. Premachandran, S. Bileschi, G. Evangelopoulos, T. Mensink, J. Pavagadhi, D. Teplyashin, P. Chang, L. Xue, G. Tanzer, S. Goldman, K. Patel, S. Li, J. Wiesner, I. Zheng, I. Stewart-Binks, J. Han, Z. Li, L. Luo, K. Lenc, M. Lučić, F. Xue, R. Mullins, A. Guseynov, C. Chang, I. Galatzer-Levy, A. Zhang, G. Bingham, G. Hu, A. Hartman, Y. Ma, J. Griffith, A. Irpan, C. Radebaugh, S. Yue, L. Fan, V. Ungureanu, C. Sorokin, H. Teufel, P. Li, R. Anil, D. Paparas, T. Wang, C. Lin, H. Peng, M. Shum, G. Petrovic, D. Brady, R. Nguyen, K. Macherey, Z. Li, H. Singh, M. Yenugula, M. Iinuma, X. Chen, K. Kopparapu, A. Stern, S. Dave, C. Thekkath, F. Perot, A. Kumar, F. Li, Y. Xiao, M. Bilotti, M. H. Bateni, I. Noble, L. Lee, A. Vázquez-Reina, J. Salazar, X. Yang, B. Wang, E. Gruzewska, A. Rao, S. Raghuram, Z. Xu, E. Ben-David, J. Mei, S. Dalmia, Z. Zhang, Y. Liu, G. Bansal, H. Pankov, S. Schwarcz, A. Burns, C. Chan, S. Sanghai, R. Liang, E. Liang, A. He, A. Stuart, A. Narayanan, Y. Zhu, C. Frank, B. Fatemi, A. Sabne, O. Lang, I. Bhattacharya, S. Settle, M. Wang, B. McMahan, A. Tacchetti, L. B. Soares, M. Hadian, S. Cabi, T. Chung, N. Putikhin, G. Li, J. Chen, A. Tarango, H. Michalewski, M. Kazemi, H. Masoom, H. Sheftel, R. Shivanna, A. Vadali, R. Comanescu, D. Reid, J. Moore, A. Neelakantan, M. Sander, J. Herzig, A. Rosenberg, M. Dehghani, J. Choi, M. Fink, R. Hayes, E. Ge, S. Weng, C. Ho, J. Karro, K. Krishna, L. N. Thiet, A. Skerry-Ryan, D. Eppens, M. Andreetto, N. Sarma, S. Bonacina, B. K. Ayan, M. Nawhal, Z. Shan, M. Dusenberry, S. Thakoor, S. Gubbi, D. D. Nguyen, R. Tsarfaty, S. Albanie, J. Mitrović, M. Gandhi, B. Chen, A. Epasto, G. Stephanov, Y. Jin, S. Gehman, A. Amini, J. Weber, F. Behbahani, S. Xu, M. Allamanis, X. Chen, M. Ott, C. Sha, M. Jastrzebski, H. Qi, D. Greene, X. Wu, A. Toki, D. Vlasic, J. Shapiro, R. Kotikalapudi, Z. Shen, T. Saeki, S. Xie, A. Cassirer, S. Bharadwaj, T. Kiyono, S. Bhojanapalli, E. Rosenfeld, S. Ritter, J. Mao, J. G. Oliveira, Z. Egyed, B. Bandemer, E. Parisotto, K. Kinoshita, J. Pluto, P. Maniatis, S. Li, Y. Guo, G. Ghiasi, J. Tarbouriech, S. Chatterjee, J. Jin, Katrina, Xu, J. Palomaki, S. Arnold, M. Sewak, F. Piccinini, M. Sharma, B. Albrecht, S. Purser-haskell, A. Vaswani, C. Chen, M. Wisniewski, Q. Cao, J. Aslanides, N. M. Phu, M. Sieb, L. Agubuzu, A. Zheng, D. Sohn, M. Selvi, A. Andreassen, K. Subudhi, P. Eruvbetine, O. Woodman, T. Mery, S. Krause, X. Ren, X. Ma, J. Luo, D. Chen, W. Fan, H. Griffiths, C. Schuler, A. Li, S. Zhang, J. Sarr, S. Luo, R. Patana, M. Watson, D. Naboulsi, M. Collins, S. Sidhwani, E. Hoogeboom, S. Silver, E. Caveness, X. Zhao, M. Rodriguez, M. Deines, L. Bai, P. Griffin, M. Tagliasacchi, E. Xue, S. R. Babbula, B. Pang, N. Ding, G. Shen, E. Peake, R. Crocker, S. S. Raghvendra, D. Swisher, W. Han, R. Singh, L. Wu, V. Pchelin, T. Munkhdalai, D. Alon, G. Bacon, E. Robles, J. Bulian, M. Johnson, G. Powell, F. T. Ferreira, Y. Li, F. Benzing, M. Velimirović, H. Soyer, W. Kong, Tony, Nguyên, Z. Yang, J. Liu, J. van Amersfoort, D. Gillick, B. Sun, N. Rauschmayr, K. Zhang, S. Zhan, T. Zhou, A. Frolov, C. Yang, D. Vnukov, L. Rouillard, H. Li, A. Mandhane, N. Fallen, R. Venkataraman, C. H. Hu, J. Brennan, J. Lee, J. Chang, M. Sundermeyer, Z. Pan, R. Ke, S. Tong, A. Fabrikant, W. Bono, J. Gu, R. Foley, Y. Mao, M. Delakis, D. Bhaswar, R. Frostig, N. Li, A. Zipori, C. Hope, O. Kozlova, S. Mishra, J. Djolonga, C. Schiff, M. A. Merey, E. Briakou, P. Morgan, A. Wan, A. Hassidim, R. Skerry-Ryan, K. Sengupta, M. Jasarevic, P. Kallakuri, P. Kunkle, H. Brennan, T. Lieber, H. Mansoor, J. Walker, B. Zhang, A. Xie, G. Žužić, A. Chukwuka, A. Druinsky, D. Cho, R. Yao, F. Naeem, S. Butt, E. Kim, Z. Jia, M. Jordan, A. Lelkes, M. Kurzeja, S. Wang, J. Zhao, A. Over, A. Chakladar, M. Prasetya, N. Jha, S. Ganapathy, Y. Cong, P. Shroff, C. Saroufim, S. Miryoosefi, M. Hammad, T. Nasir, W. Xi, Y. Gao, Y. Maeng, B. Hora, C. Cheng, P. Haghani, Y. Lewenberg, C. Lu, M. Matysiak, N. Raisinghani, H. Wang, L. Baugher, R. Sukthankar, M. Giang, J. Schultz, N. Fiedel, M. Chen, C. Lee, T. Dey, H. Zheng, S. Paul, C. Smith, A. Ly, Y. Wang, R. Bansal, B. Perz, S. Ricco, S. Blank, V. Keshava, D. Sharma, M. Chow, K. Lad, K. Jalan, S. Osindero, C. Swanson, J. Scott, A. Ilić, X. Li, S. R. Jonnalagadda, A. S. Soudagar, Y. Xiong, B. Batsaikhan, D. Jarrett, N. Kumar, M. Shah, M. Lawlor, A. Waters, M. Graham, R. May, S. Ramos, S. Lefdal, Z. Cankara, N. Cano, B. O’Donoghue, J. Borovik, F. Liu, J. Grimstad, M. Alnahlawi, K. Tsihlas, T. Hudson, N. Grigorev, Y. Jia, T. Huang, T. P. Igwe, S. Lebedev, X. Tang, I. Krivokon, F. Garcia, M. Tan, E. Jia, P. Stys, S. Vashishth, Y. Liang, B. Venkatraman, C. Gu, A. Kementsietsidis, C. Zhu, J. Jung, Y. Bai, M. J. Hosseini, F. Ahmed, A. Gupta, X. Yuan, S. Ashraf, S. Nigam, G. Vasudevan, P. Awasthi, A. M. Gilady, Z. Mariet, R. Eskander, H. Li, H. Hu, G. Garrido, P. Schlattner, G. Zhang, R. Saxena, P. Dević, K. Muralidharan, A. Murthy, Y. Zhou, M. Choi, A. Wongpanich, Z. Wang, P. Shah, Y. Xu, Y. Huang, S. Spencer, A. Chen, J. Cohan, J. Wang, J. Tompson, J. Wu, R. Haroun, H. Li, B. Huergo, F. Yang, T. Yin, J. Wendt, M. Bendersky, R. Chaabouni, J. Snaider, J. Ferret, A. Jindal, T. Thompson, A. Xue, W. Bishop, S. M. Phal, A. Sharma, Y. Sung, P. Radhakrishnan, M. Shomrat, R. Ingle, R. Vij, J. Gilmer, M. D. Istin, S. Sobell, Y. Lu, E. Nottage, D. Sadigh, J. Willcock, T. Zhang, S. Xu, S. Brown, K. Lee, G. Wang, Y. Zhu, Y. Tay, C. Kim, A. Gutierrez, A. Sharma, Y. Xian, S. Seo, C. Cui, E. Pochernina, C. Baetu, K. Jastrzębski, M. Ly, M. Elhawaty, D. Suh, E. Sezener, P. Wang, N. Yuen, G. Tucker, J. Cai, Z. Yang, C. Wang, A. Muzio, H. Qian, J. Yoo, D. Lockhart, K. R. McKee, M. Guo, M. Mehrotra, A. Mendonça, S. V. Mehta, S. Ben, C. Tekur, J. Mu, M. Zhu, V. Krakovna, H. Lee, A. Maschinot, S. Cevey, H. Choe, A. Bai, H. Srinivasan, D. Gasaway, N. Young, P. Siegler, D. Holtmann-Rice, V. Piratla, K. Baumli, R. Yogev, A. Hofer, H. van Hasselt, S. Grant, Y. Chervonyi, D. Silver, A. Hogue, A. Agarwal, K. Wang, P. Singh, F. Flynn, J. Lipschultz, R. David, L. Bellot, Y. Yang, L. Le, F. Graziano, K. Olszewska, K. Hui, A. Maurya, N. Parotsidis, W. Chen, T. Oguntebi, J. Kelley, A. Baddepudi, J. Mauerer, G. Shaw, A. Siegman, L. Yang, S. Shetty, S. Roy, Y. Song, W. Stokowiec, R. Burnell, O. Savant, R. Busa-Fekete, J. Miao, S. Ghosh, L. MacDermed, P. Lippe, M. Dektiarev, Z. Behrman, F. Mentzer, K. Nguyen, M. Wei, S. Verma, C. Knutsen, S. Dasari, Z. Yan, P. Mitrichev, X. Wang, V. Shejwalkar, J. Austin, S. Sunkara, N. Potti, Y. Virin, C. Wright, G. Liu, O. Riva, E. Pot, G. Kochanski, Q. Le, G. Balasubramaniam, A. Dhar, Y. Liao, A. Bloniarz, D. Shukla, E. Cole, J. Lee, S. Zhang, S. Kafle, S. Vashishtha, P. Mahmoudieh, G. Chen, R. Hoffmann, P. Srinivasan, A. D. Lago, Y. B. Shalom, Z. Wang, M. Elabd, A. Sharma, J. Oh, S. Kothawade, M. Le, M. Monteiro, S. Yang, K. Alarakyia, R. Geirhos, D. Mincu, H. Garnes, H. Kobayashi, S. Mariooryad, K. Krasowiak, Zhixin, Lai, S. Mourad, M. Wang, F. Bu, O. Aharoni, G. Chen, A. Goyal, V. Zubov, A. Bapna, E. Dabir, N. Kothari, K. Lamerigts, N. D. Cao, J. Shar, C. Yew, N. Kulkarni, D. Mahaarachchi, M. Joshi, Z. Zhu, J. Lichtarge, Y. Zhou, H. Muckenhirn, V. Selo, O. Vinyals, P. Chen, A. Brohan, V. Mehta, S. Cogan, R. Wang, T. Geri, W. Ko, W. Chen, F. Viola, K. Shivam, L. Wang, M. C. Elish, R. A. Popa, S. Pereira, J. Liu, R. Koster, D. Kim, G. Zhang, S. Ebrahimi, P. Talukdar, Y. Zheng, P. Poklukar, A. Mikhalap, D. Johnson, A. Vijayakumar, M. Omernick, M. Dibb, A. Dubey, Q. Hu, A. Suman, V. Aggarwal, I. Kornakov, F. Xia, W. Lowe, A. Kolganov, T. Xiao, V. Nikolaev, S. Hemingray, B. Li, J. Iljazi, M. Rybiński, B. Sandhu, P. Lu, T. Luong, R. Jenatton, V. Govindaraj, Hui, Li, G. Dulac-Arnold, W. Park, H. Wang, A. Modi, J. Pouget-Abadie, K. Greller, R. Gupta, R. Berry, P. Ramachandran, J. Xie, L. McCafferty, J. Wang, K. Gupta, H. Lim, B. Bratanič, A. Brock, I. Akolzin, J. Sproch, D. Karliner, D. Kim, A. Goedeckemeyer, N. Shazeer, C. Schmid, D. Calandriello, P. Bhatia, K. Choromanski, C. Montgomery, D. Dua, A. Ramalho, H. King, Y. Gao, L. Nguyen, D. Lindner, D. Pitta, O. Johnson, K. Salama, D. Ardila, M. Han, E. Farnese, S. Odoom, Z. Wang, X. Ding, N. Rink, R. Smith, H. T. Lehri, E. Cohen, N. Vats, T. He, P. Gopavarapu, A. Paszke, M. Patel, W. V. Gansbeke, L. Loher, L. Castro, M. Voitovich, T. von Glehn, N. George, S. Niklaus, Z. Eaton-Rosen, N. Rakićević, E. Jue, S. Perel, C. Zhang, Y. Bahat, A. Pouget, Z. Xing, F. Huot, A. Shenoy, T. Bos, V. Coriou, B. Richter, N. Noy, Y. Wang, S. Ontanon, S. Qin, G. Makarchuk, D. Hassabis, Z. Li, M. Sharma, K. Venkatesan, I. Kemaev, R. Daniel, S. Huang, S. Shah, O. Ponce, Warren, Chen, M. Faruqui, J. Wu, S. Andačić, S. Payrits, D. McDuff, T. Hume, Y. Cao, M. Tessler, Q. Wang, Y. Wang, I. Rendulic, E. Agustsson, M. Johnson, T. Lando, A. Howard, S. G. S. Padmanabhan, M. Daswani, A. Banino, M. Kilgore, J. Heek, Z. Ji, A. Caceres, C. Li, N. Kassner, A. Vlaskin, Z. Liu, A. Grills, Y. Hou, R. Sukkerd, G. Cheon, N. Shetty, L. Markeeva, P. Stanczyk, T. Iyer, Y. Gong, S. Gao, K. Gopalakrishnan, T. Blyth, M. Reynolds, A. Bhoopchand, M. Bilenko, D. Gharibian, V. Zayats, A. Faust, A. Singh, M. Ma, H. Jiao, S. Vijayanarasimhan, L. Aroyo, V. Yadav, S. Chakera, A. Kakarla, V. Meshram, K. Gregor, G. Botea, E. Senter, D. Jia, G. Kovacs, N. Sharma, S. Baur, K. Kang, Y. He, L. Zhuo, M. Kostelac, I. Laish, S. Peng, L. O’Bryan, D. Kasenberg, G. R. Rao, E. Leurent, B. Zhang, S. Stevens, A. Salazar, Y. Zhang, I. Lobov, J. Walker, A. Porter, M. Redshaw, H. Ke, A. Rao, A. Lee, H. Lam, M. Moffitt, J. Kim, S. Qiao, T. Koo, R. Dadashi, X. Song, M. Sundararajan, P. Xu, C. Kawamoto, Y. Zhong, C. Barbu, A. Reddy, M. Verzetti, L. Li, G. Papamakarios, H. Klimczak-Plucińska, M. Cassin, K. Kavukcuoglu, R. Swavely, A. Vaucher, J. Zhao, R. Hemsley, M. Tschannen, H. Ge, G. Menghani, Y. Yu, N. Ha, W. He, X. Wu, M. Song, R. Sterneck, S. Zinke, D. A. Calian, A. Marsden, A. C. Ruiz, M. Hessel, A. Gueta, B. Lee, B. Farris, M. Gupta, Y. Li, M. Saleh, V. Misra, K. Xiao, P. Mendolicchio, G. Buttimore, V. Krayvanova, N. Nayakanti, M. Wiethoff, Y. Pande, A. Mirhoseini, N. Lao, J. Liu, Y. Hua, A. Chen, Y. Malkov, D. Kalashnikov, S. Gupta, K. Audhkhasi, Y. Zhai, S. Kopalle, P. Jain, E. Ofek, C. Meyer, K. Baatarsukh, H. Strejček, J. Qian, J. Freedman, R. Figueira, M. Sokolik, O. Bachem, R. Lin, D. Kharrat, C. Hidey, P. Xu, D. Duan, Y. Li, M. Ersoy, R. Everett, K. Cen, R. Santamaria-Fernandez, A. Taubenfeld, I. Mackinnon, L. Deng, P. Zablotskaia, S. Viswanadha, S. Goel, D. Yates, Y. Deng, P. Choy, M. Chen, A. Sinha, A. Mossin, Y. Wang, A. Szlam, S. Hao, P. K. Rubenstein, M. Toksoz-Exley, M. Aperghis, Y. Zhong, J. Ahn, M. Isard, O. Lacombe, F. Luisier, C. Anastasiou, Y. Kalley, U. Prabhu, E. Dunleavy, S. Bijwadia, J. Mao-Jones, K. Chen, R. Pasumarthi, E. Wood, A. Dostmohamed, N. Hurley, J. Simsa, A. Parrish, M. Pajarskas, M. Harvey, O. Skopek, Y. Kochinski, J. Rey, V. Rieser, D. Zhou, S. J. Lee, T. Acharya, G. Li, J. Jiang, X. Zhang, B. Gipson, E. Mahintorabi, M. Gelmi, N. Khajehnouri, A. Yeh, K. Lee, L. Matthey, L. Baker, T. Pham, H. Fu, A. Pak, P. Gupta, C. Vasconcelos, A. Sadovsky, B. Walker, S. Hsiao, P. Zochbauer, A. Marzoca, N. Velan, J. Zeng, G. Baechler, D. Driess, D. Jain, Y. Huang, L. Tao, J. Maggs, N. Levine, J. Schneider, E. Gemzer, S. Petit, S. Han, Z. Fisher, D. Zelle, C. Biles, E. Ie, A. Fadeeva, C. Liu, J. V. Franco, A. Collister, H. Zhang, R. Wang, R. Zhao, L. Kieliger, K. Shuster, R. Zhu, B. Gong, L. Chan, R. Sun, S. Basu, R. Zimmermann, J. Hayes, A. Bapna, J. Snoek, W. Yang, P. Datta, J. A. Abdallah, K. Kilgour, L. Li, S. Mah, Y. Jun, M. Rivière, A. Karmarkar, T. Spalink, T. Huang, L. Gonzalez, D. Tran, A. Nowak, J. Palowitch, M. Chadwick, E. Talius, H. Mehta, T. Sellam, P. Fränken, M. Nicosia, K. He, A. Kini, D. Amos, S. Basu, H. Jobe, E. Shaw, Q. Xu, C. Evans, D. Ikeda, C. Yan, L. Jin, L. Wang, S. Yadav, I. Labzovsky, R. Sampath, A. Ma, C. Schumann, A. Siddhant, R. Shah, J. Youssef, R. Agarwal, N. Dabney, A. Tonioni, M. Ambar, J. Li, I. Guyon, B. Li, D. Soergel, B. Fang, G. Karadzhov, C. Udrescu, T. Trinh, V. Raunak, S. Noury, D. Guo, S. Gupta, M. Finkelstein, D. Petek, L. Liang, G. Billock, P. Sun, D. Wood, Y. Song, X. Yu, T. Matejovicova, R. Cohen, K. Andra, D. D’Ambrosio, Z. Deng, V. Nallatamby, E. Songhori, R. Dangovski, A. Lampinen, P. Botadra, A. Hillier, J. Cao, N. Baddi, A. Kuncoro, T. Yoshino, A. Bhagatwala, M. Ranzato, R. Schaeffer, T. Liu, S. Ye, O. Sarvana, J. Nham, C. Kuang, I. Gao, J. Baek, S. Mittal, A. Wahid, A. Gergely, B. Ni, J. Feldman, C. Muir, P. Lamblin, W. Macherey, E. Dyer, L. Kilpatrick, V. Campos, M. Bhutani, S. Fort, Y. Ahmad, A. Severyn, K. Chatziprimou, O. Ferludin, M. Dimarco, A. Kusupati, J. Heyward, D. Bahir, K. Villela, K. Millican, D. Marcus, S. Bahargam, C. Unlu, N. Roth, Z. Wei, S. Gopal, D. Ghoshal, E. Lee, S. Lin, J. Lees, D. Lee, A. Hosseini, C. Fan, S. Neel, M. Wu, Y. Altun, H. Cai, E. Piqueras, J. Woodward, A. Bissacco, S. Haykal, M. Bordbar, P. Sundaram, S. Hodkinson, D. Toyama, G. Polovets, A. Myers, A. Sinha, T. Levinboim, K. Krishnakumar, R. Chhaparia, T. Sholokhova, N. B. Gundavarapu, G. Jawahar, H. Qureshi, J. Hu, N. Momchev, M. Rahtz, R. Wu, A. P. S, K. Dhamdhere, M. Guo, U. Gupta, A. Eslami, M. Schain, M. Blokzijl, D. Welling, D. Orr, L. Bolelli, N. Perez-Nieves, M. Sirotenko, A. Prasad, A. Kar, B. D. B. Pigem, T. Terzi, G. Weisz, D. Ghosh, A. Mavalankar, D. Madeka, K. Daugaard, H. Adam, V. Shah, D. Berman, M. Tran, S. Baker, E. Andrejczuk, G. Chole, G. Raboshchuk, M. Mirzazadeh, T. Kagohara, S. Wu, C. Schallhart, B. Orlando, C. Wang, A. Rrustemi, H. Xiong, H. Liu, A. Vezer, N. Ramsden, S. Chang, S. Mudgal, Y. Li, N. Vieillard, Y. Hoshen, F. Ahmad, A. Slone, A. Hua, N. Potikha, M. Rossini, J. Stritar, S. Prakash, Z. Wang, X. Dong, A. Nazari, E. Nehoran, K. Tekelioglu, Y. Li, K. Badola, T. Funkhouser, Y. Li, V. Yerram, R. Ganeshan, D. Formoso, K. Langner, T. Shi, H. Li, Y. Yamamori, A. Panda, A. Saade, A. S. Scarpati, C. Breaux, C. Carey, Z. Zhou, C. Hsieh, S. Bridgers, A. Butryna, N. Gupta, V. Tulsyan, S. Woo, E. Eltyshev, W. Grathwohl, C. Parks, S. Benjamin, R. Panigrahy, S. Dodhia, D. D. Freitas, C. Sauer, W. Song, F. Alet, J. Tolins, C. Paduraru, X. Zhou, B. Albert, Z. Zhang, L. Shu, M. Bansal, S. Nguyen, A. Globerson, O. Xiao, J. Manyika, T. Hennigan, R. Rong, J. Matak, A. Bakalov, A. Sharma, D. Sinopalnikov, A. Pierson, S. Roller, G. Brown, M. Gao, T. Fukuzawa, A. Ghafouri, K. Vassigh, I. Barr, Z. Wang, A. Korsun, R. Jayaram, L. Ren, T. Zaman, S. Khan, Y. Lunts, D. Deutsch, D. Uthus, N. Katz, M. Samsikova, A. Khalifa, N. Sethi, J. Sun, L. Tang, U. Alon, X. Luo, D. Yu, A. Nayyar, B. Petrini, W. Truong, V. Hellendoorn, N. Chinaev, C. Alberti, W. Wang, J. Hu, V. Mirrokni, A. Balashankar, A. Aharon, A. Mehta, A. Iscen, J. Kready, L. Manning, A. Mohananey, Y. Chen, A. Tripathi, A. Wu, I. Petrovski, D. Hwang, M. Baeuml, S. Chandrakaladharan, Y. Liu, R. Coaguila, M. Chen, S. Ma, P. Tafti, S. Tatineni, T. Spitz, J. Ye, P. Vicol, M. Rosca, A. Puigdomènech, Z. Yahav, S. Ghemawat, H. Lin, P. Kirk, Z. Nabulsi, S. Brin, B. Bohnet, K. Caluwaerts, A. S. Veerubhotla, D. Zheng, Z. Dai, P. Petrov, Y. Xu, R. Mehran, Z. Xu, L. Zintgraf, J. Choi, S. A. Hombaiah, R. Thoppilan, S. Reddi, L. Lew, L. Li, K. Webster, K. Sawhney, L. Lamprou, S. Shakeri, M. Lunayach, J. Chen, S. Bagri, A. Salcianu, Y. Chen, Y. Donchev, C. Magister, S. Nørly, V. Rodrigues, T. Izo, H. Noga, J. Zou, T. Köppe, W. Zhou, K. Lee, X. Long, D. Eisenbud, A. Chen, C. Schenck, C. M. To, P. Zhong, E. Taropa, M. Truong, O. Levy, D. Martins, Z. Zhang, C. Semturs, K. Zhang, A. Yakubovich, P. Moreno, L. McConnaughey, D. Lu, S. Redmond, L. Weerts, Y. Bitton, T. Refice, N. Lacasse, A. Conmy, C. Tallec, J. Odell, H. Forbes-Pollard, A. Socala, J. Hoech, P. Kohli, A. Walton, R. Wang, M. Sazanovich, K. Zhu, A. Kapishnikov, R. Galt, M. Denton, B. Murdoch, C. Sikora, K. Mohamed, W. Wei, U. First, T. McConnell, L. C. Cobo, J. Qin, T. Avrahami, D. Balle, Y. Watanabe, A. Louis, A. Kraft, S. Ariafar, Y. Gu, E. Rives, C. Yoon, A. Rusu, J. Cobon-Kerr, C. Hahn, J. Luo, Yuvein, Zhu, N. Ahuja, R. Benenson, R. L. Kaufman, H. Yu, L. Hightower, J. Zhang, D. Ni, L. A. Hendricks, G. Wang, G. Yona, L. Jain, P. Barrio, S. Bhupatiraju, S. Velusamy, A. Dafoe, S. Riedel, T. Thomas, Z. Yuan, M. Bellaiche, S. Panthaplackel, K. Kloboves, S. Jauhari, C. Akbulut, T. Davchev, E. Gladchenko, D. Madras, A. Chuklin, T. Hill, Q. Yuan, M. Madhavan, L. Leonhard, D. Scandinaro, Q. Chen, N. Niu, A. Douillard, B. Damoc, Y. Onoe, F. Pedregosa, F. Bertsch, C. Leichner, J. Pagadora, J. Malmaud, S. Ponda, A. Twigg, O. Duzhyi, J. Shen, M. Wang, R. Garg, J. Chen, U. Evci, J. Lee, L. Liu, K. Kojima, M. Yamaguchi, A. Rajendran, A. Piergiovanni, V. K. Rajendran, M. Fornoni, G. Ibagon, H. Ragan, S. M. Khan, J. Blitzer, A. Bunner, G. Sun, T. Kosakai, S. Lundberg, N. Elue, K. Guu, S. Park, J. Park, A. Narayanaswamy, C. Wu, J. Mudigonda, T. Cohn, H. Mu, R. Kumar, L. Graesser, Y. Zhang, R. Killam, V. Zhuang, M. Giménez, W. A. Jishi, R. Ley-Wild, A. Zhai, K. Osawa, D. Cedillo, J. Liu, M. Upadhyay, M. Sieniek, R. Sharma, T. Paine, A. Angelova, S. Addepalli, C. Parada, K. Majumder, A. Lamp, S. Kumar, X. Deng, A. Myaskovsky, T. Sabolić, J. Dudek, S. York, F. de Chaumont Quitry, J. Nie, D. Cattle, A. Gunjan, B. Piot, W. Khawaja, S. Bang, S. Wang, S. Khodadadeh, R. R, P. Rawlani, R. Powell, K. Lee, J. Griesser, G. Oh, C. Magalhaes, Y. Li, S. Tokumine, H. N. Vogel, D. Hsu, A. BC, D. Jindal, M. Cohen, Z. Yang, J. Yuan, D. de Cesare, T. Bruguier, J. Xu, M. Roy, A. Jacovi, D. Belov, R. Arya, P. Meadowlark, S. Cohen-Ganor, W. Ye, P. Morris-Suzuki, P. Banzal, G. Song, P. Ponnuramu, F. Zhang, G. Scrivener, S. Zaiem, A. R. Rochman, K. Han, B. Ghazi, K. Lee, S. Drath, D. Suo, A. Girgis, P. Shenoy, D. Nguyen, D. Eck, S. Gupta, L. Yan, J. Carreira, A. Gulati, R. Sang, D. Mirylenka, E. Cooney, E. Chou, M. Ling, C. Fan, B. Coleman, G. Tubone, R. Kumar, J. Baldridge, F. Hernandez-Campos, A. Lazaridou, J. Besley, I. Yona, N. Bulut, Q. Wellens, A. Pierigiovanni, J. George, R. Green, P. Han, C. Tao, G. Clark, C. You, A. Abdolmaleki, J. Fu, T. Chen, A. Chaugule, A. Chandorkar, A. Rahman, W. Thompson, P. Koanantakool, M. Bernico, J. Ren, A. Vlasov, S. Vassilvitskii, M. Kula, Y. Liang, D. Kim, Y. Huang, C. Ye, D. Lepikhin, and W. Helmholz Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. External Links: 2507.06261, Link Cited by: item •, §4.1. de Haan et al. (2019) P. de Haan, D. Jayaraman, and S. Levine Causal confusion in imitation learning. In Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32, p. . External Links: Link Cited by: §A.4. Fang et al. (2026) R. Fang, Y. Liang, X. Wang, J. Wu, S. Qiao, P. Xie, F. Huang, H. Chen, and N. Zhang Memp: exploring agent procedural memory. External Links: 2508.06433, Link Cited by: §1, §5. Feng et al. (2025) L. Feng, Z. Xue, T. Liu, and B. An Group-in-group policy optimization for llm agent training. In Advances in Neural Information Processing Systems, D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen (Eds.), Vol. 38, p. 46375–46408. External Links: Link Cited by: §B.2, §1, §5. Hübotter et al. (2026) J. Hübotter, F. Lübeck, L. Behric, A. Baumann, M. Bagatella, D. Marta, I. Hakimi, I. Shenfeld, T. K. Buening, C. Guestrin, and A. Krause Reinforcement learning via self-distillation. External Links: 2601.20802, Link Cited by: §5. Jin et al. (2025) B. Jin, H. Zeng, Z. Yue, J. Yoon, S. Arik, D. Wang, H. Zamani, and J. Han Search-r1: training llms to reason and leverage search engines with reinforcement learning. External Links: 2503.09516, Link Cited by: §1, §5. Kaelbling et al. (1998) L. P. Kaelbling, M. L. Littman, and A. R. Cassandra Planning and acting in partially observable stochastic domains. Artificial Intelligence 101 (1–2), p. 99–134. Cited by: §2. Liu et al. (2026) Z. Liu, J. Kim, X. Luo, D. Li, and Y. Yang Exploratory memory-augmented llm agent via hybrid on- and off-policy optimization. External Links: 2602.23008, Link Cited by: item •, §1, §4.1, §5, §5. Lu et al. (2026a) Z. Lu, Z. Yao, Z. Han, Z. Wang, J. Wu, Q. Gu, X. Cai, W. Lu, J. Xiao, Y. Zhuang, and Y. Shen Self-distilled agentic reinforcement learning. External Links: 2605.15155, Link Cited by: Table 1, §5. Lu et al. (2026b) Z. Lu, Z. Yao, J. Wu, C. Han, Q. Gu, X. Cai, W. Lu, J. Xiao, Y. Zhuang, and Y. Shen SKILL0: in-context agentic reinforcement learning for skill internalization. External Links: 2604.02268, Link Cited by: §5. Ma et al. (2026) W. Ma, Y. Zeng, Y. Song, X. Cui, J. Zhao, X. Liu, and M. Elhoseiny Freshness-aware prioritized experience replay for llm/vlm reinforcement learning. External Links: 2604.16918, Link Cited by: §5. OpenAI et al. (2024) OpenAI, :, A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, A. Mądry, A. Baker-Whitcomb, A. Beutel, A. Borzunov, A. Carney, A. Chow, A. Kirillov, A. Nichol, A. Paino, A. Renzin, A. T. Passos, A. Kirillov, A. Christakis, A. Conneau, A. Kamali, A. Jabri, A. Moyer, A. Tam, A. Crookes, A. Tootoochian, A. Tootoonchian, A. Kumar, A. Vallone, A. Karpathy, A. Braunstein, A. Cann, A. Codispoti, A. Galu, A. Kondrich, A. Tulloch, A. Mishchenko, A. Baek, A. Jiang, A. Pelisse, A. Woodford, A. Gosalia, A. Dhar, A. Pantuliano, A. Nayak, A. Oliver, B. Zoph, B. Ghorbani, B. Leimberger, B. Rossen, B. Sokolowsky, B. Wang, B. Zweig, B. Hoover, B. Samic, B. McGrew, B. Spero, B. Giertler, B. Cheng, B. Lightcap, B. Walkin, B. Quinn, B. Guarraci, B. Hsu, B. Kellogg, B. Eastman, C. Lugaresi, C. Wainwright, C. Bassin, C. Hudson, C. Chu, C. Nelson, C. Li, C. J. Shern, C. Conger, C. Barette, C. Voss, C. Ding, C. Lu, C. Zhang, C. Beaumont, C. Hallacy, C. Koch, C. Gibson, C. Kim, C. Choi, C. McLeavey, C. Hesse, C. Fischer, C. Winter, C. Czarnecki, C. Jarvis, C. Wei, C. Koumouzelis, D. Sherburn, D. Kappler, D. Levin, D. Levy, D. Carr, D. Farhi, D. Mely, D. Robinson, D. Sasaki, D. Jin, D. Valladares, D. Tsipras, D. Li, D. P. Nguyen, D. Findlay, E. Oiwoh, E. Wong, E. Asdar, E. Proehl, E. Yang, E. Antonow, E. Kramer, E. Peterson, E. Sigler, E. Wallace, E. Brevdo, E. Mays, F. Khorasani, F. P. Such, F. Raso, F. Zhang, F. von Lohmann, F. Sulit, G. Goh, G. Oden, G. Salmon, G. Starace, G. Brockman, H. Salman, H. Bao, H. Hu, H. Wong, H. Wang, H. Schmidt, H. Whitney, H. Jun, H. Kirchner, H. P. de Oliveira Pinto, H. Ren, H. Chang, H. W. Chung, I. Kivlichan, I. O’Connell, I. O’Connell, I. Osband, I. Silber, I. Sohl, I. Okuyucu, I. Lan, I. Kostrikov, I. Sutskever, I. Kanitscheider, I. Gulrajani, J. Coxon, J. Menick, J. Pachocki, J. Aung, J. Betker, J. Crooks, J. Lennon, J. Kiros, J. Leike, J. Park, J. Kwon, J. Phang, J. Teplitz, J. Wei, J. Wolfe, J. Chen, J. Harris, J. Varavva, J. G. Lee, J. Shieh, J. Lin, J. Yu, J. Weng, J. Tang, J. Yu, J. Jang, J. Q. Candela, J. Beutler, J. Landers, J. Parish, J. Heidecke, J. Schulman, J. Lachman, J. McKay, J. Uesato, J. Ward, J. W. Kim, J. Huizinga, J. Sitkin, J. Kraaijeveld, J. Gross, J. Kaplan, J. Snyder, J. Achiam, J. Jiao, J. Lee, J. Zhuang, J. Harriman, K. Fricke, K. Hayashi, K. Singhal, K. Shi, K. Karthik, K. Wood, K. Rimbach, K. Hsu, K. Nguyen, K. Gu-Lemberg, K. Button, K. Liu, K. Howe, K. Muthukumar, K. Luther, L. Ahmad, L. Kai, L. Itow, L. Workman, L. Pathak, L. Chen, L. Jing, L. Guy, L. Fedus, L. Zhou, L. Mamitsuka, L. Weng, L. McCallum, L. Held, L. Ouyang, L. Feuvrier, L. Zhang, L. Kondraciuk, L. Kaiser, L. Hewitt, L. Metz, L. Doshi, M. Aflak, M. Simens, M. Boyd, M. Thompson, M. Dukhan, M. Chen, M. Gray, M. Hudnall, M. Zhang, M. Aljubeh, M. Litwin, M. Zeng, M. Johnson, M. Shetty, M. Gupta, M. Shah, M. Yatbaz, M. J. Yang, M. Zhong, M. Glaese, M. Chen, M. Janner, M. Lampe, M. Petrov, M. Wu, M. Wang, M. Fradin, M. Pokrass, M. Castro, M. O. T. de Castro, M. Pavlov, M. Brundage, M. Wang, M. Khan, M. Murati, M. Bavarian, M. Lin, M. Yesildal, N. Soto, N. Gimelshein, N. Cone, N. Staudacher, N. Summers, N. LaFontaine, N. Chowdhury, N. Ryder, N. Stathas, N. Turley, N. Tezak, N. Felix, N. Kudige, N. Keskar, N. Deutsch, N. Bundick, N. Puckett, O. Nachum, O. Okelola, O. Boiko, O. Murk, O. Jaffe, O. Watkins, O. Godement, O. Campbell-Moore, P. Chao, P. McMillan, P. Belov, P. Su, P. Bak, P. Bakkum, P. Deng, P. Dolan, P. Hoeschele, P. Welinder, P. Tillet, P. Pronin, P. Tillet, P. Dhariwal, Q. Yuan, R. Dias, R. Lim, R. Arora, R. Troll, R. Lin, R. G. Lopes, R. Puri, R. Miyara, R. Leike, R. Gaubert, R. Zamani, R. Wang, R. Donnelly, R. Honsby, R. Smith, R. Sahai, R. Ramchandani, R. Huet, R. Carmichael, R. Zellers, R. Chen, R. Chen, R. Nigmatullin, R. Cheu, S. Jain, S. Altman, S. Schoenholz, S. Toizer, S. Miserendino, S. Agarwal, S. Culver, S. Ethersmith, S. Gray, S. Grove, S. Metzger, S. Hermani, S. Jain, S. Zhao, S. Wu, S. Jomoto, S. Wu, Shuaiqi, Xia, S. Phene, S. Papay, S. Narayanan, S. Coffey, S. Lee, S. Hall, S. Balaji, T. Broda, T. Stramer, T. Xu, T. Gogineni, T. Christianson, T. Sanders, T. Patwardhan, T. Cunninghman, T. Degry, T. Dimson, T. Raoux, T. Shadwell, T. Zheng, T. Underwood, T. Markov, T. Sherbakov, T. Rubin, T. Stasi, T. Kaftan, T. Heywood, T. Peterson, T. Walters, T. Eloundou, V. Qi, V. Moeller, V. Monaco, V. Kuo, V. Fomenko, W. Chang, W. Zheng, W. Zhou, W. Manassra, W. Sheu, W. Zaremba, Y. Patil, Y. Qian, Y. Kim, Y. Cheng, Y. Zhang, Y. He, Y. Zhang, Y. Jin, Y. Dai, and Y. Malkov GPT-4o system card. External Links: 2410.21276, Link Cited by: item •, §B.2, §4.1, §4.1. Ouyang et al. (2026) S. Ouyang, J. Yan, I. Hsu, Y. Chen, K. Jiang, Z. Wang, R. Han, L. T. Le, S. Daruki, X. Tang, V. Tirumalashetty, G. Lee, M. Rofouei, H. Lin, J. Han, C. Lee, and T. Pfister ReasoningBank: scaling agent self-evolving with reasoning memory. External Links: 2509.25140, Link Cited by: §5. Pinto et al. (2017) L. Pinto, M. Andrychowicz, P. Welinder, W. Zaremba, and P. Abbeel Asymmetric actor critic for image-based robot learning. External Links: 1710.06542, Link Cited by: §5. Qin et al. (2025) Y. Qin, X. Tan, Z. He, G. Li, H. Lin, Z. Li, Z. Xu, Y. Shi, S. Cai, R. Rui, S. Cai, Y. Cai, X. Zhang, S. Ye, K. Li, and X. Sun Learn the ropes, then trust the wins: self-imitation with progressive exploration for agentic reinforcement learning. External Links: 2509.22601, Link Cited by: §5. Qwen et al. (2025) Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu Qwen2.5 technical report. External Links: 2412.15115, Link Cited by: §4.1. Ross et al. (2011) S. Ross, G. Gordon, and D. Bagnell A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, G. Gordon, D. Dunson, and M. Dudík (Eds.), Proceedings of Machine Learning Research, Vol. 15, Fort Lauderdale, FL, USA, p. 627–635. External Links: Link Cited by: §A.4. Schulman et al. (2017) J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov Proximal policy optimization algorithms. External Links: 1707.06347, Link Cited by: §5. Shao et al. (2024) Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo DeepSeekMath: pushing the limits of mathematical reasoning in open language models. External Links: 2402.03300, Link Cited by: item •, §1, §2, §4.1, §5. Shinn et al. (2023) N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao Reflexion: language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), External Links: Link Cited by: item •, §1, §1, §4.1, §5. Shridhar et al. (2021) M. Shridhar, X. Yuan, M. Côté, Y. Bisk, A. Trischler, and M. J. Hausknecht ALFWorld: aligning text and embodied environments for interactive learning. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, External Links: Link Cited by: Appendix F, §1, §1, §4.1, §5. Snell et al. (2022) C. Snell, D. Klein, and R. Zhong Learning by distilling context. External Links: 2209.15189, Link Cited by: §5. Vapnik and Vashist (2009) V. Vapnik and A. Vashist A new learning paradigm: learning using privileged information. Neural Networks 22 (5), p. 544–557. Note: Advances in Neural Networks Research: IJCNN2009 External Links: ISSN 0893-6080, Document, Link Cited by: §5. Wang et al. (2025a) H. Wang, C. T. Leong, J. Wang, J. Wang, and W. Li SPA-rl: reinforcing llm agents via stepwise progress attribution. External Links: 2505.20732, Link Cited by: §5. Wang et al. (2026) J. Wang, Q. Yan, Y. Wang, Y. Tian, S. S. Mishra, Z. Xu, M. Gandhi, P. Xu, and L. L. Cheong Reinforcement learning for self-improving agent with skill library. External Links: 2512.17102, Link Cited by: §5. Wang et al. (2025b) Z. Wang, K. Wang, Q. Wang, P. Zhang, L. Li, Z. Yang, X. Jin, K. Yu, M. N. Nguyen, L. Liu, E. Gottlieb, Y. Lu, K. Cho, J. Wu, L. Fei-Fei, L. Wang, Y. Choi, and M. Li RAGEN: understanding self-evolution in llm agents via multi-turn reinforcement learning. External Links: 2504.20073, Link Cited by: §1, §5. Wei et al. (2025) Q. Wei, S. Zeng, C. Li, W. Brown, O. Frunza, W. Deng, A. Schneider, Y. Nevmyvaka, Y. K. Zhao, A. Garcia, and M. Hong Reinforcing multi-turn reasoning in llm agents via turn-level reward design. External Links: 2505.11821, Link Cited by: §5. Wu et al. (2026) R. Wu, X. Wang, J. Mei, P. Cai, D. Fu, C. Yang, L. Wen, X. Yang, Y. Shen, Y. Wang, and B. Shi EvolveR: self-evolving llm agents through an experience-driven lifecycle. External Links: 2510.16079, Link Cited by: item •, §4.1, §5. Xia et al. (2026) P. Xia, J. Chen, H. Wang, J. Liu, K. Zeng, Y. Wang, S. Han, Y. Zhou, X. Zhao, H. Chen, Z. Zheng, C. Xie, and H. Yao SkillRL: evolving agents via recursive skill-augmented reinforcement learning. External Links: 2602.08234, Link Cited by: item •, §1, §3.3, Table 1, §4.1, §5. Yan et al. (2025) J. Yan, Y. Li, Z. Hu, Z. Wang, G. Cui, X. Qu, Y. Cheng, and Y. Zhang Learning to reason under off-policy guidance. External Links: 2504.14945, Link Cited by: §5. Yang et al. (2024) L. Yang, Z. Yu, T. Zhang, S. Cao, M. Xu, W. Zhang, J. E. Gonzalez, and B. Cui Buffer of thoughts: thought-augmented reasoning with large language models. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), External Links: Link Cited by: §5. Yao et al. (2022) S. Yao, H. Chen, J. Yang, and K. Narasimhan WebShop: towards scalable real-world web interaction with grounded language agents. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), External Links: Link Cited by: Appendix F, §1, §1, §4.1, §5. Yao et al. (2023) S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao ReAct: synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, External Links: Link Cited by: item •, §1, §4.1. Ye et al. (2026) T. Ye, L. Dong, X. Wu, S. Huang, and F. Wei On-policy context distillation for language models. External Links: 2602.12275, Link Cited by: §5. Zhan et al. (2026) R. Zhan, Y. Li, Z. Wang, X. Qu, D. Liu, J. Shao, D. F. Wong, and Y. Cheng ExGRPO: learning to reason from experience. External Links: 2510.02245, Link Cited by: §5. Zhang et al. (2026a) H. Zhang, Q. Long, J. Bao, T. Feng, W. Zhang, H. Yue, and W. Wang MemSkill: learning and evolving memory skills for self-evolving agents. External Links: 2602.02474, Link Cited by: §5. Zhang et al. (2025a) K. Zhang, A. Lv, J. Li, Y. Wang, F. Wang, H. Hu, and R. Yan StepHint: multi-level stepwise hints enhance reinforcement learning to reason. External Links: 2507.02841, Link Cited by: §5. Zhang et al. (2026b) S. Zhang, Y. Xiong, X. Chen, Z. Jia, R. Huang, J. Xu, and J. Zhang RAPO: expanding exploration for llm agents via retrieval-augmented policy optimization. External Links: 2603.03078, Link Cited by: §5. Zhang et al. (2025b) Y. Zhang, M. Li, D. Long, X. Zhang, H. Lin, B. Yang, P. Xie, A. Yang, D. Liu, J. Lin, F. Huang, and J. Zhou Qwen3 embedding: advancing text embedding and reranking through foundation models. External Links: 2506.05176, Link Cited by: §B.2, §4.1. Zhao et al. (2024) A. Zhao, D. Huang, Q. Xu, M. Lin, Y. Liu, and G. Huang ExpeL: LLM agents are experiential learners. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada, M. J. Wooldridge, J. G. Dy, and S. Natarajan (Eds.), p. 19632–19642. External Links: Link, Document Cited by: §1, §5. Zhao et al. (2026) S. Zhao, Z. Xie, M. Liu, J. Huang, G. Pang, F. Chen, and A. Grover Self-distilled reasoner: on-policy self-distillation for large language models. External Links: 2601.18734, Link Cited by: item •, §4.1, §5. Appendix A Theoretical Analysis of EDGE This section provides an analytical justification for the four core components of EDGE, following the pipeline through which a retrieved experience e influences the policy update: its marginal utility is estimated (§A.1), smoothed across training steps (§A.2), used to gate and calibrate the RL advantage (§A.3), and channeled through on-support distillation (§A.4). Throughout, let c S denote the standard context, c=c⊕ec T=c S e the privileged context, πθ _θ the current policy, and R(τ)∈0,1R(τ)∈\0,1\ the binary outcome reward. Teacher and student rollout subsets ,T T,T S each have size K=G/2K=G/2. All expectations and probabilities condition on fixed πθ _θ and e. A.1 Marginal-Gain Estimation and Gating Define the value of privileged information as the expected return gap between the privileged and standard contexts: e(θ)=τ∼πθ(⋅∣c)[R(τ)]−τ∼πθ(⋅∣c)[R(τ)].V_e(θ)\;=\;E_τ _θ(· c T)\! [R(τ) ]\;-\;E_τ _θ(· c S)\! [R(τ) ]. (10) The empirical marginal gain is Δe=1K∑τ∈R(τ)−1K∑τ∈R(τ). _e\;=\; 1K\! _τ T\!R(τ)\;-\; 1K\! _τ S\!R(τ). (11) Under the assumption that the two subsets are conditionally independent and identically distributed up to the presence of e, Δe _e is unbiased for e(θ)V_e(θ) with variance (pT(1−pT)+pS(1−pS))/K(p_T(1-p_T)+p_S(1-p_S))/K, where pTp_T and pSp_S are the respective success probabilities. Decomposing Δe _e as a sum of 2K2K independent terms each bounded in an interval of length 1/K1/K, Hoeffding’s inequality yields Pr(|Δe−e(θ)|≥ϵ)≤ 2exp(−Kϵ2). \! (| _e-V_e(θ)|≥ε )\;≤\;2 \! (-Kε^2 ). (12) EDGE converts this estimate into a binary gate: Me=(Δe>0).M_e\;=\;I( _e>0). (13) The gate serves as a probabilistic risk-control mechanism. By the one-sided form of (12), if an experience is truly harmful with margin γ>0γ>0 (e(θ)≤−γV_e(θ)≤-γ), then Pr(Me=1)≤exp(−Kγ2) (M_e=1)≤ (-Kγ^2); the symmetric bound holds for missed activations when e(θ)≥γV_e(θ)≥γ. When Me=0M_e=0, teacher rollouts are masked from both the RL and distillation losses, and the update falls back to standard GRPO on the student subset. A.2 EMA Smoothing Because Δe _e is high-variance for small K and the true utility e(θ)V_e(θ) drifts as the policy evolves, EDGE tracks each experience’s utility with an exponential moving average: Ue(t)=(1−μ)Ue(t−1)+μΔe(t),μ∈(0,1).U_e^(t)\;=\;(1-μ)\,U_e^(t-1)\;+\;μ\, _e^(t), μ∈(0,1). (14) Under a local-stationarity approximation (successive Δe(t) _e^(t) treated as i.i.d. with variance σe2 _e^2), the steady-state variance is Var(Ue)=μ2−μσe2Var(U_e)= μ2-μ\, _e^2, which is strictly less than σe2 _e^2 for any μ<1μ<1. Smaller μ yields greater smoothing at the cost of slower adaptation to genuine utility shifts. The pruning threshold η thus operates on a smoothed estimate of recent marginal utility rather than a single noisy contrast. A.3 Pooled Advantage Calibration When the gate activates, teacher rollouts enter the RL update and alter the advantage baseline. Let bS,bTb_S,b_T be the student and teacher mean returns, so Δe=bT−bS _e=b_T-b_S. The active set is =∪,Me=1,,Me=0,A\;=\; casesT S T,&M_e=1,\\[2.0pt] T S,&M_e=0, cases (15) with baseline b=meanj∈(Rj)b_A=mean_j (R_j). When Me=1M_e=1, the pooled baseline bpool=12(bS+bT)b_pool= 12(b_S+b_T) shifts the raw advantage numerators (prior to GRPO’s standard-deviation normalization, which rescales uniformly without changing signs) as follows: A~SEDGE A_S EDGE =A~Swithin−12Δe, = A_S^within- 12 _e, (16) A~TEDGE A_T EDGE =A~Twithin+12Δe, = A_T^within+ 12 _e, (17) where the superscript “within” denotes baselines computed from each subset alone. Since Me=1M_e=1 implies Δe>0 _e>0, student advantages are shifted downward and teacher advantages upward: privileged successes receive stronger reinforcement, while unguided successes are tempered when the scaffold demonstrates superior performance. When Me=0M_e=0, the teacher subset is excluded and no shift is applied. A.4 On-Support Reverse-KL Distillation Directly imitating teacher-generated trajectories risks covariate shift (22), as the teacher may visit prefixes whose rationale depends on information absent from the student context (causal misidentification; 7). EDGE mitigates this by evaluating the teacher only on student-generated prefixes. For a student trajectory τ=(y1,…,ym)∈τ=(y_1,…,y_m) S, let ht=(c,y<t)h_t S=(c S,y_<t) and ht=(c,y<t)h_t T=(c T,y_<t). The on-support distillation loss is ℒ^distill(θ)=Me∑τ∈∑t=1|τ|DKL(πθ(⋅∣ht)∥πsg(θ)(⋅∣ht)), split L_distill(θ)\;=\;\;&M_e\! _τ S _t=1^|τ|\\ &D_KL\! ( _θ(· h_t S)\; \|\; _sg(θ)(· h_t T) ), split (18) where sg(⋅)sg(·) denotes stop-gradient. Unlike forward-KL imitation on teacher rollouts, this objective compares distributions only at prefixes the student has actually reached, avoiding training on teacher-only states. The reverse-KL direction is mode-seeking with respect to the teacher, concentrating student mass on teacher-preferred actions rather than spreading to cover the full teacher distribution, and thereby keeping the update conservative on the student support. The final actor loss combines both pathways: ℒactor(θ)=ℒRL(θ,)+λℒ^distill(θ).L_actor(θ)\;=\;L_RL(θ;A)\;+\;λ\, L_distill(θ). (19) When Me=0M_e=0, both terms reduce to unprivileged-only updates (=A=T S, ℒ^distill=0 L_distill=0). When Me=1M_e=1, the teacher subset raises the RL baseline (§A.3) and the reverse KL transfers privileged behavior into the standard-context policy. Appendix B Implementation Details B.1 Baselines • GPT-4o (17): A closed-source multimodal LLM from OpenAI, used as a strong proprietary agent baseline with standard prompting. • Gemini-2.5-Pro (6): A closed-source reasoning model from Google DeepMind, serving as another proprietary agent baseline with standard prompting. • ReAct (38): A prompting framework that interleaves chain-of-thought reasoning with environment actions, enabling LLMs to plan and act in a synergistic loop. • Reflexion (25): Extends ReAct by appending verbal self-reflection after task failures, allowing the agent to refine its strategy across successive trials without weight updates. • GRPO (24): A group-relative policy optimization algorithm that estimates advantages from a group of sampled rollouts, eliminating the need for a separate critic network. • OPSD (46): An on-policy self-distillation method that distills the model’s own high-quality rollouts back into itself to improve reasoning without external supervision. • GRPO+OPSD: A hybrid baseline that combines GRPO’s group-relative advantage estimation with OPSD’s on-policy self-distillation objective. • EvolveR (33): An experience-augmented training method that iteratively evolves a retrieval-augmented memory of past trajectories to guide policy learning. • GRPO+Mem0 (4): Augments GRPO with Mem0, a memory module that stores and retrieves past interaction experiences as additional context during both training and inference. • SkillRL (34): A skill-based RL framework that extracts reusable skills from successful trajectories and conditions policy optimization on retrieved skill demonstrations. • EMPO2 (13): An exploratory memory-augmented policy optimization method that leverages curated past experiences to enhance exploration during RL training. These baselines span closed-source LLM agents, prompting-based reasoning frameworks, standard RL algorithms, self-distillation methods, and experience-augmented training approaches, enabling a comprehensive evaluation of EDGE from multiple perspectives. B.2 RL-Training Configuration We implement EDGE on top of the verl-agent framework (9) and train the model with the joint optimization objective in Eq. 19. To reduce the computational cost of reverse-KL distillation, we approximate the vocabulary-level loss using only the top-k student tokens with the highest log probabilities. The experience bank is initialized as empty, and each newly inserted experience is assigned an initial utility score of 0. For retrieval, we use the task instruction and initial environment observations as the query, encode experiences with Qwen3-Embedding-0.6B (44), and rank candidates by cosine similarity. We retrieve a top-m candidate pool and use the highest-scoring experience as the scaffold for the current rollout. After each training step, we update the experience bank through both expansion and pruning. For expansion, if the success rate of a task category falls below the expansion threshold ξ, GPT-4o (17) is used as the reflector LLM to contrast successful and failed trajectories from the same category and synthesize up to three new experiences. For pruning, we update the EMA utility score of each experience according to Eq. 14 and remove experiences whose utility falls below the pruning threshold η. For both ALFWorld and WebShop, we use the hyperparameters in Table 3. All training experiments are conducted on 8×808× 80GB GPUs. Configuration Value RL-Training Actor learning rate 1e−61e^-6 Maximum prompt length 4096 Maximum response length 512 Training batch size 16 Rollout group size (G) 8 Training mini-batch size 128 Rollout temperature 1.0 Training steps 200 EDGE configuration Distillation weight λ 0.1 Distillation top-k tokens 20 Utility pruning threshold η -0.1 Retrieval pool size (top-m) 6 Expansion success threshold ξ 0.4 EMA momentum μ 0.5 Maximum new experiences per step 3 Table 3: RL Hyperparameters Case Failure mode Reflected experience Later evidence 1. Clean bowl Navigates to destination before finding object Locate object before placing Retrieved on a related task at step 155; U=0.125U=0.125 2. Pillow on sofa Issues take from source lacking target Verify source before taking Utility rises to 0.5780.578 by step 200 (78 retrievals) 3. Pick2 search Unstructured search before collecting targets Search systematically before storing Retrieved on a related Pick2 task at step 166 4. No-op movement Repeats navigation after arrival Avoid repeated no-op moves Near-threshold utility; recovers to U=0.452U=0.452 by step 200 Table 4: Experience-bank co-evolution examples. Each case is extracted from saved failure trajectories, LLM reflection logs, retrieval logs, and utility traces. Appendix C Further Analysis Figure 5: Sensitivity of validation success rate to distillation weight λ. C.1 Sensitivity to Distillation Weight λ. Figure 5 sweeps the distillation coefficient λ on Qwen2.5-1.5B-Instruct, revealing a clear trade-off between the RL and distillation objectives. At λ=1λ=1, the distillation term dominates the gradient and effectively freezes the policy near its initial performance (∼ 15%), preventing autonomous exploration beyond the scaffold-prescribed behavior. At λ=0.01λ=0.01, the policy recovers RL-driven improvement but internalizes scaffold-induced patterns too slowly, converging roughly 5 points below the best setting. The moderate value λ=0.1λ=0.1 balances both pressures, sustaining steady improvement to approximately 80%—enough distillation to accelerate internalization without suppressing the RL objective’s exploratory signal. We adopt this value for all remaining experiments. C.2 Case Study: Experience Bank Evolution Dynamics This section complements the aggregate gain-tracking curves in Figure 4 with entry-level evidence from ALFWorld training logs (Qwen2.5-7B-Instruct). We trace four representative cases (Table 4), each linking a rollout failure to the experience it produces, its later retrieval, and the resulting change in agent behavior. Cases 1–2 (Figure 6) illustrate the expansion phase, while Cases 3–4 (Figure 7) illustrate late-stage refinement and non-stationary utility. Case 1: From destination-first wandering to object-first execution Failure (step 150). The task is put a clean bowl in shelf. The failed rollout immediately navigates to the destination and loops around empty shelves: go to shelf 3 → go to cabinet 1 → go to shelf 3 → go to cabinet 3 → examine shelf 3. Successful contrast. The agent first searches likely sources, finds a bowl in the fridge, takes and cleans it, then navigates to a shelf: go to fridge 1 → open fridge 1 → take bowl 1 from fridge 1 → clean bowl 1 with sinkbasin 1 → go to shelf 1. Reflected experience. Entry task_943: “First identify where the target bowl is, retrieve it, clean it if the goal requires a clean bowl, and only then move it to a shelf.” Later retrieval. At step 155, retrieved for put a clean bowl in diningtable (similarity 0.8860.886). The rollout no longer starts by visiting the table; it locates, cleans, and places the bowl. Utility: U=0.125U=0.125. Case 2: From hallucinated source to source verification Failure (step 58). The task is put some pillow on sofa. The agent goes to the sofa, observes box, creditcard, keychain, and newspaper—but no pillow. It issues take pillow from sofa 1, receives Nothing happens, and repeats similar invalid source assumptions. Successful contrast. The paired trajectory searches alternative sources, reaches armchair 1, observes pillow 1, takes it, and moves it to the sofa—only issuing take when the observation lists the target. Reflected experience. Entry step_1067: “Only use take <object> from <container> when the object is actually listed at that location; if not observed, do not repeat the same invalid action.” Later retrieval. At step 60, retrieved for put some keychain on ottoman, where the observation lacks the keychain—the same structural error. Utility evolves from 0.0000.000 at creation to 0.5780.578 at step 200 after 78 retrievals. Figure 6: Experience bank evolution: early-stage expansion. Case 1 illustrates destination-first wandering corrected by an object-first scaffold; Case 2 illustrates hallucinated source actions corrected by an observation-gated rule. Case 3: From unstructured Pick2 search to systematic collection Failure (step 165). For find two ladle and put them in drawer, the policy spends actions on destination drawers or repeated navigation before confirming where the two targets are. Unlike early scaffolds, this failure involves managing object count, source search, and final placement simultaneously. Reflected experience. Entry task_995: “Identify likely locations of the target item, inspect those locations methodically, pick up each required object, and then place them into the specified container. Avoid random movement or opening unrelated containers before locating the targets.” Later retrieval. At step 166, retrieved for find two spatula and put them in drawer (similarity 0.8540.854). The rollout locates spatula 3 on a countertop, picks it up, checks other locations, and deposits it in a drawer. Utility reaches U=0.375U=0.375 by steps 195–200. Case 4: Utility tracking separates useful retrieval from harmful repetition Creation (step 144). Entry step_914, “Avoid repeating a no-op movement,” is reflected from a find two tissuebox and put them in drawer failure. The rule: if the agent is already at the target location, it should inspect or act rather than re-issuing the same navigation action. Non-stationary utility. The entry is retrieved frequently but its utility is initially unstable: U=−0.026U=-0.026 at step 150 and −0.085-0.085 at step 165—hovering near the pruning threshold η=−0.1η=-0.1 but remaining above it. Frequency-based retention would treat this entry as important despite its near-zero or negative marginal gain; conversely, a positive threshold would have discarded it prematurely. Recovery. As the policy reaches more states where repeated no-op movements are the dominant error, the entry becomes useful: utility turns positive at U=0.048U=0.048 by step 180 and rises to U=0.452U=0.452 at step 200 after 142 retrievals. This illustrates why EDGE combines smoothed EMA tracking (Eq. (14)) with a mildly negative pruning threshold: experience utility can shift as the policy distribution changes, and premature pruning would discard entries whose value has not yet materialized. Figure 7: Experience bank evolution: late-stage refinement and non-stationary utility. Case 3 illustrates specialized multi-object coordination guidance; Case 4 illustrates an entry whose utility is initially near the pruning threshold but recovers as the policy’s failure distribution shifts. The four cases reveal a natural curriculum driven by the policy’s evolving failure distribution. Early failures are structural and broadly shared: destination-first wandering (Case 1) and hallucinated object presence (Case 2) affect many task variants, producing generic scaffolds with sustained high utility—these entries drive the initial bank growth visible in Figure 4. As the policy internalizes these broad strategies, the remaining failures narrow in scope. By step 165, single-object manipulation is reliable but coordinating two targets remains fragile; Case 3’s search-then-collect scaffold addresses precisely this gap, and its moderate final utility (U=0.375U=0.375) is consistent with Pick2 remaining the hardest subtask in Table 2. Case 4 provides the most direct evidence for non-stationary utility. Entry step_914 (“avoid repeating no-op movements”) hovers near the pruning boundary for roughly 20 steps (U=−0.026U=-0.026 at step 150, −0.085-0.085 at step 165) before recovering to U=0.452U=0.452 by step 200. The mechanism is interpretable: once destination-first errors (Case 1) are resolved, no-op navigation loops become the dominant Pick2 failure mode, re-activating a previously marginal entry. A positive pruning threshold would have discarded it; the adopted η=−0.1η=-0.1 retains such latent-value entries until the policy’s distribution shifts in their favor. Together, these cases explain the aggregate bank trajectory in Figure 4—initial growth as generic scaffolds accumulate, followed by contraction as the policy absorbs broad patterns and the bank converges to a compact set of specialized, currently relevant entries. Appendix D Pseudocode Algorithm 1 outlines the full EDGE training loop, which alternates among three stages per iteration: experience-guided exploration scaffolding (§3.1), gain-gated privileged distillation (§3.2), and experience bank evolution (§3.3). Algorithm 1 EDGE: Experience-Distillation for Guided Exploration 1: 2: Policy πθ _θ; experience bank ℰE; 3: Rollout group size G; distillation weight λ; 4: EMA momentum μ; pruning threshold η; 5: Success-rate threshold ξ. 6: for each training iteration do 7: for each task x in batch do 8: Stage 1: Experience-Guided Exploration Scaffolding 9: Retrieve top experience e∈ℰe by embedding similarity. 10: Sample G trajectories from πθ _θ: 11: T T (G/2G/2): conditioned on c=x⊕ec T=x e 12: T S (G/2G/2): conditioned on c=xc S=x 13: Compute marginal gain Δe _e via Eq. (5). 14: Stage 2: Gain-Gated Privileged Distillation 15: if Δe>0 _e>0 then 16: Compute advantages over all G trajectories. 17: for each τ∈τ S do 18: Forward πsg(θ) _sg(θ) on τ with c T. 19: Compute ℒdistillL_distill via Eq. (18). 20: end for 21: else 22: Compute advantages over T S only. 23: ℒdistill←0L_distill← 0. 24: end if 25: ℒactor←ℒRL+λℒdistillL_actor _RL+λ\,L_distill. 26: Stage 3: Experience Bank Evolution 27: Update EMA utility: Ue(t)←(−μ)Ue(t−1)+μΔeU_e^(t)\!←\!(1\!-\!μ)\,U_e^(t-1)+μ\, _e. 28: end for 29: Update πθ _θ with ℒactorL_actor. 30: Prune experiences with Ue(t)<ηU_e^(t)<η from ℰE. 31: Identify task categories with success rate <ξ<ξ. 32: Generate new experiences via freflect(τ+,τ−)f_reflect(τ^+,τ^-); insert into ℰE. 33: end for 34: Trained policy πθ _θ (deployed without ℰE). Appendix E Prompts This section provides the prompt templates used in our experiments. We include both rollout prompts and experience update prompts for ALFWorld and WebShop. The rollout prompts are used for action generation, while the update prompts are used to generate new state-aware experiences from contrasted trajectories. E.1 ALFWorld Figure 8 shows the ALFWorld rollout prompts. The first prompt provides retrieved experiences as additional context, while the second prompt removes external experiences and asks the agent to act only based on the current task, history, observation, and admissible actions. Prompt: ALFWorld Agent Execution with Experience System Prompt: You are an expert agent operating in the ALFRED Embodied Environment. Your task is to: task_description # Retrieved Relevant Experience retrieved_experiences Warning: These experiences may be outdated. Use them only if they align with your current observation. # Current Progress Prior to this step, you have already taken step_count step(s). Below are the most recent history_length observations and the corresponding actions you took: action_history You are now at step current_step and your current observation is: current_observation Your admissible actions of the current situation are: [admissible_actions]. Now it is your turn to take an action. You should first reason step-by-step about the current situation. This reasoning process MUST be enclosed within <think> </think> tags. Once you have finished your reasoning, you should choose an admissible action for the current step and present it within <action> </action> tags. Prompt: ALFWorld Agent Execution without Experience System Prompt: You are an expert agent operating in the ALFRED Embodied Environment. Your task is to: task_description # Retrieved Relevant Experience No external general and task-specific experiences are provided. Use your learned strategy and current observation. # Current Progress Prior to this step, you have already taken step_count step(s). Below are the most recent history_length observations and the corresponding actions you took: action_history You are now at step current_step and your current observation is: current_observation Your admissible actions of the current situation are: [admissible_actions]. Now it is your turn to take an action. You should first reason step-by-step about the current situation. This reasoning process MUST be enclosed within <think> </think> tags. Once you have finished your reasoning, you should choose an admissible action for the current step and present it within <action> </action> tags. Figure 8: ALFWorld rollout prompts with and without retrieved experience. Figure 9 shows the prompt used to update the ALFWorld experience bank. Given a failed trajectory and a successful reference trajectory for the same task, the model generates compact state-aware experiences in JSON format. Prompt: ALFWorld Experience Update System Prompt: You are an expert updating an ALFWorld state-aware experience bank. You are given one failed trajectory and one successful trajectory for the same task. Analyze the contrasted rollout snippets below and propose NEW or revised experiences that help the agent act better in the same environment state. # Task Information Task: task_text Task Type: task_type The task was successfully/unsuccessfully completed. # Contrasted Rollout Snippets Failed trajectory: failed_text Successful trajectory (reference): success_text # Requirements • Each experience must be state-aware, not a generic tip. • Focus on what the successful no-experience rollout did, and what the retrieved experience may have caused the agent to do incorrectly. • Prefer compact trigger language that can be matched at retrieval time. • Use actionable wording tied to admissible actions and visible state cues. • Return only JSON. Generate 1–max_new_skills_per_update new experiences. # Output Format Example ⬇ "title": "Plan object location before acting", "principle": "For this task type, first identify where the object is, then plan the sequence of actions.", "when_to_apply": "When the task involves finding or moving a specific object" Figure 9: Prompt for updating the ALFWorld state-aware experience bank from contrasted failed and successful trajectories. The model is instructed to generate compact, state-aware experiences that can guide future retrieval and decision making. E.2 WebShop Figure 10 shows the WebShop rollout prompts. The experience-conditioned prompt includes retrieved memories, while the experience-free prompt relies only on the shopping instruction, interaction history, current observation, and admissible actions. Prompt: WebShop Agent Execution with Experience System Prompt: You are an expert autonomous agent operating in the WebShop e-commerce environment. Your task is to: task_description. # Retrieved Relevant Experience retrieved_memories Warning: These experiences may be outdated. Use them only if they align with your current observation. # Current Progress Prior to this step, you have already taken step_count step(s). Below are the most recent history_length observations and the corresponding actions you took: action_history You are now at step current_step and your current observation is: current_observation. Your admissible actions of the current situation are: available_actions Now it is your turn to take one action for the current step. You should first reason step-by-step about the current situation, then think carefully which admissible action best advances the shopping goal. This reasoning process MUST be enclosed within <think> </think> tags. Once you have finished your reasoning, you should choose an admissible action for current step and present it within <action> </action> tags. Prompt: WebShop Agent Execution without Experience System Prompt: You are an expert autonomous agent operating in the WebShop e-commerce environment. Your task is to: task_description. # Retrieved Relevant Experience No external general and task-specific experiences are provided. Use your learned strategy and current observation. # Current Progress Prior to this step, you have already taken step_count step(s). Below are the most recent history_length observations and the corresponding actions you took: action_history You are now at step current_step and your current observation is: current_observation. Your admissible actions of the current situation are: available_actions Now it is your turn to take one action for the current step. You should first reason step-by-step about the current situation, then think carefully which admissible action best advances the shopping goal. This reasoning process MUST be enclosed within <think> </think> tags. Once you have finished your reasoning, you should choose an admissible action for current step and present it within <action> </action> tags. Figure 10: WebShop rollout prompts with and without retrieved experience. Figure 11 shows the prompt used to update the WebShop experience bank. The model compares failed and successful shopping trajectories and produces new state-aware experiences tied to visible state cues and available actions. Prompt: WebShop Experience Update System Prompt: You are an expert updating a WebShop shopping state-aware experience bank. You are given one failed trajectory and one successful trajectory for the same task. Analyze the contrasted rollout snippets below and propose NEW or revised experiences that help the agent act better in the same environment state. # Task Information Task: task_text Task Type: task_type The task was successfully/unsuccessfully completed. # Contrasted Rollout Snippets Failed trajectory: failed_text Successful trajectory (reference): success_text # Requirements • Each experience must be state-aware, not a generic tip. • Focus on what the successful no-experience rollout did, and what the retrieved experience may have caused the agent to do incorrectly. • Prefer compact trigger language that can be matched at retrieval time. • Use actionable wording tied to available actions and visible state cues. • Return only JSON. Generate 1–max_new_skills_per_update new experiences. # Output Format Example ⬇ "title": "Verify Early, Abort Fast", "principle": "On the product page, immediately check category, core attributes, and price; if a key constraint is violated, leave the page at once.", "when_to_apply": "Within the first observation on every product detail page." Figure 11: Prompt for updating the WebShop shopping state-aware experience bank from contrasted failed and successful trajectories. The model is instructed to generate compact, state-aware experiences that can guide future retrieval and shopping decisions. Appendix F Dataset License Our experiments are based on the publicly available ALFWorld (26) and WebShop (37) environments. Training, evaluation, and experience-bank construction are performed using trajectories generated from the task instances and interaction interfaces provided by these environments. We strictly follow the licenses and usage terms of ALFWorld and WebShop, and use the resulting data only for academic research purposes. No private, personally identifiable, or proprietary user data is used in any stage of training, evaluation, or experience construction. Appendix G LLMs Usage Statement We employed a Large Language Model (LLM) to assist exclusively in the editorial stage of manuscript preparation. Its role was limited to refining phrasing, correcting grammar, and enhancing clarity and readability across different sections. The LLM had no involvement in formulating research ideas, designing experiments, or conducting analyses. All scientific contributions and findings are entirely the work of the authors. The authors have ensured that the use of the LLM complies with ethical standards, avoiding plagiarism and scientific misconduct.