Paper deep dive
EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning
Can Xie, Yuyi Zhou, Wen Yang, Ziyi zhang, Siyao Song, Yingzhuo Deng, Shuo Ren, Jiajun Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/27/2026, 3:53:45 AM
Summary
The paper introduces EDGE (Experience-Distillation for Guided Exploration), a framework for agentic reinforcement learning that treats external experiences as temporary training-time scaffolds rather than persistent inference-time memory. EDGE partitions rollouts to estimate marginal gains from retrieved experiences, distills beneficial behaviors into the policy via reverse-KL divergence only when gains are positive, and manages an experience bank that co-evolves with the policy by synthesizing new experiences from failures and pruning obsolete ones. This approach improves performance on embodied, web, and search-based QA tasks while eliminating the need for external retrieval at deployment.
Entities (9)
Relation Signals (7)
EDGE → improvesperformanceon → AlfWorld
confidence 95% · On ALFWorld (Shridhar et al., 2021) and WebShop (Yao et al., 2022), EDGE consistently outperforms skill-free RL
EDGE → improvesperformanceon → WebShop
confidence 95% · On ALFWorld (Shridhar et al., 2021) and WebShop (Yao et al., 2022), EDGE consistently outperforms skill-free RL
EDGE → manages → Experience Bank
confidence 95% · A co-evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries
EDGE → uses → Reverse-KL
confidence 95% · distills the induced behavior into the base policy via a reverse-KL objective on its own empirical support
EDGE → uses → GRPO
confidence 95% · EDGE partitions each rollout group... Concretely, EDGE partitions each rollout group into experience-conditioned and experience-free trajectories... We build on Group Relative Policy Optimization (GRPO)
EDGE → outperforms → SkillRL
confidence 90% · EDGE retains 96.0% of its scaffolded performance—compared with 82.9% for SkillRL
EDGE → outperforms → EMPO2
confidence 90% · compared with 82.9% for SkillRL and 92.3% for EMPO2
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy update. Existing experience-augmented approaches retrieve historical guidance at inference time, but they apply experiences without accounting for the policy's evolving capability and create persistent dependencies on external retrieval. We propose EDGE (Experience-Distillation for Guided Exploration), a framework that treats retrieved experiences as temporary training-time scaffolds and progressively internalizes their benefits into the parametric policy. Concretely, EDGE partitions each rollout group into experience-conditioned and experience-free trajectories to estimate and admit only positive marginal gains without extra sampling, then distills the induced behavior into the base policy via a reverse-KL objective on its own empirical support. A co-evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries as the policy evolves. Across embodied, web, and search-based QA tasks, EDGE improves over strong RL baselines by up to 12.5 points and remains effective without inference-time scaffolds or a proprietary reflector. The code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.21946v2
- Canonical: https://arxiv.org/abs/2608.21946v2
Trouble viewing inline? Open PDF directly →
Full Text
127,523 characters extracted from source content.
Expand or collapse full text
EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning Can Xie Affiliation: School of Artificial Intelligence, University of Chinese Academy of Sciences Affiliation: Institute of Automation, Chinese Academy of Sciences Yuyi Zhou Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences Wen Yang Affiliation: School of Artificial Intelligence, University of Chinese Academy of Sciences Affiliation: Institute of Automation, Chinese Academy of Sciences Ziyi Zhang Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences Siyao Song, Yingzhuo Deng, Shuo Ren, Jiajun Zhang11footnotemark: 1 †thanks: Corresponding authors Affiliation: School of Artificial Intelligence, University of Chinese Academy of Sciences Affiliation: Institute of Automation, Chinese Academy of Sciences Affiliation: Wuhan AI Researchxiecan2024,shuo.ren@ia.ac.cn, jjzhang@nlpr.ia.ac.cn Abstract Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy update. Existing experience-augmented approaches retrieve historical guidance at inference time, but they apply experiences without accounting for the policy’s evolving capability and create persistent dependencies on external retrieval. We propose EDGE (Experience-Distillation for Guided Exploration), a framework that treats retrieved experiences as temporary training-time scaffolds and progressively internalizes their benefits into the parametric policy. Concretely, EDGE partitions each rollout group into experience-conditioned and experience-free trajectories to estimate and admit only positive marginal gains without extra sampling, then distills the induced behavior into the base policy via a reverse-KL objective on its own empirical support. A co-evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries as the policy evolves. Across embodied, web, and search-based QA tasks, EDGE improves over strong RL baselines by up to 12.5 points and remains effective without inference-time scaffolds or a proprietary reflector. The code is available at https://github.com/xvolcano02/EDGE. Figure 1: Overview of the EDGE framework. (a) Experience-guided exploration scaffolding uses contrastive rollouts to estimate the marginal gain Δe _e of a retrieved experience. (b) Gain-gated distillation internalizes beneficial scaffold behavior into the base policy only when Δe>0 _e>0. (c) The experience bank co-evolves with the policy via failure-driven expansion and utility-based pruning, yielding a retrieval-free policy at deployment. 1 Introduction Reinforcement learning (RL) with outcome-based objectives such as GRPO (Shao et al., 2024) has become a standard post-training paradigm for LLM-based agents that reason, plan, and act through multi-turn interactions (Yao et al., 2023; Shinn et al., 2023; Yao et al., 2022; Shridhar et al., 2021; Feng et al., 2025; Wang et al., 2025b; Jin et al., 2025). Yet current agentic RL uses its own experience poorly. A single rollout may contain reusable patterns—effective task decompositions, recovery strategies, dead-end avoidance, but once a trajectory contributes a scalar advantage to one policy update, these patterns are largely discarded. The agent must therefore rediscover them from scratch, a particularly costly failure mode in sparse-reward, long-horizon environments where agentic RL is most needed. A growing body of work addresses this inefficiency by augmenting agents with external experience at inference time, through episodic reflections (Shinn et al., 2023; Zhao et al., 2024), persistent memory (Chhikara et al., 2025; Fang et al., 2026), or skill libraries (Xia et al., 2026; Liu et al., 2026). While effective, these approaches share two structural limitations that become apparent as training progresses. Experience utility is inherently policy-dependent: guidance that accelerates an undertrained policy can become redundant---or actively harmful---once the corresponding behavior has been internalized, yet most methods apply experience unconditionally or filter it by static heuristics. Moreover, if the agent must retrieve experience at deployment, part of its competence resides in the context window rather than in the model parameters, incurring persistent token overhead and sensitivity to retrieval noise. 11 1 These are not hypothetical concerns: in our experiments, naive memory augmentation (e.g., EvolveR, GRPO+Mem0) degrades performance well below vanilla GRPO, confirming that unfiltered experience injection is unreliable. These observations suggest that experience reuse in agentic RL should be treated as a dynamic lifecycle rather than a static retrieval mechanism. External experience could first guide exploration, as a temporary training-time scaffold, then be exploited by consolidating its useful behavioral effect into the experience-free policy, and finally be retired once its utility vanishes. Under this view, experience reuse becomes a scaffold-to-parameter learning problem rather than a retrieve-and-prompt augmentation problem. In this paper, we instantiate this idea in EDGE (Experience-Distillation for Guided Exploration), a framework that turns retrieved experience from persistent inference-time memory into dynamically validated training-time scaffolding. EDGE uses online marginal-gain estimation to decide when an experience should guide exploration, be distilled into the policy, or be retired as the policy evolves, via three core mechanisms. First, experience-guided exploration scaffolding partitions each GRPO rollout group into experience-conditioned and experience-free trajectories, estimating the marginal gain of a retrieved experience without additional environment sampling and retaining only positive-gain instances. Then, gain-gated privileged distillation serves as the exploitation step, transferring the behavior induced by the privileged context into the standard policy via reverse-KL divergence computed on the student’s own empirical support, activating only when the estimated gain is positive to avoid negative transfer. Finally, experience bank evolution and management enables the experience bank to co-evolve with the policy throughout training: new experiences are synthesized from failure-mode analysis, while obsolete ones are pruned based on tracked utility, keeping the scaffold aligned with the agent’s evolving capability. On ALFWorld (Shridhar et al., 2021) and WebShop (Yao et al., 2022), EDGE consistently outperforms skill-free RL and prior experience-augmented methods, with the largest gains on exploration-intensive subtasks (e.g., Heat, Cool, Pick2). When external experiences are withheld at inference time, EDGE retains 96.0% of its scaffolded performance—compared with 82.9% for SkillRL and 92.3% for EMPO2—confirming that useful exploration priors have been absorbed into the parametric policy. On seven search-based QA benchmarks with Qwen3-4B, EDGE further improves over Search-R1 by 5.9 points on average; using the policy itself as the reflector preserves 97.3% of the performance obtained with GPT-4o. Our contributions are as follows: • We identify two failure modes of experience-augmented agentic RL—policy-dependent experience utility and persistent inference-time retrieval dependence—and reframe experience reuse as a dynamic scaffold-to-parameter transition. • We introduce experience-guided exploration scaffolding, which partitions rollout groups to estimate the marginal value of each retrieved experience under the current policy without requiring additional environment sampling. • We propose gain-gated privileged distillation, which internalizes scaffold-induced behavior via reverse-KL on the student’s own empirical support, gated by the estimated marginal gain to prevent negative transfer, together with a co-evolutionary experience bank that expands and prunes in response to the policy’s evolving needs. • Experiments across embodied, web, and search-based QA tasks show that EDGE improves task performance and training efficiency, transfers to a newer Qwen3 backbone, and remains effective with a self-reflector and without experience retrieval at deployment. 2 Preliminaries We formalize the multi-turn decision-making process of an LLM-based agent as a Partially Observable Markov Decision Process (POMDP) (Kaelbling et al., 1998) ⟨,,,,ℛ⟩ ,A,O,T,R . The observation space O consists of natural-language strings emitted by the environment, and the action space A comprises token sequences generated autoregressively by the policy πθ _θ. Given a task instruction x, the agent interacts with the environment over a sequence of turns: at step t it receives observation ot∈o_t and produces action at∈a_t according to at∼πθ(⋅∣x,ht,ot),a_t _θ(· x,h_t,o_t), (1) where ht=(o1,a1,…,ot−1,at−1)h_t=(o_1,a_1,…,o_t-1,a_t-1) is the interaction history. An episode terminates upon task completion or at a maximum step limit, yielding a sparse binary reward R(τ)∈0,1R(τ)∈\0,1\ and a full trajectory τ=(x,o1,a1,…,oT,aT,R(τ)).τ=(x,\;o_1,a_1,\;…,\;o_T,a_T,\;R(τ)). (2) We build on Group Relative Policy Optimization (GRPO) (Shao et al., 2024), which samples a group of G parallel trajectories τii=1G\ _i\_i=1^G per task and computes group-normalized advantages: A^i=R(τi)−mean(R(τj)j=1G)std(R(τj)j=1G). A_i= R( _i)-mean(\R( _j)\_j=1^G)std(\R( _j)\_j=1^G). (3) The policy is updated by maximizing a clipped surrogate objective with KL regularization: ℒRL(θ)=−[1G∑i=1G1|τi|∑t=1|τi|(min(ρi,tA^i, _RL(θ)=-E\! [ 1G _i=1^G 1| _i| _t=1^| _i| ( \! ( _i,t\, A_i,\; clip(ρi,t,−ϵ,+ϵ)A^i)−βDKL(πθ∥πref))], ( _i,t,1\!-\!ε,1\!+\!ε)\, A_i )-β\,D_KL ( _θ\| _ref ) ) ], (4) where ρi,t=πθ(at∣x,ht,ot)/πθold(at∣x,ht,ot) _i,t= _θ(a_t x,h_t,o_t)\,/\, _ _old(a_t x,h_t,o_t) is the importance sampling ratio, πref _ref is the reference policy, and β controls KL penalty strength. By contrasting outcomes within each group, GRPO steers θ toward successful action sequences without a learned value function. However, in sparse-reward, partially observable environments, unguided exploration often yields groups in which few or no trajectories succeed, rendering the advantage estimate uninformative—a limitation we address in the following section. 3 Method: EDGE We present EDGE (Experience-Distillation for Guided Exploration), a framework that iteratively strengthens the multi-turn reasoning capability of LLM-based agents (Figure 1). Central to our approach is treating retrieved external experience not as a static inference-time prompt—which inflates the context window and creates persistent retrieval dependence—but as a training-time scaffold that guides exploration and is discarded at deployment. The agent first explores with privileged access to experience, then distills only the empirically beneficial behavior into its own parameters, so the deployed policy requires no external scaffold. Section 3.1 introduces the experience-guided exploration scaffolding, Section 3.2 details the gain-gated privileged distillation mechanism, and Section 3.3 describes the experience bank evolution and management strategy. 3.1 Experience-Guided Exploration Scaffolding To provide exploratory guidance in the sparse-reward, partially observable environments typical of agentic tasks, EDGE introduces a controlled information asymmetry within the standard GRPO rollout group. Given a task instruction x and an experience bank ℰE, we retrieve the top-m most relevant experiences by embedding similarity, using the task instruction and initial environmental observations as the query, and select the highest-scoring entry e∈ℰe . Recall that GRPO samples a group of G parallel trajectories per task to estimate relative advantages (Eq. (3)). We partition this group into two equal subsets without adding extra rollouts: • Teacher rollouts (T T, G/2G/2 trajectories): conditioned on the privileged context c=x⊕ec T=x e, where ⊕ denotes concatenation under a unified chat template. • Student rollouts (T S, G/2G/2 trajectories): conditioned on the standard context c=xc S=x alone. Because both subsets share the same policy πθ _θ and differ only in whether the retrieved experience is visible, any performance gap can be attributed to the informational advantage provided by e. We quantify this gap via the instantaneous marginal gain: Δe=1||∑τ∈R(τ)−1||∑τ∈R(τ), _e\;=\; 1|T T| _τ TR(τ)\;-\; 1|T S| _τ SR(τ), (5) where R(τ)R(τ) is the binary outcome reward defined in §2. A positive Δe _e signals that the experience provides useful guidance beyond the agent’s current capability; a non-positive value indicates it is redundant or harmful. This estimate gates the distillation objective (§3.2) and drives experience bank updates (§3.3). Beyond gating distillation, Δe _e controls which rollouts enter the RL loss itself. When Δe>0 _e>0, advantages (Eq. (3)) are computed over all G trajectories. The pooled baseline therefore reflects the performance level achievable under the scaffold, giving student rollouts a stronger calibrated reference than a within-subset baseline would provide. When Δe≤0 _e≤ 0, teacher-conditioned trajectories are excluded from the RL loss entirely: advantages are computed over T S alone, reducing the update to a standard GRPO step. Without this masking, non-beneficial teacher rollouts would still shift the group baseline and distort student advantages—contaminating the policy gradient even though distillation is gated off. 3.2 Gain-Gated Privileged Distillation The scaffolding in §3.1 exposes the agent to successful reasoning patterns it may not discover on its own, but because the retrieved experience is unavailable at deployment, these gains remain ephemeral unless consolidated into the policy itself. This motivates a complementary mechanism: distilling the scaffold-induced improvements into the base policy parameters so that the agent can reproduce them from the standard context alone. Our approach realizes this through asymmetric self-distillation. Rather than relying on a separate, unconditionally superior teacher, we use the same policy πθ _θ under two informational conditions: the teacher is πθ _θ evaluated with the privileged context c T (gradients stopped), while the student operates under c S. The asymmetry therefore lies in the information available to each role, not in model capacity. The gain gate (Δe>0)I( _e>0) further ensures that distillation is triggered only when the scaffold yields a verified performance advantage, allowing the policy to selectively internalize reusable reasoning patterns while avoiding negative transfer. To guard against covariate shift, we construct the distillation target on the student’s own empirical support. For a gain-gated task instance (Δe>0 _e>0) and a student trajectory τ∈τ S with token sequence (y1,…,ym)(y_1,…,y_m) generated under c S, we perform a no-gradient forward pass of πθ _θ over the same tokens conditioned on c T. Because both passes share identical actions, the comparison isolates the informational advantage of e without introducing out-of-distribution transitions. We adopt the reverse KL divergence DKL(πstudent∥πteacher)D_KL( _student\| _teacher) as the distillation objective. Unlike the forward KL, which compels the student to cover the full teacher support and can induce mode-covering artifacts, the reverse KL is mode-seeking: it encourages the policy to concentrate on the most effective reasoning mode under the scaffold. The token-level scaffold internalization loss is: ℒdistill(θ) _distill(θ) =τ∼[(Δe>0) =E_τ S\! [I( _e\!>\!0) ∑t=1m∑y∈πθ(y)δt(y)], _t=1^m _y _θ S(y)\, _t(y) ], δt(y) _t(y) =logπθ(y∣c<t)πsg(θ)(y∣c<t), = _θ(y c S_<t) _sg(θ)(y c T_<t), (6) where πθ(y)≜πθ(y∣c<t) _θ S(y) _θ(y c S_<t) for brevity, δt(y) _t(y) is the per-token log-ratio between the student and the privileged teacher, and sg(⋅)sg(·) denotes stop-gradient. This loss is combined with the GRPO objective (Eq. (4)) to form the joint actor loss: ℒactor(θ)=ℒRL(θ)+λℒdistill(θ),L_actor(θ)=L_RL(θ)+λ\,L_distill(θ), (7) where λ controls the relative weight of scaffold internalization versus the RL signal. Through this joint optimization, the deployed agent reproduces privileged reasoning without any external scaffold at test time. 3.3 Experience Bank Evolution and Management A static experience library cannot adequately support the agent throughout training: as πθ _θ improves, it encounters situations where existing experiences provide insufficient guidance; conversely, previously beneficial experiences may become redundant once the policy has internalized the corresponding behavior, or harmful if they conflict with newly discovered strategies. EDGE therefore treats ℰE as a living repository that co-evolves with the policy through both expansion and pruning. Following Xia et al. (2026), we generate new experiences from the agent’s own rollout trajectories. After each training step, we identify task categories whose success rate falls below a threshold ξ and collect representative failed and successful trajectories from these categories. A reflector LLM then analyzes the contrast between them to synthesize new experiential guidance: enew=freflect(τ+,τ−),e_new=f_reflect\! (τ^+,\,τ^- ), (8) where τ+τ^+ and τ−τ^- denote a successful and a failed trajectory from the same task category, respectively. The generated experiences are deduplicated against existing entries and inserted into ℰE for retrieval in subsequent training iterations. Meanwhile, the utility of existing experiences is continuously tracked via an Exponential Moving Average (EMA) score for each experience e: Ue(t)=(1−μ)Ue(t−1)+μΔe,U_e^(t)=(1-μ)\,U_e^(t-1)+μ\, _e, (9) where μ∈(0,1)μ∈(0,1) is the momentum coefficient and Δe _e is the instantaneous marginal gain from Eq. (5). The EMA smooths over stochastic fluctuations while remaining responsive to genuine shifts in experience utility. Experiences whose tracked utility Ue(t)U_e^(t) falls below a threshold η are removed from ℰE, retiring scaffolds that have been absorbed into the policy and reducing overhead during retrieval and rollout. This yields a co-evolutionary dynamic: the policy improves by internalizing useful experiences, exposing new failure modes that drive bank expansion, while obsolete experiences are simultaneously pruned away. Type Method ALFWorld WebShop Pick Look Clean Heat Cool Pick2 All Score Succ. Closed-Source Model Prompting GPT-4o 75.3 60.8 31.2 56.7 21.6 49.8 48.0 31.8 23.7 Prompting Gemini-2.5-Pro 92.8 63.3 62.1 69.0 26.6 58.7 60.3 42.5 35.9 Qwen2.5-1.5B-Instruct Prompting Base Model 5.9 5.5 3.3 9.7 4.2 0.0 4.1 23.1 5.2 Prompting ReAct 17.4 20.5 15.7 6.2 7.7 2.0 12.8 40.1 11.3 Prompting Reflexion 35.3 22.2 21.7 13.6 19.4 3.7 21.8 55.8 21.9 Post-Training GRPO 85.3 53.7 84.5 78.2 59.7 53.5 72.8 75.8 56.8 Post-Training SkillRL 91.2±4.3 64.3±4.6 78.1±5.4 76.9±6.3 70.8±6.5 54.6±6.1 74.2±4.7 77.3±3.5 60.9±3.7 Post-Training EMPO2 86.9±3.3 66.2±5.2 79.3±4.9 79.1±4.5 75.3±5.7 64.8±6.6 76.8±3.7 78.2±3.7 63.2±3.3 Post-Training EDGE 80.6±2.2 73.7±5.5 76.0±4.3 87.3±4.0 85.6±6.1 68.4±5.6 79.7±2.3 78.8±2.1 65.6±3.2 Qwen2.5-7B-Instruct Prompting Base Model 33.4 21.6 19.3 6.9 2.8 3.2 14.8 26.4 7.8 Prompting ReAct 48.5 35.4 34.3 13.2 18.2 17.6 31.2 46.2 19.5 Prompting Reflexion 62.0 41.6 44.9 30.9 36.3 23.8 42.7 58.1 28.8 Post-Training EvolveR† 64.9 33.3 46.4 13.3 33.3 33.3 43.8 42.5 17.6 Post-Training GRPO+Mem0† 78.1 54.8 56.1 31.0 65.0 26.9 54.7 58.1 37.5 Post-Training OPSD† 50.0 60.0 22.7 21.4 17.6 9.5 32.8 4.5 2.3 Post-Training GRPO 92.8 85.7 89.3 75.7 74.5 67.7 82.1 80.3 70.1 Post-Training GRPO+OPSD† 91.4 61.5 100 87.5 76.5 52.2 80.4 86.8 76.5 Post-Training SkillRL 94.1±2.4 83.3±4.3 88.4±3.6 85.6±5.1 90.2±5.6 78.6±4.9 88.2±3.7 86.2±3.1 76.7±3.3 Post-Training EMPO2 93.3±3.7 88.9±3.9 91.2±4.4 88.5±5.3 89.4±4.6 79.8±5.1 89.1±2.8 88.3±2.6 77.1±4.0 Post-Training EDGE 96.0±1.9 85.1±4.8 93.4±3.2 90.0±2.6 92.9±4.5 81.6±6.0 90.4±2.2 89.6±2.8 82.6±3.8 Table 1: Performance on ALFWorld and WebShop. We report the average success rate (%) per subtask and overall for ALFWorld, and both the average score and success rate (%) for WebShop. Results of EDGE and other experience-augmented training methods are averaged over 3 random seeds (mean± ). † denotes results replicated from Xia et al. (2026) and Lu et al. (2026a). 4 Experiments We evaluate EDGE to examine whether experience scaffolding can improve agentic RL while being progressively internalized by the policy. Our experiments address five questions: (1) Does EDGE improve over prompting, vanilla RL, and prior experience-augmented methods, particularly on exploration-intensive tasks? (2) Does the trained policy retain its performance when external experiences are removed at inference time? (3) How much does each component—gain gating, distillation, and pruning—contribute, and how do they interact? (4) Is experience utility truly non-stationary, and does the co-evolutionary bank adapt accordingly during training? (5) Does EDGE generalize to a newer backbone and search-based QA, and remain effective without a stronger proprietary reflector? 4.1 Experiment Setup Environments. We evaluate EDGE across two complementary agentic settings: interactive decision making and search-based QA. For the former, we use two environments with sparse outcome-level feedback. ALFWorld (Shridhar et al., 2021) comprises 3,827 text-based household tasks across six types (Pick, Look, Clean, Heat, Cool, and Pick2), while WebShop (Yao et al., 2022) requires agents to search and purchase products matching user instructions from a catalog of over 1.1M items. To test generalization beyond interactive environments, we further evaluate search-based QA on NQ, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, MuSiQue, and Bamboogle, covering single- and multi-hop evidence retrieval and reasoning. Baselines. We compare against three families of methods: closed-source LLM agents (GPT-4o (OpenAI et al., 2024) and Gemini-2.5-Pro (Comanici et al., 2025)), prompting agents (ReAct (Yao et al., 2023) and Reflexion (Shinn et al., 2023)), and training-based agents. The post-training baselines include GRPO (Shao et al., 2024), OPSD (Zhao et al., 2026), EvolveR (Wu et al., 2026), GRPO augmented with Mem0 (Chhikara et al., 2025), SkillRL (Xia et al., 2026), and EMPO2 (Liu et al., 2026). Detailed descriptions are provided in Appendix B.1. Training details. We use Qwen2.5-1.5B/7B-Instruct (Qwen et al., 2025) as the base models for ALFWorld and WebShop, and Qwen3-4B-Instruct-2507 (Yang et al., 2025) for search-based QA. For ALFWorld and WebShop, all post-training methods use exactly the same hyperparameter configurations. The rollout group size G for group-based RL methods is set to 8. For experience retrieval, we encode experiences with Qwen3-Embedding-0.6B (Zhang et al., 2025b) and rank candidates by cosine similarity against the task instruction and initial observations. We use GPT-4o (OpenAI et al., 2024) as the default reflector. To test whether EDGE depends on a stronger proprietary model, the QA experiments also replace GPT-4o with the Qwen3 policy itself while leaving the remaining framework unchanged. Full training settings and hyperparameter details are provided in Appendix B.2, with a wall-clock breakdown in Appendix B.3. 4.2 Main Results Table 1 presents results on ALFWorld and WebShop across two model scales. At 7B, EDGE achieves 90.4% success rate on ALFWorld and 82.6% on WebShop, improving over GRPO by 8.3 and 12.5 points and over the strongest experience-augmented baseline EMPO2 by 1.3 and 5.5 points, with consistent gains at 1.5B. The subtask breakdown reveals that EDGE’s advantage concentrates on exploration-heavy tasks—Heat, Cool, and Pick2—which demand longer action sequences with fewer intermediate rewards. On WebShop, the success-rate improvement over GRPO (+12.5) exceeds the score improvement (+9.3), indicating that EDGE helps agents complete full decision chains rather than merely accumulate partial credit. These results highlight that experience is most effective as selective exploration guidance rather than unconditional augmentation. Naive memory approaches (EvolveR, GRPO+Mem0) degrade well below vanilla GRPO, and standalone self-distillation (OPSD) provides insufficient signal without effective exploration. Combining the two (GRPO+OPSD) recovers competitive aggregate performance but remains brittle across subtasks, whereas EDGE’s marginal-gain gating ensures only beneficial experiences contribute, yielding both higher overall success and more uniform subtask coverage. Method NQ TriviaQA PopQA HotpotQA 2Wiki MuSiQue Bamboogle Avg. Qwen3-4B-Instruct-2507 17.8 42.3 27.2 23.6 21.5 5.6 7.4 20.8 Search-R1 40.6 61.2 43.7 35.2 36.3 12.4 41.6 38.7 SkillRL (GPT-4o) 42.4 59.6 44.3 42.1 43.4 15.9 43.2 41.5 EDGE (GPT-4o) 48.0 63.4 46.9 42.4 48.1 17.8 45.6 44.6 EDGE (self) 44.8 64.3 45.4 43.6 45.5 16.2 44.2 43.4 Table 2: Performance (%) on seven search-based QA benchmarks with Qwen3-4B-Instruct-2507. “GPT-4o” and “self” denote the reflector used to synthesize experiences. Best results are bolded. Generalization to Search-Based QA and Self-Reflection. Table 2 tests whether EDGE transfers to a newer backbone and search-based QA without relying on external reflection. With GPT-4o, EDGE leads on five of seven benchmarks. More importantly, self-reflection retains 97.3% of this performance while surpassing both baselines on every benchmark. Its particularly strong results on TriviaQA and HotpotQA further suggest that self-generated experience remains effective across both fact-oriented and multi-hop search settings. Together, these findings indicate that EDGE’s gains arise primarily from its experience construction mechanism rather than privileged access to a stronger reflector; GPT-4o improves the quality of the resulting experience, but is not essential for successful transfer. Figure 2: Performance retention after scaffold removal at inference time. EDGE preserves 96.0% of its scaffolded performance without external experiences, compared with 82.9% for SkillRL and 92.3% for EMPO2, indicating more effective internalization into the parametric policy at the 7B scale. Figure 3: Training dynamics of EDGE vs. GRPO on ALFWorld with Qwen2.5-7B-Instruct. EDGE achieves higher validation success while reducing both environment steps and trajectory length more rapidly, indicating that scaffolded exploration is progressively internalized into more efficient experience-free behavior. Inference without External Scaffolds. Figure 2 poses a stricter test: how much performance survives when all external experiences are withheld at inference time? EDGE preserves 96.0% of its scaffolded performance, compared with 92.3% for EMPO2 and 82.9% for SkillRL, confirming that the reverse-KL distillation stage successfully transfers scaffold-induced behavior into the parametric policy. This near-complete retention supports the intended scaffold-to-parameter transition: experience-guided behaviors are largely internalized by the policy and remain available without inference-time retrieval. 4.3 Analysis Ablation Studies. Method Pick Look Clean Heat Cool Pick2 All GRPO 92.8 85.7 89.3 75.7 74.5 67.7 82.1 EDGE 96.0 85.1 93.4 90.0 92.9 81.6 90.4 w/o ExpPruning 90.5 82.2 90.1 84.3 87.6 76.8 86.7 w/o Distillation 91.2 81.5 91.8 78.7 77.2 71.1 83.6 w/o Gain-Gating 83.1 72.5 79.7 65.3 61.8 61.2 72.3 Table 3: Ablation results on ALFWorld with Qwen2.5-7B-Instruct. We report success rate (%) per subtask and overall. Table 3 isolates each component on ALFWorld with Qwen2.5-7B-Instruct. The most striking finding is that removing the gain gate drops overall success to 72.3%—9.8 points below vanilla GRPO—revealing that unfiltered experience injection does not merely fail to help but actively harms the policy, particularly on exploration-heavy subtasks where misleading guidance compounds over long horizons. This result also addresses a natural concern about the G/2+G/2G/2+G/2 rollout partition: since all other ablated variants retain the same split yet outperform GRPO, the partition itself does not dilute the RL signal; the degradation is attributable entirely to distilling low-quality experiences. The remaining two components contribute complementary benefits. Without distillation, performance falls to 83.6%, only 1.5 points above GRPO, with losses concentrated on the same exploration-heavy subtasks—confirming that scaffolded exploration provides transient guidance but does not durably reshape the policy without explicit behavioral transfer. Without experience pruning, performance declines more modestly to 86.7% with losses spread evenly across subtasks, indicating that pruning acts as a maintenance mechanism that keeps the experience bank aligned with the policy’s evolving capability rather than targeting any specific failure mode. Training Dynamics. Figure 3 reveals a clear divergence between EDGE and GRPO after approximately step 100: GRPO plateaus and begins to regress, whereas EDGE continues to improve steadily toward 90% validation success. We attribute GRPO’s decline to an exploration–exploitation collapse: once the policy commits to locally successful strategies, it loses the diversity needed to solve the remaining hard tasks, and further optimization erodes earlier gains. EDGE’s experience scaffolds counteract this by continually injecting structured exploration guidance, while gain gating ensures that this guidance remains beneficial as the policy strengthens. The efficiency panels corroborate this interpretation—EDGE reduces both environment steps and trajectory length more rapidly than GRPO, indicating that the policy internalizes increasingly direct action sequences rather than relying on extended trial-and-error, consistent with the scaffold-removal results in Figure 2. Figure 4: Experience gain dynamics during training. Experience Gain Tracking. Figure 4 tracks the mean marginal gain Δe _e (Eq. (5)) across training to test a key premise of §3.3: that experience utility is non-stationary. With a static bank, the initially positive gain decays and turns negative after roughly step 100—confirming that once-useful experiences become actively harmful as the policy outgrows them, and explaining why removing gain gating in Table 3 degrades performance below vanilla GRPO. Experience evolution delays this decay, but unchecked bank growth (reaching over 650 entries) introduces retrieval noise that keeps the gain signal volatile. The full co-evolutionary configuration sustains the most stable positive gain: pruning not only curbs bank growth but causes the bank to shrink after step 100, indicating that the policy absorbs existing experiences faster than new failure modes generate replacements—a direct signature of successful internalization. Further details on bank evolution dynamics appear in Appendix C.2. 5 Related Work Reinforcement Learning for LLM Agents. RL has become a standard post-training paradigm for LLM agents in multi-turn environments (Yao et al., 2022; Shridhar et al., 2021; Feng et al., 2025; Wang et al., 2025b; Jin et al., 2025; Li et al., 2026), progressing from PPO-based methods (Schulman et al., 2017) to critic-free objectives such as GRPO (Shao et al., 2024) and RLOO (Ahmadian et al., 2024), with further improvements in multi-turn credit assignment through turn-level or stepwise signals (Feng et al., 2025; Wei et al., 2025; Wang et al., 2025a; Xie et al., 2026; Zhang et al., 2026b). However, the reusable exploration patterns within trajectories are still consumed once and discarded. EDGE retains the group-sampling backbone of GRPO but repurposes it to estimate which retrieved experiences currently improve exploration and to transfer their effects into the policy. Experience-Augmented LLM Agents. External memory and experience reuse have been explored through prompting-based reflections (Shinn et al., 2023; Zhao et al., 2024; Yang et al., 2024; Fang et al., 2026; Ouyang et al., 2026) and persistent retrieval during interaction (Chhikara et al., 2025; Xia et al., 2026; Wu et al., 2026; Ma et al., 2026; Wang et al., 2026; Zhang et al., 2026a). Recent methods selectively replay past reasoning traces (Yan et al., 2025; Zhang et al., 2025a; Qin et al., 2025; Zhan et al., 2026), but primarily target single-turn tasks without the multi-turn, partially observable interaction loops of agentic settings. EMPO2 (Liu et al., 2026) pursues memory-free behavior through heuristic experience selection and off-policy distillation. Concurrent SKILL0 (Lu et al., 2026b) progressively withdraws a preconstructed skill inventory under a decaying curriculum. In contrast, EDGE uses paired-rollout gains to validate and distill online experience while co-evolving its bank through utility-based pruning. Privileged Information and Self-Distillation. Leveraging training-time information unavailable at deployment spans LUPI (Vapnik and Vashist, 2009), asymmetric actor-critic methods (Pinto et al., 2017), and context distillation in LLMs (Snell et al., 2022; Choudhury and Sodhi, 2025; Liu et al., 2026; Cheng et al., 2026), though these approaches are largely off-policy and suffer from distribution mismatch with the student’s own visitation. On-policy distillation (Agarwal et al., 2024; Ye et al., 2026) addresses this by supervising the student on its own sequences, and on-policy self-distillation (OPSD) (Zhao et al., 2026; Hübotter et al., 2026; Lu et al., 2026a) further removes the need for a separate teacher. EDGE shares this structure but adds gain-based distillation gating to activate distillation only under verified positive gain and co-evolves the experience bank through utility-driven expansion and pruning. 6 Conclusion We presented EDGE, a framework for using retrieved experience as a temporary training-time scaffold rather than a persistent inference-time dependency. EDGE estimates the marginal utility of retrieved experiences under the current policy, admits only positive-gain scaffolds, and distills their behavioral effect into the standard experience-free policy. Across ALFWorld, WebShop, and seven search-based QA benchmarks, EDGE improves over RL and experience-augmented baselines, with especially large gains on exploration-intensive subtasks and strong retention after scaffold removal. The Qwen3 results further show that these gains transfer to a newer backbone and largely persist when the policy itself replaces GPT-4o as the reflector. These results suggest a practical principle for agentic RL: external experience is most useful when it is dynamically validated, selectively applied, and ultimately internalized. Limitations While EDGE requires no extra environment rollouts, it introduces training-time overhead for maintaining the experience bank, computing teacher–student comparisons, and invoking the reflector LLM. The effectiveness of gain-gating also depends on reward quality; as with all outcome-based RL methods, noisy or misspecified rewards would reduce the reliability of the marginal-gain estimates. In terms of empirical scope, our evaluation covers two interactive environments and seven search-based QA benchmarks with Qwen2.5 models at the 1.5B and 7B scales and Qwen3 at the 4B scale; broader validation across model families, larger scales, and multimodal or real-world environments remains future work. More broadly, EDGE assumes experiences are discrete textual artifacts; extending the scaffold-to-parameter principle to latent memory representations is an open problem. Acknowledgements This work is supported by the Strategic Priority Research Program of Chinese Academy of Sciences under Grant XDA04080400 and Beijing Natural Science Foundation L259016. References Agarwal et al. (2024) R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. R. Garea, M. Geist, and O. Bachem On-policy distillation of language models: learning from self-generated mistakes. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, External Links: Link Cited by: §5. Ahmadian et al. (2024) A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, O. Pietquin, A. Üstün, and S. Hooker Back to basics: revisiting reinforce-style optimization for learning from human feedback in llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), p. 12248–12267. External Links: Link, Document Cited by: §5. Cheng et al. (2026) Z. Cheng, Z. Liu, Y. Shan, X. Wang, X. Zhu, Y. Ma, H. Wang, Y. Guo, W. Lin, and Y. Wang Mem2^2evolve: towards self-evolving agents via co-evolutionary capability expansion and experience distillation. External Links: 2604.10923, Link Cited by: §5. Chhikara et al. (2025) P. Chhikara, D. Khant, S. Aryan, T. Singh, and D. Yadav Mem0: building production-ready ai agents with scalable long-term memory. External Links: 2504.19413, Link Cited by: item •, §1, §4.1, §5. Choudhury and Sodhi (2025) S. Choudhury and P. Sodhi Better than your teacher: LLM agents that learn from privileged AI feedback. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, External Links: Link Cited by: §5. Comanici et al. (2025) G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, L. Marris, S. Petulla, C. Gaffney, A. Aharoni, N. Lintz, T. C. Pais, H. Jacobsson, I. Szpektor, N. Jiang, K. Haridasan, A. Omran, N. Saunshi, D. Bahri, G. Mishra, E. Chu, T. Boyd, B. Hekman, A. Parisi, C. Zhang, K. Kawintiranon, T. Bedrax-Weiss, O. Wang, Y. Xu, O. Purkiss, U. Mendlovic, I. Deutel, N. Nguyen, A. Langley, F. Korn, L. Rossazza, A. Ramé, S. Waghmare, H. Miller, N. Byrd, A. Sheshan, R. Hadsell, S. Bhardwaj, P. Janus, T. Rissa, D. Horgan, A. Abdagic, L. Belenki, J. Allingham, A. Singh, T. Guidroz, S. Srinivasan, H. Schmit, K. Chiafullo, A. Elisseeff, N. Jha, P. Kolhar, L. Berrada, F. Ding, X. Si, S. B. Mallick, F. Och, S. Erell, E. Ni, T. Latkar, S. Yang, P. Sirkovic, Z. Feng, R. Leland, R. Hornung, G. Wu, C. Blundell, H. Alvari, P. Huang, C. Yip, S. Deur, L. Liu, G. Surita, P. Duque, D. Damen, J. Jia, A. Guez, M. Mircea, A. Sinha, A. Magni, P. Stradomski, T. Marian, V. Galić, W. Chen, H. Husain, A. Singhal, D. Grewe, F. Aubet, S. Song, L. Blanco, L. Rechis, L. Ho, R. Munoz, K. Zheng, J. Hamrick, K. Mather, H. Taitelbaum, E. Rutherford, Y. Lei, K. Chen, A. Shukla, E. Moreira, E. Doi, B. Isik, N. Shabat, D. Rogozińska, K. Kolipaka, J. Chang, E. Vušak, S. Venkatachary, S. Noghabi, T. Bharti, Y. Jun, A. Zaks, S. Green, J. Challagundla, W. Wong, M. Mohammad, D. Hirsch, Y. Cheng, I. Naim, L. Proleev, D. Vincent, A. Singh, M. Krikun, D. Krishnan, Z. Ghahramani, A. Atias, R. Aggarwal, C. Kirov, D. Vytiniotis, C. Koh, A. Chronopoulou, P. Dogra, V. Ion, G. Tyen, J. Lee, F. Weissenberger, T. Strohman, A. Balakrishna, J. Rae, M. Velic, R. de Liedekerke, O. Elyada, W. Yuan, C. Liu, L. Shani, S. Kishchenko, B. Alessio, Y. Li, R. Song, S. Kwei, O. Jankowski, A. Pappu, Y. Namiki, Y. Ma, N. Tripuraneni, C. Cherry, M. Ikonomidis, Y. Ling, C. Ji, B. Westberg, A. Wright, D. Yu, D. Parkinson, S. Ramaswamy, J. Connor, S. H. Yeganeh, S. Grover, G. Kenwright, L. Litchev, C. Apps, A. Tomala, F. Halim, A. Castro-Ros, Z. Li, A. Boral, P. Sho, M. Yarom, E. Malmi, D. Klinghoffer, R. Lin, A. Ansell, P. K. S, S. Zhao, S. Zuo, A. Santoro, H. Cheng, S. Demmessie, Y. Liu, N. Brichtova, A. Culp, N. Braun, D. Graur, W. Ng, N. Mehta, A. Phillips, P. Sundberg, V. Godbole, F. Liu, Y. Katariya, D. Rim, M. Seyedhosseini, S. Ammirati, J. Valfridsson, M. Malihi, T. Knight, A. Toor, T. Lampe, A. Ittycheriah, L. Chiang, C. Yeung, A. Fréchette, J. Rao, H. Wang, H. Srivastava, R. Zhang, R. Rhodes, A. Brand, D. Weesner, I. Figotin, F. Gimeno, R. Fellinger, P. Marcenac, J. Leal, E. Marcus, V. Cotruta, R. Cabrera, S. Luo, D. Garrette, V. Axelrod, S. Baltateanu, D. Barker, D. Chen, H. Toma, B. Ingram, J. Riesa, C. Kulkarni, Y. Zhang, H. Liu, C. Wang, M. Polacek, W. Wu, K. Hui, A. N. Reyes, Y. Su, M. Barnes, I. Malhi, A. Siddiqui, Q. Feng, M. Damaschin, D. Pighin, A. Steiner, S. Yang, R. S. Boppana, S. Ivanov, A. Kandoor, A. Shah, A. Mujika, D. Huang, C. A. Choquette-Choo, M. Patel, T. Yu, T. Creswell, Jerry, Liu, C. Barros, Y. Razeghi, A. Roy, P. Culliton, B. Xiong, J. Pan, T. Strohmann, T. Powell, B. Seal, D. DeCarlo, P. Shyam, K. Katircioglu, X. Wang, C. Hardin, I. Odisho, J. Broder, O. Chang, A. Nair, A. Shtefan, M. O’Brien, M. Agarwal, S. Potluri, S. Goyal, A. Jhindal, S. Thakur, Y. Stuken, J. Lyon, K. Toutanova, F. Feng, A. Wu, B. Horn, A. Wang, A. Cullum, G. Taubman, D. Shrivastava, C. Shi, H. Tomlinson, R. Patel, T. Tu, A. M. Oflazer, F. Pongetti, M. Yang, A. A. Taïga, V. Perot, N. W. Pierse, F. Han, Y. Drori, I. Iturrate, A. Chakrabarti, L. Yeung, D. Dopson, Y. Chen, A. Kulshreshtha, T. Guo, P. Pham, T. Schuster, J. Chen, A. Polozov, J. Xing, H. Zhou, P. Kacham, D. Kukliansky, A. Miech, S. Yaroshenko, E. Chi, S. Douglas, H. Fei, M. Blondel, P. Myla, L. Madmoni, X. Wu, D. Keysers, K. Kjems, I. Albuquerque, L. Yu, J. D’sa, M. Plantan, V. Ionescu, J. S. Elias, A. Gupta, M. R. Vuyyuru, F. Alcober, T. Zhou, K. Ji, F. Hartmann, S. Puttagunta, H. Song, E. Amid, A. Stefanoiu, A. Lee, P. Pucciarelli, E. Wang, A. Raul, S. Petrov, I. Tian, V. Anklin, N. Nti, V. Gomes, M. Schumacher, G. Vesom, A. Panagopoulos, K. Bousmalis, D. Andor, J. Jacob, Y. Zhang, B. Rosgen, M. Kecman, M. Tung, A. Belias, N. Goodman, P. Covington, B. Wieder, N. Saxena, E. Davoodi, M. Huang, S. Maddineni, V. Roulet, F. Campbell-Ajala, P. G. Sessa, Xintian, Wu, G. Lai, P. Collins, A. Haig, V. Sakenas, X. Xu, M. Giustina, L. E. Shafey, P. Charoenpanit, S. Garg, J. Ainslie, B. Severson, M. G. Arenas, S. Pathak, S. Rajayogam, J. Feng, M. Bakker, S. Li, N. Wichers, J. Rogers, X. Geng, Y. Li, R. Jagerman, C. Jia, N. Olmert, D. Sharon, M. Mauger, S. Mariserla, H. Ma, M. Mohabey, K. Kim, A. Andreev, S. Pollom, J. Love, V. Jain, P. Agrawal, Y. Schroecker, A. Fortin, M. Warmuth, J. Liu, A. Leach, I. Blok, G. P. Girirajan, R. Aharoni, B. Uria, A. Sozanschi, D. Goldberg, L. Ionita, M. T. Ribeiro, M. Zlocha, V. Birodkar, S. Lachgar, L. Yuan, H. Choudhury, M. Ginsberg, F. Zheng, G. Dibb, E. Graves, S. Lokhande, G. Rasskin, G. Muraru, C. Quick, S. Tata, P. Sermanet, A. Chawla, I. Karo, Y. Wang, S. Zhang, O. Keller, A. Dragan, G. Su, I. Chou, X. Liu, Y. Tao, S. Prabhakara, M. Wilson, R. Liu, S. Wang, G. Evans, D. Du, A. Castaño, G. Prasad, M. E. Mahdy, S. Gerlach, M. Reid, J. Kahn, A. Zait, T. S. Pillai, T. Ulrich, G. Wang, J. Wassenberg, E. Farkash, K. Yalasangi, C. Wang, M. Bauza, S. Bucher, T. Liu, J. Yan, G. Leung, V. Sindhwani, P. Barnes, A. Singh, I. Jurin, J. Chang, N. K. Bhumihar, S. Eiger, G. Citovsky, B. Withbroe, Z. Li, S. Xue, N. D. Santo, G. Stoyanov, Y. Raimond, S. Zheng, Y. Gao, V. Listík, S. Kwasiborski, R. Saputro, A. Ozturel, G. Mallya, K. Majmundar, R. West, P. Caron, J. Wei, L. Castrejon, S. Vikram, D. Ramachandran, N. Dhawan, J. Park, S. Smoot, G. van den Driessche, Y. Blau, C. Malik, W. Liang, R. Hirsch, C. N. dos Santos, E. Weinstein, A. van den Oord, S. Lall, N. FitzGerald, Z. Jiang, X. Yang, D. Webster, A. Elqursh, A. Pope, G. Rotival, D. Raposo, W. Zhu, J. Dean, S. Alabed, D. Tran, A. Gupta, Z. Gleicher, J. Austin, E. Rosseel, M. Umekar, D. Das, Y. Sun, K. Chen, K. Misiunas, X. Zhou, Y. Di, A. Loo, J. Newlan, B. Li, V. Ramasesh, Y. Xu, A. Chen, S. Gandhe, R. Soricut, N. Gupta, S. Hu, S. El-Sayed, X. Garcia, I. Brusilovsky, P. Chen, A. Bolt, L. Huang, A. Gurney, Z. Zhang, A. Pritzel, J. Wilkiewicz, B. Seybold, B. K. Shamanna, F. Fischer, J. Dean, K. Gill, R. Mcilroy, A. Bhowmick, J. Selier, A. Yang, D. Cheng, V. Magay, J. Tan, D. Varma, C. Walder, T. Kocisky, R. Nakashima, P. Natsev, M. Kwong, I. Gog, C. Zhang, S. Dieleman, T. Jimma, A. Ryabtsev, S. Brahma, D. Steiner, D. Du, A. Žužul, M. Žanić, M. Raghavachari, W. Gierke, Z. Zheng, D. Petrova, Y. Dauphin, Y. Liu, I. Kessler, S. Hand, C. Duvarney, S. Kim, H. Lee, L. Hussenot, J. Hui, J. Smith, D. Jain, J. Xia, G. S. Tomar, K. Amiri, D. Phan, F. Fuchs, T. Weyand, N. Tomasev, A. Cordell, X. Liu, J. Mallinson, P. Joshi, A. Crawford, A. Suggala, S. Chien, N. Fernando, M. Sanchez-Vargas, D. Williams, P. Crone, X. Luo, I. Karpov, J. Shan, T. Thurk, R. Strudel, P. Voigtlaender, P. Patil, T. Dozat, A. Khodaei, S. Singla, P. Ambroszczyk, Q. Wu, Y. Chang, B. Roark, C. Hegde, T. Ding, A. Filos, Z. Wu, A. S. Pinto, S. Liu, S. Khanna, A. Pandey, S. Mcloughlin, Q. Li, S. Haves, A. Zhou, E. Buchatskaya, I. Leal, P. de Boursac, N. Akazawa, N. Anderson, T. Chen, K. Somandepalli, C. Liang, S. Goenka, S. Winkler, A. Grushetsky, Y. Ding, J. Smith, F. Ye, J. Pont-Tuset, E. Li, R. Li, T. Golany, D. Wegner, T. Jiang, O. Barak, Y. Shangguan, E. Vértes, R. Wong, J. Bornschein, A. Tudor, M. Bevilacqua, T. Schaul, A. S. Rawat, Y. Zhao, K. Axiotis, L. Meng, C. McLean, J. Lai, J. Beattie, N. Kushman, Y. Liu, B. Kutzman, F. Lang, J. Ye, P. Netrapalli, P. Mishra, M. Khan, M. Goel, R. Willoughby, D. Tian, H. Zhuang, J. Chen, Z. Tsai, T. Kementsietsidis, A. Khare, J. Keeling, K. Xu, N. Waters, F. Altché, A. Popat, B. Mittal, D. Saxton, D. E. Badawy, M. Mathieu, Z. Zheng, H. Zhou, N. Ranka, R. Shin, Q. Duan, T. Salimans, I. Mihailescu, U. Shaham, M. Chang, Y. Assael, N. Dikkala, M. Izzard, V. Cohen-Addad, C. Graves, V. Feinberg, G. Chung, D. Strouse, D. Karmon, S. Sharifzadeh, Z. Ashwood, K. Pham, J. Blanton, A. Vasiloff, J. Barber, M. Geller, A. Zhou, F. Zubach, T. Huang, L. Zhang, H. Gupta, M. Young, J. Proskurnia, R. Votel, V. Gabeur, G. Barcik, A. Tripathi, H. Yu, G. Yan, B. Changpinyo, F. Pavetić, A. Coyle, Y. Fujii, J. G. Mendez, T. Zhou, H. Rajamani, B. Hechtman, E. Cao, D. Juan, Y. Tan, V. Dalibard, Y. Du, N. Clay, K. Yao, W. Jia, D. Vijaykumar, Y. Zhou, X. Bai, W. Hung, S. Pecht, G. Todorov, N. Khadke, P. Gupta, P. Lahoti, A. Autef, K. Duddu, J. Lee-Thorp, A. Bykovsky, T. Misiunas, S. Flennerhag, S. Thangaraj, J. McGiffin, Z. Nado, M. Kunesch, A. Noever, A. Hertz, M. Liang, V. Stone, E. Palmer, S. Daruki, A. Pramanik, S. Põder, A. Kyker, M. Khan, E. Sluzhaev, M. Ritter, A. Ruderman, W. Zhou, C. Nagpal, K. Vodrahalli, G. Necula, P. Barham, E. Pavlick, J. Hartford, I. Shafran, L. Zhao, M. Mikuła, T. Eccles, H. Shimokawa, K. Garg, L. Vilnis, H. Chen, I. Shumailov, K. Lee, A. Abdelhamed, M. Xie, V. Cohen, E. Hlavnova, D. Malkin, C. Sitawarin, J. Lottes, P. Coquinot, T. Yu, S. Kumar, J. Zhang, A. Mahendru, Z. Ahmed, J. Martens, T. Chen, A. Boag, D. Peng, C. Devin, A. Klimovskiy, M. Phuong, D. Vainstein, J. Xie, B. Ramabhadran, N. Howard, X. Yu, G. Goswami, J. Cui, S. Shleifer, M. Pinto, C. Yeh, M. Yang, S. Javanmardi, D. Ethier, C. Lee, J. Orbay, S. Kotecha, C. Bromberg, P. Shaw, J. Thornton, A. G. Rosenthal, S. Gu, M. Thomas, I. Gemp, A. Ayyar, A. Ushio, A. Selvan, J. Wee, C. Liu, M. Majzoubi, W. Yu, J. Abernethy, T. Liechty, R. Pan, H. Nguyen, Qiong, Hu, S. Perrin, A. Arora, E. Pitler, W. Wang, K. Shivakumar, F. Prost, B. Limonchik, J. Wang, Y. Gao, T. Cour, S. Buch, H. Gui, M. Ivanova, P. Neubeck, K. Chan, L. Kim, H. Chen, N. Goyal, D. Chung, L. Liu, Y. Su, A. Petrushkina, J. Shen, A. Joulin, Y. Xu, S. X. Lin, Y. Kulizhskaya, C. Chelba, S. Vasudevan, E. Collins, V. Bashlovkina, T. Lu, D. Fritz, J. Park, Y. Zhou, C. Su, R. Tanburn, M. Sushkov, M. Rasquinha, J. Li, J. Prendki, Y. Li, P. LV, S. Sharma, H. Fitoussi, H. Huang, A. Dai, P. Dao, M. Burrows, H. Prior, D. Qin, G. Pundak, L. L. Sjoesund, A. Khurshudov, Z. Zhu, A. Webson, E. Kemp, T. Tan, S. Agrawal, S. Sargsyan, L. Cheng, J. Stephan, T. Kwiatkowski, D. Reid, A. Byravan, A. H. Michaely, N. Heess, L. Zhou, S. Goenka, V. Carpenter, A. Levskaya, B. Wang, R. Roberts, R. Leblond, S. Chikkerur, S. Ginzburg, M. Chang, R. Riachi, Chuqiao, Xu, Z. Borsos, M. Pliskin, J. Pawar, M. Lustman, H. Kirkwood, A. Anand, A. Chaudhary, N. Kalb, K. Milan, S. Augenstein, A. Goldie, L. Prince, K. Raman, Y. Sun, V. Xia, A. Cohen, Z. Huo, J. Camp, S. Ellis, L. Zilka, D. V. Torres, L. Patel, S. Arora, B. Chan, J. Adler, K. Ayoub, J. Liang, F. Jamil, J. Jiang, S. Baumgartner, H. Sun, Y. Karov, Y. Akulov, H. Zheng, I. Cai, C. Fantacci, J. Rubin, A. R. Acha, M. Wang, N. D’Souza, R. Sathyanarayana, S. Dai, S. Rowe, A. Simanovsky, O. Goldman, Y. Kuang, X. Pan, A. Rosenberg, T. Rojas-Esponda, P. Dutta, A. Zeng, I. Jurenka, G. Farquhar, Y. Bansal, S. Iqbal, B. Roelofs, G. Joung, P. Beak, C. Ryu, R. Poplin, Y. Wu, J. Alayrac, S. Buthpitiya, O. Ronneberger, C. Habtegebriel, W. Li, P. Cavallaro, A. Wei, G. Bensky, T. Denk, H. Ganapathy, J. Stanway, P. Joshi, F. Bertolini, J. Lo, O. Ma, Z. Charles, G. Sampemane, H. Sahni, X. Chen, H. Askham, D. Gaddy, P. Young, J. Tan, M. Eyal, A. Bražinskas, L. Zhong, Z. Wu, M. Epstein, K. Bailey, A. Hard, K. Lee, S. Goldshtein, A. Ruiz, M. Badawi, M. Lochbrunner, J. Kearns, A. Brown, F. Pardo, T. Weber, H. Yang, P. Jiang, B. Akin, Z. Fu, M. Wainwright, C. Zou, M. Gaba, P. Manzagol, W. Kan, Y. Song, K. Zainullina, R. Lin, J. Ko, S. Deshmukh, A. Jindal, J. Svensson, D. Tyam, H. Zhao, C. Kaeser-Chen, S. Baird, P. Moradi, J. Hall, Q. Guo, V. Tsang, B. Liang, F. Pereira, S. Ganesh, I. Korotkov, J. Adamek, S. Thiagarajan, V. Tran, C. Chen, C. Tar, S. Jain, I. Dasgupta, T. Bilal, D. Reitter, K. Zhao, G. Vezzani, Y. Gehman, P. Mehta, L. Beltrone, X. Dotiwalla, S. Guadarrama, Z. Abbas, S. Karp, P. Georgiev, C. Ferng, M. Brockschmidt, L. Peng, C. Hirnschall, V. Verma, Y. Bi, Y. Xiao, A. Dabush, K. Xu, P. Wallis, R. Parker, Q. Wang, Y. Xu, I. Safarli, D. Tewari, Y. Zhang, S. Kim, A. Gesmundo, M. Thomas, S. Levi, A. Chowdhury, K. Rao, P. Garst, S. Conway-Rahman, H. Ran, K. McKinney, Z. Xiao, W. Yu, R. Agrawal, A. Stjerngren, C. Ionescu, J. Chen, V. Sharma, J. Chiu, F. Liu, K. Franko, C. Sanford, X. Cai, P. Michel, S. Ganapathy, J. Labanowski, Z. Garrett, B. Vargas, S. Sun, B. Gale, T. Buschmann, G. Desjardins, N. Ghelani, P. Jain, M. Verma, C. Asawaroengchai, J. Eisenschlos, J. Harlalka, H. Kazawa, D. Metzler, J. Howland, Y. Jian, J. Ades, V. Shah, T. Gangwani, S. Lee, R. Ring, S. M. Hernandez, D. Reich, A. Sinha, A. Sathe, J. Kovac, A. Gill, A. Kannan, A. D’olimpio, M. Sevenich, J. Whang, B. Kim, K. C. Sim, J. Chen, J. Zhang, S. Lall, Y. Matias, B. Jia, A. Friesen, S. Nasso, A. Thapliyal, B. Perozzi, T. Yu, A. Shekhawat, S. Huda, P. Grabowski, E. Wang, A. Sreevatsa, H. Dib, M. Hassen, P. Schuh, V. Milutinovic, C. Welty, M. Quinn, A. Shah, B. Wang, G. Barth-Maron, J. Frye, N. Axelsson, T. Zhu, Y. Ma, I. Giannoumis, H. Sedghi, C. Ye, Y. Luan, K. Aydin, B. Chandra, V. Sampathkumar, R. Huang, V. Lavrenko, A. Eleryan, Z. Hong, S. Hansen, S. M. Carthy, B. Samanta, D. Ćevid, X. Wang, F. Li, M. Voznesensky, M. Hoffman, A. Terzis, V. Sehwag, G. Fidel, L. He, M. Cai, Y. He, A. Feng, M. Nikoltchev, S. Phatale, J. Chase, R. Lawton, M. Zhang, T. Ouyang, M. Tragut, M. H. Manshadi, A. Narayanan, J. Shen, X. Gao, T. Bolukbasi, N. Roy, X. Li, D. Golovin, L. Panait, Z. Qin, G. Han, T. Anthony, S. Kudugunta, V. Patraucean, A. Ray, X. Chen, X. Yang, T. Bhatia, P. Talluri, A. Morris, A. Ražnatović, B. Brownfield, J. An, S. Peng, P. Kane, C. Zheng, N. Duduta, J. Kessinger, J. Noraky, S. Liu, K. Rong, P. Veličković, K. Rush, A. Goldin, F. Wei, S. M. R. Garlapati, C. Pantofaru, O. Kwon, J. Ni, E. Noland, J. D. Trapani, F. Beaufays, A. G. Roy, Y. Chow, A. Turker, G. Cideron, L. Mei, J. Clark, Q. Dou, M. Bošnjak, R. Leith, Y. Du, A. Yazdanbakhsh, M. Nasr, C. Kwak, S. S. Sheth, A. Kaskasoli, A. Anand, B. Lakshminarayanan, S. Jerome, D. Bieber, C. Chu, A. Senges, T. Shen, M. Sridhar, N. Ndebele, B. Beyret, S. Mohamed, M. Chen, M. Freitag, J. Guo, L. Liu, P. Roit, H. Chen, S. Yan, T. Stone, J. Co-Reyes, J. Cole, S. Scellato, S. Azizi, H. Hashemi, A. Jin, A. Iyer, M. Valentine, A. György, A. Ahuja, D. H. Diaz, C. Lee, N. Clement, W. Kong, D. Garmon, I. Watts, K. Bhatia, K. Gupta, M. Miecnikowski, H. Vallet, A. Taly, E. Loper, S. Joshi, J. Atwood, J. Chick, M. Collier, F. Iliopoulos, R. Trostle, B. Gunel, R. Leal-Cavazos, A. M. Hrafnkelsson, M. Guzman, X. Ju, A. Forbes, J. Emond, K. Chauhan, B. Caine, L. Xiao, W. Zeng, A. Moufarek, D. Murphy, M. Meng, N. Gupta, F. Riedel, A. Das, E. Lawal, S. Narayan, T. Sosea, J. Swirhun, L. Friso, B. Neyshabur, J. Lu, S. Girgin, M. Wunder, E. Yvinec, A. Pyne, V. Carbune, S. Rijhwani, Y. Guo, T. Doshi, A. Briukhov, M. Bain, A. Hitron, X. Wang, A. Gupta, K. Chen, C. Du, W. Zhang, D. Shah, A. Akula, M. Dylla, A. Kachra, W. Kuo, T. Zou, L. Wang, L. Xu, J. Zhu, J. Snyder, S. Menon, O. Firat, I. Mordatch, Y. Yuan, N. Ponomareva, R. Blevins, L. Moore, W. Wang, P. Chen, M. Scholz, A. Dwornik, J. Lin, S. Li, D. Antognini, T. I, X. Song, M. Miller, U. Kalra, A. Raveret, O. Akerlund, F. Wu, A. Nystrom, N. Godbole, T. Liu, H. DeBalsi, J. Zhao, B. Liu, A. Caciularu, L. Lax, U. Khandelwal, V. Langston, E. Bailey, S. Lattanzi, Y. Wang, N. Kovelamudi, S. Mondal, G. Guruganesh, N. Hua, O. Roval, P. Wesołowski, R. Ingale, J. Halcrow, T. Sohn, C. Angermueller, B. Raad, E. Stickgold, E. Lu, A. Kosik, J. Xie, T. Lillicrap, A. Huang, L. L. Zhang, D. Paulus, C. Farabet, A. Wertheim, B. Wang, R. Joshi, C. Ko, Y. Wu, S. Agrawal, L. Lin, X. Sheng, P. Sung, T. Breland-King, C. Butterfield, S. Gawde, S. Singh, Q. Zhang, R. Apte, S. Shetty, A. Hutter, T. Li, E. Salesky, F. Lebron, J. Kanerva, M. Paganini, A. Nguyen, R. Vallu, J. Peter, S. Velury, D. Kao, J. Hoover, A. Bortsova, C. Bishop, S. Jakobovits, A. Agostini, A. Agarwal, C. Liu, C. Kwong, S. Tavakkol, I. Bica, A. Greve, A. GP, J. Marcus, L. Hou, T. Duerig, R. Moroshko, D. Lacey, A. Davis, J. Amelot, G. Wang, F. Kim, T. Strinopoulos, H. Wan, C. L. Lan, S. Krishnan, H. Tang, P. Humphreys, J. Bai, I. H. Shtacher, D. Machado, C. Pang, K. Burke, D. Liu, R. Aravamudhan, Y. Song, E. Hirst, A. Singh, B. Jou, L. Bai, F. Piccinno, C. K. Fu, R. Alazard, B. Meiri, D. Winter, C. Chen, M. Zhang, J. Heitkaemper, J. Lambert, J. Lee, A. Frömmgen, S. Rogulenko, P. Nair, P. Niemczyk, A. Bulyenov, B. Xu, H. Shemtov, M. Zadimoghaddam, S. Toropov, M. Wirth, H. Dai, S. Gollapudi, D. Zheng, A. Kurakin, C. Lee, K. Bullard, N. Serrano, I. Balazevic, Y. Li, J. Schalkwyk, M. Murphy, M. Zhang, K. Sequeira, R. Datta, N. Agrawal, C. Sutton, N. Attaluri, M. Chiang, W. Farhan, G. Thornton, K. Lin, T. Choma, H. Nguyen, K. Dasgupta, D. Robinson, I. Comşa, M. Riley, A. Pillai, B. Mustafa, B. Golan, A. Zandieh, J. Lespiau, B. Porter, D. Ross, S. Rajayogam, M. Agarwal, S. Venugopalan, B. Shahriari, Q. Yan, H. Xu, T. Tobin, P. Dubov, H. Shi, A. Recasens, A. Kovsharov, S. Borgeaud, L. Dery, S. Vasanth, E. Gribovskaya, L. Qiu, M. Mahdieh, W. Skut, E. Nielsen, C. Zheng, A. Yu, C. G. Bostock, S. Gupta, A. Archer, C. Rawles, E. Davies, A. Svyatkovskiy, T. Tsai, Y. Halpern, C. Reisswig, B. Wydrowski, B. Chang, J. Puigcerver, M. H. Taege, J. Li, E. Schnider, X. Li, D. Dena, Y. Xu, U. Telang, T. Shi, H. Zen, K. Kastner, Y. Ko, N. Subramaniam, A. Kumar, P. Blois, Z. Dai, J. Wieting, Y. Lu, Y. Zeldes, T. Xie, A. Hauth, A. Ţifrea, Y. Li, S. El-Husseini, D. Abolafia, H. Zhou, W. Ding, S. Ghalebikesabi, C. Guía, A. Maksai, Á. Weisz, S. Arik, N. Sukhanov, A. Świetlik, X. Jia, L. Yu, W. Wang, M. Brand, D. Bloxwich, S. Kirmani, Z. Chen, A. Go, P. Sprechmann, N. Kannen, A. Carin, P. Sandhu, I. Edkins, L. Nooteboom, J. Gupta, L. Maggiore, J. Azizi, Y. Pritch, P. Yin, M. Gupta, D. Tarlow, D. Smith, D. Ivanov, M. Babaeizadeh, A. Goel, S. Kambala, G. Chu, M. Kastelic, M. Liu, H. Soltau, A. Stone, S. Agrawal, M. Kim, K. Soparkar, S. Tadepalli, O. Bunyan, R. Soh, A. Kannan, D. Kim, B. J. Chen, A. Halumi, S. Roy, Y. Wang, O. Sercinoglu, G. Gibson, S. Bhatnagar, M. Sano, D. von Dincklage, Q. Ren, B. Mitrevski, M. Olšák, J. She, C. Doersch, Jilei, Wang, B. Liu, Q. Tan, T. Yakar, T. Warkentin, A. Ramirez, C. Lebsack, J. Dillon, R. Mathews, T. Cobley, Z. Wu, Z. Chen, J. Simon, S. Nath, T. Sainath, A. Bendebury, R. Julian, B. Mankalale, D. Ćurko, P. Zacchello, A. R. Brown, K. Sodhia, H. Howard, S. Caelles, A. Gupta, G. Evans, A. Bulanova, L. Katzen, R. Goldenberg, A. Tsitsulin, J. Stanton, B. Schillings, V. Kovalev, C. Fry, R. Shah, K. Lin, S. Upadhyay, C. Li, S. Radpour, M. Maggioni, J. Xiong, L. Haas, J. Brennan, A. Kamath, N. Savinov, A. Nagrani, T. Yacovone, R. Kappedal, K. Andriopoulos, L. Lao, Y. Li, G. Rozhdestvenskiy, K. Hashimoto, A. Audibert, S. Austin, D. Rodriguez, A. Ruoss, G. Honke, D. Karkhanis, X. Xiong, Q. Wei, J. Huang, Z. Leng, V. Premachandran, S. Bileschi, G. Evangelopoulos, T. Mensink, J. Pavagadhi, D. Teplyashin, P. Chang, L. Xue, G. Tanzer, S. Goldman, K. Patel, S. Li, J. Wiesner, I. Zheng, I. Stewart-Binks, J. Han, Z. Li, L. Luo, K. Lenc, M. Lučić, F. Xue, R. Mullins, A. Guseynov, C. Chang, I. Galatzer-Levy, A. Zhang, G. Bingham, G. Hu, A. Hartman, Y. Ma, J. Griffith, A. Irpan, C. Radebaugh, S. Yue, L. Fan, V. Ungureanu, C. Sorokin, H. Teufel, P. Li, R. Anil, D. Paparas, T. Wang, C. Lin, H. Peng, M. Shum, G. Petrovic, D. Brady, R. Nguyen, K. Macherey, Z. Li, H. Singh, M. Yenugula, M. Iinuma, X. Chen, K. Kopparapu, A. Stern, S. Dave, C. Thekkath, F. Perot, A. Kumar, F. Li, Y. Xiao, M. Bilotti, M. H. Bateni, I. Noble, L. Lee, A. Vázquez-Reina, J. Salazar, X. Yang, B. Wang, E. Gruzewska, A. Rao, S. Raghuram, Z. Xu, E. Ben-David, J. Mei, S. Dalmia, Z. Zhang, Y. Liu, G. Bansal, H. Pankov, S. Schwarcz, A. Burns, C. Chan, S. Sanghai, R. Liang, E. Liang, A. He, A. Stuart, A. Narayanan, Y. Zhu, C. Frank, B. Fatemi, A. Sabne, O. Lang, I. Bhattacharya, S. Settle, M. Wang, B. McMahan, A. Tacchetti, L. B. Soares, M. Hadian, S. Cabi, T. Chung, N. Putikhin, G. Li, J. Chen, A. Tarango, H. Michalewski, M. Kazemi, H. Masoom, H. Sheftel, R. Shivanna, A. Vadali, R. Comanescu, D. Reid, J. Moore, A. Neelakantan, M. Sander, J. Herzig, A. Rosenberg, M. Dehghani, J. Choi, M. Fink, R. Hayes, E. Ge, S. Weng, C. Ho, J. Karro, K. Krishna, L. N. Thiet, A. Skerry-Ryan, D. Eppens, M. Andreetto, N. Sarma, S. Bonacina, B. K. Ayan, M. Nawhal, Z. Shan, M. Dusenberry, S. Thakoor, S. Gubbi, D. D. Nguyen, R. Tsarfaty, S. Albanie, J. Mitrović, M. Gandhi, B. Chen, A. Epasto, G. Stephanov, Y. Jin, S. Gehman, A. Amini, J. Weber, F. Behbahani, S. Xu, M. Allamanis, X. Chen, M. Ott, C. Sha, M. Jastrzebski, H. Qi, D. Greene, X. Wu, A. Toki, D. Vlasic, J. Shapiro, R. Kotikalapudi, Z. Shen, T. Saeki, S. Xie, A. Cassirer, S. Bharadwaj, T. Kiyono, S. Bhojanapalli, E. Rosenfeld, S. Ritter, J. Mao, J. G. Oliveira, Z. Egyed, B. Bandemer, E. Parisotto, K. Kinoshita, J. Pluto, P. Maniatis, S. Li, Y. Guo, G. Ghiasi, J. Tarbouriech, S. Chatterjee, J. Jin, Katrina, Xu, J. Palomaki, S. Arnold, M. Sewak, F. Piccinini, M. Sharma, B. Albrecht, S. Purser-haskell, A. Vaswani, C. Chen, M. Wisniewski, Q. Cao, J. Aslanides, N. M. Phu, M. Sieb, L. Agubuzu, A. Zheng, D. Sohn, M. Selvi, A. Andreassen, K. Subudhi, P. Eruvbetine, O. Woodman, T. Mery, S. Krause, X. Ren, X. Ma, J. Luo, D. Chen, W. Fan, H. Griffiths, C. Schuler, A. Li, S. Zhang, J. Sarr, S. Luo, R. Patana, M. Watson, D. Naboulsi, M. Collins, S. Sidhwani, E. Hoogeboom, S. Silver, E. Caveness, X. Zhao, M. Rodriguez, M. Deines, L. Bai, P. Griffin, M. Tagliasacchi, E. Xue, S. R. Babbula, B. Pang, N. Ding, G. Shen, E. Peake, R. Crocker, S. S. Raghvendra, D. Swisher, W. Han, R. Singh, L. Wu, V. Pchelin, T. Munkhdalai, D. Alon, G. Bacon, E. Robles, J. Bulian, M. Johnson, G. Powell, F. T. Ferreira, Y. Li, F. Benzing, M. Velimirović, H. Soyer, W. Kong, Tony, Nguyên, Z. Yang, J. Liu, J. van Amersfoort, D. Gillick, B. Sun, N. Rauschmayr, K. Zhang, S. Zhan, T. Zhou, A. Frolov, C. Yang, D. Vnukov, L. Rouillard, H. Li, A. Mandhane, N. Fallen, R. Venkataraman, C. H. Hu, J. Brennan, J. Lee, J. Chang, M. Sundermeyer, Z. Pan, R. Ke, S. Tong, A. Fabrikant, W. Bono, J. Gu, R. Foley, Y. Mao, M. Delakis, D. Bhaswar, R. Frostig, N. Li, A. Zipori, C. Hope, O. Kozlova, S. Mishra, J. Djolonga, C. Schiff, M. A. Merey, E. Briakou, P. Morgan, A. Wan, A. Hassidim, R. Skerry-Ryan, K. Sengupta, M. Jasarevic, P. Kallakuri, P. Kunkle, H. Brennan, T. Lieber, H. Mansoor, J. Walker, B. Zhang, A. Xie, G. Žužić, A. Chukwuka, A. Druinsky, D. Cho, R. Yao, F. Naeem, S. Butt, E. Kim, Z. Jia, M. Jordan, A. Lelkes, M. Kurzeja, S. Wang, J. Zhao, A. Over, A. Chakladar, M. Prasetya, N. Jha, S. Ganapathy, Y. Cong, P. Shroff, C. Saroufim, S. Miryoosefi, M. Hammad, T. Nasir, W. Xi, Y. Gao, Y. Maeng, B. Hora, C. Cheng, P. Haghani, Y. Lewenberg, C. Lu, M. Matysiak, N. Raisinghani, H. Wang, L. Baugher, R. Sukthankar, M. Giang, J. Schultz, N. Fiedel, M. Chen, C. Lee, T. Dey, H. Zheng, S. Paul, C. Smith, A. Ly, Y. Wang, R. Bansal, B. Perz, S. Ricco, S. Blank, V. Keshava, D. Sharma, M. Chow, K. Lad, K. Jalan, S. Osindero, C. Swanson, J. Scott, A. Ilić, X. Li, S. R. Jonnalagadda, A. S. Soudagar, Y. Xiong, B. Batsaikhan, D. Jarrett, N. Kumar, M. Shah, M. Lawlor, A. Waters, M. Graham, R. May, S. Ramos, S. Lefdal, Z. Cankara, N. Cano, B. O’Donoghue, J. Borovik, F. Liu, J. Grimstad, M. Alnahlawi, K. Tsihlas, T. Hudson, N. Grigorev, Y. Jia, T. Huang, T. P. Igwe, S. Lebedev, X. Tang, I. Krivokon, F. Garcia, M. Tan, E. Jia, P. Stys, S. Vashishth, Y. Liang, B. Venkatraman, C. Gu, A. Kementsietsidis, C. Zhu, J. Jung, Y. Bai, M. J. Hosseini, F. Ahmed, A. Gupta, X. Yuan, S. Ashraf, S. Nigam, G. Vasudevan, P. Awasthi, A. M. Gilady, Z. Mariet, R. Eskander, H. Li, H. Hu, G. Garrido, P. Schlattner, G. Zhang, R. Saxena, P. Dević, K. Muralidharan, A. Murthy, Y. Zhou, M. Choi, A. Wongpanich, Z. Wang, P. Shah, Y. Xu, Y. Huang, S. Spencer, A. Chen, J. Cohan, J. Wang, J. Tompson, J. Wu, R. Haroun, H. Li, B. Huergo, F. Yang, T. Yin, J. Wendt, M. Bendersky, R. Chaabouni, J. Snaider, J. Ferret, A. Jindal, T. Thompson, A. Xue, W. Bishop, S. M. Phal, A. Sharma, Y. Sung, P. Radhakrishnan, M. Shomrat, R. Ingle, R. Vij, J. Gilmer, M. D. Istin, S. Sobell, Y. Lu, E. Nottage, D. Sadigh, J. Willcock, T. Zhang, S. Xu, S. Brown, K. Lee, G. Wang, Y. Zhu, Y. Tay, C. Kim, A. Gutierrez, A. Sharma, Y. Xian, S. Seo, C. Cui, E. Pochernina, C. Baetu, K. Jastrzębski, M. Ly, M. Elhawaty, D. Suh, E. Sezener, P. Wang, N. Yuen, G. Tucker, J. Cai, Z. Yang, C. Wang, A. Muzio, H. Qian, J. Yoo, D. Lockhart, K. R. McKee, M. Guo, M. Mehrotra, A. Mendonça, S. V. Mehta, S. Ben, C. Tekur, J. Mu, M. Zhu, V. Krakovna, H. Lee, A. Maschinot, S. Cevey, H. Choe, A. Bai, H. Srinivasan, D. Gasaway, N. Young, P. Siegler, D. Holtmann-Rice, V. Piratla, K. Baumli, R. Yogev, A. Hofer, H. van Hasselt, S. Grant, Y. Chervonyi, D. Silver, A. Hogue, A. Agarwal, K. Wang, P. Singh, F. Flynn, J. Lipschultz, R. David, L. Bellot, Y. Yang, L. Le, F. Graziano, K. Olszewska, K. Hui, A. Maurya, N. Parotsidis, W. Chen, T. Oguntebi, J. Kelley, A. Baddepudi, J. Mauerer, G. Shaw, A. Siegman, L. Yang, S. Shetty, S. Roy, Y. Song, W. Stokowiec, R. Burnell, O. Savant, R. Busa-Fekete, J. Miao, S. Ghosh, L. MacDermed, P. Lippe, M. Dektiarev, Z. Behrman, F. Mentzer, K. Nguyen, M. Wei, S. Verma, C. Knutsen, S. Dasari, Z. Yan, P. Mitrichev, X. Wang, V. Shejwalkar, J. Austin, S. Sunkara, N. Potti, Y. Virin, C. Wright, G. Liu, O. Riva, E. Pot, G. Kochanski, Q. Le, G. Balasubramaniam, A. Dhar, Y. Liao, A. Bloniarz, D. Shukla, E. Cole, J. Lee, S. Zhang, S. Kafle, S. Vashishtha, P. Mahmoudieh, G. Chen, R. Hoffmann, P. Srinivasan, A. D. Lago, Y. B. Shalom, Z. Wang, M. Elabd, A. Sharma, J. Oh, S. Kothawade, M. Le, M. Monteiro, S. Yang, K. Alarakyia, R. Geirhos, D. Mincu, H. Garnes, H. Kobayashi, S. Mariooryad, K. Krasowiak, Zhixin, Lai, S. Mourad, M. Wang, F. Bu, O. Aharoni, G. Chen, A. Goyal, V. Zubov, A. Bapna, E. Dabir, N. Kothari, K. Lamerigts, N. D. Cao, J. Shar, C. Yew, N. Kulkarni, D. Mahaarachchi, M. Joshi, Z. Zhu, J. Lichtarge, Y. Zhou, H. Muckenhirn, V. Selo, O. Vinyals, P. Chen, A. Brohan, V. Mehta, S. Cogan, R. Wang, T. Geri, W. Ko, W. Chen, F. Viola, K. Shivam, L. Wang, M. C. Elish, R. A. Popa, S. Pereira, J. Liu, R. Koster, D. Kim, G. Zhang, S. Ebrahimi, P. Talukdar, Y. Zheng, P. Poklukar, A. Mikhalap, D. Johnson, A. Vijayakumar, M. Omernick, M. Dibb, A. Dubey, Q. Hu, A. Suman, V. Aggarwal, I. Kornakov, F. Xia, W. Lowe, A. Kolganov, T. Xiao, V. Nikolaev, S. Hemingray, B. Li, J. Iljazi, M. Rybiński, B. Sandhu, P. Lu, T. Luong, R. Jenatton, V. Govindaraj, Hui, Li, G. Dulac-Arnold, W. Park, H. Wang, A. Modi, J. Pouget-Abadie, K. Greller, R. Gupta, R. Berry, P. Ramachandran, J. Xie, L. McCafferty, J. Wang, K. Gupta, H. Lim, B. Bratanič, A. Brock, I. Akolzin, J. Sproch, D. Karliner, D. Kim, A. Goedeckemeyer, N. Shazeer, C. Schmid, D. Calandriello, P. Bhatia, K. Choromanski, C. Montgomery, D. Dua, A. Ramalho, H. King, Y. Gao, L. Nguyen, D. Lindner, D. Pitta, O. Johnson, K. Salama, D. Ardila, M. Han, E. Farnese, S. Odoom, Z. Wang, X. Ding, N. Rink, R. Smith, H. T. Lehri, E. Cohen, N. Vats, T. He, P. Gopavarapu, A. Paszke, M. Patel, W. V. Gansbeke, L. Loher, L. Castro, M. Voitovich, T. von Glehn, N. George, S. Niklaus, Z. Eaton-Rosen, N. Rakićević, E. Jue, S. Perel, C. Zhang, Y. Bahat, A. Pouget, Z. Xing, F. Huot, A. Shenoy, T. Bos, V. Coriou, B. Richter, N. Noy, Y. Wang, S. Ontanon, S. Qin, G. Makarchuk, D. Hassabis, Z. Li, M. Sharma, K. Venkatesan, I. Kemaev, R. Daniel, S. Huang, S. Shah, O. Ponce, Warren, Chen, M. Faruqui, J. Wu, S. Andačić, S. Payrits, D. McDuff, T. Hume, Y. Cao, M. Tessler, Q. Wang, Y. Wang, I. Rendulic, E. Agustsson, M. Johnson, T. Lando, A. Howard, S. G. S. Padmanabhan, M. Daswani, A. Banino, M. Kilgore, J. Heek, Z. Ji, A. Caceres, C. Li, N. Kassner, A. Vlaskin, Z. Liu, A. Grills, Y. Hou, R. Sukkerd, G. Cheon, N. Shetty, L. Markeeva, P. Stanczyk, T. Iyer, Y. Gong, S. Gao, K. Gopalakrishnan, T. Blyth, M. Reynolds, A. Bhoopchand, M. Bilenko, D. Gharibian, V. Zayats, A. Faust, A. Singh, M. Ma, H. Jiao, S. Vijayanarasimhan, L. Aroyo, V. Yadav, S. Chakera, A. Kakarla, V. Meshram, K. Gregor, G. Botea, E. Senter, D. Jia, G. Kovacs, N. Sharma, S. Baur, K. Kang, Y. He, L. Zhuo, M. Kostelac, I. Laish, S. Peng, L. O’Bryan, D. Kasenberg, G. R. Rao, E. Leurent, B. Zhang, S. Stevens, A. Salazar, Y. Zhang, I. Lobov, J. Walker, A. Porter, M. Redshaw, H. Ke, A. Rao, A. Lee, H. Lam, M. Moffitt, J. Kim, S. Qiao, T. Koo, R. Dadashi, X. Song, M. Sundararajan, P. Xu, C. Kawamoto, Y. Zhong, C. Barbu, A. Reddy, M. Verzetti, L. Li, G. Papamakarios, H. Klimczak-Plucińska, M. Cassin, K. Kavukcuoglu, R. Swavely, A. Vaucher, J. Zhao, R. Hemsley, M. Tschannen, H. Ge, G. Menghani, Y. Yu, N. Ha, W. He, X. Wu, M. Song, R. Sterneck, S. Zinke, D. A. Calian, A. Marsden, A. C. Ruiz, M. Hessel, A. Gueta, B. Lee, B. Farris, M. Gupta, Y. Li, M. Saleh, V. Misra, K. Xiao, P. Mendolicchio, G. Buttimore, V. Krayvanova, N. Nayakanti, M. Wiethoff, Y. Pande, A. Mirhoseini, N. Lao, J. Liu, Y. Hua, A. Chen, Y. Malkov, D. Kalashnikov, S. Gupta, K. Audhkhasi, Y. Zhai, S. Kopalle, P. Jain, E. Ofek, C. Meyer, K. Baatarsukh, H. Strejček, J. Qian, J. Freedman, R. Figueira, M. Sokolik, O. Bachem, R. Lin, D. Kharrat, C. Hidey, P. Xu, D. Duan, Y. Li, M. Ersoy, R. Everett, K. Cen, R. Santamaria-Fernandez, A. Taubenfeld, I. Mackinnon, L. Deng, P. Zablotskaia, S. Viswanadha, S. Goel, D. Yates, Y. Deng, P. Choy, M. Chen, A. Sinha, A. Mossin, Y. Wang, A. Szlam, S. Hao, P. K. Rubenstein, M. Toksoz-Exley, M. Aperghis, Y. Zhong, J. Ahn, M. Isard, O. Lacombe, F. Luisier, C. Anastasiou, Y. Kalley, U. Prabhu, E. Dunleavy, S. Bijwadia, J. Mao-Jones, K. Chen, R. Pasumarthi, E. Wood, A. Dostmohamed, N. Hurley, J. Simsa, A. Parrish, M. Pajarskas, M. Harvey, O. Skopek, Y. Kochinski, J. Rey, V. Rieser, D. Zhou, S. J. Lee, T. Acharya, G. Li, J. Jiang, X. Zhang, B. Gipson, E. Mahintorabi, M. Gelmi, N. Khajehnouri, A. Yeh, K. Lee, L. Matthey, L. Baker, T. Pham, H. Fu, A. Pak, P. Gupta, C. Vasconcelos, A. Sadovsky, B. Walker, S. Hsiao, P. Zochbauer, A. Marzoca, N. Velan, J. Zeng, G. Baechler, D. Driess, D. Jain, Y. Huang, L. Tao, J. Maggs, N. Levine, J. Schneider, E. Gemzer, S. Petit, S. Han, Z. Fisher, D. Zelle, C. Biles, E. Ie, A. Fadeeva, C. Liu, J. V. Franco, A. Collister, H. Zhang, R. Wang, R. Zhao, L. Kieliger, K. Shuster, R. Zhu, B. Gong, L. Chan, R. Sun, S. Basu, R. Zimmermann, J. Hayes, A. Bapna, J. Snoek, W. Yang, P. Datta, J. A. Abdallah, K. Kilgour, L. Li, S. Mah, Y. Jun, M. Rivière, A. Karmarkar, T. Spalink, T. Huang, L. Gonzalez, D. Tran, A. Nowak, J. Palowitch, M. Chadwick, E. Talius, H. Mehta, T. Sellam, P. Fränken, M. Nicosia, K. He, A. Kini, D. Amos, S. Basu, H. Jobe, E. Shaw, Q. Xu, C. Evans, D. Ikeda, C. Yan, L. Jin, L. Wang, S. Yadav, I. Labzovsky, R. Sampath, A. Ma, C. Schumann, A. Siddhant, R. Shah, J. Youssef, R. Agarwal, N. Dabney, A. Tonioni, M. Ambar, J. Li, I. Guyon, B. Li, D. Soergel, B. Fang, G. Karadzhov, C. Udrescu, T. Trinh, V. Raunak, S. Noury, D. Guo, S. Gupta, M. Finkelstein, D. Petek, L. Liang, G. Billock, P. Sun, D. Wood, Y. Song, X. Yu, T. Matejovicova, R. Cohen, K. Andra, D. D’Ambrosio, Z. Deng, V. Nallatamby, E. Songhori, R. Dangovski, A. Lampinen, P. Botadra, A. Hillier, J. Cao, N. Baddi, A. Kuncoro, T. Yoshino, A. Bhagatwala, M. Ranzato, R. Schaeffer, T. Liu, S. Ye, O. Sarvana, J. Nham, C. Kuang, I. Gao, J. Baek, S. Mittal, A. Wahid, A. Gergely, B. Ni, J. Feldman, C. Muir, P. Lamblin, W. Macherey, E. Dyer, L. Kilpatrick, V. Campos, M. Bhutani, S. Fort, Y. Ahmad, A. Severyn, K. Chatziprimou, O. Ferludin, M. Dimarco, A. Kusupati, J. Heyward, D. Bahir, K. Villela, K. Millican, D. Marcus, S. Bahargam, C. Unlu, N. Roth, Z. Wei, S. Gopal, D. Ghoshal, E. Lee, S. Lin, J. Lees, D. Lee, A. Hosseini, C. Fan, S. Neel, M. Wu, Y. Altun, H. Cai, E. Piqueras, J. Woodward, A. Bissacco, S. Haykal, M. Bordbar, P. Sundaram, S. Hodkinson, D. Toyama, G. Polovets, A. Myers, A. Sinha, T. Levinboim, K. Krishnakumar, R. Chhaparia, T. Sholokhova, N. B. Gundavarapu, G. Jawahar, H. Qureshi, J. Hu, N. Momchev, M. Rahtz, R. Wu, A. P. S, K. Dhamdhere, M. Guo, U. Gupta, A. Eslami, M. Schain, M. Blokzijl, D. Welling, D. Orr, L. Bolelli, N. Perez-Nieves, M. Sirotenko, A. Prasad, A. Kar, B. D. B. Pigem, T. Terzi, G. Weisz, D. Ghosh, A. Mavalankar, D. Madeka, K. Daugaard, H. Adam, V. Shah, D. Berman, M. Tran, S. Baker, E. Andrejczuk, G. Chole, G. Raboshchuk, M. Mirzazadeh, T. Kagohara, S. Wu, C. Schallhart, B. Orlando, C. Wang, A. Rrustemi, H. Xiong, H. Liu, A. Vezer, N. Ramsden, S. Chang, S. Mudgal, Y. Li, N. Vieillard, Y. Hoshen, F. Ahmad, A. Slone, A. Hua, N. Potikha, M. Rossini, J. Stritar, S. Prakash, Z. Wang, X. Dong, A. Nazari, E. Nehoran, K. Tekelioglu, Y. Li, K. Badola, T. Funkhouser, Y. Li, V. Yerram, R. Ganeshan, D. Formoso, K. Langner, T. Shi, H. Li, Y. Yamamori, A. Panda, A. Saade, A. S. Scarpati, C. Breaux, C. Carey, Z. Zhou, C. Hsieh, S. Bridgers, A. Butryna, N. Gupta, V. Tulsyan, S. Woo, E. Eltyshev, W. Grathwohl, C. Parks, S. Benjamin, R. Panigrahy, S. Dodhia, D. D. Freitas, C. Sauer, W. Song, F. Alet, J. Tolins, C. Paduraru, X. Zhou, B. Albert, Z. Zhang, L. Shu, M. Bansal, S. Nguyen, A. Globerson, O. Xiao, J. Manyika, T. Hennigan, R. Rong, J. Matak, A. Bakalov, A. Sharma, D. Sinopalnikov, A. Pierson, S. Roller, G. Brown, M. Gao, T. Fukuzawa, A. Ghafouri, K. Vassigh, I. Barr, Z. Wang, A. Korsun, R. Jayaram, L. Ren, T. Zaman, S. Khan, Y. Lunts, D. Deutsch, D. Uthus, N. Katz, M. Samsikova, A. Khalifa, N. Sethi, J. Sun, L. Tang, U. Alon, X. Luo, D. Yu, A. Nayyar, B. Petrini, W. Truong, V. Hellendoorn, N. Chinaev, C. Alberti, W. Wang, J. Hu, V. Mirrokni, A. Balashankar, A. Aharon, A. Mehta, A. Iscen, J. Kready, L. Manning, A. Mohananey, Y. Chen, A. Tripathi, A. Wu, I. Petrovski, D. Hwang, M. Baeuml, S. Chandrakaladharan, Y. Liu, R. Coaguila, M. Chen, S. Ma, P. Tafti, S. Tatineni, T. Spitz, J. Ye, P. Vicol, M. Rosca, A. Puigdomènech, Z. Yahav, S. Ghemawat, H. Lin, P. Kirk, Z. Nabulsi, S. Brin, B. Bohnet, K. Caluwaerts, A. S. Veerubhotla, D. Zheng, Z. Dai, P. Petrov, Y. Xu, R. Mehran, Z. Xu, L. Zintgraf, J. Choi, S. A. Hombaiah, R. Thoppilan, S. Reddi, L. Lew, L. Li, K. Webster, K. Sawhney, L. Lamprou, S. Shakeri, M. Lunayach, J. Chen, S. Bagri, A. Salcianu, Y. Chen, Y. Donchev, C. Magister, S. Nørly, V. Rodrigues, T. Izo, H. Noga, J. Zou, T. Köppe, W. Zhou, K. Lee, X. Long, D. Eisenbud, A. Chen, C. Schenck, C. M. To, P. Zhong, E. Taropa, M. Truong, O. Levy, D. Martins, Z. Zhang, C. Semturs, K. Zhang, A. Yakubovich, P. Moreno, L. McConnaughey, D. Lu, S. Redmond, L. Weerts, Y. Bitton, T. Refice, N. Lacasse, A. Conmy, C. Tallec, J. Odell, H. Forbes-Pollard, A. Socala, J. Hoech, P. Kohli, A. Walton, R. Wang, M. Sazanovich, K. Zhu, A. Kapishnikov, R. Galt, M. Denton, B. Murdoch, C. Sikora, K. Mohamed, W. Wei, U. First, T. McConnell, L. C. Cobo, J. Qin, T. Avrahami, D. Balle, Y. Watanabe, A. Louis, A. Kraft, S. Ariafar, Y. Gu, E. Rives, C. Yoon, A. Rusu, J. Cobon-Kerr, C. Hahn, J. Luo, Yuvein, Zhu, N. Ahuja, R. Benenson, R. L. Kaufman, H. Yu, L. Hightower, J. Zhang, D. Ni, L. A. Hendricks, G. Wang, G. Yona, L. Jain, P. Barrio, S. Bhupatiraju, S. Velusamy, A. Dafoe, S. Riedel, T. Thomas, Z. Yuan, M. Bellaiche, S. Panthaplackel, K. Kloboves, S. Jauhari, C. Akbulut, T. Davchev, E. Gladchenko, D. Madras, A. Chuklin, T. Hill, Q. Yuan, M. Madhavan, L. Leonhard, D. Scandinaro, Q. Chen, N. Niu, A. Douillard, B. Damoc, Y. Onoe, F. Pedregosa, F. Bertsch, C. Leichner, J. Pagadora, J. Malmaud, S. Ponda, A. Twigg, O. Duzhyi, J. Shen, M. Wang, R. Garg, J. Chen, U. Evci, J. Lee, L. Liu, K. Kojima, M. Yamaguchi, A. Rajendran, A. Piergiovanni, V. K. Rajendran, M. Fornoni, G. Ibagon, H. Ragan, S. M. Khan, J. Blitzer, A. Bunner, G. Sun, T. Kosakai, S. Lundberg, N. Elue, K. Guu, S. Park, J. Park, A. Narayanaswamy, C. Wu, J. Mudigonda, T. Cohn, H. Mu, R. Kumar, L. Graesser, Y. Zhang, R. Killam, V. Zhuang, M. Giménez, W. A. Jishi, R. Ley-Wild, A. Zhai, K. Osawa, D. Cedillo, J. Liu, M. Upadhyay, M. Sieniek, R. Sharma, T. Paine, A. Angelova, S. Addepalli, C. Parada, K. Majumder, A. Lamp, S. Kumar, X. Deng, A. Myaskovsky, T. Sabolić, J. Dudek, S. York, F. de Chaumont Quitry, J. Nie, D. Cattle, A. Gunjan, B. Piot, W. Khawaja, S. Bang, S. Wang, S. Khodadadeh, R. R, P. Rawlani, R. Powell, K. Lee, J. Griesser, G. Oh, C. Magalhaes, Y. Li, S. Tokumine, H. N. Vogel, D. Hsu, A. BC, D. Jindal, M. Cohen, Z. Yang, J. Yuan, D. de Cesare, T. Bruguier, J. Xu, M. Roy, A. Jacovi, D. Belov, R. Arya, P. Meadowlark, S. Cohen-Ganor, W. Ye, P. Morris-Suzuki, P. Banzal, G. Song, P. Ponnuramu, F. Zhang, G. Scrivener, S. Zaiem, A. R. Rochman, K. Han, B. Ghazi, K. Lee, S. Drath, D. Suo, A. Girgis, P. Shenoy, D. Nguyen, D. Eck, S. Gupta, L. Yan, J. Carreira, A. Gulati, R. Sang, D. Mirylenka, E. Cooney, E. Chou, M. Ling, C. Fan, B. Coleman, G. Tubone, R. Kumar, J. Baldridge, F. Hernandez-Campos, A. Lazaridou, J. Besley, I. Yona, N. Bulut, Q. Wellens, A. Pierigiovanni, J. George, R. Green, P. Han, C. Tao, G. Clark, C. You, A. Abdolmaleki, J. Fu, T. Chen, A. Chaugule, A. Chandorkar, A. Rahman, W. Thompson, P. Koanantakool, M. Bernico, J. Ren, A. Vlasov, S. Vassilvitskii, M. Kula, Y. Liang, D. Kim, Y. Huang, C. Ye, D. Lepikhin, and W. Helmholz Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. External Links: 2507.06261, Link Cited by: item •, §4.1. de Haan et al. (2019) P. de Haan, D. Jayaraman, and S. Levine Causal confusion in imitation learning. In Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32, p. . External Links: Link Cited by: §A.4. Fang et al. (2026) R. Fang, Y. Liang, X. Wang, J. Wu, S. Qiao, P. Xie, F. Huang, H. Chen, and N. Zhang Memp: exploring agent procedural memory. External Links: 2508.06433, Link Cited by: §1, §5. Feng et al. (2025) L. Feng, Z. Xue, T. Liu, and B. An Group-in-group policy optimization for llm agent training. In Advances in Neural Information Processing Systems, D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen (Eds.), Vol. 38, p. 46375–46408. External Links: Link Cited by: §B.2, §1, §5. Hübotter et al. (2026) J. Hübotter, F. Lübeck, L. Behric, A. Baumann, M. Bagatella, D. Marta, I. Hakimi, I. Shenfeld, T. K. Buening, C. Guestrin, and A. Krause Reinforcement learning via self-distillation. External Links: 2601.20802, Link Cited by: §5. Jin et al. (2025) B. Jin, H. Zeng, Z. Yue, J. Yoon, S. Arik, D. Wang, H. Zamani, and J. Han Search-r1: training llms to reason and leverage search engines with reinforcement learning. External Links: 2503.09516, Link Cited by: item •, §1, §5. Kaelbling et al. (1998) L. P. Kaelbling, M. L. Littman, and A. R. Cassandra Planning and acting in partially observable stochastic domains. Artificial Intelligence 101 (1–2), p. 99–134. Cited by: §2. Li et al. (2026) S. Li, X. Guo, T. Liu, B. Yi, Z. Gong, Z. Liu, H. Chen, and W. Zhang What’s missing in screen-to-action? towards a UI-in-the-loop paradigm for multimodal GUI reasoning. In Findings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, p. 17674–17690. External Links: Link, Document, ISBN 979-8-89176-395-1 Cited by: §5. Liu et al. (2026) Z. Liu, J. Kim, X. Luo, D. Li, and Y. Yang Exploratory memory-augmented llm agent via hybrid on- and off-policy optimization. External Links: 2602.23008, Link Cited by: item •, §1, §4.1, §5, §5. Lu et al. (2026a) Z. Lu, Z. Yao, Z. Han, Z. Wang, J. Wu, Q. Gu, X. Cai, W. Lu, J. Xiao, Y. Zhuang, and Y. Shen Self-distilled agentic reinforcement learning. External Links: 2605.15155, Link Cited by: Table 1, §5. Lu et al. (2026b) Z. Lu, Z. Yao, J. Wu, C. Han, Q. Gu, X. Cai, W. Lu, J. Xiao, Y. Zhuang, and Y. Shen SKILL0: in-context agentic reinforcement learning for skill internalization. External Links: 2604.02268, Link Cited by: §5. Ma et al. (2026) W. Ma, Y. Zeng, Y. Song, X. Cui, J. Zhao, X. Liu, and M. Elhoseiny Freshness-aware prioritized experience replay for llm/vlm reinforcement learning. External Links: 2604.16918, Link Cited by: §5. OpenAI et al. (2024) OpenAI, :, A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, A. Mądry, A. Baker-Whitcomb, A. Beutel, A. Borzunov, A. Carney, A. Chow, A. Kirillov, A. Nichol, A. Paino, A. Renzin, A. T. Passos, A. Kirillov, A. Christakis, A. Conneau, A. Kamali, A. Jabri, A. Moyer, A. Tam, A. Crookes, A. Tootoochian, A. Tootoonchian, A. Kumar, A. Vallone, A. Karpathy, A. Braunstein, A. Cann, A. Codispoti, A. Galu, A. Kondrich, A. Tulloch, A. Mishchenko, A. Baek, A. Jiang, A. Pelisse, A. Woodford, A. Gosalia, A. Dhar, A. Pantuliano, A. Nayak, A. Oliver, B. Zoph, B. Ghorbani, B. Leimberger, B. Rossen, B. Sokolowsky, B. Wang, B. Zweig, B. Hoover, B. Samic, B. McGrew, B. Spero, B. Giertler, B. Cheng, B. Lightcap, B. Walkin, B. Quinn, B. Guarraci, B. Hsu, B. Kellogg, B. Eastman, C. Lugaresi, C. Wainwright, C. Bassin, C. Hudson, C. Chu, C. Nelson, C. Li, C. J. Shern, C. Conger, C. Barette, C. Voss, C. Ding, C. Lu, C. Zhang, C. Beaumont, C. Hallacy, C. Koch, C. Gibson, C. Kim, C. Choi, C. McLeavey, C. Hesse, C. Fischer, C. Winter, C. Czarnecki, C. Jarvis, C. Wei, C. Koumouzelis, D. Sherburn, D. Kappler, D. Levin, D. Levy, D. Carr, D. Farhi, D. Mely, D. Robinson, D. Sasaki, D. Jin, D. Valladares, D. Tsipras, D. Li, D. P. Nguyen, D. Findlay, E. Oiwoh, E. Wong, E. Asdar, E. Proehl, E. Yang, E. Antonow, E. Kramer, E. Peterson, E. Sigler, E. Wallace, E. Brevdo, E. Mays, F. Khorasani, F. P. Such, F. Raso, F. Zhang, F. von Lohmann, F. Sulit, G. Goh, G. Oden, G. Salmon, G. Starace, G. Brockman, H. Salman, H. Bao, H. Hu, H. Wong, H. Wang, H. Schmidt, H. Whitney, H. Jun, H. Kirchner, H. P. de Oliveira Pinto, H. Ren, H. Chang, H. W. Chung, I. Kivlichan, I. O’Connell, I. O’Connell, I. Osband, I. Silber, I. Sohl, I. Okuyucu, I. Lan, I. Kostrikov, I. Sutskever, I. Kanitscheider, I. Gulrajani, J. Coxon, J. Menick, J. Pachocki, J. Aung, J. Betker, J. Crooks, J. Lennon, J. Kiros, J. Leike, J. Park, J. Kwon, J. Phang, J. Teplitz, J. Wei, J. Wolfe, J. Chen, J. Harris, J. Varavva, J. G. Lee, J. Shieh, J. Lin, J. Yu, J. Weng, J. Tang, J. Yu, J. Jang, J. Q. Candela, J. Beutler, J. Landers, J. Parish, J. Heidecke, J. Schulman, J. Lachman, J. McKay, J. Uesato, J. Ward, J. W. Kim, J. Huizinga, J. Sitkin, J. Kraaijeveld, J. Gross, J. Kaplan, J. Snyder, J. Achiam, J. Jiao, J. Lee, J. Zhuang, J. Harriman, K. Fricke, K. Hayashi, K. Singhal, K. Shi, K. Karthik, K. Wood, K. Rimbach, K. Hsu, K. Nguyen, K. Gu-Lemberg, K. Button, K. Liu, K. Howe, K. Muthukumar, K. Luther, L. Ahmad, L. Kai, L. Itow, L. Workman, L. Pathak, L. Chen, L. Jing, L. Guy, L. Fedus, L. Zhou, L. Mamitsuka, L. Weng, L. McCallum, L. Held, L. Ouyang, L. Feuvrier, L. Zhang, L. Kondraciuk, L. Kaiser, L. Hewitt, L. Metz, L. Doshi, M. Aflak, M. Simens, M. Boyd, M. Thompson, M. Dukhan, M. Chen, M. Gray, M. Hudnall, M. Zhang, M. Aljubeh, M. Litwin, M. Zeng, M. Johnson, M. Shetty, M. Gupta, M. Shah, M. Yatbaz, M. J. Yang, M. Zhong, M. Glaese, M. Chen, M. Janner, M. Lampe, M. Petrov, M. Wu, M. Wang, M. Fradin, M. Pokrass, M. Castro, M. O. T. de Castro, M. Pavlov, M. Brundage, M. Wang, M. Khan, M. Murati, M. Bavarian, M. Lin, M. Yesildal, N. Soto, N. Gimelshein, N. Cone, N. Staudacher, N. Summers, N. LaFontaine, N. Chowdhury, N. Ryder, N. Stathas, N. Turley, N. Tezak, N. Felix, N. Kudige, N. Keskar, N. Deutsch, N. Bundick, N. Puckett, O. Nachum, O. Okelola, O. Boiko, O. Murk, O. Jaffe, O. Watkins, O. Godement, O. Campbell-Moore, P. Chao, P. McMillan, P. Belov, P. Su, P. Bak, P. Bakkum, P. Deng, P. Dolan, P. Hoeschele, P. Welinder, P. Tillet, P. Pronin, P. Tillet, P. Dhariwal, Q. Yuan, R. Dias, R. Lim, R. Arora, R. Troll, R. Lin, R. G. Lopes, R. Puri, R. Miyara, R. Leike, R. Gaubert, R. Zamani, R. Wang, R. Donnelly, R. Honsby, R. Smith, R. Sahai, R. Ramchandani, R. Huet, R. Carmichael, R. Zellers, R. Chen, R. Chen, R. Nigmatullin, R. Cheu, S. Jain, S. Altman, S. Schoenholz, S. Toizer, S. Miserendino, S. Agarwal, S. Culver, S. Ethersmith, S. Gray, S. Grove, S. Metzger, S. Hermani, S. Jain, S. Zhao, S. Wu, S. Jomoto, S. Wu, Shuaiqi, Xia, S. Phene, S. Papay, S. Narayanan, S. Coffey, S. Lee, S. Hall, S. Balaji, T. Broda, T. Stramer, T. Xu, T. Gogineni, T. Christianson, T. Sanders, T. Patwardhan, T. Cunninghman, T. Degry, T. Dimson, T. Raoux, T. Shadwell, T. Zheng, T. Underwood, T. Markov, T. Sherbakov, T. Rubin, T. Stasi, T. Kaftan, T. Heywood, T. Peterson, T. Walters, T. Eloundou, V. Qi, V. Moeller, V. Monaco, V. Kuo, V. Fomenko, W. Chang, W. Zheng, W. Zhou, W. Manassra, W. Sheu, W. Zaremba, Y. Patil, Y. Qian, Y. Kim, Y. Cheng, Y. Zhang, Y. He, Y. Zhang, Y. Jin, Y. Dai, and Y. Malkov GPT-4o system card. External Links: 2410.21276, Link Cited by: item •, §B.2, §4.1, §4.1. Ouyang et al. (2026) S. Ouyang, J. Yan, I. Hsu, Y. Chen, K. Jiang, Z. Wang, R. Han, L. T. Le, S. Daruki, X. Tang, V. Tirumalashetty, G. Lee, M. Rofouei, H. Lin, J. Han, C. Lee, and T. Pfister ReasoningBank: scaling agent self-evolving with reasoning memory. External Links: 2509.25140, Link Cited by: §5. Pinto et al. (2017) L. Pinto, M. Andrychowicz, P. Welinder, W. Zaremba, and P. Abbeel Asymmetric actor critic for image-based robot learning. External Links: 1710.06542, Link Cited by: §5. Qin et al. (2025) Y. Qin, X. Tan, Z. He, G. Li, H. Lin, Z. Li, Z. Xu, Y. Shi, S. Cai, R. Rui, S. Cai, Y. Cai, X. Zhang, S. Ye, K. Li, and X. Sun Learn the ropes, then trust the wins: self-imitation with progressive exploration for agentic reinforcement learning. External Links: 2509.22601, Link Cited by: §5. Qwen et al. (2025) Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu Qwen2.5 technical report. External Links: 2412.15115, Link Cited by: §4.1. Ross et al. (2011) S. Ross, G. Gordon, and D. Bagnell A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, G. Gordon, D. Dunson, and M. Dudík (Eds.), Proceedings of Machine Learning Research, Vol. 15, Fort Lauderdale, FL, USA, p. 627–635. External Links: Link Cited by: §A.4. Schulman et al. (2017) J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov Proximal policy optimization algorithms. External Links: 1707.06347, Link Cited by: §5. Shao et al. (2024) Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo DeepSeekMath: pushing the limits of mathematical reasoning in open language models. External Links: 2402.03300, Link Cited by: item •, §1, §2, §4.1, §5. Shinn et al. (2023) N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao Reflexion: language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), External Links: Link Cited by: item •, §1, §1, §4.1, §5. Shridhar et al. (2021) M. Shridhar, X. Yuan, M. Côté, Y. Bisk, A. Trischler, and M. J. Hausknecht ALFWorld: aligning text and embodied environments for interactive learning. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, External Links: Link Cited by: Appendix F, §1, §1, §4.1, §5. Snell et al. (2022) C. Snell, D. Klein, and R. Zhong Learning by distilling context. External Links: 2209.15189, Link Cited by: §5. Vapnik and Vashist (2009) V. Vapnik and A. Vashist A new learning paradigm: learning using privileged information. Neural Networks 22 (5), p. 544–557. Note: Advances in Neural Networks Research: IJCNN2009 External Links: ISSN 0893-6080, Document, Link Cited by: §5. Wang et al. (2025a) H. Wang, C. T. Leong, J. Wang, J. Wang, and W. Li SPA-rl: reinforcing llm agents via stepwise progress attribution. External Links: 2505.20732, Link Cited by: §5. Wang et al. (2026) J. Wang, Q. Yan, Y. Wang, Y. Tian, S. S. Mishra, Z. Xu, M. Gandhi, P. Xu, and L. L. Cheong Reinforcement learning for self-improving agent with skill library. External Links: 2512.17102, Link Cited by: §5. Wang et al. (2025b) Z. Wang, K. Wang, Q. Wang, P. Zhang, L. Li, Z. Yang, X. Jin, K. Yu, M. N. Nguyen, L. Liu, E. Gottlieb, Y. Lu, K. Cho, J. Wu, L. Fei-Fei, L. Wang, Y. Choi, and M. Li RAGEN: understanding self-evolution in llm agents via multi-turn reinforcement learning. External Links: 2504.20073, Link Cited by: §1, §5. Wei et al. (2025) Q. Wei, S. Zeng, C. Li, W. Brown, O. Frunza, W. Deng, A. Schneider, Y. Nevmyvaka, Y. K. Zhao, A. Garcia, and M. Hong Reinforcing multi-turn reasoning in llm agents via turn-level reward design. External Links: 2505.11821, Link Cited by: §5. Wu et al. (2026) R. Wu, X. Wang, J. Mei, P. Cai, D. Fu, C. Yang, L. Wen, X. Yang, Y. Shen, Y. Wang, and B. Shi EvolveR: self-evolving llm agents through an experience-driven lifecycle. External Links: 2510.16079, Link Cited by: item •, §4.1, §5. Xia et al. (2026) P. Xia, J. Chen, H. Wang, J. Liu, K. Zeng, Y. Wang, S. Han, Y. Zhou, X. Zhao, H. Chen, Z. Zheng, C. Xie, and H. Yao SkillRL: evolving agents via recursive skill-augmented reinforcement learning. External Links: 2602.08234, Link Cited by: item •, §1, §3.3, Table 1, §4.1, §5. Xie et al. (2026) C. Xie, R. Pan, X. Wu, Z. Yunfei, J. Fu, T. Gao, and G. Zhou Unlocking exploration in RLVR: uncertainty-aware advantage shaping for deeper reasoning. In Findings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, p. 19057–19076. External Links: Link, Document, ISBN 979-8-89176-395-1 Cited by: §5. Yan et al. (2025) J. Yan, Y. Li, Z. Hu, Z. Wang, G. Cui, X. Qu, Y. Cheng, and Y. Zhang Learning to reason under off-policy guidance. External Links: 2504.14945, Link Cited by: §5. Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu Qwen3 technical report. External Links: 2505.09388, Link Cited by: §4.1. Yang et al. (2024) L. Yang, Z. Yu, T. Zhang, S. Cao, M. Xu, W. Zhang, J. E. Gonzalez, and B. Cui Buffer of thoughts: thought-augmented reasoning with large language models. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), External Links: Link Cited by: §5. Yao et al. (2022) S. Yao, H. Chen, J. Yang, and K. Narasimhan WebShop: towards scalable real-world web interaction with grounded language agents. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), External Links: Link Cited by: Appendix F, §1, §1, §4.1, §5. Yao et al. (2023) S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao ReAct: synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, External Links: Link Cited by: item •, §1, §4.1. Ye et al. (2026) T. Ye, L. Dong, X. Wu, S. Huang, and F. Wei On-policy context distillation for language models. External Links: 2602.12275, Link Cited by: §5. Zhan et al. (2026) R. Zhan, Y. Li, Z. Wang, X. Qu, D. Liu, J. Shao, D. F. Wong, and Y. Cheng ExGRPO: learning to reason from experience. External Links: 2510.02245, Link Cited by: §5. Zhang et al. (2026a) H. Zhang, Q. Long, J. Bao, T. Feng, W. Zhang, H. Yue, and W. Wang MemSkill: learning and evolving memory skills for self-evolving agents. External Links: 2602.02474, Link Cited by: §5. Zhang et al. (2025a) K. Zhang, A. Lv, J. Li, Y. Wang, F. Wang, H. Hu, and R. Yan StepHint: multi-level stepwise hints enhance reinforcement learning to reason. External Links: 2507.02841, Link Cited by: §5. Zhang et al. (2026b) S. Zhang, Y. Xiong, X. Chen, Z. Jia, R. Huang, J. Xu, and J. Zhang RAPO: expanding exploration for llm agents via retrieval-augmented policy optimization. External Links: 2603.03078, Link Cited by: §5. Zhang et al. (2025b) Y. Zhang, M. Li, D. Long, X. Zhang, H. Lin, B. Yang, P. Xie, A. Yang, D. Liu, J. Lin, F. Huang, and J. Zhou Qwen3 embedding: advancing text embedding and reranking through foundation models. External Links: 2506.05176, Link Cited by: §B.2, §4.1. Zhao et al. (2024) A. Zhao, D. Huang, Q. Xu, M. Lin, Y. Liu, and G. Huang ExpeL: LLM agents are experiential learners. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada, M. J. Wooldridge, J. G. Dy, and S. Natarajan (Eds.), p. 19632–19642. External Links: Link, Document Cited by: §1, §5. Zhao et al. (2026) S. Zhao, Z. Xie, M. Liu, J. Huang, G. Pang, F. Chen, and A. Grover Self-distilled reasoner: on-policy self-distillation for large language models. External Links: 2601.18734, Link Cited by: item •, §4.1, §5. Appendix A Theoretical Analysis of EDGE This section provides an analytical justification for the four core components of EDGE, following the pipeline through which a retrieved experience e influences the policy update: its marginal utility is estimated (§A.1), smoothed across training steps (§A.2), used to gate and calibrate the RL advantage (§A.3), and channeled through on-support distillation (§A.4). Throughout, let c S denote the standard context, c=c⊕ec T=c S e the privileged context, πθ _θ the current policy, and R(τ)∈0,1R(τ)∈\0,1\ the binary outcome reward. Teacher and student rollout subsets ,T T,T S each have size K=G/2K=G/2. All expectations and probabilities condition on fixed πθ _θ and e. A.1 Marginal-Gain Estimation and Gating Define the value of privileged information as the expected return gap between the privileged and standard contexts: e(θ)=τ∼πθ(⋅∣c)[R(τ)]−τ∼πθ(⋅∣c)[R(τ)].V_e(θ)\;=\;E_τ _θ(· c T)\! [R(τ) ]\;-\;E_τ _θ(· c S)\! [R(τ) ]. (10) The empirical marginal gain is Δe=1K∑τ∈R(τ)−1K∑τ∈R(τ). _e\;=\; 1K\! _τ T\!R(τ)\;-\; 1K\! _τ S\!R(τ). (11) Under the assumption that the two subsets are conditionally independent and identically distributed up to the presence of e, Δe _e is unbiased for e(θ)V_e(θ) with variance (pT(1−pT)+pS(1−pS))/K(p_T(1-p_T)+p_S(1-p_S))/K, where pTp_T and pSp_S are the respective success probabilities. Decomposing Δe _e as a sum of 2K2K independent terms each bounded in an interval of length 1/K1/K, Hoeffding’s inequality yields Pr(|Δe−e(θ)|≥ϵ)≤ 2exp(−Kϵ2). \! (| _e-V_e(θ)|≥ε )\;≤\;2 \! (-Kε^2 ). (12) EDGE converts this estimate into a binary gate: Me=(Δe>0).M_e\;=\;I( _e>0). (13) The gate serves as a probabilistic risk-control mechanism. By the one-sided form of (12), if an experience is truly harmful with margin γ>0γ>0 (e(θ)≤−γV_e(θ)≤-γ), then Pr(Me=1)≤exp(−Kγ2) (M_e=1)≤ (-Kγ^2); the symmetric bound holds for missed activations when e(θ)≥γV_e(θ)≥γ. When Me=0M_e=0, teacher rollouts are masked from both the RL and distillation losses, and the update falls back to standard GRPO on the student subset. A.2 EMA Smoothing Because Δe _e is high-variance for small K and the true utility e(θ)V_e(θ) drifts as the policy evolves, EDGE tracks each experience’s utility with an exponential moving average: Ue(t)=(1−μ)Ue(t−1)+μΔe(t),μ∈(0,1).U_e^(t)\;=\;(1-μ)\,U_e^(t-1)\;+\;μ\, _e^(t), μ∈(0,1). (14) Under a local-stationarity approximation (successive Δe(t) _e^(t) treated as i.i.d. with variance σe2 _e^2), the steady-state variance is Var(Ue)=μ2−μσe2Var(U_e)= μ2-μ\, _e^2, which is strictly less than σe2 _e^2 for any μ<1μ<1. Smaller μ yields greater smoothing at the cost of slower adaptation to genuine utility shifts. The pruning threshold η thus operates on a smoothed estimate of recent marginal utility rather than a single noisy contrast. A.3 Pooled Advantage Calibration When the gate activates, teacher rollouts enter the RL update and alter the advantage baseline. Let bS,bTb_S,b_T be the student and teacher mean returns, so Δe=bT−bS _e=b_T-b_S. The active set is =∪,Me=1,,Me=0,A\;=\; casesT S T,&M_e=1,\\[2.0pt] T S,&M_e=0, cases (15) with baseline b=meanj∈(Rj)b_A=mean_j (R_j). When Me=1M_e=1, the pooled baseline bpool=12(bS+bT)b_pool= 12(b_S+b_T) shifts the raw advantage numerators (prior to GRPO’s standard-deviation normalization, which rescales uniformly without changing signs) as follows: A~SEDGE A_S EDGE =A~Swithin−12Δe, = A_S^within- 12 _e, (16) A~TEDGE A_T EDGE =A~Twithin+12Δe, = A_T^within+ 12 _e, (17) where the superscript “within” denotes baselines computed from each subset alone. Since Me=1M_e=1 implies Δe>0 _e>0, student advantages are shifted downward and teacher advantages upward: privileged successes receive stronger reinforcement, while unguided successes are tempered when the scaffold demonstrates superior performance. When Me=0M_e=0, the teacher subset is excluded and no shift is applied. A.4 On-Support Reverse-KL Distillation Directly imitating teacher-generated trajectories risks covariate shift (Ross et al., 2011), as the teacher may visit prefixes whose rationale depends on information absent from the student context (causal misidentification; de Haan et al., 2019). EDGE mitigates this by evaluating the teacher only on student-generated prefixes. For a student trajectory τ=(y1,…,ym)∈τ=(y_1,…,y_m) S, let ht=(c,y<t)h_t S=(c S,y_<t) and ht=(c,y<t)h_t T=(c T,y_<t). The on-support distillation loss is ℒ^distill(θ)=Me∑τ∈∑t=1|τ|DKL(πθ(⋅∣ht)∥πsg(θ)(⋅∣ht)), split L_distill(θ)\;=\;\;&M_e\! _τ S _t=1^|τ|\\ &D_KL\! ( _θ(· h_t S)\; \|\; _sg(θ)(· h_t T) ), split (18) where sg(⋅)sg(·) denotes stop-gradient. Unlike forward-KL imitation on teacher rollouts, this objective compares distributions only at prefixes the student has actually reached, avoiding training on teacher-only states. The reverse-KL direction is mode-seeking with respect to the teacher, concentrating student mass on teacher-preferred actions rather than spreading to cover the full teacher distribution, and thereby keeping the update conservative on the student support. The final actor loss combines both pathways: ℒactor(θ)=ℒRL(θ,)+λℒ^distill(θ).L_actor(θ)\;=\;L_RL(θ;A)\;+\;λ\, L_distill(θ). (19) When Me=0M_e=0, both terms reduce to unprivileged-only updates (=A=T S, ℒ^distill=0 L_distill=0). When Me=1M_e=1, the teacher subset raises the RL baseline (§A.3) and the reverse KL transfers privileged behavior into the standard-context policy. Appendix B Implementation Details B.1 Baselines • GPT-4o (OpenAI et al., 2024): A closed-source multimodal LLM from OpenAI, used as a strong proprietary agent baseline with standard prompting. • Gemini-2.5-Pro (Comanici et al., 2025): A closed-source reasoning model from Google DeepMind, serving as another proprietary agent baseline with standard prompting. • ReAct (Yao et al., 2023): A prompting framework that interleaves chain-of-thought reasoning with environment actions, enabling LLMs to plan and act in a synergistic loop. • Reflexion (Shinn et al., 2023): Extends ReAct by appending verbal self-reflection after task failures, allowing the agent to refine its strategy across successive trials without weight updates. • GRPO (Shao et al., 2024): A group-relative policy optimization algorithm that estimates advantages from a group of sampled rollouts, eliminating the need for a separate critic network. • OPSD (Zhao et al., 2026): An on-policy self-distillation method that distills the model’s own high-quality rollouts back into itself to improve reasoning without external supervision. • GRPO+OPSD: A hybrid baseline that combines GRPO’s group-relative advantage estimation with OPSD’s on-policy self-distillation objective. • EvolveR (Wu et al., 2026): An experience-augmented training method that iteratively evolves a retrieval-augmented memory of past trajectories to guide policy learning. • GRPO+Mem0 (Chhikara et al., 2025): Augments GRPO with Mem0, a memory module that stores and retrieves past interaction experiences as additional context during both training and inference. • Search-R1 (Jin et al., 2025): Trains language models with reinforcement learning to interleave reasoning and search, serving as the standard search-agent baseline in our QA evaluation. • SkillRL (Xia et al., 2026): A skill-based RL framework that extracts reusable skills from successful trajectories and conditions policy optimization on retrieved skill demonstrations. • EMPO2 (Liu et al., 2026): An exploratory memory-augmented policy optimization method that leverages curated past experiences to enhance exploration during RL training. These baselines span closed-source LLM agents, prompting-based reasoning frameworks, standard RL algorithms, self-distillation methods, and experience-augmented training approaches, enabling a comprehensive evaluation of EDGE from multiple perspectives. B.2 RL-Training Configuration We implement EDGE on top of the verl-agent framework (Feng et al., 2025) and train the model with the joint optimization objective in Eq. 19. To reduce the computational cost of reverse-KL distillation, we approximate the vocabulary-level loss using only the top-k student tokens with the highest log probabilities. The experience bank is initialized as empty, and each newly inserted experience is assigned an initial utility score of 0. For retrieval, we use the task instruction and initial environment observations as the query, encode experiences with Qwen3-Embedding-0.6B (Zhang et al., 2025b), and rank candidates by cosine similarity. We retrieve a top-m candidate pool and use the highest-scoring experience as the scaffold for the current rollout. After each training step, we update the experience bank through both expansion and pruning. For expansion, if the success rate of a task category falls below the expansion threshold ξ, the reflector contrasts successful and failed trajectories from the same category and synthesizes up to three new experiences. We use GPT-4o (OpenAI et al., 2024) by default; in the EDGE (self) QA variant, the Qwen3 policy itself serves as the reflector, with the remaining framework unchanged. For pruning, we update the EMA utility score of each experience according to Eq. 14 and remove experiences whose utility falls below the pruning threshold η. For both ALFWorld and WebShop, we use the hyperparameters in Table 4. All training and profiling experiments are conducted on 8×8× A100 80GB GPUs. Configuration Value RL-Training Actor learning rate 1e−61e^-6 Maximum prompt length 4096 Maximum response length 512 Training batch size 16 Rollout group size (G) 8 Training mini-batch size 128 Rollout temperature 1.0 Training steps 200 EDGE configuration Distillation weight λ 0.1 Distillation top-k tokens 20 Utility pruning threshold η -0.1 Retrieval pool size (top-m) 6 Expansion success threshold ξ 0.4 EMA momentum μ 0.5 Maximum new experiences per step 3 Table 4: RL Hyperparameters B.3 Computational Cost We profile all methods under the same hardware and training configuration using Qwen2.5-7B-Instruct on ALFWorld. Table 5 reports wall-clock seconds per training step. Rollout time includes experience retrieval and prompt construction, while reflector latency is reported separately. Method Rollout Old Prob. Ref. Prob. Update Reflector GRPO 263.41 14.10 14.23 55.38 – SkillRL 334.13 18.45 18.51 64.74 17.12 EMPO2 318.69 17.37 17.65 63.17 27.56 EDGE 304.56 17.26 17.67 61.45 16.78 Table 5: Per-step wall-clock time (seconds) on ALFWorld with Qwen2.5-7B-Instruct. Summing all components, EDGE requires 417.72 seconds per step, a 20.3% overhead over GRPO, mainly from retrieval, privileged-context forward computation, and reflection. Nevertheless, it is 7.8% faster than SkillRL and 6.0% faster than EMPO2. All methods use the same rollout budget; EDGE partitions the original group into experience-conditioned and experience-free trajectories without additional environment interactions. Case Failure mode Reflected experience Later evidence 1. Clean bowl Navigates to destination before finding object Locate object before placing Retrieved on a related task at step 155; U=0.125U=0.125 2. Pillow on sofa Issues take from source lacking target Verify source before taking Utility rises to 0.5780.578 by step 200 (78 retrievals) 3. Pick2 search Unstructured search before collecting targets Search systematically before storing Retrieved on a related Pick2 task at step 166 4. No-op movement Repeats navigation after arrival Avoid repeated no-op moves Near-threshold utility; recovers to U=0.452U=0.452 by step 200 Table 6: Experience-bank co-evolution examples. Each case is extracted from saved failure trajectories, LLM reflection logs, retrieval logs, and utility traces. Appendix C Further Analysis Figure 5: Sensitivity of validation success rate to distillation weight λ. C.1 Sensitivity to Distillation Weight λ. Figure 5 sweeps the distillation coefficient λ on Qwen2.5-1.5B-Instruct, revealing a clear trade-off between the RL and distillation objectives. At λ=1λ=1, the distillation term dominates the gradient and effectively freezes the policy near its initial performance (∼ 15%), preventing autonomous exploration beyond the scaffold-prescribed behavior. At λ=0.01λ=0.01, the policy recovers RL-driven improvement but internalizes scaffold-induced patterns too slowly, converging roughly 5 points below the best setting. The moderate value λ=0.1λ=0.1 balances both pressures, sustaining steady improvement to approximately 80%—enough distillation to accelerate internalization without suppressing the RL objective’s exploratory signal. We adopt this value for all remaining experiments. C.2 Case Study: Experience Bank Evolution Dynamics This section complements the aggregate gain-tracking curves in Figure 4 with entry-level evidence from ALFWorld training logs (Qwen2.5-7B-Instruct). We trace four representative cases (Table 6), each linking a rollout failure to the experience it produces, its later retrieval, and the resulting change in agent behavior. Cases 1–2 (Figure 6) illustrate the expansion phase, while Cases 3–4 (Figure 7) illustrate late-stage refinement and non-stationary utility. Case 1: From destination-first wandering to object-first execution Failure (step 150). The task is put a clean bowl in shelf. The failed rollout immediately navigates to the destination and loops around empty shelves: go to shelf 3 → go to cabinet 1 → go to shelf 3 → go to cabinet 3 → examine shelf 3. Successful contrast. The agent first searches likely sources, finds a bowl in the fridge, takes and cleans it, then navigates to a shelf: go to fridge 1 → open fridge 1 → take bowl 1 from fridge 1 → clean bowl 1 with sinkbasin 1 → go to shelf 1. Reflected experience. Entry task_943: “First identify where the target bowl is, retrieve it, clean it if the goal requires a clean bowl, and only then move it to a shelf.” Later retrieval. At step 155, retrieved for put a clean bowl in diningtable (similarity 0.8860.886). The rollout no longer starts by visiting the table; it locates, cleans, and places the bowl. Utility: U=0.125U=0.125. Case 2: From hallucinated source to source verification Failure (step 58). The task is put some pillow on sofa. The agent goes to the sofa, observes box, creditcard, keychain, and newspaper—but no pillow. It issues take pillow from sofa 1, receives Nothing happens, and repeats similar invalid source assumptions. Successful contrast. The paired trajectory searches alternative sources, reaches armchair 1, observes pillow 1, takes it, and moves it to the sofa—only issuing take when the observation lists the target. Reflected experience. Entry step_1067: “Only use take <object> from <container> when the object is actually listed at that location; if not observed, do not repeat the same invalid action.” Later retrieval. At step 60, retrieved for put some keychain on ottoman, where the observation lacks the keychain—the same structural error. Utility evolves from 0.0000.000 at creation to 0.5780.578 at step 200 after 78 retrievals. Figure 6: Experience bank evolution: early-stage expansion. Case 1 illustrates destination-first wandering corrected by an object-first scaffold; Case 2 illustrates hallucinated source actions corrected by an observation-gated rule. Case 3: From unstructured Pick2 search to systematic collection Failure (step 165). For find two ladle and put them in drawer, the policy spends actions on destination drawers or repeated navigation before confirming where the two targets are. Unlike early scaffolds, this failure involves managing object count, source search, and final placement simultaneously. Reflected experience. Entry task_995: “Identify likely locations of the target item, inspect those locations methodically, pick up each required object, and then place them into the specified container. Avoid random movement or opening unrelated containers before locating the targets.” Later retrieval. At step 166, retrieved for find two spatula and put them in drawer (similarity 0.8540.854). The rollout locates spatula 3 on a countertop, picks it up, checks other locations, and deposits it in a drawer. Utility reaches U=0.375U=0.375 by steps 195–200. Case 4: Utility tracking separates useful retrieval from harmful repetition Creation (step 144). Entry step_914, “Avoid repeating a no-op movement,” is reflected from a find two tissuebox and put them in drawer failure. The rule: if the agent is already at the target location, it should inspect or act rather than re-issuing the same navigation action. Non-stationary utility. The entry is retrieved frequently but its utility is initially unstable: U=−0.026U=-0.026 at step 150 and −0.085-0.085 at step 165—hovering near the pruning threshold η=−0.1η=-0.1 but remaining above it. Frequency-based retention would treat this entry as important despite its near-zero or negative marginal gain; conversely, a positive threshold would have discarded it prematurely. Recovery. As the policy reaches more states where repeated no-op movements are the dominant error, the entry becomes useful: utility turns positive at U=0.048U=0.048 by step 180 and rises to U=0.452U=0.452 at step 200 after 142 retrievals. This illustrates why EDGE combines smoothed EMA tracking (Eq. (14)) with a mildly negative pruning threshold: experience utility can shift as the policy distribution changes, and premature pruning would discard entries whose value has not yet materialized. Figure 7: Experience bank evolution: late-stage refinement and non-stationary utility. Case 3 illustrates specialized multi-object coordination guidance; Case 4 illustrates an entry whose utility is initially near the pruning threshold but recovers as the policy’s failure distribution shifts. The four cases reveal a natural curriculum driven by the policy’s evolving failure distribution. Early failures are structural and broadly shared: destination-first wandering (Case 1) and hallucinated object presence (Case 2) affect many task variants, producing generic scaffolds with sustained high utility—these entries drive the initial bank growth visible in Figure 4. As the policy internalizes these broad strategies, the remaining failures narrow in scope. By step 165, single-object manipulation is reliable but coordinating two targets remains fragile; Case 3’s search-then-collect scaffold addresses precisely this gap, and its moderate final utility (U=0.375U=0.375) is consistent with Pick2 remaining the hardest subtask in Table 3. Case 4 provides the most direct evidence for non-stationary utility. Entry step_914 (“avoid repeating no-op movements”) hovers near the pruning boundary for roughly 20 steps (U=−0.026U=-0.026 at step 150, −0.085-0.085 at step 165) before recovering to U=0.452U=0.452 by step 200. The mechanism is interpretable: once destination-first errors (Case 1) are resolved, no-op navigation loops become the dominant Pick2 failure mode, re-activating a previously marginal entry. A positive pruning threshold would have discarded it; the adopted η=−0.1η=-0.1 retains such latent-value entries until the policy’s distribution shifts in their favor. Together, these cases explain the aggregate bank trajectory in Figure 4—initial growth as generic scaffolds accumulate, followed by contraction as the policy absorbs broad patterns and the bank converges to a compact set of specialized, currently relevant entries. Appendix D Pseudocode Algorithm 1 outlines the full EDGE training loop, which alternates among three stages per iteration: experience-guided exploration scaffolding (§3.1), gain-gated privileged distillation (§3.2), and experience bank evolution (§3.3). Algorithm 1 EDGE: Experience-Distillation for Guided Exploration 1: 2: Policy πθ _θ; experience bank ℰE; 3: Rollout group size G; distillation weight λ; 4: EMA momentum μ; pruning threshold η; 5: Success-rate threshold ξ. 6: for each training iteration do 7: for each task x in batch do 8: Stage 1: Experience-Guided Exploration Scaffolding 9: Retrieve top experience e∈ℰe by embedding similarity. 10: Sample G trajectories from πθ _θ: 11: T T (G/2G/2): conditioned on c=x⊕ec T=x e 12: T S (G/2G/2): conditioned on c=xc S=x 13: Compute marginal gain Δe _e via Eq. (5). 14: Stage 2: Gain-Gated Privileged Distillation 15: if Δe>0 _e>0 then 16: Compute advantages over all G trajectories. 17: for each τ∈τ S do 18: Forward πsg(θ) _sg(θ) on τ with c T. 19: Compute ℒdistillL_distill via Eq. (18). 20: end for 21: else 22: Compute advantages over T S only. 23: ℒdistill←0L_distill← 0. 24: end if 25: ℒactor←ℒRL+λℒdistillL_actor _RL+λ\,L_distill. 26: Stage 3: Experience Bank Evolution 27: Update EMA utility: Ue(t)←(−μ)Ue(t−1)+μΔeU_e^(t)\!←\!(1\!-\!μ)\,U_e^(t-1)+μ\, _e. 28: end for 29: Update πθ _θ with ℒactorL_actor. 30: Prune experiences with Ue(t)<ηU_e^(t)<η from ℰE. 31: Identify task categories with success rate <ξ<ξ. 32: Generate new experiences via freflect(τ+,τ−)f_reflect(τ^+,τ^-); insert into ℰE. 33: end for 34: Trained policy πθ _θ (deployed without ℰE). Appendix E Prompts This section provides the prompt templates used in our experiments. We include both rollout prompts and experience update prompts for ALFWorld and WebShop. The rollout prompts are used for action generation, while the update prompts are used to generate new state-aware experiences from contrasted trajectories. E.1 ALFWorld Figure 8 shows the ALFWorld rollout prompts. The first prompt provides retrieved experiences as additional context, while the second prompt removes external experiences and asks the agent to act only based on the current task, history, observation, and admissible actions. Prompt: ALFWorld Agent Execution with Experience System Prompt: You are an expert agent operating in the ALFRED Embodied Environment. Your task is to: task_description # Retrieved Relevant Experience retrieved_experiences Warning: These experiences may be outdated. Use them only if they align with your current observation. # Current Progress Prior to this step, you have already taken step_count step(s). Below are the most recent history_length observations and the corresponding actions you took: action_history You are now at step current_step and your current observation is: current_observation Your admissible actions of the current situation are: [admissible_actions]. Now it is your turn to take an action. You should first reason step-by-step about the current situation. This reasoning process MUST be enclosed within <think> </think> tags. Once you have finished your reasoning, you should choose an admissible action for the current step and present it within <action> </action> tags. Prompt: ALFWorld Agent Execution without Experience System Prompt: You are an expert agent operating in the ALFRED Embodied Environment. Your task is to: task_description # Retrieved Relevant Experience No external general and task-specific experiences are provided. Use your learned strategy and current observation. # Current Progress Prior to this step, you have already taken step_count step(s). Below are the most recent history_length observations and the corresponding actions you took: action_history You are now at step current_step and your current observation is: current_observation Your admissible actions of the current situation are: [admissible_actions]. Now it is your turn to take an action. You should first reason step-by-step about the current situation. This reasoning process MUST be enclosed within <think> </think> tags. Once you have finished your reasoning, you should choose an admissible action for the current step and present it within <action> </action> tags. Figure 8: ALFWorld rollout prompts with and without retrieved experience. Figure 9 shows the prompt used to update the ALFWorld experience bank. Given a failed trajectory and a successful reference trajectory for the same task, the model generates compact state-aware experiences in JSON format. Prompt: ALFWorld Experience Update System Prompt: You are an expert updating an ALFWorld state-aware experience bank. You are given one failed trajectory and one successful trajectory for the same task. Analyze the contrasted rollout snippets below and propose NEW or revised experiences that help the agent act better in the same environment state. # Task Information Task: task_text Task Type: task_type The task was successfully/unsuccessfully completed. # Contrasted Rollout Snippets Failed trajectory: failed_text Successful trajectory (reference): success_text # Requirements • Each experience must be state-aware, not a generic tip. • Focus on what the successful no-experience rollout did, and what the retrieved experience may have caused the agent to do incorrectly. • Prefer compact trigger language that can be matched at retrieval time. • Use actionable wording tied to admissible actions and visible state cues. • Return only JSON. Generate 1–max_new_skills_per_update new experiences. # Output Format Example ⬇ "title": "Plan object location before acting", "principle": "For this task type, first identify where the object is, then plan the sequence of actions.", "when_to_apply": "When the task involves finding or moving a specific object" Figure 9: Prompt for updating the ALFWorld state-aware experience bank from contrasted failed and successful trajectories. The model is instructed to generate compact, state-aware experiences that can guide future retrieval and decision making. E.2 WebShop Figure 10 shows the WebShop rollout prompts. The experience-conditioned prompt includes retrieved memories, while the experience-free prompt relies only on the shopping instruction, interaction history, current observation, and admissible actions. Prompt: WebShop Agent Execution with Experience System Prompt: You are an expert autonomous agent operating in the WebShop e-commerce environment. Your task is to: task_description. # Retrieved Relevant Experience retrieved_memories Warning: These experiences may be outdated. Use them only if they align with your current observation. # Current Progress Prior to this step, you have already taken step_count step(s). Below are the most recent history_length observations and the corresponding actions you took: action_history You are now at step current_step and your current observation is: current_observation. Your admissible actions of the current situation are: available_actions Now it is your turn to take one action for the current step. You should first reason step-by-step about the current situation, then think carefully which admissible action best advances the shopping goal. This reasoning process MUST be enclosed within <think> </think> tags. Once you have finished your reasoning, you should choose an admissible action for current step and present it within <action> </action> tags. Prompt: WebShop Agent Execution without Experience System Prompt: You are an expert autonomous agent operating in the WebShop e-commerce environment. Your task is to: task_description. # Retrieved Relevant Experience No external general and task-specific experiences are provided. Use your learned strategy and current observation. # Current Progress Prior to this step, you have already taken step_count step(s). Below are the most recent history_length observations and the corresponding actions you took: action_history You are now at step current_step and your current observation is: current_observation. Your admissible actions of the current situation are: available_actions Now it is your turn to take one action for the current step. You should first reason step-by-step about the current situation, then think carefully which admissible action best advances the shopping goal. This reasoning process MUST be enclosed within <think> </think> tags. Once you have finished your reasoning, you should choose an admissible action for current step and present it within <action> </action> tags. Figure 10: WebShop rollout prompts with and without retrieved experience. Figure 11 shows the prompt used to update the WebShop experience bank. The model compares failed and successful shopping trajectories and produces new state-aware experiences tied to visible state cues and available actions. Prompt: WebShop Experience Update System Prompt: You are an expert updating a WebShop shopping state-aware experience bank. You are given one failed trajectory and one successful trajectory for the same task. Analyze the contrasted rollout snippets below and propose NEW or revised experiences that help the agent act better in the same environment state. # Task Information Task: task_text Task Type: task_type The task was successfully/unsuccessfully completed. # Contrasted Rollout Snippets Failed trajectory: failed_text Successful trajectory (reference): success_text # Requirements • Each experience must be state-aware, not a generic tip. • Focus on what the successful no-experience rollout did, and what the retrieved experience may have caused the agent to do incorrectly. • Prefer compact trigger language that can be matched at retrieval time. • Use actionable wording tied to available actions and visible state cues. • Return only JSON. Generate 1–max_new_skills_per_update new experiences. # Output Format Example ⬇ "title": "Verify Early, Abort Fast", "principle": "On the product page, immediately check category, core attributes, and price; if a key constraint is violated, leave the page at once.", "when_to_apply": "Within the first observation on every product detail page." Figure 11: Prompt for updating the WebShop shopping state-aware experience bank from contrasted failed and successful trajectories. The model is instructed to generate compact, state-aware experiences that can guide future retrieval and shopping decisions. Appendix F Dataset License Our experiments are based on the publicly available ALFWorld (Shridhar et al., 2021) and WebShop (Yao et al., 2022) environments. Training, evaluation, and experience-bank construction are performed using trajectories generated from the task instances and interaction interfaces provided by these environments. We strictly follow the licenses and usage terms of ALFWorld and WebShop, and use the resulting data only for academic research purposes. No private, personally identifiable, or proprietary user data is used in any stage of training, evaluation, or experience construction. Appendix G LLMs Usage Statement We employed a Large Language Model (LLM) to assist exclusively in the editorial stage of manuscript preparation. Its role was limited to refining phrasing, correcting grammar, and enhancing clarity and readability across different sections. The LLM had no involvement in formulating research ideas, designing experiments, or conducting analyses. All scientific contributions and findings are entirely the work of the authors. The authors have ensured that the use of the LLM complies with ethical standards, avoiding plagiarism and scientific misconduct.