Paper deep dive
RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback
Xiaoying Zhang, Zichen Liu, Yipeng Zhang, Xia Hu, Wenqi Shao
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 3/13/2026, 12:51:26 AM
Summary
RetroAgent is an online reinforcement learning framework for LLM-based agents that improves experiential learning through a hindsight self-reflection mechanism. It utilizes dual intrinsic feedback: numerical feedback to reward incremental subtask progress and language feedback stored in a memory buffer, retrieved via a novel Similarity & Utility-Aware Upper Confidence Bound (SimUtil-UCB) strategy. The framework demonstrates state-of-the-art performance across four agentic benchmarks (ALFWorld, WebShop, Sokoban, MineSweeper) by balancing exploration and exploitation.
Entities (7)
Relation Signals (3)
RetroAgent â evaluatedon â ALFWorld
confidence 100% ¡ Extensive experiments across four challenging agentic tasks show that RetroAgent achieves state-of-the-art (SOTA) performance... on ALFWorld
RetroAgent â outperforms â GRPO
confidence 100% ¡ exceeding Group Relative Policy Optimization (GRPO)-trained agents by +18.3% on ALFWorld
RetroAgent â utilizes â SimUtil-UCB
confidence 100% ¡ retrieved via our proposed Similarity & Utility-Aware Upper Confidence Bound (SimUtil-UCB) strategy
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Standard reinforcement learning (RL) for large language model (LLM)-based agents typically optimizes extrinsic task-success rewards, prioritizing one-off task solving over continual adaptation. As a result, agents may converge to suboptimal policies due to limited exploration, and accumulated experience remains implicitly stored in model parameters, hindering efficient experiential learning. Inspired by humans' capacity for retrospective self-improvement, we introduce RetroAgent, an online RL framework that enables agents to master complex interactive environments not only by solving, but also by evolving under the joint guidance of extrinsic task-success rewards and retrospective dual intrinsic feedback. Concretely, RetroAgent features a hindsight self-reflection mechanism that produces: (1) intrinsic numerical feedback, which tracks incremental subtask completion relative to prior attempts to reward promising exploration; and (2) intrinsic language feedback, which distills reusable lessons into a memory buffer retrieved via our proposed Similarity & Utility-Aware Upper Confidence Bound (SimUtil-UCB) strategy, jointly balancing relevance, utility, and exploration. Extensive experiments across four challenging agentic tasks show that RetroAgent achieves state-of-the-art (SOTA) performance, substantially outperforming RL fine-tuning, memory-augmented RL, exploration-guided RL, and meta-RL methods -- e.g., exceeding Group Relative Policy Optimization (GRPO)-trained agents by +18.3% on ALFWorld, +15.4% on WebShop, +27.1% on Sokoban, and +8.9% on MineSweeper -- while maintaining strong test-time adaptation and out-of-distribution generalization.
Tags
Links
- Source: https://arxiv.org/abs/2603.08561v3
- Canonical: https://arxiv.org/abs/2603.08561v3
Trouble viewing inline? Open PDF directly â
Full Text
133,707 characters extracted from source content.
Expand or collapse full text
RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback Xiaoying Zhang1, Zichen Liu2, Yipeng Zhang, Xia Hu1, Wenqi Shao1 1Shanghai AI Lab 2National University of Singapore zhangxycuhk@gmail.com https://github.com/zhangxy-2019/RetroAgent Corresponding author. Abstract Standard reinforcement learning (RL) for large language model (LLM)-based agents typically optimizes extrinsic task-success rewards, prioritizing one-off task solving over continual adaptation. As a result, agents may converge to suboptimal policies due to limited exploration, and accumulated experience remains implicitly stored in model parameters, hindering efficient experiential learning. Inspired by humansâ capacity for retrospective self-improvement, we introduce RetroAgent, an online RL framework that enables agents to master complex interactive environments not only by solving tasks, but also by evolving under the joint guidance of extrinsic task-success rewards and retrospective dual intrinsic feedback. Concretely, RetroAgent features a hindsight self-reflection mechanism that produces: (i)( i) intrinsic numerical feedback, which tracks incremental subtask completion relative to prior attempts to reward promising exploration; and (i)( i) intrinsic language feedback, which distills reusable lessons into a memory buffer retrieved via our proposed Similarity & Utility-Aware Upper Confidence Bound (SimUtil-UCB) strategy, jointly balancing relevance, utility, and exploration. Extensive experiments across four challenging agentic tasks show that RetroAgent achieves state-of-the-art (SOTA) performance, substantially outperforming RL fine-tuning, memory-augmented RL, exploration-guided RL, and meta-RL methodsâe.g., exceeding Group Relative Policy Optimization (GRPO)-trained agents by +18.3%+18.3\% on ALFWorld, +15.4%+15.4\% on WebShop, +27.1%+27.1\% on Sokoban, and +8.9%+8.9\% on MineSweeperâwhile maintaining strong test-time adaptation and out-of-distribution generalization. Figure 1: (a) Overview of the RetroAgent framework. After each episode, the agent analyzes its trajectory via a self-reflection mechanism to produce dual intrinsic feedback, enabling effective learning from past experiences. (b) Initialized from Qwen-2.5-7B-Instruct, RetroAgent substantially outperforms the GRPO-trained baseline (Shao et al., 2024b) and achieves SOTA results across four challenging agentic benchmarks. 1 Introduction Reinforcement learning (RL) (Sutton et al., 1998) has become a central paradigm for training large language model (LLM)-based agents (Comanici et al., 2025; Singh et al., 2025) to master complex interactive environments through direct interaction (Ouyang et al., 2022; Zhang et al., 2022; Liu et al., 2025). Despite recent progress, standard RL training frameworks optimize agents primarily for extrinsic task-success rewards, thereby prioritizing one-off task solving over continuous adaptation (Abel et al., 2023). For example, in embodied benchmarks, training often terminates once the agent discovers a single valid action sequence. This focus introduces two key limitations. First, agents tend to over-exploit and may converge to suboptimal policies rather than exploring diverse alternatives (Kirk et al., 2024). Second, experience is typically stored implicitly in model parameters, making relevant past interactions difficult to retrieve and reuse at decision time (Lin, 1992; Graves et al., 2014). This can slow learning and impair generalization (Goyal et al., 2022). Prior work addresses these limitations along two largely separate directions. One line promotes exploration: Jiang et al. (2025) use meta-RL (Beck et al., 2025) with cross-episode training to optimize long-horizon returns, and Wang et al. (2025b) calibrate rewards using step-wise uncertainty in sparse-reward settings. A second line equips agents with explicit memory, storing either raw interaction histories (Goyal et al., 2022; Wu et al., 2025; Liu et al., 2026b) or distilled skills and lessons (Anthropic, 2025; Wang et al., 2025c; Liu et al., 2026b; Xia et al., 2026). While effective in isolation, these approaches do not tightly couple exploration and memory to support continuous adaptation. We aim to close this gap by drawing inspiration from human retrospective reflection (Lyons and Zelazo, 2011; Liu and van der Schaar, 2025), where individuals evaluate prior actions, compare outcomes across attempts, diagnose success and failure, identify promising directions despite setbacks, and plan improvements. We introduce RetroAgent (Figure 1), an online RL framework that enables agents to master complex interactive environments not merely by solving tasks, but by evolving across episodes under the joint guidance of extrinsic task-success rewards and retrospective dual intrinsic feedback. At its core, RetroAgent features a hindsight self-reflection mechanism: after each episode, the agent analyzes its trajectory to generate dual intrinsic feedback: (i)( i) Intrinsic Numerical Feedback, which measures incremental subtask progress relative to prior attempts (e.g., locating a target item even if purchase fails) and converts it into a scalar reward that reinforces promising exploration and mitigates premature convergence; and (i)( i) Intrinsic Language Feedback, which distills actionable lessons from successes and failures into an explicit memory buffer that can be retrieved to guide subsequent decisions. To retrieve lessons effectively, we propose the Similarity & Utility-Aware Upper Confidence Bound (SimUtil-UCB) strategy, which combines semantic relevance with historical utility and applies the UCB algorithm (Auer et al., 2002) to balance exploiting high-utility lessons and exploring under-used ones. We study two variants of RetroAgent: (i)( i) an in-context self-reflection mechanism, and (i)( i) an RL-trained self-reflection mechanism whose reflective capability is jointly optimized with the decision-making policy. RetroAgent is compatible with a range of RL algorithms; in our implementation, we optimize the decision policy with GRPO (Shao et al., 2024b) and the self-reflection policy with REINFORCE (Williams, 1992). We evaluate both variants using Qwen-2.5-7B-Instruct (Qwen et al., 2025) and Llama-3.1-8B-Instruct (Grattafiori et al., 2024) on four agentic benchmarks: ALFWorld (Shridhar et al., 2021), WebShop (Yao et al., 2022b), Sokoban (Racanière et al., 2017), and MineSweeper (Li et al., 2024). Across all environments, RetroAgent consistently outperforms prior methodsâincluding RL fine-tuning, memory-augmented RL, exploration-guided RL, and meta-RL methodsâimproving SOTA success rates by approximately +10%+10\% on WebShop and +16%+16\% on Sokoban, while exhibiting strong test-time adaptation and out-of-distribution generalization. In summary, our contributions are threefold: (i)( i) We introduce RetroAgent, an online RL framework with a hindsight self-reflection mechanism that enables efficient experiential learning via dual intrinsic feedback. (i)( i) We propose SimUtil-UCB, a retrieval strategy that balances semantic similarity, historical utility, and exploration to effectively leverage accumulated lessons. (i)( i) We evaluate RetroAgent on four challenging agentic benchmarks and show that it substantially outperforms all baselines, achieving SOTA performance in both in-distribution and out-of-distribution settings. 2 Related Work LLMs as Decision-Making Agents. The reasoning capabilities of LLMs have driven their deployment as autonomous decision-making agents. An initial line of research prompts frozen LLMs: ReAct (Yao et al., 2022c), Reflexion (Shinn et al., 2023), and related methods (Park et al., 2023; Wang et al., 2024a) leverage in-context examples, structured prompts, memory retrieval (Wang et al., 2024b), and external tools (Schick et al., 2023; Xie et al., 2024; Zhang et al., 2025a) to tackle complex tasks. However, these approaches are inherently bounded by the capabilities of the underlying foundation model. This ceiling has motivated a second line of work that trains LLM agents directlyâthrough supervised fine-tuning (Tajwar et al., 2025; Xi et al., 2025) or RL (Song et al., 2024; Zhang et al., 2025b; Feng et al., 2025; Jiang et al., 2025)âenabling them to improve from environmental interactions rather than relying on static prompts or handcrafted workflows. Reinforcement Learning for LLM Agents. RL has become a central paradigm for training agents in multi-turn, dynamic environments (Wang et al., 2025d; Putta et al., 2025). ArCHer (Zhou et al., 2024) employs hierarchical value functions for WebShop (Yao et al., 2022a), while LOOP (Chen et al., 2025) integrates PPO (Schulman et al., 2017) with Leave-One-Out advantage estimation for long-horizon tasks in AppWorld (Trivedi et al., 2024). Group-based RL methods have further refined credit assignment: building on GRPO (Shao et al., 2024a), GiGPO (Feng et al., 2025) introduces two-level advantage estimation, while other works investigate turn-level reward shaping (Wei et al., 2025) and stepwise progress attribution (Wang et al., 2025a). Meta-RL (Beck et al., 2025) offers a complementary perspective; notably, LAMER (Jiang et al., 2025) uses cross-episode training to enable active test-time exploration. However, these methods optimize primarily against extrinsic environmental feedback, and recent analyses argue that genuine self-improvement requires intrinsic signals beyond sparse task rewards (Liu and van der Schaar, 2025). Although prior works have explored intrinsic motivation (Gao et al., 2025) or entropy-modulated policies (Wang et al., 2025b), RetroAgent takes a fundamentally different path: a hindsight self-reflection mechanism produces dual intrinsic feedback, shifting the objective from isolated problem-solving toward continuous adaptation. Learning from Experience through Retrospection. A growing body of work moves beyond scalar rewards by leveraging verbal feedback and retrospective memory for agent self-improvement. Early approaches (Shinn et al., 2023; Madaan et al., 2023; Yao et al., 2024) generate natural-language critiques or lessons from interactions, iteratively refining same-task performance via in-context learning. Subsequent work internalizes such feedback into model parameters: Jiang et al. (2025) use reflections to guide cross-episode adaptation within a meta-RL framework, while Zhang et al. (2025c); hĂźbotter2026reinforcementlearningselfdistillation refine failed trajectories into high-quality data for policy optimization through RL or distillation. A complementary direction adopts memory-based architectures (Goyal et al., 2022; Wu et al., 2025; Wang et al., 2025c; Zhang et al., 2026; Zhou et al., 2025; Fang et al., 2026; Liu et al., 2026b) that store trajectories, lessons, or skills (Xia et al., 2026) in a retrieval buffer to assist similar future tasks in context. RetroAgent advances this paradigm along a new axis: the agent reflects on its trajectories to produce both intrinsic numerical rewards that guide exploration and intrinsic language feedback that facilitates exploiting past experiences, with these dual signals jointly driving policy optimization. 3 RetroAgent In this section, we introduce RetroAgent (Figure 2), an online RL training framework that employs a hindsight self-reflection mechanism to foster efficient learning from experiences. We begin with the problem formulation and an overview of the self-reflection mechanism in Section 3.1. Section 3.2 then details our strategy for encouraging exploration via intrinsic numerical feedback. Section 3.3 describes how intrinsic language feedback facilitates the exploitation of past experiences. Finally, Section 3.4 presents the policy optimization objectives for both variants of RetroAgent. Figure 2: Overview of the RetroAgent framework. After each episode, a self-reflection mechanism analyzes the trajectory to produce two forms of intrinsic feedback: (i)( i) Intrinsic Numerical Feedback, which quantifies incremental subtask completion relative to prior attempts, rewarding promising exploratory behaviors that may not yet yield task success; and (i)( i) Intrinsic Language Feedback, which distills actionable lessons from past successes and failures into a memory buffer, retrieved via the proposed SimUtil-UCB strategy to effectively leverage accumulated experiences on similar tasks. 3.1 General Overview Problem Formulation. We model the LLM agentâs multi-turn interaction with its environment as a Markov Decision Process (MDP) (Sutton et al., 1998), defined by âł=(,,P,R,Îł)M=(S,A,P,R,Îł), where S is the state space, A the action space, Pâ(st+1âŁst,at)P(s_t+1 s_t,a_t) the environmentâs transition dynamics, Râ(st,at)R(s_t,a_t) the reward function, and Îłâ[0,1]Îłâ[0,1] the discount factor. At each step t=0,âŚ,Tâ1t=0,âŚ,T-1, the agent observes state stâs_t and samples action atâa_t from its policy Ďθ(â âŁst) _θ(¡ s_t). In the LLM agent setting, the state is the concatenation of all preceding observations and actions: st=(o0,a0,âŚ,atâ1,ot)s_t=(o_0,a_0,âŚ,a_t-1,o_t). Executing ata_t yields reward rt+1=Râ(st,at)r_t+1=R(s_t,a_t) and successor state st+1âźP(â âŁst,at)s_t+1 P(¡ s_t,a_t), producing a trajectory Ď=(s0,a0,r1,âŚ,sTâ1,aTâ1,rT)Ď=(s_0,a_0,r_1,âŚ,s_T-1,a_T-1,r_T). With purely extrinsic rewards rt+1=rt+1extr_t+1=r_t+1^ext, the standard objective is to maximize the expected discounted return: Standardâ(θ)=ĎâźĎθ(â âŁx)ĂPâ[G0]=ĎâźĎθ(â âŁx)ĂPâ[ât=0Tâ1Îłtârt+1ext],J_Standard(θ)=E_Ď _θ(¡ x)Ă P\! [\,G_0\, ]=E_Ď _θ(¡ x)Ă P\! [\, _t=0^T-1Îł^t\,r_t+1^ext ], (1) where x=o0x=o_0 is the task instruction drawn from the training set D, and ĎâźĎθ(â âŁx)ĂPĎ _θ(¡ x)Ă P denotes that trajectories are generated jointly by the policy and the environment dynamics. In practice, extrinsic rewards are sparse: a non-zero terminal reward RextR^ext is provided only when the episode ends, either upon successful task completion or upon exceeding the allowed number of steps. To simplify credit assignment, we redistribute this terminal reward uniformly across all steps, setting rt+1ext=Rextr_t+1^ext=R^ext for every t. RetroAgent augments this objective with intrinsic feedback from a hindsight self-reflection mechanism. An intrinsic reward RintR^int (Section 3.2) is likewise assigned uniformly to every step, yielding the composite objective: RetroAgentâ(θ)=ĎâźÎ θ(â âŁx)ĂPâ[ât=0Tâ1Îłtâ(Rext+Rint)],J_RetroAgent(θ)=E_Ď _θ(¡ x)Ă P [\, _t=0^T-1Îł^t (R^ext+R^int ) ], (2) where Πθ(â âŁx) _θ(¡ x) denotes a mixture distribution over trajectories induced by two policies: the base policy Ďθ(â âŁx) _θ(¡ x) and a memory-augmented policy Ďθ(â âŁfmemory(x,âŹ)) _θ\! (¡ f_memory(x,B) ). Here, fmemoryâ(x,âŹ)f_memory(x,B) is the proposed SimUtil-UCB retrieval strategy (Section 3.3), which retrieves a relevant lesson from the memory buffer âŹB to augment the task instruction x. Self-Reflection Mechanism. At its core, RetroAgent incorporates a hindsight self-reflection mechanism for efficient experiential learning. At the conclusion of each episode, the agent evaluates its trajectory via a reflection function z=freflectâ(Ď)z=f_reflect(Ď), leveraging in-context learning (Wei et al., 2022).111For notational simplicity, we reuse Ď to denote the agentâenvironment interaction history, consisting of interleaved observations and actions. This function produces a reflection tuple z=(Ď(x,Ď),c,m)z=( _(x,Ď),c,m) comprising three components: (i)( i) a scalar potential score Ď(x,Ď)â[0,1] _(x,Ď)â[0,1] estimating the subtask completion rate, from which the intrinsic numerical reward RintR^int is derived (Section 3.2); (i)( i) a binary success prediction câsuccess,failurecâ\success,failure\; and (i)( i) a natural-language retrospective lesson m distilled from the trajectory. The lesson m is stored in a memory buffer âŹB and subsequently retrieved to provide in-context guidance as intrinsic language feedback via fmemoryâ(x,âŹ)f_memory(x,B) (Section 3.3). The central challenge of this mechanism lies in eliciting high-quality intrinsic feedback. To this end, we propose two variants: an in-context variant and an RL-trained variant. In-Context Variant. We employ pairwise induction by augmenting the reflection function with two additional inputs: (i)(i) a binary outcome indicator Iextâsuccess,failureI^extâ\success,failure\, and (i)(i) a contrastive reference trajectory Ďref _ref collected from an earlier training step whose outcome differs from that of the current episode. Contrasting successful and failed trajectories enables the model to more precisely isolate behavioral strengths and deficiencies, yielding higher-quality potential scores and lessons (Lee et al., 2023). The resulting reflection function takes the form z=freflectâ(Ďref,Iext,Ď)z=f_reflect( _ref,\,I^ext,\,Ď). RL-Trained Variant. In this variant, the agent is jointly optimized so that its self-reflection capability co-evolves with its decision-making policy. We introduce a reflection reward RreflectR^reflect that quantifies the accuracy of the agentâs self-assessment: Rreflect:=Rextâ â(c=Iext),R^reflect:=R^ext¡1(c=I^ext), (3) where â(â )1(¡) is the indicator function and c is the success prediction produced by the reflection. Scaling by RextR^ext aligns the magnitude of the reflection reward with that of the extrinsic signal.222Alternative reward-scaling strategies are possible but are left for future work. Let Ďθ _θ denote the reflection policy, which generates the reflection tuple z=(Ď(x,Ď),c,m)z=( _(x,Ď),\,c,\,m) conditioned on the trajectory Ď. The composite training objective generalizes Equation 2 by incorporating a self-reflection term: RetroAgentâ(θ)=ĎâźÎ θ(â âŁx)ĂPâ[ât=0Tâ1Îłtâ(Rext+Rint)]âDecision-Making+Îťreflectâ zâźĎθ(â âŁĎ)â[Rreflect]âSelf-Reflection,J_RetroAgent(θ)= E_Ď _θ(¡ x)Ă P [\, _t=0^T-1Îł^t (R^ext+R^int ) ]_Decision-Making\;+\; _reflect¡E_z _θ(¡ Ď) [\,R^reflect ]_Self-Reflection, (4) where ÎťreflectâĽ0 _reflect⼠0 is a coefficient controlling the relative weight of the self-reflection objective; Equation 2 is recovered when Îťreflect=0 _reflect=0. Prompt templates for both variants are provided in Appendix A, and optimization details are discussed in Section 3.4. 3.2 Encouraging Exploration with Intrinsic Numerical Feedback We now describe how the potential score Ď(x,Ď) _(x,Ď) produced by the self-reflection mechanism is transformed into a shaped intrinsic rewardâthe capability-evolution reward RintR^intâthat captures incremental subtask completion relative to prior attempts, thereby encouraging promising exploratory behaviors that do not yet yield full task success. For each task x, we maintain a historical baseline ÎŚx _x equal to the highest group-mean success rate observed across all prior training iterations, where per-episode success is the environment-provided binary indicator IextI^ext. At training iteration k, the intrinsic reward is the rectified gain of the potential score over this baseline: Rkint:=maxâĄ(0,Ď(x,Ď),kâÎŚx).R^int_k:= \! (0,\; _(x,Ď),k- _x ). (5) The baseline is updated using the mean success rate IÂŻkext I^ext_k of the current group of N rollouts: IÂŻkext=1Nââj=1NIkextâ(j),ÎŚx:=maxâĄ(ÎŚx,IÂŻkext). I^ext_k= 1N _j=1^NI^ext(j)_k, _x:= \! ( _x,\; I^ext_k ). (6) Because the max operator can only raise the threshold, ÎŚx _x is monotonically non-decreasing and anchored in demonstrated performance. Requiring the potential score to exceed this historical best to earn intrinsic reward promotes consistent policy improvement and prevents the optimization from being dominated by isolated, non-replicable successes. For simplicity, we omit the iteration index k in all subsequent formulations. The composite per-trajectory reward is the sum of the extrinsic and intrinsic components: Râ(Ď)=Rextâ(Ď)+Rintâ(Ď)R(Ď)=R^ext(Ď)+R^int(Ď). 3.3 Facilitating Experience Exploitation via Intrinsic Language Feedback While the capability-evolution reward (Section 3.2) encourages promising exploration, it lacks the semantic richness required to guide the agent on how to improve or avoid unpromising behaviors (Xu et al., 2025). To provide such guidance, RetroAgent maintains a retrieval-augmented reflection memory that distills past experiences into actionable textual lessons and injects them into the policyâs context at decision time during subsequent training steps. We describe the memory structure and the retrieval strategy below. Reflection Memory Buffer. We maintain a persistent buffer âŹ=bii=1|âŹ|B=\b_i\_i=1^|B| in which each entry is a tuple bi=(xi,mi,Ďi,ui,ni,di),b_i= (x_i,\;m_i,\; _i,\;u_i,\;n_i,\;d_i ), where xix_i is the task instruction, mim_i the natural-language lesson produced by the self-reflection mechanism (Section 3.1), Ďi _i the trajectory from which the lesson was derived, uiâ[0,1]u_iâ[0,1] a utility score estimating the lessonâs helpfulness for subsequent task completion, niâân_i the number of times the entry has been retrieved, and diâsuccess,failured_iâ\success,\,failure\ the outcome indicator (IextI^ext) of the originating episode. To enable efficient retrieval, every task instruction is mapped to a shared embedding space by a frozen sentence encoder â°E (specifically all-MiniLM-L6-v2333https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2), yielding the embedding i=â°â(xi)v_i=E(x_i). Similarity- and Utility-Aware UCB Memory Retrieval (SimUtil-UCB). Given a current task x, the retrieval procedure selects the top-k most valuable lessons from âŹB by jointly considering three criteria: semantic relevance, which ensures that retrieved lessons pertain to tasks similar to the current one; reflection utility, which favors lessons that have historically contributed to successful outcomes; and exploration coverage, which prevents the agent from repeatedly exploiting a narrow subset of entries while neglecting potentially valuable but under-accessed ones. Criterion 1: Semantic relevance. The relevance between the current task and each stored entry is measured via cosine similarity between their embeddings (Lewis et al., 2020): srelâ(x,xi)=â°â(x)â iââ°â(x)âââiâ.s_rel(x,\,x_i)\;=\; E(x)¡v_i \|E(x) \|\; \|v_i \|. (7) Candidates with srel<0.4s_rel<0.4 are discarded to guarantee a minimum level of contextual relevance. Criterion 2: Reflection utility. Among the surviving candidates, lesson quality is assessed through the utility score uiu_i, which captures how helpful a lesson has historically been for task completion. Each utility score is initialized to 0.50.5 and updated after every episode in which the lesson is retrieved. Concretely, if at training step t the entry bib_i was retrieved and the resulting episode achieved a task success score u^tâ[0,1] u_tâ[0,1], the utility is updated via an exponential moving average (Klinker, 2011): ui:=(1âβutil)âui+βutilâu^t,u_i\;:=(1- _util)\,u_i\;+\; _util\, u_t, (8) where βutilâ(0,1) _utilâ(0,1) is a smoothing coefficient. Criterion 3: Exploration coverage. To balance exploitation of high-utility lessons with exploration of under-accessed ones, we adopt an Upper Confidence Bound (UCB) formulation (Auer et al., 2002) that augments the utility with an exploration bonus: uUCB(i):=ui+ÎşâlnâĄNni,u_UCB^(i)\;:=\;u_i\;+\;Îş\, Nn_i, (9) where N=âjnjN= _jn_j is the total retrieval count across the buffer and Îş>0Îş>0 is a scaling constant controlling the degree of exploration (set to 1.01.0 in our experiments). Combined retrieval score. The final retrieval score integrates semantic relevance and UCB-augmented utility through a convex combination: Sâ(biâŁx):=Îąâsrelâ(x,xi)+(1âÎą)âuUCB(i),S(b_i x)\;:=\;Îą\;s_rel(x,\,x_i)\;+\;(1-Îą)\;u_UCB^(i), (10) where Îąâ[0,1]Îąâ[0,1] governs the trade-off between relevance and utility. The top-k (e.g., k=1k=1) entries ranked by S are selected, and their lessons mi\m_i\ are concatenated with the task prompt to form the memory-augmented input fmemoryâ(x,âŹ)=xâmretrievedf_memory(x,\,B)=x m_retrieved, which is supplied to the policy as Ďθ(â âŁfmemory(x,âŹ)) _θ\! (¡ f_memory(x,\,B) ) (cf. Equation 2). Upon retrieval, the access count of each selected entry is incremented: ni:=ni+1n_i:=n_i+1. 3.4 Policy Optimization with Dual Intrinsic Feedback RetroAgent is compatible with a broad class of RL algorithms. In this work, we instantiate it with GRPO (Shao et al., 2024b), adapted to incorporate dual intrinsic feedback into multi-turn trajectory optimization. We describe the trajectory generation procedure, the decision-making objective, and the optional self-reflection objective in turn. Trajectory Generation with Memory Augmentation. For each task instruction x from D, we generate N trajectories under Πθold(â âŁx)ĂP _ _old(¡ x)Ă P (Equation (2)). The first N/2N/2 are sampled from the base policy, Ď(i)âźĎθold(â âŁx)ĂPĎ^(i) _ _old(¡ x)Ă P, and the remaining N/2N/2 from the memory-augmented policy, Ď(i)âźĎθold(â âŁfmemory(x,âŹ))ĂPĎ^(i) _ _old\! (¡ f_memory(x,B) )Ă P. Each trajectory Ď(i)=(s0(i),a0(i),âŚ,sTiâ1(i),aTiâ1(i))Ď^(i)=(s_0^(i),a_0^(i),âŚ,s_T_i-1^(i),a_T_i-1^(i)) is a stateâaction sequence of length TiT_i. This partition lets the agent exploit past experience via memory retrieval while retaining the capacity for independent exploration, facilitating continuous policy adaptation. Decision-Making Objective. Since both RextR^ext and RintR^int are uniform across time steps (Section 3.1), the discounted return reduces to a trajectory-level scalar G(i)=ât=0Tiâ1Îłtâ(Rext,(i)+Rint,(i))G^(i)= _t=0^T_i-1Îł^t (R^ext,(i)+R^int,(i) ), and every step within a trajectory shares the same group-relative advantage: A^(i)=G(i)âmeanâĄ(G(1),âŚ,G(N))stdâĄ(G(1),âŚ,G(N)). A^(i)= G^(i)-mean\! (\G^(1),âŚ,G^(N)\ )std\! (\G^(1),âŚ,G^(N)\ ). Defining the per-token importance ratio as Ďt,j(i)â(θ)=Ďθâ(at,j(i)âŁst(i),at,<j(i))Ďθoldâ(at,j(i)âŁst(i),at,<j(i)) _t,j^(i)(θ)= _θ(a_t,j^(i) s_t^(i),\,a_t,<j^(i)) _ _old(a_t,j^(i) s_t^(i),\,a_t,<j^(i)), the decision-making objective is formulated as: Decision-Makingâ(θ) _Decision-Making(θ) =xâź,Ď(i)âźÎ θold(â âŁx)ĂP[1Nâi=1N1Tiât=0Tiâ11|at(i)|âj=1|at(i)|(ât,jclip(θ,A^(i)) =E_x ,\,\Ď^(i)\ _ _old(¡ x)Ă P [ 1N _i=1^N 1T_i _t=0^T_i-1 1|a_t^(i)| _j=1^|a_t^(i)| (L_t,j^clip\! (θ,\, A^(i) ) (11) âβDKL[Ďθ(â âŁst(i))âĽĎref(â âŁst(i))])], -β\,D_KL\! [ _θ(¡ s_t^(i))\,\|\, _ref(¡ s_t^(i)) ] ) ], where |at(i)||a_t^(i)| denotes the number of tokens in action at(i)a_t^(i). The clipped surrogate function is defined as ât,jclipâ(θ,A^(i))=minâĄ(Ďt,j(i)â(θ)âA^(i),clipâĄ(Ďt,j(i)â(θ), 1âĎľclip, 1+Ďľclip)âA^(i))L^clip_t,j\! (θ,\, A^(i) )= ( _t,j^(i)(θ)\, A^(i),\;clip\! ( _t,j^(i)(θ),\,1- _clip,\,1+ _clip ) A^(i) ), where Ďľclip _clip bounds the policy update and β controls the KL divergence regularization toward the reference policy Ďref _ref. For the in-context self-reflection variant, the total objective is simply RetroAgentâ(θ)=Decision-Makingâ(θ)J_RetroAgent(θ)=J_Decision-Making(θ). Self-Reflection Objective (for RL-Trained Variant). The RL-trained variant additionally optimizes the reflection policy Ďθ _θ. For each trajectory Ď(i)Ď^(i), Ďθ _θ generates a reflection sequence z(i)=(Ď(x,Ď)(i),c(i),m(i))z^(i)=(Ď^(i)_(x,Ď),\,c^(i),\,m^(i)). The success prediction component c(i)c^(i) is scored by Rreflect,(i)R^reflect,(i) (Equation (3)). We optimize Ďθ _θ using REINFORCE (Williams, 1992): Reflectionâ(θ)=z(i)âźĎθold(â âŁĎ(i))â[1Nââi=1N1|z(i)|ââj=1|z(i)|logâĄĎθâ(zj(i)âŁĎ(i),z<j(i))â Rreflect,(i)],J_Reflection(θ)=E_\z^(i) _ _old(¡ Ď^(i))\\! [ 1N _i=1^N 1|z^(i)| _j=1^|z^(i)| _θ\! (z_j^(i) Ď^(i),\,z_<j^(i) )¡ R^reflect,(i) ], (12) where |z(i)||z^(i)| is the token length of the reflection sequence. 4 Experiments 4.1 Experimental Setup Environments. We evaluate RetroAgent across four distinct agentic tasks: (i)( i) ALFWorld (Shridhar et al., 2021), a text-based embodied environment where agents complete household tasks through navigation and object interaction. We assess both in-distribution (seen rooms) and out-of-distribution (unseen rooms) generalization. (i)( i) Webshop (Yao et al., 2022b), a simulated e-commerce environment requiring agents to navigate a web interface to purchase products matching user specifications. (i)( i) Sokoban (Racanière et al., 2017), a planning-heavy puzzle task where agents must push boxes to target locations. Due to the irreversible nature of pushing actions, errors often render puzzles unsolvable. Complexity is governed by board size and box count; we train on 6Ă66Ă 6 boards with 2 boxes, following Jiang et al. (2025). (iv)( iv) MineSweeper (Li et al., 2024), a logic-based puzzle requiring agents to identify mine locations using numerical clues. Difficulty is controlled by board size and mine density; we train on 6Ă66Ă 6 boards with 3 mines. We report Success Rate across all tasks, supplemented by Task Score for WebShop. Compared Methods. We evaluate RetroAgentgent against four categories of competitive baselines, reporting results averaged over three independent runs: (i)( i) Prompting-based methods: We compare against ReAct (Yao et al., 2022c) and Reflexion (Shinn et al., 2023), the latter of which incorporates an in-context self-reflection mechanism for iterative refinement. (i)( i) RL algorithms: We include REINFORCE Leave-One-Out (RLOO) (Kool et al., 2019; Ahmadian et al., 2024), GRPO (Shao et al., 2024b), and Group-in-Group Policy Optimization (GiGPO) (Feng et al., 2025). GiGPO represents the current state-of-the-art by utilizing anchor-state grouping for fine-grained credit assignment. (i)( i) RL-based frameworks: This category includes memory-augmented methods such as MemRL (Zhang et al., 2026) (which updates a memory bank while keeping the policy frozen), EvolveR (Wu et al., 2025) (which integrates raw trajectories into optimization), and Mem0 (Chhikara et al., 2025)+GRPO and SimpleMem (Liu et al., 2026a)+GRPO, (which incorporate persistent memory into the training process). We also compare against SkillRL (Xia et al., 2026), a hybrid approach (supervised finetuning and RL) that induces actionable skills via a teacher model to guide the studentâs policy optimization, and GRPO with EMPG (Wang et al., 2025b), which utilizes entropy-modulated policy gradients for long-horizon optimization. (iv)( iv) A Meta-RL framework (Beck et al., 2025): We compare against LaMer (Jiang et al., 2025), which leverages a multi-episode structure to foster active exploration and robust adaptation within a meta-learning context. Implementation Details. We evaluate RetroAgent on Qwen-2.5-7B-Instruct (Qwen et al., 2025) and Llama-3.1-8B-Instruct (Grattafiori et al., 2024). Although RetroAgent is generally compatible with various RL algorithms, we adopt GRPO as the default and implement our framework by adapting the open-source Verl training library (Sheng et al., 2024). We employ the task prompts from Feng et al. (2025) to enable decision-making via the ReAct format (Yao et al., 2022c), in which the model generates step-by-step reasoning before its corresponding action. At training time, the agent distills lessons as memories from trajectories on the training set; at test time, the agent leverages these memories for task completion on the test set. Detailed hyperparameter settings and training configurations are provided in Appendix B. 4.2 Main Results Method ALFWorld WebShop Sokoban MineSweeper Success (%) Score (%) Success (%) Success (%) Success (%) Qwen-2.5-7B-Instruct (Zero-Shot) 16.9Âą1.816.9_Âą 1.8 4.5Âą1.84.5_Âą 1.8 0.8Âą0.00.8_Âą 0.0 2.6Âą0.52.6_Âą 0.5 6.5Âą1.66.5_Âą 1.6 Prompting-based Methods ReActâ (Yao et al., 2022c) 31.2 46.2 19.5 3.9 7.0 Reflexionâ (Shinn et al., 2023) 42.7 58.1 28.8 4.3 7.4 Fine-tuning with RL RLOOâ (Kool et al., 2019) 75.5Âą4.675.5_Âą 4.6 80.3Âą3.280.3_Âą 3.2 65.7Âą4.065.7_Âą 4.0 9.9Âą1.69.9_Âą 1.6 32.8Âą4.832.8_Âą 4.8 GRPO (Shao et al., 2024b) 77.3Âą4.377.3_Âą 4.3 75.5Âą3.675.5_Âą 3.6 66.9Âą1.266.9_Âą 1.2 11.2Âą2.511.2_Âą 2.5 39.3Âą2.739.3_Âą 2.7 GiGPOâ (Feng et al., 2025) 90.8Âą1.390.8_Âą 1.3 84.4Âą2.984.4_Âą 2.9 72.8Âą3.272.8_Âą 3.2 21.9Âą2.821.9_Âą 2.8 41.1Âą1.241.1_Âą 1.2 Fine-tuning with RL-based Frameworks MemRLâ (Zhang et al., 2026) 21.4 29.5 9.2 4.2Âą3.24.2_Âą 3.2 7.0Âą1.47.0_Âą 1.4 EvolveRâ (Wu et al., 2025) 43.8 42.5 17.6 6.0Âą3.26.0_Âą 3.2 11.7Âą3.111.7_Âą 3.1 Mem0 (Chhikara et al., 2025)+GRPOâ 54.7 58.1 37.5 â â SimpleMem (Liu et al., 2026a)+GRPOâ 62.5 67.8 46.9 â â SkillRLâ (Xia et al., 2026) w/ Teacher Model 89.9 85.2 72.7 â â GRPO w/ EMPGâ (Wang et al., 2025b) 78.5 81.0 69.3 12.8Âą2.312.8_Âą 2.3 40.1Âą3.640.1_Âą 3.6 Fine-tuning with Meta-RL Frameworks LaMer (Jiang et al., 2025) 82.3Âą3.682.3_Âą 3.6 â 61.7Âą4.761.7_Âą 4.7 14.3Âą1.214.3_Âą 1.2 33.3Âą1.833.3_Âą 1.8 RL Training with Extrinsic and Dual Intrinsic Feedback RetroAgent (In-Context Reflection) 91.7Âą1.291.7_Âą 1.2 87.6Âą2.187.6_Âą 2.1 78.9Âą3.678.9_Âą 3.6 32.6Âą4.632.6_Âą 4.6 47.9Âą2.047.9_Âą 2.0 RetroAgent (RL-Trained Reflection) 95.6Âą2.395.6_Âą 2.3 88.9Âą1.388.9_Âą 1.3 82.3Âą1.682.3_Âą 1.6 38.3Âą3.438.3_Âą 3.4 48.2Âą2.048.2_Âą 2.0 Table 1: Main results across four benchmarks, averaged over three independent runs (mean Âą standard deviation). All improvements are statistically significant with p<0.01p<0.01. Results marked with â are cited from prior work (Xia et al., 2026; Feng et al., 2025; Wang et al., 2025b). Unless otherwise specified, all training frameworks use the GRPO algorithm. âSuccessâ and âScoreâ denote Success Rate and Task Score, respectively. w/ Teacher Model indicates methods that require a teacher model for skill induction (Xia et al., 2026). We present the main evaluation results in Table 1. Our key findings are summarized below: Dual intrinsic feedback effectively facilitates agentic reasoning. As shown in Table 1, RetroAgent consistently achieves SOTA performance across all four benchmarks, outperforming the GRPO baseline by +14.4, +12.0, +21.4, and +8.6 percentage points on ALFWorld, WebShop, Sokoban, and MineSweeper, respectively. Notably, on WebShop, RetroAgent surpasses the most competitive baselines, GiGPO and SkillRL, by approximately +6.1â6.2%, confirming that equipping agents with self-generated signals for assessing progress and distilling reusable knowledge yields a more effective learning paradigm than relying on extrinsic rewards alone. Dual feedback outperforms either form of intrinsic signal in isolation. RetroAgent significantly outperforms existing memory-augmented RL frameworksâincluding MemRL, EvolveR, SimpleMem+GRPO, and SkillRLâacross all environments, indicating that supplementing intrinsic language feedback with numerical signals grounded in capability evolution produces a more effective learning signal than language-based guidance alone. Conversely, RetroAgent substantially outperforms GRPO w/ EMPG, which leverages uncertainty as intrinsic numerical feedback for long-horizon policy optimization, confirming that intrinsic language feedback provides complementary experiential guidance that numerical signals alone cannot capture. Together, these results validate the complementary nature of our dual intrinsic feedback design. Distilled lessons outperform raw trajectories. RetroAgent markedly outperforms EvolveR (e.g., 78.9â82.3% vs. 17.6% success rate on WebShop), which incorporates raw trajectories as in-context demonstrations to guide policy optimization. We attribute this gap to the fact that raw trajectories may contain noise that hinders exploration, whereas the actionable lessons distilled by RetroAgentâs self-reflection mechanism provide cleaner and more transferable guidance for subsequent decision-making. RL-trained self-reflection might further boost performance. Equipping RetroAgent with the RL-trained self-reflection variant, whose reflective capability is jointly refined with the decision-making policy, yields additional gainsâincreasing success rates to 95.6% on ALFWorld, 82.3% on WebShop, and 38.3% on Sokoban. 4.3 Test-Time Adaptation and Generalization (a) Test-time adaptation on WebShop (ID). (b) Test-time adaptation on ALFWorld (OOD). Figure 3: Test-time adaptation in an in-distribution (ID) setting on WebShop and an out-of-distribution (OOD) setting on ALFWorld. Test-Time Adaptation. Following Jiang et al. (2025), we evaluate test-time adaptation using the Discoveryâ@âkDiscovery@k metric (hĂźbotter2026reinforcementlearningselfdistillation), which measures the probability of completing a task within k attempts: Discoveryâ@âk:=Pâ(âi=1krâ(yiâŁx)=1)Discovery@k:=P\! ( _i=1^kr(y_i x)=1 ). Results are presented in Figure 3. Dual intrinsic feedback enables rapid and consistent test-time adaptation. RetroAgent achieves near-perfect discovery rates within three attempts in both in-distribution (WebShop: 82.3%â99.0%82.3\%â 99.0\%) and out-of-distribution (ALFWorld: 92.9%â100.0%92.9\%â 100.0\%) settings, consistently outperforming the Meta-RL baseline LaMer across both. Notably, the margin over LaMer widens with increasing k in OOD environments, indicating that retrospective reasoning scales more favorably with additional attempts. Method Memory Retrieval WebShop Discovery@1 (%) Discovery@2 (%) Discovery@3 (%) GRPO (Baseline) â 66.9Âą1.266.9_Âą 1.2 87.8Âą1.887.8_Âą 1.8 97.1Âą0.597.1_Âą 0.5 RetroAgent (In-Context) Ă 76.8Âą1.676.8_Âą 1.6 91.9Âą1.291.9_Âą 1.2 98.4Âą0.098.4_Âą 0.0 RetroAgent (RL-Trained) Ă 77.1Âą1.677.1_Âą 1.6 91.7Âą1.291.7_Âą 1.2 99.0Âą0.599.0_Âą 0.5 RetroAgent (In-Context) â 78.9Âą3.678.9_Âą 3.6 93.0Âą1.493.0_Âą 1.4 97.9Âą0.597.9_Âą 0.5 RetroAgent (RL-Trained) â 82.3Âą1.682.3_Âą 1.6 93.0Âą0.893.0_Âą 0.8 99.0Âą0.599.0_Âą 0.5 Table 2: Impact of memory retrieval on test-time adaptation. RetroAgent effectively internalizes dual intrinsic feedback during training. Table 2 isolates the contribution of memory retrieval at test time. Removing memory augmentation causes only a marginal drop on Discoveryâ@â1Discovery@1 (e.g., 78.9%â76.8%78.9\%â 76.8\% for RetroAgent with in-context self-reflection) and Discoveryâ@â2Discovery@2, while Discoveryâ@â3Discovery@3 is fully preserved. This indicates that the benefits of dual intrinsic feedback are largely absorbed into the policy weights during training, rather than being contingent on retrieval at inference time. (a) Test-time adaptation using Discoveryâ@âkDiscovery@k on harder instances (trained with 3 mines, evaluated with 4 mines). (b) Generalization across increasing difficulty levels (evaluated with the number of mines ranging from 3 to 5). Figure 4: Robustness to challenging tasks on MineSweeper. Robustness to Challenging Tasks. We assess robustness on MineSweeper by constructing two evaluation scenarios that exceed the training difficulty (Figure 4), following Jiang et al. (2025): (i)( i) increasing the mine count from 3 (training) to 4 to test adaptation to harder instances, and (i)( i) varying the mine count from 3 to 5 to measure degradation under progressively increasing difficulty. RetroAgent shows strong robustness on harder tasks. RetroAgent consistently outperforms all baselines in both scenarios, demonstrating rapid adaptation to harder instances (Figure 4(a)) and graceful degradation under increasing difficulty (Figure 4(b)). 4.4 Analysis of In-Context Self-Reflection (a) Completion scores via single induction. (b) Completion scores via pairwise induction. Figure 5: Accuracy of subtask completion scores generated via single-trajectory (single) vs. pairwise-trajectory (pairwise) induction for Qwen-2.5-7B-Instruct on WebShop. Method Hallucination Rate (%) Estimated Utility Score (%) Failure (â ) Success (â ) Failure Success Low (â ) Med (â-) High (â ) Low (â ) Med (â-) High (â ) Single Induction 8.8 15.1 8.8 78.2 12.9 12.2 75.6 12.2 Pairwise Induction 3.8 11.9 3.1 76.7 20.1 6.2 76.2 17.6 Table 3: Quality of lessons (i.e., memories) generated via single-trajectory vs. pairwise-trajectory induction, as assessed by GPT-4o. Method Augmentation Ratio WebShop Task Score (%) Success Rate (%) GRPO â 75.5Âą3.6 66.9Âą1.2 + Single Induction 100% (Full Group) 81.3Âą2.6 70.3Âą2.1 + Pairwise Induction 100% (Full Group) 82.3Âą1.3 72.9Âą1.6 + Pairwise Induction 050% (Half Group) 82.4Âą2.982.4_Âą 2.9 75.3Âą4.375.3_Âą 4.3 Table 4: Effect of induction method and augmentation ratio on GRPO performance. Augmentation Ratio denotes the fraction of sampled trajectories per prompt that receive memory-augmented generation; the remaining trajectories are sampled without augmentation. The effectiveness of RetroAgent depends on its self-reflection mechanism, which governs both the accuracy of intrinsic numerical feedback (how precisely capability evolution is quantified) and the quality of intrinsic language feedback (how valuable the distilled lessons are). We compare single-trajectory and pairwise-trajectory induction within the in-context self-reflection mechanism. To evaluate intrinsic numerical feedback, we treat subtask completion scores produced by GPT-4o (OpenAI et al., 2024) as oracle values and measure each induction methodâs correlation against them. For intrinsic language feedback, we prompt GPT-4o to assess lesson quality. We further quantify downstream impact in Table 4 by augmenting GRPO with lessons from each method, retrieved by semantic relevance to the task prompt. Additional details are provided in Appendix C. Intrinsic language feedback improves RL policy optimization, with pairwise induction yielding the most accurate self-reflection. Pairwise-trajectory induction produces more accurate intrinsic numerical feedback, as evidenced by its higher correlation with oracle subtask completion scoresâthe red line tracks the dashed oracle line more closely in Figure 5. It also yields higher-quality language feedback, with lower hallucination rates and higher estimated utility scores (Table 3). Consistent with these improvements, GRPO augmented with pairwise-induction lessons outperforms its single-induction counterpart in downstream success rate (72.9%72.9\% vs. 70.3%70.3\%; Table 4). Retaining unaugmented exploration is essential. In Table 4, half-group memory augmentation outperforms full-group augmentation (75.3%75.3\% vs. 72.9%72.9\% in success rate), indicating that applying memory-guided generation to the entire sampling group reduces trajectory diversity and risks premature convergence on suboptimal strategies. 4.5 Impact of Intrinsic Numerical Feedback Method Discounted Returns Reward Type WebShop Task Score (%) Success Rate (%) GRPO (Baseline) â Extrinsic 75.5Âą3.675.5_Âą 3.6 66.9Âą1.266.9_Âą 1.2 GRPO â Extrinsic 84.2Âą0.284.2_Âą 0.2 74.7Âą2.774.7_Âą 2.7 + Progress-Guided Rewards â Extrinsic 84.2Âą1.784.2_Âą 1.7 75.0Âą3.175.0_Âą 3.1 + Capability-Evolution Rewards â Extrinsic & Intrinsic 88.2Âą2.188.2_Âą 2.1 79.7Âą3.179.7_Âą 3.1 Table 5: Impact of discounted returns and intrinsic reward shaping on GRPO. Capability-evolution rewards denote the intrinsic numerical feedback described in Section 3.2. (a) Impact of capability-evolution rewards. (b) Impact of memory-retrieval strategies. Figure 6: Valid-set performance dynamics on WebShop when augmenting GRPO with intrinsic numerical feedback (a) or intrinsic language feedback (b). We investigate the impact of discounted returns and intrinsic reward shaping on GRPO, reporting evaluation results in Table 5 and valid-set performance dynamics in Figure 6(a). As an additional baseline, we compare against progress-guided rewards, which replace the potential score Ď(x,Ď) _(x,Ď) in Equation 5 with the binary environment success score IExtI^Ext, grounding the rectified gain in extrinsic outcomes rather than intrinsic self-assessment. Intrinsic numerical feedback enhances agentic reasoning. Table 5 shows that applying discounted returns to derive trajectory-level advantages improves GRPO by approximately +8.7 percentage points in task score and +7.8 in success rate on WebShop. Adding capability-evolution rewards (Equation 5) further raises the task score and success rate to 88.2% and 79.7%, respectively, with Figure 6(a) confirming that this gain emerges consistently from step 25 onward. Moreover, capability-evolution rewards outperform progress-guided rewards, confirming that potential scores from self-reflection provide richer shaping signals than binary extrinsic outcomes alone. 4.6 Impact of Intrinsic Language Feedback Method Discounted Returns Retrieval Strategy WebShop Performance Task Score (%) Success Rate (%) GRPO (Baseline) â â 75.5Âą3.675.5_Âą 3.6 66.9Âą1.266.9_Âą 1.2 GRPO â â 84.2Âą0.284.2_Âą 0.2 74.7Âą2.774.7_Âą 2.7 + Memory Retrieval â Similarity 79.1Âą7.179.1_Âą 7.1 70.1Âą5.570.1_Âą 5.5 + Memory Retrieval â Similarity & Utility 78.4Âą11.478.4_Âą 11.4 69.5Âą8.769.5_Âą 8.7 + Memory Retrieval â SimUtil-UCB 86.4Âą1.886.4Âą 1.8 78.6Âą1.678.6Âą 1.6 Table 6: Impact of intrinsic language feedback on GRPO using different memory-retrieval strategies. SimUtil-UCB denotes the our proposed memory retrieval strategy (Section 3.3). Having established in Section 4.4 that intrinsic language feedback improves RL policy optimization, we now evaluate the proposed SimUtil-UCB retrieval strategy against two ablated variants: similarity-based retrieval (Criterion 1 only) and similarity & utility-based retrieval (Criteria 1â2, without the exploration bonus). Results are reported in Table 6, with valid-set performance dynamics in Figure 6(b). All experiments use half-group memory augmentation. Balancing semantic relevance, utility, and exploration is critical. As shown in Table 6, although discounted returns alone improve GRPO, augmenting training with memory instances retrieved via either similarity-based or similarity & utility-based strategies leads to performance degradation. This is surprising because similarity-based retrieval does improve performance when combined with standard GRPO without discounted returns (Table 4, Section 4.4), suggesting that discounted returns may amplify low-quality memory-guided exploration behaviors. In contrast, SimUtil-UCB consistently yields improvements (Table 6 and Figure 6(b)), raising the task score to 86.4% and the success rate to 78.6%. By incorporating the exploration bonus (Equation 9), SimUtil-UCB balances the exploitation of high-utility lessons with the coverage of underutilized entries, mitigating over-reliance on semantically similar or high-utility memories that may reinforce suboptimal behaviors. Figure 7 shows the distribution of accumulated retrieval counts across memory instances during training under each strategy, where each instance is initialized with a count of 1 that increments upon retrieval. SimUtil-UCB (Figure 7(c)) distributes access more uniformly, with most instances accessed around 5 times, whereas similarity-based retrieval (Figure 7(a)) concentrates access on a narrow subset, many of which exceed 15 retrievals. This confirms that the UCB exploration bonus effectively diversifies memory usage, contributing to the stronger final performance of SimUtil-UCB. (a) Similarity-based retrieval. (b) Similarity & utility. (c) SimUtil-UCB. Figure 7: Distribution of accumulated memory usage counts across retrieval strategies on WebShop, estimated via kernel density estimation (KDE) (Chen, 2017). Each panel shows how frequently stored memory instances are accessed under a given strategy. 4.7 Analysis of Combining Dual Intrinsic Feedback Method Intrinsic Feedback Self-Reflection Mechanism WebShop Task Score (%) Success Rate (%) GRPO (Baseline) â â 75.5Âą3.675.5_Âą 3.6 66.9Âą1.266.9_Âą 1.2 + Capability-Evolution Rewards Numerical â 88.2Âą2.188.2_Âą 2.1 79.7Âą3.179.7_Âą 3.1 + SimUtil-UCB Memory Retrieval Language â 86.4Âą1.886.4_Âą 1.8 78.6Âą1.678.6_Âą 1.6 RetroAgent (In-Context) Dual Pairwise Induction 87.6Âą2.187.6_Âą 2.1 78.9Âą3.678.9_Âą 3.6 RetroAgent (RL-Trained) Dual Pairwise Induction 87.0Âą1.487.0_Âą 1.4 77.1Âą1.077.1_Âą 1.0 RetroAgent (RL-Trained) Dual Single Induction 88.9Âą1.388.9_Âą 1.3 82.3Âą1.682.3_Âą 1.6 Table 7: Individual and combined effects of intrinsic numerical and language feedback under different self-reflection mechanisms on WebShop. Rows above the dashed line ablate each feedback type in isolation; rows below combine both (Dual). (a) Valid-set performance over the course of training. (b) Reflection accuracy over the course of training, smoothed with exponential moving average (EMA) (Klinker, 2011). Figure 8: In-context vs. RL-trained self-reflection mechanisms in RetroAgent on WebShop. We present results for combining intrinsic numerical and language feedback in Table 7 and compare in-context versus RL-trained reflection mechanisms in Figure 8. Combining dual intrinsic feedback facilitates superior agentic reasoning. As shown in Table 7, RetroAgent achieves notable performance gains (e.g., â+3%â+3\% success rate) by integrating dual intrinsic feedback compared to using either capability-evolution rewards or SimUtil-UCB memory retrieval in isolation. The in-context variant, however, slightly underperforms GRPO with capability-evolution rewards only, suggesting that simultaneous exploration signals from both feedback channels might interfere with each other during action selection. Joint optimization maintains reflection capability and enhances RL training. In Figure 8(b), the reflection accuracy of the in-context variant declines steadily as the policy improves (orange curve), even though extrinsic success signals remain available. In contrast, the RL-trained self-reflection mechanism maintains accuracy throughout training (blue curve). Although accuracy dips slightly before step 75âlikely because decision-making policy improvement temporarily outpaces reflection adaptationâit recovers and increases steadily thereafter. The initial gap relative to the in-context baseline arises because the RL-trained variant uses single induction, which is less informative than pairwise induction (consistent with Section 4.4). We validate the choice of single induction by comparing it against a pairwise variant that conditions on a reference trajectory: z=freflectâ(Ďref,Ď)z=f_reflect( _ref,Ď). Although including Ďref _ref yields the highest reflection accuracy (green curve, Figure 8(b)), it does not improve task performance (Table 7). This discrepancy suggests that contrastive comparison enables the reflector to infer outcomes from relative differences between trajectories rather than developing robust standalone evaluation capability. 4.8 Analysis of Training Efficiency Figure 9: Training time (wall-clock hours) on WebShop. âTime to Match GRPOâ denotes the time required for each RetroAgent variant to reach the peak performance of the GRPO baseline. We evaluate training efficiency by comparing the training time of RetroAgent against the GRPO baseline (Figure 9). Intrinsic feedback significantly accelerates training convergence. Although RetroAgent requires more total training time than the GRPO baseline, it reaches the baselineâs peak performance substantially faster. The in-context variant matches GRPOâs peak at step 65, and the RL-trained variant does so at step 73 (Figure 8(a)), corresponding to training time reductions of 46% and 32%, respectively. The slightly slower convergence of the RL-trained variant is likely due to the additional cost of optimizing the reflection objective. 4.9 Encouraging Exploration with Intrinsic Feedback Both numerical and language-based intrinsic feedback are intended to improve RL performance by guiding exploration: capability-evolution rewards steer the agent toward promising action sequences, while retrieved lessons from past experience discourage previously failed behaviors and reinforce successful ones. We validate this hypothesis by measuring trajectory diversity on the WebShop test set across three configurations: (i)( i) GRPO with capability-evolution rewards (numerical feedback only); (i)( i) GRPO with SimUtil-UCB memory retrieval (language feedback only); and (i)( i) RetroAgent with in-context or RL-trained self-reflection (dual feedback). Diversity is quantified with the Vendi Score (Friedman and Dieng, 2023) over both successful and failed trajectories. Method Intrinsic Feedback Vendi Score (â ) Successful Traj. Failed Traj. Qwen-2.5-7B-Instruct â 0.00* 1.89 GRPO (Baseline) â 1.85 1.71 + Capability-Evolution Rewards Numerical 2.04 1.82 + SimUtil-UCB Memory Retrieval Language 2.13 1.97 RetroAgent (In-Context Self-Reflection) Dual 2.01 1.78 RetroAgent (RL-Trained Self-Reflection) Dual 2.20 1.94 Table 8: Impact of intrinsic feedback on trajectory diversity on WebShop, measured by the Vendi Score (Friedman and Dieng, 2023). A score of 0.00 for Qwen-2.5-7B-Instruct indicates that fewer than two successful trajectories were generated, precluding diversity measurement. Intrinsic feedback encourages valuable exploration. All methods incorporating intrinsic feedback achieve higher Vendi Scores on successful trajectories than the GRPO baseline. The in-context RetroAgent variant, however, exhibits slightly lower diversity than either single-feedback ablation, consistent with the observation that simultaneous exploration signals from the two feedback channels can interfere and reduce the net exploration incentive (Table 7). 4.10 Impact of Relevance-Utility Trade-off on RetroAgent Figure 10: Impact of the relevanceâutility tradeoff coefficient Îą on RetroAgent (in-context self-reflection) in terms of task score and success rate on WebShop. We examine the impact of the relevanceâutility tradeoff in memory retrieval on the RetroAgent equipped with a in-context self-reflection mechanism. Specifically, we vary the coefficient Îą, which governs this tradeoff, ranging from 0.30.3 (emphasizing utility) to 0.70.7 (emphasizing relevance). As shown in Figure 10, the RetroAgent achieves higher task scores and success rates on WebShop when utility is prioritized (Îą=0.3Îą=0.3). This underscores the importance of considering memory utility rather than relying solely on semantic relevance. 4.11 Generalization Across Model Architectures Method ALFWorld WebShop Sokoban MineSweeper Success (%) Task Score (%) Success (%) Success (%) Success (%) GRPO (Baseline) 72.7Âą2.372.7_Âą 2.3 78.0Âą2.378.0_Âą 2.3 67.6Âą2.867.6_Âą 2.8 12.2Âą1.212.2_Âą 1.2 42.4Âą2.542.4_Âą 2.5 LaMer (Jiang et al., 2025) 76.0Âą1.876.0_Âą 1.8 - 70.3Âą3.670.3_Âą 3.6 15.9Âą2.415.9_Âą 2.4 32.0Âą3.432.0_Âą 3.4 RetroAgent (In-Context) 93.1Âą1.593.1_Âą 1.5 87.8Âą1.887.8_Âą 1.8 71.9Âą3.671.9_Âą 3.6 39.1Âą1.339.1_Âą 1.3 52.3Âą1.652.3_Âą 1.6 RetroAgent (RL-Trained) 91.4Âą1.491.4_Âą 1.4 89.5Âą2.189.5_Âą 2.1 80.5Âą2.280.5_Âą 2.2 24.5Âą2.824.5_Âą 2.8 59.9Âą3.259.9_Âą 3.2 Table 9: Performance of RetroAgent on Llama-3.1-8B-Instruct across four agentic benchmarks. All improvements are statistically significant (p<0.01p<0.01). To validate the generalizability of RetroAgent across model architectures, we further evaluate it with Llama-3.1-8B-Instruct (Grattafiori et al., 2024). As shown in Table 9, both variants of RetroAgent consistently outperform GRPO and LaMer by substantial margins across all four tasks. Notably, RetroAgent with RL-trained self-reflection underperforms its in-context counterpart on ALFWorld and Sokoban, which we attribute to interference between the self-reflection and decision-making objectives during joint optimization: the auxiliary reflection loss may impede the primary policy gradient signal, slightly degrading task performance. We leave the exploration of more effective multi-objective balancing strategies to future work. 4.12 Qualitative Analysis Figure 11: Qualitative comparison of RetroAgent (in-context self-reflection) on the WebShop validation set between training step 65 (failed trajectory, left) and training step 150 (successful trajectory, right). For conciseness, only action tokens and their corresponding probabilities are shown at each decision step. We qualitatively examine RetroAgentâs continuous adaptation by analyzing how lessons distilled from past experiences on similar tasks inform decision-making as training progresses. Specifically, we compare a failed trajectory generated by RetroAgent (with in-context self-reflection) at an early training step (step 65) with a successful trajectory produced at a later step (step 150) on the WebShop validation set. As shown in Figure 11, at step 65 RetroAgent selects an incorrect item at decision Step 1 and subsequently fails to choose the required pink variant. Moreover, it exhibits notably lower token-level confidence when selecting the correct category âyouth.â In contrast, at step 150 RetroAgent accurately and confidently selects the appropriate item with the correct attributes by leveraging lessons retrieved from its memory buffer. Complete trajectories are presented in Appendix D. 5 Conclusion We present RetroAgent, an online RL framework that bridges one-off task solving and continuous adaptation. Through a hindsight self-reflection mechanism, RetroAgent generates dual intrinsic feedback: (i)( i) intrinsic numerical feedback that rewards promising exploration by tracking incremental subtask completion, and (i)( i) intrinsic language feedback that distills reusable lessons into a memory buffer. This memory is retrieved via SimUtil-UCB, which balances relevance, utility, and exploration to leverage prior experience effectively. By jointly learning from extrinsic task-success rewards and retrospective dual intrinsic feedback, RetroAgent enables efficient experiential learning. Experiments across four diverse agentic tasks show that RetroAgent consistently achieves SOTA performance while exhibiting strong test-time adaptation and out-of-distribution generalization. These results suggest that dual intrinsic feedback is a promising direction for building continuously adaptive agents. Future work includes developing more effective multi-objective optimization strategies for jointly training self-reflection and decision-making, and extending RetroAgent to multi-agent and open-ended settings. References D. Abel, A. Barreto, B. Van Roy, D. Precup, H. P. van Hasselt, and S. Singh (2023) A definition of continual reinforcement learning. Advances in Neural Information Processing Systems 36, p. 50377â50407. Cited by: §1. A. Ahmadian, C. Cremer, M. GallĂŠ, M. Fadaee, J. Kreutzer, O. Pietquin, A. ĂstĂźn, and S. Hooker (2024) Back to basics: revisiting reinforce-style optimization for learning from human feedback in llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 12248â12267. Cited by: §4.1. Anthropic (2025) Introducing agent skills. Claude Blog. Cited by: §1. P. Auer, N. Cesa-Bianchi, and P. Fischer (2002) Finite-time analysis of the multiarmed bandit problem. Machine learning 47 (2), p. 235â256. Cited by: §1, §3.3. J. Beck, R. Vuorio, E. Zheran Liu, Z. Xiong, L. Zintgraf, C. Finn, and S. Whiteson (2025) A tutorial on meta-reinforcement learning. Foundations and Trends in Machine Learning 18 (2-3), p. 224â384. Cited by: §1, §2, §4.1. K. Chen, M. Cusumano-Towner, B. Huval, A. Petrenko, J. Hamburger, V. Koltun, and P. KrähenbĂźhl (2025) Reinforcement learning for long-horizon interactive llm agents. External Links: 2502.01600, Link Cited by: §2. Y. Chen (2017) A tutorial on kernel density estimation and recent advances. Biostatistics & Epidemiology 1 (1), p. 161â187. Cited by: Figure 7. P. Chhikara, D. Khant, S. Aryan, T. Singh, and D. Yadav (2025) Mem0: building production-ready ai agents with scalable long-term memory. External Links: 2504.19413, Link Cited by: §4.1, Table 1. G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, et al. (2025) Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Cited by: §1. R. Fang, Y. Liang, X. Wang, J. Wu, S. Qiao, P. Xie, F. Huang, H. Chen, and N. Zhang (2026) Memp: exploring agent procedural memory. External Links: 2508.06433, Link Cited by: §2. L. Feng, Z. Xue, T. Liu, and B. An (2025) Group-in-group policy optimization for LLM agent training. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §2, §2, §4.1, §4.1, Table 1, Table 1. D. Friedman and A. B. Dieng (2023) The vendi score: a diversity evaluation metric for machine learning. Transactions on Machine Learning Research. External Links: ISSN 2835-8856 Cited by: §4.9, Table 8. J. Gao, L. Pan, Y. Wang, R. Zhong, C. Lu, Q. Cai, P. Jiang, and X. Zhao (2025) Navigate the unknown: enhancing llm reasoning with intrinsic motivation guided exploration. External Links: 2505.17621, Link Cited by: §2. A. Goyal, A. Friesen, A. Banino, T. Weber, N. R. Ke, A. P. Badia, A. Guez, M. Mirza, P. C. Humphreys, K. Konyushova, et al. (2022) Retrieval-augmented reinforcement learning. In International Conference on Machine Learning, p. 7740â7765. Cited by: §1, §1, §2. A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. GuzmĂĄn, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Ăelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma (2024) The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §1, §4.1, §4.11. A. Graves, G. Wayne, and I. Danihelka (2014) Neural turing machines. arXiv preprint arXiv:1410.5401. Cited by: §1. Y. Jiang, L. Jiang, D. Teney, M. Moor, and M. Brbic (2025) Meta-rl induces exploration in language agents. External Links: 2512.16848, Link Cited by: §1, §2, §2, §2, §4.1, §4.1, §4.3, §4.3, Table 1, Table 9. R. Kirk, I. Mediratta, C. Nalmpantis, J. Luketina, E. Hambro, E. Grefenstette, and R. Raileanu (2024) Understanding the effects of RLHF on LLM generalisation and diversity. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §1. F. Klinker (2011) Exponential moving average versus moving exponential average. Mathematische Semesterberichte 58 (1), p. 97â107. Cited by: §3.3, 8(b). W. Kool, H. van Hoof, and M. Welling (2019) Buy 4 REINFORCE samples, get a baseline for free!. External Links: Link Cited by: §4.1, Table 1. H. Lee, S. Phatale, H. Mansoor, K. R. Lu, T. Mesnard, J. Ferret, C. Bishop, E. Hall, V. Carbune, and A. Rastogi (2023) Rlaif: scaling reinforcement learning from human feedback with ai feedback. Cited by: §3.1. P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. KĂźttler, M. Lewis, W. Yih, T. Rocktäschel, et al. (2020) Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems 33, p. 9459â9474. Cited by: §3.3. Y. Li, H. Wang, and C. Zhang (2024) Assessing logical puzzle solving in large language models: insights from a minesweeper case study. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 59â81. Cited by: §1, §4.1. L. Lin (1992) Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine learning 8 (3), p. 293â321. Cited by: §1. J. Liu, Y. Su, P. Xia, S. Han, Z. Zheng, C. Xie, M. Ding, and H. Yao (2026a) SimpleMem: efficient lifelong memory for llm agents. External Links: 2601.02553, Link Cited by: §4.1, Table 1. T. Liu and M. van der Schaar (2025) Position: truly self-improving agents require intrinsic metacognitive learning. In Forty-second International Conference on Machine Learning Position Paper Track, External Links: Link Cited by: §1, §2. Z. Liu, J. Kim, X. Luo, D. Li, and Y. Yang (2026b) Exploratory memory-augmented LLM agent via hybrid on- and off-policy optimization. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1, §2. Z. Liu, A. Sims, K. Duan, C. Chen, S. Yu, X. Zhou, H. Xu, S. Xiong, B. Liu, C. Tan, et al. (2025) Gem: a gym for agentic llms. arXiv preprint arXiv:2510.01051. Cited by: §1. K. E. Lyons and P. D. Zelazo (2011) Monitoring, metacognition, and executive function: elucidating the role of self-reflection in the development of self-regulation. Advances in child development and behavior 40, p. 379â412. Cited by: §1. A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al. (2023) Self-refine: iterative refinement with self-feedback. Advances in Neural Information Processing Systems 36, p. 46534â46594. Cited by: §2. OpenAI, :, A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, A. MÄ dry, A. Baker-Whitcomb, A. Beutel, A. Borzunov, A. Carney, A. Chow, A. Kirillov, A. Nichol, A. Paino, A. Renzin, A. T. Passos, A. Kirillov, A. Christakis, A. Conneau, A. Kamali, A. Jabri, A. Moyer, A. Tam, A. Crookes, A. Tootoochian, A. Tootoonchian, A. Kumar, A. Vallone, A. Karpathy, A. Braunstein, A. Cann, A. Codispoti, A. Galu, A. Kondrich, A. Tulloch, A. Mishchenko, A. Baek, A. Jiang, A. Pelisse, A. Woodford, A. Gosalia, A. Dhar, A. Pantuliano, A. Nayak, A. Oliver, B. Zoph, B. Ghorbani, B. Leimberger, B. Rossen, B. Sokolowsky, B. Wang, B. Zweig, B. Hoover, B. Samic, B. McGrew, B. Spero, B. Giertler, B. Cheng, B. Lightcap, B. Walkin, B. Quinn, B. Guarraci, B. Hsu, B. Kellogg, B. Eastman, C. Lugaresi, C. Wainwright, C. Bassin, C. Hudson, C. Chu, C. Nelson, C. Li, C. J. Shern, C. Conger, C. Barette, C. Voss, C. Ding, C. Lu, C. Zhang, C. Beaumont, C. Hallacy, C. Koch, C. Gibson, C. Kim, C. Choi, C. McLeavey, C. Hesse, C. Fischer, C. Winter, C. Czarnecki, C. Jarvis, C. Wei, C. Koumouzelis, D. Sherburn, D. Kappler, D. Levin, D. Levy, D. Carr, D. Farhi, D. Mely, D. Robinson, D. Sasaki, D. Jin, D. Valladares, D. Tsipras, D. Li, D. P. Nguyen, D. Findlay, E. Oiwoh, E. Wong, E. Asdar, E. Proehl, E. Yang, E. Antonow, E. Kramer, E. Peterson, E. Sigler, E. Wallace, E. Brevdo, E. Mays, F. Khorasani, F. P. Such, F. Raso, F. Zhang, F. von Lohmann, F. Sulit, G. Goh, G. Oden, G. Salmon, G. Starace, G. Brockman, H. Salman, H. Bao, H. Hu, H. Wong, H. Wang, H. Schmidt, H. Whitney, H. Jun, H. Kirchner, H. P. de Oliveira Pinto, H. Ren, H. Chang, H. W. Chung, I. Kivlichan, I. OâConnell, I. OâConnell, I. Osband, I. Silber, I. Sohl, I. Okuyucu, I. Lan, I. Kostrikov, I. Sutskever, I. Kanitscheider, I. Gulrajani, J. Coxon, J. Menick, J. Pachocki, J. Aung, J. Betker, J. Crooks, J. Lennon, J. Kiros, J. Leike, J. Park, J. Kwon, J. Phang, J. Teplitz, J. Wei, J. Wolfe, J. Chen, J. Harris, J. Varavva, J. G. Lee, J. Shieh, J. Lin, J. Yu, J. Weng, J. Tang, J. Yu, J. Jang, J. Q. Candela, J. Beutler, J. Landers, J. Parish, J. Heidecke, J. Schulman, J. Lachman, J. McKay, J. Uesato, J. Ward, J. W. Kim, J. Huizinga, J. Sitkin, J. Kraaijeveld, J. Gross, J. Kaplan, J. Snyder, J. Achiam, J. Jiao, J. Lee, J. Zhuang, J. Harriman, K. Fricke, K. Hayashi, K. Singhal, K. Shi, K. Karthik, K. Wood, K. Rimbach, K. Hsu, K. Nguyen, K. Gu-Lemberg, K. Button, K. Liu, K. Howe, K. Muthukumar, K. Luther, L. Ahmad, L. Kai, L. Itow, L. Workman, L. Pathak, L. Chen, L. Jing, L. Guy, L. Fedus, L. Zhou, L. Mamitsuka, L. Weng, L. McCallum, L. Held, L. Ouyang, L. Feuvrier, L. Zhang, L. Kondraciuk, L. Kaiser, L. Hewitt, L. Metz, L. Doshi, M. Aflak, M. Simens, M. Boyd, M. Thompson, M. Dukhan, M. Chen, M. Gray, M. Hudnall, M. Zhang, M. Aljubeh, M. Litwin, M. Zeng, M. Johnson, M. Shetty, M. Gupta, M. Shah, M. Yatbaz, M. J. Yang, M. Zhong, M. Glaese, M. Chen, M. Janner, M. Lampe, M. Petrov, M. Wu, M. Wang, M. Fradin, M. Pokrass, M. Castro, M. O. T. de Castro, M. Pavlov, M. Brundage, M. Wang, M. Khan, M. Murati, M. Bavarian, M. Lin, M. Yesildal, N. Soto, N. Gimelshein, N. Cone, N. Staudacher, N. Summers, N. LaFontaine, N. Chowdhury, N. Ryder, N. Stathas, N. Turley, N. Tezak, N. Felix, N. Kudige, N. Keskar, N. Deutsch, N. Bundick, N. Puckett, O. Nachum, O. Okelola, O. Boiko, O. Murk, O. Jaffe, O. Watkins, O. Godement, O. Campbell-Moore, P. Chao, P. McMillan, P. Belov, P. Su, P. Bak, P. Bakkum, P. Deng, P. Dolan, P. Hoeschele, P. Welinder, P. Tillet, P. Pronin, P. Tillet, P. Dhariwal, Q. Yuan, R. Dias, R. Lim, R. Arora, R. Troll, R. Lin, R. G. Lopes, R. Puri, R. Miyara, R. Leike, R. Gaubert, R. Zamani, R. Wang, R. Donnelly, R. Honsby, R. Smith, R. Sahai, R. Ramchandani, R. Huet, R. Carmichael, R. Zellers, R. Chen, R. Chen, R. Nigmatullin, R. Cheu, S. Jain, S. Altman, S. Schoenholz, S. Toizer, S. Miserendino, S. Agarwal, S. Culver, S. Ethersmith, S. Gray, S. Grove, S. Metzger, S. Hermani, S. Jain, S. Zhao, S. Wu, S. Jomoto, S. Wu, Shuaiqi, Xia, S. Phene, S. Papay, S. Narayanan, S. Coffey, S. Lee, S. Hall, S. Balaji, T. Broda, T. Stramer, T. Xu, T. Gogineni, T. Christianson, T. Sanders, T. Patwardhan, T. Cunninghman, T. Degry, T. Dimson, T. Raoux, T. Shadwell, T. Zheng, T. Underwood, T. Markov, T. Sherbakov, T. Rubin, T. Stasi, T. Kaftan, T. Heywood, T. Peterson, T. Walters, T. Eloundou, V. Qi, V. Moeller, V. Monaco, V. Kuo, V. Fomenko, W. Chang, W. Zheng, W. Zhou, W. Manassra, W. Sheu, W. Zaremba, Y. Patil, Y. Qian, Y. Kim, Y. Cheng, Y. Zhang, Y. He, Y. Zhang, Y. Jin, Y. Dai, and Y. Malkov (2024) GPT-4o system card. External Links: 2410.21276, Link Cited by: §A.3, §4.4. L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe (2022) Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), External Links: Link Cited by: §1. J. S. Park, J. OâBrien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein (2023) Generative agents: interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, p. 1â22. Cited by: §2. P. Putta, E. Mills, N. Garg, S. R. Motwani, E. S. Markowitz, J. Kiseleva, C. Finn, D. Garg, and R. Rafailov (2025) Agent q: advanced reasoning and learning for autonomous AI agents. External Links: Link Cited by: §2. Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu (2025) Qwen2.5 technical report. External Links: 2412.15115, Link Cited by: §1, §4.1. S. Racanière, T. Weber, D. Reichert, L. Buesing, A. Guez, D. Jimenez Rezende, A. Puigdomènech Badia, O. Vinyals, N. Heess, Y. Li, et al. (2017) Imagination-augmented agents for deep reinforcement learning. Advances in neural information processing systems 30. Cited by: §1, §4.1. T. Schick, J. Dwivedi-Yu, R. DessĂŹ, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom (2023) Toolformer: language models can teach themselves to use tools. Advances in Neural Information Processing Systems 36, p. 68539â68551. Cited by: §2. J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017) Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347. Cited by: §2. Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo (2024a) DeepSeekMath: pushing the limits of mathematical reasoning in open language models. External Links: 2402.03300, Link Cited by: §2. Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al. (2024b) Deepseekmath: pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300. Cited by: Figure 1, §1, §3.4, §4.1, Table 1. G. Sheng, C. Zhang, Z. Ye, X. Wu, W. Zhang, R. Zhang, Y. Peng, H. Lin, and C. Wu (2024) HybridFlow: a flexible and efficient rlhf framework. arXiv preprint arXiv: 2409.19256. Cited by: §4.1. N. Shinn, F. Cassano, A. Gopinath, K. R. Narasimhan, and S. Yao (2023) Reflexion: language agents with verbal reinforcement learning. In Thirty-seventh Conference on Neural Information Processing Systems, External Links: Link Cited by: §2, §2, §4.1, Table 1. M. Shridhar, X. Yuan, M. CĂ´tĂŠ, Y. Bisk, A. Trischler, and M. J. Hausknecht (2021) ALFWorld: aligning text and embodied environments for interactive learning. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, External Links: Link Cited by: §1, §4.1. A. Singh, A. Fry, A. Perelman, A. Tart, A. Ganesh, A. El-Kishky, A. McLaughlin, A. Low, A. Ostrow, A. Ananthram, et al. (2025) Openai gpt-5 system card. arXiv preprint arXiv:2601.03267. Cited by: §1. Y. Song, D. Yin, X. Yue, J. Huang, S. Li, and B. Y. Lin (2024) Trial and error: exploration-based trajectory optimization of LLM agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 7584â7600. External Links: Link, Document Cited by: §2. R. S. Sutton, A. G. Barto, et al. (1998) Reinforcement learning: an introduction. Vol. 1, MIT press Cambridge. Cited by: §1, §3.1. F. Tajwar, Y. Jiang, A. Thankaraj, S. S. Rahman, J. Z. Kolter, J. Schneider, and R. Salakhutdinov (2025) Training a generally curious agent. In Forty-second International Conference on Machine Learning, External Links: Link Cited by: §2. H. Trivedi, T. Khot, M. Hartmann, R. Manku, V. Dong, E. Li, S. Gupta, A. Sabharwal, and N. Balasubramanian (2024) Appworld: a controllable world of apps and people for benchmarking interactive coding agents. arXiv preprint arXiv:2407.18901. Cited by: §2. G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar (2024a) Voyager: an open-ended embodied agent with large language models. Transactions on Machine Learning Research. Note: External Links: ISSN 2835-8856, Link Cited by: §2. H. Wang, C. T. Leong, J. Wang, J. Wang, and W. Li (2025a) SPA-rl: reinforcing llm agents via stepwise progress attribution. External Links: 2505.20732, Link Cited by: §2. J. Wang, J. Liu, Y. Fu, Y. Li, X. Wang, Y. Lin, Y. Yue, L. Zhang, Y. Wang, and K. Wang (2025b) Harnessing uncertainty: entropy-modulated policy gradients for long-horizon llm agents. External Links: 2509.09265, Link Cited by: §1, §2, §4.1, Table 1, Table 1. J. Wang, H. Xu, H. Jia, X. Zhang, M. Yan, W. Shen, J. Zhang, F. Huang, and J. Sang (2024b) Mobile-agent-v2: mobile device operation assistant with effective navigation via multi-agent collaboration. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §2. S. Wang, Y. Wu, and Z. Xu (2025c) Cogito, ergo ludo: an agent that learns to play by reasoning and planning. External Links: 2509.25052, Link Cited by: §1, §2. Z. Wang, K. Wang, Q. Wang, P. Zhang, L. Li, Z. Yang, X. Jin, K. Yu, M. N. Nguyen, L. Liu, et al. (2025d) Ragen: understanding self-evolution in llm agents via multi-turn reinforcement learning. arXiv preprint arXiv:2504.20073. Cited by: §2. J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus (2022) Emergent abilities of large language models. Trans. Mach. Learn. Res. 2022. External Links: Link Cited by: §3.1. Q. Wei, S. Zeng, C. Li, W. Brown, O. Frunza, W. Deng, A. Schneider, Y. Nevmyvaka, Y. K. Zhao, A. Garcia, and M. Hong (2025) Reinforcing multi-turn reasoning in llm agents via turn-level reward design. External Links: 2505.11821, Link Cited by: §2. R. J. Williams (1992) Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 8 (3), p. 229â256. Cited by: §1, §3.4. R. Wu, X. Wang, J. Mei, P. Cai, D. Fu, C. Yang, L. Wen, X. Yang, Y. Shen, Y. Wang, and B. Shi (2025) EvolveR: self-evolving llm agents through an experience-driven lifecycle. External Links: 2510.16079, Link Cited by: §1, §2, §4.1, Table 1. Z. Xi, Y. Ding, W. Chen, B. Hong, H. Guo, J. Wang, X. Guo, D. Yang, C. Liao, W. He, S. Gao, L. Chen, R. Zheng, Y. Zou, T. Gui, Q. Zhang, X. Qiu, X. Huang, Z. Wu, and Y. Jiang (2025) AgentGym: evaluating and training large language model-based agents across diverse environments. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 27914â27961. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §2. P. Xia, J. Chen, H. Wang, J. Liu, K. Zeng, Y. Wang, S. Han, Y. Zhou, X. Zhao, H. Chen, Z. Zheng, C. Xie, and H. Yao (2026) SkillRL: evolving agents via recursive skill-augmented reinforcement learning. External Links: 2602.08234, Link Cited by: §1, §2, §4.1, Table 1, Table 1. T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, et al. (2024) Osworld: benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems 37, p. 52040â52094. Cited by: §2. W. Xu, A. Nie, R. Zheng, A. Modi, A. Swaminathan, and C. Cheng (2025) Provably learning from language feedback. External Links: 2506.10341, Link Cited by: §3.3. S. Yao, H. Chen, J. Yang, and K. R. Narasimhan (2022a) WebShop: towards scalable real-world web interaction with grounded language agents. In Advances in Neural Information Processing Systems, A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho (Eds.), External Links: Link Cited by: §2. S. Yao, H. Chen, J. Yang, and K. Narasimhan (2022b) WebShop: towards scalable real-world web interaction with grounded language agents. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), External Links: Link Cited by: §1, §4.1. S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao (2022c) React: synergizing reasoning and acting in language models. In The eleventh international conference on learning representations, Cited by: §2, §4.1, §4.1, Table 1. W. Yao, S. Heinecke, J. C. Niebles, Z. Liu, Y. Feng, L. Xue, R. R. N, Z. Chen, J. Zhang, D. Arpit, R. Xu, P. L. Mui, H. Wang, C. Xiong, and S. Savarese (2024) Retroformer: retrospective large language agents with policy gradient optimization. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §2. C. Zhang, L. Li, S. He, X. Zhang, B. Qiao, S. Qin, M. Ma, Y. Kang, Q. Lin, S. Rajmohan, et al. (2025a) Ufo: a ui-focused agent for windows os interaction. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 597â622. Cited by: §2. H. Zhang, X. Liu, B. Lv, X. Sun, B. Jing, I. L. Iong, Z. Hou, Z. Qi, H. Lai, Y. Xu, R. Lu, H. Wang, J. Tang, and Y. Dong (2025b) AgentRL: scaling agentic reinforcement learning with a multi-turn, multi-task framework. External Links: 2510.04206, Link Cited by: §2. S. Zhang, J. Wang, R. Zhou, J. Liao, Y. Feng, Z. Li, Y. Zheng, W. Zhang, Y. Wen, Z. Li, F. Xiong, Y. Qi, B. Tang, and M. Wen (2026) MemRL: self-evolving agents via runtime reinforcement learning on episodic memory. External Links: 2601.03192, Link Cited by: §2, §4.1, Table 1. X. Zhang, B. Peng, J. Gao, and H. Meng (2022) Toward self-learning end-to-end task-oriented dialog systems. In Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue, O. Lemon, D. Hakkani-Tur, J. J. Li, A. Ashrafzadeh, D. H. Garcia, M. Alikhani, D. Vandyke, and O. DuĹĄek (Eds.), Edinburgh, UK, p. 516â530. External Links: Link, Document Cited by: §1. X. Zhang, Y. Zhang, H. Sun, K. Feng, C. Lu, C. Yang, and H. Meng (2025c) Critique-grpo: advancing llm reasoning with natural language and numerical feedback. arXiv preprint arXiv:2506.03106. Cited by: §2. H. Zhou, Y. Chen, S. Guo, X. Yan, K. H. Lee, Z. Wang, K. Y. Lee, G. Zhang, K. Shao, L. Yang, and J. Wang (2025) Memento: fine-tuning llm agents without fine-tuning llms. External Links: 2508.16153, Link Cited by: §2. Y. Zhou, A. Zanette, J. Pan, S. Levine, and A. Kumar (2024) ArCHer: training language model agents via hierarchical multi-turn RL. In Forty-first International Conference on Machine Learning, External Links: Link Cited by: §2. Appendix A Task Prompts A.1 Prompt Templates for In-Context Self-Reflection Prompt for Single Induction on Webshop You are an expert evaluating a WebShop shopping attempt. Your task is to: task_description You have just completed an attempt at this shopping task. The task was success completed. Trajectory of the attempt: current_trajectory <think> Given the task outcome, analyze the trajectory to understand: 1. What subtasks were attempted? (search, filter, select, purchase) 2. Which subtasks succeeded vs failed based on the observations? 3. What specific actions or decisions led to this outcome? 4. What are the 1-2 most valuable lessons from this attempt? </think> Output your evaluation as JSON: "subtasks": [ "name": "search_product", "description": "[describe actual search]", "status": "[completed or incomplete]", "name": "apply_filters", "description": "[describe filters used]", "status": "[completed or incomplete]", "name": "select_item", "description": "[describe selection]", "status": "[completed or incomplete]", "name": "complete_purchase", "description": "[describe purchase]", "status": "[completed or incomplete]" ], "task_success": [true if successfully completed, false if unsuccessfully completed], "action_lesson": "[key action insight, e.g., âPrecise search with brand+model found exact matchâ OR âGeneric search missed required featuresâ]", "navigation_lesson": "[navigation insight, e.g., âEfficient use of filters saved timeâ OR âFailed to check additional pages with better optionsâ]" EVALUATION GUIDELINES: ⢠The task outcome has been provided - use it to set task_success accordingly ⢠Focus on WHY the attempt had this outcome: â If successful: What strategies worked well? â If unsuccessful: What went wrong and where? ⢠Each subtask status must reflect actual trajectory events ⢠Lessons should explain factors that led to the outcome ⢠Reference specific elements from trajectory (item IDs, pages, search terms) ⢠Use null for lessons only if truly not applicable Output ONLY the JSON evaluation. Prompt for Pairwise Induction on Webshop You are an expert evaluating a WebShop shopping attempt. Your task is to: task_description You have just completed an attempt at this shopping task. The task was success completed. reference_trajectory Trajectory of the attempt: current_trajectory <think> If a reference trajectory exists, compare it with the current trajectory. Given the task outcome, analyze the trajectory to understand: 1. What subtasks were attempted? (search, filter, select, purchase) 2. Which subtasks succeeded vs failed based on the observations? 3. What specific actions or decisions led to this outcome? 4. What are the 1-2 most valuable lessons from this attempt? </think> Output your evaluation as JSON: "subtasks": [ "name": "search_product", "description": "[describe actual search]", "status": "[completed or incomplete]", "name": "apply_filters", "description": "[describe filters used]", "status": "[completed or incomplete]", "name": "select_item", "description": "[describe selection]", "status": "[completed or incomplete]", "name": "complete_purchase", "description": "[describe purchase]", "status": "[completed or incomplete]" ], "task_success": [true if successfully completed, false if unsuccessfully completed], "action_lesson": "[key action insight, e.g., âPrecise search with brand+model found exact matchâ OR âGeneric search missed required featuresâ]", "navigation_lesson": "[navigation insight, e.g., âEfficient use of filters saved timeâ OR âFailed to check additional pages with better optionsâ]" EVALUATION GUIDELINES: ⢠The task outcome has been provided - use it to set task_success accordingly ⢠Focus on WHY the attempt had this outcome: â If successful: What strategies worked well? â If unsuccessful: What went wrong and where? ⢠Each subtask status must reflect actual trajectory events ⢠Lessons should explain factors that led to the outcome ⢠Reference specific elements from trajectory (item IDs, pages, search terms) ⢠Use null for lessons only if truly not applicable Output ONLY the JSON evaluation. Prompt for Pairwise Induction on Alfworld You are an expert evaluating an ALFRED Embodied Environment task attempt. Your task is to: task_description You have just completed an attempt at this task. The task was success completed. reference_trajectory Trajectory of the attempt: current_trajectory <think> If a reference trajectory exists, compare it with the current trajectory. Given the task outcome, analyze the trajectory to understand: 1. What subtasks were attempted? (pick up, navigate, use appliance, place object) 2. Which subtasks succeeded vs failed based on the observations? 3. What specific actions or decisions led to this outcome? 4. What is the most valuable lesson from this attempt? </think> Output your evaluation as JSON: "subtasks": [ "name": "pick_up_object", "description": "[describe pickup action, e.g., âPick up mug from countertopâ]", "status": "[completed or incomplete]", "name": "navigate_to_location", "description": "[describe navigation, e.g., âGo to microwave 1â]", "status": "[completed or incomplete]", "name": "use_appliance", "description": "[describe appliance use, e.g., âHeat mug in microwaveâ]", "status": "[completed or incomplete]", "name": "place_object", "description": "[describe placement, e.g., âPlace heated mug in cabinetâ]", "status": "[completed or incomplete]" ], "task_success": [true if successfully completed task goal, false if failed], "action_lesson": "[key action insight, e.g., âAttempted to place mug 1 directly in cabinet 2 without heating - must use microwave 1 firstâ OR âSuccessfully found knife in drawer 3 after checking wrong locationsâ]", "navigation_lesson": "[spatial insight, e.g., âMicrowave 1 located in kitchen area, not near cabinetsâ OR âMultiple sinkbasins exist - must check all for target objectâ]" EVALUATION GUIDELINES: ⢠The task outcome has been provided - use it to set task_success accordingly ⢠Focus on WHY the attempt had this outcome: â If successful: What sequence or strategy worked well? â If unsuccessful: What step failed or was missed? ⢠Each subtask status must reflect actual trajectory events ⢠Lessons should explain factors that led to the outcome ⢠Reference specific elements from trajectory (object IDs, locations, appliances) ⢠Use null for lessons only if truly not applicable Output ONLY the JSON evaluation. Prompt for Pairwise Induction on Minesweeper (1/2) You are an expert evaluating a Minesweeper game attempt. Task Requirements: Reveal all non-mine cells on a board_sizexboard_size board with n_mines mines without detonating any mine. You have just completed an attempt at this Minesweeper game. The game was success completed. reference_trajectory Current Trajectory of the attempt: current_trajectory <think> If a reference trajectory exists, compare it with the current trajectory. Analyze the current trajectory to determine: 1. Which subtasks were attempted and their completion status 2. Specific actions/decisions that caused the outcome 3. What went wrong (if failed) or right (if succeeded) 4. Devise a concise, new plan of action that accounts for any mistakes with reference to specific actions that should be taken in the next trial Game notation for reference: ⢠Cell states: ? (unopened), . (blank/no neighbors), 1-8 (mine count), * (mine) ⢠Coordinates: rows/columns indexed 1 to board_size ⢠Valid actions: (row, col) where 1â¤row,colâ¤1 ,col⤠board_size ⢠Blank cells auto-cascade to reveal connected blanks + borders Subtask Completion Criteria (binary evaluation for failed trajectories too): ⢠valid_moves: COMPLETED if made at least 2 valid format moves; INCOMPLETE if mostly invalid actions ⢠exploration_progress: COMPLETED if revealed >10% of board; INCOMPLETE if revealed <10% ⢠logical_attempt: COMPLETED if attempted any deduction (even if wrong); INCOMPLETE if only random/invalid moves ⢠error_recovery: COMPLETED if corrected any error within 3 attempts; INCOMPLETE if repeated same errors ⢠cascade_usage: COMPLETED if triggered or attempted any cascade; INCOMPLETE if only single cell reveals ⢠systematic_approach: COMPLETED if showed any pattern in move selection; INCOMPLETE if purely random </think> Prompt for Pairwise Induction on Minesweeper (2/2) Required JSON Output: "subtasks": [ "name": "valid_moves", "description": "[e.g., âMade 5 valid moves like (1,1), (2,3)â or âOnly invalid formats like (-1,-1)â]", "status": "[completed/incomplete]", "name": "exploration_progress", "description": "[e.g., âRevealed 15 cells (25% of board)â or âOnly revealed 2 cellsâ]", "status": "[completed/incomplete]", "name": "logical_attempt", "description": "[e.g., âTried to use (3,3)=1 constraintâ or âNo deduction attemptsâ]", "status": "[completed/incomplete]", "name": "error_recovery", "description": "[e.g., âFixed format after 2 attemptsâ or âRepeated invalid action 10 timesâ]", "status": "[completed/incomplete]", "name": "cascade_usage", "description": "[e.g., â(1,1) triggered 8-cell cascadeâ or âNo cascade attemptsâ]", "status": "[completed/incomplete]", "name": "systematic_approach", "description": "[e.g., âChecked corners firstâ or âRandom clickingâ]", "status": "[completed/incomplete]" ], "trajectory_value": [count of completed subtasks out of 6], "task_success": [true if successfully completed, false if unsuccessfully completed], "next_priority": "[Most important fix, e.g., âUse valid (row,col) formatâ or âWhen cell shows 1, count unopened neighborsâ]" Evaluation Rules: ⢠Award COMPLETED for ANY positive demonstration, even in failed games ⢠valid_moves: Just need 2+ correctly formatted moves anywhere in trajectory ⢠exploration_progress: 10% is roughly 6 cells on 8x8 board - achievable even if hit mine ⢠logical_attempt: Credit for trying logic, even if conclusion was wrong ⢠error_recovery: Credit for any correction, even if made new errors later ⢠cascade_usage: Credit for choosing corners/edges that could cascade ⢠systematic_approach: Credit for any non-random pattern in moves ⢠trajectory_value helps distinguish quality among failed attempts (0-6 scale) Output JSON only. Prompt for Pairwise Induction on Sokoban (1/2) You are an expert evaluating a Sokoban game attempt. Task Requirements: Push all boxes (âXâ) onto target spots (âOâ) in the grid without getting them stuck against walls (â#â) or in corners. You have just completed an attempt at this Sokoban level. The game was success completed. reference_trajectory Current Trajectory of the attempt: current_trajectory <think> If a reference trajectory exists, compare it with the current trajectory. Given the task outcome, analyze the trajectory to understand: 1. Which subtasks were attempted and their completion status 2. Specific actions/decisions that caused the outcome 3. What went wrong (if failed) or right (if succeeded) 4. Devise a concise, new plan of action that accounts for any mistakes with reference to specific actions that should be taken in the next trial Game notation for reference: ⢠Symbols: # (wall), _ (floor), O (target), X (box), P (player), â (box on target) ⢠Coordinates: (row, col) ⢠Valid actions: ["up", "down", "left", "right"] ⢠Rules: Push only (no pull), one box at a time, walls block movement. Subtask Completion Criteria (binary evaluation for failed trajectories too): ⢠valid_moves: COMPLETED if made at least 2 valid directional moves; INCOMPLETE if mostly invalid formats/hallucinations ⢠navigation_logic: COMPLETED if player successfully navigated to a box; INCOMPLETE if stuck hitting walls/looping ⢠box_interaction: COMPLETED if at least one box was pushed to a new coordinate; INCOMPLETE if no boxes moved ⢠deadlock_avoidance: COMPLETED if avoided pushing boxes into unrecoverable corners/walls; INCOMPLETE if immediate deadlock created ⢠goal_progress: COMPLETED if at least one box was placed on a target; INCOMPLETE if 0 boxes on targets ⢠systematic_approach: COMPLETED if moves showed clear intent (e.g., moving behind a box to push); INCOMPLETE if random walking </think> Prompt for Pairwise Induction on Sokoban (2/2) Required JSON Output: "subtasks": [ "name": "valid_moves", "description": "[e.g., âOutputted valid directions like up, downâ or âUsed invalid commandsâ]", "status": "[completed/incomplete]", "name": "navigation_logic", "description": "[e.g., âReached box at (3,2)â or âWalked into wall at (1,1) repeatedlyâ]", "status": "[completed/incomplete]", "name": "box_interaction", "description": "[e.g., âPushed box from (2,2) to (2,3)â or âNo boxes movedâ]", "status": "[completed/incomplete]", "name": "deadlock_avoidance", "description": "[e.g., âKept boxes away from cornersâ or âPushed box into corner (1,1)â]", "status": "[completed/incomplete]", "name": "goal_progress", "description": "[e.g., â1/3 boxes placed on targetâ or âNo boxes on targetsâ]", "status": "[completed/incomplete]", "name": "systematic_approach", "description": "[e.g., âCleared path for second boxâ or âRandom movementâ]", "status": "[completed/incomplete]" ], "trajectory_value": [count of completed subtasks out of 6], "task_success": [true if successfully completed, false if unsuccessfully completed], "next_priority": "[Most important fix, e.g., âDonât push box into corner at (1,1)â or âMove to (2,3) to push downâ]" Evaluation Rules: ⢠Award COMPLETED for ANY positive demonstration, even in failed games ⢠valid_moves: Just need 2+ correctly formatted actions ⢠navigation_logic: Credit for traversing the map without getting stuck on walls immediately ⢠box_interaction: Credit for changing the state of the board (moving a box) ⢠deadlock_avoidance: Credit if the first box move didnât result in an immediate game-over state ⢠goal_progress: Credit for securing at least one objective, even if others failed ⢠systematic_approach: Credit for positioning the player specifically to push a box ⢠trajectory_value helps distinguish quality among failed attempts (0-6 scale) Output JSON only. A.2 Prompt Templates for RL-Trained Self-Reflection Prompt for Pairwise Induction on Webshop You are an expert evaluating a WebShop shopping attempt. Target Task: task_description You have just completed an attempt at this shopping task. Trajectory of the attempt: current_trajectory <think> If a reference trajectory exists, compare it with the current trajectory. Analyze the trajectory to determine if the task was successful: 1. Identify the specific requirements in the âTarget Taskâ (attributes, type, options). 2. Examine the final action in the trajectory. Did it end in a âclick[buy]â? 3. If a purchase was made, compare the purchased itemâs details against the âTarget Taskâ requirements. 4. Did the purchased item match ALL requirements? (If no purchase was made, it is a failure). 5. What specific actions or decisions led to this outcome? 6. What are the 1-2 most valuable lessons from this attempt? </think> Output your evaluation as JSON: "subtasks": [ "name": "search_product", "description": "[describe actual search]", "status": "[completed or incomplete]", "name": "apply_filters", "description": "[describe filters used]", "status": "[completed or incomplete]", "name": "select_item", "description": "[describe selection]", "status": "[completed or incomplete]", "name": "complete_purchase", "description": "[describe purchase]", "status": "[completed or incomplete]" ], "task_success": [true if the correct item was purchased, false otherwise], "action_lesson": "[key action insight, e.g., âPrecise search with brand+model found exact matchâ OR âGeneric search missed required featuresâ]", "navigation_lesson": "[navigation insight, e.g., âEfficient use of filters saved timeâ OR âFailed to check additional pages with better optionsâ]" EVALUATION GUIDELINES: ⢠Determine Success Yourself: You must judge âtask_successâ by comparing the purchased item in the trajectory to the Target Task. ⢠Criteria for Success: The task is ONLY true if the agent successfully clicked âbuyâ on an item that matches all required attributes (color, size, flavor, etc.). ⢠Criteria for Failure: If the trajectory ends without a purchase, or if the wrong item was bought, âtask_successâ is false. ⢠Each subtask status must reflect actual trajectory events. ⢠Lessons should explain factors that led to the outcome. ⢠Reference specific elements from trajectory (item IDs, pages, search terms). ⢠Use null for lessons only if truly not applicable. Output ONLY the JSON evaluation. Prompt for Pairwise Induction on ALFWorld (1/2) You are an expert evaluating an ALFWorld embodied agent attempt. Target Task: task_description You have just completed an attempt at this household task. Trajectory of the attempt: current_trajectory <think> 1. If a reference trajectory exists, compare it with the current trajectory. 2. Analyze the trajectory to determine if the task was successful: (a) Identify the specific requirements in the âTarget Taskâ (target object, required state change, final destination). (b) Examine the sequence of actions. Did the agent successfully locate the correct object? (c) If a state change was required (clean, heat, cool, slice), was the correct appliance or tool used? (d) Did the agent place the object in the correct final receptacle? (e) Did the trajectory end with the âstopâ action after achieving the goal state? (If the agent stopped prematurely or failed to stop, it is a failure). (f) What specific actions or decisions led to this outcome? (g) What are the 1-2 most valuable lessons from this attempt? </think> Output your evaluation as JSON: "subtasks": [ "name": "locate_object", "description": "[describe search for target object]", "status": "[completed or incomplete]", "name": "acquire_object", "description": "[describe picking up target]", "status": "[completed or incomplete]", "name": "modify_state", "description": "[describe heating/cleaning /cooling/slicing if applicable, else âN/Aâ]", "status": "[completed, incomplete, or N/A]", "name": "place_object", "description": "[describe final placement]", "status": "[completed or incomplete]" ], "task_success": [true if the goal state was achieved and âstopâ was called, false otherwise], "action_lesson": "[key action insight, e.g., âUsed microwave to heat apple instead of fridgeâ OR âFailed to slice bread before platingâ]", "navigation_lesson": "[spatial/search insight, e.g., âsystematically checked all cabinet receptaclesâ OR âwasted steps revisiting empty drawersâ]" Prompt for Pairwise Induction on ALFWorld (2/2) EVALUATION GUIDELINES: ⢠Determine Success Yourself: You must judge âtask_successâ by comparing the final state in the trajectory to the Target Task. ⢠Criteria for Success: The task is ONLY true if the agent manipulated the correct object, achieved the correct state (e.g., hot, clean), placed it in the correct target, and issued the âstopâ command. ⢠Criteria for Failure: If the trajectory ends without the âstopâ command, or if the agent stopped without completing the goal (e.g., holding the object instead of placing it), âtask_successâ is false. ⢠Each subtask status must reflect actual trajectory events. ⢠Lessons should explain factors that led to the outcome. ⢠Reference specific elements from trajectory (object IDs like âapple 1â, receptacle IDs like âcountertop 2â). ⢠Use null for lessons only if truly not applicable. Output ONLY the JSON evaluation. Prompt for Pairwise Induction on Sokoban (1/2) You are an expert evaluating a Sokoban game attempt. Task Requirements: Push all boxes (âXâ) onto target spots (âOâ) in the grid without getting them stuck against walls (â#â) or in corners. You have just completed an attempt at this Sokoban level. Current Trajectory of the attempt: current_trajectory <think> 1. If a reference trajectory exists, compare it with the current trajectory. 2. Analyze the trajectory to determine if the task was successful: (a) Identify the grid layout and target locations in the âTarget Taskâ. (b) Examine the final board state in the trajectory. Are ALL boxes (âXâ) placed on targets (âOâ) resulting in ââ â? (c) If the game ended without success, check for deadlocks (boxes stuck in corners or against walls). (d) Did the player successfully navigate the player (âPâ) to push positions without hitting walls repeatedly? (e) What specific logic or movement behavior led to this outcome? (f) What are the 1-2 most valuable lessons from this attempt? (g) Devise a concise, new plan of action that accounts for any mistakes with reference to specific actions that should be taken in the next trial Game notation for reference: ⢠Symbols: # (wall), _ (floor), O (target), X (box), P (player), â (box on target) ⢠Coordinates: (row, col) ⢠Valid actions: ["up", "down", "left", "right"] ⢠Rules: Push only (no pull), one box at a time, walls block movement. Subtask Completion Criteria (binary evaluation for failed trajectories too): ⢠valid_moves: COMPLETED if made at least 2 valid directional moves; INCOMPLETE if mostly invalid formats/hallucinations ⢠navigation_logic: COMPLETED if player successfully navigated to a box; INCOMPLETE if stuck hitting walls/looping ⢠box_interaction: COMPLETED if at least one box was pushed to a new coordinate; INCOMPLETE if no boxes moved ⢠deadlock_avoidance: COMPLETED if avoided pushing boxes into unrecoverable corners/walls; INCOMPLETE if immediate deadlock created ⢠goal_progress: COMPLETED if at least one box was placed on a target; INCOMPLETE if 0 boxes on targets ⢠systematic_approach: COMPLETED if moves showed clear intent (e.g., moving behind a box to push); INCOMPLETE if random walking </think> Prompt for Pairwise Induction on Sokoban (2/2) Required JSON Output: "subtasks": [ "name": "valid_moves", "description": "[e.g., âOutputted valid directions like up, downâ or âUsed invalid commandsâ]", "status": "[completed/incomplete]", "name": "navigation_logic", "description": "[e.g., âReached box at (3,2)â or âWalked into wall at (1,1) repeatedlyâ]", "status": "[completed/incomplete]", "name": "box_interaction", "description": "[e.g., âPushed box from (2,2) to (2,3)â or âNo boxes movedâ]", "status": "[completed/incomplete]", "name": "deadlock_avoidance", "description": "[e.g., âKept boxes away from cornersâ or âPushed box into corner (1,1)â]", "status": "[completed/incomplete]", "name": "goal_progress", "description": "[e.g., â1/3 boxes placed on targetâ or âNo boxes on targetsâ]", "status": "[completed/incomplete]", "name": "systematic_approach", "description": "[e.g., âCleared path for second boxâ or âRandom movementâ]", "status": "[completed/incomplete]" ], "trajectory_value": [count of completed subtasks out of 6], "task_success": [true if successfully placed all boxes on targets, false if deadlock or incomplete], "next_priority": "[Most important fix, e.g., âDonât push box into corner at (1,1)â or âMove to (2,3) to push downâ]" Evaluation Rules: ⢠Determine Success Yourself: You must judge âtask_successâ by comparing the final board state in the trajectory to the Target Task. ⢠Criteria for Success: The task is ONLY true if ALL boxes are on target spots (ââ â). ⢠Criteria for Failure: If the trajectory ends with a deadlock, or if the agent stopped before placing all boxes, âtask_successâ is false. ⢠Each subtask status must reflect actual trajectory events. ⢠Lessons should explain factors that led to the outcome (planning vs. random). ⢠Reference specific elements from trajectory (coordinates, symbols). ⢠Use null for lessons only if truly not applicable. Output ONLY the JSON evaluation. Prompt for Pairwise Induction on Minesweeper (1/2) You are an expert evaluating a Minesweeper game attempt. Task Requirements: Reveal all non-mine cells on a board_sizexboard_size board with n_mines mines without detonating any mine. You have just completed an attempt at this Minesweeper game. Current Trajectory of the attempt: current_trajectory <think> 1. If a reference trajectory exists, compare it with the current trajectory. 2. Analyze the trajectory to determine if the task was successful: (a) Identify the board constraints (size, mine count) in the âTarget Taskâ. (b) Examine the final action in the trajectory. Did it result in a mine detonation (loss) or a cleared board (win)? (c) If the game ended without a mine detonation, check if ALL safe cells were revealed. (d) Did the player successfully flag mines (optional but helpful) and reveal all safe spots? (If a mine was hit or safe cells remain hidden, it is a failure). (e) What specific logic or guessing behavior led to this outcome? (f) What are the 1-2 most valuable lessons from this attempt? (g) Devise a concise, new plan of action that accounts for any mistakes with reference to specific actions that should be taken in the next trial Game notation for reference: ⢠Cell states: ? (unopened), . (blank/no neighbors), 1-8 (mine count), * (mine) ⢠Coordinates: rows/columns indexed 1 to board_size ⢠Valid actions: (row, col) where 1 ⤠row,col ⤠board_size ⢠Blank cells auto-cascade to reveal connected blanks + borders Subtask Completion Criteria (binary evaluation for failed trajectories too): ⢠valid_moves: COMPLETED if made at least 2 valid format moves; INCOMPLETE if mostly invalid actions ⢠exploration_progress: COMPLETED if revealed >10% of board; INCOMPLETE if revealed <10% ⢠logical_attempt: COMPLETED if attempted any deduction (even if wrong); INCOMPLETE if only random/invalid moves ⢠error_recovery: COMPLETED if corrected any error within 3 attempts; INCOMPLETE if repeated same errors ⢠cascade_usage: COMPLETED if triggered or attempted any cascade; INCOMPLETE if only single cell reveals ⢠systematic_approach: COMPLETED if showed any pattern in move selection; INCOMPLETE if purely random </think> Prompt for Pairwise Induction on Minesweeper (2/2) Required JSON Output: "subtasks": [ "name": "valid_moves", "description": "[e.g., âMade 5 valid moves like (1,1), (2,3)â or âOnly invalid formats like (-1,-1)â]", "status": "[completed/incomplete]", "name": "exploration_progress", "description": "[e.g., âRevealed 15 cells (25% of board)â or âOnly revealed 2 cellsâ]", "status": "[completed/incomplete]", "name": "logical_attempt", "description": "[e.g., âTried to use (3,3)=1 constraintâ or âNo deduction attemptsâ]", "status": "[completed/incomplete]", "name": "error_recovery", "description": "[e.g., âFixed format after 2 attemptsâ or âRepeated invalid action 10 timesâ]", "status": "[completed/incomplete]", "name": "cascade_usage", "description": "[e.g., â(1,1) triggered 8-cell cascadeâ or âNo cascade attemptsâ]", "status": "[completed/incomplete]", "name": "systematic_approach", "description": "[e.g., âChecked corners firstâ or âRandom clickingâ]", "status": "[completed/incomplete]" ], "trajectory_value": [count of completed subtasks out of 6], "task_success": [true if successfully cleared all safe cells, false if detonated mine or incomplete], "next_priority": "[Most important fix, e.g., âUse valid (row,col) formatâ or âWhen cell shows 1, count unopened neighborsâ]" Evaluation Rules: ⢠Determine Success Yourself: You must judge âtask_successâ by comparing the final board state in the trajectory to the Target Task. ⢠Criteria for Success: The task is ONLY true if the agent successfully revealed ALL safe cells without detonating a mine. ⢠Criteria for Failure: If the trajectory ends with a mine detonation, or if the agent stopped before revealing all safe cells, âtask_successâ is false. ⢠Each subtask status must reflect actual trajectory events. ⢠Lessons should explain factors that led to the outcome (logic vs. guessing). ⢠Reference specific elements from trajectory (coordinates, cell values). ⢠Use null for lessons only if truly not applicable. Output ONLY the JSON evaluation. A.3 Prompts for Analyzing the Quality of Intrinsic Feedback To assess the fidelity of the intrinsic feedback generated via self-reflection, we employ GPT-4o (OpenAI et al., 2024) as an external judge. Our evaluation focuses on two key components: the accuracy of the induced subtask completion scores (intrinsic rewards) and the quality of the summarized lessons (intrinsic feedback). To verify the accuracy of the subtask completion scores, we utilize the prompt detailed in Section A.1. To evaluate the quality of the summarized lessons derived from the agentâs trajectories, we use the prompt presented below. Prompt for Evaluating Summarized Lessons System Prompt: You are an expert evaluator of AI Memory Systems. Your goal is to determine the âInformation Gainâ and âCrucialityâ of lessons generated by an agent. You must distinguish between generic fluff (low quality) and specific, actionable insights (high quality). User Prompt: # Context The agent performed a task in a web environment. Actual Outcome: actual_outcome # Trajectory (History of Actions) trajectory # Agentâs Generated Reflection (containing Lessons) reflection # Evaluation Task Analyze the action_lesson and navigation_lesson in the reflection above. 1. Specificity: Is the lesson specific to the UI elements/errors encountered? (e.g., âClicking âSubmitâ failed because the form was emptyâ vs. âI failed to clickâ). 2. Causal Accuracy: Does the lesson correctly identify the root cause of the actual_outcome? 3. Utility: If the agent retrieves this lesson in a future attempt, will it significantly improve the success rate? # Output Format (JSON Only) "lesson_quality_score": <int 1-10>, "specificity_rating": <"High"|"Medium"|"Low">, "utility_rating": <"High"|"Medium"|"Low">, "reasoning": "<Short explanation of why this lesson is useful/useless>", "is_hallucination": <bool, true if lesson mentions events not in trajectory> Appendix B Implementation Details Hyperparameter Qwen-2.5-7B-Instruct Llama-3.1-8B-Instruct Description Training Configuration Training batch size 16 16 Accumulated batch size per update Validation batch size 128 128 Batch size for validation Learning rate 10â610^-6 10â610^-6 Optimizer learning rate Max prompt length 16 384 16 384 Maximum input context length (tokens) Max response length 2 048 2 048 Maximum generated response length (tokens) Group size (N) 8 8 Number of rollouts per prompt Total steps 150 / 300 150 / 300 Training epochs (150 for ALFWorld and WebShop; 300 for Sokoban and Minesweeper) Evaluation frequency 5 5 Epochs between consecutive evaluations Reward and Regularization Extrinsic reward (RextR^ext) 0, 10\0,\,10\ 0, 10\0,\,10\ Scalar reward from the environment Intrinsic reward (RintR^int) [0, 1][0,\,1] [0, 1][0,\,1] Capability-evolution intrinsic reward KL coefficient (β) 0.01 0.01 KL-divergence regularization weight Discount factor (Îł) 0.95 0.95 Discount factor for multi-step returns Memory and Sampling Training temperature 0.4 0.4 Sampling temperature during rollouts Validation temperature 0.4 0.4 Sampling temperature during validation Initial utility score 0.5 0.5 Initial utility assigned to each memory entry Utility smoothing (βutil _util) 0.05 0.05 Exponential moving average coefficient for utility updates UCB exploration constant (c) 1.0 1.0 Exploration coefficient in UCB-based retrieval Relevanceâutility weight (Îą) 0.7 0.7 Trade-off coefficient in retrieval scoring Memory-augmented ratio 1:11:1 1:11:1 Ratio of memory-augmented to base rollouts Self-Reflection (RL-Trained Variant) Reflection reward (RreflectR^reflect) 0, 10\0,\,10\ 0, 10\0,\,10\ Scalar reward for reflection accuracy Reflection weight (Îťreflect _reflect) 1.0 1.0 Weight of the self-reflection objective relative to the decision-making objective Evaluation Configuration Evaluation temperature 0.4 0.4 Sampling temperature during evaluation Max inference tokens 2 048 2 048 Maximum token budget per inference step Table 10: Default hyperparameters and training configurations for RetroAgent across all environments. Appendix C Superiority of Pairwise Induction over Single Induction We analyze reflection sequences generated during GRPO training augmented with either single-trajectory or pairwise-trajectory induction. Specifically, we sample four sequences per task prompt every 10 training steps, randomly selecting 800 sequences from this pool. Using GPT-4o as an external evaluator, we assess two metrics: the alignment between the agentâs induced subtask completion scores and oracle outcomes (induced by GPT-4o), and the quality of the agentâs summarized lessons (see Appendix A.3 for evaluation prompts). (a) Single induction. (b) Pairwise induction. Figure 12: Comparison of intrinsic discrimination capability for Qwen-2.5-7B-Instruct using single vs. pairwise induction, measured by subtask completion scores on WebShop. Figure 12 compares the ability of single- and pairwise-trajectory induction to discriminate trajectory quality within the training-free self-reflection mechanism. For this analysis, we define a successful trajectory as having a completion score of 1. While both methods correctly assign minimal scores to âTotal Failureâ cases, pairwise induction demonstrates superior discrimination for intermediate outcomes. For example, as shown in Figure 12(a), single induction struggles to distinguish âEarly Progressâ from âTotal Failure,â resulting in a distribution heavily skewed toward zero for partial successes. In contrast, pairwise induction (Figure 12(b)) effectively recognizes partial progress, assigning distinctly higher scores to âEarly Progressâ trajectories and establishing a clear separation from total failures. Appendix D Generated Trajectories by RetroAgent Instruction: You are an expert autonomous agent operating in the WebShop eâcommerce environment. Your task is to: Find me officially licensed, machine wash menâs t-shirts with polyester heathers, heathers cotton, cotton heather, needle sleeve, classic fit with color: pink, and fit type: youth, and size: small, and price lower than 50.00 dollars. Past reflections on similar tasks: Action Insight: The search terms used did not match the available products, leading to no suitable options being found. | Navigation Insight: Efficient use of filters could have potentially narrowed down the search, but the specific dimensions and color were not met in any of the results. Warning: These lessons may be outdated. Use them only if they align with your current observation. Observation 1: Your current observation is: âSearchâ. Your admissible actions of the current situation are: [âsearch[<your query>]â, âclick[search]â] Now itâs your turn to take one action for the current step. You should first reason step-by-step about the current situation, then think carefully which admissible action best advances the shopping goal. This reasoning process MUST be enclosed within <think> </think> tags. Once youâve finished your reasoning, you should choose an admissible action for current step and present it within <action> </action> tags. RetroAgent: <think> To achieve the goal of finding officially licensed, machine wash menâs t-shirts with specific attributes, I need to construct a search query that includes all the required attributes⌠I will construct the search query to include all these attributes and then execute the search. </think> Action: search[officially licensed menâs t-shirts polyester heathers heathers cotton cotton heather needle sleeve classic fit color pink fit type youth fit type size small price < 50.00] Observation 2: You are now at step 2 and your current observation is: âBack to Searchâ [SEP] âPage 1 (Total results: 50)â [SEP] âNext >>â [SEP] âB09GL561XHâ [SEP] âDisney Mickey And Friends Happy Birthday Mickey Confetti T-Shirtâ [SEP] â$22.99â [SEP] âB07XPR3R7Nâ [SEP] âMarvel Avengers: Endgame Captain America Americaâs Language T-Shirtâ [SEP] â$22.99â ⌠[SEP] âB09Q8RD8YNâ [SEP] âBungo Stray Anime Dogs Anime Character, Long Sleeve, Sweatshirt, Hoodie, T shirtâ [SEP] â$19.99â. RetroAgent: <think> The current page does not list any t-shirts that are officially licensed⌠The best action at this point is to click on the first product listed to see if it is the correct product⌠</think> Action: click[b09gl561xh] Observation 3: You are now at step 3⌠âfit typeâ [SEP] âmenâ [SEP] âwomenâ [SEP] âyouthâ [SEP] âcolorâ [SEP] âwhiteâ ⌠[SEP] âpinkâ [SEP] âsizeâ [SEP] âsmallâ ⌠âDisney Mickey And Friends Happy Birthday Mickey Confetti T-Shirtâ [SEP] âPrice: $22.99â. RetroAgent: <think> The current page is not displaying the results for the search query that was previously conducted⌠The fit type âyouthâ is one of the specific fit types mentioned in the search query, so I will click on âyouthâ to ensure that the filter is set correctly. </think> Action: click[youth] Observation 4: You are now at step 4⌠âfit typeâ [SEP] âmenâ [SEP] âwomenâ [SEP] âyouthâ [SEP] âcolorâ ⌠[SEP] âpinkâ [SEP] âsizeâ [SEP] âsmallâ ⌠âDisney Mickey And Friends Happy Birthday Mickey Confetti T-Shirtâ [SEP] âPrice: $22.99â. RetroAgent: <think> The current situation shows that the search has narrowed down to a specific product⌠The next step is to select the color pink to ensure that the color is officially licensed and meets the criteria. </think> Action: click[pink] Observation 5: You are now at step 5⌠âDisney Mickey And Friends Happy Birthday Mickey Confetti T-Shirtâ [SEP] âPrice: $22.99â ⌠RetroAgent: <think> The current situation shows that the search has narrowed down to a specific product⌠with the color being pink. The size is not yet selected⌠The next step is to select the size small to confirm that it meets the criteria. </think> Action: click[small] Observation 6: You are now at step 6⌠âDisney Mickey And Friends Happy Birthday Mickey Confetti T-Shirtâ [SEP] âPrice: $22.99â ⌠RetroAgent: <think> The current situation shows that the search has narrowed down to a specific product⌠The price is $22.99, which is within the budget⌠I can proceed with clicking on âBuy Nowâ to confirm the selection. </think> Action: click[buy now] Table 11: Successful trajectory generated by RetroAgent (in-context self-reflection) at training step 150 on the WebShop validation set. Instruction: You are an expert autonomous agent operating in the WebShop eâcommerce environment. Your task is to: Find me officially licensed, machine wash menâs t-shirts with polyester heathers, heathers cotton, cotton heather, needle sleeve, classic fit with color: pink, and fit type: youth, and size: small, and price lower than 50.00 dollars. Observation 1: Your current observation is: âSearchâ. Your admissible actions of the current situation are: [âsearch[<your query>]â, âclick[search]â]. Now itâs your turn to take one action for the current step. You should first reason step-by-step about the current situation, then think carefully which admissible action best advances the shopping goal. This reasoning process MUST be enclosed within <think> </think> tags. Once youâve finished your reasoning, you should choose an admissible action for current step and present it within <action> </action> tags. GRPO: <think> To achieve the goal of finding officially licensed, machine wash menâs t-shirts with specific attributes, I need to construct a search query that includes all the required attributes⌠I will construct the search query to include all these attributes and then execute the search. </think> Action: search[officially licensed menâs t-shirts polyester heathers heathers cotton cotton heather needle sleeve classic fit color pink fit type youth fit type size small price < 50.00] Observation 2: You are now at step 2 and your current observation is: âBack to Searchâ [SEP] âPage 1 (Total results: 50)â [SEP] âNext >â [SEP] âB08HF132PNâ [SEP] âFlamingo Pink waterbird costume Gift Premium T-Shirtâ [SEP] â$19.99â [SEP] âB07JVVDJ6Lâ [SEP] âCanadian Tuxedo T Shirt - Novelty Denim Tux T-Shirtâ [SEP] â$19.99â [SEP] âB09GL561XHâ [SEP] âDisney Mickey And Friends Happy Birthday Mickey Confetti T-Shirtâ [SEP] â$22.99â ⌠[SEP] âB07HRFSNL4â [SEP] âNautica Menâs Solid Crew Neck Short-Sleeve Pocket T-Shirtâ [SEP] â$16.05 to $40.98â. GRPO: <think> The current observation indicates that the search results are being displayed, and the first result is "B08HF132PN" with the title "Flamingo Pink waterbird costume Gift Premium T-Shirt" and a price of $19.99. However, this product does not match the specific criteria we are looking for, as it is a costume t-shirt rather than a menâs t-shirt, and it is not officially licensed. Given the current search results, the next logical step is to click on the first result to see if it matches the criteria⌠The most appropriate action here is to click on the first result, "B08HF132PN", to check if it matches the criteria. </think> Action: click[b08hf132pn] Observation 3: You are now at step 3 and your current observation is: âBack to Searchâ [SEP] â< Prevâ [SEP] âfit typeâ [SEP] âmenâ [SEP] âwomenâ [SEP] âyouthâ [SEP] âsizeâ [SEP] âsmallâ ⌠âFlamingo Pink waterbird costume Gift Premium T-Shirtâ [SEP] âPrice: $19.99â [SEP] âRating: N.A.â [SEP] âDescriptionâ [SEP] âFeaturesâ [SEP] âReviewsâ [SEP] âBuy Nowâ. GRPO: <think> 1. The current observation shows that the search results are being displayed⌠2. The product does not match the search criteria⌠3. The next step is to refine the search results to ensure that the search criteria are met. The most appropriate action here is to click on "youth" to filter the results to show only products that are intended for youth. </think> Action: click[youth] Observation 4: You are now at step 4⌠âfit typeâ [SEP] âmenâ [SEP] âwomenâ [SEP] âyouthâ [SEP] âsizeâ [SEP] âsmallâ ⌠âFlamingo Pink waterbird costume Gift Premium T-Shirtâ [SEP] âPrice: $19.99â ⌠GRPO: <think> The current observation shows that the search results are being displayed, and the last action was to click on "youth" fit type⌠However, the product does not match the criteria we are looking for⌠Given the current situation, the best action to take is to click on "small" size, as we are looking for a menâs t-shirt with specific attributes, and the size is part of the search criteria. </think> Action: click[small] Observation 5: You are now at step 5⌠âFlamingo Pink waterbird costume Gift Premium T-Shirtâ [SEP] âPrice: $19.99â ⌠GRPO: <think> The current observation indicates that the search results are filtered by fit type, size, and price. The last action was to click on "small" size⌠However, this product does not match the specific criteria we are looking for⌠Given the current search results, the best action to take is to click on "Buy Now" for the current product to see if it meets the criteria or if it is a close match. </think> Action: click[Buy Now] Table 12: Failed trajectory generated by RetroAgent (in-context self-reflection) at training step 65 on the WebShop validation set.