Paper deep dive
LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
Jingya Wang, Yuyang Gao, Liuzhenghao Lv, Yonghong Tian, Yuyang Liu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/3/2026, 2:02:12 AM
Summary
The paper introduces LabEvolver, a training-free framework for safe and grounded wet-lab agents that utilizes a nested dual-loop architecture. The inner loop handles state-driven safe execution via an Observer, Operator, and Supervisor with a tri-layer safety gate. The outer loop employs a Strategist to distill completed trajectories into reusable skills, strategies, and safety experiences, stored in a global memory bank. Evaluated on robotic wet-lab tasks and the ALFWorld benchmark, LabEvolver demonstrates significant improvements in success rates, reduced completion times, and efficient memory management compared to baselines like ReAct and Act.
Entities (10)
Relation Signals (9)
LabEvolver → containsmodule → Operator
confidence 95% · the Operator maps goals to executable LabSkill actions
LabEvolver → containsmodule → Supervisor
confidence 95% · the Supervisor manages runtime context for closed-loop replanning
LabEvolver → containsmodule → Strategist
confidence 95% · the Strategist distills completed execution traces into reusable skill, strategy, and safety experience
LabEvolver → containsmodule → Observer
confidence 95% · LabEvolver couples a state-grounded inner trial loop... the Observer maintains the hierarchical laboratory state
LabEvolver → evaluatedon → ALFWorld
confidence 95% · On ALFWorld, it further improves cumulative success rate
LabEvolver → outperforms → ReAct
confidence 92% · improves cumulative success rate within 20 steps from 76.2% with ReAct to 91.4%
Strategist → manages → Global_Memory_Bank
confidence 90% · updating the global experience bank M for cross-experiment reuse
LabEvolver → reducesmetric → pH-regulation_completion_time
confidence 90% · reducing pH-regulation completion time... by 48.2%
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic memory from execution experience. LabEvolver couples a state-grounded inner trial loop for adaptive perception, online planning, and safety validation with an outer evolution loop that distills completed trajectories into reusable skill, strategy, and safety experience. On robotic solution-preparation tasks, LabEvolver demonstrates real-world feasibility, reducing pH-regulation completion time and safety-gate intercepts by 48.2% and 60.0%, respectively. On ALFWorld, it further improves cumulative success rate within 20 steps from 76.2% with ReAct to 91.4% over 500 continual tasks, showing generality beyond wet-lab settings. These results support learn-by-doing experience evolution as a feasible path toward closed-loop automated scientific discovery. The project page is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.27690v2
- Canonical: https://arxiv.org/abs/2607.27690v2
Trouble viewing inline? Open PDF directly →
Full Text
43,850 characters extracted from source content.
Expand or collapse full text
LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents Jingya Wang 2∗ Yuyang Gao 3∗ Liuzhenghao Lv 3 Yonghong Tian 1,2,3† Yuyang Liu 1,2† 1 School of AI for Science, Peking University 2 School of Electronic and Computer Engineering, Peking University 3 School of Computer Science, Peking University liuyuyang13,yhtian@pku.edu.cn Abstract We introduce LabEvolver, a training-free frame- work that equips safe and grounded wet-lab agents with episodic memory from execution experience. LabEvolver couples a state-grounded inner trial loop for adaptive perception, online planning, and safety validation with an outer evolution loop that dis- tills completed trajectories into reusable skill, strat- egy, and safety experience. On robotic solution- preparation tasks, LabEvolver demonstrates real- world feasibility, reducing pH-regulation completion time and safety-gate intercepts by 48.2% and 60.0%, respectively. On ALFWorld, it further improves cu- mulative success rate within 20 steps from 76.2% with ReAct to 91.4% over 500 continual tasks, show- ing generality beyond wet-lab settings. These re- sults support learn-by-doing experience evolution as a feasible path toward closed-loop automated sci- entific discovery. The project page is available at https://andygao6186.github.io/LabEvolver/. 1 Introduction From empirical observation and theoretical mod- eling to computational simulation and data-driven science, successive paradigms have expanded the scale and speed of scientific discovery. The Fifth Paradigm, autonomous scientific discovery, is now becoming increasingly plausible [1]. Self-driving lab- oratories (SDLs) advance this vision by integrat- ing robotics for high-throughput autonomous ex- periments [2–5]. However, most SDLs still depend on predefined interfaces and workflows, which lack ∗ Equal contribution. † Corresponding author. Figure 1: Our proposed LabEvolver. In a pH- regulation task, LabEvolver surpasses one-shot ac- tion generation and purely within-trial feedback cor- rection by leveraging accumulated experience. a decision-making mechanism to understand high- level scientific goals and dynamically organize low- level operations [6, 7]. In this context, foundation models offer a cogni- tive layer for SDLs by interpreting goals, generat- ing plans, and connecting scientific knowledge with physical execution [8–10]. However, many existing systems still represent physical actions as abstract symbols, leaving the planning and execution lay- ers without a unified experimental state [8, 9, 11]. Vision-language-action (VLA) models further inte- grate perception, language, and robot actions, of- fering a promising route toward open-ended exper- imental operation [12, 13]. Yet their deployment in real scientific experiments remains bounded by prac- tical limitations: 1) End-to-end policies often require costly scenario-specific data and are difficult to scale across diverse laboratory settings [12–14]; 2) implic- itly generated actions are difficult to trace, making 1 arXiv:2607.27690v2 [cs.RO] 31 Jul 2026 Figure 2: Overview of the LabEvolver Framework. Given a high-level experimental goal, LabEvolver executes autonomous wet-lab workflows through a nested dual-loop process. (a) Inner Trial Loop: the Observer maintains the hierarchical laboratory state s t , the Operator maps goals to executable LabSkill actions under the tri-layer Safety Gate Γ, and the Supervisor manages runtime context for closed-loop replanning. (b) Outer Evolution Loop: the Strategist distills completed execution traces into reusable skill, strategy, and safety experience, updating the global experience bank M for cross-experiment reuse. their physical effects hard to predict or control under real-world constraints [15, 16]. These limitations highlight that embodied agents cannot rely solely on end-to-end action generation. Instead, they require grounded state anchors to plan physical actions prior to costly trial-and-error. This motivates a training-free, bounded framework that preserves flexible goal-directed control while enforc- ing explicit constraints in lab environments. Beyond bounded execution within a single trial, autonomous laboratories must also improve across repeated experimental practice. Current embodied systems lack post-deployment practice loops, forc- ing each run to begin as a cold start without distill- ing trajectories into reusable knowledge [10, 16, 17]. Consequently, despite automating execution, they remain close to static intelligence. By contrast, hu- man scientific expertise is not innate, but devel- ops through conceptual understanding and hands-on practice [1]. To move toward the Fifth Paradigm of scientific discovery, autonomous laboratories should acquire a learn-by-doing capability that accumu- lates episodic memory from concrete operations [18– 20]. This shifts the goal from static automation toward continuously learnable scientific discovery, where planning, cognition, and action improve to- gether through practice. To operationalize this paradigm, we intro- duce LabEvolver, an experience-driven and state- grounded autonomous experimentation framework that organizes laboratory workflows into a nested dual-loop process. In the inner trial loop, the sys- tem supports online execution through multi-agent collaboration under strict runtime safety gates. In the outer evolution loop, LabEvolver uses Strate- gist to distill completed experiments into reusable skills, strategies, and safety mitigation rules. Across real-world experiments, LabEvolver improves robust planning, safe execution, and cross-scenario evolu- tion in long-horizon scientific tasks. Evaluation on ALFWorld further shows that LabEvolver provides a more general embodied cognitive architecture for experience reuse [21]. Our contributions are summarized as follows: • We propose LabEvolver, a training-free dual- loop framework that integrates online closed- loop execution with post-trial experience evolu- tion for autonomous scientific experimentation, without updating model weights. • We introduce Strategist, a hierarchical experi- ence manager that distills completed laboratory 2 trajectories into reusable skill-level, strategy- level, and safety-level experience for better planning and execution. • We validate LabEvolver through physical wet- lab experiments and ALFWorld, showing its ef- fectiveness in safe execution, experience reuse, and cross-scenario transfer. 2 Related Work Self-driving laboratories. Self-driving laborato- ries integrate robotic platforms, experimental de- sign, and closed-loop optimization to accelerate sci- entific discovery. Prior work spans autonomous hy- pothesis testing [2], programmable robotic work- flows [6, 7], closed-loop discovery systems [3–5], and recent wet-lab robotic platforms [10, 15, 17]. While existing systems establish the hardware foundation for automated experimentation, they typically re- quire structured interfaces and rigid layouts. In contrast, LabEvolver addresses the complementary problem of flexible execution, enabling safe, state- grounded execution under changing wet-lab environ- ments. Foundation models for embodied experimen- tation. Foundation models increasingly serve as decision layers for robotic agents, supporting goal interpretation, tool use, and embodied planning [8, 9, 22–25]. Robot foundation models and VLA poli- cies connect language and vision to executable con- trol [12–14, 26–29], while laboratory-oriented agents extend these capabilities to scientific protocol exe- cution [11, 16]. However, many systems still depend on specialized interfaces or costly data. LabEvolver instead uses foundation models for high-level per- ception and planning, while grounding execution in structured laboratory states and LabSkill actions. Self-evolving agents. Language and embodied agents can improve from past trials without up- dating model weights through reflection, proce- dural memory, and reusable skills [18, 19, 30– 35]. In embodied settings, systems such as Voy- ager, EmbodiSkill, and ELITE accumulate experi- ential knowledge across tasks [20, 36, 37], with re- cent work extending self-evolution toward labora- tory robotics [38–40]. Most methods, however, re- main in digital or general robotic settings where feedback is detached from deployed wet-lab con- straints. LabEvolver addresses this gap by convert- ing state-based wet-lab trajectories into practical skill, strategy, and safety experience. 3 Methodology We formulate autonomous wet-lab experimentation as a constrained closed-loop problem. Given a high- level experimental goal g, LabEvolver seeks a phys- ically viable trajectory τ that reaches the target ex- perimental condition under real-world constraints. As shown in Figure 2, LabEvolver adopts a nested dual-loop architecture. The inner trial loop con- structs a hierarchical laboratory state, grounds goals into LabSkill actions, and validates each action through a tri-layer safety gate for state-grounded execution. The outer evolution loop leverages the Strategist to extract, maintain, and retrieve state- paired experience for future trials. Throughout the following exposition, we use the solution prepara- tion task as a running example, where the goal g is to bring the terminal solution pH to around 5. 3.1 Inner Trial Loop: State-Driven Safe Execution The inner trial loop consists of three coupled mech- anisms for hierarchical state construction, state- conditioned action planning, and runtime safety val- idation to support safe and robust long-horizon em- bodied wet-lab execution. Dynamic Hierarchical Laboratory State The shared laboratory state is the control anchor of the inner trial loop. LabEvolver maintains this state through the Observer, a perception module that converts free-form environmental feedback into structured variables for planning and safety check- ing. At step t, the system maintains s t = (A t ,E t ,O t ). Here, A t denotes the agent embodiment state, in- cluding robot pose, gripper status, and interacting instruments. E t denotes the environment state, in- cluding object identities, poses, and attributes. O t denotes multimodal observations, including visual perception, sensor readings, and actuator feedback. Unlike prior robotic laboratory systems that rely on manually specified layouts [10, 15], the Observer dynamically constructs this state from visual obser- vations and runtime feedback. • State Initialization. Before execution, the Observer reconstructs the initial state s 0 from RGB-D observations, as illustrated in Figure 3. It first identifies task-relevant objects through vision-language scene understanding, then lo- calizes and segments these objects to estimate 3 Figure 3: State Initialization Pipeline. A four- stage process for the Observer to reconstruct the ini- tial laboratory state s 0 from RGB-D data. their spatial poses. The resulting identities, poses, and attributes are compiled into E 0 = e i 0 = (ι i 0 ,p i 0 ,x i 0 ) N 0 i=1 , while robot configuration initializes A 0 = (q 0 ,h 0 ). For the running exam- ple, the perceived mixing beaker is represented as e b 0 = (ι b 0 ,p b 0 ,x b 0 ), with x b 0 initialized as empty before preparation. Further robustness analysis of the Observer is provided in the supplemen- tary document. • State Tracking. During execution, the Ob- server updates s t from environmental feedback after each robotic action. These updates pro- vide a reliable reference to the current labora- tory state for subsequent planning, validation, and execution. During pH regulation, balance and pH-meter feedback update the beaker state x b t , including its solution mass and pH value. State-Conditioned Action Planning Based on the current laboratory state, the Op- erator uses LabSkill to ground high-level reason- ing into executable laboratory actions while hiding hardware-specific control details. LabSkill is orga- nized as a three-level action interface. The atomic ac- tion layer defines basic robot-executable primitives, the digital tool layer connects the Operator with ex- ternal programs and runtime records, and the se- mantic skill layer provides task-level templates that compose atomic actions and tools into unified guide- lines. This hierarchy allows the Operator to compose complex operations from reusable primitives, rather than following a hard-coded action sequence. See the supplementary document for more LabSkill details. At step t, the Operator receives the planning con- text c t = (g,s t ,I LS ), where g represents the experi- mental goal, s t the current laboratory state, and I LS the LabSkill instructions. Conditioned on c t , the Op- erator first generates a high-level intent u t and sub- sequently grounds it into a parameterized action a t : u t ∼ π(c t ), a t = (k t ,θ t )∼ π(·| u t ,c t ), where k t denotes the action type, and θ t specifies the target objects, waypoints, and interaction pa- rameters dynamically determined by the Operator according to s t to match the laboratory environ- ment. For example, when the latest pH measurement is above the pH 5 target range, the Operator may instantiate an acid-pouring action, with θ t specify- ing the target beaker, pouring amount, and control parameters. For further long-horizon execution, the Supervi- sor periodically compresses accumulated state tran- sitions into structured progress checkpoints, reduc- ing context saturation and goal drift. Tri-layer Runtime Safety Gate Because flexible LabSkill composition may produce actions that violate physical or procedural con- straints, LabEvolver employs a Tri-layer Runtime Safety Gate as a safety boundary. Given the cur- rent state s t and a proposed action a t , the gate val- idates the action at three levels before dispatch: Γ(s t ,a t ) = Γ intf (s t ,a t )∧ Γ proc (s t ,a t )∧ Γ phys (s t ,a t ). Here, Γ intf audits whether the LabSkill call is well formed, Γ proc checks consistency with the current ex- perimental progress, and Γ phys verifies physical and measurement feasibility under s t . For instance, Γ proc blocks taking the pH meter when the robotic arm is still holding a test tube. When Γ(s t ,a t ) = 1, the action is dispatched to the physical platform, and the resulting feedback y t+1 updates the state via s t+1 = U (s t ,a t ,y t+1 ). Con- versely, if Γ(s t ,a t ) = 0, physical execution is halted immediately, and the gate generates a structured ex- ception log identifying the violation to guide Op- erator replanning. By unifying pre-execution filter- ing with post-execution state updates, this mecha- nism establishes a reliable safety boundary for long- horizon wet-lab automation. A detailed analysis is provided in the supplementary document. 3.2 Outer Evolution Loop: State- Paired Experience Refinement The outer evolution loop turns completed trials into state-paired experience fragments for future reuse. Instead of storing trajectories as stateless text logs, the Strategist links each reusable experience to the state condition under which it was useful, making retrieval sensitive to the current physical context rather than only to task keywords. 4 Experience extraction. After a trial terminates, the Strategist maps the trajectory τ and goal g to a structured experience tuple e = Ψ(τ,g) = (e skill ,e strategy ,e safety ), where e skill records effec- tive LabSkill parameters and boundary conditions, e strategy captures reusable procedural knowledge, and e safety stores blocked-action diagnoses and failure-avoidance rules. In the running example, e strategy can record that an initial pH of 7.74 suggests first adding about 8 g of acid before re-measurement. Experience maintenance. The extracted expe- rience is merged into a global memory bank: M← Maintain(M,e,o), o∈Add, Update, Upvote, Downvote. By comparing newly extracted experience e with existing records in M, the Strategist inserts novel items, merges redundant entries, and adjusts mem- ory priorities to retain useful records while suppress- ing unreliable ones. To keep long-term experience accumulation bounded, LabEvolver further applies an Ebbinghaus-style soft forgetting mechanism [41], periodically decaying the scores of experience not reused in successful trials, and experience records with scores below zero are deleted. Details are pro- vided in the supplementary document. Experience retrieval. Before each trial, the Strategist retrieves top-K relevant experience frag- ments from M based on the goal and initial state, forming a retrieved context r to initialize the inner execution loop. During long-horizon tasks, retrieval is dynamically refreshed at Supervisor checkpoints. This state-based retrieval mechanism enables prior experience to guide both initial planning and mid- course replanning without replaying entire trajecto- ries. 4 Experiments 4.1 Experimental Setup We evaluate LabEvolver in two complementary regimes, progressing from simulation to real-world wet-lab validation. ALFWorld isolates outer-loop ex- perience evolution in a scalable long-horizon bench- mark to test algorithmic scalability and cross-task generalization. Robotic wet-lab tasks further as- sess the full dual-loop system under noisy physical feedback, state changes, and real-world safety con- straints. Together, these evaluations test whether ex- ecution feedback becomes reusable experience while preserving state-valid laboratory operation. Compared methods and backbones. Unless otherwise specified, we use DeepSeek-V4-Pro [42] as the default backbone. Cross-model comparisons ad- ditionally include Claude-Sonnet-4.6 [43], Qwen3.5- 35B-A3B [44], and GPT-5 [45]. We compare four methods: Act, which directly predicts actions; Re- Act [46], which adds within-task reasoning; Inner- only, which runs the complete inner loop without historical experience; and LabEvolver, which further implements outer-loop experience evolution. ALFWorld benchmark. We evaluate scalable outer-loop experience evolution on long-horizon tasks using ALFWorld [21]. All experiments are training-free and conducted without any prior mem- ory, treating the train and valid_unseen splits purely as task streams. LabEvolver adopts the Re- Act style prompt, incorporating retrieved experience as its sole additional context. We report the suc- cess rate within 20 steps (Success@20) across 134 valid_unseen tasks, while tracking both task per- formance and experience accumulation metrics over a stream of 500 shuffled train tasks. Robotic wet-lab platform. To validate the real- world effectiveness of LabEvolver, we conduct physi- cal experiments on a wet-lab robotic platform shown in Figure 1. The setup consists of a RealMan RM75- B arm, an Inspire Robots EG2-4B gripper, an Intel RealSense D435 RGB-D camera, an electronic bal- ance, and a water-quality meter. Wet-lab tasks and metrics. We evaluate three wet-lab task families: quantitative pouring, single- objective pH regulation, and coupled pH–EC reg- ulation. As a core challenge in embodied manip- ulation, quantitative pouring tests whether physi- cal feedback can be converted into reusable skill- level guidance. Serving as an indispensable founda- tion for scientific experimentation, solution prepara- tion in pH and coupled pH–EC regulation evaluates experiment-level planning from sequential measure- ments, where the agent must select operations and action arguments online rather than follow a prede- fined protocol. With 11 action types, a T-step tra- jectory admits up to 11 T nominal action-type se- quences, reaching 1.38 × 10 104 at T = 100 before considering continuous parameters and state con- straints. We report task success (Succ.), reagent ad- 5 Figure 4: ALFWorld results. (a) Cumulative Success@20 over 500 continual tasks. (b) Success@20 across backbone models on 134 valid_unseen tasks. (c) Available-memory growth over the 500-task continual stream. ditions (Add.), safety-gate intercepts (Gate), com- pletion time, and task-specific accuracy. 4.2 ALFWorld Results: Scalable Ex- perience Improves Decision Mak- ing LabEvolver improves training-free long- horizon decision making. We first evaluate the outer evolution loop on 500 ALFWorld train- split episodes with DeepSeek-V4-Pro, comparing LabEvolver against Act, ReAct, and MemP [31], an online baseline that proceduralizes trajectories into workflows. As shown in Figure 4(a), LabEvolver achieves the highest cumulative success rate of 91.4%, outperforming MemP by 3.8 points, ReAct by 15.2 points, and Act by 18.8 points. The large margins over non-memory baselines highlight the efficacy of test-time experience evolution, while the gain over MemP demonstrates that post-episode reflection and structured memory organization provide more actionable guidance than generic procedural workflows. Furthermore, LabEvolver’s sustained advantage beyond early exploration indicates that its accumulated experience becomes increasingly reusable for downstream decisions. LabEvolver generalizes across diverse back- bone models. We next examine whether outer- loop experience improves decision making across different backbone models. On 134 valid_unseen ALFWorld tasks, LabEvolver achieves the highest Success@20 for every tested backbone, as shown in Figure 4(b). The gain is smaller for stronger executors with already competitive action genera- tion, such as Claude-Sonnet-4.6 and GPT-5, but be- comes more pronounced for budget-sensitive mod- els. For example, LabEvolver improves Qwen3.5- 35B-A3B by 8.2 percentage points over ReAct and raises DeepSeek-V4-Pro from 72.4% with Act to 85.8%. These results suggest that retrieved experi- ence reduces short-budget exploration and recovery costs, while the final gain depends on how effectively the executor interprets and applies the guidance. Full results, including cross-model and cross-method comparisons along with Success@50 metrics, are re- ported in the supplementary document. LabEvolver prevents memory explosion. We further analyze memory-bank growth during contin- ual ALFWorld evaluation. As shown in Figure 4(c), MemP grows almost linearly by proceduralizing ev- ery trajectory, reaching 1177 memories at episode 500. LabEvolver limits this growth to 372 memo- ries through Add, Update, Upvote, and Down- vote, which merge shared experience and admit only useful new fragments. With further long-term memory forgetting, the bank stays nearly flat after early accumulation and contains only 281 memories at episode 500. These results validate the memory- maintenance mechanism of the Strategist, which prevents memory explosion while compacting accu- mulated experience into a bounded memory bank. See the supplementary document for more detailed memory analysis. 4.3 Pouring: Feedback-Driven Skill Evolution Quantitative pouring tests whether physical feed- back can be converted into transferable LabSkill pa- rameters. We implement pouring with a PD-based 6 Figure 5: Representative pH regulation process showing safe closed-loop operation with state-feedback-driven replanning. Yellow boxes indicate the instruments currently manipulated by the robotic arm. MethodMAE (g)↓ SD (g)↓ Time (s)↓ Fixed PD0.5000.10021.035 Adaptive pre-stop PD0.1330.15320.360 CLAIRify0.8670.37961.930 LabEvolver (Ours)0.0830.117 11.179 Table 1: Controller comparison for 20 g water pour- ing. LabSkill whose k p , k d , and pre_stop control wrist velocity and anticipatory stopping. LabEvolver converts physical feedback into accurate pouring skills. We first evaluate the pouring LabSkill on a controlled 20 g water-transfer task. We compare LabEvolver against three non- evolving controllers (fixed PD, manually tuned adaptive pre-stop PD, and CLAIRify [11]), repeating all comparisons three times to ensure consistency. Unlike these baselines, LabEvolver optimizes con- trol parameters via outer-loop experience evolution across 40 self-exploration trials. As shown in Table 1, LabEvolver achieves the lowest MAE and shortest execution time. Compared with the manually tuned adaptive pre-stop PD controller, LabEvolver reduces MAE by 37.5% and time by 45.1%. This result shows that the outer loop can convert physical trials into a reusable control prior, while the inner loop dynam- ically grounds the evolved skill based on real-time mass measurements. Figure 6: Quantitative-pouring transfer across target masses and liquids. (a) Mean absolute error and val- idation count for each locked-parameter condition. (b) Adapted anticipatory stopping parameter across conditions. Gray hatched regions denote conditions where the target mass approaches the test tube ca- pacity. LabEvolver enables autonomous skill experi- ence transfer. We next evaluate whether pour- ing experience transfers beyond the initial 20 g wa- ter setting to water, sucrose solution, and olive oil at target masses of 10, 20, 30, and 40 g. These condi- tions vary both target quantity and fluid dynamics, requiring adaptive pouring parameters. As shown in Figure 6, 41 of 43 locked-parameter validation tri- als fall within ±0.2 g of the target mass, yielding a 95.3% success rate and a trial-weighted MAE of 0.081 g. LabEvolver automatically assigns condition- specific values of k p , k d , and pre_stop, avoiding the manual retuning required by non-evolving con- trollers for each liquid-target pair. Full per-condition parameter details are provided in the supplementary 7 Backbone MethodSucc. Add.↓ Gate↓ Time (s)↓ DeepSeek ActNo401130 DeepSeek ReActYes811875 DeepSeek Inner-only Yes32872 DeepSeek LabEvolver Yes20652 QwenActNo21408 QwenReActYes44706 QwenInner-only Yes7152311 QwenLabEvolver Yes22609 Claude ActNo21476 Claude ReActYes502068 Claude Inner-only Yes411317 Claude LabEvolver Yes301145 Table 2: pH 5 regulation across DeepSeek-V4-Pro, Qwen3.5-35B-A3B, and Claude-Sonnet-4.6. document. 4.4 pH Regulation: State-Grounded Strategy Evolution pH regulation tests whether state tracking and run- time safety gates keep long-horizon solution prepa- ration valid, and whether strategy-level experience reduces unnecessary additions and invalid proposals. LabEvolver achieves robust solution prepa- ration across backbones. We compare the four agent configurations for pH 5 regulation across three backbone models. As shown in Table 2, Act fails to reach the target across the tested backbones, in- dicating that one-shot action generation from the model’s prior knowledge is insufficient for this task setting. ReAct and Inner-only recover task comple- tion by using feedback in different ways, but both still incur extra additions or safety-gate interven- tions. By adding outer-loop experience, LabEvolver reduces the mean additions from 5.67 to 2.33 and completion time from 25.83 to 13.37 minutes relative to ReAct, while reducing mean gate intercepts from 6.00 to 0.67 relative to Inner-only under the same state mechanism. The improvement is larger for less reliable executors such as Qwen3.5-35B-A3B, indi- cating that retrieved experience can compensate for weaker planning. Overall, the inner loop supports ro- bust online execution, while the outer loop improves efficiency and proposal quality. LabEvolver enables strategy experience transfer. We further examine repeated pH 5 regulation and transfer to pH 6 and pH 9 targets. Figure 7: Coupled pH–EC regulation results. (a) Planning cost decreases as experience becomes more task-relevant. (b) Performance on coupled pH–EC targets. Rectangles denote admissible target regions, and stars denote final prepared states. LabEvolver consistently reaches the target ranges with limited operations, indicating that retrieved strategy experience can guide reagent selection and dose estimation while preserving measurement- conditioned replanning. Detailed results are provided in the supplementary document. 4.5 Coupled pH–EC Regulation: Multi-Objective Evolution Unlike single-objective pH regulation, coupled pH– EC regulation requires the agent to coordinate two interacting solution properties, where different reagents can shift pH and EC in different directions with nonlinear and target-dependent effects. This task therefore evaluates how accumulated experi- ence supports online planning under coupled multi- objective solution dynamics. LabEvolver supports complex planning with task-relevant experience. We compare three experience settings on the same pH 6–EC 300 tar- get to isolate the effect of experience relevance. The no-experience setting relies only on online feedback, the partial setting reuses experience from quanti- tative pouring and single-objective pH regulation, and the full setting further includes previous joint pH–EC trajectories while excluding the same tar- get. As shown in Figure 7(a), partial experience already reduces exploratory operations, indicating 8 that skill-level and single-objective strategy expe- rience provide transferable priors for dosing and reagent selection. Full experience further reduces ad- ditions from 21 to 11, eliminates gate intercepts, and shortens completion time by 55.0%, showing that joint pH–EC trajectories provide task-specific guidance for handling interacting objectives. Under the full experience setting, LabEvolver also reaches all four admissible pH–EC target regions in Fig- ure 7(b), requiring 8 additions on average despite target-dependent and non-monotonic trajectories. These results suggest that beyond simply accumu- lating experience, LabEvolver further benefits from reusing history with higher task relevance, while re- lying on measurement-conditioned replanning dur- ing execution. Full per-target results are provided in the supplementary document. 5 Conclusion We present LabEvolver, a training-free, dual-loop framework that enables flexible and safe robotic execution grounded in real-time laboratory states. By continually accumulating physical experience, it advances wet-lab agents from state-valid control to learn-by-doing autonomy, a practical step to- ward automated scientific discovery. Our wet-lab experiments demonstrate the real-world feasibility of this approach, while further evaluation in ALF- World shows that its outer-loop experience distilla- tion generalizes as a universal embodied cognition mechanism. Looking ahead, the long-horizon physi- cal traces accumulated by LabEvolver offer valuable data to power future scientific world models. De- spite these capabilities, current limitations mainly arise from relying on human-designed LabSkills and safety rules, which may constrain adaptation to un- seen operations and novel failure modes. References [1] Yihang Chen, Zhongwei Yu, Yang Li, Xingyu Lu, Qiang Li, Hongxia Yu, and Jun Wang. Position: Autonomous scientific discov- ery needs embodied experimentation with learnability.TechRxiv preprint, 2026. URLhttps://w.techrxiv.org/doi/abs/ 10.36227/techrxiv.177100993.33363174/v1. https://doi.org/10.36227/techrxiv.177100993. 33363174/v1. [2] Ross D King, Jem Rowland, Stephen G Oliver, Michael Young, Wayne Aubrey, Emma Byrne, Maria Liakata, Magdalena Markham, Pinar Pir, Larisa N Soldatova, et al. The automation of science. Science, 324(5923):85–89, 2009. [3] Benjamin P MacLeod, Fraser GL Parlane, Thomas D Morrissey, Florian Häse, Loïc M Roch, Kevan E Dettelbach, Raphaell Moreira, Lars PE Yunker, Michael B Rooney, Joseph R Deeth, et al. Self-driving laboratory for accel- erated discovery of thin-film materials. Science Advances, 6(20):eaaz8867, 2020. [4] Benjamin Burger, Phillip M Maffettone, Vladimir V Gusev, Catherine M Aitchison, Yang Bai, Xiaoyan Wang, Xiaobo Li, Ben M Alston, Buyi Li, Rob Clowes, et al. A mobile robotic chemist. Nature, 583(7815):237–241, 2020. [5] Nathan J Szymanski, Bernardus Rendy, Yuxing Fei, Rishi E Kumar, Tanjin He, David Milsted, Matthew J McDermott, Max Gallant, Ekin Do- gus Cubuk, Amil Merchant, et al. An au- tonomous laboratory for the accelerated synthe- sis of novel materials. Nature, 624(7990):86–91, 2023. [6] Sebastian Steiner, Jakob Wolf, Stefan Glatzel, Anna Andreou, Jarosław M Granda, Gra- ham Keenan, Trevor Hinkley, Gerardo Aragon- Camarasa, Philip J Kitson, Davide Angelone, et al. Organic synthesis in a modular robotic system driven by a chemical programming lan- guage. Science, 363(6423):eaav2211, 2019. [7] Loic M. Roch, Florian Hase, Christoph Kreis- beck, Teresa Tamayo-Mendoza, Lars P. E. Yunker, Jason E. Hein, and Alán Aspuru- Guzik. ChemOS: Orchestrating autonomous experimentation. Science Robotics, 3(19): eaat5559, 2018.doi: 10.1126/scirobotics. aat5559. [8] Daniil A Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. Autonomous chemical re- search with large language models. Nature, 624 (7992):570–578, 2023. [9] Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. Augmenting large language models with chemistry tools. Nature machine intelligence, 6(5):525–535, 2024. [10] Yibo Qiu, Zan Huang, Zhiyu Wang, Handi Liu, Yiling Qiao, Yifeng Hu, Shu’ang Sun, Hangke Peng, Ronald X Xu, and Mingzhai Sun. 9 Biomars: A multi-agent robotic system for au- tonomous biological experiments, 2025. URL https://arxiv.org/abs/2507.01485. [11] Naruki Yoshikawa, Marta Skreta, Kourosh Darvish, Sebastian Arellano-Rubach, Zhi Ji, Lasse Bjørn Kristensen, Andrew Zou Li, Yuchi Zhao, Haoping Xu, Artur Kuramshin, et al. Large language models for chemistry robotics. Autonomous Robots, 47(8):1057–1086, 2023. [12] Anthony Brohan, Noah Brown, Justice Car- bajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Flo- rence, Chuyuan Fu, Montse Gonzalez Are- nas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Lisa Lee, Tsang-Wei Ed- ward Lee, Sergey Levine, Yao Lu, Henryk Michalewski, Igor Mordatch, Karl Pertsch, Kanishka Rao, Krista Reymann, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Pierre Ser- manet, Jaspiar Singh, Anikait Singh, Radu Soricut, Huong Tran, Vincent Vanhoucke, Quan Vuong, Ayzaan Wahid, Stefan Welker, Paul Wohlhart, Jialin Wu, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich. Rt-2: Vision-language-action mod- els transfer web knowledge to robotic control, 2023. URL https://arxiv.org/abs/2307.15818. [13] Moo Jin Kim, Karl Pertsch, Siddharth Karam- cheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. Openvla: An open-source vision- language-action model, 2024. URL https:// arxiv.org/abs/2406.09246. [14] Homanga Bharadhwaj, Jay Vakil, Mohit Sharma, Abhinav Gupta, Shubham Tulsiani, and Vikash Kumar. Roboagent: Generalization and efficiency in robot manipulation via seman- tic augmentations and action chunking. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 4788–4795. IEEE, 2024. [15] Kourosh Darvish, Marta Skreta, Yuchi Zhao, Naruki Yoshikawa, Sagnik Som, Miroslav Bog- danovic, Yang Cao, Han Hao, Haoping Xu, Alán Aspuru-Guzik, et al. Organa: A robotic assis- tant for automated chemistry experimentation and characterization. Matter, 8(2), 2025. [16] Baochang Ren, Xinjie Liu, Xi Chen, Yanshuo Liu, Chenxi Li, Daqi Gao, Zeqin Su, Jin- tao Xing, Zirui Xue, Rui Li, Xiangyu Zhao, Shuofei Qiao, Minting Pan, Wangmeng Zuo, Lei Bai, Dongzhan Zhou, Ningyu Zhang, and Hua- jun Chen. Labvla: Grounding vision-language- action models in scientific laboratories, 2026. URL https://arxiv.org/abs/2606.13578. [17] Kevin Angers, Kourosh Darvish, Naruki Yoshikawa, Sargol Okhovatian, Dawn Ban- nerman, Ilya Yakavets, Florian Shkurti, Alán Aspuru-Guzik, and Milica Radisic. Robo- culture: A robotics platform for automated biological experimentation, 2025.URL https://arxiv.org/abs/2505.14941. [18] Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. Advances in neural information processing systems, 36:8634–8652, 2023. [19] Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. Expel: Llm agents are experiential learners. In Proceedings of the AAAI Conference on Artifi- cial Intelligence, volume 38, pages 19632–19642, 2024. [20] Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large lan- guage models, 2023. URL https://arxiv.org/ abs/2305.16291. [21] Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning, 2021. URL https://arxiv.org/abs/ 2010.03768. [22] Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691, 2022. 10 [23] Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. In 2023 IEEE International conference on robotics and automation (ICRA), pages 9493–9500. IEEE, 2023. [24] Ishika Singh, Valts Blukis, Arsalan Mousa- vian, Ankit Goyal, Danfei Xu, Jonathan Trem- blay, Dieter Fox, Jesse Thomason, and Animesh Garg. Progprompt: Generating situated robot task plans using large language models, 2022. URL https://arxiv.org/abs/2209.11302. [25] Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023. [26] Anthony Brohan, Noah Brown, Justice Carba- jal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Haus- man, Alex Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Kuang-Huei Lee, Sergey Levine, Yao Lu, Utsav Malla, Deeksha Manjunath, Igor Mor- datch, Ofir Nachum, Carolina Parada, Jodi- lyn Peralta, Emily Perez, Karl Pertsch, Jor- nell Quiambao, Kanishka Rao, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Kevin Sayed, Jaspiar Singh, Sumedh Sontakke, Austin Stone, Clayton Tan, Huong Tran, Vincent Vanhoucke, Steve Vega, Quan Vuong, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich. Rt-1: Robotics transformer for real- world control at scale, 2023. URL https://arxiv. org/abs/2212.06817. [27] Abby O’Neill, Abdul Rehman, Abhiram Mad- dukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892– 6903. IEEE, 2024. [28] Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty El- lis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. arXiv preprint arXiv:2403.12945, 2024. [29] Kevin Black, Noah Brown, Danny Driess, Ad- nan Esmail, Michael Equi, Chelsea Finn, Nic- colo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky. π 0 : A vision-language-action flow model for general robot control, 2026. URL https://arxiv.org/abs/2410.24164. [30] Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refine- ment with self-feedback. Advances in neural in- formation processing systems, 36:46534–46594, 2023. [31] Runnan Fang, Yuan Liang, Xiaobin Wang, Jia- long Wu, Shuofei Qiao, Pengjun Xie, Fei Huang, Huajun Chen, and Ningyu Zhang. Memp: Ex- ploring agent procedural memory. In Findings of the Association for Computational Linguis- tics: ACL 2026, pages 17490–17502, 2026. [32] Haiteng Zhao, Chang Ma, Guoyin Wang, Jing Su, Lingpeng Kong, Jingjing Xu, Zhi-Hong Deng, and Hongxia Yang. Empowering large language model agents through action learning, 2024. URL https://arxiv.org/abs/2402.15809. [33] Guibin Zhang, Muxin Fu, Kun Wang, Frank Wan, Miao Yu, and Shuicheng Yan. G-memory: Tracing hierarchical memory for multi-agent systems. Advances in Neural Information Pro- cessing Systems, 38:12988–13018, 2026. [34] Peng Xia, Jianwen Chen, Hanyang Wang, Ji- aqi Liu, Kaide Zeng, Yu Wang, Siwei Han, Yiyang Zhou, Xujiang Zhao, Haifeng Chen, Zeyu Zheng, Cihang Xie, and Huaxiu Yao. Skillrl: Evolving agents via recursive skill- augmented reinforcement learning, 2026. URL https://arxiv.org/abs/2602.08234. [35] Jingwei Ni, Yihao Liu, Xinpeng Liu, Yutao Sun, Mengyu Zhou, Pengyu Cheng, Dexin Wang, Er- chao Zhao, Xiaoxi Jiang, and Guanjun Jiang. Trace2skill: Distill trajectory-local lessons into 11 transferable agent skills, 2026. URL https: //arxiv.org/abs/2603.25158. [36] Ruofei Ju, Xinrui Wang, Xin Ding, Yifan Yang, Hao Wu, Shiqi Jiang, Qianxi Zhang, Hao Wen, Xiangyu Li, Weijun Wang, Kun Li, Yunxin Liu, Haipeng Dai, Wei Wang, and Ting Cao. Em- bodiskill: Skill-aware reflection for self-evolving embodied agents, 2026. URL https://arxiv.org/ abs/2605.10332. [37] Bingqing Wei, Zhongyu Xia, Dingai Liu, Xiaoyu Zhou, Zhiwei Lin, and Yongtao Wang. Elite: Experiential learning and intent-aware transfer for self-improving embodied agents, 2026. URL https://arxiv.org/abs/2603.24018. [38] Ruiying Li, Yunlang Zhou, YuYao Zhu, Kylin Chen, Jingyuan Wang, Sukai Wang, Kongtao Hu, Minhui Yu, Bowen Jiang, Zhan Su, Ji- ayao Ma, Xin He, Yongjian Shen, Yang Yang, Guanghui Ren, Maoqing Yao, Wenhao Wang, and Yao Mu. Roboclaw: An agentic framework for scalable long-horizon robotic tasks, 2026. URL https://arxiv.org/abs/2603.11558. [39] Dongjie Huo, Haoyun Liu, Guoqing Liu, Dekang Qi, Zhiming Sun, Maoguo Gao, Jianxin He, Yandan Yang, Xinyuan Chang, Feng Xiong, Xing Wei, Zhiheng Ma, and Mu Xu. Abot- claw: A foundation for persistent, cooperative, and self-evolving robotic agents, 2026. URL https://arxiv.org/abs/2604.10096. [40] Yunfei Wang, Xiaohao Xu, Yang Li, and Xi- aonan Huang. When search becomes mem- ory: Turning robot design trials into transfer- able skills, 2026. URL https://arxiv.org/abs/ 2605.25832. [41] Hermann Ebbinghaus. Memory: A Contribu- tion to Experimental Psychology. Teachers Col- lege, Columbia University, New York, 1913. Original work published 1885. [42] Anyi Xu, Bangcai Lin, Bing Xue, Bingx- uan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, et al. Deepseek-v4: Towards highly effi- cient million-token context intelligence. arXiv preprint arXiv:2606.19348, 2026. [43] Anthropic.Claude Sonnet 4.6 system card.https://w.anthropic.com/system- cards, 2026. Accessed: 2026-07-28. [44] Qwen Team. Qwen3.5: Towards native multi- modal agents. https://qwen.ai/blog?id=qwen3. 5, February 2026. Model card: https:// huggingface.co/Qwen/Qwen3.5-35B-A3B; Ac- cessed: 2026-07-28. [45] OpenAI. GPT-5 system card. https://openai. com/index/gpt-5-system-card/, 2025.Ac- cessed: 2026-07-28. [46] Shunyu Yao, Jeffrey Zhao, Dian Yu, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. In NeurIPS 2022 Foundation Models for Decision Making Workshop, 2022. 12