Paper deep dive
Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use
Yi Ding, Yanzhao Yu, Xili Dai, Xianbiao Qi, Peiwen Sun, Xueqian Wang, Xiangyu Yue, Jianan Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/17/2026, 4:54:10 AM
Summary
The paper introduces ART (Agentic Robot with Tool-use), a framework that integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use capabilities. ART employs a tool-injection mechanism using adaptive LoRA fine-tuning to allow VLA models to leverage off-the-shelf tools for vision, affordance, and embodiment enhancement without catastrophic forgetting. The authors propose a data generation pipeline to create a 30K trajectory dataset for training and demonstrate that ART achieves a 20% higher success rate than baselines in simulation and real-world tasks.
Entities (8)
Relation Signals (7)
ART → buildsupon → VLA
confidence 95% · This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose Agentic Robot with Tool-use (ART).
ART → employs → Tool Injection
confidence 95% · ART is a tool-injection framework that tunes any VLA model...
ART → uses → LoRA
confidence 95% · we define a two-stage LoRA fine-tuning method... only the LoRA modules dedicated to tool reasoning are fine-tuned.
ART → evaluatedon → LIBERO
confidence 90% · Through experiments in the LIBERO simulation and real-world settings, we demonstrate that our method enhances the model’s robustness...
ART → finetunes → Pi-FAST
confidence 90% · We use the pre-trained 3B π-FAST [35] model... we fine-tuned the model for 1 epoch...
ART → outperforms → RoboTool
confidence 85% · ART achieves a 20% higher success rate than mainstream baselines... Compared to vanilla VLA models... ART reduces the complexity...
ART → outperforms → OpenVLA
confidence 85% · ART achieves a 20% higher success rate than mainstream baselines... Compared to vanilla VLA models... ART reduces the complexity...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose Agentic Robot with Tool-use (ART). ART is a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for low-level vision, high-level affordance, and embodiment enhancement. Compared to vanilla VLA models with a whole continuous action solution space, ART reduces the complexity of the action solution space through tool-use, which not only improves generalizability across different tasks but also reduces data dependency. To demonstrate the advantages (high generalizability and low data dependency) of this framework, we first built a dataset of 30K tool-use trajectories and action demonstrations, which is much smaller than those used by baseline methods. We then designed a training regimen for long-trajectory tool-use reasoning in challenging environments. Experiments show that ART achieves a 20% higher success rate than mainstream baselines on simulation and real-world tasks, such as pick-and-place in the dark at novel viewpoints. Empirical results highlight the benefits of an agent-based approach: modular tool utilization enables more efficient training, lightweight deployment, and scalable integration of new tools. This design fosters robustness, adaptability, and extensibility, paving the way for the practical deployment of VLA systems in complex real-world scenarios.
Tags
Links
- Source: https://arxiv.org/abs/2608.14047v1
- Canonical: https://arxiv.org/abs/2608.14047v1
Trouble viewing inline? Open PDF directly →
Full Text
56,617 characters extracted from source content.
Expand or collapse full text
Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use Yi Ding 1 * Yanzhao Yu 1,2 * Xili Dai 3 Xianbiao Qi 4 Peiwen Sun 5 Xueqian Wang 2 Xiangyu Yue 5 Jianan Wang 1† 1 Astribot 2 Tsinghua University 3 Juxi Tech 4 IntelliFusion Inc. 5 The Chinese University of Hong Kong onedonedone@sjtu.edu.cn, yuyz24@mails.tsinghua.edu.cn, jiananwang@astribot.com Abstract This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose Agentic Robot with Tool-use (ART). ART is a tool-injection frame- work that tunes any VLA model to leverage off-the-shelf tool modules for low-level vision, high-level affordance, and em- bodiment enhancement. Compared to vanilla VLA models with a whole continuous action solution space, ART reduces the complexity of the action solution space through tool- use, which not only improves generalizability across dif- ferent tasks but also reduces data dependency. To demon- strate the advantages (high generalizability and low data dependency) of this framework, we first built a dataset of 30K tool-use trajectories and action demonstrations, which is much smaller than those used by baseline methods. We then designed a training regimen for long-trajectory tool- use reasoning in challenging environments. Experiments show that ART achieves a 20% higher success rate than mainstream baselines on simulation and real-world tasks, such as pick-and-place in the dark at novel viewpoints. Empirical results highlight the benefits of an agent-based approach: modular tool utilization enables more efficient training, lightweight deployment, and scalable integration of new tools. This design fosters robustness, adaptability, and extensibility, paving the way for the practical deploy- ment of VLA systems in complex real-world scenarios. 1. Introduction The field of Vision-Language-Action (VLA) models has seen significant advancements, broadly categorized into two main development paths: modular and end-to-end ap- proaches. The modular approach, exemplified by methods such as RoboTool [45], RoboScript [7], and RoboCodeX [34], encapsulates specific functionalities like action exe- * Equal contribution. Work done during internship at Astribot. † Corresponding author. cution and goal recognition within distinct tool modules. Models are then trained to call the APIs of these predefined tools. While this approach offers complete decoupling of different capabilities, its reliance on fixed tool sets and ac- tion functions inherently limits its applicability, primarily to simpler action outputs. In contrast, the end-to-end ap- proach, as demonstrated by models like OpenVLA [28], ECoT [50], and π [3, 4, 15, 35], involves training a single base model on extensive multi-task datasets. This paradigm allows models to produce highly precise and detailed ac- tions in these trained tasks. However, a significant draw- back arises when encountering novel scenarios and tasks: the base model typically necessitates complex and costly post-training. Furthermore, the inherent coupling of all ca- pabilities in end-to-end models makes them susceptible to catastrophic forgetting when adapted to new data. A critical question thus emerges: can we harness the pow- erful, precise action output capabilities of end-to-end VLA models while simultaneously enabling them to swiftly inte- grate and utilize external tools for adapting to new scenar- ios and tasks? Naturally, the action solution space of a VLA model would be significantly constrained if continuous actions were replaced by predefined discrete tools. Hence, in this paper, we propose ART (Agentic Robot with Tool-use), a novel fine-tuning tool-injection framework for VLA models that seamlessly integrates multi-modal tool usage. Through ART, VLA models can acquire new, plug-and-play capabili- ties—such as visual enhancement, sophisticated spatial rea- soning, and expanded action abilities—without compromis- ing their original action output fidelity. Compared to vanilla VLA models with a whole continuous action solution space, ART framework reduces the complexity of the action solu- tion space through tool-use, which not only improves gen- eralizability across different tasks but also reduces data de- pendency. To demonstrate the advantages (high generalizability and low data dependency) of this framework, we first built a dataset of 30K tool-use trajectories and action demonstra- 1 arXiv:2608.14047v1 [cs.RO] 14 Aug 2026 Figure 1. Comparison of VLA Paradigms, including (a) standard end-to-end VLA [4, 6, 28], (b) VLA enhanced with Chain-of-Thought reasoning [50], (c) modular robotic behavior synthesis [31, 34], and (d) our Agentic Robot with Tool use (ART) framework, which intDegrates tool use into the end-to-end action generation process. The ART model demonstrates enhanced adaptability to complex environments (Task 1) and correct hallucinated predictions in foundation models (Task 2) through dynamic tool use. tions, which is much smaller than those used by base- line methods. Inspired by T3-Agent [18], we introduce a novel pipeline to extend existing VLA datasets. This three- step process—task design, reasoning generation, and tool- trajectory synthesis—efficiently generates large volumes of tool-usage data. Initially, we systematically ”degrade” ex- isting VLA data by introducing complexities into task in- structions, thereby creating tasks that explicitly necessitate tool usage for their resolution. Subsequently, based on these modified tasks, we leverage the base VLA model to gener- ate the requisite reasoning processes for invoking necessary tools. Finally, these reasoning processes are synthesized into comprehensive tool usage trajectories. This innova- tive approach allows for the integration of tool reasoning into existing datasets without the need for costly manual collection of new action data. Then, to mitigate the risk of catastrophic forgetting during fine-tuning with new tool-use data, we define a two-stage LoRA fine-tuning method. Dur- ing the tool usage training phase, the original VLA model weights are frozen, and only the LoRA modules dedicated to tool reasoning are fine-tuned. Crucially, once the VLA model has utilized the tools, we mask out the LoRA layer’s outputs during action generation, allowing the model to re- vert to its original action output capabilities. This strate- gic decoupling effectively separates tool reasoning from the core VLA model, thereby preventing data conflicts and pre- serving high-quality action generation during training. In summary, our key contributions are as follows: • We introduce a data generation method that integrates multi-modal tools into existing VLA datasets, enabling complex tool usage in long task chains. • We propose ART (Agentic Robot with Tool-use), a fine- tuning tool-injection framework for VLA models that preserves the pre-trained VLA action generation ability while enabling new tool reasoning capabilities. • Through experiments in the LIBERO simulation and real- world settings, we demonstrate that our method enhances the model’s robustness in handling new scenarios, ob- jects, and action instructions. 2 2. Related Works We first provide an overview of two research directions in VLA models, followed by a brief discussion of tool use in multi-modal agents, another cornerstone of ART. Modular Robotic Behavior Synthesis. The fundamental challenge in early-stage research lies in bridging the gap be- tween high-level human instructions and robotic behaviors. A mainstream research line breaks down action generation into modular subtasks [7, 24, 31, 32, 34, 38, 40, 45]. Given the observation o t at time t, the model π synthesizes a se- quence of external modules that progressively address the task, rather than directly generating actions: π : o t 7→ (f,f act ), f : o t 7→ o ∗ t , f act : o ∗ t 7→ A t (1) where given the environment observations o t , f returns in- termediate representations, referred to as affordances, such as object positions [20, 26] and segmentation masks [29]. f act denotes the embodied functions that translate the geo- metric position to robot movements [40]. For example, CaP [31] invokes the function pick obj and placeatpos in sequence to complete a pick-and-place task. The method packs up different robotic abilities as modules, and then the model can flexibly utilize off-the-shelf tools. However, the method oversimplifies the language-to-action mapping into handcrafted functions, and it struggles with complex tasks requiring complicated language understanding and dexter- ous action generation [54]. Vision-Language-Action Model. The success of large- scale multi-modal models [1, 2, 10, 14, 43] has motivated an end-to-end method to generate generalist robot policies [3–6, 15, 28, 30, 35, 41], where a foundation model π is finetuned to directly predict an action chunkA t conditioned on the observation o t at time t, written as π : o t 7→ A t , A t = [a t ,a t+1 ,· ,a t+N ],(2) where the observation o t consists of the language instruc- tionsl, visual dataI t , and the proprioceptive statesq t of the robot, anda t is a continuous control signal. For a robot with d degrees of freedom, a t ∈R d . Pre-trained on multi-modal web data [9, 11, 19, 42, 48], and co-trained on several robot datasets [13, 27, 33, 44], this end-to-end model can directly map human instructions to robotic actions, and generalize to multiple scenes [3, 6, 12]. To improve planning ability and action accuracy, follow- ing methods tailor the model π to describe intermediate af- fordance o ∗ t , such as object bounding boxes [8, 50] or ma- nipulation keypoints [25]. These predictions serve as guid- ance for subsequent action generation, formalized as π = π 2 ◦ π 1 , π 1 : o t 7→ o ∗ t , π 2 : (o t ,o ∗ t )7→ A t , (3) However, the approach has several limitations. First, it re- lies on large-scale data with handcrafted affordance anno- tations that are hard to come by in robotics [13]. Second, it depends on high-quality observations and is vulnerable to visual perturbations which significantly degrade their per- formance in realistic scenarios [16]. Multi-modal Agents with Tool-use. Research on large language models demonstrates their ability as multi-modal agents to solve complex tasks with tool use [17, 18, 21, 39, 46, 49, 52, 53]. The agent leverages powerful language models to generate pseudocode [17, 21], scripts [39], or a API list [37] that invokes external tools in sequence. Given the model’s generalist reasoning ability and the external tools’ enhancement, agents show superior performance in the web search [52], image editing [21], and game play [53]. Recently, several works have incorporated tool-use into the end-to-end VLA framework. TIGeR [22] incorporates geo- metric computation tools into end-to-end action generation to enhance the VLA’s spatial perception and reasoning ca- pabilities, and VLA 2 [51] employs web knowledge retrieval to generalize the VLA model on tasks involving unseen tar- gets. These methods demonstrate the scalability and flex- ibility of combining the end-to-end architecture with on- the-fly external tools. However, they are limited on specific tasks, object identification [51] and 3D geometrics [22]. In contrast, ART takes on the challenge of tool use in long, multi-task trajectories, thereby facilitating the solution of a broader range of tasks with tool-integrated reasoning. 3. ART: Agentic Robot with Tool-use In this section, we present our framework for an agentic robot with tool use. We begin with the formulation of op- timization objective (Sec. 3.1). Next, we detail the system architecture designed for tool injection (Sec. 3.2), followed by a brief description of our training procedures (Sec. 3.3). 3.1. Optimization with Tool-use Reasoning We begin with the standard setup of a vanilla VLA model. At each timestep t, the model π receives a clear and high- quality observation o t as input, from which it predicts an embodied action a t ∈ A that governs the robot’s move- ment. However, in real-world scenarios, observations are often ambiguous and prone to various forms of corruption, which impedes the model to accurately learn the mapping from observations to embodied actions during training. The idea behind ART is simple: we augment the model’s action space toA ∗ = A×R×T , whereR represents the space for language reasoning, andT is the action space for tool usage. An action a r,t ∈ R aims to compose informa- tion by reasoning over the observations, while a t,t ∈ T is intended to activate external tools to enhance the observa- tions. Thus, a solution to a task is represented as a trajectory 3 Figure 2. Overview of the ART architecture. With a fine-tuned LoRA module, ART first reasons over raw observations and predicts tool use actions that activate external tools, thereby improving the input data. Specifically, the external tools provide vision, affordance, and embodiment enhancement. Given the enhanced observations, the model generates the final embodied actions. of augmented actions, denoted asa ∗ 1:T = (a ∗ 1 ,a ∗ 2 ,...,a ∗ T ), where each a ∗ t ∈ A ∗ . Consequently, the overall optimiza- tion objective becomes: π ∗ = arg max π P θ (a ∗ 1:T | ̃ o 1:T ),(4) where the model π is parameterized by θ. Our primary insight is that an accurate embodied ac- tion a t should be independent of the noise and corruptions in raw observations ̃ o t , and thus depend solely on the en- hanced observations o t . Thus, the overall optimization ob- jective is decomposed as π ∗ = arg max π T Y t=1 P θ,1 (a t |o 1:t )P θ,2 (a r,t ,a t,t | ̃ o 1:t ). (5) Therefore, we decompose the overall objective into two: P θ,1 , the training objective of a vanilla VLA, and P θ,2 , the fine-tuning objective of reasoning and tool use, which we refer to as tool injection. 3.2. Tool Injection Framework The decomposition in Sec. 3.1 allows us to train ART from fine-tuning an existing VLA model with an extra optimiza- tion objective. However, previous works prove that a na ̈ ıve coupling of training recipes significantly degrades the capa- bility of the backbone model and causes catastrophic for- getting [15, 54]. To avoid it and realize the factorization in Eq. 6, we introduce a non-destructive modification within the standard VLA architecture. Tool Token Injection. Unlike an embodied action a t that encodes a continuous control signal, the tool use action a t,t is discrete, as the state of each tool is naturally binary: ei- ther active or inactive. Given a tool function set F of size n, an action a t,t ∈ 0, 1 n , where each element a t,t,i indi- cates the state of the i-th tool function f i ∈F . This discrete representation enables us to map tool use actions to tokens in the vocabulary, allowing the model to treat tool activation as a form of discrete decision-making. By representing the tool states as tokens, we treat tool use similarly to language generation, where each token corresponds to the activation or deactivation of a specific tool function. To minimize the impact on the original vocabulary, following RT-2 [6] and OpenVLA [28], we leverage the last N tokens in the vo- cabulary for this purpose. Consequently, tool use tokens are seamlessly integrated into the standard next-token predic- tion training, with the cross-entropy loss calculated on the predicted actions. Adaptive LoRA Fine-tuning. To prevent interference be- tween embodied action generation and tool use, we adopt a fine-tuning strategy with a dynamically activated LoRA module [23]. Starting from a pre-trained VLA model, we freeze the backbone and fine-tune the LoRA module to en- hance tool-use reasoning. This decouples the two compo- nents: the VLA backbone without LoRA generates embod- ied actions, while with the LoRA module, the model gener- ates tool use actions. In the inference process, the model first generates reasoning and tool use actions (a r,t , a t,t ) based on observations, activating tools with LoRA; then, the enhanced observation o t is passed to the model to gen- erate the final embodied action a t . This independent opti- mization of tool use and action generation ensures efficient learning and allows ART to be easily integrated into various VLA architectures with minimal modifications. Tool-use with Action Chunk. Modern VLA architec- tures increasingly adopt an action-chunking strategy to im- prove efficiency, where the model predicts a chunk of H future embodied actions in one inference call, thereby re- 4 ducing the frequency of heavy model inference [28]. To ac- commodate this in the ART framework, we execute tool-use reasoning at every H timesteps: rather than invoking tool selection at every single control step, the model predicts reasoning tokens a r,t and tool-use decision a t,t once per chunk. Then the selected tools are activated for next H timestep observations ̃ o t:t+H . 3.3. Training Procedures We use the pre-trained 3B π-FAST [35] model. During training, we optimize two components, the VLA backbone, and the finetuning LoRA module. As is shown in Sec. 3.1, our training objective has two key parts: tool use reasoning token prediction and embodied action token prediction. Tool use reasoning token prediction. For tool use rea- soning, the model first predicts the discrete language to- kens of reasoning a r,t and generates a discrete set of tool use actions a t,t based on the current raw observations ̃ o t . This prediction is treated as a token selection problem and trained using next-token prediction loss, formulated as L tool =− T X t=1 logP θ L (a r,t ,a t,t | ̃ o 1:t ),(6) where θ L denotes parameters of the LoRA modules. Embodied action token prediction. For embodied action prediction, the model generates a sequence of embodied ac- tions a t in an auto-regressive manner, using the enhanced observations o t . Since FAST [35] discretizes the continu- ous signal using discrete cosine transform, the transformed action tokens are trained with auto-regressive next-token prediction. The loss function for embodied actions is given by: L action =− T X t=1 logP θ (a t |f t−1 ( ̃ o t )),(7) where θ represents the parameters of the VLA model without LoRA modules, and f t−1 is the activated function at the previous timestamp. This formulation allows for ef- ficient training, ensuring that each action token is predicted based on the enhanced observations, resulting in improved robot behavior. 4. Data Collection Compared to the Internet-scale multi-modal datasets used in other multi-modal model research, a common challenge in VLA research is the high cost of collecting robotic datasets. Following the approach of T3-Agent, we propose an exten- sion to the currently available VLA datasets, AT (Action with Tool), incorporating reasoning and tool use into the existing datasets. To the best of our knowledge, we are the first to introduce a VLA dataset with long-trajectory tool- use reasoning. Below, we will detail the definition of tools (4.1) and the process of long-trajectory generation (4.2). Statistics and the examples of our dataset will be listed in Appendix. 4.1. Multi-modal Tool Enhancement We define tools as modifications and enhancements to ob- servations. As shown in Eq. 2, the VLA model accepts three types of inputs: visual data I t , language instructions l, and the proprioceptive states q t of the robot. We cat- egorize the tools into visual, affordance, and embodiment enhancements, which correspond to the three input modali- ties: (1) Visual Enhancement. We define 10 types of visual enhancement tools, including low-light enhancement, de- noising, jitter correction, and deblurring. These tools assist the model in adapting to more robust and complex scenar- ios. (2) Affordance Enhancement. We provide the ART model with tools for depth estimation [47] and object detec- tion [36]. These tools help the model effectively determine affordance information. For example, when handling the task of identifying the ”farthest object,” the model can use affordance enhancement to pinpoint the specific target. (3) Embodiment Enhancement.We consider a simple robotic setup, including the movement of the head-mounted cam- era and the initial state of the robotic arm. As pointed out by LIBERO-Plus [16], the performance of models signif- icantly degrades when there is a shift in the viewpoint or initial state. We provide the robot with tools for camera rotation, zooming, and resetting the robot’s bodily state, al- lowing the model to learn how to use these tools to solve problems. 4.2. Long-Trajectory Tool Reasoning Generation In this subsection, we detail the process of generating long- trajectory tool-use reasoning tasks. The task generation pro- cess begins by creating a series of challenges that require the model to reason about its environment and the appro- priate tool to use. These challenges are designed using the three types of tool enhancements: visual, affordance, and embodiment, which correspond to different task types. Task Generation. In the task generation stage, the model is tasked with solving specific challenges that require tool use and reasoning. This includes three primary categories: Vision Reduction, where random visual degradations (such as noise or light transitions) are applied to simulate difficult environmental conditions, and the model uses visual en- hancement tools to compensate; Affordance Query, where the model must identify objects or properties, with the help of tools like depth estimation and object detection; and Em- bodiment Disturbance, where disturbances in the robot’s 5 Figure 3. Overview of ART dataset collection pipeline. The core idea is to fully utilize existing robotic dataset. (1) Task generation. For the vision and embodiment modalities, we randomly select some disturbances. For the language modality, we use GPT to transform simple tasks from the original dataset into reasoning tasks that require affordance information. (2) Tool chain generation. The tools needed to resolve disturbances and affordance reasoning are recorded into a reasoning chain. (3) Trajectory generation. A GPT is prompted to chain these models with the reasoning together. embodiment (e.g., shifts in viewpoint or initial state) are simulated, and the model uses tools like camera rotations or position resetting to address these changes. Tool Chain Generation. Once a task is generated, it is paired with the corresponding tools necessary to solve the task. This involves creating a tool chain, a sequence of tasks and the tools required to complete them. For example, a task might require a vision enhancement tool followed by an embodiment adjustment tool, creating a specific tool chain that the model must follow. Trajectory Generation. The final stage in the process is the generation of a long-trajectory reasoning process. Here, the task-tool pair is used to prompt a model (such as GPT) to generate a detailed reasoning trajectory. This trajectory explains why each tool is needed and how it helps complete the task. The model’s reasoning is then consolidated into a coherent, step-by-step process that forms a long trajectory, guiding the robot through each action and decision point. 5. Experiment The ART model is designed to generalize broadly to new environments, tasks, and actions with extendable tools. In this section, we evaluate the effectiveness of ART in utiliz- ing external tools through tool-use reasoning. 5.1. Experiment Setup In this section, we introduce our training settings and de- scribe the three benchmark environments used for model evaluation and comparison. These include a simulated en- vironment, a closed-loop real-world scenario, and an open- loop real-world testing setup. Training Settings. Following the previously designed Training Injection method, we fine-tuned the model for 1 epoch on a 30k tool trajectory dataset introduced in Sec. 4, starting from a pre-trained model. The fine-tuning was con- ducted on 8× A800 GPUs with a batch size of 24. An initial learning rate of 5×10 −5 was used, with a 1k-step warm-up period. Astribot S1. We evaluated our model on the Astribot S1 humanoid robot, a dual-arm robot with 16 degrees of free- dom. Initially, we trained a baseline FAST model on a 16k pick-and-place action trajectory dataset, which includes 80 different objects and 10 container types, following standard VLA training formats. AT Dataset (LIBERO). We tested our model in the LIBERO environment, a simulated platform that allows for modifications in visual factors such as lighting and noise. The precision of task instructions was also adjusted based on object positioning (e.g., changing ”place on the plate” to ”place 30 cm to the left of the drawer”). In addition, distur- 6 Figure 4. Rollouts of experiments. (a) The results ofπ 0 -FAST and ART on low-light environment. (b) The performance of ART on affordance reasoning task, whereas the robot controlled byπ 0 -FAST predicts no action tokens and remains static here. Tasks AT LIBERO (Success Rate)↑AT Astribot S1 (Success Rate)↑ VisionAffordanceEmbodimentAverageVisionAffordanceEmbodimentAverage OpenVLA20%10%7%12%5%5%10%6.7% π 0 65%15%10%30%40%30%40%37% π 0 -FAST60%12%45%39%40%30%60%43% ART-FAST81%62%82%75%70%55%70%62% Table 1. Performance improvement by on-the-fly tool use. We compare the success rate of generalized tasks between ART and different mainstream VLA models on different generalized tasks. The LIBERO experiments show that tool use significantly improve the model’s ability in complex scenarios. The Astribot S1 experiments show the tool reasoning ability tuned on AT can generalize to new robotic scene. bances were introduced to the initial positions of the camera and the robot’s body. These modifications provided a con- trolled environment for comparing the performance of ART with current SOTA baselines in simulation. 5.2. Performance of Tool Integration Table 1 shows the straightforward improvement brought by AT. As reported by LIBERO-Plus [16], state-of-the-art (SOTA) Vision-Language-Action (VLA) models struggle significantly with visual, semantic, and embodiment dis- turbances. These disturbances reduce their performance on complex tasks, especially when compared to general robotics datasets that contain high-quality, clean observa- tions. In our experiments, models like OpenVLA and π 0 exhibit poor performance, with average success rates sig- nificantly below 30% across both the LIBERO simulated environment and real-world Astribot S1 robot tasks. This demonstrates the difficulty of applying these models to real- world scenarios, where noise and variability in observations can dramatically affect the results. Performance of on-the-fly tool use. The key advantage of ART is its ability to integrate external tools on-the-fly, enabling it to adapt to a variety of complex tasks without the need for retraining the entire model. This tool-based reason- ing capability allows ART to address the disturbances that hinder the performance of traditional VLA models. As seen in Table 1, ART-FAST outperforms other models in both the LIBERO environment and the Astribot S1 robot tasks, with an average success rate of 75% and 62%, respectively. This significant improvement suggests that external tools can be leveraged effectively for real-time decision-making, allow- ing ART to handle a wider range of tasks, even in envi- ronments with noisy or incomplete observations. This ca- 7 pability demonstrates the potential of on-the-fly tool use in enhancing the generalization of VLA models to more com- plex and dynamic scenarios. Generalization of tool reasoning ability. Another no- table result is the ability of ART to generalize tool-use reasoning across multiple environments. While previous SOTA models struggle to generalize beyond their train- ing domains, ART shows robust performance in both the LIBERO simulation and real-world robotic tasks. The fine- tuning on the AT dataset, which includes data from LIBERO [33], Bridge v2 [44], and DROID [27], enables the model to effectively apply tools across different datasets, achiev- ing impressive results in new tasks without requiring task- specific re-training. This confirms that the tool reason- ing ability of ART is not only effective in controlled en- vironments but also transferable to real-world applications, where the robot must adapt to new objects, goals, and con- ditions. The ability to generalize this tool-use capability is a crucial step towards building more adaptable and robust robotic systems. 5.3. Performance of Offloading Model Tasks Affordance Tasks AT LIBERO (Success Rate)↑ w/ Vision Corruption w/o Vision Corruption Average ECoT12%58%36% ART-FAST72%62%67% Table 2. Performance improvement compared to the reason- only model. We test the success rate of ECoT and ART on the affordance part of AT, with or without vision corruptions. Comparison with End-to-End VLA Models. In this section, we compare the tool-use approach with traditional end-to-end VLA models. Specifically, we consider two types of baseline models: Action-only models (Fig. 1 (1a)) and Reason-only models (Fig. 1 (1b) & (2b)). Tasks AT LIBERO (Success Rate)↑ VisionEmbodimentAffordance FAST (Post-trained)71%61%65% ART-FAST81%62%82% Table 3. Performance improvement compared to end-to-end training. We test the success rate of ART and the finetuned FAST model with the same training consumption. Comparison with end-to-end learning. A standard VLA model typically relies on an end-to-end learning paradigm. This approach assumes that, given enough data, the model can learn the input-output mapping directly from raw obser- vations to actions through a data-driven process. To evalu- ate the effectiveness of ART in comparison to such models, we fine-tuned a FAST model with the same amount of train- ing data, but removed the intermediate tool-invocation rea- soning. In this case, the model was tasked with learning to generate embodied actions directly from raw observations. While the FAST model exhibited some improvement, it still lagged behind ART in terms of its ability to handle tasks using external tools. As shown in Table 3, ART signifi- cantly outperforms the post-trained FAST model, demon- strating that incorporating external tools allows for more robust task-solving and better generalization to unseen sce- narios. Comparison with ECoT. We also compared ART with another common reasoning-only model in VLA research, ECoT [50].As shown in Table 2, ART performed slightly better than ECoT when no vision corruption was intro- duced, validating that ART can handle affordance tasks more effectively when the input data is clean. ficant find- ing emerges when vision disturbances are added. ECoT’s performance significantly drops, as it was only trained on high-quality, clear data without exposure to such vi- sual distortions. This causes the model to encounter out- of-distribution data during inference, severely limiting its generalization capability. In contrast, ART maintains ro- bust performance by leveraging external vision enhance- ment tools, which help to mitigate the effect of these distur- bances. This approach improves task generalization with- out requiring retraining, thus offering a flexible solution to complex, dynamic environments. 6. Conclusion This paper presents ART (Agentic Robot with Tool-use), a fine-tuning framework that bridges modular and end-to-end VLA models. It enables VLA models to leverage external tools for novel tasks without compromising their precision or continuous action capabilities. Our method integrates multi-modal tool usage into existing VLA datasets, reduc- ing the need for new data collection, and employs a two- stage LoRA fine-tuning process to decouple tool reasoning from core VLA tasks, mitigating catastrophic forgetting. Experimental results in both the LIBERO simulation and real-world environments show that ART enhances model robustness and generalizability, enabling effective handling of new objects, scenarios, and action instructions. ART re- duces data dependency, paving the way for more adaptable robotic agents in dynamic environments. 8 References [1] Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Men- sch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikołaj Bi ́ nkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ́ en Simonyan. Flamingo: A Visual Language Model for Few-Shot Learn- ing. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 3 [2] Lucas Beyer, Andreas Steiner, Andr ́ e Susano Pinto, Alexan- der Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisensch- los, Rishabh Kabra, Matthias Bauer, Matko Bo ˇ snjak, Xi Chen, Matthias Minderer, Paul Voigtlaender, Ioana Bica, Ivana Balazevic, Joan Puigcerver, Pinelopi Papalampidi, Olivier Henaff, Xi Xiong, Radu Soricut, Jeremiah Harmsen, and Xiaohua Zhai. PaliGemma: A versatile 3B VLM for transfer, 2024. 3 [3] Kevin Black, Noah Brown, James Darpinian, Karan Dha- balia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, James Tanner, Quan Vuong, Homer Walke, Anna Walling, Haohuan Wang, Lili Yu, and Ury Zhilinsky.π0.5: A Vision-Language-Action Model with Open-World Generalization, 2025. 1, 3 [4] Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mo- hith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky.π0: A Vision-Language-Action Flow Model for General Robot Control. In Proceedings of Robotics: Science and System (RSS), 2025. 1, 2 [5] Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakr- ishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, Ju- lian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil J. Joshi, Ryan Julian, Dmitry Kalash- nikov, Yuheng Kuang, Isabel Leal, Kuang-Huei Lee, Sergey Levine, Yao Lu, Utsav Malla, Deeksha Manjunath, Igor Mordatch, Ofir Nachum, Carolina Parada, Jodilyn Peralta, Emily Perez, Karl Pertsch, Jornell Quiambao, Kanishka Rao, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Kevin Sayed, Jaspiar Singh, Sumedh Sontakke, Austin Stone, Clay- ton Tan, Huong Tran, Vincent Vanhoucke, Steve Vega, Quan Vuong, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich. RT-1: Robotics Transformer for Real-World Control at Scale, 2022. [6] Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakr- ishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Lisa Lee, Tsang-Wei Edward Lee, Sergey Levine, Yao Lu, Henryk Michalewski, Igor Mordatch, Karl Pertsch, Kan- ishka Rao, Krista Reymann, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Pierre Sermanet, Jaspiar Singh, Anikait Singh, Radu Soricut, Huong Tran, Vincent Vanhoucke, Quan Vuong, Ayzaan Wahid, Stefan Welker, Paul Wohlhart, Jialin Wu, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. In Proceed- ings of the Annual Conference on Robot Learning (CoRL). arXiv, 2023. 2, 3, 4 [7] Junting Chen, Yao Mu, Qiaojun Yu, Tianming Wei, Silang Wu, Zhecheng Yuan, Zhixuan Liang, Chao Yang, Kaipeng Zhang, Wenqi Shao, Yu Qiao, Huazhe Xu, Mingyu Ding, and Ping Luo. RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation, 2024. 1, 3 [8] William Chen, Suneel Belkhale, Suvir Mirchandani, Oier Mees, Danny Driess, Karl Pertsch, and Sergey Levine. Train- ing Strategies for Efficient Embodied Reasoning, 2025. 3 [9] Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedan- tam, Saurabh Gupta, Piotr Dollar, and C. Lawrence Zitnick. Microsoft COCO Captions: Data Collection and Evaluation Server, 2015. 3 [10] Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Jialin Wu, Carlos Riquelme Ruiz, Sebastian Goodman, Xiao Wang, Yi Tay, Siamak Shak- eri, Mostafa Dehghani, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang, Ceslee Montgomery, Paulina Pietrzyk, Marvin Rit- ter, A. J. Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li, Ibrahim Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Peter Steiner, Yang Li, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut. PaLI-X: On Scaling up a Multilingual Vision and Language Model, 2023. 3 [11] Xi Chen, Xiao Wang, Soravit Changpinyo, A. J. Piergio- vanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Has- san Akbari, Gaurav Mishra, Linting Xue, Ashish V. Thap- liyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhos- seini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme Ruiz, Andreas Peter Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut. PaLI: A Jointly-Scaled Mul- tilingual Language-Image Model. In Proceedings of the In- ternational Conference on Learning Representations (ICLR), 2023. 3 9 [12] Jaden Clark, Suvir Mirchandani, Dorsa Sadigh, and Suneel Belkhale.Action-Free Reasoning for Policy Generaliza- tion. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) Workshop, 2025. 3 [13] Open X.-Embodiment Collaboration, Abby O’Neill, Abdul Rehman, Abhinav Gupta, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khaz- atsky, Anant Rai, Anchit Gupta, Andrew Wang, Andrey Kolobov, Anikait Singh, Animesh Garg, Aniruddha Kem- bhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Ar- chit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakr- ishna, Ayzaan Wahid, Ben Burgess-Limerick, Beomjoon Kim, Bernhard Sch ̈ olkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu, Charlotte Le, Chelsea Finn, Chen Wang, Chenfeng Xu, Cheng Chi, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Driess, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter B ̈ uchler, Di- nesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Ed- ward Johns, Ethan Foster, Fangchen Liu, Federico Ceola, Fei Xia, Feiyu Zhao, Felipe Vieira Frujeri, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Ge Yan, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guangwen Yang, Guanzhi Wang, Hao Su, Hao-Shu Fang, Haochen Shi, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homanga Bharadhwaj, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Ra- dosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jae- hyung Kim, Jaimyn Drake, Jan Peters, Jan Schneider, Jas- mine Hsu, Jay Vakil, Jeannette Bohg, Jeffrey Bingham, Jef- frey Wu, Jensen Gao, Jiaheng Hu, Jiajun Wu, Jialin Wu, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jingyun Yang, Jitendra Malik, Jo ̃ ao Silv ́ erio, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kaiyuan Wang, Kanishka Rao, Karl Pertsch, Karol Haus- man, Keegan Go, Keerthana Gopalakrishnan, Ken Gold- berg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Lin, Kevin Zhang, Kiana Ehsani, Ki- ran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srini- vasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Lawrence Yunliang Chen, Lerrel Pinto, Li Fei-Fei, Liam Tan, Linxi ”Jim” Fan, Li- onel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itk- ina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma, Moo Jin Kim, Muhammad Zubair Irshad, Naoaki Kanazawa, Nicklas Hansen, Nicolas Heess, Nikhil J. Joshi, Niko Suen- derhauf, Ning Liu, Norman Di Palo, Nur Muhammad Mahi Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pan- nag R. Sanketi, Patrick ”Tree” Miller, Patrick Yin, Paul Wohlhart, Peng Xu, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Quan Vuong, Rafael Rafailov, Ran Tian, Ria Doshi, Roberto Mart ́ ın-Mart ́ ın, Rohan Baijal, Rosario Scalise, Rose Hen- drix, Roy Lin, Runjia Qian, Ruohan Zhang, Russell Men- donca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bus- tamante, Sean Kirmani, Sergey Levine, Shan Lin, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham Sonawani, Shubham Tulsiani, Shuran Song, Sichun Xu, Siddhant Hal- dar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Mat- sushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vikash Kumar, Vincent Vanhoucke, Vitor Guizilini, Wei Zhan, Wenxuan Zhou, Wol- fram Burgard, Xi Chen, Xiangyu Chen, Xiaolong Wang, Xinghao Zhu, Xinyang Geng, Xiyuan Liu, Xu Liangwei, Xuanlin Li, Yansong Pang, Yao Lu, Yecheng Jason Ma, Yejin Kim, Yevgen Chebotar, Yifan Zhou, Yifeng Zhu, Yilin Wu, Ying Xu, Yixuan Wang, Yonatan Bisk, Yongqiang Dou, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yue Cao, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yun- fan Jiang, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yu- taka Matsuo, Zehan Ma, Zhuo Xu, Zichen Jeff Cui, Zichen Zhang, Zipeng Fu, and Zipeng Lin. Open X-Embodiment: Robotic Learning Datasets and RT-X Models, 2025. 3 [14] Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duck- worth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. PaLM-E: An Embodied Multimodal Lan- guage Model. In Proceedings of the International Confer- ence on Machine Learning (ICML), 2023. 3 [15] Danny Driess, Jost Tobias Springenberg, Brian Ichter, Lili Yu, Adrian Li-Bell, Karl Pertsch, Allen Z. Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine.Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better, 2025. 1, 3, 4 [16] Senyu Fei, Siyin Wang, Junhao Shi, Zihao Dai, Jikun Cai, Pengfang Qian, Li Ji, Xinzhe He, Shiduo Zhang, Zhaoye Fei, Jinlan Fu, Jingjing Gong, and Xipeng Qiu. LIBERO- Plus: In-depth Robustness Analysis of Vision-Language- Action Models, 2025. 3, 5, 7 [17] Zhi Gao, Yuntao Du, Xintong Zhang, Xiaojian Ma, Wen- juan Han, Song-Chun Zhu, and Qing Li. CLOVA: A Closed- LOop Visual Assistant with Tool Usage and Update. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 3 [18] Zhi Gao, Bofei Zhang, Pengxiang Li, Xiaojian Ma, Tao Yuan, Yue Fan, Yuwei Wu, Yunde Jia, Song-Chun Zhu, and Qing Li. Multi-modal Agent Tuning: Building a VLM- Driven Agent for Efficient Tool Usage. In Proceedings of the International Conference on Learning Representations (ICLR), 2025. 2, 3 [19] Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Ba- 10 tra, and Devi Parikh. Making the V in VQA Matter: Ele- vating the Role of Image Understanding in Visual Question Answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 3 [20] Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary Object Detection via Vision and Language Knowledge Distillation. In Proceedings of the International Conference on Learning Representations (ICLR), 2022. 3 [21] Tanmay Gupta and Aniruddha Kembhavi. Visual Program- ming: Compositional Visual Reasoning Without Training. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2023. 3 [22] Yi Han, Cheng Chi, Enshen Zhou, Shanyu Rong, Jingkun An, Pengwei Wang, Zhongyuan Wang, Lu Sheng, and Shanghang Zhang. TIGeR: Tool-Integrated Geometric Rea- soning in Vision-Language Models for Robotics, 2025. 3 [23] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the International Conference on Learning Representations (ICLR), 2022. 4 [24] Siyuan Huang, Zhengkai Jiang, Hao Dong, Yu Qiao, Peng Gao, and Hongsheng Li.Instruct2Act: Mapping Multi- modality Instructions to Robotic Actions with Large Lan- guage Model, 2023. 3 [25] Yuheng Ji, Huajie Tan, Jiayu Shi, Xiaoshuai Hao, Yuan Zhang, Hengyuan Zhang, Pengwei Wang, Mengdi Zhao, Yao Mu, Pengju An, Xinda Xue, Qinghang Su, Huaihai Lyu, Xi- aolong Zheng, Jiaming Liu, Zhongyuan Wang, and Shang- hang Zhang. RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 3 [26] Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. MDETR - Mod- ulated Detection for End-to-End Multi-Modal Understand- ing. In Proceedings of the IEEE/CVF International Confer- ence on Computer Vision (ICCV), 2021. 3 [27] Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Bal- akrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, Peter David Fagan, Joey Hejna, Masha Itkina, Marion Lepert, Yecheng Jason Ma, Patrick Tree Miller, Jimmy Wu, Suneel Belkhale, Shivin Dass, Huy Ha, Arhan Jain, Abraham Lee, Youngwoon Lee, Marius Mem- mel, Sungjae Park, Ilija Radosavovic, Kaiyuan Wang, Al- bert Zhan, Kevin Black, Cheng Chi, Kyle Beltran Hatch, Shan Lin, Jingpei Lu, Jean Mercat, Abdul Rehman, Pan- nag R. Sanketi, Archit Sharma, Cody Simpson, Quan Vuong, Homer Rich Walke, Blake Wulfe, Ted Xiao, Jonathan Hee- won Yang, Arefeh Yavary, Tony Z. Zhao, Christopher Agia, Rohan Baijal, Mateo Guaman Castro, Daphne Chen, Qi- uyu Chen, Trinity Chung, Jaimyn Drake, Ethan Paul Fos- ter, Jensen Gao, David Antonio Herrera, Minho Heo, Kyle Hsu, Jiaheng Hu, Donovon Jackson, Charlotte Le, Yun- shuang Li, Xinyu Lin, Zehan Ma, Abhiram Maddukuri, Suvir Mirchandani, Daniel Morton, Tony Khuong Nguyen, Abigail O’Neill, Rosario Scalise, Derick Seale, Victor Son, Stephen Tian, Emi Tran, Andrew E. Wang, Yilin Wu, An- nie Xie, Jingyun Yang, Patrick Yin, Yunchu Zhang, Os- bert Bastani, Glen Berseth, Jeannette Bohg, Ken Gold- berg, Abhinav Gupta, Abhishek Gupta, Dinesh Jayaraman, Joseph J. Lim, Jitendra Malik, Roberto Mart ́ ın-Mart ́ ın, Sub- ramanian Ramamoorthy, Dorsa Sadigh, Shuran Song, Jia- jun Wu, Michael C. Yip, Yuke Zhu, Thomas Kollar, Sergey Levine, and Chelsea Finn.DROID: A Large-Scale In- The-Wild Robot Manipulation Dataset. In Proceedings of Robotics: Science and System (RSS) Workshop, 2024. 3, 8 [28] Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P. Foster, Pannag R. Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn.OpenVLA: An Open-Source Vision-Language-Action Model. In Proceed- ings of the Annual Conference on Robot Learning (CoRL), 2024. 1, 2, 3, 4, 5 [29] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. Segment Anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 3 [30] Xinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu, Jie Xu, Hongtao Wu, Chilam Cheang, Ya Jing, Weinan Zhang, Huaping Liu, Hang Li, and Tao Kong. Vision-Language Foundation Models as Effective Robot Imitators. In Pro- ceedings of the International Conference on Learning Rep- resentations (ICLR), 2024. 3 [31] Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as Policies: Language Model Programs for Embodied Con- trol. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), 2023. 2, 3 [32] Kevin Lin, Christopher Agia, Toki Migimatsu, Marco Pavone, and Jeannette Bohg. Text2Motion: From natural language instructions to feasible plans. Autonomous Robots, 2023. 3 [33] Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. LIBERO: Benchmarking Knowl- edge Transfer for Lifelong Robot Learning. In Advances in Neural Information Processing Systems (NeurIPS), 2023. 3, 8 [34] Yao Mu, Junting Chen, Qing-Long Zhang, Shoufa Chen, Qiaojun Yu, Chongjian Ge, Runjian Chen, Zhixuan Liang, Mengkang Hu, Chaofan Tao, Peize Sun, Haibao Yu, Chao Yang, Wenqi Shao, Wenhai Wang, Jifeng Dai, Yu Qiao, Mingyu Ding, and Ping Luo. RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis. In Proceed- ings of the International Conference on Machine Learning (ICML), 2024. 1, 2, 3 [35] Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine.FAST: Efficient Action Tokenization for Vision-Language-Action Models.In Proceedings of Robotics: Science and System (RSS), 2025. 1, 3, 5 11 [36] Tianhe Ren, Yihao Chen, Qing Jiang, Zhaoyang Zeng, Yuda Xiong, Wenlong Liu, Zhengyu Ma, Junyi Shen, Yuan Gao, Xiaoke Jiang, Xingyu Chen, Zhuheng Song, Yuhong Zhang, Hongjie Huang, Han Gao, Shilong Liu, Hao Zhang, Feng Li, Kent Yu, and Lei Zhang. Dino-x: A unified vision model for open-world object detection and understanding, 2024. 5 [37] Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. In Ad- vances in Neural Information Processing Systems (NeurIPS), 2023. 3 [38] Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. ProgPrompt: Program generation for sit- uated robot task planning using large language models. Au- tonomous Robots, 2023. 3 [39] D ́ ıdac Sur ́ ıs, Sachit Menon, and Carl Vondrick. ViperGPT: Visual Inference via Python Execution for Reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2023. 3 [40] Huajie Tan, Xiaoshuai Hao, Cheng Chi, Minglan Lin, Yaoxu Lyu, Mingyu Cao, Dong Liang, Zhuo Chen, Mengsi Lyu, Cheng Peng, Chenrui He, Yulong Ao, Yonghua Lin, Peng- wei Wang, Zhongyuan Wang, and Shanghang Zhang. Ro- boOS: A Hierarchical Embodied Framework for Cross- Embodiment and Multi-Agent Collaboration, 2025. 3 [41] Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Lawrence Yunliang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An Open-Source Generalist Robot Policy. In Proceedings of Robotics: Science and System (RSS), 2024. 3 [42] Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai C. Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann LeCun, and Saining Xie. Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs. In Ad- vances in Neural Information Processing Systems (NeurIPS), 2024. 3 [43] Hugo Touvron, Louis Martin, Kevin Stone, Peter Al- bert, Amjad Almahairi, Yasmine Babaei, Nikolay Bash- lykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fer- nandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Vik- tor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Ko- renev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizen- stein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiao- qing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open Foundation and Fine- Tuned Chat Models, 2023. 3 [44] Homer Rich Walke, Kevin Black, Tony Z. Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen-Estruch, An- dre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abra- ham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine. BridgeData V2: A Dataset for Robot Learning at Scale. In Proceedings of the Annual Conference on Robot Learning (CoRL), 2023. 3, 8 [45] Mengdi Xu, Peide Huang, Wenhao Yu, Shiqi Liu, Xilun Zhang, Yaru Niu, Tingnan Zhang, Fei Xia, Jie Tan, and Ding Zhao. Creative Robot Tool Use with Large Language Mod- els, 2023. 1, 3 [46] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the International Conference on Learning Representations (ICLR), 2023. 3 [47] Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chunhua Shen. Met- ric3D: Towards Zero-shot Metric 3D Prediction from A Sin- gle Image. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 5 [48] Qiying Yu, Quan Sun, Xiaosong Zhang, Yufeng Cui, Fan Zhang, Yue Cao, Xinlong Wang, and Jingjing Liu. CapsFu- sion: Rethinking Image-Text Data at Scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), 2024. 3 [49] Lifan Yuan, Yangyi Chen, Xingyao Wang, Yi Fung, Hao Peng, and Heng Ji. CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets. In Proceedings of the International Conference on Learning Representations (ICLR), 2024. 3 [50] Michał Zawalski, William Chen, Karl Pertsch, Oier Mees, Chelsea Finn, and Sergey Levine. Robotic Control via Em- bodied Chain-of-Thought Reasoning. In Proceedings of the Annual Conference on Robot Learning (CoRL), 2024. 1, 2, 3, 8 [51] Han Zhao, Jiaxuan Zhang, Wenxuan Song, Pengxiang Ding, and Donglin Wang. VLAˆ2: Empowering Vision-Language- Action Models with an Agentic Framework for Unseen Con- cept Manipulation, 2025. 3 [52] Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. GPT-4V(ision) is a Generalist Web Agent, if Grounded. In Proceedings of the International Conference on Machine Learning (ICML), 2024. 3 [53] Sipeng Zheng, Jiazheng Liu, Yicheng Feng, and Zongqing Lu. Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds. In Proceedings of the International Conference on Learning Representations (ICLR), 2024. 3 [54] Yifan Zhong, Fengshuo Bai, Shaofei Cai, Xuchuan Huang, Zhang Chen, Xiaowei Zhang, Yuanfei Wang, Shaoyang Guo, Tianrui Guan, Ka Nam Lui, Zhiquan Qi, Yitao Liang, Yuanpei Chen, and Yaodong Yang. A Survey on Vision- Language-Action Models: An Action Tokenization Perspec- tive, 2025. 3, 4 12