Paper deep dive
Adversarial Attacks on Robotic Vision Language Action Models
Eliot Krzysztof Jones, Alexander Robey, Andy Zou, Zachary Ravichandran, George J. Pappas, Hamed Hassani, Matt Fredrikson, J. Zico Kolter
Models: OpenVLA, Pi-0, RT-1, RT-2
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/12/2026, 6:54:33 PM
Summary
This paper investigates the vulnerability of Vision-Language-Action (VLA) models to adversarial attacks. The authors adapt Greedy Coordinate Gradient (GCG) jailbreaking techniques from Large Language Models (LLMs) to gain control over robotic systems. They demonstrate that textual adversarial prompts can force VLAs to execute targeted actions with high success rates, persist across multiple rollout steps, and transfer across different robotic environments, highlighting a critical security gap in current robotic foundation models.
Entities (5)
Relation Signals (3)
GCG → isadaptedfor → VLA
confidence 100% · Our main algorithmic contribution is the adaptation and application of LLM jailbreaking attacks to obtain complete control authority over VLAs.
OpenVLA → istrainedon → LIBERO
confidence 95% · Across various fine-tunes of OpenVLA on distinct tasks from the LIBERO dataset
VLA → isvulnerableto → Adversarial Attacks
confidence 95% · We find that textual attacks... facilitate full reachability of the action space of commonly used VLAs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The emergence of vision-language-action models (VLAs) for end-to-end control is reshaping the field of robotics by enabling the fusion of multimodal sensory inputs at the billion-parameter scale. The capabilities of VLAs stem primarily from their architectures, which are often based on frontier large language models (LLMs). However, LLMs are known to be susceptible to adversarial misuse, and given the significant physical risks inherent to robotics, questions remain regarding the extent to which VLAs inherit these vulnerabilities. Motivated by these concerns, in this work we initiate the study of adversarial attacks on VLA-controlled robots. Our main algorithmic contribution is the adaptation and application of LLM jailbreaking attacks to obtain complete control authority over VLAs. We find that textual attacks, which are applied once at the beginning of a rollout, facilitate full reachability of the action space of commonly used VLAs and often persist over longer horizons. This differs significantly from LLM jailbreaking literature, as attacks in the real world do not have to be semantically linked to notions of harm. We make all code available at this https URL .
Tags
Links
Trouble viewing inline? Open PDF directly →
Full Text
77,739 characters extracted from source content.
Expand or collapse full text
arXiv:2506.03350v1 [cs.RO] 3 Jun 2025 Adversarial Attacks on Robotic Vision-Language-Action Models Eliot Krzysztof Jones 1∗ Alexander Robey 1,2 Andy Zou 1,2 Zachary Ravichandran 3 George J. Pappas 3 Hamed Hassani 3 Matt Fredrikson 1,2 J. Zico Kolter 1,2 1 Gray Swan AI 2 Carnegie Mellon University 3 University of Pennsylvania Abstract The emergence of vision-language-action models (VLAs) for end-to-end control is reshaping the field of robotics by enabling the fusion of multimodal sensory inputs at the billion-parameter scale. The capabilities of VLAs stem primarily from their architectures, which are often based on frontier large language models (LLMs). However, LLMs are known to be susceptible to adversarial misuse, and given the significant physical risks inherent to robotics, questions remain regarding the extent to which VLAs inherit these vulnerabilities. Motivated by these concerns, in this work we initiate the study of adversarial attacks on VLA-controlled robots. Our main algorithmic contribution is the adaptation and application of LLM jailbreaking attacks to obtain complete control authority over VLAs. We find that textual attacks, which are applied once at the beginning of a rollout, facilitate full reachability of the action space of commonly used VLAs and often persist over longer horizons. This differs significantly from LLM jailbreaking literature, as attacks in the real world do not have to be semantically linked to notions of harm. We make all code available athttps://github.com/eliotjones1/robogcg. 1 Introduction The emergence of robotic foundation models (RFMs) has transformed the field of robotics, driving progress in domains as diverse as robot-assisted surgery [1,2], autonomous driving [3,4], and agri- culture [5,6]. Rapid industry progress has resulted in the mass-production of commercially available AI-enabled robots [7,8]. Moreover, production-ready systems such as Physical Intelligence’sπ 0 model [9] and Google’s fleet of Gemini-controlled robots [10,11] excel at dynamic manipulation and multi-agent coordination [12]. Taken together, this accelerating landscape of capable RFMs has added to a growing belief: AI-enabled robots will soon collaborate in society alongside humans. Motivated in part by the growing deployment of RFMs and analogous agentic systems in real-world settings, the AI safety community has begun to anticipate new risks posed by stronger capabilities [13– 15]. Traditional AI security efforts have primarily focused on model-level threats (e.g., prompt injection [16–18] and jailbreaking [19–21]). These attacks target a model’s outputs, but stop short of obtaining control over the larger system governing reasoning, long-term planning, and actuation. In contrast, emerging concerns—such as deceptive alignment [22–24] and self-replication [25,26]— anticipate risks that may arise as models become increasingly autonomous and agentic. In line with this agenda, we consider the possibility that AI-enabled robots will one day interact with humans in open-world environments. While such systems are not yet deployed at scale, their rapid progress suggests that the risks surrounding these models may soon become highly relevant—and we ∗ Correspondence toeliot@grayswan.ai Preprint. Under review. Figure 1:Adversarial attacks on VLAs.VLA architectures fuse input images and textual task descriptions to produce low-level actuation. In this paper, we show that we can subvert the actions produced by an unattacked VLA (left) by adversarially attacking the textual prompt, resulting in the elicitation of a targeted action or sequence of actions (right). believe that it is critical to understand thembeforedeployment becomes widespread. To this end, in this paper, we initiate the study of adversarial attacks on vision-language-action (VLA) models. Existing VLAs cast robotic control through the lens of autoregressive prediction by fusing textual and visual inputs [9,27,28]. While recent works on jailbreaking RFMs [29,30] develop notions of semantic safety analogous to traditional, language-based alignment, the attacks we propose are designed to obtaincomplete control authority—i.e., the ability to continuously drive a VLA-controlled robot to any targeted action regardless of its sensory inputs or prior training—over a targeted VLA via textual prompting. Whereas LLM alignment tends to block non-adversarial generations of harmful actions by robotic planners, there is no analogous notion of refusal training or preference optimization for VLAs, which is corroborated by our finding that token-based attacks are more effective and efficient when applied to VLAs relative to chatbots. These differences contribute to a distinct attack landscape for VLAs, which we characterize in this paper. Contributions.Our main contributions are as follows. •VLA threat models.We identify realistic threat models for VLA-controlled robots, which concern the elicitation of targeted robotic actions via textual prompting. • VLA attack algorithms.We propose a family of token-level attacks that elicit fixed actions or sequences of actions from targeted VLA-controlled robots. •Targeted action elicitation. We show that VLA action spaces are reachable, in that adversarial prompting suffices to elicit nearly any targeted action. Across various fine-tunes of OpenVLA on distinct tasks from the LIBERO dataset, we repeatedly achieve upwards of 90% success rates at eliciting targeted actions. •Attack persistence. We show that our attacks tend to persist across multiple VLA rollout steps even as the model observes new visual inputs. Relative to nominal operation, our attacks increase the number of targeted persistence rollout steps by up to 28×. •Universality.We demonstrate that our attacks can beenvironment-agnostic, meaning that they can be successfully deployed across multiple robotic environments, both in simulation and in the real world. 2 Related work 2.1 Foundation models for robotic applications Over the past decade, advances in deep learning have been responsible for remarkable progress in robotics. Early efforts at the intersection of these two areas centered on training end-to-end policy networks from scratch [31,32], although progress slowed given computational costs coupled with the challenges inherent to generalizing to unseen robotic tasks [33] and fusing together multiple data modalities. However, the rise of transformer-based architectures, which capture long-term dependen- cies in sequential data [34], has rejuvenated the field of robotic control. As these technologies have matured, two dominant paradigms of AI-enabled robotic control have emerged: 1.High-level planners.RFMs employed to control a robot via a pre-defined API containing high-level primitives (e.g., “walk_forward” or “find_object”). 2 2.Low-level actuators.RFMs employed to output sequences of low-level control actions, such as regulating the torques and velocities of robotics arms. In the remainder of this subsection, we further describe these two paradigms of robotic control. High-level planners.High-level planners have gained popularity due to their versatility, meaning that they are designed to be drop-in components within existing pipelines [35]. Initial inroads include code-as-policies and related algorithms [36–38], which facilitate short-horizon interactions between a robot and an LLM chatbot instructed to produce API code. Similarly, both unimodal models (e.g., LLMs) and multimodal models (e.g., VLLMs) have been successfully deployed as robotic planners in domains spanning self-driving cars [39,40], service robots [41,42], and robot-assisted surgery [1]. Planners that combine pre-trained language and vision backbones into a unified architecture, followed by domain-specific fine-tuning, have also shown promise, albeit at a slightly higher computational cost. For instance, PaLM-E combines a PaLM backbone [43] with vision transformers [44] to facilitate high-level planning and reasoning about real-world environments [45]. Similarly, both [46] and [47] improve robotic task planning by incorporating 3D scene graphs. Finally, some high-level planners—LEO [48] and SayCan [49]—rely on the use of low-level features for end-to-end planning. Low-level actuators.Low-level actuators, often termed vision-language-action models (VLAs), are trained to generate continuous actions given textual goal descriptions and visual inputs. More traditional architectures, such as Octo [50], use transformers to map embeddings to actions at smaller parameters scales. More powerful VLAs utilize pre-trained language models as their backbone, with prominent VLA architectures including Google’s RT-1, RT-2, and RT-X models, each of which built upon its predecessor by training on larger and more diverse datasets to advance toward a generalist robotic policy [28,51,52], and the best open-source alternative, OpenVLA [27]. Notably, Physical Intelligence recently proposed a new flow-matching VLA calledπ 0 [9]. This model uses a pre-training/post-training recipe along with a so-called “action expert,” which is inspired by the ubiquitous mixture-of-experts paradigm [53], to improve generalization. Separately, diffusion-based action architectures such as CogACT [54] have also shown recent promise due to their ability to effectively capture the continuous nature of robotic actions in the physical world. 2.2 Adversarial attacks and defenses A central goal of the AI safety community is to understand and mitigate the potential misuse of AI and AI-enabled systems. At the heart of these efforts is the belief that the actions taken by AIs should align with human values [55–57]. Given the broad scope of this goal, early efforts primarily targeted immediate sources of misalignment, such as the generation of harmful content [58,59]. More recent research, however, has expanded focus toward anticipating the long-term risks associated with deploying highly capable AI-powered agents [60–62]. Special attention has also been paid to regulating the use of frontier AIs, particularly as they are used into human-facing domains [63, 64]. Increased interest in AI safety has led to a broad array of technical methods that asses the propensity of AIs to cause harm. Much of this research has focused on jailbreaking attacks on large language models (LLMs) and vision-enabled LLMs (VLLMs), wherein the goal is to elicit objectionable text [19,20,65] or toxic visual media [66]. And while numerous models are known to remain susceptible to state-of-the-art attacks [67,68], existing defenses [69,70], which are often informed by third-party red-teaming efforts, have contributed to a relatively robust suite of frontier models [71,72]. More recently, researchers have designed attacks to probe the vulnerabilities of AIs deployed for specific, downstream tasks, such as web-based agents [16,73,74] and AI-powered search engines [75]. Most related to our study is the recent work of Robey et al.[29], which demonstrates that LLM-based high-level planners are susceptible to jailbreaking attacks. Concurrent studies have corroborated this finding by showing that rephrasing instructions can lead to dangerous robotic actions [30,76]. However, to the best of our knowledge, our study is the first to consider attacks on low-level VLAs. 3 Jailbreaking attacks on VLAs To anticipate how VLA-integrated systems might enable misuse or unsafe behavior in future deploy- ments, we next seek to formalize a set of plausible, yet forward-looking threat models targeting VLAs. Our approach is grounded in the evolving literature on jailbreaking attacks, in which adversaries seek 3 to elicit objectionable responses from chatbots. After reviewing several preliminaries, we show that these attacks can be adapted to obtain complete control authority over a targeted VLA. 3.1 Jailbreaking LLM chatbots We start by reviewing the greedy coordinate gradient (GCG) chatbot jailbreaking attack [20], which underpins our approach to attacking VLAs. Given a goal stringG(e.g., “Tell me how to build a bomb”), the objective of GCG is to elicit a response from a targeted LLM that begins with a concomitant targetTstring (e.g., “Sure, here is how to build a bomb”). And because directly prompting the model withGmay result in a refusal (e.g., “I’m sorry, I cannot help you with that”), GCG’s threat model permits an attacker to modifyGby appending a fixed-length suffixS. In this way, whereas passingGas an input prompt may result in a refusal, the expectation is that prompting the model with[G;S](which denotes the concatenation ofGandS) will result in a jailbroken response. For more detailed derivations of the objective, please refer to equations (1)–(4) in [20]. 3.2 Threat models for VLAs Unlike LLMs, VLAs fuse two distinct sources of input: a textual prompt describing a robotic task, and an image showing the robot’s current scene. To roll out a VLA-based policy, the user supplies an initial prompt—which is fixed for all steps—and the VLA captures an image of its surroundings— which is updated at each step. These two inputs are concatenated in a joint embedding space, passed through an LLM backbone, and then processed by a bespoke action detokenizer. To detokenize actions, architectures tend to use an approach known as “symbol tuning,” wherein a setA⊊V containing the least used tokens in the LLM backbone’s vocabulary are identified with points in discretized version of the robot’s action space. In general, these values are uniformly distributed between the1 st and99 th quantiles seen in the training dataset for each degree of freedom [27, 51]. A forward pass through a VLA constitutes the autoregressive generation ofdtokens fromA, whered denotes the robot’s number of degrees of freedom. We write the probability of a VLA generating a length-dsequencex n+1:n+d fromAgiven the concatenated textx 1:n and image embeddingszas Pr [x n+1:n+d |x 1:n ;z] = Y d j=1 Pr[x n+j |x 1:n j −1 ;z].(1) To parallel Zou et al.[20], we consider attacks that aim to elicit a targeted action or sequence of actions. Given the notation in(1), one could consider two possible attack surfaces: the task description and the input image, both of which can be attacked in their semantic spaces (i.e., language or image pixels) or their representation spaces (i.e., textual or image embeddings). In this paper, we consider a threat model in which the adversary can modify the textual prompt, either by adding tokens to the end of a nominal instruction, or else replacing the prompt with an adversarially chosen sequence of tokens. We anticipate extending this threat model to include vision-based attacks in future work. Implications of this threat model.This threat model reframes safety in VLA-integrated systems as a matter ofcontrol authority, rather than harm-centric definitions typically associated with jailbreaking. That is, unlike traditional chatbot jailbreaks that elicit dangerous responses, our attacks aim to grant an adversary effective control over a robot’s low-level actions via input prompt manipulation. This perspective avoids the ambiguity of labeling individual actions as “harmful,” since identical actions may be safe in one context and dangerous in another. In other words, a robust VLA should resist adversarial takeover and simultaneously ensure that, even under adversarial control, generated actions should remain within or close to the distribution of actions seen during training. 3.3 Adversarial attacks on VLAs Having restricted our attention to attacks on a targeted VLA’s textual embeddings, we now seek an efficient, performant attack algorithm to stress test their robustness. Throughout, we consider an analogous loss function to the loss defined in equation (3) of [20], with the only difference being the additional image embedding input: ℓ(x 1:n ;z j )≜−log Pr [x n+1:n+d |x 1:n ;z j ].(2) Here,x n+1:n+d denotes the targeted action. 4 Table 1:Single step attacks.We report the attack success rates of the single step attack on four variants of OpenVLA, each of which is fine-tuned on a different subset of the LIBERO benchmark. We consider a sparse gridding of the action space for each model: For each model and each of the seven action dimensions, we consider one-hot targets for each of the 256 discrete bins, resulting in 256×7 = 1792distinct target actions per model. This table reports the per-dimension success rates for these one-hot targets, as well as the overall success rate, which requires the elicitation of each of the seven dimensional targets simultaneously. Model Per-dimension success rateOverall success rate Avg. computation per success 0123456Optim. stepsTime (sec.) Libero-Goal98.198.598.398.798.198.596.696.553.2304.6 Libero-Object98.297.898.398.697.297.093.793.873.6461.6 Libero-Spatial99.398.399.399.497.998.497.797.532.8185.2 Libero-1093.890.891.791.092.094.077.477.3109.7604.3 Single-step attacks.We first consider single-step attacks, which target the generation of a single fixed action. The performance of such attacks speak to the “reachability” of a VLA’s action space, in the sense that single-step attack algorithms seek to determine whether there exists an input prompt that will drive a VLA to a specific, targeted action. We operationalize single-step attacks by adapting the GCG algorithm introduced in §3.1 and [20] to the setting of VLAs. Specifically, consider the following optimization problem: minimize x i ∈V:i∈I ℓ(x 1:n ;z).(3) Here,zis the image embedding from the first rollout step. By taking the length-dtarget action x n+1:n+d as being analogous to the target stringTin §3.1, we directly adapt GCG for VLAs. Persistence attacks.We next consider a more sophisticated attack in which the attacker’s goal is to cause an action topersistfor a longer horizon. That is, the attack should elicit a targeted action across VLA inference steps despite evolving image representations. We implement this idea by modifying the objective in (3) to encourage invariance to the image representations: minimize x i ∈V:i∈I X r j=1 ℓ(x 1:n ;z j ).(4) Here, the objective is aggregated over the losses corresponding tordistinct image embeddingsz j . Obtaining the image embeddingsz j can be accomplished in various ways (e.g., performing data augmentation on the first-step image or collecting multiple random initializations). We compare the efficacy of different strategies in the experiments in §4.2. Transfer attacks.GCG is a white-box attack, meaning that it requires access to the weights of the target model to craft jailbreaks. Therefore, assessing the robustness of closed-weight chatbots (e.g., OpenAI’s o1 or Anthropic’s Claude models) via GCG necessitates the paradigm oftransfer, wherein attack strings are optimized on an open-weight source model and then inputted into a closed-weight model. Given the effectiveness of transfer in the LLM setting, we also consider such attacks in the context of VLAs. Specifically, when transferring attacks between VLAs, we first solve(3)on one or more source models and then apply the corresponding attack string to a distinct target model. 3.4 Implementation details VLA templates.As described in §3.3, GCG appends a suffixSto the nominal instructionG. In this paper, we take two approaches to inserting the adversarially-chosen tokensx i fori∈Vinto the VLA prompt. In the nominal case, text is inputted into a VLA using the following template: In: What action should the robot take to [INSTRUCTION]? Out: where[INSTRUCTION]is replaced with a short piece of text (e.g., “pick coke can”). In practice, we consider attacks with and without this nominal instruction; when we include the instruction, the adversarial string is appended at the end. An example attack (highlighted in red) is as follows: 5 In: What action should the robot take to bra x pill tin door f=db Eric swe bottom left m N x xtheless in extension x? Out: Normalization statistics.In general, VLA architectures are fine-tuned on downstream task datasets [1]. As a result, each model we consider has a distinct set of statistics that normalize the model’s predicted actions based on the distribution of actions seen during training. Thus, although the discretized actions are generally mapped to[−1,1], when constructing the target tokensx n+1:n+d , we use the normalization statistics to normalize the action for its particular, task-specific environment. 3.5 Attacking chatbots versus VLAs While the threat models and algorithms discussed in this section are adapted from the chatbot jailbreak literature, the VLA setting admits several key differences. Firstly, as the severity of a jailbroken response can be subjective, the performance of chatbot jailbreaking is heavily dependent on the choice of the evaluation judge (c.f., [19, Table 1]). In contrast, attacks on VLAs do not require a judge. Success is evaluated solely on whether the attack elicits the numerical target action, which is more reminiscent of more attacks in the literature surrounding adversarial examples [77,78]. Another byproduct of this difference is that semantic jailbreaks—e.g., prompts that embody human personas, invent new contexts, or mask harmful words [19,79]—are less applicable to VLAs than to chatbots. A second key difference lies in the role of model alignment. In the context of chatbots, the difficulty of jailbreaking is tightly coupled to the strength of safety-oriented post-training: models with more robust internal representations (see, e.g., [69]) are significantly harder to jailbreak than those with less involved post-training recipes. However, for VLAs, these internal representations are less relevant. Because VLA outputs correspond to low-level actuation, it is less meaningful to “align” them to a semantic notion of safety, especially given that the interaction between a robot’s environment and generated actions is more critical when determining overall safety. As such, we focus not on semantic notions of harm, but on the adversary’s ability to gain control authority: the capacity to drive the robot to a specific target action, independent of what that action means or whether it is harmful. 4 Experiments In this section, we evaluate the adversarial attacks proposed in §3.3 across a range of VLA architec- tures. In keeping with the norms in the VLA literature, all of the architectures that we consider target the control of a seven degree-of-freedom robotic arm with an attached gripper. Each action dimension is discretized into 256 distinct bins, and thus each action space comprises7 256 distinct actions. 4.1 Single-step attacks Given the effectiveness of VLAs fine-tuned on downstream task data, we begin our evaluation with four fine-tuned versions of OpenVLA [1], the most widely used open-source VLA. Each variant is fine-tuned on a different Libero subset: Libero-Goal, Libero-Object, Libero-Spatial, and Libero-10. To evaluate the single-step attack introduced in §3.3, we consider a sparse gridding of the action space comprising all7×256 = 1792one-hot target vectors. This is motivated both by the combinatorial size of the full action space and the tendency for actions containing many nonzero dimensions to be physically unrealizable or out-of-distribution. In Table 1, we report two metrics: (1) the overall success rate, which requires that each of the seven dimensions match the target action, and (2) the per-dimension success rate, which measures the success rate for each of the seven dimensions individually. We find that the Libero-Goal, Libero-Object, and Libero-Spatial models all achieve well above 90% overall success rates, whereas Libero-10 achieves a slightly reduced 77.4% success rate. This indicates that adversarial prompting is sufficient to drive a VLA to nearly any targeted action. Efficiency analysis.In keeping with the original implementation of GCG [20], we run the single step attacks for a maximum of 500 steps; the algorithm terminates if an exact match for every dimension in the target is found.The rightmost columns in Table 1 indicate that successful matches are found in between 30-110 steps, which stands in contrast to the chatbot jailbreaking literature, wherein jailbreaks often require optimization for all 500 steps. 6 Burn-inRollout 0.0 1.0 2.0 3.0 4.0 Persistence steps Libero-Goal Burn-inRollout Libero-Object Burn-inRollout Libero-Spatial Burn-inRollout Libero-10 Number of attack images 123 Persistence type SeenUnseenNominal Figure 2:Persistence attacks.For each of the four OpenVLA fine-tunes considered in Table 1, we measure the tendency of the persistence attack outlined in §3.3 to elicit a targeted action over the course of a full rollout. We run this attack withr∈1,2,3images in the objective in(4). Each bar is shaded to indicate whether a persistence step corresponded to an image seen while solving(4), or else corresponded to an unseen image at a later point in the rollout. Thex-axis denotes whether ther seed images were taking from the a “burn-in” period before the rollout begins—during which we actuate via randomly selected actions—or else from the firstrsteps of the rollout. And finally, the red dashed line denotes the frequency with which 50 non-attacked rollouts elicit the targeted action. 51015 0 20 40 60 80 100 Success Rate (%) Success Rate by Model 51015 Number of Tokens 0 20 40 60 80 Average Steps per Success Computation Steps by Model GoalObjectSpatial10 Figure 3:Token budget ablation.We observe that as the attacker’s token bud- get increases, the success rate also tends to increase. However, there is not a clear correlation between the token budget and the average number of steps per success. On H100 GPUs, this translates to between 3-10 minutes per success on average depending on the model, in contrast to an average of over an hour on the same hardware to generate chatbot jailbreaks. The attacker’s token budget.Another component of the computation complexity of single step attacks is the number of tokens|I|the adversary can manipulate. In Figure 3, we run single step attacks on each of the four fine-tuned models for|I| ∈ 5,10,15; the full results in Table 1 correspond to|I|= 20. This figure indicates that success rate tends to improve as the adversary’s token budget increases, and similarly, as the budget increases, the number of steps required to find a match tends to decrease. Relative to Table 1, the middle bar ground indicates that halving the budget also serves to reduce the success rate by at least a factor of two. This indicates that an attacker can reliably obtain stronger control authority over a targeted VLA by increasing its token budget. Attacks are environmentally agnostic.The so-called “sim-to-real” gap in robotics is a phenomenon whereby policies trained in simulated environments struggle to gen- eralize to real-world environments. This challenge extends to VLA policies, which—despite improved generalization from pretraining on large-scale embodied data—still exhibit performance gaps when transferred to real-world settings. To assess how well our attacks optimized in a simulated environment transfer to real-world settings, we evaluate single-step attacks on two environments from the Open-X- Embodiment [52] set that OpenVLA was trained on: HYDRA [80], a real-world environment, and SIMPLER [81], a simulated environment. As shown in Table 2, our attack is successful across both of these environments, indicating that such attacks also yield control authority in more realistic, open-world settings. 7 Table 2:Attacks on real-world images.We find that our attacks exhibit relatively strong perfor- mance when optimized on images drawn from SIMPLER, a simulated environment, and SIM- PLER, a real-world environment. DimensionHYDRA (%)SIMPLER (%) 093.450.4 186.048.1 273.647.3 385.148.1 488.445.0 588.448.8 663.641.9 Overall ASR61.238.0 0.00.51.01.52.02.5 Original action distance to target 0.0 0.5 1.0 1.5 2.0 2.5 Output action distance from target after optimization Both Failed: 320 Opt Success Only: 78 Both Succeeded: 13 Transfer Success Only: 9 Points below line: 400/420 (95.2%) Avg improvement: 54.7% Original vs. Optimized Distance to Target (All Dimensions) Both Failed Optimization Success Only Both Succeeded Transfer Success Only Figure 4:Ensemble transfer results.We find that ensemble attacks have a relatively uncorre- lated, yet nontrivial effect on transferability. 4.2 Persistence attacks We next consider persistent attacks, for which the goal is to elicit a targeted action over a longer horizon relative to single step attacks. In this setting, the attacker is given access torimages, where r∈1,2,3, which are collected in one of two ways: (1) images are taken from a “burn-in” period before the rollout begins, during which the VLA is actuated with randomly generated actions; and (2) images are taken from the firstrsteps of the rollout. In both settings, we play the VLA policy for 80 steps after applying the attack. We use hatches to denote persistence steps corresponding to therimages seen by the attacker, and non-hatched boxes to denote persistence for future, unseen images. The red dashed line indicates the frequency with which 50 independent, non-attacked rollouts elicit the targeted action. Our results in Figure 2 indicate that we consistently persist across the seen images, and as we increase the attacker’s image budget, generalization to unseen images also tends to increase, particularly on the Libero-Spatial fine-tune. Moreover, for this model, both variants achieve a nearly 28×improvement in the number of persistence steps relative to the nominal baseline. 4.3 Transfer attacks In the setting of transfer attacks, our goals are to (a) evaluate the extent to which attacks optimized for one VLA architecture transfer to other VLA architectures and (b) evaluate whether one can obtain auniversalattack acrossnOpenVLA fine-tunes, and assess transfer on the remaining4−nmodels. Architecture transfer.We optimize single-step attack strings for the OpenVLA base model, and then transfer these strings to three models: TraceVLA [82], CogACT [54], and OpenPi0 [83]. The architectural differences between the chosen models and OpenVLA are discussed in detail in Appendix B.1. To the best of our knowledge, SIMPLER [81] is the only benchmark on which all of these models have been evaluated, and we therefore adopt it for our comparison. To assess transfer, we compare four different prompting methods: the nominal instruction, a randomly chosen string of tokens from the downstream model’s vocabulary, and transferred strings both when the optimization successfully and unsuccessfully resulted in a match on the source model. In these experiments, we didnotobserve any exact matches across the seven action dimensions. This is unsurprising, given that textual attacks on standard VLLMs are known to exhibit little, if any, transferability [84]. We therefore compare theℓ 2 distance between the target action and the elicited action (discussed in greater detail in Appendix B.1). While we find that some attacks resulted in lowℓ 2 distances, this was more attributable to the fact that these targets happened to be easier to hit, which is evinced by the fact that random prompts tend to do well for these actions. In other words, optimized instructions tend to dono betterthan random instructions for target action elicitation. Ensemble transfer.Due to the computational cost, we only assess ensemble transfer forn= 2, meaning we attempt to optimize a universal suffix across two models, and test transfer across the other two. We report our results in Figure 4. We note that transfer has a nontrivial effect on control 8 Table 3:Candidate defenses against VLA attacks.Attack Success Rate (ASR) comparison across models with two modes of perplexity filtering (abbreviated as PF) and smoothing applied. DefenseLibero-10Libero-GoalLibero-ObjectLibero-Spatial No Defense63.3100.096.7100.0 Multimodal PF63.3100.096.7100.0 LLM-Only PF0.00.00.00.0 Smoothing0.00.00.00.0 authority. However, it is uncorrelated with whether the initial optimization over the source models was successful. Further, given that the majority of successful transfers occur at the bottom left of the plot, we observe a relationship between the nominal action’s distance from the target and the transfer success rate. Given the proven ineffectiveness of transfer in multi-modal models, we do not claim to have found conclusive proof of ensemble transfer. However, a perhaps more important takeaway is that we are able to train universal (atn= 2) attacks that are successful across multiple models. 5 Discussion VLA defenses.Given the connections drawn between chatbot jailbreaking and VLA attacks in §3.3, a natural question is whether jailbreaking defenses for LLMs and VLLMs can be extended to VLAs. While several defenses—including those that rely on modified system prompts and in-context demonstrations [85]—are inapplicable given that VLAs do not generally use a system prompt, several earlier defenses can be applied to VLAs. In particular, we consider the impact of the perplexity filter defense [86], which rejects queries if they match a user-defined preplexity threshold, and the smoothing defense [70]. In Table 3, we record the effectiveness of a text-based and multimodal perplexity filter, and smoothing, for 120 randomly selected one-hot target actions. As shown, there is a clear difference between the filters when we consider the image embeddings in the loss calculation. The ineffectiveness of the multimodal perplexity filter (where perplexity is calculated over the vision embeddings and the instruction) is due to the dominance of the image inputs over the loss term. We arrive at a similar conclusion as Jain et al.[86], who find the language-only perplexity filter effective against suffix-style attacks. However, in practice, this defense proves infeasible, primarily because the perplexity threshold depends entirely on the maximum perplexity of instructions seen on a held-out set, which cannot be known beforehand in open-world robotics applications. Further, we find that smoothing results in a 0% success rate, but also corrupts the instructions, resulting in a 0% success rate on non-attacked tasks. It is possible that this will become a more viable defense as models scale in both parameters and capabilities in the future. Safety mechanisms for VLAs.In the field of language modeling, a broad array of techniques are used to align outputs with human intentions, including supervised fine-tuning [87], RLHF [88,89], and adversarial training [90]. However, due to a mix of limited capability sets and decreased attention relative to chatbots, the field of AI-enabled robotics has not yet seen the proposal of analogous notions of alignment. Initial efforts, such the GRAPE robotic alignment algorithm [91], align VLAs at the trajectory-level toward improving generalization and task completion. In a similar spirit, several works design RL-style approaches to improve robot safety [92,93]. However, these works do not consider adversarial attacks on VLAs. This points to the need for notions of VLA refusal when attempts to subvert control are detected. Moreover, future defenses could incorporate tools from classical control (e.g., control barrier functions or formal methods), which have recently shown effectiveness against attacks on robotic planners [94]. 6 Conclusion VLAs are gaining momentum in the field of robotics due to their ability to fuse the textual and visual understanding of VLMs with the low-level actuation. In this paper, we attempt to anticipate future threat models that may impact robotic foundation models as they are deployed commercially. In particular, we present the first study of adversarial attacks on low-level VLA actuators, showing that by optimizing instructions we can obtain complete control authority over a target VLA. These results 9 underline the necessity for new forms of defenses that are reflective of the unique output format VLAs pose, as these systems become more powerful and widely used in society. Limitations and future work.We recognize that our attack may be difficult to employ in practice, due to the white-box nature and relative cost of the GCG algorithm. Further, while our attack is designed to work on any autoregressive VLA, diffusion-based models are also very prevalent throughout the field. Extending attack frameworks to black-box scenarios and diffusion-based models will be a critical step in the pursuit of fully assessing the risks these models pose. 10 References [1] Ji Woong Kim, Tony Z Zhao, Samuel Schmidgall, Anton Deguet, Marin Kobilarov, Chelsea Finn, and Axel Krieger. Surgical robot transformer (srt): Imitation learning for surgical tasks. arXiv preprint arXiv:2407.12998, 2024. 1, 3, 6 [2]Samuel Schmidgall, Ji Woong Kim, Alan Kuntz, Ahmed Ezzat Ghazi, and Axel Krieger. General-purpose foundation models for increased autonomy in robot-assisted surgery.Nature Machine Intelligence, pages 1–9, 2024. 1 [3]Rohan Sinha, Amine Elhafsi, Christopher Agia, Matthew Foutter, Edward Schmerling, and Marco Pavone. Real-time anomaly detection and reactive planning with large language models. arXiv preprint arXiv:2407.08735, 2024. 1 [4] Yingzi Ma, Yulong Cao, Jiachen Sun, Marco Pavone, and Chaowei Xiao. Dolphins: Multimodal language model for driving. InEuropean Conference on Computer Vision, pages 403–420. Springer, 2025. 1 [5] Bruno Silva, Leonardo Nunes, Roberto Estevão, Vijay Aski, and Ranveer Chandra. Gpt-4 as an agronomist assistant? answering agriculture exams using large language models.arXiv preprint arXiv:2310.06225, 2023. 1 [6]Djavan De Clercq, Elias Nehring, Harry Mayne, and Adam Mahdi. Large language models can help boost food production, but be mindful of their risks.Frontiers in Artificial Intelligence, 7: 1326153, 2024. 1 [7]Figure. Master plan.https://w.figure.ai/master-plan, 2022. Accessed: 2025-01-06. 1 [8]Unitree Robotics. Unitree go2.https://shop.unitree.com/products/unitree-go2, 2023. Accessed: 2025-01-06. 1 [9]Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky.π 0 : A vision-language-action flow model for general robot control, 2024. URLhttps://arxiv. org/abs/2410.24164. 1, 2, 3, 20 [10] Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model.arXiv preprint arXiv:2303.03378, 2023. 1 [11] Hao-Tien Lewis Chiang, Zhuo Xu, Zipeng Fu, Mithun George Jacob, Tingnan Zhang, Tsang- Wei Edward Lee, Wenhao Yu, Connor Schenck, David Rendleman, Dhruv Shah, et al. Mobility vla: Multimodal instruction navigation with long-context vlms and topological graphs.arXiv preprint arXiv:2407.07775, 2024. 1 [12]Michael Ahn, Debidatta Dwibedi, Chelsea Finn, Montse Gonzalez Arenas, Keerthana Gopalakr- ishnan, Karol Hausman, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, et al. Autort: Embodied foundation models for large scale orchestration of robotic agents.arXiv preprint arXiv:2401.12963, 2024. 1 [13]Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. Ai control: Improving safety despite intentional subversion.arXiv preprint arXiv:2312.06942, 2023. 1 [14] Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, et al. Alignment faking in large language models.arXiv preprint arXiv:2412.14093, 2024. [15] Erik Jones, Anca Dragan, and Jacob Steinhardt. Adversaries can misuse combinations of safe models.arXiv preprint arXiv:2406.14595, 2024. 1 11 [16]Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. 1, 3 [17] Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang, Xiaofeng Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, et al. Prompt injection attack against llm-integrated applications.arXiv preprint arXiv:2306.05499, 2023. [18] Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis, and Florian Tramèr. Defeating prompt injections by design.arXiv preprint arXiv:2503.18813, 2025. 1 [19]Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries, 2024. URL https://arxiv.org/abs/2310.08419. 1, 3, 6 [20]Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models, 2023. URLhttps: //arxiv.org/abs/2307.15043. 3, 4, 5, 6 [21]Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. arXiv preprint arXiv:2404.01318, 2024. 1 [22]Joseph Carlsmith. Is power-seeking ai an existential risk?arXiv preprint arXiv:2206.13353, 2022. 1 [23]Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Risks from learned optimization in advanced machine learning systems.arXiv preprint arXiv:1906.01820, 2019. [24]Andres Carranza, Dhruv Pai, Rylan Schaeffer, Arnuv Tandon, and Sanmi Koyejo. Deceptive alignment monitoring.arXiv preprint arXiv:2307.10569, 2023. 1 [25] Sid Black, Asa Cooper Stickland, Jake Pencharz, Oliver Sourbut, Michael Schmatz, Jay Bailey, Ollie Matthews, Ben Millwood, Alex Remedios, and Alan Cooney. Replibench: Evaluating the autonomous replication capabilities of language model agents.arXiv preprint arXiv:2504.18565, 2025. 1 [26]Xudong Pan, Jiarun Dai, Yihe Fan, and Min Yang. Frontier ai systems have surpassed the self-replicating red line.arXiv preprint arXiv:2412.12140, 2024. 1 [27] Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024. 2, 3, 4, 19 [28]Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choro- manski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al. Rt-2: Vision-language- action models transfer web knowledge to robotic control.arXiv preprint arXiv:2307.15818, 2023. 2, 3 [29]Alexander Robey, Zachary Ravichandran, Vijay Kumar, Hamed Hassani, and George J Pappas. Jailbreaking llm-controlled robots.arXiv preprint arXiv:2410.13691, 2024. 2, 3 [30]Hangtao Zhang, Chenyu Zhu, Xianlong Wang, Ziqi Zhou, Changgan Yin, Minghui Li, Lulu Xue, Yichen Wang, Shengshan Hu, Aishan Liu, et al. Badrobot: Manipulating embodied llms in the physical world.arXiv preprint arXiv:2407.20242, 2024. 2, 3 12 [31]Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine. Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In2017 IEEE International Conference on Robotics and Automation (ICRA), pages 3389–3396, 2017. doi: 10.1109/ICRA. 2017.7989385. 2 [32] Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. End-to-end training of deep visuomotor policies, 2016. URLhttps://arxiv.org/abs/1504.00702. 2 [33]Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta. R3m: A universal visual representation for robot manipulation, 2022. URLhttps://arxiv.org/abs/ 2203.12601. 2 [34]Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need.Advances in Neural Information Processing Systems, 30, 2017. 2 [35]Sai Vemprala, Rogerio Bonatti, Arthur Bucker, and Ashish Kapoor. Chatgpt for robotics: Design principles and model abilities, 2023. URLhttps://arxiv.org/abs/2306.17582. 3 [36]Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. In2023 IEEE International Conference on Robotics and Automation (ICRA), pages 9493–9500. IEEE, 2023. 3 [37]Montserrat Gonzalez Arenas, Ted Xiao, Sumeet Singh, Vidhi Jain, Allen Ren, Quan Vuong, Jake Varley, Alexander Herzog, Isabel Leal, Sean Kirmani, et al. How to prompt your robot: A promptbook for manipulation skills with code as policies. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 4340–4348. IEEE, 2024. [38]Sai H Vemprala, Rogerio Bonatti, Arthur Bucker, and Ashish Kapoor. Chatgpt for robotics: Design principles and model abilities.IEEE Access, 2024. 3 [39]Boyi Li, Yue Wang, Jiageng Mao, Boris Ivanovic, Sushant Veer, Karen Leung, and Marco Pavone. Driving everywhere with large language model policy adaptation, 2024. URLhttps: //arxiv.org/abs/2402.05932. 3 [40]Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, Yang Zhou, Kaizhao Liang, Jintai Chen, Juanwu Lu, Zichong Yang, Kuei-Da Liao, Tianren Gao, Erlong Li, Kun Tang, Zhipeng Cao, Tong Zhou, Ao Liu, Xinrui Yan, Shuqi Mei, Jianguo Cao, Ziran Wang, and Chao Zheng. A survey on multimodal large language models for autonomous driving, 2023. URLhttps://arxiv.org/ abs/2311.12320. 3 [41]Zichao Hu, Francesca Lucchetti, Claire Schlesinger, Yash Saxena, Anders Freeman, Sadanand Modak, Arjun Guha, and Joydeep Biswas. Deploying and evaluating llms to program service mobile robots.IEEE Robotics and Automation Letters, 9(3):2853–2860, March 2024. ISSN 2377-3774. doi: 10.1109/lra.2024.3360020. URLhttp://dx.doi.org/10.1109/LRA.2024. 3360020. 3 [42] Krishan Rana, Jesse Haviland, Sourav Garg, Jad Abou-Chakra, Ian Reid, and Niko Suenderhauf. SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Task Planning. In7th Annual Conference on Robot Learning, 2023. URLhttps://openreview.net/ forum?id=wMpOMO0Ss7a. 3 [43]Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240): 1–113, 2023. 3 [44]Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Je- natton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Manoj Kumar, Sjoerd van Steenkiste, Gamaleldin F. Elsayed, Aravindh Mahendran, Fisher Yu, Avital Oliver, Fantine Huot, Jasmijn Bastings, 13 Mark Patrick Collier, Alexey Gritsenko, Vighnesh Birodkar, Cristina Vasconcelos, Yi Tay, Thomas Mensink, Alexander Kolesnikov, Filip Paveti ́ c, Dustin Tran, Thomas Kipf, Mario Lu ˇ ci ́ c, Xiaohua Zhai, Daniel Keysers, Jeremiah Harmsen, and Neil Houlsby. Scaling vision transformers to 22 billion parameters, 2023. URLhttps://arxiv.org/abs/2302.05442. 3 [45]Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. Palm-e: An embodied multimodal language model. InarXiv preprint arXiv:2303.03378, 2023. 3, 19 [46]Krishan Rana, Jesse Haviland, Sourav Garg, Jad Abou-Chakra, Ian Reid, and Niko Suenderhauf. Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning, 2023. URLhttps://arxiv.org/abs/2307.06135. 3 [47]Qiao Gu, Alihusein Kuwajerwala, Sacha Morin, Krishna Murthy Jatavallabhula, Bipasha Sen, Aditya Agarwal, Corban Rivera, William Paul, Kirsty Ellis, Rama Chellappa, Chuang Gan, Celso Miguel de Melo, Joshua B. Tenenbaum, Antonio Torralba, Florian Shkurti, and Liam Paull. Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning.arXiv, 2023. 3 [48] Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An embodied generalist agent in 3d world, 2024. URLhttps://arxiv.org/abs/2311.12871. 3 [49] Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, and Andy Zeng. Do as i can and not as i say: Grounding language in robotic affordances. InarXiv preprint arXiv:2204.01691, 2022. 3 [50]Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Charles Xu, Jianlan Luo, Tobias Kreiman, You Liang Tan, Lawrence Yun- liang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An open-source generalist robot policy. InProceedings of Robotics: Science and Systems, Delft, Netherlands, 2024. 3 [51]Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choro- manski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalash- nikov, Yuheng Kuang, Isabel Leal, Lisa Lee, Tsang-Wei Edward Lee, Sergey Levine, Yao Lu, Henryk Michalewski, Igor Mordatch, Karl Pertsch, Kanishka Rao, Krista Reymann, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Pierre Sermanet, Jaspiar Singh, Anikait Singh, Radu Soricut, Huong Tran, Vincent Vanhoucke, Quan Vuong, Ayzaan Wahid, Stefan Welker, Paul Wohlhart, Jialin Wu, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich. Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023. URL https://arxiv.org/abs/2307.15818. 3, 4 [52]Open X-Embodiment Collaboration, Abby O’Neill, Abdul Rehman, Abhinav Gupta, Abhi- ram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew Wang, Andrey Kolobov, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess- Limerick, Beomjoon Kim, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu, Charlotte Le, Chelsea Finn, Chen Wang, Chenfeng Xu, Cheng Chi, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel 14 Morton, Danny Driess, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Foster, Fangchen Liu, Federico Ceola, Fei Xia, Feiyu Zhao, Felipe Vieira Frujeri, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Ge Yan, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guangwen Yang, Guanzhi Wang, Hao Su, Hao-Shu Fang, Haochen Shi, Henghui Bao, Heni Ben Amor, Henrik I Christensen, Hiroki Furuta, Homanga Bharadhwaj, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim, Jaimyn Drake, Jan Peters, Jan Schneider, Jasmine Hsu, Jay Vakil, Jeannette Bohg, Jeffrey Bingham, Jeffrey Wu, Jensen Gao, Jiaheng Hu, Jiajun Wu, Jialin Wu, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jingyun Yang, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kaiyuan Wang, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Ken Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Lin, Kevin Zhang, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Lawrence Yunliang Chen, Lerrel Pinto, Li Fei-Fei, Liam Tan, Linxi "Jim" Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen, Nicolas Heess, Nikhil J Joshi, Niko Suenderhauf, Ning Liu, Norman Di Palo, Nur Muhammad Mahi Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R Sanketi, Patrick "Tree" Miller, Patrick Yin, Paul Wohlhart, Peng Xu, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Quan Vuong, Rafael Rafailov, Ran Tian, Ria Doshi, Roberto Mart’in-Mart’in, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Ruohan Zhang, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante, Sean Kirmani, Sergey Levine, Shan Lin, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham Sonawani, Shubham Tulsiani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vikash Kumar, Vincent Vanhoucke, Wei Zhan, Wenxuan Zhou, Wolfram Burgard, Xi Chen, Xiangyu Chen, Xiaolong Wang, Xinghao Zhu, Xinyang Geng, Xiyuan Liu, Xu Liangwei, Xuanlin Li, Yansong Pang, Yao Lu, Yecheng Jason Ma, Yejin Kim, Yevgen Chebotar, Yifan Zhou, Yifeng Zhu, Yilin Wu, Ying Xu, Yixuan Wang, Yonatan Bisk, Yongqiang Dou, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yue Cao, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zhuo Xu, Zichen Jeff Cui, Zichen Zhang, Zipeng Fu, and Zipeng Lin. Open X-Embodiment: Robotic learning datasets and RT-X models.https://arxiv.org/abs/2310.08864, 2023. 3, 7, 19 [53] Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.arXiv preprint arXiv:1701.06538, 2017. 3 [54] Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizhong Zhang, Xiaofan Wang, Bei Liu, Jianlong Fu, Jianmin Bao, Dong Chen, Yuanchun Shi, Jiaolong Yang, and Baining Guo. Cogact: A foundational vision-language- action model for synergizing cognition and action in robotic manipulation, 2024. URLhttps: //arxiv.org/abs/2411.19650. 3, 8, 20 [55] Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. A general language assistant as a laboratory for alignment.arXiv preprint arXiv:2112.00861, 2021. 3 [56]Philipp Hacker, Andreas Engel, and Marco Mauer. Regulating chatgpt and other large generative ai models. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 1112–1123, 2023. 15 [57]Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022. 3 [58] Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail?Advances in Neural Information Processing Systems, 36, 2024. 3 [59]Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei W Koh, Daphne Ippolito, Florian Tramer, and Ludwig Schmidt. Are aligned neural networks adversarially aligned?Advances in Neural Information Processing Systems, 36, 2024. 3 [60]Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Zico Kolter, Matt Fredrikson, Eric Winsor, Jerome Wynne, Yarin Gal, and Xander Davies. Agentharm: A benchmark for measuring harmfulness of llm agents, 2024. URLhttps://arxiv.org/abs/2410.09024. 3 [61]Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. Frontier models are capable of in-context scheming.arXiv preprint arXiv:2412.04984, 2024. [62] Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger. Stress-testing capability elicitation with password-locked models.arXiv preprint arXiv:2405.19550, 2024. 3 [63]Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Zheng-Xin Yong, Suhas Kotha, et al. A safe harbor for ai evaluation and red teaming.arXiv preprint arXiv:2403.04893, 2024. 3 [64]Anka Reuel, Ben Bucknall, Stephen Casper, Tim Fist, Lisa Soder, Onni Aarne, Lewis Hammond, Lujain Ibrahim, Alan Chan, Peter Wills, et al. Open problems in technical ai governance.arXiv preprint arXiv:2407.14981, 2024. 3 [65]Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. Autodan: Generating stealthy jailbreak prompts on aligned large language models.arXiv preprint arXiv:2310.04451, 2023. 3 [66]Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. Visual adversarial examples jailbreak aligned large language models. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 21527–21536, 2024. 3 [67]Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue. Llm defenses are not robust to multi-turn human jailbreaks yet.arXiv preprint arXiv:2408.15221, 2024. 3 [68] Mark Russinovich, Ahmed Salem, and Ronen Eldan. Great, now write an article about that: The crescendo multi-turn llm jailbreak attack.arXiv preprint arXiv:2404.01833, 2024. 3 [69]Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, J Zico Kolter, Matt Fredrikson, and Dan Hendrycks. Improving alignment and robustness with circuit breakers. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. 3, 6 [70]Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas. Smoothllm: Defending large language models against jailbreaking attacks.arXiv preprint arXiv:2310.03684, 2023. 3, 9 [71]Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card.arXiv preprint arXiv:2412.16720, 2024. 3 [72]Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024. 3 16 [73]Chen Henry Wu, Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried, and Aditi Raghunathan. Adversarial attacks on multimodal agents.arXiv preprint arXiv:2406.12814, 2024. 3 [74] Yifei Wang, Dizhan Xue, Shengjie Zhang, and Shengsheng Qian. Badagent: Inserting and activating backdoor attacks in llm agents.arXiv preprint arXiv:2406.03007, 2024. 3 [75]Fredrik Nestaas, Edoardo Debenedetti, and Florian Tramèr. Adversarial search engine optimiza- tion for large language models.arXiv preprint arXiv:2406.18382, 2024. 3 [76] Sathwik Karnik, Zhang-Wei Hong, Nishant Abhangi, Yen-Chen Lin, Tsun-Hsuan Wang, and Pulkit Agrawal. Embodied red teaming for auditing robotic foundation models, 2024. URL https://arxiv.org/abs/2411.18676. 3 [77]C Szegedy. Intriguing properties of neural networks.arXiv preprint arXiv:1312.6199, 2013. 6 [78]Aleksander Madry. Towards deep learning models resistant to adversarial attacks.arXiv preprint arXiv:1706.06083, 2017. 6 [79] Rusheb Shah, Soroush Pour, Arush Tagade, Stephen Casper, Javier Rando, et al. Scalable and transferable black-box jailbreaks for language models via persona modulation.arXiv preprint arXiv:2311.03348, 2023. 6 [80]Suneel Belkhale, Yuchen Cui, and Dorsa Sadigh. Hydra: Hybrid robot actions for imitation learning, 2023. URLhttps://arxiv.org/abs/2306.17237. 7 [81]Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch, Oier Mees, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat, Isabel Sieh, Sean Kirmani, Sergey Levine, Jiajun Wu, Chelsea Finn, Hao Su, Quan Vuong, and Ted Xiao. Evaluating real-world robot manipulation policies in simulation. arXiv preprint arXiv:2405.05941, 2024. 7, 8 [82]Ruijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao, Hal Daumé I, Andrey Kolobov, Furong Huang, and Jianwei Yang. Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies, 2024. URLhttps://arxiv.org/abs/2412.10345. 8, 20 [83]Allen Ren. GitHub - allenzren/open-pi-zero: Re-implementation of pi0 vision-language-action (VLA) model from Physical Intelligence — github.com.https://github.com/allenzren/ open-pi-zero, 2024. [Accessed 29-01-2025]. 8, 20 [84]Rylan Schaeffer, Dan Valentine, Luke Bailey, James Chua, Cristobal Eyzaguirre, Zane Durante, Joe Benton, Brando Miranda, Henry Sleight, Tony Tong Wang, et al. Failures to find transferable image jailbreaks between vision-language models. InThe Thirteenth International Conference on Learning Representations, 2024. 8 [85]Zeming Wei, Yifei Wang, Ang Li, Yichuan Mo, and Yisen Wang. Jailbreak and guard aligned language models with only few in-context demonstrations.arXiv preprint arXiv:2310.06387, 2023. 9 [86]Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. Baseline de- fenses for adversarial attacks against aligned language models.arXiv preprint arXiv:2309.00614, 2023. 9 [87]Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022. URLhttps://arxiv.org/abs/2204.05862. 9 [88]Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences, 2023. URLhttps://arxiv.org/abs/ 1706.03741. 9 17 [89]Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. Scalable agent alignment via reward modeling: a research direction, 2018. URLhttps://arxiv.org/ abs/1811.07871. 9 [90]Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024. URL https://arxiv.org/abs/2402.04249. 9 [91]Zijian Zhang, Kaiyuan Zheng, Zhaorun Chen, Joel Jang, Yi Li, Chaoqi Wang, Mingyu Ding, Dieter Fox, and Huaxiu Yao. Grape: Generalizing robot policy via preference alignment.arXiv preprint arXiv:2411.19309, 2024. 9 [92]Ammar N. Abbas, Shakra Mehak, Georgios C. Chasparis, John D. Kelleher, Michael Guilfoyle, Maria Chiara Leva, and Aswin K Ramasubramanian. Safety-driven deep reinforcement learning framework for cobots: A sim2real approach, 2024. URLhttps://arxiv.org/abs/2407. 02231. 9 [93]Borong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei, Josef Dai, Yuanpei Chen, and Yaodong Yang. Safevla: Towards safety alignment of vision-language-action model via safe reinforcement learning.arXiv preprint arXiv:2503.03480, 2025. 9 [94]Zachary Ravichandran, Alexander Robey, Vijay Kumar, George J Pappas, and Hamed Hassani. Safety guardrails for llm-enabled robots.arXiv preprint arXiv:2503.07885, 2025. 9 [95] Kyle Stachowicz, Lydia Ignatova, and Sergey Levine. Lifelong autonomous improvement of navigation foundation models in the wild. In8th Annual Conference on Robot Learning, 2024. 19 18 A On the consequences of fine-tuning 0123456 Action index 0 20 40 60 Percent change GoalObjectSpatial10 Overall Success rate discrepancy OpenVLA base vs. fine-tunes Figure 5:Attacking fine-tuned vs. base VLAs.Each bar shows the percent change in the single step access rate of each OpenVLA fine-tune relative to analogous attacks on the OpenVLA base model. Both the overall and per-dimension success rates show a similar trend: single step attack effectiveness increases by 40-60% for each fine-tune relative to the base model. 0123456 Action Index -1 0 1 Target value Libero-Goal 0123456 Action Index Libero-Object 0123456 Action Index Libero-Spatial 0123456 Action Index Libero-10 Figure 6:Visualizing single step success rates.As we use one-hot target actions in the evaluation of single step attacks in §4.1, we can visualize the locations at which the attack fail. In particular, thex-axis of these plots shows the action dimensions, and they-axis shows the value of the of the one-hot component of the target. The blue dots denote the locations at which the single step attacks failed. We observe significant clustering of failures, particularly for the Libero-10 fine-tune, which had the lowest overall success rate. While the paradigm of fine-tuning VLAs on downstream tasks has gained traction of late [27,95], earlier works at the intersection of VLAs and robotics tended to focus on employing generalist, non-fine-tuned policies [45,52]. Thus, a natural question is whether task-specific fine-tuning has an impact on the effectiveness of attacks on VLAs. In Figure 5, we find that the OpenVLA base model is significantly more resistant to adversarial attacks relative to the four fine-tunes in Table 1. More specifically, across a uniform gridding of the action space of the OpenVLA-Base model, we ran the single step attack on 120 different target actions, recording our overall success rate of only 38% (see Table 2). In Figure 5, we plot the percent increase in the effectiveness of our single step attack on each of the four OpenVLA fine-tunes relative to the base model. We find that on average, attack effectiveness increases by nearly 40-60% for the fine-tuned models relative to the base model. This could indicate that fine-tuning in in some sense concentrates the actions in a smaller region (after normalization), making attacks easier to carry out. This point is corroborated by the evidence in 19 Figure 6, wherein we visualize the action spaces of the four fine-tuned models; we observe noticeable clustering of the locations of single step attacks, particularly for the lowest performing Libero-10 model. B Continued discussion of architectural transfer TraceVLACogACTOpenPi0 0.0 0.2 0.4 0.6 0.8 Distance from Target Average Distance from Target by Model Nominal Random Optim(Fail) Optim(Succ) Figure 7:Transfer attacks.By optimizing an instruction for OpenVLA-Base, and then attempt transfer to each of TraceVLA, CogACT, and OpenPi0, we evaluate the effectiveness of transferable attacks. Our results show that there is little to no transfer between OpenVLA-Base and other models. B.1 Architectural details We begin by discussing the differences between OpenVLA and the models chosen for our testing of transfer. For our motivations behind choosing models with different architectures, please refer to appendix A Of the three models chosen to test transfer against,TraceVLA[82] is the closest architectural match. It is a direct finetune of OpenVLA, however at each step, the model is given two images: the original image, as well as a version of the original image overlaid with a visual trace of active point trajectories. Therefore, at each stept, the model requires two images instead of one, therefore drastically shrinking the percentage of the embedding space occupied by the text embeddings.CogACT[54] introduces a diffusion action module onto the end of the language model outputs in order to predict a sequence of actions. Critically, the diffusion-based architecture makes text-based optimizations difficult due to the more continuous nature of the outputs. Finally,OpenPi0[83], an open-source implementation ofπ 0 [9], implements anaction experton top of the language model, which utilizes both a diffusion-based process and a separate set of weights for improved performance. All three models output actions in seven dimensions, identically to OpenVLA. B.2 Random vs. optimized instruction action distributions In this subsection, we further discuss our conclusion that transfer between models of different architectures has not occurred during our experiments. All experiments were run offline using an image taken from the SIMPLER environment for the "pick coke can" objective. For a given rollout, the image used for both GCG optimization and our tests for transfer was taken from timestep 10, which is the first image after the burn-in period. In the figures below: • "Nominal" refers to the nominal action that is predicted by each of the models when provided the "nominal" instruction, which is "pick coke can". 20 •"Optimized" refers to the action that is predicted by each of the models when provided the "optimized" string that is a result of running GCG on OpenVLA, under the nominal "pick coke can" environment. We use "optimized instruction" to refer to the instruction itself, and "optimized action" to refer to the resulting action. •"Random" refers to the action that is predicted by each of the models when provided with a "random" string sampled from each model’s vocabulary as the instruction. •"Target" refers to the action that we optimized fore when running the GCG optimization algorithm on OpenVLA. This was also the target action that we compared each of the three models’ actions to when computingℓ 2 distance. dim 0 dim 1dim 2 dim 3 dim 4dim 5 0.2 0.4 0.6 0.8 1.0 1.2 TraceVLA dim 0 dim 1dim 2 dim 3 dim 4dim 5 0.2 0.4 0.6 0.8 1.0 1.2 CogACT NominalOptimizedRandomTarget dim 0 dim 1dim 2 dim 3 dim 4dim 5 0.2 0.4 0.6 0.8 1.0 1.2 OpenPi0 Average Action Transfer by Model Figure 8:Average action for samples where GCGdid notconverge Figure 8 presents the average action across the first six dimensions (disregarding the gripper, which skews the distributions of the actions due to its relatively large magnitude) for each model, comparing the nominal, unsuccessfully optimized, and random instructions to the target action. Figure 9 presents the average action across the first six dimensions, except for all runs where GCG converged to obtain an instruction that is optimized for the target action. The plots have been normalized such that there are no negative actions, despite the fact that the action space lies in[−1,1]. dim 0 dim 1dim 2 dim 3 dim 4dim 5 0.2 0.4 0.6 0.8 1.0 1.2 TraceVLA dim 0 dim 1dim 2 dim 3 dim 4dim 5 0.2 0.4 0.6 0.8 1.0 1.2 CogACT NominalOptimizedRandomTarget dim 0 dim 1dim 2 dim 3 dim 4dim 5 0.2 0.4 0.6 0.8 1.0 1.2 OpenPi0 Average Action Transfer by Model Figure 9:Average action for samples where GCGdidconverge Figure 7 alone appears to demonstrate that the optimized instructions play an integral role in the success of transfer. However, in examining these figures, a different conclusion can be made. First, across both successful and unsuccessful optimizations, the average action elicited from the optimized instruction is incredibly similar to the average action elicited from random instructions. This suggests that the optimized strings are being interpreted no differently than random strings of tokens, and result in similar action outputs. Secondly, the average target action, for samples where GCG did converge, is far closer to the average actions for the optimized and random instructions than it is for samples were GCG did not converge. This is best evidenced in the TraceVLA plots, where we can clearly see the degree to which the target action follows a similar distribution to the optimized and 21 random actions in the successful case, compared to the unsuccessful one. Further, we notice that for samples where GCG converged, the average target action is much more uniform compared to the samples where GCG did not converge. These facts yield the following conclusions: •Whether or not a GCG run converged is correlated with the uniformity of the target action about the ⃗ 0vector.For the OpenVLA base model that was not finetuned on downstream LIBERO tasks, we found that one of the most accessible target actions was ⃗ 0, and as such, the randomly selected actions that we used to test transfer were most likely to converge the closer they were to that vector. • Optimized instructions are seen no differently to random instructions in the vocabu- laries of the models that we tested.This is demonstrated by the almost identical average actions for optimized instructions and random instructions across both figure 9 and figure 8. •The seemingly positive relationship between GCG convergence and ability for the optimized instructions to transfer is due to the randomly selected target action’s proximity to the distribution of random and optimized actions.It is not reflective of successful transfer, despite at first glance appearing to be. C Ablations C.1 Action elicitation We notice that in figure 6, target actions where the optimization failed are likely to be grouped together. Most notably, for Libero-10, the worst-performing of the four finetunes, we notice that there are large sections of the action space that are difficult to attain. In part, this can be attributed to the normalization factor: across the first four action indices, the q01 values are−0.6348214149475098, −0.7741071581840515,−0.7633928656578064, and−0.09749999642372131respectively. Best illustrated at index 3, since the q01 value is so close to 0, almost all 128 actions between 0 and -1 get normalized to this value. As a result, we are effectively repeating the same optimization trial for a singular value repeatedly until we reach−0.09749999642372131, an action which is clearly difficult to achieve in the action space. This is represented in the figure, as nearly all of those trials are failures. Similarly for dimensions 0 and 1, the majority of failures on that half of the action space occur before we reach the q01 value. This is not always the case, however. Note that, aside from the fourth dimension of Libero-10 and the 0th dimension of Libero-Object, we are able to achieve maximum values across all other dimensions of all other models. Table 4:Suffix Trials.We report the per-dimension ASR and overall ASR for each of the four models across 12 different actions, the−1and1one-hot vectors across the first six dimensoins (excluding the gripper), when utilizing the optimized instruction as a suffix for the nominal instruction. Model Per-dimension success rateOverall success rate Avg. computation per success 012345Optim. stepsTime (sec.) Libero-Goal100.0100.091.7100.0100.0100.091.7170.2942.7 Libero-Object100.091.791.791.775.091.750.0184.71242.5 Libero-Spatial91.783.391.7100.091.791.775.0121.0678.0 Libero-1083.375.091.783.391.775.041.7238.21358.3 C.2 Suffix trials Table 4 demonstrates the success of our method when being utilized akin to standard GCG, where we optimize a suffix to be placed at the end of an instruction. The nominal instructions for the chosen tasks are as follows: •Libero-Goal:“push the plate to the front of the stove” •Libero-Object:“pick up the alphabet soup and place it in the basket” 22 •Libero-Spatial:“pick up the black bowl between the plate and the ramekin and place it on the plate” •Libero-10:“pick up the book and place it in the back compartment of the caddy” For the first 6 action dimensions, we perform GCG optimization on the maximum (1) and minimum (−1) one-hot vectors, as we find these to be the most difficult actions to achieve within the action space. Our results show with certainty that our method works not only as an optimized instruction, but also as an optimized suffix, given that additional trials on easier actions will result in a much higher true ASR, provided an increase in budget. Finally, in figure 3, we show that as an attacker’s token budget increases, so do overall success rates, and that increasing a budget also leads to fewer optimization steps being necessary for convergence. 23