Paper deep dive
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
Jihoon Hong, Julian Skifstad, Qiyue Dai, Alice Chan, Glen Chou
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/18/2026, 9:49:39 AM
Summary
This paper introduces World-Action Linear Quadratic Regulator (WA-LQR), a training-free method to improve the robustness of World Action Models (WAMs) against distribution shifts. By using mechanistic interpretability to identify low-dimensional linear separability in activation spaces, the authors construct contrastive steering directions. WA-LQR leverages local linearity in WAM dynamics to synthesize a reduced-order LQR controller that provides closed-loop feedback steering, significantly improving robustness to camera, gripper, and visual noise perturbations in models like Cosmos-Policy and DiT4DiT.
Entities (8)
Relation Signals (5)
Cosmos Policy → exhibits → strong steerability
confidence 95% · we predict strong steerability in the Cosmos-Policy and DiT4DiT models
LingBot-VA → exhibits → weak steerability
confidence 95% · weak steerability in LingBot-VA, consistent with steering intervention results
WA-LQR → improvesrobustnessof → World Action Models
confidence 95% · WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations
Mechanistic Interpretability → enablesidentificationof → low-dimensional linear separability
confidence 92% · we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features
WA-LQR → isbasedon → Linear Quadratic Regulator
confidence 90% · yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This motivates the use of contrastive activation directions for training-free WAM steering. We also show that local linearity in WAM activation dynamics enables efficient feedback steering via model-based optimal control, yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller. Via mechanistic evaluations, we predict strong steerability in the Cosmos-Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations over unsteered and prompt steering baselines.
Tags
Links
- Source: https://arxiv.org/abs/2607.14943v1
- Canonical: https://arxiv.org/abs/2607.14943v1
Trouble viewing inline? Open PDF directly →
Full Text
88,827 characters extracted from source content.
Expand or collapse full text
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control Jihoon Hong ⋆ , Julian Skifstad ⋆ , Qiyue Dai, Alice Chan, Glen Chou Georgia Institute of Technology jhong392, jskifstad3, qdai41, ichan30, chou@gatech.edu ⋆ Equal contribution Website§ CodeÅ Video Abstract: World Action Models (WAMs) enable semantically- and physically- informed control but are brittle under distribution shift. In this work, we use mech- anistic interpretability to study how robustness-relevant perturbations are repre- sented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This moti- vates the use of contrastive activation directions for training-free WAM steering. We also show that local linearity in WAM activation dynamics enables efficient feedback steering via model-based optimal control, yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR con- troller. Via mechanistic evaluations, we predict strong steerability in the Cosmos- Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR general- izes contrastive directions to new tasks and improves robustness to camera, grip- per, and visual-noise perturbations over unsteered and prompt steering baselines. Keywords: mechanistic interpretability, world action models, optimal control 퐂퐚퐦퐞퐫퐚퐎퐫퐢퐞퐧퐭퐚퐭퐢퐨퐧퐆퐚퐮퐬퐢퐚퐧퐍퐨퐢퐬퐞퐆퐫퐢퐩퐞퐫퐏퐨퐬퐢퐭퐢퐨퐧 퐍 퐨 퐭 퐒 퐭 퐞 퐞 퐫 퐞 퐝 퐖 퐀 - 퐋 퐐 퐑 퐅 퐚 퐢 퐥 퐮 퐫 퐞 퐒 퐮 퐜 퐜 퐞 퐬 퐬 Figure 1: WA-LQR makes World Action Models more robust to perturbations including gripper- position changes, camera-orientation shifts, and Gaussian sensor noise. In these Cosmos-Policy examples from LIBERO-10, the yellow and green boxes mark the two objects that must be placed in the basket. Without steering, the WAM fails; with WA-LQR, it succeeds. 1 Introduction Foundation models have advanced robot learning through policies that generalize across objects, tasks, and environments. While VLA models map observations and instructions directly to ac- tions, World Action Models (WAMs) additionally couple action prediction with future-state mod- arXiv:2607.14943v1 [cs.RO] 16 Jul 2026 eling through video-generative backbones [1–5]. This makes WAMs a promising route toward dynamically-coherent robot policies. However, they remain brittle under out-of-distribution shifts, including camera changes, robot initial-state perturbations, and visual corruption [5]. Because such nuisance factors are unavoidable in deployment and cannot be exhaustively covered with training data, it is essential to improve robustness without extensive new data collection or slow retraining. In this paper, we ask if the internal activations of WAMs reveal why they fail under such perturba- tions, and if these representations can be used to improve robustness without finetuning. We study perturbations to camera position, initial gripper position, and Gaussian image noise. Using mecha- nistic interpretability (MI), we compare activations from nominal and perturbed rollouts to test for simple linear geometric activation space structure. By evaluating the degree of linear separability across models, we find this structure is perturbation- and architecture-dependent, suggesting that some WAMs contain steerable representations of robustness-critical features, while others may not. By leveraging the degree of linear separability as a predictor of steerability, we can identify models that are well suited to training-free WAM activation steering, in which contrastive examples are used to construct directions that distinguish nominal from perturbed behavior. We use these directions for open-loop and closed-loop robustness interventions. In particular, we introduce World-Action Linear Quadratic Regulator (WA-LQR), a reduced-order optimal-control method that projects ac- tivations into a low-dimensional contrastive subspace and uses local linear dynamics to synthesize a closed-loop LQR steering controller. Unlike open-loop activation addition, WA-LQR adapts on- line, steering only when activations deviate from the target feature strength while penalizing large perturbations. We evaluate our mechanistic predictions and steering method on Cosmos-Policy [2], DiT4DiT [6], and LingBot-VA [7]. Our low-dimensional linear separability analysis predicts strong steerability for Cosmos-Policy and DiT4DiT and weak steerability for LingBot-VA, which empiri- cal intervention results validate. On Cosmos-Policy, contrastive directions transfer across LIBERO tasks, and WA-LQR improves robustness to camera, gripper, and visual-noise perturbations over non-steered, prompt-steering, and several open-loop steering baselines. These results show that MI can diagnose WAM robustness and guide inference-time interventions. Our contributions are: • We conduct a mechanistic analysis of WAM activations under robustness-relevant perturbations, identifying when nuisance features exhibit low-dimensional linear separability. • We show that steerability is architecture-dependent: Cosmos-Policy [2] and DiT4DiT [6] exhibit clear linear structure across multiple perturbations, whereas LingBot-VA [7] exhibits substantially weaker separability. Moreover, we show that separability is strongly correlated with steerability, making it a useful diagnostic for identifying models that are amenable to activation steering. • We leverage these insights to construct contrastive activation directions for training-free WAM steering. First, we adapt open-loop activation addition techniques from the LLM literature to WAMs. We then show that the local linearity of the diffusion transformer dynamics in a reduced WAM activation space enables the efficient closed-form synthesis of closed-loop steering con- trollers via WA-LQR. To the best of our knowledge, these are the first open- and closed-loop activation steering methods for WAMs. • We evaluate WA-LQR on robustness benchmarks, showing that MI-based steerability predictions strongly correlate with intervention outcomes and that WA-LQR improves robustness on Cosmos- Policy across multiple perturbation types. 2 Related Work WAMs. Foundation models have enabled general-purpose robotic policies that map observations to actions [1, 8–14]. VLA models are often reactive action predictors and do not explicitly model how the world evolves under robot interventions, limiting long-horizon reasoning and robustness under distribution shift [5, 15]. WAMs address this by coupling action generation with future-state prediction, often adapting video generative backbones to robotic data [2–4, 16]. Existing WAMs in- clude inverse-dynamics-style models, where predicted futures are decoded into actions [17–21], and unified video-action models, where future states and actions are generated jointly [2–4]. We study both families and show that steerable, robot-relevant activation structure is architecture-dependent. 2 Robustness of Action Models. Despite progress in general-purpose robot policies, robustness re- mains a major deployment obstacle. Benchmarks such as VLATest, COLOSSEUM, and LIBERO- Plus expose brittleness to camera viewpoint, robot initial state, object layout, lighting, background texture, sensor noise, and language phrasing [22–26]. Prior studies suggest that robustness gains of- ten require data diversity, wrist-camera observations, RL post-training, or robustness-oriented fine- tuning rather than arising from inherently stable representations [27–30]. WAMs aim to improve robustness with future-state modeling and video-based priors [2–4, 19, 21, 31], but still fail under shifts such as camera viewpoint and robot initial-state changes [5]. This motivates inference-time methods that improve robustness without exhaustive perturbation data or costly retraining. MI and Activation Steering. MI identifies internal representations that modulate model behavior [32, 33]. In LLMs, many semantic features appear approximately linear in activation space, moti- vating activation steering: inference-time hidden-state modifications that change behavior without retraining [34–38]. Most methods compute contrastive directions from examples with and without a target concept, then add or transform activations to steer behaviors [39–46]. Because these interven- tions are often layer-local and open-loop, recent work models activations as dynamical systems and uses feedback for more targeted steering [47–51]. Activation steering has also begun to extend to image and video generators [52–55], where semantics are distributed across text, spatial, temporal, timestep, and layer representations. WAMs add a further challenge: their activations affect not only generated visual content, but also action-relevant predictions that determine robot behavior. Relatively little work applies MI to robotic foundation models, and existing studies focus on VLAs. Prior work identifies steerable VLA activation directions [56], formalizes feature observability and controllability [57], finetunes task-relevant attention heads [58], discovers SAE-based motion prim- itives [59], measures causal reliance on visual regions [60], and uses activation injection, probes, sparse latents, or conceptor subspaces to analyze and steer VLA behavior [61–63]. In contrast, we study WAMs, whose DiT-style video backbones jointly encode future visual states, action dynam- ics, and control outputs. We show that WAM steerability is architecture-dependent and develop feedback-based steering that treats WAM inference as a reduced-order dynamical system. 3 Preliminaries and Problem Statement Linear Quadratic Regulator (LQR) The LQR problem (1) [64] seeks a controller that minimizes a quadratic state-control cost (1) for linear time-varying dynamics (1b) min u k H−1 k=1 L := H−1 X k=1 z ⊤ k Q k z k + u ⊤ k R k u k + z ⊤ H Q H z H (1a) subject to z k+1 = A k z k + B k u k , ∀k = 1,...,H − 1.(1b) where Q k ⪰ 0 penalizes state error and R k ≻ 0 penalizes control effort. The optimal policy has closed form u ∗ k = −K k z k , where K k is obtained efficiently via Riccati recursions. Given a set of nominal setpoints( ̄z k , ̄u k ) k=1,·,H , (1) can be generalized to penalize deviations δz k := z k − ̄z k and δu k := u k − ̄u k from the setpoints. By modifying (1) to min δu k H−1 k=1 H−1 X k=1 δz ⊤ k Q k δz k + δu ⊤ k R k δu k + δz ⊤ H Q H δz H (2a) subject to δz k+1 = A k δz k + B k δu k , k = 1,...,H − 1,(2b) the solution admits a closed-form tracking controller u ∗ k := ̄u k − K k δz k . When the dynam- ics z k+1 = f (z k ,u k ) are non-linear, similar formulations to (2) are possible by letting A k = ∇ z k f (z k ,u k ) and B k = ∇ u k f (z k ,u k ), and approximating δz k+1 ≈ A k δz k + B k δu k assuming local linearity. WAM Architectures While the two types of WAMs vary in architecture, they both build on Dif- fusion Transformer (DiT) based video generation models [65]. Starting from a random latent action representation ˆx T ∼N (0,I), these models are used to gradually denoise it to a clean x out by ˆx t−1 = STEP(ˆx t ,t,M (ˆx t ,t,h)), x out = DEC(ˆx 1 ),(3) 3 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Cosmos-Policy First Condition positive negative PC1 PC2 PC3 Block 2 avg loss: 0.0000 Best PC1 PC2 PC3 Block 27 avg loss: 0.1645 Last PC1 PC2 PC3 Block 0 avg loss: 0.8877 DiT4DiT PC1 PC2 PC3 Block 4 avg loss: 0.0000 PC1 PC2 PC3 Block 15 avg loss: 0.0000 PC1 PC2 PC3 Block 0 avg loss: 0.5501 First PC1 PC2 PC3 Block 21 avg loss: 0.0854 Best PC1 PC2 PC3 Block 27 avg loss: 0.3930 Last PC1 PC2 PC3 Block 0 avg loss: 0.2626 PC1 PC2 PC3 Block 8 avg loss: 0.2626 PC1 PC2 PC3 Block 15 avg loss: 0.2626 (a)(b) Figure 2: On Task 0 of LIBERO-10 [22], we evaluate (a) activations corresponding to noise- perturbed and clean inputs on the first, best intermediate, and final DiT block residual stream for Cosmos-Policy 2B and DiT4DiT [2, 6]. Activations are projected onto the top three principal com- ponents, with the reported SVM (gray) and hinge loss. (b) Repeated for camera perturbations. where h is the embedding of the task prompt p, M is the model, STEP is a choice of ODE solver, and DEC decodes latent actions. Typically, M is a sequence of L DiT layers φ (l) for l = 0,· ,L− 1: x 0,t := Tok(ˆx t ), x l+1,t = φ (l) (x l,t ,t,h), M (ˆx t ,t,h) := Detok(x L,t ).(4) After block L, the WAM output is detokenized and passed to the model scheduler (3). We treat the scheduler transition in (3) as the mechanism that chains together T independent L-horizon within- denoising-timestep controllers. In this work, we will perform activation steering by adding control inputs inside the DiT blocks. Let u l,t denote these perturbations. The steered WAM block is x l+1,t = f l,t (x l,t ,u l,t ) := φ (l) (x l,t ,t,h) + u l,t . (5) Problem Statement In this paper, we use MI to identify robustness-relevant features in WAM acti- vation space and use these features for training-free robustness improvement via activation steering. Problem 1: Linear feature discovery in WAM activation space (Sec. 4). Given a WAM, determine whether a target nuisance feature is represented by approximately linear directions in WAM activa- tion space. Concretely, for each layer l and denoising timestep t, we seek low-dimensional projec- tions P l,t and linear directions v l,t such that the projected activations of ξ + and ξ − are separable. Problem 2: Training-free robustness steering (Sec. 5). Given a WAM, synthesize an open- or closed- loop inference-time steering policy that modifies activations and modulates the features discovered in Problem 1. 4 A Mechanistic Study for Interpreting WAMs A central assumption in the activation steering literature is that model activation spaces exhibit sim- ple, interpretable geometric structure, across domains including LLMs, VLAs, and video generation models [41, 42, 51, 55, 56]. Specifically, most steering algorithms assume linear semantic feature directions, as determined by the emergent linear separability of activations corresponding to semanti- cally contrastive inputs [43, 66]. However, the existence of this structure in WAMs is not guaranteed, despite the steerable foundation model backbone. Indeed, prior work has shown that such represen- tations are fragile to the finetuning process undergone by robotics foundation models [56, 67]. Setup We study the emergent geometry in robustness features for WAMs, specifically Cosmos- Policy 2B, DiT4DiT, and LingBot-VA [2, 6, 7]. We consider contrastive datasets related to three key sources of sensitivity in WAM manipulation tasks: perturbation to initial gripper position, ini- tial camera position, and corruption of camera inputs with Gaussian noise. We collect activations corresponding to each dataset through a full model forward pass. That is, for prompts p + ∈ D + , 4 p − ∈ D − where, e.g., D + = Clean inputs, D − = Noised inputs, we collect DiT acti- vations x p + t,ℓ and x p − t,ℓ for all t,ℓ. Typically, DiT activations have a shape of (F,H,W,D), for some token frame, height, width and hidden dimension F < H ≤ W ≪ D, making direct storage in- tractable for a substantial sample size. Thus, for this section, we consider an average pooling over the token position, resulting in a summarized activation ̄x ∈R D for all sampled activations. We also perform an average pooling over robot action chunk timesteps; see Sec. 5.1 for details. Given contrastive sets of activations ̄x + k k and ̄x − k k , we define a contrastive direction d k = ̄x + k − ̄x − k .(6) Inspired by [66], we consider a simple geometric evaluation of the separability of contrastive datasets via principal component analysis (PCA) conducted on the set of contrastive directions,d k k . With the objective of enabling linear feature-based steering, we seek to identify linear separability of contrastive datasets. Note that in the high dimensionality of the original mean-pooled system, N samples with N ≪ D are trivially linearly separable, rendering analysis in the full-dimensional space uninformative. Hence, we consider a low-dimensional approximation of the data. In fact, we find in a later section (Sec. 6.1) that the contrastive directions are well summarized by as few as three principal components. Projecting the contrastive activations ̄x + and ̄x − onto the top three PCs, we construct a quantitative metric for the linear separability of the pairs of contrastive datasets. In particular, we fit a linear support vector machine (SVM) to the contrastive activations in three dimensions, and measure the classification loss as the average hinge loss per sample [68], loss(D + ,D − ) = 1 N X i∈[N ] max(0, 1− y i (w T ̄x i + b)),(7) where x | w T x + b = 0 denotes the SVM hyperplane and y i ∈ −1, 1. Intuitively, this measures how well the three-dimensional linear SVM distinguishes the two datasets, where perfect classification yields a loss of 0, and random classification yields a loss of 1. PC1 PC2 PC3 Block 0 avg loss: 0.0000 Tasks 6&0 First Cond. / Task task 6 positive task 6 negative task 0 positive task 0 negative PC1 PC2 PC3 Block 1 avg loss: 0.0000 Best PC1 PC2 PC3 Block 27 avg loss: 0.1068 Last PC1 PC2 PC3 Block 0 avg loss: 0.0000 Tasks 6&1 PC1 PC2 PC3 Block 1 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.1012 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Tasks 6&4 PC1 PC2 PC3 Block 2 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.1829 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Tasks 6&7 PC1 PC2 PC3 Block 1 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.0309 Figure 3: Pairwise separability for Cosmos-Policy under Gaussian noise corruption. Across task pairs, high linear separability and shared feature clusters suggest reusable representations that can be exploited for activation steering. Results We perform this analysis across the perturbations for all 10 tasks of the LIBERO-10 dataset [22], and visualize the results for Cosmos- Policy in Fig. 2 (further evaluations are provided in App. E, including for DiT4DiT [6]). Notably, we find that the separability of different features is highly task-dependent, indicating that the latent representation correspond- ing to the same perturbation is not always shared across different scenes and tasks.A key observation of this study is that certain combinations of tasks share feature representations corresponding to the same input per- turbation – see Fig. 3 for an exam- ple with Cosmos-Policy. Using these clustered features, we are able to con- struct steering objectives which gener- alize across tasks, as discussed in the following sections. This observation enables the activation steering formu- lation in Sec. 5. Applying the same analysis to the ac- tion inference activations for LingBot- 5 Figure 4:Activations from camera-perturbedandclean LingBot-VA inputs at the first, best intermediate, and final trans- former block residual streams. We observe much weaker separa- bility than in the Cosmos-Policy setting. VA yields substantially milder results. Across tasks and perturbations, we observe little to no sepa- rability as reported by the SVM loss, as well as qualitatively by inspection (see Fig. 4). 5 Activation Steering for Robustifying WAMs Motivated by the MI results of Sec. 4, we give an overview of our steering method. We first construct contrastive vectors isolating the desired robustness feature and discuss a simple open-loop method for steering WAMs with the contrastive vectors (Sec. 5.1). We also propose a reduced-order LQR-based steering approach (Sec. 5.2) to enable scalable control-theoretic steering for WAMs. 5.1 Open-Loop Activation Addition from Contrastive Vectors Given paired inputs (ξ + n ,ξ − n ) that differ primarily in a desired feature, e.g., nominal versus perturbed camera pose or clean versus corrupted observations, we compute contrastive activation directions by subtracting hidden states from the two forward passes (App. A). For layer l, denoising timestep t, and action-chunk timestep τ ∈ [H a ], this gives d (n) l,t,τ := x (ξ + n ) l,t,τ − x (ξ − n ) l,t,τ .(8) We thus describe the simplest contrastive steering method: activation addition (ActAdd) [41]. Ac- tAdd assumes that the target feature is roughly linear in activation space. A steering vector is formed by averaging contrastive directions (8) over pairs/action-chunk positions a l,t := 1 NH a N X n=1 H a X τ =1 d (n) l,t,τ .(9) At inference time, ActAdd steers by adding a l,t with strength γ: x l,t,τ ← x l,t,τ + γa l,t , ∀τ ∈ [H a ].(10) The same direction is applied across all action-chunk timesteps, with γ setting steering strength: positive values push toward ξ + , while negative values reverse the effect. ActAdd is training-free and weight-preserving, but open-loop: it ignores the current activation, WAM dynamics, and whether the feature is already at the desired strength, which can cause oversteering and action degradation. 5.2 Steering Latent World Activation Dynamics via WA-LQR WA-LQR generalizes ActAdd with feedback control, adapting the T2V steering method of [55] to WAMs. Instead of adding a fixedγa l,t , it projects contrastive vectors into a latent feature space, mea- sures deviation from a latent setpoint, and computes a minimum-cost LQR intervention. Thus, it pre- serves ActAdd’s training-free nature while adapting interventions to the realized WAM activation. Dimensionality Reduction As in T2V models [55], full activation-space LQR is infeasible due to the high dimensionality of WAM activations. We instead assume (and validate in Sec. 6) that robustness-relevant factors lie largely in a low-dimensional contrastive subspace. For each layer- denoising pair (l,t), we construct this subspace by pooling contrastive directions (8) over prompt pairs and action-chunk timesteps, then applying streaming randomized singular value decomposition (SVD) [69] to the matrix whose rows are d (n) l,t,τ : n ∈ [N ], τ ∈ [H a ]. This yields a compact 6 orthonormal basis V l,t ∈R d x ×d z with d z ≪ d x . Define the projection matrix P l,t := V ⊤ l,t ∈ R d z ×d x , which is shared across robot action-chunk timesteps τ ∈ [H a ]. For each (l,t,τ ), define the latent activation z l,t,τ := P l,t x l,t,τ ∈R d z . Let A l,t and B l,t denote the raw Jacobians of the controlled block dynamics in (5) along a nominal trajectory: A l,t := ∂f l,t ∂x ̄x l,t , ̄u l,t , B l,t := ∂f l,t ∂u ̄x l,t , ̄u l,t . The reduced latent dynamics are δz l+1,t,τ ≈ e A l,t δz l,t,τ + e B l,t δu l,t,τ , l = 0,...,L− 1,(11) where e A l,t := P l+1,t A l,t P ⊤ l,t ∈R d z ×d z and e B l,t := P l+1,t B l,t ∈R d z ×d u . Neither A l,t nor B l,t is materialized explicitly; products with e A l,t and e B l,t are computed efficiently using Jacobian-vector products (JVPs) or vector-Jacobian products (VJPs). Defining Feature Setpoints We define feature setpoints in the latent space for robustness-relevant WAM features. We first compute a latent contrastive direction averaged over action chunk indices e z l,t := 1 NH a N X n=1 H a X τ =1 P l,t (x (ξ + n ) l,t,τ − x (ξ − n ) l,t,τ ), v z l,t := e z l,t ∥e z l,t ∥ 2 .(12) For a realized latent activation z l,t,τ , the feature strength is β z l,t,τ := (v z l,t ) ⊤ z l,t,τ . We set the desired feature strength as β z,∗ l,t := λ∥e z l,t ∥ 2 , where λ controls steering strength. The feature tracking error for action-chunk timestep τ is α l,t,τ := β z,∗ l,t − (v z l,t ) ⊤ z l,t,τ , with δz l,t,τ :=−α l,t,τ v z l,t . Thus, δz l,t,τ is the latent tracking error that, if corrected, would bring the chunk-τ activation to the desired feature setpoint along the contrastive direction. Reaching Feature Setpoints via WA-LQR Using the latent dynamics in (11), we compute LQR controllers that steer WAM activations toward the feature setpoints. Unlike the T2V setting [55], we do not solve one T×L-step LQR over both transformer blocks and scheduler transitions. Instead, for each chunk index τ and denoising timestep t, we solve an independent L-step LQR over the transformer blocks: min δu l,t L−1 l=0 L−1 X l=0 δz ⊤ l,t Q l,t δz l,t + δu ⊤ l,t R chunk l,t (τ )δu l,t + δz ⊤ L,t Q L,t δz L,t s.t. δz l+1,t = e A l,t δz l,t + e B l,t δu l,t , l = 0,...,L− 1. (13) The control penalty incorporates an action-decay schedule over robot action chunk indices. Let τ ∈ 0,...,H a − 1 index the action chunks executed over time. For chunk τ , we set r(τ ) := min(R final ,R init exp(τ/τ R )), where R init is the initial steering penalty, R final is a large satu- ration value, and τ R controls the decay rate. The LQR control-cost matrix at chunk τ is then R chunk l,t (τ ) := r(τ )I d u , which is directly used in LQR (13). Since r(τ ) increases with τ and sat- urates at R final , steering is strongest for early chunks and gradually decays as the robot proceeds. A larger τ R slows this growth, causing more chunks to be steered before saturation. In our experi- ments, R final is chosen large enough that saturated chunks receive negligible steering. Notably, (13) can be efficiently solved inO(Ld 3 z ) time [70] on the CPU orO(logL· log 2 d z ) on the GPU [71]. For each action chunk index τ and denoising timestep t, solving (13) yields gains K l,t,τ ∈R d u ×d z , which are used only within the corresponding denoising pass. At inference time, the controller is u ∗ l,t,τ := ̄u l,t,τ − K l,t,τ δz l,t,τ = ̄u l,t,τ + K l,t,τ α l,t,τ v z l,t .(14) When ̄u l,t,τ = 0, the intervention magnitude is proportional to the online feature-tracking error for robot action chunk τ . The chunk-specific policy K l,t,τ is then applied to the activation at the cor- responding action-chunk index. The full WA-LQR procedure chains the T per-timestep controllers through WAM inference. For action chunk index τ and denoising timestep t, we apply the L feed- back gains K 0,t,τ ,...,K L−1,t,τ inside the DiT blocks, as in (5). At each chunk index τ , the steered block output x L,t is then passed to the scheduler, which produces ˆx t−1 and initializes the 7 next denoising pass, yielding T chained L-step LQR controllers. Compared with open-loop Ac- tAdd, WA-LQR adapts to the realized latent activation at each layer and denoising timestep, steering only when the WAM deviates from the desired robustness feature setpoint. 6 Results In this section, we evaluate whether the mechanistic structure identified above can be used to im- prove WAM robustness through activation steering. We first assess the validity of the assumptions made by WA-LQR to justify its applicability to WAM steering (Sec. 6.1). Next, we compare open- loop activation addition (ActAdd) and closed-loop WA-LQR across Cosmos-Policy, DiT4DiT, and LingBot-VA under OOD perturbations to camera orientation, initial gripper position, and camera noise (Sec. 6.2). Overall, the results show that steering is effective when feature-relevant activations are linearly separable, improving success rates by up to 41%. WA-LQR further improves robustness while helping avoid oversteering in settings where open-loop control is less reliable. See App. D for further experimental details. 6.1 Activation Properties ퟏퟎퟐퟎ ퟎ . ퟓ @ 퐑 퐚 퐧 퐤 = ퟑ 퐑퐚퐧퐤 퐄 퐱 퐩 퐥 퐚 퐢 퐧 퐞 퐝 퐕 퐚 퐫 퐢 퐚 퐧 퐜 퐞 ퟎ . ퟗ @ 퐑 퐚 퐧 퐤 = ퟔ ퟎ . ퟗ ퟗ @ 퐑 퐚 퐧 퐤 = ퟏ ퟗ ퟎ.ퟎ ퟎ.ퟐ ퟎ.ퟒ ퟎ.ퟔ ퟎ.ퟖ ퟏ.ퟎ ퟑퟎퟒퟎퟓퟎퟔퟎ Figure 6: Cumulative variance of latent contrastive vec- tors for camera orientation perturbation on Cosmos-Policy 2B [2] explained by the top-k singular vectors. We first empirically verify the assump- tions that enable LQR-based steering of latent activations: 1) preservation of the information in the contrastive vectors in the latent subspace and 2) linearity of the WAM dynamics within the sub- space. We first study contrastive infor- mation preservation. In Fig. 6 we show that the variance of contrastive vectors projected to a 64-dimensional latent subspace, P l,t (x (ξ + n ) l,t,τ −x (ξ − n ) l,t,τ ), is mostly captured by the top few dimensions. This is caused by rapid decay of singu- lar values, which suggests that the influence of perturbations such as the change in camera orientation on activations manifests in a significantly lower dimensional subspace compared to raw activations. To assess dynamics linearity, we observe that WAMs are locally linear (aligned with recent discov- eries in LLM [51] and T2V [55] models), which allows them to be effectively steered using LQR. Fig. 5(a) shows the cosine similarity and the magnitude ratio between P l+1,t φ (l) x l,t + P T l,t ε,t,h and P l+1,t φ (l) (x l,t ,t,h) + ̃ A l,t ε for random perturbation ε whose norm is proportional to that of ∥x l,t ∥. The cosine similarities and magnitude ratios remain close to 1 throughout inference, demon- (퐚)(퐛) ퟏ.ퟎ ퟎ.ퟖ ퟎ.ퟔ ퟎ.ퟒ ퟎ.ퟐ ퟎ.ퟎ (퐥,퐭)=(ퟔ,ퟎ)(퐥,퐭)=(ퟏퟒ,ퟐ) 퐑퐚퐧퐝퐨퐦 (퐥,퐭)=(ퟐퟓ,ퟒ) 퐂 퐨 퐬 퐢 퐧 퐞 퐒 퐢 퐦 퐢 퐥 퐚 퐫 퐢 퐭 퐲 퐓퐚퐬퐤ퟎ퐓퐚퐬퐤ퟐ퐓퐚퐬퐤ퟒ퐓퐚퐬퐤ퟓ퐓퐚퐬퐤ퟗ 퐌 퐚 퐠 퐧 퐢 퐭 퐮 퐝 퐞 퐑 퐚 퐭 퐢 퐨 (퐥,퐭) ퟎ.ퟗퟐ ퟎ.ퟗퟒ ퟎ.ퟗퟔ ퟎ.ퟗퟖ ퟏ.ퟎ ퟎ.ퟗퟓ ퟏ.ퟎ ퟏ.ퟎퟓ ퟏ.ퟏퟎ ퟏ.ퟏퟓ ퟏ.ퟐퟎ Figure 5: (a) The cosine similarity/magnitude ratio between the linear approximation using ̃ A l,t and actual latent activation, under change in camera orientation. (b) Overlap between subspaces spanned by 16 top right singular vectors of ̃ A l,t matrices, obtained from 25 random inputs across 5 tasks. 8 Table 1: LIBERO-10 [22] success rates on Cosmos-Policy 2B [2]. Task i → Task j denotes all P l,t , ̃ A l,t , ̃ B l,t matrices, e z l,t vectors are computed with task i, and used to steer task j. 30 trials/task. PerturbationTasksNo Steering Prompt Steering [72] ActAdd [41] WA-LQR Camera Orientation Task 0→ Task 033.3%±8.6% 40.0%±8.9% 43.3%±9.0% 60.0%±8.9% Task 0→ Task 263.3%±8.8% 60.0%±8.9% 73.3%±8.1% 70.0%±8.4% Task 0→ Task 410.0%±5.5% 6.7%±4.6%16.7%±6.8% 13.3%±6.2% Task 0→ Task 556.7%±9.0% 70.0%±8.4% 56.7%±9.0% 83.3%±6.8% Task 0→ Task 966.7%±8.6% 56.7%±9.0% 56.7%±9.0% 70.0%±8.4% Average46.0%±4.1% 46.7%±4.1% 49.3%±4.1% 59.3%±4.0% Initial Gripper Position Task 1→ Task 146.7%±9.1% 46.7%±9.1% 66.7%±8.6% 60.0%±8.9% Task 1→ Task 283.3%±6.8% 83.3%±6.8% 80.0%±7.3% 100.0%±0.0% Task 1→ Task 346.7%±9.1% 53.3%±9.1% 56.7%±9.0% 60.0%±8.9% Task 1→ Task 780.0%±7.3% 63.3%±8.8% 60.0%±8.9% 83.3%±6.8% Task 1→ Task 950.0%±9.1% 56.7%±9.0% 53.3%±9.1% 60.0%±8.9% Average61.3%±4.0% 60.7%±4.0% 63.3%±3.9% 72.7%±3.6% Camera Gaussian Noise Task 6→ Task 03.3%±3.3%0.0%±0.0%43.3%±9.0% 33.3%±8.6% Task 6→ Task 153.3%±9.1% 0.0%±0.0%70.0%±8.4% 73.3%±8.1% Task 6→ Task 410.0%±5.5% 0.0%±0.0%80.0%±7.3% 53.3%±9.1% Task 6→ Task 636.7%±8.8% 3.3%±3.3%76.7%±7.7% 73.3%±8.1% Task 6→ Task 730.0%±8.4% 0.0%±0.0%66.7%±8.6% 60.0%±8.9% Average26.7%±3.6% 0.7%±0.7%67.3%±3.8% 58.7%±4.0% strating that the latent dynamics are well captured by a first-order approximation. Furthermore, we find that the matrices ̃ A l,t are highly similar across different inputs. Motivated by the concentration of variance on the top few singular vectors, we visualize in Fig. 5(b) the overlap between the subspaces spanned by the top 16 right singular vectors of ̃ A l,t over 25 random inputs from 5 different tasks, measured by the mean squared cosine of principal angles (as inspired by a similar metric in [51]). While not shown here, the left singular vectors also demonstrate significant overlap, enabling the reuse of ̃ A l,t computed from a single input for steering model behavior on other tasks. 6.2 Steering to Improve Robustness 200204060 Success rate (Steered Unsteered) 0.2 0.0 0.2 0.4 0.6 0.8 1.0 Hinge Loss r =0.63 WA-LQR ActAdd Figure 7: Hinge loss vs. steering performance across tasks and models, with reported line of best fit (red) and correlation coefficient. We leverage the mechanistic insights from Sec. 4 to inform steering objectives designed to improve robustness to OOD perturbations in Cosmos-Policy 2B [2], DiT4DiT [6], and LingBot-VA [7], using WA-LQR (Sec. 5.2) and a variant of simple activation addition adapted to WAMs (Sec. 5.1) [41]. Specifically, we seek to improve success rate under perturbations to initial gripper position, camera orientation, and Gaussian noise corruption in the image data, applied to LIBERO-10 evaluation tasks [22]. According to the mechanistic analysis, it is in- feasible to construct linear feature directions that apply for all scenes and tasks (see App. E). Thus, we consider groups of activations which we empirically find have shared representation as suggested by our analysis in Sec. 4. For baselines that do not involve finetuning, we compare to the original model under perturbation and a baseline that decomposes the prompt into subtasks, inspired by [72]. 9 Table 2: Success rates on DiT4DiT [6]. TasksNo Steer.ActAddWA-LQR Initial Gripper Position T0→T053.3%±9.1%60.0%±8.9%66.7%±8.6% T0→T169.6%76.7%±7.7%70.0%±8.4% T0→T270.0%±8.4%63.3%±8.8%73.3%±8.1% T0→T370.0%±8.4%63.3%±8.8%76.7%±7.7% Avg.65.7%±4.3%65.8%±4.3%71.7%±4.1% Camera Gaussian Noise T1→T016.7%±6.8%40.0%±8.9%73.3%±8.1% T1→T120.0%±7.3%63.3%±8.8%43.3%±9.0% T1→T210.0%±5.5%26.7%±8.1%30.0%±8.4% Avg.15.6%±3.8%43.3%±5.2%48.9%±5.3% Table 3: LingBot-VA [7] success rates. TasksNo Steer.ActAddWA-LQR Camera Orientation T0→T065.0%±10.7% 35.0%±10.7% 70.0%±10.2% T0→T255.0%±11.1% 45.0%±11.1% 45.0%±11.1% T0→T415.0%±8.0%10.0%±6.7%30.0%±10.2% T0→T570.0%±10.2% 75.0%±9.7%60.0%±11.0% T0→T935.0%±10.7% 40.0%±11.0% 50.0%±11.2% Avg.48.0%±5.0%41.0%±4.9%51.0%±5.0% Initial Gripper Position T1→T165.0%±10.7% 75.0%±9.7%75.0%±9.7% T1→T270.0%±10.3% 70.0%±10.3% 80.0%±8.9% T1→T370.0%±10.3% 85.0%±8.0%75.0%±9.7% T1→T780.0%±8.9%60.0%±11.0% 80.0%±8.9% T1→T975.0%±9.7%70.0%±10.3% 65.0%±10.7% Avg.72.0%±4.5%72.0%±4.5%75.0%±4.3% Camera Gaussian Noise T6→T035.0%±10.7% 15.0%±8.0%20.0%±8.9% T6→T190.0%±6.7%60.0%±11.0% 95.0%±4.9% T6→T475.0%±9.7%45.0%±11.1% 75.0%±9.7% T6→T685.0%±8.0%55.0%±11.1% 70.0%±10.3% T6→T710.0%±6.7%15.0%±8.0%20.0%±8.9% Avg.59.0%±4.9%38.0%±4.9%56.0%±5.0% Results The grouping and steering results for Cosmos-Policy are summarized in Tab. 1, and we provide qualitative examples of unsteered and steered trajectory rollouts in Fig. 1. We provide a full set of snapshots from quali- tative unsteered and steered rollouts in Fig. 8. WA-LQR is the most effective method in improving robustness to camera orientation and gripper position perturbations, with both steer- ing methods outperforming the unsteered or prompt-steering baselines. As one exception, ActAdd is more effective than WA-LQR on camera Gaussian noise corruption, which we hypothesize results from the strong separability of this feature, enabling simple open-loop con- trol directly in activation space to be viable. We also show in Sec. C that ActAdd is highly sen- sitive to the strength parameter γ, highlighting the benefit of the closed-loop steering provided by WA-LQR, which adaptively modulates the steering magnitude based on alignment with the desired robustness feature direction. On DiT4DiT (Tab. 2), ActAdd yields substantial improvements on Gaussian noise but does not change performance on initial gripper position, while WA-LQR outperforms all baselines and ActAdd on Gaussian noise and initial gripper position. Across evaluations, we find that prompt steering is ineffective, either matching unsteered performance or strongly degrading performance. Task groupings did not emerge on LingBot-VA due to the poor alignment results in Sec. 4, so we in- stead map the groups discovered for Cosmos-Policy onto LingBot-VA. The results are summarized in Tab. 3 (more results are in App. B). Consistent with Sec. 4, we do not observe major robustness improvements from steering for LingBot-VA. In fact, ActAdd often degrades performance due to over-steering, while WA-LQR avoids this because of its closed-loop modulation of steering mag- nitude. We summarize the relationship between separability and steering performance in Fig. 7, where we observe a negative correlation between hinge loss and robustness improvements through steering, supporting the use of hinge loss as a predictor for steering success. Overall, these results suggest that hinge loss is an effective predictor of steering success, that both open-loop and closed- loop activation steering are effective in improving success rate across models with low separability loss, and that closed-loop steering can be more effective than open-loop perturbations. 7 Discussion, Limitations, and Conclusion We investigate mechanistic interpretability and activation steering in WAMs. Our analysis reveals clear linear structure in feature-relevant activations under OOD perturbations for Cosmos-Policy [2] and DiT4DiT [6], but not for LingBot-VA [7]. We further find that linear separability loss is a strong predictor of steering performance. Motivated by these observations, we adapt activation addition to the WAM setting and introduce WA-LQR, a novel activation steering framework that exploits locally linear dynamics within a feature-relevant activation subspace to improve OOD robustness without finetuning, achieving success-rate improvements of up to 41%. Overall, our results suggest that mechanistic control of hidden model representations is promising for improving WAM robustness. Limitations. A key limitation of WA-LQR is its limited applicability across tasks and models. To our knowledge, there is currently no interpretable method for predicting when tasks or environments share transferable representations, requiring mechanistic analysis on a per-setting basis. We also ob- serve strong dependence on model architecture. These results highlight the need to better understand how steerability emerges in robotics foundation models and motivate the design of WAMs that are both steerable and able to preserve representations from their base foundation models. 10 Camera pert. Gaussian pert.Gripper pert. WA-LQR Unsteered “put both the alphabet soup and the tomato sauce in the basket” WA-LQR Unsteered WA-LQR Unsteered Cosmos -Policy “put the white mug on the plate and put the chocolate pudding to the right of the plate” “put both the cream cheese box and the butter in the basket” Camera pert. Gaussian pert.Gripper pert. “turn on the stove and put the moka pot on it” “put the black bowl in the bottom drawer of the cabinet and close it” DiT4DiT Gaussian pert. WA-LQR Unsteered WA-LQR Unsteered Gripper pert. WA-LQR Unsteered WA-LQR Unsteered WA-LQR Unsteered “put both the alphabet soup and the tomato sauce in the basket” “put both the cream cheese box and the butter in the basket” “put both the alphabet soup and the tomato sauce in the basket” LingBot -VA Figure 8: Snapshots from rollouts where steering enables task success despite unsteered failure, across perturbations, LIBERO-10 tasks, and WAM architectures. Each rollout shows six equally spaced snapshots, with time increasing from left to right. 11 References [1] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024. [2] M. J. Kim, Y. Gao, T.-Y. Lin, Y.-C. Lin, Y. Ge, G. Lam, P. Liang, S. Song, M.-Y. Liu, C. Finn, et al. Cosmos policy: Fine-tuning video models for visuomotor control and planning. arXiv preprint arXiv:2601.16163, 2026. [3] S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang, et al. World action models are zero-shot policies. arXiv preprint arXiv:2602.15922, 2026. [4] A. Ye, B. Wang, C. Ni, G. Huang, G. Zhao, H. Li, H. Li, J. Li, J. Lv, J. Liu, et al. Gigaworld- policy: An efficient action-centered world–action model. arXiv preprint arXiv:2603.17240, 2026. [5] Z. Zhang, Z. Li, B. Rahmati, R. H. Yang, Y. Ma, A. Rasouli, S. Pakdamansavoji, Y. Wu, L. Zhang, T. Cao, et al. Do world action models generalize better than vlas? a robustness study. arXiv preprint arXiv:2603.22078, 2026. [6] T. Ma, J. Zheng, Z. Wang, C. Jiang, A. Cui, J. Liang, and S. Yang. Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control. arXiv preprint arXiv:2603.10448, 2026. [7] L. Li, Q. Zhang, Y. Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, Y. Shen, and Y. Xu. Causal world modeling for robot control. arXiv preprint arXiv:2601.21998, 2026. [8] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022. [9] B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In Conference on Robot Learning, pages 2165–2183. PMLR, 2023. [10] O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al. Octo: An open-source generalist robot policy. arXiv preprint arXiv:2405.12213, 2024. [11] A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. IEEE, 2024. [12] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. π 0 : A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024. [13] M. J. Kim, C. Finn, and P. Liang. Fine-tuning vision-language-action models: Optimizing speed and success. arXiv preprint arXiv:2502.19645, 2025. [14] K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. Fast: Efficient action tokenization for vision-language-action models. arXiv preprint arXiv:2501.09747, 2025. [15] B. Hou, G. Li, J. Jia, T. An, X. Guo, S. Leng, H. Geng, Y. Ze, T. Harada, P. Torr, et al. World model for robot learning: A comprehensive survey. arXiv preprint arXiv:2605.00080, 2026. 12 [16] S. Wang, J. Shi, Z. Fu, X. He, F. Liu, C. Yang, Y. Zhou, Z. Fei, J. Gong, J. Fu, et al. World action models: The next frontier in embodied ai. arXiv preprint arXiv:2605.12090, 2026. [17] Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel. Learning universal policies via text-guided video generation. Advances in neural information processing systems, 36:9156–9172, 2023. [18] Y. Wen, J. Lin, Y. Zhu, J. Han, H. Xu, S. Zhao, and X. Liang. Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation. Advances in Neural Information Processing Systems, 37:41051–41075, 2024. [19] Y. Hu, Y. Guo, P. Wang, X. Chen, Y.-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen. Video prediction policy: A generalist robot policy with predictive visual representations. arXiv preprint arXiv:2412.14803, 2024. [20] Y. Jia, J. Liu, S. Liu, R. Zhou, W. Yu, Y. Yan, X. Chi, Y. Guo, B. Shi, and S. Zhang. Video2act: A dual-system video diffusion policy with robotic spatio-motional modeling. arXiv preprint arXiv:2512.03044, 2025. [21] L. Li, Q. Zhang, Y. Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, et al. Causal world modeling for robot control. arXiv preprint arXiv:2601.21998, 2026. [22] B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone. Libero: Benchmarking knowl- edge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36:44776–44791, 2023. [23] X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kir- mani, et al. Evaluating real-world robot manipulation policies in simulation. arXiv preprint arXiv:2405.05941, 2024. [24] Z. Wang, Z. Zhou, J. Song, Y. Huang, Z. Shu, and L. Ma. Vlatest: Testing and evaluating vision-language-action models for robotic manipulation. Proceedings of the ACM on Software Engineering, 2(FSE):1615–1638, 2025. [25] W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox. The colosseum: A bench- mark for evaluating generalization for robotic manipulation. arXiv preprint arXiv:2402.08191, 2024. [26] S. Fei, S. Wang, J. Shi, Z. Dai, J. Cai, P. Qian, L. Ji, X. He, S. Zhang, Z. Fei, et al. Libero-plus: In-depth robustness analysis of vision-language-action models. arXiv preprint arXiv:2510.13626, 2025. [27] J. Zhou, K. Ye, J. Liu, T. Ma, Z. Wang, R. Qiu, K.-Y. Lin, Z. Zhao, and J. Liang. Exploring the limits of vision-language-action manipulations in cross-task generalization, 2025. URL https://arxiv.org/abs/2505.15660. [28] J. Liu, F. Gao, B. Wei, X. Chen, Q. Liao, Y. Wu, C. Yu, and Y. Wang. What can rl bring to vla generalization? an empirical study. Advances in Neural Information Processing Systems, 38: 97121–97151, 2026. [29] S. Tan, K. Dou, Y. Zhao, and P. Kr ̈ ahenb ̈ uhl. Interactive post-training for vision-language- action models. arXiv preprint arXiv:2505.17016, 2025. [30] Y. Guo, J. Zhang, X. Chen, X. Ji, Y.-J. Wang, Y. Hu, and J. Chen. Improving vision-language- action model with online reinforcement learning. In 2025 IEEE International Conference on Robotics and Automation (ICRA), pages 15665–15672. IEEE, 2025. [31] T. Yuan, Z. Dong, Y. Liu, and H. Zhao. Fast-wam: Do world action models need test-time future imagination? arXiv preprint arXiv:2603.16666, 2026. 13 [32] L. Bereska and S. Gavves. Mechanistic interpretability for AI safety - a review. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https://openreview.net/ forum?id=ePUVetPKu6. Survey Certification, Expert Certification. [33] L. Sharkey, B. Chughtai, J. Batson, J. Lindsey, J. Wu, L. Bushnaq, N. Goldowsky-Dill, S. Heimersheim, A. Ortega, J. Bloom, et al. Open problems in mechanistic interpretability. arXiv preprint arXiv:2501.16496, 2025. [34] N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, R. Grosse, S. McCandlish, J. Kaplan, D. Amodei, M. Wat- tenberg, and C. Olah. Toy models of superposition. (arXiv:2209.10652), 2022. doi:10.48550/ arXiv.2209.10652. URL http://arxiv.org/abs/2209.10652. arXiv:2209.10652. [35] K. Park, Y. J. Choe, and V. Veitch. The linear representation hypothesis and the geometry of large language models. arXiv preprint arXiv:2311.03658, 2023. [36] S. Marks and M. Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=aajyHYjjsk. [37] A. Zou, L. Phan, S. Chen, J. Campbell, P. Guo, R. Ren, A. Pan, X. Yin, M. Mazeika, A.- K. Dombrowski, et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023. [38] B. W. Lee, I. Padhi, K. N. Ramamurthy, E. Miehling, P. Dognin, M. Nagireddy, and A. Dhu- randhar. Programming refusal with conditional activation steering. In The Thirteenth Inter- national Conference on Learning Representations, 2025. URL https://openreview.net/ forum?id=Oi47wc10sm. [39] S. Dathathri, A. Madotto, J. Lan, J. Hung, E. Frank, P. Molino, J. Yosinski, and R. Liu. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum? id=H1edEyBKDS. [40] K. Li, O. Patel, F. Vi ́ egas, H. Pfister, and M. Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36:41451–41530, 2023. [41] A. M. Turner, L. Thiergart, G. Leech, D. Udell, U. Mini, and M. MacDiarmid. Activation addition: Steering language models without optimization. 2024. [42] N. Rimsky, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. Turner. Steering llama 2 via contrastive activation addition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15504–15522, 2024. [43] A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda. Refusal in language models is mediated by a single direction. Advances in Neural Information Processing Systems, 37:136037–136083, 2024. [44] P. Rodriguez, A. Blaas, M. Klein, L. Zappella, N. Apostoloff, X. Suau, et al. Controlling language and diffusion models by transporting activations. In International Conference on Learning Representations, volume 2025, pages 89812–89855, 2025. [45] Z. Wu, A. Arora, Z. Wang, A. Geiger, D. Jurafsky, C. D. Manning, and C. Potts. Reft: Repre- sentation finetuning for language models. Advances in Neural Information Processing Systems, 37:63908–63962, 2024. [46] H. M. Vu and T. M. Nguyen. Angular steering: Behavior control via rotation in activa- tion space. (arXiv:2510.26243), Oct. 2025. doi:10.48550/arXiv.2510.26243. URL http: //arxiv.org/abs/2510.26243. arXiv:2510.26243. 14 [47] A. Bhargava, C. Witkowski, S.-Z. Looi, and M. Thomson. What’s the magic word? a control theory of llm prompting. arXiv preprint arXiv:2310.04444, 2023. [48] L. Kong, H. Wang, W. Mu, Y. Du, Y. Zhuang, Y. Zhou, Y. Song, R. Zhang, K. Wang, and C. Zhang. Aligning large language models with representation editing: A control perspective. (arXiv:2406.05954), Nov. 2024. doi:10.48550/arXiv.2406.05954. URL http://arxiv.org/ abs/2406.05954. arXiv:2406.05954. [49] E. Cheng and C. A. Alonso. Linearly controlled language generation with performative guarantees. (arXiv:2405.15454), Sept. 2025. doi:10.48550/arXiv.2405.15454. URL http: //arxiv.org/abs/2405.15454. arXiv:2405.15454. [50] D. V. Nguyen, H. M. Vu, N. Y. Pham, L. Zhang, and T. M. Nguyen. Activation steering with a feedback controller. (arXiv:2510.04309), Oct. 2025. doi:10.48550/arXiv.2510.04309. URL http://arxiv.org/abs/2510.04309. arXiv:2510.04309. [51] J. Skifstad, X. A. Yang, and G. Chou. Local linearity of llms enables activation steering via model-based linear optimal control. arXiv preprint arXiv:2604.19018, 2026. [52] P. Rodriguez, M. Klein, E. Gualdoni, V. Maiorca, A. Blaas, L. Zappella, M. Cuturi, and X. Suau. Lineas: End-to-end learning of activation steering with a distributional loss. arXiv preprint arXiv:2503.10679, 2025. [53] S. Facchiano, S. Saravalle, M. Migliarini, E. De Matteis, A. Sampieri, A. Pilzer, E. Rodol ` a, I. Spinelli, L. Franco, and F. Galasso. Video unlearning via low-rank refusal vector. arXiv preprint arXiv:2506.07891, 2025. [54] Y. Ekin and Y. Gandelsman. The unreasonable effectiveness of text embedding interpolation for continuous image steering. arXiv preprint arXiv:2603.17998, 2026. [55] J. Hong, A. Chan, Q. Dai, J. Skifstad, and G. Chou. Activation steering of video generation models via reduced-order linear optimal control. 2026. [56] B. H ̈ aon, K. C. Stocking, I. Chuang, and C. Tomlin. Mechanistic interpretability for steering vision-language-action models. In J. Lim, S. Song, and H.-W. Park, editors, Proceedings of The 9th Conference on Robot Learning, volume 305 of Proceedings of Machine Learning Research, pages 2743–2762. PMLR, 27–30 Sep 2025. URL https://proceedings.mlr. press/v305/haon25a.html. [57] H. Buurmeijer, C. A. Alonso, A. Swann, and M. Pavone. Observing and controlling features in vision-language-action models. arXiv preprint arXiv:2603.05487, 2026. [58] C. Mitra, Y. Luo, R. Saravanan, D. Niu, A. Pai, J. Thomason, T. Darrell, A. Anwar, D. Ra- manan, and R. Herzig. Mechanistic finetuning of vision-language-action models via few-shot demonstrations. arXiv preprint arXiv:2511.22697, 2025. [59] A. Swann, L. McGranahan, H. Buurmeijer, M. Kennedy I, and M. Schwager.Sparse autoencoders reveal interpretable and steerable features in vla models. arXiv preprint arXiv:2603.19183, 2026. [60] H. Zhang, M. Xu, A. Dhafer, S. Yue, H. Dong, and Z. D. Hao. Embodied interpretability: Link- ing causal understanding to generalization in vision-language-action models. arXiv preprint arXiv:2605.00321, 2026. [61] B. Grant, X. Zhao, and P. Wang. Not all features are created equal: A mechanistic study of vision-language-action models. arXiv preprint arXiv:2603.19233, 2026. [62] M. A. Khan, N. Boskov, F. M. Anwar, and M. A. Khan. Controlling vision–language–action policies through sparse latent directions. In Mechanistic Interpretability Workshop at NeurIPS 2025. 15 [63] M. M. Miao, S. Kim, B. Yang, and L. Ungar. Contrastive conceptor activation steering (coast): Unlocking vision-language-action models through hidden states. arXiv preprint arXiv:2605.17144, 2026. [64] J. P. Hespanha. Linear systems theory. Princeton university press, 2018. [65] W. Peebles and S. Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023. [66] S. Marks and M. Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. arXiv preprint arXiv:2310.06824, Aug. 2024. doi:10.48550/arXiv.2310.06824. URL http://arxiv.org/abs/2310.06824. [67] C. Huang, M. M. Zhang, R. Azarcon, G. Chou, and Z. Kira. Maps: Preserving vision-language representations via module-wise proximity scheduling for better vision-language-action gener- alization. arXiv preprint arXiv:2511.19878, Nov. 2025. doi:10.48550/arXiv.2511.19878. URL http://arxiv.org/abs/2511.19878. [68] V. Kecman. Support Vector Machines – An Introduction, page 1–47. Springer, Berlin, Heidel- berg, 2005. ISBN 9783540323846. doi:10.1007/10984697 1. URL https://doi.org/10. 1007/10984697_1. [69] N. Halko, P.-G. Martinsson, and J. A. Tropp. Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions. SIAM review, 53(2):217–288, 2011. [70] J. B. Rawlings, D. Q. Mayne, M. Diehl, et al. Model predictive control: theory, computation, and design, volume 2. Nob Hill Publishing Madison, WI, 2020. [71] J. Fang and G. Chou. Safe large-scale robust nonlinear mpc in milliseconds via reachability- constrained system level synthesis on the gpu. arXiv preprint arXiv:2604.07644, 2026. [72] Z. Wang, Y. Chen, Y. Liu, J. Ye, P. Chen, C. Lu, S. Liu, B. Yu, and J. Jia.Vp- vla: Visual prompting as an interface for vision-language-action models. arXiv preprint arXiv:2603.22003, May 2026. doi:10.48550/arXiv.2603.22003. URL http://arxiv.org/ abs/2603.22003. arXiv:2603.22003. 16 Appendices In the following, we provide an overview of our appendices. In App. A, we describe how contrastive vectors are constructed for Gaussian noise, camera perturbations, and gripper-position perturbations, including the definitions of the desirable and undesirable activation sets used for steering. In App. B, we provide supplemental experimental results for LingBot-VA, including additional WA-LQR evaluations on camera orientation, initial gripper-position, and Gaussian-noise perturbations in both the action and video modules. In App. C, we provide a parameter sensitivity analysis on ActAdd. In App. D, we provide additional experimental details for the mechanistic study and for the model evaluations, including activation collection, dimensionality reduction, SVM fitting, and LQR steer- ing implementation details. In App. E, we provide the complete low-dimensional separability results across all LIBERO-10 tasks for all models, including both task-level and pairwise separability plots. In App. F, we provide full model separation plots across layers for Cosmos-Policy and DiT4DiT. A Contrastive Vectors To construct contrastive vectors we define a desirable or target set of inputsD + , with rollouts em- blematic of behavior we seek to induce via steering, and undesirable set of inputs D − , which are perturbed by some nuisance and result in failed rollouts. The definition of this set is consistent between models, for each perturbation. For Gaussian noise, we consider noised and clean camera images, i.e.,D + =Clean inputs and D − = Noised inputs. For pairs of inputs in ξ + ∈ D + and ξ − ∈ D − , we perform a model forward pass on each input and collect their corresponding activations x + ,x − . Similarly, the camera perturbation contrastive input sets are defined asD + =No perturbation andD − =Perturbed camera view. The corresponding activations x + ,x − are collected from a forward pass of the model. For gripper position perturbation, the definition ofD + andD − differs slightly. Rather than defining the sets as unperturbed and perturbed, respectively, we letD + be inputs corresponding to successful rollouts and D − be inputs corresponding to unsuccessful inputs, both under gripper perturbation. This is to account for an observed variable sensitivity to different perturbations, i.e., directly includ- ing perturbed inputs in the negative dataset would result in many successful rollouts in the negative activations, resulting intuitively in “steering away” from desired behavior. Note, however, that we do not observe a meaningful difference in the mechanistic analysis when making the distinction between successful and unsuccessful rollouts in our dataset configuration (see App. D.1). Given sets of contrastive vectorsx + k andx − k , we compute the contrastive direction simply as the difference between pairs of positive and negative activations, with different pooling and processing as described in Sec. 5. B Supplemental Experimental Results on LingBot-VA Although we show some results of WA-LQR on LingBot in Table 3, we conducted a more compre- hensive evaluation of WA-LQR applied in both the action module and the video module separately. Specifically, relative to Table 3, we further evaluate on Gaussian noise perturbations, on more task transfers, and on steering in the video module (denoted “(Video)” in Table 4). To allow a fair comparison with Cosmos, we evaluate our method on LingBot-VA on the same set of tasks and perturbations as shown in Table 4. We did not observe significant effectiveness of steering, and this is supported by the result of our mechanistic analysis on the extracted activations from LingBot- VA’s action and video modules, where it is shown that there is poor linear separability (with high numerical classification losses, especially relative to Cosmos), as presented in Fig 15, 16, and 17. 17 Table 4: LIBERO-10 [22] success rates on LingBot-VA [7].Task i → Task j denotes all P l,t , ̃ A l,t , ̃ B l,t matrices, e z l,t vectors are computed with task i, and used to steer task j. 20 trial- s/task. PerturbationTasksNo Steering ActAddOurs (Action) Ours (Video) Camera Orientation Task 0→ Task 065.0%±10.7% 35.0%±10.7% 70.0%±10.3% 65.0%±10.7% Task 0→ Task 255.0%±11.1% 45.0%±11.1% 45.0%±11.1% 55.0%±11.1% Task 0→ Task 415.0%±8.0% 10.0%±6.7% 30.0%±10.3% 10.0%±6.7% Task 0→ Task 570.0%±10.3% 75.0%±9.7% 60.0%±11.0% 65.0%±10.7% Task 0→ Task 935.0%±10.7% 40.0%±11.0% 50.0%±11.2% 60.0%±11.0% Average48.0%±5.0% 41.0%±4.9% 51.0%±5.0% 51.0%±5.0% Initial Gripper Position Task 1→ Task 165.0%±10.7% 75.0%±9.7% 75.0%±9.7% 70.0%±10.3% Task 1→ Task 270.0%±10.3% 70.0%±10.3% 80.0%±8.9% 70.0%±10.3% Task 1→ Task 370.0%±10.3% 85.0%±8.0% 75.0%±9.7% 85.0%±8.0% Task 1→ Task 780.0%±8.9% 60.0%±11.0% 80.0%±8.9% 80.0%±8.9% Task 1→ Task 975.0%±9.7% 70.0%±10.3% 65.0%±10.7% 70.0%±10.3% Average72.0%±4.5% 72.0%±4.5% 75.0%±4.3% 75.0%±4.3% Camera Gaussian Noise Task 6→ Task 035.0%±10.7% 15.0%±8.0% 20.0%±8.9% 65.0%±10.7% Task 6→ Task 190.0%±6.7% 60.0%±11.0% 95.0%±4.9% 20.0%±8.9% Task 6→ Task 475.0%±9.7% 45.0%±11.1% 75.0%±9.7% 100.0%±0.0% Task 6→ Task 685.0%±8.0% 55.0%±11.1% 70.0%±10.3% 75.0%±9.7% Task 6→ Task 710.0%±6.7% 15.0%±8.0% 20.0%±8.9% 20.0%±8.9% Average59.0%±4.9% 38.0%±4.9% 56.0%±5.0% 56.0%±5.0% 18 C ActAdd Sensitivity Analysis We perform a sensitive analysis of the effectiveness of ActAdd on improving the performance of Cosmos-Policy 2B [2] under Gaussian noise perturbation in the sensor input. The steering vectors are acquired from Task06 and used to steer the model over 5 tasks. As shown in Tab. 5, ActAdd demonstrates a sharp peak in success rates at a specific hyperparameter (γ = 0.1), and this perfor- mance boost quickly vanishes with smaller or larger γ. Table 5: Performance across different values of γ under sensor input perturbed with Gaussian noise. γMeanTask06Task00Task01Task04Task07 0.026.6%37%3%53%10%30% 0.0236.2%37%17%67%13%47% 0.0550.8%57%37%73%27%60% 0.07552.6%70%37%70%43%43% 0.167.4%77%43%70%80%67% 0.1532.6%13%10%17%80%43% 0.20.0%0%0%0%0%0% 0.250.0%0%0%0%0%0% 0.50.0%0%0%0%0%0% 19 D Experimental Details D.1 Mechanistic Study To analyze the geometry of each model’s latent activations, we operate directly in the full activation space, using mean-pooled activations across token positions as described in Sec. 4. For Cosmos- Policy, which utilizes a unified single-backbone architecture, we consider the activations across all latent frames. For LingBot-VA and DiT4DiT, which adopt inverse-dynamics style architectures with distinct DiTs for different modalities, we consider the activations from the action-generation DiT. This is consistent with our steering formulation (App. D.2-D.3). To collect contrastive examples for each perturbation type, we conduct a similar procedure to the contrastive vector collection process in App. A. That is, we consider successful vs. unsuc- cessful rollout as D + and D − , respectively, under camera or gripper position perturbation. For Gaussian noise, we let D + = Clean inputs, and D − = Noised inputs. We also eval- uated the gripper and camera perturbation in terms of D + = Without perturbation, and D − =With perturbation, but did not observe a meaningful difference. GivenD + andD − , we compute contrastive directions, and conduct Principal Component Analysis on the set of contrastive directions, as described in Sec. 4. To fit the support vector machine (SVM), we initialize the separatingx| ˆw T x + ˆ b = 0 hyperplane as the zero vector and perform gradient descent with the loss function: loss SVM ( ˆw, ˆ b) = (1/2) ˆw ⊤ ˆw + C· X i∈[N ] max(0, 1− y i ( ˆw T ̄x i + ˆ b)),(15) where C is some constant (in our experiments, we set C = 10). Note the second term in Eq. 15 resembles the hinge loss in Eq. 7, evaluated on the intermediate hyperplane parameterized by ˆw, ˆ b. The SVM is always computed in the three-dimensional subspace defined by the top three principal components of the contrastive directions. D.2 Cosmos-Policy Evaluations on LIBERO-10 Cosmos-Policy is a diffusion-based world-action model with a diffusion transformer backbone, where it jointly denoises a robot action chunk and future state predictions (future proprioception, wrist image, and third-person image). In our method, steering is applied at all latent outputs except for the action chunks. To construct contrastive vectors, activations are collected from all 28 trans- former blocks across 5 denoising timesteps for matched pairs of clean and perturbed observations of the same task and scene. For each layer and timestep, a randomized SVD is run on the paired matrix to produce a rank- 64 basis to reduce dimensionality. The compact subspace is used to find the linearized dynamics that approximates how activations propagate from one layer to the next. The LQR gains are pre- computed via backward Riccati recursion, where the control cost grows exponentially with the action chunk index. An LQR hyperparameter search is conducted before settling on the best-performing parameters for all other tasks of the same scene. During inference time, the projected activation error relative to the nominal trajectory is used to inject a correction through the precomputed gains. D.3 LingBot-VA Evaluations on LIBERO-10 LingBot-VA is a mixture-of-transformers (MoT) based world-action model. It first uses a video module to predict future visual frames, which are then passed to a lightweight action module with a smaller hidden dimension to generate robot actions. In our experiments, we apply our method to either the video module or the action module while keeping the other component unchanged. For both modules, we collect activations from the outputs of all 30 transformer blocks. Since the video and action modules use different denoising schedules, we select module-specific denoising timesteps. For the action module, we use timesteps (t∈ 0, 10, 20, 30, 40). For the video module, we 20 use timesteps (t∈ 0, 4, 9, 14, 19). These timesteps are chosen to cover the corresponding denoising trajectory of each module. To reduce dimensionality, we partition the 30 transformer layers into three groups: layers 0–9, 10–19, and 20–29. For each layer partition and denoising timestep, we pool activation differences of contrastive pairs from all layers within the corresponding partition and compute a shared rank- 64 SVD basis. The produced compact subspace for each layer group and timestep pair is used by our LQR injector during evaluation. At evaluation time, the LQR injector follows the same layer partitioning and timestep mapping. For each activation, we project it onto the corresponding low-dimensional SVD subspace, compute the LQR correction in this reduced space, and map the correction back to the full activation space before applying it to the model. D.4 DiT4DiT Evaluations on LIBERO-10 Similar to LingBot-VA, DiT4DiT is a mixture-of-transformers (MoT) based world-action model with distinct video and action modules. In our experiments, we apply our method to the action module, and do not intervene on the video generation module. Otherwise, our evaluation proce- dure structurally follows Cosmos-Policy on all token positions, intervening across all 4 diffusion timesteps and 16 model layers. The dimensionality reduction procedure also follows the other models, with action generation layers partitioned into three groups: layers 0-5, 6-10, and 11-15. For each layer pool and diffusion timestep, we pool contrastive vectors and analogously construct a 64-dimensional SVD basis, which defines the projection onto the low-dimensional latent space. Jacobian construction, LQR, and control syn- thesis follow the same procedure as the other two models. 21 E Complete Low-Dimensional Separability Results We report the feature separation for three considered robustness features: perturbation to initial gripper position, initial camera position, and corruption of camera inputs with Gaussian noise, for all tasks in the LIBERO-10 dataset [22]. The results are summarized in Figs. 9-14. As reported in Sec. 4, we observe linear separation across tasks and perturbations in Cosmos-Policy (Figs. 9-11), but do not observe such separation in LingBot-VA (Figs. 12-17). For Cosmos-Policy, where we observe meaningful separation, we also present pairwise separability plots (Fig. 18-20), generated by aggregating the activations for pairs of tasks and otherwise following the same procedure as before. 22 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Task 0 First Condition positive negative PC1 PC2 PC3 Block 2 avg loss: 0.0000 Best PC1 PC2 PC3 Block 27 avg loss: 0.1645 Last PC1 PC2 PC3 Block 0 avg loss: 0.0000 Task 1 PC1 PC2 PC3 Block 2 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.1715 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Task 2 PC1 PC2 PC3 Block 1 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.1712 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Task 3 PC1 PC2 PC3 Block 1 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.0000 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Task 4 PC1 PC2 PC3 Block 1 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.2411 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Task 5 First PC1 PC2 PC3 Block 2 avg loss: 0.0000 Best PC1 PC2 PC3 Block 27 avg loss: 0.4458 Last PC1 PC2 PC3 Block 0 avg loss: 0.0000 Task 6 PC1 PC2 PC3 Block 2 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.1190 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Task 7 PC1 PC2 PC3 Block 2 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.0787 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Task 8 PC1 PC2 PC3 Block 1 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.0702 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Task 9 PC1 PC2 PC3 Block 1 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.0183 Figure 9: Cosmos-Policy noise corruption separation for all LIBERO-10 tasks PC1 PC2 PC3 Block 0 avg loss: 0.8049 Task 0 First Condition positive negative PC1 PC2 PC3 Block 19 avg loss: 0.0979 Best PC1 PC2 PC3 Block 27 avg loss: 0.6671 Last PC1 PC2 PC3 Block 0 avg loss: 0.7786 Task 1 PC1 PC2 PC3 Block 17 avg loss: 0.1119 PC1 PC2 PC3 Block 27 avg loss: 0.5929 PC1 PC2 PC3 Block 0 avg loss: 0.6296 Task 2 PC1 PC2 PC3 Block 1 avg loss: 0.3670 PC1 PC2 PC3 Block 27 avg loss: 0.7934 PC1 PC2 PC3 Block 0 avg loss: 0.6728 Task 3 PC1 PC2 PC3 Block 21 avg loss: 0.6263 PC1 PC2 PC3 Block 27 avg loss: 0.8211 PC1 PC2 PC3 Block 0 avg loss: 0.6275 Task 4 PC1 PC2 PC3 Block 16 avg loss: 0.1449 PC1 PC2 PC3 Block 27 avg loss: 0.5281 PC1 PC2 PC3 Block 0 avg loss: 0.4680 Task 5 First PC1 PC2 PC3 Block 12 avg loss: 0.3881 Best PC1 PC2 PC3 Block 27 avg loss: 0.6923 Last PC1 PC2 PC3 Block 0 avg loss: 0.6096 Task 6 PC1 PC2 PC3 Block 20 avg loss: 0.2609 PC1 PC2 PC3 Block 27 avg loss: 0.5963 PC1 PC2 PC3 Block 0 avg loss: 0.7709 Task 7 PC1 PC2 PC3 Block 12 avg loss: 0.1952 PC1 PC2 PC3 Block 27 avg loss: 0.4725 PC1 PC2 PC3 Block 0 avg loss: 0.6446 Task 8 PC1 PC2 PC3 Block 10 avg loss: 0.6777 PC1 PC2 PC3 Block 27 avg loss: 0.7774 PC1 PC2 PC3 Block 0 avg loss: 0.6799 Task 9 PC1 PC2 PC3 Block 2 avg loss: 0.4546 PC1 PC2 PC3 Block 27 avg loss: 0.8781 Figure 10: Cosmos-Policy camera perturbation separation for all LIBERO-10 tasks. 23 PC1 PC2 PC3 Block 0 avg loss: 0.3878 Task 0 First Condition positive negative PC1 PC2 PC3 Block 6 avg loss: 0.2885 Best PC1 PC2 PC3 Block 27 avg loss: 0.2864 Last PC1 PC2 PC3 Block 0 avg loss: 0.1746 Task 1 PC1 PC2 PC3 Block 25 avg loss: 0.0998 PC1 PC2 PC3 Block 27 avg loss: 0.1084 PC1 PC2 PC3 Block 0 avg loss: 0.1679 Task 2 PC1 PC2 PC3 Block 26 avg loss: 0.1106 PC1 PC2 PC3 Block 27 avg loss: 0.1718 PC1 PC2 PC3 Block 0 avg loss: 0.1397 Task 3 PC1 PC2 PC3 Block 25 avg loss: 0.1365 PC1 PC2 PC3 Block 27 avg loss: 0.1351 PC1 PC2 PC3 Block 0 avg loss: 0.3931 Task 4 PC1 PC2 PC3 Block 3 avg loss: 0.3329 PC1 PC2 PC3 Block 27 avg loss: 0.4239 PC1 PC2 PC3 Block 0 avg loss: 0.0829 Task 5 First PC1 PC2 PC3 Block 25 avg loss: 0.0415 Best PC1 PC2 PC3 Block 27 avg loss: 0.0616 Last PC1 PC2 PC3 Block 0 avg loss: 0.2048 Task 6 PC1 PC2 PC3 Block 26 avg loss: 0.1345 PC1 PC2 PC3 Block 27 avg loss: 0.2082 PC1 PC2 PC3 Block 0 avg loss: 0.3548 Task 7 PC1 PC2 PC3 Block 25 avg loss: 0.2738 PC1 PC2 PC3 Block 27 avg loss: 0.3276 PC1 PC2 PC3 Block 0 avg loss: 0.6899 Task 8 PC1 PC2 PC3 Block 25 avg loss: 0.6883 PC1 PC2 PC3 Block 27 avg loss: 0.6973 PC1 PC2 PC3 Block 0 avg loss: 0.1921 Task 9 PC1 PC2 PC3 Block 21 avg loss: 0.1429 PC1 PC2 PC3 Block 27 avg loss: 0.1595 Figure 11: Cosmos-Policy gripper perturbation separation for all LIBERO-10 tasks. PC1 PC2 PC3 Block 0 avg loss: 0.8571 Task 0 First Condition positive negative PC1 PC2 PC3 Block 16 avg loss: 0.8550 Best PC1 PC2 PC3 Block 29 avg loss: 0.8571 Last PC1 PC2 PC3 Block 0 avg loss: 0.8571 Task 1 PC1 PC2 PC3 Block 19 avg loss: 0.8571 PC1 PC2 PC3 Block 29 avg loss: 0.8571 PC1 PC2 PC3 Block 0 avg loss: 0.8770 Task 2 PC1 PC2 PC3 Block 2 avg loss: 0.8462 PC1 PC2 PC3 Block 29 avg loss: 0.8462 PC1 PC2 PC3 Block 0 avg loss: 0.9271 Task 3 PC1 PC2 PC3 Block 20 avg loss: 0.9189 PC1 PC2 PC3 Block 29 avg loss: 0.9189 PC1 PC2 PC3 Block 0 avg loss: 0.8859 Task 4 PC1 PC2 PC3 Block 6 avg loss: 0.8571 PC1 PC2 PC3 Block 29 avg loss: 0.8571 PC1 PC2 PC3 Block 0 avg loss: 0.9715 Task 5 First PC1 PC2 PC3 Block 18 avg loss: 0.9612 Best PC1 PC2 PC3 Block 29 avg loss: 0.9655 Last PC1 PC2 PC3 Block 0 avg loss: 0.9104 Task 6 PC1 PC2 PC3 Block 3 avg loss: 0.8889 PC1 PC2 PC3 Block 29 avg loss: 0.9200 PC1 PC2 PC3 Block 0 avg loss: 0.7310 Task 7 PC1 PC2 PC3 Block 1 avg loss: 0.6923 PC1 PC2 PC3 Block 29 avg loss: 0.6923 PC1 PC2 PC3 Block 0 avg loss: 0.9546 Task 8 PC1 PC2 PC3 Block 13 avg loss: 0.9372 PC1 PC2 PC3 Block 29 avg loss: 0.9478 PC1 PC2 PC3 Block 0 avg loss: 0.8571 Task 9 PC1 PC2 PC3 Block 11 avg loss: 0.8571 PC1 PC2 PC3 Block 29 avg loss: 0.8571 Figure 12: LingBot-VA camera perturbation separation for all LIBERO-10 tasks (Action Module). 24 PC1 PC2 PC3 Block 0 avg loss: 0.8597 Task 0 First Condition positive negative PC1 PC2 PC3 Block 10 avg loss: 0.8530 Best PC1 PC2 PC3 Block 29 avg loss: 0.8777 Last PC1 PC2 PC3 Block 0 avg loss: 0.8571 Task 1 PC1 PC2 PC3 Block 6 avg loss: 0.8571 PC1 PC2 PC3 Block 29 avg loss: 0.8571 PC1 PC2 PC3 Block 0 avg loss: 0.9303 Task 2 PC1 PC2 PC3 Block 1 avg loss: 0.9288 PC1 PC2 PC3 Block 29 avg loss: 0.9300 PC1 PC2 PC3 Block 0 avg loss: 0.8235 Task 3 PC1 PC2 PC3 Block 12 avg loss: 0.8235 PC1 PC2 PC3 Block 29 avg loss: 0.8235 PC1 PC2 PC3 Block 0 avg loss: 0.9474 Task 4 PC1 PC2 PC3 Block 15 avg loss: 0.9474 PC1 PC2 PC3 Block 29 avg loss: 0.9474 PC1 PC2 PC3 Block 0 avg loss: 0.8571 Task 5 First PC1 PC2 PC3 Block 15 avg loss: 0.8213 Best PC1 PC2 PC3 Block 29 avg loss: 0.8571 Last PC1 PC2 PC3 Block 0 avg loss: 0.9046 Task 6 PC1 PC2 PC3 Block 16 avg loss: 0.8778 PC1 PC2 PC3 Block 29 avg loss: 0.9012 PC1 PC2 PC3 Block 0 avg loss: 0.6856 Task 7 PC1 PC2 PC3 Block 4 avg loss: 0.6400 PC1 PC2 PC3 Block 29 avg loss: 0.6400 PC1 PC2 PC3 Block 0 avg loss: 0.8767 Task 8 PC1 PC2 PC3 Block 6 avg loss: 0.8738 PC1 PC2 PC3 Block 29 avg loss: 0.9464 PC1 PC2 PC3 Block 0 avg loss: 0.7746 Task 9 PC1 PC2 PC3 Block 1 avg loss: 0.7753 PC1 PC2 PC3 Block 29 avg loss: 0.7843 Figure 13: LingBot-VA gripper perturbation separation for all LIBERO-10 tasks (Action Module). PC1 PC2 PC3 Block 0 avg loss: 0.9907 Task 0 First Condition positive negative PC1 PC2 PC3 Block 14 avg loss: 0.9881 Best PC1 PC2 PC3 Block 29 avg loss: 0.9982 Last PC1 PC2 PC3 Block 0 avg loss: 0.9441 Task 1 PC1 PC2 PC3 Block 27 avg loss: 0.9189 PC1 PC2 PC3 Block 29 avg loss: 0.9189 PC1 PC2 PC3 Block 0 avg loss: 0.8571 Task 2 PC1 PC2 PC3 Block 5 avg loss: 0.8571 PC1 PC2 PC3 Block 29 avg loss: 0.8571 PC1 PC2 PC3 Block 0 avg loss: 0.9911 Task 3 PC1 PC2 PC3 Block 18 avg loss: 0.9874 PC1 PC2 PC3 Block 29 avg loss: 0.9882 PC1 PC2 PC3 Block 0 avg loss: 0.8235 Task 4 PC1 PC2 PC3 Block 6 avg loss: 0.8235 PC1 PC2 PC3 Block 29 avg loss: 0.8235 PC1 PC2 PC3 Block 0 avg loss: 0.9470 Task 5 First PC1 PC2 PC3 Block 25 avg loss: 0.9375 Best PC1 PC2 PC3 Block 29 avg loss: 0.9375 Last PC1 PC2 PC3 Block 0 avg loss: 0.9072 Task 6 PC1 PC2 PC3 Block 13 avg loss: 0.8889 PC1 PC2 PC3 Block 29 avg loss: 0.8889 PC1 PC2 PC3 Block 0 avg loss: 0.9375 Task 7 PC1 PC2 PC3 Block 21 avg loss: 0.9375 PC1 PC2 PC3 Block 29 avg loss: 0.9375 PC1 PC2 PC3 Block 0 avg loss: 0.9741 Task 8 PC1 PC2 PC3 Block 7 avg loss: 0.9697 PC1 PC2 PC3 Block 29 avg loss: 0.9804 PC1 PC2 PC3 Block 0 avg loss: 0.9015 Task 9 PC1 PC2 PC3 Block 27 avg loss: 0.8889 PC1 PC2 PC3 Block 29 avg loss: 0.8889 Figure 14: LingBot-VA noise corruption separation for all LIBERO-10 tasks (Action Module). 25 PC1 PC2 PC3 Block 0 avg loss: 0.8168 Task 0 First Condition positive negative PC1 PC2 PC3 Block 14 avg loss: 0.6073 Best PC1 PC2 PC3 Block 29 avg loss: 0.7372 Last PC1 PC2 PC3 Block 0 avg loss: 0.8549 Task 2 PC1 PC2 PC3 Block 15 avg loss: 0.6008 PC1 PC2 PC3 Block 29 avg loss: 0.6602 PC1 PC2 PC3 Block 0 avg loss: 0.8784 Task 4 PC1 PC2 PC3 Block 21 avg loss: 0.4518 PC1 PC2 PC3 Block 29 avg loss: 0.7223 PC1 PC2 PC3 Block 0 avg loss: 0.8876 Task 5 PC1 PC2 PC3 Block 28 avg loss: 0.7135 PC1 PC2 PC3 Block 29 avg loss: 0.7155 PC1 PC2 PC3 Block 0 avg loss: 0.8542 Task 9 PC1 PC2 PC3 Block 7 avg loss: 0.7163 PC1 PC2 PC3 Block 29 avg loss: 0.7455 Figure 15: LingBot-VA camera orientation perturbation separation for LIBERO-10 task 0, 2, 4, 5, 9 (Video Module). 26 PC1 PC2 PC3 Block 0 avg loss: 0.7879 Task 6 First Condition positive negative PC1 PC2 PC3 Block 26 avg loss: 0.7879 Best PC1 PC2 PC3 Block 29 avg loss: 0.7879 Last PC1 PC2 PC3 Block 0 avg loss: 0.9958 Task 0 PC1 PC2 PC3 Block 28 avg loss: 0.9820 PC1 PC2 PC3 Block 29 avg loss: 0.9859 PC1 PC2 PC3 Block 0 avg loss: 0.8889 Task 1 PC1 PC2 PC3 Block 5 avg loss: 0.8889 PC1 PC2 PC3 Block 29 avg loss: 0.8889 PC1 PC2 PC3 Block 0 avg loss: 0.9773 Task 4 PC1 PC2 PC3 Block 24 avg loss: 0.9140 PC1 PC2 PC3 Block 29 avg loss: 0.9589 PC1 PC2 PC3 Block 0 avg loss: 0.8276 Task 7 PC1 PC2 PC3 Block 3 avg loss: 0.8276 PC1 PC2 PC3 Block 29 avg loss: 0.8276 Figure 16: LingBot-VA noise corruption separation for LIBERO-10 task 0, 1, 4, 6, 7 (Video Mod- ule). 27 PC1 PC2 PC3 Block 0 avg loss: 0.6207 Task 1 First Condition positive negative PC1 PC2 PC3 Block 16 avg loss: 0.6207 Best PC1 PC2 PC3 Block 29 avg loss: 0.6207 Last PC1 PC2 PC3 Block 0 avg loss: 0.9326 Task 2 PC1 PC2 PC3 Block 28 avg loss: 0.9134 PC1 PC2 PC3 Block 29 avg loss: 0.9132 PC1 PC2 PC3 Block 0 avg loss: 0.9249 Task 3 PC1 PC2 PC3 Block 15 avg loss: 0.8848 PC1 PC2 PC3 Block 29 avg loss: 0.8960 PC1 PC2 PC3 Block 0 avg loss: 0.9189 Task 7 PC1 PC2 PC3 Block 21 avg loss: 0.8584 PC1 PC2 PC3 Block 29 avg loss: 0.9032 PC1 PC2 PC3 Block 0 avg loss: 0.8815 Task 9 PC1 PC2 PC3 Block 10 avg loss: 0.8211 PC1 PC2 PC3 Block 29 avg loss: 0.8425 Figure 17: LingBot-VA gripper perturbation separation for LIBERO-10 task 1, 2, 3, 7, 9 (Video Module). 28 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Tasks 6&0 First Cond. / Task task 6 positive task 6 negative task 0 positive task 0 negative PC1 PC2 PC3 Block 1 avg loss: 0.0000 Best PC1 PC2 PC3 Block 27 avg loss: 0.1068 Last PC1 PC2 PC3 Block 0 avg loss: 0.0000 Tasks 6&1 PC1 PC2 PC3 Block 1 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.1012 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Tasks 6&4 PC1 PC2 PC3 Block 2 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.1829 PC1 PC2 PC3 Block 0 avg loss: 0.0000 Tasks 6&7 PC1 PC2 PC3 Block 1 avg loss: 0.0000 PC1 PC2 PC3 Block 27 avg loss: 0.0309 Figure 18: Pairwise separability for Cosmos-Policy with Gaussian noise corruption. 29 PC1 PC2 PC3 Block 0 avg loss: 0.1973 Tasks 1&2 First Cond. / Task task 1 positive task 1 negative task 2 positive task 2 negative PC1 PC2 PC3 Block 23 avg loss: 0.1275 Best PC1 PC2 PC3 Block 27 avg loss: 0.2020 Last PC1 PC2 PC3 Block 0 avg loss: 0.1471 Tasks 1&3 PC1 PC2 PC3 Block 25 avg loss: 0.1207 PC1 PC2 PC3 Block 27 avg loss: 0.1379 PC1 PC2 PC3 Block 0 avg loss: 0.2340 Tasks 1&7 PC1 PC2 PC3 Block 25 avg loss: 0.1592 PC1 PC2 PC3 Block 27 avg loss: 0.1827 PC1 PC2 PC3 Block 0 avg loss: 0.1836 Tasks 1&9 PC1 PC2 PC3 Block 25 avg loss: 0.1498 PC1 PC2 PC3 Block 27 avg loss: 0.1799 Figure 19: Pairwise separability for Cosmos-Policy with gripper position perturbation. 30 PC1 PC2 PC3 Block 0 avg loss: 0.6072 Tasks 0&2 First Cond. / Task task 0 positive task 0 negative task 2 positive task 2 negative PC1 PC2 PC3 Block 21 avg loss: 0.2191 Best PC1 PC2 PC3 Block 27 avg loss: 0.3805 Last PC1 PC2 PC3 Block 0 avg loss: 0.5156 Tasks 0&4 PC1 PC2 PC3 Block 22 avg loss: 0.1272 PC1 PC2 PC3 Block 27 avg loss: 0.3306 PC1 PC2 PC3 Block 0 avg loss: 0.6505 Tasks 0&5 PC1 PC2 PC3 Block 24 avg loss: 0.1662 PC1 PC2 PC3 Block 27 avg loss: 0.4288 PC1 PC2 PC3 Block 0 avg loss: 0.6539 Tasks 0&9 PC1 PC2 PC3 Block 21 avg loss: 0.1858 PC1 PC2 PC3 Block 27 avg loss: 0.4459 Figure 20: Pairwise separability for Cosmos-Policy with camera position perturbation. 31 F Full Model Separation For completeness, we report the full mechanistic separation plots across all layers in the following. PC1 PC2 PC3 loss: 0.0000 Block 0 Condition positive negative PC1 PC2 PC3 loss: 0.0000 Block 1 PC1 PC2 PC3 loss: 0.0000 Block 2 PC1 PC2 PC3 loss: 0.0000 Block 3 PC1 PC2 PC3 loss: 0.0000 Block 4 PC1 PC2 PC3 loss: 0.0000 Block 5 PC1 PC2 PC3 loss: 0.0000 Block 6 PC1 PC2 PC3 loss: 0.0000 Block 7 PC1 PC2 PC3 loss: 0.0000 Block 8 PC1 PC2 PC3 loss: 0.0000 Block 9 PC1 PC2 PC3 loss: 0.0000 Block 10 PC1 PC2 PC3 loss: 0.0469 Block 11 PC1 PC2 PC3 loss: 0.0000 Block 12 PC1 PC2 PC3 loss: 0.2298 Block 13 PC1 PC2 PC3 loss: 0.1853 Block 14 PC1 PC2 PC3 loss: 0.0835 Block 15 Task 0 DiT blocks 0 15 Figure 21: Cosmos-Policy Task 0 noise corruption 32 PC1 PC2 PC3 loss: 0.0261 Block 16 Condition positive negative PC1 PC2 PC3 loss: 0.0348 Block 17 PC1 PC2 PC3 loss: 0.0039 Block 18 PC1 PC2 PC3 loss: 0.0000 Block 19 PC1 PC2 PC3 loss: 0.0027 Block 20 PC1 PC2 PC3 loss: 0.0000 Block 21 PC1 PC2 PC3 loss: 0.0200 Block 22 PC1 PC2 PC3 loss: 0.0000 Block 23 PC1 PC2 PC3 loss: 0.0000 Block 24 PC1 PC2 PC3 loss: 0.0063 Block 25 PC1 PC2 PC3 loss: 0.1182 Block 26 PC1 PC2 PC3 loss: 0.1645 Block 27 Task 0 DiT blocks 16 27 Figure 22: Cosmos-Policy Task 0 noise corruption 33 PC1 PC2 PC3 loss: 0.0000 Block 0 Condition positive negative PC1 PC2 PC3 loss: 0.0000 Block 1 PC1 PC2 PC3 loss: 0.0000 Block 2 PC1 PC2 PC3 loss: 0.0000 Block 3 PC1 PC2 PC3 loss: 0.0000 Block 4 PC1 PC2 PC3 loss: 0.0000 Block 5 PC1 PC2 PC3 loss: 0.0000 Block 6 PC1 PC2 PC3 loss: 0.0000 Block 7 PC1 PC2 PC3 loss: 0.0000 Block 8 PC1 PC2 PC3 loss: 0.0000 Block 9 PC1 PC2 PC3 loss: 0.0000 Block 10 PC1 PC2 PC3 loss: 0.0000 Block 11 PC1 PC2 PC3 loss: 0.0000 Block 12 PC1 PC2 PC3 loss: 0.1949 Block 13 PC1 PC2 PC3 loss: 0.4516 Block 14 PC1 PC2 PC3 loss: 0.1000 Block 15 Task 1 DiT blocks 0 15 Figure 23: Cosmos-Policy Task 1 noise corruption 34 PC1 PC2 PC3 loss: 0.0002 Block 16 Condition positive negative PC1 PC2 PC3 loss: 0.0000 Block 17 PC1 PC2 PC3 loss: 0.0069 Block 18 PC1 PC2 PC3 loss: 0.0000 Block 19 PC1 PC2 PC3 loss: 0.0000 Block 20 PC1 PC2 PC3 loss: 0.0000 Block 21 PC1 PC2 PC3 loss: 0.0000 Block 22 PC1 PC2 PC3 loss: 0.0000 Block 23 PC1 PC2 PC3 loss: 0.0047 Block 24 PC1 PC2 PC3 loss: 0.0492 Block 25 PC1 PC2 PC3 loss: 0.2189 Block 26 PC1 PC2 PC3 loss: 0.1715 Block 27 Task 1 DiT blocks 16 27 Figure 24: Cosmos-Policy Task 1 noise corruption 35 PC1 PC2 PC3 loss: 0.8049 Block 0 Condition positive negative PC1 PC2 PC3 loss: 0.8044 Block 1 PC1 PC2 PC3 loss: 0.7876 Block 2 PC1 PC2 PC3 loss: 0.7814 Block 3 PC1 PC2 PC3 loss: 0.7832 Block 4 PC1 PC2 PC3 loss: 0.7485 Block 5 PC1 PC2 PC3 loss: 0.7564 Block 6 PC1 PC2 PC3 loss: 0.7584 Block 7 PC1 PC2 PC3 loss: 0.7442 Block 8 PC1 PC2 PC3 loss: 0.7763 Block 9 PC1 PC2 PC3 loss: 0.7719 Block 10 PC1 PC2 PC3 loss: 0.7869 Block 11 PC1 PC2 PC3 loss: 0.3500 Block 12 PC1 PC2 PC3 loss: 0.4208 Block 13 PC1 PC2 PC3 loss: 0.4855 Block 14 PC1 PC2 PC3 loss: 0.3126 Block 15 Task 0 DiT blocks 0 15 Figure 25: Cosmos-Policy Task 0 camera perturbation 36 PC1 PC2 PC3 loss: 0.1219 Block 16 Condition positive negative PC1 PC2 PC3 loss: 0.2819 Block 17 PC1 PC2 PC3 loss: 0.1020 Block 18 PC1 PC2 PC3 loss: 0.0980 Block 19 PC1 PC2 PC3 loss: 0.1698 Block 20 PC1 PC2 PC3 loss: 0.1166 Block 21 PC1 PC2 PC3 loss: 0.1322 Block 22 PC1 PC2 PC3 loss: 0.1455 Block 23 PC1 PC2 PC3 loss: 0.1614 Block 24 PC1 PC2 PC3 loss: 0.3327 Block 25 PC1 PC2 PC3 loss: 0.9824 Block 26 PC1 PC2 PC3 loss: 0.6671 Block 27 Task 0 DiT blocks 16 27 Figure 26: Cosmos-Policy Task 0 camera perturbation 37 PC1 PC2 PC3 loss: 0.7785 Block 0 Condition positive negative PC1 PC2 PC3 loss: 0.6822 Block 1 PC1 PC2 PC3 loss: 0.7031 Block 2 PC1 PC2 PC3 loss: 0.7157 Block 3 PC1 PC2 PC3 loss: 0.7269 Block 4 PC1 PC2 PC3 loss: 0.6185 Block 5 PC1 PC2 PC3 loss: 0.6497 Block 6 PC1 PC2 PC3 loss: 0.5879 Block 7 PC1 PC2 PC3 loss: 0.6774 Block 8 PC1 PC2 PC3 loss: 0.6694 Block 9 PC1 PC2 PC3 loss: 0.7148 Block 10 PC1 PC2 PC3 loss: 0.7620 Block 11 PC1 PC2 PC3 loss: 0.3978 Block 12 PC1 PC2 PC3 loss: 0.3679 Block 13 PC1 PC2 PC3 loss: 0.6062 Block 14 PC1 PC2 PC3 loss: 0.6071 Block 15 Task 1 DiT blocks 0 15 Figure 27: Cosmos-Policy Task 1 camera perturbation 38 PC1 PC2 PC3 loss: 0.1164 Block 16 Condition positive negative PC1 PC2 PC3 loss: 0.1119 Block 17 PC1 PC2 PC3 loss: 0.1945 Block 18 PC1 PC2 PC3 loss: 0.1451 Block 19 PC1 PC2 PC3 loss: 0.2587 Block 20 PC1 PC2 PC3 loss: 0.2586 Block 21 PC1 PC2 PC3 loss: 0.2669 Block 22 PC1 PC2 PC3 loss: 0.3127 Block 23 PC1 PC2 PC3 loss: 0.9066 Block 24 PC1 PC2 PC3 loss: 0.9851 Block 25 PC1 PC2 PC3 loss: 0.7562 Block 26 PC1 PC2 PC3 loss: 0.5930 Block 27 Task 1 DiT blocks 16 27 Figure 28: Cosmos-Policy Task 1 camera perturbation 39 PC1 PC2 PC3 Block 0 avg loss: 0.8877 Task 0 First PC1 PC2 PC3 Block 4 avg loss: 0.0000 Best PC1 PC2 PC3 Block 15 avg loss: 0.0000 Last PC1 PC2 PC3 Block 0 avg loss: 0.7908 Task 1 PC1 PC2 PC3 Block 5 avg loss: 0.0000 PC1 PC2 PC3 Block 15 avg loss: 0.4720 PC1 PC2 PC3 Block 0 avg loss: 0.7742 Task 2 PC1 PC2 PC3 Block 4 avg loss: 0.0000 PC1 PC2 PC3 Block 15 avg loss: 0.0000 PC1 PC2 PC3 Block 0 avg loss: 0.8395 Task 3 PC1 PC2 PC3 Block 8 avg loss: 0.0000 PC1 PC2 PC3 Block 15 avg loss: 0.3054 PC1 PC2 PC3 Block 0 avg loss: 0.8239 Task 4 PC1 PC2 PC3 Block 14 avg loss: 0.0240 PC1 PC2 PC3 Block 15 avg loss: 0.2146 PC1 PC2 PC3 Block 0 avg loss: 0.8035 Task 5 PC1 PC2 PC3 Block 8 avg loss: 0.0000 PC1 PC2 PC3 Block 15 avg loss: 0.7159 PC1 PC2 PC3 Block 0 avg loss: 0.8144 Task 6 PC1 PC2 PC3 Block 4 avg loss: 0.0000 PC1 PC2 PC3 Block 15 avg loss: 0.0000 PC1 PC2 PC3 Block 0 avg loss: 0.8688 Task 7 PC1 PC2 PC3 Block 8 avg loss: 0.4721 PC1 PC2 PC3 Block 15 avg loss: 0.6711 PC1 PC2 PC3 Block 0 avg loss: 0.9393 Task 8 PC1 PC2 PC3 Block 14 avg loss: 0.3843 PC1 PC2 PC3 Block 15 avg loss: 0.3244 PC1 PC2 PC3 Block 0 avg loss: 0.9560 Task 9 Condition positive negative PC1 PC2 PC3 Block 7 avg loss: 0.2993 PC1 PC2 PC3 Block 15 avg loss: 0.3259 Figure 29: Feature separability: DiT4DiT Gaussian noise corruption for tasks 0-2. 40 PC1 PC2 PC3 Block 0 avg loss: 0.7222 Task 0 First PC1 PC2 PC3 Block 5 avg loss: 0.5830 Best PC1 PC2 PC3 Block 15 avg loss: 0.6357 Last PC1 PC2 PC3 Block 0 avg loss: 0.8488 Task 1 PC1 PC2 PC3 Block 3 avg loss: 0.3751 PC1 PC2 PC3 Block 15 avg loss: 0.4193 PC1 PC2 PC3 Block 0 avg loss: 0.7213 Task 2 PC1 PC2 PC3 Block 8 avg loss: 0.6472 PC1 PC2 PC3 Block 15 avg loss: 0.6676 PC1 PC2 PC3 Block 0 avg loss: 0.5951 Task 3 PC1 PC2 PC3 Block 9 avg loss: 0.5951 PC1 PC2 PC3 Block 15 avg loss: 0.5951 PC1 PC2 PC3 Block 0 avg loss: 0.5197 Task 4 PC1 PC2 PC3 Block 14 avg loss: 0.3090 PC1 PC2 PC3 Block 15 avg loss: 0.3284 PC1 PC2 PC3 Block 0 avg loss: 0.4237 Task 5 PC1 PC2 PC3 Block 13 avg loss: 0.3815 PC1 PC2 PC3 Block 15 avg loss: 0.4237 PC1 PC2 PC3 Block 0 avg loss: 0.8030 Task 6 PC1 PC2 PC3 Block 8 avg loss: 0.6859 PC1 PC2 PC3 Block 15 avg loss: 0.7754 PC1 PC2 PC3 Block 0 avg loss: 0.9331 Task 7 PC1 PC2 PC3 Block 11 avg loss: 0.7988 PC1 PC2 PC3 Block 15 avg loss: 0.8864 PC1 PC2 PC3 Block 0 avg loss: 0.5503 Task 8 PC1 PC2 PC3 Block 7 avg loss: 0.5479 PC1 PC2 PC3 Block 15 avg loss: 0.5503 PC1 PC2 PC3 Block 0 avg loss: 0.6610 Task 9 Condition positive negative PC1 PC2 PC3 Block 7 avg loss: 0.6263 PC1 PC2 PC3 Block 15 avg loss: 0.6610 Figure 30: Feature separability: DiT4DiT gripper perturbation for tasks 0-3. 41 PC1 PC2 PC3 loss: 0.8877 Block 0 Condition positive negative PC1 PC2 PC3 loss: 0.8995 Block 1 PC1 PC2 PC3 loss: 0.7761 Block 2 PC1 PC2 PC3 loss: 0.7490 Block 3 PC1 PC2 PC3 loss: 0.0000 Block 4 PC1 PC2 PC3 loss: 0.0000 Block 5 PC1 PC2 PC3 loss: 0.0000 Block 6 PC1 PC2 PC3 loss: 0.0000 Block 7 PC1 PC2 PC3 loss: 0.0006 Block 8 PC1 PC2 PC3 loss: 0.0133 Block 9 PC1 PC2 PC3 loss: 0.0203 Block 10 PC1 PC2 PC3 loss: 0.0124 Block 11 PC1 PC2 PC3 loss: 0.0000 Block 12 PC1 PC2 PC3 loss: 0.0000 Block 13 PC1 PC2 PC3 loss: 0.0000 Block 14 PC1 PC2 PC3 loss: 0.0000 Block 15 Task 0 DiT blocks 0 15 Figure 31: DiT4DiT Task 0 noise corruption 42 PC1 PC2 PC3 loss: 0.7908 Block 0 Condition positive negative PC1 PC2 PC3 loss: 0.7985 Block 1 PC1 PC2 PC3 loss: 0.7052 Block 2 PC1 PC2 PC3 loss: 0.6874 Block 3 PC1 PC2 PC3 loss: 0.0033 Block 4 PC1 PC2 PC3 loss: 0.0000 Block 5 PC1 PC2 PC3 loss: 0.0563 Block 6 PC1 PC2 PC3 loss: 0.1486 Block 7 PC1 PC2 PC3 loss: 0.1936 Block 8 PC1 PC2 PC3 loss: 0.2992 Block 9 PC1 PC2 PC3 loss: 0.3965 Block 10 PC1 PC2 PC3 loss: 0.4297 Block 11 PC1 PC2 PC3 loss: 0.3947 Block 12 PC1 PC2 PC3 loss: 0.4285 Block 13 PC1 PC2 PC3 loss: 0.3617 Block 14 PC1 PC2 PC3 loss: 0.4720 Block 15 Task 1 DiT blocks 0 15 Figure 32: DiT4DiT Task 1 noise corruption PC1 PC2 PC3 Block 0 avg loss: 0.8109 Tasks 1&2 First Cond. / Task task 1 positive task 1 negative task 2 positive task 2 negative PC1 PC2 PC3 Block 5 avg loss: 0.0000 Best PC1 PC2 PC3 Block 15 avg loss: 0.0113 Last PC1 PC2 PC3 Block 0 avg loss: 0.8542 Tasks 1&0 PC1 PC2 PC3 Block 4 avg loss: 0.0000 PC1 PC2 PC3 Block 15 avg loss: 0.0000 PC1 PC2 PC3 Block 0 avg loss: 0.9041 Tasks 3&1 PC1 PC2 PC3 Block 10 avg loss: 0.0117 PC1 PC2 PC3 Block 15 avg loss: 0.0750 Figure 33: Pairwise Gaussian noise corruption feature separation on DiT4DiT. 43 PC1 PC2 PC3 Block 0 avg loss: 0.7503 Tasks 1&0 First Cond. / Task task 1 positive task 1 negative task 0 positive task 0 negative PC1 PC2 PC3 Block 13 avg loss: 0.3592 Best PC1 PC2 PC3 Block 15 avg loss: 0.3732 Last PC1 PC2 PC3 Block 0 avg loss: 0.7431 Tasks 1&2 PC1 PC2 PC3 Block 4 avg loss: 0.3167 PC1 PC2 PC3 Block 15 avg loss: 0.4142 PC1 PC2 PC3 Block 0 avg loss: 0.8160 Tasks 1&3 PC1 PC2 PC3 Block 5 avg loss: 0.4485 PC1 PC2 PC3 Block 15 avg loss: 0.4959 Figure 34: Pairwise gripper perturbation feature separation on DiT4DiT. 44