Paper deep dive
Learning Additively Compositional Latent Actions for Embodied AI
Hangxing Wei, Xiaoyu Chen, Chuheng Zhang, Tim Pearce, Jianyu Chen, Alex Lamb, Li Zhao, Jiang Bian
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/10/2026, 2:02:18 AM
Summary
AC-LAM (Additively Compositional Latent Action Model) is a novel framework for embodied AI that enforces a scene-wise additive composition prior on latent action spaces. By regularizing latent actions to satisfy zi,k ≈ zi,j + zj,k, the model improves displacement calibration, suppresses non-compositional noise (like scene background or future leakage), and enhances downstream policy learning performance compared to state-of-the-art latent action models.
Entities (5)
Relation Signals (3)
AC-LAM → enforces → additive composition prior
confidence 100% · AC-LAM, which enforces scene-wise additive composition structure over short horizons on the latent action space.
AC-LAM → improves → downstream policy learning
confidence 95% · Using latents generated by AC-LAM as training supervision targets improves downstream policy learning compared to state-of-the-art latent action models
AC-LAM → outperforms → LAPA
confidence 90% · AC-LAM... outperforming state-of-the-art LAMs across simulated and real-world tabletop tasks.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Latent action learning infers pseudo-action labels from visual transitions, providing an approach to leverage internet-scale video for embodied AI. However, most methods learn latent actions without structural priors that encode the additive, compositional structure of physical motion. As a result, latents often entangle irrelevant scene details or information about future observations with true state changes and miscalibrate motion magnitude. We introduce Additively Compositional Latent Action Model (AC-LAM), which enforces scene-wise additive composition structure over short horizons on the latent action space. These AC constraints encourage simple algebraic structure in the latent action space~(identity, inverse, cycle consistency) and suppress information that does not compose additively. Empirically, AC-LAM learns more structured, motion-specific, and displacement-calibrated latent actions and provides stronger supervision for downstream policy learning, outperforming state-of-the-art LAMs across simulated and real-world tabletop tasks.
Tags
Links
- Source: https://arxiv.org/abs/2604.03340v1
- Canonical: https://arxiv.org/abs/2604.03340v1
Trouble viewing inline? Open PDF directly →
Full Text
70,824 characters extracted from source content.
Expand or collapse full text
Learning Additively Compositional Latent Actions for Embodied AI Hangxing Wei Xiaoyu Chen Chuheng Zhang Tim Pearce Jianyu Chen Alex Lamb Li Zhao Jiang Bian Abstract Latent action learning infers pseudo-action labels from visual transitions, providing an approach to leverage internet-scale video for embodied AI. However, most methods learn latent actions without structural priors that encode the additive, compositional structure of physical motion. As a result, latents often entangle irrelevant scene details or information about future observations with true state changes and miscalibrate motion magnitude. We introduce Additively Compositional Latent Action Model (AC-LAM), which enforces scene-wise additive composition structure over short horizons on the latent action space. These AC constraints encourage simple algebraic structure in the latent action space (identity, inverse, cycle consistency) and suppress information that does not compose additively. Empirically, AC-LAM learns more structured, motion-specific, and displacement-calibrated latent actions and provides stronger supervision for downstream policy learning, outperforming state-of-the-art LAMs across simulated and real-world tabletop tasks. Machine Learning, ICML 1 Introduction Latent action learning has emerged as a scalable paradigm in embodied AI, enabling pretraining from internet-scale video by deriving pseudo-action labels from visual transitions (Bruce et al., 2024; Ye et al., 2024). However, existing methods lack structural priors that reflect the compositional nature of physical motion. Consequently, learned latents often entangle irrelevant information (e.g., scene details, future observations) with pure state changes and miscalibrate motion magnitude, weakening transfer, planning, and generalization. As latent actions encode motion semantics, it is natural to expect their norm to reflect motion magnitude. However, lacking explicit structural constraints, latent action learning methods usually fail to capture this relationship. As shown in Figure 1 for a real‑robot trajectory, the latent action norm ‖LAM(o0,ot)‖\|LAM(o_0,o_t)\| from LAPA (Ye et al., 2024) is systematically under‑calibrated for motion magnitude (displacement). Accurately calibrating displacement and enforcing compositional structure for latent actions is thus critical for robust control and generalization. Figure 1: Evolution of the normalized latent action norm ‖LAM(o0,ot)‖\|LAM(o_0,o_t)\| over time intervals t. The figure shows that, compared with baselines, our AC constraints (AC-LAM) induce displacement-calibrated latent actions, effectively capturing the magnitude of the transition from o0o_0 to oto_t. We address these limitations by introducing the Additively Compositional Latent Action Model (AC-LAM), which enforces scene-wise additive composition structure. Concretely, latent actions inferred between observations oi,oj,oko_i,o_j,o_k from the same scene satisfy zik=zij+zjk,z_ik=z_ij+z_jk, aligning the representation with the linear and compositional structure of short-horizon motion semantics. This structural prior regularizes the latent space toward more structured latent actions that compose additively within the same scene. While this prior may not hold for all possible actions, basic motion primitives often compose additively (e.g. ‘move right’ then ‘move up’, can be composed as ‘move diagonally right and up’). Concretely, we show that AC constraints encourage desirable algebraic structure in the latent action space: identity (zii=0z_i=0) and inverse (zji=−zijz_ji=-z_ij) elements, and cycle consistency (zij+zjk+zki=0z_ij+z_jk+z_ki=0). Moreover, under simplified assumptions, the constraints suppress information that does not compose additively: static environment terms and future leakage are suppressed because they violate additive consistency across triples. While linearity is an approximation in practice, AC acts as a regularizer that promotes motion-centric latents and reduces entanglement with non-compositional factors. Based on the AC prior, we propose AC-LAM, a novel LAM that enforces AC constraints through a compositional loss in the forward-dynamics (FDM) model, denoted ℒAC-FDML_AC -FDM. For triples (oi,oj,ok)(o_i,o_j,o_k) from the same trajectory, we reconstruct oko_k by decoding from the summed latents zij+zjkz_ij+z_jk (Eq. 5). The ℒAC-FDML_AC -FDM loss operates on post-VQ continuous embeddings, preserving displacement calibration while retaining the VQ bottleneck. We find ℒAC-FDML_AC -FDM is empirically more stable than a version of the loss implemented in the IDM (Eq. 4), and can be optimized jointly with standard reconstruction and bottleneck terms over latent actions sampled within short horizons. Our empirical results show that AC-LAM learns more interpretable, motion-specific, and displacement-calibrated latent actions: they correlate more strongly with true displacement (as shown in Figure 1), adhere more closely to identity/inverse behavior, and reduce information captured about environment aesthetics and future observations. Using latents generated by AC-LAM as training supervision targets improves downstream policy learning compared to state-of-the-art latent action models across simulated and real-world tabletop tasks. In short, by incorporating a prior reflecting the nature of physical motion, AC-LAM provides advantages over unstructured latents. Our main contributions are: • Structural prior: We introduce a scene-wise additive composition prior over short horizons, encouraging zik≈zij+zjkz_ik≈ z_ij+z_jk in the latent action space. • Analysis: We show that AC gives rise to identity and inverse elements, yields cycle consistency, and suppresses information that does not compose additively. • Practice: We enforce AC through a compositional loss in the forward-dynamics model, jointly optimized with reconstruction/bottleneck losses and short-horizon sampling. • Empirics: AC-LAM yields more structured latents which provide stronger policy supervision across simulation and real robots, surpassing state-of-the-art LAMs. 2 Related Work Latent Action Learning Latent action learning abstracts temporal dynamics in video by modeling inter-frame visual change. Early works in discrete settings (e.g., LAPO (Schmidt and Jiang, 2023), Genie (Bruce et al., 2024)) extract latent actions in 2D platformer games. Recent approaches extend to human/robot videos for continuous control (e.g., LAPA (Ye et al., 2024), IGOR (Chen et al., 2024a), MotoGPT (Chen et al., 2024b)), often using VQ-VAE to discretize latents. UniVLA (Bu et al., 2025b) learns task-centric action representations via language conditioning. Continuous latent actions with VAE regularization instead of VQ (e.g., CLAM (Liang et al., 2025), COMO (Yang et al., 2025), LAWM (Garrido et al., 2026)) offer alternative bottlenecks. Optical-flow-based methods (e.g., Motus (Bi et al., 2025), ViPRA (Routray et al., 2025), LAOF (Bu et al., 2025c)) emphasize motion-centric latents. Label-supervised designs (e.g., LAOM (Nikulin et al., 2025), Linear LAM (Zhang et al., 2025a), villa-X (Chen et al., 2025), CLAP (Zhang et al., 2026)) primarily leverage action labels and/or proprioceptive states to suppress distractor-induced variations. Temporally extended latents are explored in VideoWorld (Ren et al., 2025) and SSM-VLA (Cai et al., 2025), and viewpoint-invariant latents are explored in MVP-LAM (Authors, 2026b). While prior LAMs span discretization, conditioning, supervision, motion cues, and temporal extension, AC-LAM is the first to explicitly endow the latent action space with an additive compositional prior, yielding structured latents that strengthen downstream policy learning. Compositionality in Deep Learning Compositionality has been extensively studied in embeddings, starting with word vectors where linear semantic arithmetic holds (e.g., vec(“Russia”) + vec(“river”) ≈ vec(“Volga River”)) and extending to paraphrase and sentence embeddings (Mikolov et al., 2013b, a; Wieting et al., 2015; Arora et al., 2017), as well as vision–language representations that exhibit approximately linear subspaces and compositional generalization (Trager et al., 2024; Berasi et al., 2025). The most closely related to our work is Adaworld (Gao et al., 2025), which demonstrates latent action composition in games with a discrete action space learned without explicit structural constraints. CoLA-World (Authors, 2026a) also learns a latent-action world model but does not address latent action composition. In contrast, we focus on latent action learning for embodied AI with continuous control, enforcing scene-wise additive composition and introducing a structural prior over the latent action space. 3 Method We present the Additively Compositional Latent Action Model (AC-LAM), which enforces scene-wise additivity on the latent action space. This structural prior aligns the latent action with the physical nature of robotic motions, yielding more interpretable, motion-specific, and displacement-calibrated latent actions. Throughout, we use additive and additively compositional interchangeably. 3.1 Problem Formulation Let O denote the observation space (e.g., images), and let ⊆ℝdZ ^d denote the latent action space. Each observation o∈o belongs to a scene denoted by e(o)e(o). Definition 3.1 (Scene). Two observations oi,ojo_i,o_j are considered to be within the same scene, e(oi)=e(oj)e(o_i)=e(o_j), if some sequence of actions exists that the agent can take to get between oio_i and ojo_j. Intuitively, a scene contains information about embodiment, background, objects, and so on. Our goal is to learn a latent action model (LAM) which consists of: • An inverse dynamics model (IDM) f:×→f:O×O such that latent action zij:=f(oi,oj)z_ij:=f(o_i,o_j) encodes the transformation from oio_i to ojo_j, where e(oi)=e(oj)e(o_i)=e(o_j). • A forward dynamics model (FDM) F:×→F:Z×O such that oj=F(oi,zij)o_j=F(o_i,z_ij). We also use the notation Fz(o):=F(o,z)F_z(o):=F(o,z). LAMs are generally trained using a reconstruction loss. ℒrec=i,j‖oj−Fzij(oi)‖22L_rec\;=\;E_i,j\, \|\,o_j-F_z_ij(o_i) \|_2^2 (1) 3.2 Additively Compositional Latent Actions We enforce the (scene-wise) additive structure on the latent actions induced within the same scene. zik z_ik =zij+zjk, =z_ij+z_jk, (2) ∀oi,oj,ok,s.t.e(oi)=e(oj)=e(ok) ∀ o_i,o_j,o_k,\ s.t.\ e(o_i)=e(o_j)=e(o_k) This constraint reflects the intuition that the latent action from oio_i to oko_k can be decomposed as the sum of those from oio_i to ojo_j and from ojo_j to oko_k. The additive structure on the latent action space can also be applied to the FDM, Fzjk( F_z_jk ( Fzij(oi))=Fzij+zjk(oi), F_z_ij(o_i) )=F_z_ij+z_jk(o_i), (3) ∀oi,oj,ok,s.t.e(oi)=e(oj)=e(ok) ∀ o_i,o_j,o_k,\ s.t.\ e(o_i)=e(o_j)=e(o_k) since on the LHS we have Fzjk(oj)=okF_z_jk (o_j )=o_k and RHS Fzik(oi)=okF_z_ik(o_i)=o_k. The additive structure of the IDM form in Eq. (2) and the FDM in Eq. (3) will later be used to construct the auxiliary loss. Although the motion composition nature of the rigid-body motion described by the Lie group SE(3)SE(3) is matrix-multiplicative, small inter-frame motions captured in common LAMs can be approximated as additive in a Euclidean space, under the Baker–Campbell–Hausdorff (BCH) approximation (Barfoot, 2024). Concretely, the movement of a robot arm end-effector (say pt=[x,y,z,1]p_t=[x,y,z,1]) can be represented as a matrix Tt∈ℝ4×4T_t ^4× 4, consisting of a rotation matrix R∈ℝ3×3R ^3× 3 and translation vector t∈ℝ3t ^3. Consecutive transformations are composed through matrix multiplication. The translational aspect composes as the addition of xyz coordinates, shown below. While the the rotational composition is matrix-multiplicative, the BCH approximation models it as vector addition in the form of axis angle when the rotation is small. p3 p_3 =T2p2=T2(T1p1) =T_2p_2=T_2(T_1p_1) T2T1 T_2T_1 =[R2t201][R1t101]=[R2R1R2t1+t201] = bmatrixR_2&t_2\\ 0&1 bmatrix bmatrixR_1&t_1\\ 0&1 bmatrix= bmatrixR_2R_1&R_2t_1+t_2\\ 0&1 bmatrix Further, the additive assumption provides two practical advantages. First, it makes structural priors easy to express as differentiable training objectives. For example, this property can be turned into simple residuals (e.g., zik−zij−zjkz_ik-z_ij-z_jk) with well-behaved gradients, without additional tricks (e.g., logarithm mapping) to handle non-linearity. Second, it improves interpretability by calibrating the vector norm ‖z‖\|z\| with displacement empirically. Specifically, we observe that, in this additive space, the vector norm ‖z‖\|z\| empirically tends to correlate with physical motion magnitude over small time steps, yielding a transparent “amount of motion” signal. Figure 2: Additively Compositional Latent Action Model (AC-LAM). For triples (oi,oj,ok)(o_i,o_j,o_k) from the same scene, scene-wise additivity encourages zik≈zij+zjkz_ik≈ z_ij+z_jk, which regularizes the latent action space on top of a standard IDM–FDM architecture. The red line denotes the (i,j)(i,j) mapping with IDM encoder zij=f(oi,oj)z_ij=f(o_i,o_j) and FDM decoder o^j=Fzij(oi) o_j=F_z_ij(o_i). The blue and green lines depict the corresponding mappings for (j,k)(j,k) and (i,k)(i,k), respectively. 3.3 Analysis We now formally study what is implied by additivity in the latent action space. The below considers observations coming from the same scene e(o)e(o). Based on additivity defined in (2), we can derive the following propositions about identity, inverse consistency, and cycle consistency. Proposition 3.2 (Identity). It holds that zii=0z_i=0. Proof. Apply additivity to (i,i,k)(i,i,k) to get zik=zii+zikz_ik=z_i+z_ik. Cancel zikz_ik on both sides to obtain zii=0z_i=0. ∎ Proposition 3.3 (Inverse consistency). For any pair (i,j)(i,j) in the same scene, it holds that zji=−zijz_ji=-z_ij. Proof. Apply additivity to (i,j,i)(i,j,i) to get zii=zij+zjiz_i=z_ij+z_ji. Use Proposition 3.2 to substitute zii=0z_i=0. Rearranging yields zji=−zijz_ji=-z_ij. ∎ Proposition 3.4 (Cycle consistency). For any cycle i0→i1→⋯→im=i0i_0\!→\!i_1\!→\!·s\!→\!i_m=i_0 within a scene, we have ∑t=0m−1zitit+1=0 _t=0^m-1z_i_ti_t+1=0. Proof. Repeatedly apply additivity along the path to obtain ∑t=0m−1zitit+1=zi0im _t=0^m-1z_i_ti_t+1=z_i_0i_m. Since im=i0i_m=i_0, Proposition 3.2 gives zi0im=zi0i0=0z_i_0i_m=z_i_0i_0=0. ∎ What does additivity bring to latent actions? We will show that enforcing additivity on the latent action discourages it to capture non-additive components. Two typical examples for the non-additive components are scene-relative information (e.g., information about the background in the scene) and future leakage information (e.g., the target observation in the prediction of FDM). This information would ideally not appear in the latent action as it is irrelevant to motion and control. However, in practice it does often appear in latent action models (Li et al., 2025; Garrido et al., 2026), despite bottlenecks designed to suppress it, as it provides a shortcut to optimize for the reconstruction loss (1). First, we show that the additive structure discourages encoding static scene identifiers as a constant offset in z. Proposition 3.5 (No scene-related bias). In a fixed scene s, assume a decomposition zij=z~ij+bsz_ij= z_ij+b_s for all (i,j)(i,j) in scene s, where z~ij z_ij represents movement and bs∈b_s is constant w.r.t. (i,j)(i,j). If both z and z~ z satisfy additivity within scene s, then bs=0b_s=0. Proof. From additivity zik=zij+zjkz_ik=z_ij+z_jk, substitute zpq=z~pq+bsz_pq= z_pq+b_s to get z~ik+bs=(z~ij+bs)+(z~jk+bs) z_ik+b_s=( z_ij+b_s)+( z_jk+b_s). Use additivity of z~ z to replace z~ij+z~jk z_ij+ z_jk by z~ik z_ik. This yields z~ik+bs=z~ik+2bs z_ik+b_s= z_ik+2b_s, hence bs=0b_s=0. ∎ Next, we show that additivity also suppresses another common pitfall in LAM where zijz_ij ignores the motion and directly embeds the goal of FDM prediction ojo_j. Proposition 3.6 (No future leakage). In a fixed scene s, assume a decomposition zij=z~ij+gjz_ij= z_ij+g_j for all (i,j)(i,j) in scene s, where gj∈g_j depends only on the goal index j. If both z and z~ z satisfy additivity within scene s, then gj=0g_j=0 for all j in scene s. Proof. From additivity zik=zij+zjkz_ik=z_ij+z_jk, substitute zpq=z~pq+gqz_pq= z_pq+g_q to get z~ik+gk=(z~ij+gj)+(z~jk+gk) z_ik+g_k=( z_ij+g_j)+( z_jk+g_k). Use additivity of z~ z to replace z~ij+z~jk z_ij+ z_jk by z~ik z_ik. This yields z~ik+gk=z~ik+gj+gk z_ik+g_k= z_ik+g_j+g_k, hence gj=0g_j=0. ∎ While above results rest on simplified assumptions—e.g., approximate linear separability between scene/goal terms and motion semantics—that may not hold in practice, they clarify how additive structure regularizes latent actions: to encode additive components (which are typically related to rigid-body movement) while suppressing the non-additive components such as the scene-related and goal-only terms, reducing these common pitfalls for LAM training. 3.4 Implementation Out of convenience, we treat each trajectory as a scene (by Def. 3.1 the observations within a trajectory are reachable with the given sequence of actions). Note this results in a finer division of scenes than strictly necessary. To enforce the additive constraint on the latent action space there are two possible candidates. Firstly, one could consider applying it to the IDM in Eq. (2). ℒAC-IDM=i,j,k‖f(oi,ok)−f(oi,oj)−f(oj,ok)‖22.L_AC -IDM=E_i,j,k\,\|f(o_i,o_k)-f(o_i,o_j)-f(o_j,o_k)\|_2^2. (4) However, empirically we found that this IDM form led to instability during optimization (a trivial a solution collapses all latents to zero f(oi,oj)=0∀i,jf(o_i,o_j)=0\;∀ i,j), requiring target stabilization techniques (e.g., EMA (Grill et al., 2020; He et al., 2020) or frozen targets (Mnih et al., 2015)). The second option for enforcing additivity – through the FDM in Eq. (3) – is more stable. ℒAC-FDM=i,j,k‖ok−F(oi,f(oi,oj)+f(oj,ok))‖22.L_AC -FDM=E_i,j,k\,\|\,o_k-F(o_i,f(o_i,o_j)+f(o_j,o_k))\|_2^2. (5) Hence, the final AC-LAM objective is ℒ=ℒrec+ℒreg+λACℒAC-FDML\;=\;L_rec\;+\;L_reg\;+\; _AC\,L_AC -FDM (6) where ℒrecL_rec is from Eq. (1), ℒregL_reg comprises VQ-VAE codebook and commitment losses (Ye et al., 2024; Chen et al., 2025), and λAC _AC weights the AC-FDM term. We apply AC constraints to post‑VQ continuous embeddings used as latent actions. To align with prior LAMs and enable fair comparison, we adopt a VQ‑VAE bottleneck (Van Den Oord et al., 2017; Ye et al., 2024; Chen et al., 2025), though AC is bottleneck‑agnostic in principle. Our implementation builds on the villa‑X design (Chen et al., 2025) for its strong performance and includes a proprioceptive FDM for reconstructing robot state. Latent Action Evaluation We evaluate AC-LAM by training a policy with latent actions as supervision and testing it in closed loop. We adopt the policy architecture in villa-X (see Appendix A.2). Policies are trained with continuous post‑VQ latent embeddings from the latent action model as supervision, consistent with villa‑X. Downstream policy performance serves as a proxy for latent‑action quality. 4 Experiments Our experiments aim to answer the following questions: • Q1. Does AC-LAM learn well-structured latent actions? • Q2. Does AC-LAM outperform state-of-the-art LAMs such as those in LAPA, UniVLA, and Villa-X in downstream policy learning? • Q3. How do different design choices affect the performance of AC-LAM? 4.1 Experimental Setup Baselines We compare AC-LAM against the following baseline latent action models: • LAPA LAM (Ye et al., 2024) learns discrete latent actions via a VQ-VAE latent action tokenizer. • UniVLA LAM (Bu et al., 2025b) learns task-centric latent actions through language conditioning to extract task-relevant dynamics. • Villa-X LAM (Chen et al., 2025) learns physically grounded latent actions with an extra proprioceptive state FDM. Implementation Details We build AC-LAM on the Villa-X LAM architecture and add additive-composition (AC) regularization to structure the latent action space. The pretraining follows Villa-X the same datasets including OpenX and large-scale human video corpora. For scene-wise AC sampling, we draw triples (i,j,k)(i,j,k) from the same trajectory, bound the temporal span across i,j,ki,j,k by τ (robot: 3s3\,s; human: 2s2\,s), and filter triples with large rotations on robot data to ensure additivity is a reasonable approximation. Additional architectural, training, and dataset details are provided in Appendix A.1 and B. 4.2 Does AC-LAM learn well-structured latent actions? We assess whether AC-LAM learns latent actions with better structure along four axes: (i) adherence to additive composition, (i) alignment between latent-norm and true motion magnitude, (i) emergence of identity and inverse elements, (iv) suppression of non‑compositional leakage (environment‑specific and goal‑only terms). We perform quantitative/qualitative analyses on Fractal (Brohan et al., 2022) (in-distribution), Bridge-V2 (Walke et al., 2023) (in-distribution) and LIBERO (Liu et al., 2023a) (out-of-distribution). More details can be found in Appendix C. Adherence to additive composition We quantify adherence to the scene‑wise additive composition prior with a normalized composition residual ℒNorm-AC=i,j,k‖zik−zij−zjk‖22i,j‖zij‖22.L_Norm -AC= E_i,j,k\,\|z_ik-z_ij-z_jk\|_2^2E_i,j\,\|z_ij\|_2^2. (7) This scale‑invariant metric measures how closely latent actions compose additively within a scene; lower values indicate stronger adherence (since Eq (2) suggests zik−zij−zjk=0z_ik-z_ij-z_jk=0). As shown in Table 1, AC‑LAM achieves lower ℒNorm-ACL_Norm -AC than baselines, indicating stronger adherence to additive composition. Alignment between latent-norm and true motion magnitude To assess displacement calibration, we conduct both quantitative and qualitative evaluations. Quantitatively, we compute the Pearson correlation between the latent‑action norm ‖z‖\|z\| (using continuous post‑VQ embeddings) and the norm of the change in proprioceptive states |Δs|| s|. r(‖z‖,‖Δs‖)=cov(‖z‖,‖Δs‖)σ‖z‖σ‖Δs‖r (\|z\|,\,\| s\| )= cov\! (\|z\|,\,\| s\| ) _\|z\|\, _\| s\| (8) As shown in Table 1, AC‑LAM demonstrates strong alignment between latent magnitude and physical motion, with broadly favorable results across datasets. Qualitatively, we visualize the trajectory of the latent action norm ‖f(o0,ot)‖\|f(o_0,o_t)\| over time on real‑world tabletop manipulation in Figure 3 . Despite intermediate fluctuations, AC‑LAM tracks displacement magnitude most faithfully. In contrast, LAPA LAM and UniVLA LAM show weak correspondence to displacement, and villa‑X LAM, while more correlated due to the proprio FDM design, remains systematically under‑calibrated relative to AC‑LAM. The norm thus offers an interpretable proxy for the “amount of motion” from the initial observation. Table 1: Comparison of structured latent metrics on Fractal, Bridge and LIBERO for AC-LAM vs baselines. Lower ℒNorm-ACL_Norm -AC/ℒNorm-IdentityL_Norm -Identity/Δinv _inv indicate stronger additive/identity/inverse consistency; higher r(‖z‖,‖Δs‖)r (\|z\|,\,\| s\| ) indicate better displacement calibration. AC-LAM generally improves latent structure across datasets. Dataset Method ℒNorm-ACL_Norm -AC↓ r(‖z‖,‖Δs‖)r (\|z\|,\,\| s\| )↑ ℒNorm-IdentityL_Norm -Identity↓ Δinv _inv↓ Fractal AC-LAM 0.086 0.489 0.151 0.156 Villa-X LAM 0.374 0.289 0.358 0.874 UniVLA LAM 0.960 0.325 0.478 1.536 LAPA LAM 0.845 0.024 0.862 1.773 Bridge AC-LAM 0.135 0.408 0.020 0.267 Villa-X LAM 0.386 0.225 0.326 0.875 UniVLA LAM 0.953 0.276 0.517 1.465 LAPA LAM 0.875 -0.074 0.914 1.787 LIBERO AC-LAM 0.089 0.277 1.338 0.206 Villa-X LAM 0.491 0.074 0.516 0.955 UniVLA LAM 1.005 0.647 0.517 1.471 LAPA LAM 0.905 0.180 0.940 1.808 Emergence of identity and inverse elements The scene‑wise additive composition prior induces identity and inverse structure in the latent action space. We empirically quantify these properties using continuous post‑VQ latents. To test how accurately the identity is learned, we report: ℒNorm-Identity=i‖zii‖i,j‖zij‖L_Norm -Identity= E_i\,\|z_i\|E_i,j\,\|z_ij\| (9) where values closer to zero indicate stronger emergence of the identity property (as in Proposition 3.2). We evaluate the normalized norm for the sum of inverse elements Δinv=i,j‖zij+zji‖i,j‖zij‖ _inv= E_i,j\,\|z_ij+z_ji\|E_i,j\,\|z_ij\| (10) where values closer to zero indicate stronger emergence of the inverse property (as in Proposition 3.3). Across datasets, AC‑LAM generally trends closer to the identity and inverse ideals, suggesting a more structured and interpretable latent‑action space. Suppression of non‑compositional leakage The scene‑wise additive composition prior is intended to suppress non‑compositional signals (environment identifiers and future information leakage). To quantify environment‑specific leakage, we train an XGBoost classifier (Chen, 2016) to predict the robot dataset ID from continuous post‑VQ latent actions, treating the ID as a proxy for environment features. Lower probe accuracy AccenvmlpAcc_env^mlp indicates less environment information in the latents and, consequently, stronger cross‑environment generalization. Quantifying future information leakage is more challenging because future goals legitimately correlate with motion semantics. We therefore do not measure leakage directly; instead, we assess it indirectly via proxies—primarily the additivity residual across distinct goals within the same scene (Eq. 7), which quantifies practical adherence to additive composition. Consequently, future information leakage is evaluated through this composition‑adherence metric. Across these evaluations, AC‑LAM consistently yields lower environment‑ID probe accuracy (Table 2) and stronger proxy signals (Table LABEL:table:analysis) than baselines, indicating better suppression of non‑compositional leakage and improved generalization. Table 2: Evaluation results on probe accuracy (AccenvmlpAcc_env^mlp) on four environments (fractral, bridge, kuka, droid) with different LAMs. Lower probe accuracy indicates less environment leakage. LAM AC-LAM Villa-X UniVLA LAPA AccenvmlpAcc_env^mlp↓ 50.2% 63.0% 80.2% 52.0% Figure 3: Trajectory of the latent action norm ‖f(o0,ot)‖||f(o_0,o_t)|| in real-world tabletop manipulation, with latent actions generated by LAPA LAM, UniVLA LAM, Villa-X LAM and AC-LAM. AC‑LAM yields the most displacement‑calibrated latents, aligning with motion magnitude. 4.3 Does AC-LAM improve downstream policy learning performance? We further evaluate AC-LAM’s ability to provide effective supervision signals for downstream policy learning. Benchmarks We assess both vision–semantic generalization in simulation and accurate control on real-world tabletop manipulation (see more details in Appendix D). • Emoji Table-Top (GrinningFace) (Zhang et al., 2025b): A diagnostic simulation benchmark for vision–semantic generalization in embodied control. Each episode uses the instruction “Pick the cube and place it on [desc.]”, where [desc.] is the language description of the target emoji. Three emoji cards are placed on the tabletop; success requires grasping the cube and placing it on the correct target card. We evaluate under three protocols supported by the benchmark: ID (in-distribution combinations and order), Train (novel combinations of training-set emojis), and Val (held-out validation emojis; out-of-distribution). • Real-World Tabletop Manipulation A physical evaluation of accurate control and robustness using an AgileX Robotics Piper arm (7-DoF), focused on diverse pick tasks across varied objects and backgrounds. The dataset comprises 170 teleoperated trajectories collected under varied tabletop settings—including different tablecloth textures/colors, object layouts, and object positions—to increase scene diversity and support robustness evaluation. Performance is assessed under three regimes: in-distribution (ID) scenes, out-of-distribution distractors (OOD-D, novel or repositioned non-target objects), and out-of-distribution backgrounds (OOD-B, changes to tabletop/background appearance). Policy Training Setup We evaluate policies in the cross-dataset generalization setting, augmenting training data with Bridge-V2 to assess knowledge transfer. We use LAM to derive latent-action labels z forming robot tuples (s,z,a)(s,z,a). The policy follows the Villa-X architecture (Chen et al., 2025) and is trained end-to-end with joint supervision from both latent actions and robot actions. Baselines We compare policies trained with different LAMs (Section 3.1) to assess each LAM’s ability to provide supervision that enables downstream policy learning. We also include a baseline trained from scratch using only action-labeled data, based on the π0 _0 (Black et al., 2024) architecture (i.e., without latent-action supervision). For fairness, we align architecture and dataset across settings: Villa-X extends π0 _0 with a latent-action decoder, so π0 _0 (or w/o LAM) serves as the corresponding variant without latent-action supervision, isolating the benefit of latent labels. Table 3: Policy results on Emoji Table-Top (GrinningFace) across ID, Train, and Val splits. We report Success (S) (placement on the correct target), Any-Success (S/A) (placement on any target), and the Recognition Ratio (R=S/(S/A)), which approximates the model’s target recognition accuracy. Method ID Train Val S S/A R S S/A R S S/A R w/o LAM 22 35 0.63 26 44 0.59 19 38 0.50 LAPA LAM 19 33 0.58 7 26 0.27 9 34 0.26 UniVLA LAM 38 55 0.69 29 58 0.5 28 57 0.49 Villa-X LAM 42 53 0.79 25 52 0.48 31 58 0.53 AC-LAM 55 61 0.90 42 56 0.75 41 60 0.68 Table 4: Evaluation results on Real-World Tabletop Manipulation for policy learning with different LAMs. Method ID OOD-D OOD-B w/o LAM 33.3 26.7 6.7 LAPA LAM 13.3 6.7 0 UniVLA LAM 33.3 26.7 26.7 Villa-X LAM 40 20 26.7 AC-LAM 60 53.3 33.3 Results and Analysis Tables 3 (simulation) and 4 (real robot) report task success across protocols and tasks. Trends are consistent across simulation and real robot: adding scene-wise additive-composition (AC) constraints to latent action learning yields marked improvements over the no‑AC baseline in average task success, with gains observed on most tasks. On Emoji Table‑Top, AC‑LAM achieves higher success under ID/Train/Val. In Table 3, we report S (grasp the cube and place it on the correct target card), S/A (grasp the cube and place it on any card), and R=S/AR= SS/A, which approximates target‑card recognition. AC‑LAM notably improves R, contributing to the overall success rate. On the real‑robot suite, AC‑LAM consistently outperforms LAPA, UniVLA, villa‑X, and π0 _0 (or w/o LAM), indicating robust generalization and more effective supervision under distribution shift in real‑world robot settings. Under OOD background shifts, AC‑LAM typically preserves pick success while the w/o LAM baseline (π0 _0) and LAPA LAM fails to grasp consistently. We attribute these gains to motion‑specific, displacement‑calibrated latents induced by the additive‑composition prior, which provide stronger supervision in downstream policy learning. 4.4 How do different design choices affect AC-LAM performance? We study design choices for instantiating the scene-wise additive composition prior, targeting (i) practical adherence to additive composition, (i) improved calibration on motion magnitude, and (i) greater training stability. We evaluate above structured-latent metrics and optimization stability, and conduct ablations on AC-LAM trained on a reduced dataset comprising Bridge-V2 and Sth-Sth-V2 (Goyal et al., 2017). Design Factors Our default AC-LAM uses the FDM-form loss with constraints on post-VQ embeddings. We ablate: • AC loss form: ℒAC-IDML_AC -IDM (additivity in inverse-dynamics latents; Eq. 4) vs. ℒAC-FDML_AC -FDM (decoding from summed latents; Eq. 5). • Placement in VQ-VAE: pre-VQ (continuous encoder latents) vs. post-VQ (codebook embeddings). • Stop-gradient (sg) for ℒAC-IDML_AC -IDM: none, sg on zikz_ik, or sg on (zij+zjk)(z_ij+z_jk). Findings Enforcing the scene-wise additive prior via decoder-side ℒAC-FDML_AC -FDM is generally more stable than ℒAC-IDML_AC -IDM, as decoding from the sum of latents regularizes through the observation space and reduces optimization shocks. We observe a trade-off in where AC is applied: pre-VQ improves displacement calibration (r(‖z‖,‖Δs‖)=0.476r(\|z\|,\| s\|)=0.476) but weakens additive consistency (ℒNorm-AC=0.456L_Norm -AC=0.456), while post-VQ strengthens additive consistency (ℒNorm-AC=0.102L_Norm -AC=0.102) with stable training (and moderate r=0.256r=0.256). For ℒAC-IDML_AC -IDM, stop-gradient placement is critical: without stop-gradient the training tends to collapse; stopping gradients on (zij+zjk)(z_ij+z_jk) often drives the latent norm to blow up; stopping on zikz_ik is the most stable of the IDM variants, yet it still underperforms ℒAC-FDML_AC -FDM. In practice, both pre-VQ and post-VQ have merits; following prior LAM/policy setups that consume post-VQ codebook embeddings (Chen et al., 2025), we default to applying ℒAC-FDML_AC -FDM on post-VQ continuous embeddings. Table 5: Ablation on different design choices for AC-LAM, evaluated on Bridge. Design ℒNorm-ACL_Norm -AC↓ r(‖z‖,‖Δs‖)r (\|z\|,\,\| s\| )↑ Stability Default 0.102 0.256 stable pre-VQ 0.456 0.476 stable IDM(no sg) - - collapse sg on zikz_ik 0.141 0.098 stable sg on zij+zjkz_ij+z_jk - - explode 5 Conclusion and Future Work We introduced AC-LAM, a latent action learning framework that imposes a scene-wise additive-composition prior (zik=zij+zjkz_ik=z_ij+z_jk) aligned with short-horizon motion semantics. Our analysis and experiments show that AC-LAM yields more interpretable, motion-specific, and displacement-calibrated latents, suppresses non-compositional leakage, and provides stronger supervision for downstream policy learning. These results establish a practical and principled foundation for structured latent action learning from video. Directions for future work include stronger scene identification to form cross-trajectory triples—via scene labels or unsupervised scene recognition/clustering. We have demonstrated that additive composition is a promising direction in latent action learning; advancing scene conditioning might further align structured latents with physical motion priors. 6 Impact Statement This paper presents work whose goal is to advance the field of machine learning. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here. References S. Arora, Y. Liang, and T. Ma (2017) A simple but tough-to-beat baseline for sentence embeddings. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings, External Links: Link Cited by: §2. A. Authors (2026a) Co-evolving latent action world models. Note: Concurrent Submission to ICMLFilename: colaworld.pdf Cited by: §2. A. Authors (2026b) MVP-lam: learning action-centric latent action via cross-viewpoint reconstruction. Note: Concurrent Submission to ICMLFilename: mvplam.pdf Cited by: §2. T. D. Barfoot (2024) State estimation for robotics. Cambridge University Press. Cited by: §3.2. S. Belkhale, Y. Cui, and D. Sadigh (2023) HYDRA: hybrid robot actions for imitation learning. arxiv. Cited by: Table 6. D. Berasi, M. Farina, M. Mancini, E. Ricci, and N. Strisciuglio (2025) Not only text: exploring compositionality of visual representations in vision-language models. External Links: 2503.17142, Link Cited by: §2. H. Bi, H. Tan, S. Xie, Z. Wang, S. Huang, H. Liu, R. Zhao, Y. Feng, C. Xiang, Y. Rong, H. Zhao, H. Liu, Z. Su, L. Ma, H. Su, and J. Zhu (2025) Motus: a unified latent action world model. External Links: 2512.13030, Link Cited by: §2. K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky (2024) π0 _0: A vision-language-action flow model for general robot control. arXiv preprint arXiv: 2410.24164. Cited by: §4.3. A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. R. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich (2022) RT-1: robotics transformer for real-world control at scale. Robotics: Science and Systems. External Links: Document Cited by: Table 6, §4.2. J. Bruce, M. D. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, et al. (2024) Genie: generative interactive environments. In Forty-first International Conference on Machine Learning, Cited by: §1, §2. Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, S. Gao, X. He, X. Huang, S. Jiang, Y. Jiang, C. Jing, H. Li, J. Li, C. Liu, Y. Liu, Y. Lu, J. Luo, P. Luo, Y. Mu, Y. Niu, Y. Pan, J. Pang, Y. Qiao, G. Ren, C. Ruan, J. Shan, Y. Shen, C. Shi, M. Shi, M. Shi, C. Sima, J. Song, H. Wang, W. Wang, D. Wei, C. Xie, G. Xu, J. Yan, C. Yang, L. Yang, S. Yang, M. Yao, J. Zeng, C. Zhang, Q. Zhang, B. Zhao, C. Zhao, J. Zhao, and A. Jianchao Zhu (2025a) AgiBot world colosseo: a large-scale manipulation platform for scalable and intelligent embodied systems. arXiv preprint arXiv: 2503.06669. Cited by: §B.1, Table 6. Q. Bu, Y. Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li (2025b) UniVLA: learning to act anywhere with task-centric latent actions. External Links: 2505.06111, Link Cited by: §2, 2nd item. X. Bu, J. Lyu, F. Sun, R. Yang, Z. Ma, and W. Li (2025c) LAOF: robust latent action learning with optical flow constraints. External Links: 2511.16407, Link Cited by: §2. Z. Cai, Y. Yang, X. Chang, S. Liang, R. Chen, F. Xiong, M. Xu, and R. Huang (2025) Seeing space and motion: enhancing latent actions with spatial and dynamic awareness for vla. External Links: 2509.26251, Link Cited by: §2. [15] L. Y. Chen, S. Adebola, and K. Goldberg Berkeley UR5 demonstration dataset. Note: https://sites.google.com/view/berkeley-ur5/home Cited by: Table 6. T. Chen (2016) XGBoost: a scalable tree boosting system. Cornell University. Cited by: §4.2. X. Chen, J. Guo, T. He, C. Zhang, P. Zhang, D. C. Yang, L. Zhao, and J. Bian (2024a) IGOR: image-goal representations are the atomic control units for foundation models in embodied ai. arXiv preprint arXiv:2411.00785. Cited by: §2. X. Chen, H. Wei, P. Zhang, C. Zhang, K. Wang, Y. Guo, R. Yang, Y. Wang, X. Xiao, L. Zhao, J. Chen, and J. Bian (2025) Villa-x: enhancing latent action modeling in vision-language-action models. External Links: 2507.23682, Link Cited by: §2, §3.4, §3.4, 3rd item, §4.3, §4.4. Y. Chen, Y. Ge, Y. Li, Y. Ge, M. Ding, Y. Shan, and X. Liu (2024b) Moto: latent motion token as the bridging language for robot manipulation. arXiv preprint arXiv: 2412.04445. Cited by: §2. O. X. Collaboration, A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Gupta, A. Wang, A. Kolobov, A. Singh, A. Garg, A. Kembhavi, A. Xie, A. Brohan, A. Raffin, A. Sharma, A. Yavary, A. Jain, A. Balakrishna, A. Wahid, B. Burgess-Limerick, B. Kim, B. Schölkopf, B. Wulfe, B. Ichter, C. Lu, C. Xu, C. Le, C. Finn, C. Wang, C. Xu, C. Chi, C. Huang, C. Chan, C. Agia, C. Pan, C. Fu, C. Devin, D. Xu, D. Morton, D. Driess, D. Chen, D. Pathak, D. Shah, D. Büchler, D. Jayaraman, D. Kalashnikov, D. Sadigh, E. Johns, E. Foster, F. Liu, F. Ceola, F. Xia, F. Zhao, F. V. Frujeri, F. Stulp, G. Zhou, G. S. Sukhatme, G. Salhotra, G. Yan, G. Feng, G. Schiavi, G. Berseth, G. Kahn, G. Yang, G. Wang, H. Su, H. Fang, H. Shi, H. Bao, H. B. Amor, H. I. Christensen, H. Furuta, H. Walke, H. Fang, H. Ha, I. Mordatch, I. Radosavovic, I. Leal, J. Liang, J. Abou-Chakra, J. Kim, J. Drake, J. Peters, J. Schneider, J. Hsu, J. Bohg, J. Bingham, J. Wu, J. Gao, J. Hu, J. Wu, J. Wu, J. Sun, J. Luo, J. Gu, J. Tan, J. Oh, J. Wu, J. Lu, J. Yang, J. Malik, J. Silvério, J. Hejna, J. Booher, J. Tompson, J. Yang, J. Salvador, J. J. Lim, J. Han, K. Wang, K. Rao, K. Pertsch, K. Hausman, K. Go, K. Gopalakrishnan, K. Goldberg, K. Byrne, K. Oslund, K. Kawaharazuka, K. Black, K. Lin, K. Zhang, K. Ehsani, K. Lekkala, K. Ellis, K. Rana, K. Srinivasan, K. Fang, K. P. Singh, K. Zeng, K. Hatch, K. Hsu, L. Itti, L. Y. Chen, L. Pinto, L. Fei-Fei, L. Tan, L. ”. Fan, L. Ott, L. Lee, L. Weihs, M. Chen, M. Lepert, M. Memmel, M. Tomizuka, M. Itkina, M. G. Castro, M. Spero, M. Du, M. Ahn, M. C. Yip, M. Zhang, M. Ding, M. Heo, M. K. Srirama, M. Sharma, M. J. Kim, N. Kanazawa, N. Hansen, N. Heess, N. J. Joshi, N. Suenderhauf, N. Liu, N. D. Palo, N. M. M. Shafiullah, O. Mees, O. Kroemer, O. Bastani, P. R. Sanketi, P. ”. Miller, P. Yin, P. Wohlhart, P. Xu, P. D. Fagan, P. Mitrano, P. Sermanet, P. Abbeel, P. Sundaresan, Q. Chen, Q. Vuong, R. Rafailov, R. Tian, R. Doshi, R. Mart’in-Mart’in, R. Baijal, R. Scalise, R. Hendrix, R. Lin, R. Qian, R. Zhang, R. Mendonca, R. Shah, R. Hoque, R. Julian, S. Bustamante, S. Kirmani, S. Levine, S. Lin, S. Moore, S. Bahl, S. Dass, S. Sonawani, S. Song, S. Xu, S. Haldar, S. Karamcheti, S. Adebola, S. Guist, S. Nasiriany, S. Schaal, S. Welker, S. Tian, S. Ramamoorthy, S. Dasari, S. Belkhale, S. Park, S. Nair, S. Mirchandani, T. Osa, T. Gupta, T. Harada, T. Matsushima, T. Xiao, T. Kollar, T. Yu, T. Ding, T. Davchev, T. Z. Zhao, T. Armstrong, T. Darrell, T. Chung, V. Jain, V. Vanhoucke, W. Zhan, W. Zhou, W. Burgard, X. Chen, X. Chen, X. Wang, X. Zhu, X. Geng, X. Liu, X. Liangwei, X. Li, Y. Pang, Y. Lu, Y. J. Ma, Y. Kim, Y. Chebotar, Y. Zhou, Y. Zhu, Y. Wu, Y. Xu, Y. Wang, Y. Bisk, Y. Dou, Y. Cho, Y. Lee, Y. Cui, Y. Cao, Y. Wu, Y. Tang, Y. Zhu, Y. Zhang, Y. Jiang, Y. Li, Y. Li, Y. Iwasawa, Y. Matsuo, Z. Ma, Z. Xu, Z. J. Cui, Z. Zhang, Z. Fu, and Z. Lin (2023) Open X-Embodiment: robotic learning datasets and RT-X models. Note: https://arxiv.org/abs/2310.08864 Cited by: §B.1. Z. J. Cui, Y. Wang, N. M. M. Shafiullah, and L. Pinto (2022) From play to policy: conditional behavior generation from uncurated robot data. arXiv preprint arXiv:2210.10047. Cited by: Table 6. D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al. (2020) The epic-kitchens dataset: collection, challenges and baselines. IEEE Transactions on Pattern Analysis and Machine Intelligence 43 (11), p. 4125–4141. Cited by: §B.1, Table 6. S. Dass, J. Yapeter, J. Zhang, J. Zhang, K. Pertsch, S. Nikolaidis, and J. J. Lim (2023) CLVR jaco play dataset. External Links: Link Cited by: Table 6. F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine (2021) Bridge data: boosting generalization of robotic skills with cross-domain datasets. arXiv preprint arXiv:2109.13396. Cited by: Table 6. H. Fang, H. Fang, Z. Tang, J. Liu, J. Wang, H. Zhu, and C. Lu (2023) RH20T: a robotic dataset for learning diverse skills in one-shot. In RSS 2023 Workshop on Learning for Task and Motion Planning, Cited by: §B.1, Table 6. S. Gao, S. Zhou, Y. Du, J. Zhang, and C. Gan (2025) AdaWorld: learning adaptable world models with latent actions. External Links: 2503.18938, Link Cited by: §2. Q. Garrido, T. Nagarajan, B. Terver, N. Ballas, Y. LeCun, and M. Rabbat (2026) Learning latent action world models in the wild. External Links: 2601.05230, Link Cited by: §2, §3.3. R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic (2017) The ”something something” video database for learning and evaluating visual common sense. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Cited by: §B.1, Table 6, §4.4. K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al. (2022) Ego4d: around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 18995–19012. Cited by: §B.1, §B.2, Table 6. J. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, B. Piot, k. kavukcuoglu, R. Munos, and M. Valko (2020) Bootstrap your own latent - a new approach to self-supervised learning. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33, p. 21271–21284. External Links: Link Cited by: §3.4. K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick (2020) Momentum contrast for unsupervised visual representation learning. External Links: 1911.05722, Link Cited by: §3.4. M. Heo, Y. Lee, D. Lee, and J. J. Lim (2023) FurnitureBench: reproducible real-world benchmark for long-horizon complex manipulation. In Robotics: Science and Systems, Cited by: Table 6. E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn (2022) Bc-z: zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning, p. 991–1002. Cited by: Table 6. D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al. (2018) Qt-opt: scalable deep reinforcement learning for vision-based robotic manipulation. In CoRL, p. 651–673. Cited by: Table 6. A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y. J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y. Lee, M. Memmel, S. Park, I. Radosavovic, K. Wang, A. Zhan, K. Black, C. Chi, K. B. Hatch, S. Lin, J. Lu, J. Mercat, A. Rehman, P. R. Sanketi, A. Sharma, C. Simpson, Q. Vuong, H. R. Walke, B. Wulfe, T. Xiao, J. H. Yang, A. Yavary, T. Z. Zhao, C. Agia, R. Baijal, M. G. Castro, D. Chen, Q. Chen, T. Chung, J. Drake, E. P. Foster, J. Gao, D. A. Herrera, M. Heo, K. Hsu, J. Hu, D. Jackson, C. Le, Y. Li, K. Lin, R. Lin, Z. Ma, A. Maddukuri, S. Mirchandani, D. Morton, T. Nguyen, A. O’Neill, R. Scalise, D. Seale, V. Son, S. Tian, E. Tran, A. E. Wang, Y. Wu, A. Xie, J. Yang, P. Yin, Y. Zhang, O. Bastani, G. Berseth, J. Bohg, K. Goldberg, A. Gupta, A. Gupta, D. Jayaraman, J. J. Lim, J. Malik, R. Martín-Martín, S. Ramamoorthy, D. Sadigh, S. Song, J. Wu, M. C. Yip, Y. Zhu, T. Kollar, S. Levine, and C. Finn (2024) DROID: a large-scale in-the-wild robot manipulation dataset. Cited by: Table 6. M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. (2024) OpenVLA: an open-source vision-language-action model. arXiv preprint arXiv:2406.09246. Cited by: §B.1. Y. Li, Z. Cao, A. Liang, B. Liang, L. Chen, H. Zhao, and C. Feng (2022) Egocentric prediction of action target in 3d. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §B.1, Table 6. Y. Li, M. Liu, and J. M. Rehg (2018) In the eye of beholder: joint learning of gaze and actions in first person video. In Proceedings of the European conference on computer vision (ECCV), p. 619–635. Cited by: §B.1, Table 6. Z. Li, X. Gao, X. Wang, and J. Fu (2025) LatBot: distilling universal latent actions for vision-language-action models. arXiv preprint arXiv:2511.23034. Cited by: §3.3. A. Liang, P. Czempin, M. Hong, Y. Zhou, E. Biyik, and S. Tu (2025) Clam: continuous latent action models for robot learning from unlabeled demonstrations. arXiv preprint arXiv:2505.04999. Cited by: §2. B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone (2023a) LIBERO: benchmarking knowledge transfer for lifelong robot learning. arXiv preprint arXiv:2306.03310. Cited by: §4.2. H. Liu, S. Nasiriany, L. Zhang, Z. Bao, and Y. Zhu (2023b) Robot learning on the job: human-in-the-loop autonomy and learning during deployment. In Robotics: Science and Systems (RSS), Cited by: Table 6. Y. Liu, Y. Liu, C. Jiang, K. Lyu, W. Wan, H. Shen, B. Liang, Z. Fu, H. Wang, and L. Yi (2022) HOI4D: a 4d egocentric dataset for category-level human-object interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 21013–21022. Cited by: §B.1, Table 6. J. Luo, C. Xu, F. Liu, L. Tan, Z. Lin, J. Wu, P. Abbeel, and S. Levine (2024) FMB: a functional manipulation benchmark for generalizable robotic learning. arXiv preprint arXiv:2401.08553. Cited by: Table 6. C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence (2023) Interactive language: talking to robots in real time. IEEE Robotics and Automation Letters. Cited by: Table 6. O. Mees, J. Borja-Diaz, and W. Burgard (2023) Grounding language with visual affordances over unstructured data. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), London, UK. Cited by: Table 6. R. Mendonca, S. Bahl, and D. Pathak (2023) Structured world models from human videos. CoRL. Cited by: Table 6. T. Mikolov, K. Chen, G. Corrado, and J. Dean (2013a) Efficient estimation of word representations in vector space. External Links: 1301.3781, Link Cited by: §2. T. Mikolov, I. Sutskever, K. Chen, G. Corrado, and J. Dean (2013b) Distributed representations of words and phrases and their compositionality. External Links: 1310.4546, Link Cited by: §2. V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis (2015) Human-level control through deep reinforcement learning. Nature 518 (7540), p. 529–533. External Links: ISSN 00280836, Link Cited by: §3.4. S. Nasiriany, T. Gao, A. Mandlekar, and Y. Zhu (2022) Learning and retrieval from prior data for skill-based imitation learning. In Conference on Robot Learning (CoRL), Cited by: Table 6. A. Nikulin, I. Zisman, D. Tarasov, N. Lyubaykin, A. Polubarov, I. Kiselev, and V. Kurenkov (2025) Latent action learning requires supervision in the presence of distractors. External Links: 2502.00379, Link Cited by: §2. Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. L. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine (2024) Octo: an open-source generalist robot policy. In Proceedings of Robotics: Science and Systems, Delft, Netherlands. Cited by: §B.1. B. Pei, Y. Huang, J. Xu, G. Chen, Y. He, L. Yang, Y. Wang, W. Xie, Y. Qiao, F. Wu, and L. Wang (2025) Modeling fine-grained hand-object dynamics for egocentric video representation learning. External Links: 2503.00986, Link Cited by: §B.2, Table 6. G. Quere, A. Hagengruber, M. Iskandar, S. Bustamante, D. Leidner, F. Stulp, and J. Vogel (2020) Shared Control Templates for Assistive Robotics. In 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, p. 7 (en). Cited by: Table 6. Z. Ren, Y. Wei, X. Guo, Y. Zhao, B. Kang, J. Feng, and X. Jin (2025) VideoWorld: exploring knowledge learning from unlabeled videos. External Links: 2501.09781, Link Cited by: §2. E. Rosete-Beas, O. Mees, G. Kalweit, J. Boedecker, and W. Burgard (2022) Latent plans for task agnostic offline reinforcement learning. In Proceedings of the 6th Conference on Robot Learning (CoRL), Cited by: Table 6. S. Routray, H. Pan, U. Jain, S. Bahl, and D. Pathak (2025) ViPRA: video prediction for robot actions. External Links: 2511.07732, Link Cited by: §2. D. Schmidt and M. Jiang (2023) Learning to act without actions. arXiv preprint arXiv:2312.10812. Cited by: §2. N. M. M. Shafiullah, A. Rai, H. Etukuru, Y. Liu, I. Misra, S. Chintala, and L. Pinto (2023) On bringing robots home. External Links: 2311.16098 Cited by: Table 6. M. Trager, P. Perera, L. Zancato, A. Achille, P. Bhatia, and S. Soatto (2024) Linear spaces of meanings: compositional structures in vision-language models. External Links: 2302.14383, Link Cited by: §2. A. Van Den Oord, O. Vinyals, et al. (2017) Neural discrete representation learning. Advances in neural information processing systems 30. Cited by: §3.4. H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V. Myers, K. Fang, C. Finn, and S. Levine (2023) BridgeData v2: a dataset for robot learning at scale. In Conference on Robot Learning (CoRL), Cited by: Table 6, Appendix D, §4.2. J. Wang, Q. Zhang, Y. Chao, B. Wen, X. Guo, and Y. Xiang (2024) HO-cap: a capture system and dataset for 3d reconstruction and pose tracking of hand-object interaction. External Links: 2406.06843, Link Cited by: §B.1, Table 6. X. Wang, T. Kwon, M. Rad, B. Pan, I. Chakraborty, S. Andrist, D. Bohus, A. Feniello, B. Tekin, F. V. Frujeri, N. Joshi, and M. Pollefeys (2023) HoloAssist: an egocentric human interaction dataset for interactive ai assistants in the real world. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 20270–20281. Cited by: §B.1, Table 6. J. Wieting, M. Bansal, K. Gimpel, K. Livescu, and D. Roth (2015) From paraphrase database to compositional paraphrase model and back. External Links: 1506.03487, Link Cited by: §2. J. Yang, Y. Shi, H. Zhu, M. Liu, K. Ma, Y. Wang, G. Wu, T. He, and L. Wang (2025) CoMo: learning continuous latent motion from internet videos for scalable robot learning. External Links: 2505.17006, Link Cited by: §2. S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y. Chao, B. Y. Lin, L. Liden, K. Lee, J. Gao, L. Zettlemoyer, D. Fox, and M. Seo (2024) Latent action pretraining from videos. arXiv preprint arXiv: 2410.11758. Cited by: §1, §2, §3.4, §3.4, 1st item. C. Zhang, J. Wang, Z. Gao, Y. Su, T. Dai, C. Zhou, J. Lu, and Y. Tang (2026) CLAP: contrastive latent action pretraining for learning vision-language-action models from human videos. External Links: 2601.04061, Link Cited by: §2. C. Zhang, T. Pearce, P. Zhang, K. Wang, X. Chen, W. Shen, L. Zhao, and J. Bian (2025a) What do latent action models actually learn?. In NeurIPS 2025, External Links: Link Cited by: §2. C. Zhang, R. Yang, X. Chen, K. Wang, L. Zhao, Y. Chen, and J. Bian (2025b) How do vlas effectively inherit from vlms?. External Links: 2511.06619, Link Cited by: Appendix D, 1st item. Appendix A Training Details A.1 LAM Training Details Our LAM design largely follows villa-X LAM, augmented with an additive-composition (AC) loss and dynamic temporal intervals to support AC constraints. The IDM uses 12 Transformer encoder layers. Given an image pair (oi,oj)(o_i,o_j) (default 2×3×224×2242× 3× 224× 224), we apply a patch embedding with patch size 14, concatenate image tokens and stack 12 self-attention blocks (hidden dimension 768, 32 attention heads). The FDM is a 12-layer Vision Transformer (ViT-Base) that predicts ojo_j from (oi,zij)(o_i,z_ij). Following villa-X, we also employ a proprioceptive FDM: a 2-layer MLP with dual output heads that predict future robot states qjq_j, conditioned on (qi,zij)(q_i,z_ij). AC-LAM is trained on a mixture of human egocentric videos (e.g., Ego4D [21]) and robot trajectories (e.g., OpenX [12]). For scene-wise AC sampling, we draw triples (i,j,k)(i,j,k) from the same trajectory: robot temporal offsets are sampled uniformly from [0.1,3]s[0.1,3]\,s and human offsets from [0.1,2]s[0.1,2]\,s. We further filter robot triples exhibiting large rotations so that additive composition remains a reasonable approximation. Given the inherent temporal smoothness of robot motion, this approximation effectively captures dynamics within the proposed time range. We use a batch size of 512 and a learning rate of 1.5×10−41.5× 10^-4 with a 2000-step linear warmup. Training lasts approximately 10 days on 32 NVIDIA A100 GPUs. A.2 Policy Training Details We select villa-x as the policy architecture for downstream policy learning, to assess latent action model’s ability to provide high-quality supervision signals. The policy model in villa-x comprises three components. First, the vision–language encoder is based on PaliGemma[3], a 3B-parameter VLM pretrained with 224 × 224 images and 128-token text inputs. Second and third, the latent-action expert and the robot-action expert are each implemented as 18-layer Transformer networks, mirroring PaliGemma’s design, with a hidden dimension of 1,024 and 8 attention heads. For the latent action sequence, we select a sequence length of N = 6, and for the robot actions, we select a sequence length of M = 4. We apply the same random attention mask and random attention dropout techniques as in villa-x. We train all components jointly using a learning rate of 5e-5 with a 200-step linear warmup. We clip gradients to a maximum norm of 1.0 to ensure stable optimization. We did not pretrain the policy model on large-scale dataset. The goal here is use the model as a convenient policy learning method that can take both latent actions and robot actions as supervision signals. The policy learning follows the training data setup as mentioned in the experiment part. Each policy training with different LAMs takes 15K gradient steps, with a batch size of 512. To assess generalization under cross‑dataset transfer, we randomly form a 50%/50% mixture of the in-distribution dataset and Bridge V2 and use this combined corpus for training. Appendix B Datasets for Latent Action Learning B.1 Data Mixture We follow the data mixture in villa-x, which combines both robot data and action-free human videos for our LAM pretraining phase. For robot data, we draw primarily from OpenX (Collaboration et al., 2023) mixture and AgiBot (Bu et al., 2025a). For OpenX dataset, our base data mixture is created primarily based on (Kim et al., 2024; Octo Model Team et al., 2024). In total, we use 1.6M trajectories with 223.5M frames of robot data. For human videos, we use a mixture of Ego4D (Grauman et al., 2022), EgoPAT3D (Li et al., 2022), EGTEA Gaze+ (Li et al., 2018), EPIC-KITCHENS (Damen et al., 2020), HO-Cap (Wang et al., 2024), HOI4D (Liu et al., 2022), HoloAssist (Wang et al., 2023), RH20T (Fang et al., 2023), Something Something V2 (Goyal et al., 2017). Altogether, this yields 3.6M clips of human videos. During LAM pretraining, we exclusively utilize the primary third-person camera view. A full breakdown of our data mixture is listed in Table 6. B.2 Data Preprocessing For data cleaning, we adopt EgoHOD (Pei et al., 2025), a curated subset of Ego4D (Grauman et al., 2022), and further filter the videos based on visual quality to ensure high-quality inputs for training. For both robot data and human videos, we apply random adjustments to brightness, contrast, saturation, and hue as data augmentation. In the case of robot data, we represent both proprioceptive states and actions using euler angles. Dataset Mix Ratio (%) RT-1 Robot Action (Brohan et al., 2022) 9.70 AgiBot World Beta (Bu et al., 2025a) 20.0 Kuka (Kalashnikov et al., 2018) 1.97 Bridge (Walke et al., 2023; Ebert et al., 2021) 5.47 Taco Play (Rosete-Beas et al., 2022; Mees et al., 2023) 0.76 Jaco Play (Dass et al., 2023) 0.12 Berkely Autolab UR5 (Chen et al., ) 0.31 Language Table (Lynch et al., 2023) 0.11 Stanford Hydra Dataset (Belkhale et al., 2023) 1.61 NYU Franka Play Dataset (Cui et al., 2022) 0.22 Furniture Bench Dataset (Heo et al., 2023) 0.63 Austin Sailor Dataset (Nasiriany et al., 2022) 0.57 Austin Sirius Dataset (Liu et al., 2023b) 0.45 BC-Z (Jang et al., 2022) 3.47 DLR EDAN Shared Control (Quere et al., 2020) 0.01 CMU Stretch (Mendonca et al., 2023) 0.04 FMB Dataset (Luo et al., 2024) 0.73 DobbE (Shafiullah et al., 2023) 0.37 DROID (Khazatsky et al., 2024) 3.46 Ego4D (Grauman et al., 2022; Pei et al., 2025) 21.46 EgoPAT3D (Li et al., 2022) 0.94 EGTEA Gaze+ (Li et al., 2018) 0.89 EPIC-KITCHENS (Damen et al., 2020) 6.95 HO-Cap (Wang et al., 2024) 0.63 HOI4D (Liu et al., 2022) 1.99 HoloAssist (Wang et al., 2023) 4.77 RH20T (Fang et al., 2023) 5.56 Something-Something-V2 (Goyal et al., 2017) 6.82 Table 6: Our training data mixture used in LAM training. Appendix C More Details for Experiments on Latent Action Structure Sampling latent actions for Eq. 8 to 10 We adopt the same scene-wise (i,j,k)(i,j,k) sampling procedure used during training. For each dataset, we draw 16k latent-action instances (pairs or triplets, as required by each metric). Metrics are computed per instance and then averaged to approximate the corresponding expectations. Alignment between latent-norm and true motion magnitude We first compute the motion magnitude as the Euclidean distance between the proprioceptive states sis_i,sjs_j to obtain ‖Δsij‖\| s_ij\| Next, we rescale ‖Δsij‖\| s_ij\| per dimension using dataset quantiles: values are normalized with respect to the 1st and 99th percentiles, with clipping below the 1st percentile and above the 99th to reduce the influence of outliers. To make the Pearson correlation objective more stable and differentiable, we uniformly sampled (i,j)(i,j) pairs by motion magnitude so that the dataset spans a broad range of ‖Δsij‖\| s_ij\|. r(‖z‖,|Δs|)=cov(‖z‖,‖Δs‖)σ‖z‖σ‖Δs‖r (\|z\|,\,| s| )= cov\! (\|z\|,\,\| s\| ) _\|z\|\, _\| s\| (11) Quantifying environment-specific leakage We assess environment-specific leakage by training a simple latent probe to predict the data source (environment) from latent actions. From the latent probe dataset, we sample 100×32 latent action instances per environment across Fractal, Bridge, Kuka, and DROID. An XGBoost classifier is trained to predict the environment label from these latents, using a random 80/20 train/test split. We report test-set accuracy as the leakage metric, denoted AccenvmlpAcc_env^mlp. Higher AccenvmlpAcc_env^mlp. indicates stronger environment-identifying signals present in the latents (i.e., greater leakage), whereas lower AccenvmlpAcc_env^mlp suggests more environment-agnostic representations. All results are reported on the held-out 20% test split. More details for ablations on different design choices in AC-LAM We consider two variants of applying the stop-gradient mechanism to ℒAC-IDML_AC -IDM: ℒAC-IDM=‖sg(f(oi,ok))−(f(oi,oj)+f(oj,ok))‖22L_AC -IDM=\|sg(f(o_i,o_k))-(f(o_i,o_j)+f(o_j,o_k))\|_2^2 and ℒAC-IDM=‖f(oi,ok)−sg(f(oi,oj)+f(oj,ok))‖22.L_AC -IDM=\|f(o_i,o_k)-sg(f(o_i,o_j)+f(o_j,o_k))\|_2^2. For our pre-VQ ablations, the AC loss is applied to the continuous latent vectors immediately following the IDM encoder, prior to the discretization bottleneck of the vector quantizer. This contrasts with our default post-VQ approach, which constrains the quantized codebook embeddings. All ablation models are evaluated using the same sampling method with our main experiments. Appendix D More Details for Benchmarks (a) Emoji Table-Top (GrinningFace) (b) Real-World Tabletop Manipulation Figure 4: Two experimental environments: (a) Emoji Table-Top (GrinningFace) simulation for controlled studies of vision–semantic generalization. A robotic arm picks a cube and places it on the instructed emoji. The viewpoint is aligned with Bridge‑v2 to leverage this large-scale dataset for knowledge transfer. The initial positions of the cube, the emojis, and the robotic arm are randomized to test robustness. (b) Real‑World Tabletop Manipulation featuring diverse pick tasks across varied objects and backgrounds; evaluations cover in‑distribution scenes, OOD distractors, and OOD backgrounds to assess robustness. Emoji Table‑Top (GrinningFace) As shown in Figure 4(a) (we use the figure from (Zhang et al., 2025b)), it is a diagnostic simulation benchmark targets vision–semantic generalization in embodied control, evaluating how vision–language action models inherit priors from vision–language models. Each episode follows the instruction template “Pick the cube and place it on [desc.]”, where [desc.] is the language description of the target emoji. Three emoji cards are placed on the tabletop; success requires grasping the cube and placing it on the correct target card. To leverage Bridge dataset and enable knowledge transfer, the camera viewpoint is aligned with Bridge‑v2 (Walke et al., 2023). The initial positions of the cube, the emojis, and the robotic arm are randomized to systematically test robustness. Evaluation follows three protocols defined by the benchmark: ID (in‑distribution combinations and order), Train (novel combinations composed from training‑set emojis), and Val (held‑out validation emojis that are out‑of‑distribution). Real‑World Tabletop Manipulation A physical setup for evaluating accurate control and robustness under realistic variability. Experiments use an AgileX Robotics Piper arm featuring a 7‑DoF action space and focus on diverse pick tasks across varied objects and backgrounds. The dataset comprises 170 teleoperated trajectories, collected under varied tabletop settings—including different tablecloth textures/colors, object layouts, and object positions—to increase scene diversity and support robustness evaluation. Evaluations cover (i) in‑distribution scenes, (i) OOD distractors (novel or repositioned non‑target objects), and (i) OOD backgrounds (changes to tabletop/background appearance), enabling a comprehensive assessment of robustness. To ensure statistical reliability, we report success rates averaged over 15 rollouts for each setting. Appendix E Visualization E.1 Motion transfer demo with summed latents zij+zjkz_ij+z_jk Figure 5: Motion Transfer Demo Figure 5 shows several motion transfer examples. Input frames oio_i, ojo_j, and oko_k are sampled from the same trajectory in the BridgeV2 dataset. Latent actions zijz_ij, zjkz_jk, and zikz_ik are extracted using the IDM and then applied to another sampled frame oi′o _i via F(oi′,zik)F(o _i,z_ik) and F(oi′,zij+zjk)F(o _i,z_ij+z_jk). The results show that the FDM outputs are consistent when using either the direct latent action or the composed latent action, indicating that the semantic meanings of the two paths are well aligned. E.2 Additional Latent‑Action Norm Trajectories: Emoji Table‑Top and Real‑World Figure 6 provides extended visualizations of the latent action norm ‖LAM(o0,ot)‖\|LAM(o_0,o_t)\| evolution in both the Emoji Table-Top (GrinningFace) simulation and real-world tabletop environments, corroborating the analysis in the main text. The trends are consistent: AC-LAM demonstrates the strongest displacement calibration. Villa-X shows a correlation but remains under-calibrated, while LAPA and UniVLA fail to meaningfully track displacement. Collectively, these trajectories illustrate that the latent‑action norm generated by AC-LAM provides an interpretable proxy for the amount of motion from the initial observation. Figure 6: More trajectories of the latent action norm ‖LAM(o0,ot)‖\|LAM(o_0,o_t)\| in Emoji Table-Top simulation and real-world tabletop manipulation, with latent actions generated by LAPA LAM, UniVLA LAM, villa-X LAM and AC-LAM. AC‑LAM yields the most displacement‑calibrated latents, aligning with motion magnitude.