Paper deep dive
Inferring Action from Future Latent State for Robotic Manipulation
Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Jie Cheng, Peilin Huang, Han Fu, Zhuo Li, Xiaoxue Ren
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/25/2026, 8:07:03 AM
Summary
The paper introduces DELE-w0.5, a robotic manipulation framework that infers actions from predicted future latent states rather than generating dense video frames. It argues that video generation is an unnecessary intermediate objective for world-action modeling, as it introduces computational redundancy. DELE-w0.5 uses a dual-stream attention architecture to predict compact future states and actions simultaneously, achieving superior performance on real-robot trials compared to existing World-Action Models (WAMs) and Vision-Language-Action (VLA) models.
Entities (8)
Relation Signals (7)
DELE-w0.5 → developedby → DeepLeap Research
confidence 95% · 1DeepLeap Research, Shenzhen, China
DELE-w0.5 → proposes → Future State Inference
confidence 95% · DELE-w0.5 infers the action sequence from its corresponding compact future latent state.
Video Generation → iscritiquedby → DELE-w0.5
confidence 90% · We argue that video generation is an unnecessary intermediate objective for world-action modeling.
DELE-w0.5 → outperforms → World Action Models
confidence 90% · our DELE-w0.5 achieves the best performance among all compared policies... outperforming the strongest baseline by 47.5 and 30.7 percentage points
World Action Models → relieson → Video Generation
confidence 90% · World-Action Models (WAMs) build robot control on video-generation backbones
DELE-w0.5 → uses → Qwen3
confidence 90% · L is tokenized by Qwen3 [42]
DELE-w0.5 → uses → DINO-v3
confidence 90% · we apply DINO-v3 [31] as the vision encoder
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action modeling. For robotic manipulation, the goal of a world model is not to reproduce how the world looks at every intermediate moment, but to predict the state that the world will reach after an action is executed. The intermediate frames only describe the visual transition between physical states, which consumes substantial model capacity and computation, but do not directly specify the physical outcome that the robot action is intended to produce. In this paper, we propose DELE-w0.5, which infers robot actions from predicted future states without relying on video generation. Concretely, DELE-w0.5 infers the action sequence from its corresponding compact future latent state. The future latent state captures the action-relevant physical outcome of robot interaction and serves as an explicit bridge between world modeling and action generation. The core design principle of DELE-w0.5 is to model how the physical world changes under robot actions, rather than how its visual appearance evolves frame by frame. This formulation removes the high-dimensional visual redundancy introduced by dense video representations, and it therefore enables cheaper training and low-latency inference. Across 480 real-robot trials on four long-horizon manipulation tasks, our DELE-w0.5 achieves the best performance among all compared policies, attaining 62.5 overall full-task success and 81.3 macro ordered-stage progress, outperforming the strongest baseline by 47.5 and 30.7 percentage points, respectively.
Tags
Links
- Source: https://arxiv.org/abs/2608.22067v1
- Canonical: https://arxiv.org/abs/2608.22067v1
Trouble viewing inline? Open PDF directly →
Full Text
61,230 characters extracted from source content.
Expand or collapse full text
One Intelligence Across Embodiments Report 1DeepLeap Research, Shenzhen, China Email: yanglong0908@gmail.com -Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action modeling. For robotic manipulation, the goal of a world model is not to reproduce how the world looks at every intermediate moment, but to predict the state that the world will reach after an action is executed. The intermediate frames only describe the visual transition between physical states, which consumes substantial model capacity and computation, but do not directly specify the physical outcome that the robot action is intended to produce. In this paper, we propose DELE-.5 w0.5, which infers robot actions from predicted future states without relying on video generation. Concretely, DELE-.5 w0.5 infers the action sequence from its corresponding compact future latent state. The future latent state captures the action-relevant physical outcome of robot interaction and serves as an explicit bridge between world modeling and action generation. The core design principle of DELE-.5 w0.5 is to model how the physical world changes under robot actions, rather than how its visual appearance evolves frame by frame. This formulation removes the high-dimensional visual redundancy introduced by dense video representations, and it therefore enables cheaper training and low-latency inference. Across 480 real-robot trials on four long-horizon manipulation tasks, our DELE-.5 w0.5 achieves the best performance among all compared policies, attaining 62.5 overall full-task success and 81.3 macro ordered-stage progress, outperforming the strongest baseline by 47.5 and 30.7 percentage points, respectively. ://deepleap-x.com/research/dele Research Inferring Action from Future Latent State for Robotic Manipulation Fenghao Lei Zhixiong Huang Long Yang Jiabao Chen Jie Cheng Peilin Huang Han Fu Zhuo Li Xiaoxue Ren keywordsembodied intelligence, vision-language-action models, world model 1 Introduction Embodied intelligence has achieved rapid development and achieved strong performance for robotic manipulation [16, 18]. Vision-Language-Action (VLA) and World Action Model (WAM) are two popular model architectures. VLA is a three-step framework, see Figure 1 (a): First, the input comprises multimodal data, including human task instructions and the robot’s current observations; Then, a vision-language model (VLM) backbone performs the task planning according to the input data; Finally, an action expert decodes the task planning to the commands that are executable by the robot. For some classic work of VLA, refer to [4, 53, 14, 11]. Since the information density of language is lower than that of vision, VLA models require a large amount of data to adapt to the downstream tasks. As a result, generalization of VLA is poor, when light, color, or the task changes, the success rate decays rapidly, extensive empirical results have shown this view, see [44, 13]. WAM [9, 44, 13] fundamentally discards the framework of VLA, WAM’s core lying in the construction of a unified end-to-end learning objective, which integrates the prediction of world state evolution and action generation, see Figure 1 (b). The critical technology of WAM relies on self-supervised future state prediction on large-scale video data, then WAM develops the action generation as a conditional denoising process according to the future state prediction. 1.1 Limitation of WAM A dominant line of recent World-Action Models (WAMs) builds robot control on top of video generation and jointly predicts future visual trajectories and actions [44, 45], which brings some useful spatiotemporal priors into the policy. However, it also creates a structural mismatch between the objective of visual generation and the objective of robot control. Future videos are high-dimensional and appearance-sensitive. Most intermediate frames describe how the scene looks during a transition, rather than the physical consequence that determines the subsequent action. Therefore, the model is required to solve a substantially harder problem than control itself. It must first reconstruct dense visual evolution and then extract a compact action signal from it. Additionally, multi-frame video latents are much larger than the corresponding robot action sequence [48]. They dominate the sequence length, memory consumption, and computational budget of joint world-action learning. Consequently, WAM training inherits the heavy cost of large video generators. Recent methods introduce compact backbones, latent downsampling, or lightweight adaptation to make video-action co-training more practical [21]. These techniques reduce the cost of the existing pipeline, but do not remove its underlying redundancy. The same limitation becomes more critical during inference. Before an action can be executed, the model must perform iterative denoising over high-dimensional future video-action latents [1]. The control loop is therefore bottlenecked by synthesizing a visual trajectory that is not itself executed by the robot [17]. Therefore, the fundamental limitation of video-generation-based WAMs is not merely their computational cost. The deeper problem is that they optimize visual trajectory reconstruction as a surrogate for action-relevant world prediction. For robot control, a world model should predict how the world will become after an interaction, rather than reproduce how the world looks at every intermediate moment. This motivates a more direct formulation that models action-relevant future states and infers robot actions without relying on video generation in the control loop. ActionAction ExpertVLMLanguageVision(a) VLAVideoActionTransformerCausal DiTVisionLanguage(b) WAMFuture FrameActionProjectionProjectionCausal AttentionVisionLanguage(c) DELE-w0.5 Figure 1: Different Models for Manipulation. 1.2 Our Work We reformulate the current WAM of visual world modeling for robotic manipulation. It is different from recent video generation techinique based WAMs that model the future world through synthesizing dense visual trajectories, the proposed DELE-.5 w0.5 argues that robot control does not require reconstructing the entire visual evolution of the real world, but only requires to predict the future latent state with downstream actions. With this observation, we formulate WAM as prediction problem of a future latent state and action, where the DELE-.5 w0.5 directly infers the future visual state and corresponding robot behaviors, without any generating intermediate video frames. Rethinking World Model for Robotic Manipulation. We argue that existing video-generation-based WAMs formulate unnecessary intermediate objectives for robotic manipulation. Although predicting future videos provides rich spatiotemporal priors, it implicitly treats visual evolution as the primary target of world modeling. However, for robot robotic manipulation, a world model should not reproduce how the world looks at every intermediate moment. Instead, it should predict how the world will become after an action is executed. Most intermediate frames in generated videos only describe the visual transition between states, which provides limited information about the final physical consequence that determines subsequent actions. Therefore, video-generation-based WAMs force the model to learn the distribution of visual evolution rather than directly modeling action-relevant future states. This design introduces substantial computational overhead without necessarily improving control capability. We argue that future-state prediction with causal relevance to robot actions is a more fundamental objective for embodied world models than reconstructing complete visual trajectories. Inferring Action from Future-State. With this observation, we propose DELE-.5 w0.5 that infers robot actions from predicted future states rather than generating future visual trajectories. Different from existing video-generation-based approaches that model the continuous evolution of the visual world, DELE-.5 w0.5 treats future prediction as a state transition problem and focuses on the physical states that are causally relevant to subsequent actions. The key insight is that a world model for robotic manipulation does not need to reproduce how the world visually evolves at every intermediate moment, but should capture how the world will become after an action is executed. By directly modeling action-relevant future states, DELE-.5 w0.5 avoids the high-dimensional visual redundancy introduced by video representations, and eliminates the costly iterative generation process required by video-based WAMs. This formulation provides a more compact and efficient paradigm for world-action modeling, enabling significantly cheaper training and low-latency inference for real-world robotic deployment. Finally, we comparse the three formulation from probabilistic inference as follows, • VLA infers pθ( | , , )p_θ( _t| _t, _t, ); • WAM infers pθ( , :t+Δt| , , )p_θ( _t, _t:t+ t| _t, _t, ); • DELE-.5 w0.5 infers pθ( , +Δt| , , )p_θ( _t, _t+ t| _t, _t, ), where :t+Δt _t:t+ t is the video from current observation _t to the predicted future state +Δt _t+ t. 2 Related Work 2.1 Vision-Language-Action Models VLA models connect the semantic knowledge of pretrained vision-language models with low-level robot control. PaLM-E [8] injects continuous visual and state observations into a large language model for embodied reasoning. RT-2 [53] further casts robot actions as language-like tokens and co-trains web-scale vision-language tasks with robot trajectories. Open X-Embodiment and RT-X [26] show that heterogeneous trajectories from many robot platforms can support positive cross-embodiment transfer . Octo [24] provides an open generalist policy that supports different observations, action spaces, and robot embodiments. OpenVLA [14] adapts a pretrained VLM to autoregressive action prediction and makes large-scale VLA training accessible. TinyVLA reduces model size and data requirements, while SmolVLA further targets training and deployment on affordable hardware [38, 30]. RDT-1B [23] uses a diffusion transformer for multi-modal bimanual actions. CogACT [20] couples a VLM with a diffusion action module. π0 _0 [3] instead uses flow matching to generate continuous action chunks. DexVLA [37] adopt dual-system designs in which a vision-language component provides semantic features and a diffusion expert produces motor commands. FAST [27] compresses high-frequency action sequences into a shorter autoregressive vocabulary. OpenVLA-OFT [12] uses parallel action decoding, continuous actions, and action chunking to improve both control rate and downstream success. UniVLA [36] learns task-centric latent actions from heterogeneous videos and decodes them for different embodiments. π0.5 _0.5 [11] combines robot data, web data, semantic prediction, and heterogeneous supervision for open-world manipulation. HAMLET introduces compact historical memory, while ProgressVLA explicitly estimates task progress for long-horizon control [15, 41]. Qwen-VLA further unifies manipulation, navigation, and trajectory prediction across tasks and robot embodiments [35]. VLA still learns a direct conditional mapping from the current context to an action sequence, which is effective for behavior cloning, but it does not require the policy to represent the physical consequence of its action. The cost of large VLM backbones and iterative generative action heads also complicates real-time deployment. In contrast, DELE-.5 w0.5 proposes a predicted future state as an explicit interface between perception and control. It first estimates the action-relevant physical outcome and then infers the robot action from this state. 2.2 World-Action Models World models learn predictive structure that can support control, planning, or policy learning, and recent studies show that future prediction can provide useful structure beyond a purely reactive policy. Latent world models such as DreamerV3 [10] improve a policy through imagined state trajectories. Multi-view Masked World Models [29] learn predictive representations from multiple cameras for visual manipulation. UniPi [9] formulates sequential decision making as text-conditioned video generation and recovers actions from generated plans. Genie [5] learns action-controllable interactive environments from unlabeled videos through latent actions. Recent robotic world models connect future prediction more tightly with action generation. GR-1 [39] jointly predicts future images and robot actions after large-scale video pretraining. GR-2 [7] scales this paradigm with web videos and robot trajectories. 3D-VLA [50] predicts goal images and point clouds to connect 3D reasoning with action planning. IRASim generates action-conditioned videos with frame-level action alignment for fine-grained robot interactions [52]. Seer closes the loop through a predictive inverse-dynamics formulation, in which actions are inferred from forecast visual states [33]. CoT-VLA [49] predicts future visual goals as an explicit visual chain of thought before producing actions. WorldVLA unifies image and action generation in one autoregressive model [6]. V-JEPA 2-AC [2] instead predicts future states in a learned representation space and uses them for planning. Genie Envisioner builds policy learning and neural simulation on a shared video diffusion foundation [22]. DreamZero [44] builds a large WAM on a pretrained video diffusion backbone and demonstrates strong physical generalization. τ0 _0-WM [51] combines video-action prediction, action-conditioned simulation, and candidate evaluation in one framework. GigaWorld-Policy [43] makes future-video generation optional at deployment and places the action stream before the video stream. Flash-WAM [1] distills iterative video-action diffusion into very few sampling steps. Efficient-WAM [17] uses a compact video expert, sparse video tokens, and asymmetric denoising. AGRA [28] shows that visually plausible futures do not always yield accurate actions and aligns video features with action-relevant regions. Video-generation-based WAMs couple robot control to a substantially harder surrogate problem. A robot ultimately executes a low-dimensional action sequence, but the model must first process or synthesize dense multi-frame latents. Most intermediate frames describe visual transition. They do not directly specify the final physical state that determines the next action. This mismatch increases sequence length, memory use, and training cost. During inference, iterative denoising over future video and action latents introduces latency into the control loop. More importantly, features optimized for visual reconstruction are not necessarily organized for low-level control, as recent representation analyses have shown [28]. Distillation, token sparsification, and smaller video backbones reduce this cost, but they retain video synthesis as the underlying objective. DELE-.5 w0.5 instead formulates WAM as future-state and action prediction. It predicts a compact future state that is causally relevant to manipulation and infers the corresponding action from this state. The model does not rely on any video generation technique and does not reconstruct intermediate frames. It therefore avoids the high-dimensional visual redundancy of video WAMs and provides a more direct, cheaper, and lower-latency route from world prediction to robot control. 3 Preliminary 3.1 Robotic Manipulation We model robotic manipulation as a probabilistic inference model. For each time t, the robot obtains the multi-view RGB images _t (also known as observation), the proprioception _t, and language instruction L, we aim to predict the action according to the policy ∼π(⋅| , ,), _t π(·| _t, _t,L), where =[ +1, +2,⋯, +H] _t=[ _t+1, _t+2,·s, _t+H] corresponds to an action chunk of furture actions (we use chunk length H=60H=60 for our tasks) for the robot will play; where for the single-arm manipulation, each ∈ 7 ∈ ^7 denoted as 77-DoF action space: a 33-DoF relative positional displacement (x,y,z)(x,y,z), a 33-DoF Euler angle rotation ( roll, pitch, yaw), and a 1-DoF binary gripper state g∈0,1g∈\0,1\. For the dual-arm manipulation, it extends the setting to a 14-DoF space. 3.2 Language and Vision Tokenization The language instruction defines the task to be performed by the robot (e.g. " put~the~flowers~into~the~vase"), where in this paper, is tokenized by Qwen3 [42], we denote it as follows, ^=( ). = Text~Encoder( ). (1) For each time t, current observation ∈ ×3×h×w _t∈ ^v× 3× h× w contains multi-view images, where our robot is with three views (i.e., v=3v=3): a head-mounted view, along with a wrist-mounted view for each arm. Furthermore, a vision tokenization encoder compress the raw image _t into a latent space, denoted it as ^t=( ), _t= Vision~Encoder( _t), (2) where ^t∈ ′×h′×w′ _t∈ ^c × h × w is with the channel number of c′c , and the latent size is with h′×w′h × w . In this paper, we apply DINO-v3 [31] as the vision encoder to encode the input images. The DINO-v3 encoder provides more comprehensive representations with concurrently discern spatial and geometric relationships across observations. Such representations serve as high-quality inputs for downstream policy networks (e.g., diffusion policies or action models), enhancing the accurate perception of the relative pose between the manipulator’s end-effector and the target object, and consequently enabling the generation of more refined motion trajectories. 3.3 Timestep Embedding This section present the method of timestep embedding. For a given timestep τ, first, the scalar τ is embedded into a high-dimensional, multi-scale space that exhibits favorable distance properties; then, a two-layer MLP adapts the frequency encoding for the downstream task. Let d be the embedding dimension, k=⌊d/2⌋k= d/2 , ~∈ ∈ ^d, for each j∈0,1,⋯,d−1j∈\0,1,·s,d-1\: ~[j]=cos(τ⋅10000−j/k),if0≤j<k,sin(τ⋅10000−(j−k)/k),ifk≤j<2k,0,ifj=d−1&d is an odd number. [j]= cases (τ· 10000^-j/k ),&if~0≤ j<k,\\ (τ· 10000^-(j-k)/k ),&if~k≤ j<2k,\\ 0,&if~j=d-1~\&~d is~an~odd~number. cases (3) Furthermore, we apply a two-layer MLP to map the vector ~ from d-dimension to d-dimension as follows, =( ~), = FNN( ), (4) where FNN is a neural network that first increases the dimensiona of ~ , and then reduces it to be the d-dimension. Finally, we present above embedding as the following formulation, =(τ). = Timestep~Embedding(τ). (5) Condition Stream BlockNoise Stream Block^t A_t^t+1 O_t+1PredictionMasked Attention BlockcQ_ccK_ccV_cAdaLN & QKVTimestep τ=1τ=1Single-Stream Attention Block ^tc _t^c Encoder “put the flowerson the deskinto the vase”Language: L cat _tProjection ^t _tVision EncoderCurrent State: tO_tSelf-Attention τ ^τ_tAdd Noise[ +1, +2CLOSE[ _t+1, _t+2, ⋯, +H]·s·s, _t+H]Action Chunk: tA_tSelf-Attention +Δtτ ^τ_t+ tAdd NoiseVision EncoderNext State: t+ΔtO_t+ tnQ_nnK_nnV_nTimestep τ∼[0,1]τ _[0,1]Single-Stream Attention Block cat Figure 2: Overview of HeTu-1.0 1.0 architecture. The model consists of a condition stream and a noise stream. The condition stream encodes the language instruction L and current observation tO_t, while the noise stream jointly denoises the action chunk tA_t and future state t+ΔtO_t+ t. The two streams are coupled through AdaLN-modulated QKV projections and masked attention, enabling unified prediction of the robot action and its corresponding latent future state. 4 Framwork of HeTu-1.0 The framwork of DELE-.5 w0.5 adopts a dual-stream attention mechanism architecture, which contains two blocks: condition stream and noise stream. 4.1 Condition Stream Block 4.1.1 Input and Tokenization For each time t, we encode the language instruction input and current state input _t according to the methods (1) and (2) correspondingly, and obtain the language token , latent space of vision ^t _t. Since the output dimensions of the vision encoder and the language encoder are not necessarily the same, then we apply the projection operator to ensure that feature information and _t have the same hidden dimension, =( ^), =( ^t), = Projection( ),~~ _t= Projection( _t), (6) where in this paper, we set hidden dimension with the value of 10241024, the (⋅) Projection(·) operator we used is the linear neural networks. 4.1.2 Single-Stream Attention Block To integrate the information of language instruction with current observations, we employ a single-stream attention block [19] that concatenates the multimodal tokens (i.e., and _t), and feeds them into a shared Transformer encoder, where the attention block enables unified, source-agnostic cross-modal interactions. Concretely, firstly, we concatenate the language and vision tokens as follows, =( , )=:[ , ]. ^c_t= cat( , _t)=:[ , _t]. (7) Then we apply the standard self-attention block [34] to refine the concatenated information as follows, ^tc=( ), ^c_t= Attention( ^c_t), (8) which exploites the complementary information gain from language and vision, and shows the generalization robustness in embodied artificial intelligence. 4.1.3 AdaLN and Output of Condition Stream Furthermore, DELE-.5 w0.5 applies the AdaLN facilitates the integration of conditioning information into the feature representation, which resides the dynamic learning of intermediate representations in response to the incoming conditioning signal, and which ia a critical prerequisite for robots to attain fine-grained control. After the single-stream attention block, with the fused information ^tc ^c_t, we obtain the query, keys, values as the output of condition stream as follows, = _c= QKNorm (RMSNorm( ^tc)⊙c), \ _q(RMSNorm( ^c_t) γ_c) \, (9) = _c= QKNorm (RMSNorm( ^tc)⊙c), \ _k(RMSNorm( ^c_t) γ_c) \, (10) = _c= (RMSNorm( ^tc)⊙c), _v(RMSNorm( ^c_t) γ_c), (11) where RMSNorm(⋅)RMSNorm(·) is root mean square normalization [46]; c γ_c is the scaling factor for the condition stream with following linear structure, c=1+ c+Adaln, γ_c=1+ _Adalnt_c+b_Adaln, (12) _Adaln and Adalnb_Adaln are the learnable weights, ct_c is the timestep τ=1τ=1 for condition stream, i.e., c=(5)(1);t_c timestep-embedding= Timestep~Embedding(1); QKNorm(⋅)QKNorm(·) is a normalization operator maps any vector x as follows, ′=QKNorm()=/‖2; =QKNorm(x)=x/\|x\|_2; (13) and QKNorm(⋅)QKNorm(·) applies ℓ2 _2 normalization along the head dimension of each query and key matrix prior to multiplying them. 4.2 Noise Stream Block For each time t, we need to add noise to the clear data _t and next state +Δt _t+ t. Let τ∼[0,1]τ _[0,1], ∼ (, ) ( 0, ), where [0,1]U_[0,1] is the uniform distribution on [0,1][0,1], we obtain noise action as follows, τ=τ +(1−τ) . ^τ_t=τ _t+(1-τ) . (14) The noise addition of image is on the latent space: ^t+Δt= _t+ t= ( +Δt), Vision~Encoder( _t+ t), (15) then we obtain the noise image as follows, +Δtτ=τ ^t+Δt+(1−τ) . ^τ_t+ t=τ _t+ t+(1-τ) . (16) Furthermore, following the same steps from (7), (8), (9), (10) and (11), we obtain the query, keys, values as the output of noise stream, denoted them as _n, _n and _n. 4.3 Flow Matching Objective We concatenate the information from condition stream and noise stream as follows, =( , ), =( , ), =( , ). = cat( _c, _n),~~~ = cat( _c, _n),~~~ = cat( _c, _n). (17) With the joint attention block and projection layer, we obtain the predicted action chunk ^t _t, and predicted next latent observations ^t+Δt _t+ t as follows, ^tτ, ^t+Δtτ=(( , , )). ^τ_t,~ ^τ_t+ t= Prejection( Masked~Attention( , , )). (18) Finally, we regress the velocity field according the following way, ℒa= _a= τ∼[0,1], ∼ (, )[‖a( τ, +Δtτ, ,)−( − )‖22], _τ _[0,1], ( 0, ) [ \|u_a( ^τ_t, ^τ_t+ t, _t,l)-( _t- ) \|_2^2 ], (19) ℒo= _o= τ∼[0,1], ∼ (, )[‖a( τ, +Δtτ, ,)−( ^t+Δt− )‖22]. _τ _[0,1], ( 0, ) [ \|o_a( ^τ_t, ^τ_t+ t, _t,l)-( _t+ t- ) \|_2^2 ]. (20) The overall training objective is ℒ=ℒa+λℒo, =L_a+ _o, (21) where λ balances the action learning the next state prediction. TrainingLtO_ttA_tt+ΔtO_t+ tLtO_ttA_tt+ΔtO_t+ tInferenceLtO_ttA_tLtO_ttA_t Figure 3: Attention masks for training and inference. Rows are query-token groups and columns are key/value-token groups. L, tO_t, tA_t, and t+ΔtO_t+ t denote language, current observation, action, and predicted future observation, respectively. During training (left), L and tO_t attend bidirectionally; tA_t attends to ,t,t\L,O_t,A_t\; and t+ΔtO_t+ t attends to all groups. During inference (right), t+ΔtO_t+ t is omitted. Colored and white cells indicate permitted and masked attention, respectively. Left- and right-view tokens are merged into each observation group, and padded language positions are masked along both dimensions. 4.4 Training and Inference During training, the input tokens are arranged as language, left-view observation, right-view observation, action, observation, and future observation. The language and current-observation tokens form the conditioning group and attend bidirectionally to one another. They cannot attend to either the action tokens or the predicted future-observation tokens, which prevents information from the prediction targets from leaking into the conditioning representation. The action tokens attend to the entire conditioning group and to the action group itself, but they cannot access the predicted future observations. Therefore, robot actions are learned only from the language instruction, the current multi-view observations, and the dependencies within the action sequence. In contrast, the predicted future-observation tokens attend to all token groups, allowing the model to learn the physical state transition conditioned on the instruction, current observations, and robot actions. During inference, the two predicted future-observation groups are removed, while the visibility rules for the conditioning and action groups remain unchanged. The model therefore predicts the action sequence without generating any future visual tokens, which shortens the inference sequence, removes unnecessary visual generation, and enables efficient low-latency robot control. 5 Experiments In this section, we study the following three questions: How reliably does DELE-.5 w0.5 complete real-world manipulation tasks with different horizons? How does it compare with representative vision-language-action (VLA) policies under the same fine-tuning data and evaluation protocol? At which stages do the methods fail as task horizon and coordination difficulty increase? We answer these questions through full-task success, normalized stage progress, and task-wise stage-reach analysis. 5.1 Experimental Setup Robot and tasks. We conduct all experiments on the same Astribot S1 dual-arm robot in fixed real-world scenes. We consider four tasks with different manipulation horizons: opening a hinged door, retrieving a canned Pepsi from a refrigerator, adding ice to a cup, and loading a popcorn package into a microwave and starting the heating cycle. Figure 4 illustrates the task sequences, and Table 1 gives the ordered stages and conditions for full success. A stage is counted only when all preceding stages have been completed in order. Thus, a later action cannot compensate for a skipped prerequisite. Task-suite design. The task suite covers six sources of difficulty in long-horizon manipulation: contact-rich articulated interaction, target-object retrieval, bimanual transfer, tool use, constrained insertion, and terminal-state verification. The door task requires a stable handle grasp followed by timely handle rotation and pushing to release the latch. The Pepsi task adds target extraction, a right-to-left handoff, and refrigerator closure while retaining the can. The ice task is the longest sequence and requires coordinated use of the scoop, ice-maker lid, and cup, followed by pouring and scene restoration. The microwave task combines appliance opening, package pickup, constrained insertion, door closure, and heating activation. These tasks therefore test not only individual manipulation skills, but also whether a policy can connect the skills while preserving previously achieved task state. Figure 4: Representative stage sequences for the four real-robot tasks. Each row shows five temporally ordered key states from left to right for door opening, Pepsi retrieval, adding ice, and microwave popcorn. Table 1: Ordered task stages and complete task conditions. Stage zero denotes no effective progress, and reaching the final stage denotes full task completion. Task K Ordered stages Complete physical state Door opening 4 Handle contact → stable grasp → coordinated handle rotation and push → sustained opening Door remains beyond 45∘45 for at least 3 seconds. Pepsi retrieval 5 Handle interaction → door open → can extracted → left-hand handoff → door closed Left gripper retains the target can while the refrigerator is closed. Add ice 6 Scoop pickup → lid open → cup pickup → ice acquired → ice poured → scene restored Ice remains in the upright cup and the manipulated scene is restored. Microwave popcorn 5 Door open → package pickup → insertion → door closed → heating started Popcorn is inside the closed microwave and heating visibly starts. Baselines and training protocol. We compare DELE-.5 w0.5 with five representative VLA policies: π0.5 _0.5 [11], LingBot-VLA2 [40], Hy-Embodied-0.5-VLA (HY-VLA) [47], Spirit-v1.5 [32], and GR00T-N1.7 [25]. All methods are compared under identical fine-tuning data and evaluation protocols. Each baseline follows its official training configuration and is trained on the same task dataset for an equivalent of three epochs. At evaluation time, all methods use the same robot embodiment, task instructions, scene-reset configurations, 180-second execution budget, ordered-stage definitions, and full-task success criteria. Evaluation protocol. For each method-task pair, we run twenty formal trials, giving 6×4×20=4806× 4× 20=480 real-robot trials in total. Before each trial, we reset the robot and all task-relevant objects to predefined poses. Each policy has at most 180 seconds to complete the task. The operator provides no physical assistance after execution begins and stops a rollout only after observable success or terminal task failure, or when continued execution is unsafe. A rollout terminated for unsafe execution is counted as a failure. A trial is counted as a full success only when the task-specific complete condition in Table 1 is satisfied without human assistance. Metrics. We report two metrics. The first is the full-task success rate, which measures whether the complete physical task condition is satisfied. The second is normalized ordered-stage progress. For a task with K stages and a trial whose highest consecutively completed stage is s, the progress score is s/Ks/K. We average this score over twenty trials for each task and then macro-average over the four tasks. Therefore, every task has equal weight even though the tasks contain different numbers of stages. 5.2 Main Results Table 2 and Figure 5 compare normalized progress and full-task success. Results show that DELE-.5 w0.5 performs best on all four tasks. It obtains 81.3% macro progress and 62.5% overall success (50/80). The strongest baseline, π0.5 _0.5, obtains 50.6% macro progress and 15.0% success (12/80). Compared with this baseline, DELE-.5 w0.5 improves macro progress by 30.7 percentage points and full-task success by 47.5 percentage points. Table 2: Real-robot results over twenty trials per task. Each task entry is reported as normalized progress (Prog.) and full-task success (Succ.), in percent. Macro progress weights the four tasks equally; overall success is computed over all 80 trials per method. Door Pepsi Add ice Microwave Macro Overall Method Prog. Succ. Prog. Succ. Prog. Succ. Prog. Succ. Prog. Succ. DELE-.5 w0.5 (ours) 95.0 80.0 82.0 65.0 63.3 45.0 85.0 60.0 81.3 62.5 π0.5 _0.5 80.080.0 45.045.0 42.042.0 5.05.0 32.532.5 0.00.0 48.048.0 10.010.0 50.650.6 15.015.0 LingBot-VLA2 60.060.0 20.020.0 50.050.0 0.00.0 25.825.8 0.00.0 48.048.0 0.00.0 46.046.0 5.05.0 HY-VLA 57.557.5 0.00.0 35.035.0 0.00.0 37.537.5 10.010.0 32.032.0 15.015.0 40.540.5 6.36.3 Spirit-v1.5 57.557.5 15.015.0 35.035.0 0.00.0 14.214.2 0.00.0 29.029.0 0.00.0 33.933.9 3.83.8 GR00T-N1.7 40.040.0 0.00.0 25.025.0 0.00.0 9.29.2 0.00.0 14.014.0 0.00.0 22.022.0 0.00.0 Figure 5: Task-wise real-robot performance. The top panel reports full-task success, and the bottom panel reports normalized ordered-stage progress. Overall denotes aggregate success over 80 trials in the top panel and macro-average progress across the four tasks in the bottom panel. Figure 5 also shows a consistent gap between stage progress and full success. Several baselines complete early interactions, such as contacting a handle or opening an appliance, but rarely preserve this progress through object transfer, bimanual coordination, and terminal-state completion. DELE-.5 w0.5 converts early progress into full completion much more often. This result indicates that its advantage mainly appears when errors can accumulate over a long action sequence. 5.3 Task-wise and Stage-wise Analysis Figure 6 shows the empirical stage-reach probability P(s≥k)P(s≥ k) for each task. A single progress value measures how far a method moves on average, while the heatmaps show the exact stage at which performance drops. We make four observations from the results. Figure 6: Stage-wise long-horizon analysis. Each cell is the percentage of twenty trials whose ordered-prefix score reaches the indicated stage. Task blocks share a common model axis, and the outlined Full column in each block represents complete physical task success. Door opening. From Figure 6(a), DELE-.5 w0.5 achieves 16/20 full successes and 95.0% normalized progress, followed by π0.5 _0.5 with 9/20 successes and 80.0% progress. After grasping the handle, the robot must maintain the grasp, rotate the handle to a sufficient angle, and push at the correct time. DELE-.5 w0.5 reaches this coordinated rotation-and-push stage in all twenty trials, and its four failures occur only at the final sustained-opening condition. In comparison, most baselines reach the handle but fail to synchronize the rotation and push or fail to maintain the required opening angle. This result shows that handle perception alone is insufficient; the task depends on timely contact-rich coordination. Pepsi retrieval. Figure 6(b) shows that DELE-.5 w0.5 completes 13/20 Pepsi retrieval trials. It opens the refrigerator and extracts the target can in 18/20 trials. The remaining failures mainly occur at the right-to-left handoff and final door closure. LingBot-VLA2 reaches a stable left-hand handoff in six trials but never closes the refrigerator while retaining the can, whereas π0.5 _0.5 completes the full task once. Most other baselines fail during refrigerator opening. Thus, articulated interaction is the first bottleneck, while handoff and closure determine whether early progress becomes a complete retrieval. Adding ice. Figure 6(c) reports the longest task sequence. DELE-.5 w0.5 completes the six-stage task in 9/20 trials. It picks up the scoop in 19 trials, reaches cup pickup in 14, acquires ice in 11, and pours ice into the cup in nine. HY-VLA is the strongest competing method by progress (37.5%) and completes the task twice. π0.5 _0.5 obtains 32.5% progress but never places ice into the cup, while the remaining methods generally stop at scoop pickup or lid opening. These results show that a successful initial grasp is not enough. Repeated bimanual coordination and precise release determine the final performance. Microwave popcorn. From Figure 6(d), DELE-.5 w0.5 obtains 12/20 full successes and 85.0% normalized progress. It opens the microwave and picks up the package in every trial, closes the microwave in 16 trials, and starts heating in 12. Although π0.5 _0.5 and Spirit-v1.5 open the microwave in all twenty trials and LingBot-VLA2 does so in 19, their final success rates are only 2/20, 0/20, and 0/20. This sharp stage-wise drop explains why appliance opening alone is a weak measure of long-horizon performance. Package retention, insertion, door closure, and heating activation introduce successive opportunities for failure. 5.4 Failure Analysis For each unsuccessful trial, we record the first uncompleted ordered stage. This view identifies the main bottleneck without relying on model-specific runtime logs. The dominant failure stage differs across tasks. For door opening, most baselines grasp the handle but fail during coordinated rotation and pushing. For Pepsi retrieval, weaker methods fail before full extraction, while stronger methods lose progress during handoff or final refrigerator closure. For adding ice, the main drop occurs before ice acquisition and pouring. For microwave popcorn, many rollouts open the appliance but fail during insertion, door closure, or heating activation. All four bottlenecks require a policy to preserve previously achieved task state while changing the active arm, object, or interaction mode. DELE-.5 w0.5 reaches the late stages substantially more often, although its remaining failures still occur during object transfer, precise release, or the final task condition. The results therefore show that long-horizon performance depends both on individual skills and on the reliability of the transitions between them. 6 Conclusion In this work, we revisit the objective of World-Action Models for robotic manipulation. We show that video generation is not a necessary component of world modeling for robot control. Instead of reconstructing how the world visually evolves over time, a practical world model should predict how the physical world will change after an action is executed. Based on this observation, we introduce HeTu-1.0, a new future-state-driven world-action framework that infers robot actions from compact action-relevant future states without relying on video generation. By treating future-state prediction as the intermediate representation between world understanding and action generation, DELE removes unnecessary visual redundancy and enables efficient training and low-latency inference. Extensive real-world experiments demonstrate the effectiveness of this formulation on challenging long-horizon manipulation tasks. Across 480 robot trials on four diverse tasks, DELE consistently outperforms representative VLA baselines, achieving 62.5% full-task success and 81.3% normalized ordered-stage progress. The results show that the advantage of DELE becomes more significant as task horizons increase and errors accumulate across multiple interaction stages. These findings suggest that action-relevant future-state prediction provides a more direct and practical foundation for embodied world models, offering a promising direction toward scalable and deployable robot intelligence. References [1] A. Akbari, C. Zhang, A. Akbari, L. Zhao, Y. Chen, W. Chen, X. Zhang, G. Yuan, and Y. Wang (2026) Flash-wam: modality-aware distillation for world action models. External Links: 2606.05254, Link Cited by: §1.1, §2.2. [2] M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, Mojtaba, Komeili, M. Muckley, A. Rizvi, C. Roberts, K. Sinha, A. Zholus, S. Arnaud, A. Gejji, A. Martin, F. R. Hogan, D. Dugas, P. Bojanowski, V. Khalidov, P. Labatut, F. Massa, M. Szafraniec, K. Krishnakumar, Y. Li, X. Ma, S. Chandar, F. Meier, Y. LeCun, M. Rabbat, and N. Ballas (2025) V-jepa 2: self-supervised video models enable understanding, prediction and planning. External Links: 2506.09985, Link Cited by: §2.2. [3] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky (2026) π0 _0: A vision-language-action flow model for general robot control. External Links: 2410.24164, Link Cited by: §2.1. [4] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al. (2022) Rt-1: robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817. External Links: Link Cited by: §1. [5] J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y. Aytar, S. Bechtle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, and T. Rocktäschel (2024) Genie: generative interactive environments. External Links: 2402.15391, Link Cited by: §2.2. [6] J. Cen, C. Yu, H. Yuan, Y. Jiang, S. Huang, J. Guo, X. Li, Y. Song, H. Luo, F. Wang, D. Zhao, and H. Chen (2025) WorldVLA: towards autoregressive action world model. External Links: 2506.21539, Link Cited by: §2.2. [7] C. Cheang, G. Chen, Y. Jing, T. Kong, H. Li, Y. Li, Y. Liu, H. Wu, J. Xu, Y. Yang, H. Zhang, and M. Zhu (2024) GR-2: a generative video-language-action model with web-scale knowledge for robot manipulation. External Links: 2410.06158, Link Cited by: §2.2. [8] D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence (2023) PaLM-e: an embodied multimodal language model. In Proceedings of the 40th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 202, p. 8469–8488. External Links: Link Cited by: §2.1. [9] Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel (2023) Learning universal policies via text-guided video generation. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, p. 9156–9172. External Links: Document, Link Cited by: §1, §2.2. [10] D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap (2025) Mastering diverse control tasks through world models. Nature 640 (8059), p. 647–653. External Links: Link Cited by: §2.2. [11] P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky (2025) π0.5 _0.5: A vision-language-action model with open-world generalization. External Links: 2504.16054, Link Cited by: §1, §2.1, §5.1. [12] M. J. Kim, C. Finn, and P. Liang (2025) Fine-tuning vision-language-action models: optimizing speed and success. External Links: 2502.19645, Link Cited by: §2.1. [13] M. J. Kim, Y. Gao, T. Lin, Y. Lin, Y. Ge, G. Lam, P. Liang, S. Song, M. Liu, C. Finn, and J. Gu (2026) Cosmos policy: fine-tuning video models for visuomotor control and planning. External Links: 2601.16163, Link Cited by: §1. [14] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn (2024) OpenVLA: an open-source vision-language-action model. In 8th Annual Conference on Robot Learning, External Links: Link Cited by: §1, §2.1. [15] M. Koo, D. Choi, T. Kim, K. Lee, C. Kim, Y. Seo, and J. Shin (2026) HAMLET: switch your vision-language-action model into a history-aware policy. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §2.1. [16] G. Li, R. Wang, P. Xu, Q. Ye, and J. Chen (2025) The developments and challenges toward dexterous and embodied robotic manipulation: a survey. IEEE Robotics & Automation Magazine. External Links: Link Cited by: §1. [17] J. Li, T. Guo, Y. Ye, R. Zhang, X. Chi, Q. Sun, Y. Li, Y. Lou, Y. Huang, Z. Lu, M. Guo, and S. Zhang (2026) Efficient-wam: a 1b-parameter world-action model with low-cost future imagination. External Links: 2606.10040, Link Cited by: §1.1, §2.2. [18] K. Li, B. Hou, M. Zhu, T. Zhang, Z. Cheng, Z. Wang, S. Leng, X. Li, X. Lin, B. Yao, et al. (2026) RynnBrain 1.1: towards more capable and generalizable embodied foundation model. arXiv preprint arXiv:2607.17977. External Links: Link Cited by: §1. [19] L. H. Li, M. Yatskar, D. Yin, C. Hsieh, and K. Chang (2019) VisualBERT: a simple and performant baseline for vision and language. External Links: 1908.03557, Link Cited by: §4.1.2. [20] Q. Li, Y. Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei, Y. Deng, S. Xu, Y. Zhang, X. Wang, B. Liu, J. Fu, J. Bao, D. Chen, Y. Shi, J. Yang, and B. Guo (2024) CogACT: a foundational vision-language-action model for synergizing cognition and action in robotic manipulation. External Links: 2411.19650, Link Cited by: §2.1. [21] Z. Li, D. Cheng, Y. Wang, S. Wang, X. Xu, L. Weng, J. Wang, and J. Wang (2026) Light-wam: efficient world action models with state-fusion action decoding. External Links: 2606.08242, Link Cited by: §1.1. [22] Y. Liao, P. Zhou, S. Huang, D. Yang, S. Chen, Y. Jiang, Y. Hu, J. Cai, S. Liu, J. Luo, L. Chen, S. Yan, M. Yao, and G. Ren (2025) Genie envisioner: a unified world foundation platform for robotic manipulation. External Links: 2508.05635, Link Cited by: §2.2. [23] S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu (2025) RDT-1b: a diffusion foundation model for bimanual manipulation. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §2.1. [24] O. Mees, D. Ghosh, K. Pertsch, K. Black, H. R. Walke, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y. L. Tan, D. Sadigh, C. Finn, and S. Levine (2024) Octo: an open-source generalist robot policy. In First Workshop on Vision-Language Models for Navigation and Manipulation at ICRA 2024, External Links: Link Cited by: §2.1. [25] NVIDIA, :, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. ". Fan, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y. L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y. Xie, Y. Xu, Z. Xu, S. Ye, Z. Yu, A. Zhang, H. Zhang, Y. Zhao, R. Zheng, and Y. Zhu (2025) GR00T n1: an open foundation model for generalist humanoid robots. External Links: 2503.14734, Link Cited by: §5.1. [26] A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, Gupta, and Z. Lin (2024) Open x-embodiment: robotic learning datasets and rt-x models : open x-embodiment collaboration0. In 2024 IEEE International Conference on Robotics and Automation (ICRA), p. 6892–6903. External Links: Link Cited by: §2.1. [27] K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine (2025) FAST: efficient action tokenization for vision-language-action models. External Links: 2501.09747, Link Cited by: §2.1. [28] L. Qiu, Y. Li, Y. Chen, Y. Ge, Y. Ge, and X. Liu (2026) Making foresight actionable: repurposing representation alignment in world action models. External Links: 2606.12217, Link Cited by: §2.2, §2.2. [29] Y. Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel (2023) Multi-view masked world models for visual robotic manipulation. In Proceedings of the 40th International Conference on Machine Learning, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (Eds.), Proceedings of Machine Learning Research, Vol. 202, p. 30613–30632. External Links: Link Cited by: §2.2. [30] M. Shukor, D. Aubakirova, F. Capuano, P. Kooijmans, S. Palma, A. Zouitine, M. Aractingi, C. Pascal, M. Russi, A. Marafioti, S. Alibert, M. Cord, T. Wolf, and R. Cadene (2025) SmolVLA: a vision-language-action model for affordable and efficient robotics. External Links: 2506.01844, Link Cited by: §2.1. [31] O. Siméoni, H. V. Vo, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V. Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haziza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Couprie, J. Mairal, H. Jégou, P. Labatut, and P. Bojanowski (2025) DINOv3. External Links: 2508.10104, Link Cited by: §3.2. [32] Spirit AI (2026) Spirit-v1.5: clean data is the enemy of great robot foundation models. Note: Spirit AI BlogPublished January 11, 2026 External Links: Link Cited by: §5.1. [33] Y. Tian, S. Yang, J. Zeng, P. Wang, D. Lin, H. Dong, and J. Pang (2025) Predictive inverse dynamics models are scalable learners for robotic manipulation. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §2.2. [34] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. In Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30, p. . External Links: Link Cited by: §4.1.2. [35] Q. Wang, M. Li, J. Guan, J. Ye, S. Xie, Y. Liu, J. Chen, Z. Liang, J. Zhang, X. Hu, X. Huang, P. Lin, J. Lin, D. Liu, S. Bai, J. Zhou, J. Zhang, H. Yuan, G. Zhou, H. Yin, Y. Wang, Y. Huang, Z. Lei, W. Peng, D. Chen, Y. Zheng, J. Fan, X. Zhuang, X. Zhou, H. Li, A. Chen, T. Zhang, X. Liu, Y. Sun, R. Chen, Z. Li, C. Lü, Z. Yang, T. Yu, and X. Chen (2026) Qwen-vla: unifying vision-language-action modeling across tasks, environments, and robot embodiments. External Links: 2605.30280, Link Cited by: §2.1. [36] Y. Wang, X. Li, W. Wang, J. Zhang, Y. Li, Y. Chen, X. Wang, and Z. Zhang (2025) Unified vision-language-action model. External Links: 2506.19850, Link Cited by: §2.1. [37] J. Wen, Y. Zhu, J. Li, Z. Tang, C. Shen, and F. Feng (2025) DexVLA: vision-language model with plug-in diffusion expert for general robot control. External Links: 2502.05855, Link Cited by: §2.1. [38] J. Wen, Y. Zhu, J. Li, M. Zhu, Z. Tang, K. Wu, Z. Xu, N. Liu, R. Cheng, C. Shen, et al. (2025) TinyVLA: toward fast, data-efficient vision-language-action models for robotic manipulation. IEEE Robotics and Automation Letters 10 (4), p. 3988–3995. External Links: Link Cited by: §2.1. [39] H. Wu, Y. Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong (2023) Unleashing large-scale video generative pre-training for visual robot manipulation. External Links: 2312.13139, Link Cited by: §2.2. [40] W. Wu, F. Wang, F. Lu, H. Sun, S. Liu, Y. Wang, Y. Yan, Y. Wang, S. Ma, X. Wang, Y. Liu, S. Yang, T. Zhou, K. Zhang, L. Zhou, C. Su, N. Xue, B. Tan, H. Zhang, Y. Zhang, F. Liao, X. Zhu, Y. Shen, and K. Zheng (2026) From foundation to application: improving vla models in practice. External Links: 2607.06403, Link Cited by: §5.1. [41] H. Yan, Q. Li, J. Yang, and Y. Mu (2026) ProgressVLA: progress-guided diffusion policy for vision-language robotic manipulation. External Links: 2603.27670, Link Cited by: §2.1. [42] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu (2025) Qwen3 technical report. External Links: 2505.09388, Link Cited by: §3.2. [43] A. Ye, B. Wang, C. Ni, G. Huang, G. Zhao, H. Li, H. Li, J. Li, J. Lv, J. Liu, M. Cao, P. Li, Q. Deng, W. Mei, X. Wang, X. Chen, X. Zhou, Y. Wang, Y. Chang, Y. Li, Y. Zhou, Y. Ye, Z. Liu, and Z. Zhu (2026) GigaWorld-policy: an efficient action-centered world–action model. External Links: 2603.17240, Link Cited by: §2.2. [44] S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang, A. Malik, K. Lee, W. Liang, N. Ranawaka, J. Gu, Y. Xu, G. Wang, F. Hu, A. Narayan, J. Bjorck, J. Wang, G. Kim, D. Niu, R. Zheng, Y. Xie, J. Wu, Q. Wang, R. Julian, D. Xu, Y. Du, Y. Chebotar, S. Reed, J. Kautz, Y. Zhu, L. ". Fan, and J. Jang (2026) World action models are zero-shot policies. External Links: 2602.15922, Link Cited by: §1.1, §1, §2.2. [45] T. Yuan, Z. Dong, Y. Liu, and H. Zhao (2026) Fast-wam: do world action models need test-time future imagination?. External Links: 2603.16666, Link Cited by: §1.1. [46] B. Zhang and R. Sennrich (2019) Root mean square layer normalization. In Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32, p. . External Links: Link Cited by: §4.1.3. [47] H. Zhang, L. Xiang, H. Lin, Z. Huang, M. Wang, D. Zhong, Y. Dong, Y. Wu, Y. Rao, D. Zhang, W. He, L. Chen, K. Huang, J. Chen, S. Su, X. Yu, Z. Wang, C. Zhu, X. Teng, Y. Guo, Y. Zhang, Y. Liu, R. Wang, Z. Lu, H. Hu, and Z. Zhang (2026) Hy-embodied-0.5-vla: from vision-language-action models to a real-world robot learning stack. External Links: 2606.14409, Link Cited by: §5.1. [48] Y. Zhang, W. Zhang, Z. Qi, H. Zhang, H. Lin, J. Zhang, Y. Mu, X. Yang, W. Zeng, and X. Jin (2026) ImageWAM: do world action models really need video generation, or just image editing?. External Links: 2606.19531, Link Cited by: §1.1. [49] Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn, A. Handa, T. Lin, G. Wetzstein, M. Liu, and D. Xiang (2025) CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models . In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 1702–1713. External Links: Link Cited by: §2.2. [50] H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y. Du, Y. Hong, and C. Gan (2024) 3D-vla: a 3d vision-language-action generative world model. External Links: 2403.09631, Link Cited by: §2.2. [51] P. Zhou, S. Chen, D. Chen, J. Wang, R. Jin, B. Zhu, Y. Pan, S. Gu, K. Wang, S. Nan, X. Qiu, C. Qiu, P. Yang, Y. Cai, J. Gao, Y. Li, Y. Fu, X. Yue, Z. Chen, and J. Luo (2026) τ0 _0-WM: a unified video-action world model for robotic manipulation. External Links: 2606.01027, Link Cited by: §2.2. [52] F. Zhu, H. Wu, S. Guo, Y. Liu, C. Cheang, and T. Kong (2025) IRASim: a fine-grained world model for robot manipulation. External Links: 2406.14540, Link Cited by: §2.2. [53] B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V. Vanhoucke, H. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. Sanketi, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y. Lu, S. Levine, L. Lee, T. E. Lee, I. Leal, Y. Kuang, D. Kalashnikov, R. Julian, N. J. Joshi, A. Irpan, B. Ichter, J. Hsu, A. Herzog, K. Hausman, K. Gopalakrishnan, C. Fu, P. Florence, C. Finn, K. A. Dubey, D. Driess, T. Ding, K. M. Choromanski, X. Chen, Y. Chebotar, J. Carbajal, N. Brown, A. Brohan, M. G. Arenas, and K. Han (2023) RT-2: vision-language-action models transfer web knowledge to robotic control. In Proceedings of The 7th Conference on Robot Learning, J. Tan, M. Toussaint, and K. Darvish (Eds.), Vol. 229, p. 2165–2183. External Links: Link Cited by: §1, §2.1.