Paper deep dive
Keep the Future, Drop the Rollout: RIFT for World Action Models
Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on all 40 LIBERO tasks, paired closed-loop interventions show that masking or reassigning future-cache values changes execution and reduces success, indicating sensitivity to future values and their assigned positions. For Joint and Cosmos-2, however, replaying one fixed final-clean key/value (K/V) cache nearly preserves unmodified execution, with $1.7$ to $1.9$~cm end-effector average displacement error and $97.9\%$ to $98.2\%$ success. This separates cache consumption from production: these models can reuse a fixed cache but still require iterative rollout to construct it. We therefore propose RIFT (\emph{Rollout-free Imagination via Future Tokens}), which uses learned anticipation tokens to construct a complete future K/V cache in one backbone pass while retaining the original future-read interface. On LIBERO, RIFT achieves $98.8\%$ success, close to rollout-based Joint, IDM, and LingBot-VA at $98.4\%$ to $98.6\%$, while reducing action-chunk latency by $68.2\%$ to $89.1\%$. On RoboTwin~2.0, RIFT reaches $92.9/92.6\%$ on clean/randomized scenes, the highest observed among the evaluated methods. These results support rollout-free future conditioning without iterative video generation at deployment.
Tags
Links
- Source: https://arxiv.org/abs/2608.11521v1
- Canonical: https://arxiv.org/abs/2608.11521v1
Trouble viewing inline? Open PDF directly →
Full Text
61,166 characters extracted from source content.
Expand or collapse full text
Keep the Future, Drop the Rollout: Rift for World Action Models Chushan Zhang Jinguang Tong Xuesong Li Yikai Wang Hongdong Li Abstract World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on all 40 LIBERO tasks, paired closed-loop interventions show that masking or reassigning future-cache values changes execution and reduces success, indicating sensitivity to future values and their assigned positions. For Joint and Cosmos-2, however, replaying one fixed final-clean key/value (K/V) cache nearly preserves unmodified execution, with 1.71.7 to 1.91.9 cm end-effector average displacement error and 97.9%97.9\% to 98.2%98.2\% success. This separates cache consumption from production: these models can reuse a fixed cache but still require iterative rollout to construct it. We therefore propose Rift (Rollout-free Imagination via Future Tokens), which uses learned anticipation tokens to construct a complete future K/V cache in one backbone pass while retaining the original future-read interface. On LIBERO, Rift achieves 98.8%98.8\% success, close to rollout-based Joint, IDM, and LingBot-VA at 98.4%98.4\% to 98.6%98.6\%, while reducing action-chunk latency by 68.2%68.2\% to 89.1%89.1\%. On RoboTwin 2.0, Rift reaches 92.9/92.6%92.9/92.6\% on clean/randomized scenes, the highest observed among the evaluated methods. These results support rollout-free future conditioning without iterative video generation at deployment. 1 Introduction Figure 1: LIBERO success versus deployment latency. Points show mean success and bars show ± over three evaluation seeds of one checkpoint per method (2,0002,000 trials per seed); latency is ms per action chunk on one A800. Gray squares use no future read; blue circles read a future representation produced by rollout; Rift (orange star) reads one-pass anticipation tokens without rollout. The x-axis is logarithmic; the y-axis is truncated. Methods use their original denoising configurations. World action models (WAMs) couple future video prediction with robot control. A single network predicts a short video and conditions its next actions on that prediction (45; 24; 4). In matched settings, policies that retain this future read achieve higher success than current-only variants. However, iterative video generation makes rollout-based systems incur 3.3×3.3× to 9.6×9.6× the latency of current-only deployment (fig. 1). Existing efficient variants remove the future read at deployment. Fast-WAM drops the future branch after future-prediction co-training (45), whereas PFD distills a future-conditioned correction into the current-only path (14). Both reduce latency but retain a gap to policies that explicitly read a generated future. Yet the action expert consumes a future representation, whereas iterative rollout is the process that constructs it. Existing comparisons change both the availability of this representation and the process used to construct it, so their separate contributions remain unresolved. We therefore ask whether a WAM can preserve explicit test-time future conditioning without iterative video rollout. To separate these factors, we examine the future-position K/V cache, the per-layer channel from predicted futures to action tokens. Joint co-denoising updates this cache at every denoising step, whereas the generate-then-act inverse dynamics model (IDM) exposes only a final clean cache. Their similar success suggests that the representation, rather than its iterative trajectory, may provide the shared benefit. Success alone cannot reveal what the action reads. We therefore intervene on this cache. The attention mask prevents video tokens from attending to action tokens, so the cache is an action-independent intervention site. We record it and replay action denoising with non-target inputs fixed, either masking the future read or editing future values under the recorded keys. Each closed-loop intervention uses the same initial state and policy seed as the unmodified policy. We measure success rate (SR) and end-effector average displacement error (E-ADE), the average drift from the unmodified end-effector trajectory. Across 2,0002,000 paired trials on all 40 LIBERO tasks, masking the future read yields 18.718.7 cm E-ADE and reduces success from 98.4%98.4\% to 9.7%9.7\%. Spatial permutation and temporal swapping yield similar E-ADE (14.314.3 and 15.615.6 cm) but sharply different success (65.2%65.2\% and 0.7%0.7\%), showing that the action reads future content at its assigned positions. In contrast, replaying final clean values under the original keys yields only 1.91.9 cm E-ADE and 97.9%97.9\% success. Under the original keys, the action therefore depends strongly on future value content and its organization, but little on how the future values evolve across denoising steps. These findings provide a value-side constraint rather than a complete cache recipe. They do not show that the key trajectory can be frozen or that a complete cache can be produced without rollout. This limitation motivates the stronger hypothesis that one pass can produce the complete future interface. We propose Rift (Rollout-free Imagination via Future Tokens), which places learned anticipation tokens at future temporal positions and fills their per-layer K/V cache with one video-backbone pass. The action expert consumes this cache through the original future-read interface. We shape the anticipation states with conditional flow matching, using a distributional objective rather than direct L2 regression to the single observed future. Deployment requires one cache prefill followed by the ordinary action flow, without video diffusion or video decoding at test time. On LIBERO, Rift achieves 98.8%98.8\% success, compared with 96.8%96.8\% for current-only Fast-WAM and within the same success tier as rollout-based Joint (98.4%98.4\%) and IDM (98.6%98.6\%). It runs at 1.1×1.1× current-only latency rather than 3.3×3.3× to 9.6×9.6× for rollout-based alternatives. On RoboTwin 2.0, Rift reaches 92.9/92.6%92.9/92.6\% success on clean/randomized scenes, the highest observed among the evaluated methods. These results indicate that, for the studied WAM family, explicit future-position representations can support rollout-level success without requiring iterative video generation at deployment. We make three contributions. 1. We introduce a paired closed-loop intervention protocol for future caches. The protocol records the action-independent cache, then either masks the future read or edits its values under the recorded keys before measuring the resulting physical execution with E-ADE and SR. 2. We identify two properties of the future cache. Actions are highly sensitive to removing the future read or reassigning values across space and time, while final clean values under the original keys nearly preserve execution. 3. We design and validate a rollout-free future interface. Rift produces a complete future-position K/V cache in one backbone pass, matches the rollout-based success tier at 1.1×1.1× current-only latency, and reaches the highest observed LIBERO and RoboTwin 2.0 success. 2 Related Work Five lines of work set up the question this paper asks: policies with current-only deployment, policies that render a future at deployment, policies that use future prediction only during training, runtime uncertainty monitors, and the interpretability methods we borrow from. Vision-language-action policies. Vision-language-action policies map observations and language directly to actions through pretrained vision-language backbones (52; 23; 6; 34; 5; 38; 16), often paired with diffusion or flow-matching action heads (12; 30; 41). Because the mapping is direct, inference needs no iterative video branch, which makes these policies the natural latency reference for methods that add one. Their action experts receive no explicit future-position state. Our current-only baseline occupies this regime, and the gap it leaves is what we try to close without rollout cost. World action models. A complementary line makes future prediction explicit. Classical world-model methods learn dynamics and use them for planning or control (36; 17; 18). More recent robot policies condition on generated future video (13; 7; 3; 50) or jointly learn future and action representations (43; 9; 19; 25; 51; 15; 33; 42; 26; 21; 48; 8; 22; 27). World action models tighten this coupling so that a single network both imagines and acts (44; 45; 24; 4). Two interfaces dominate. Imagine-then-execute systems generate a future video and then apply inverse dynamics to it, so the action reads a fixed clean representation. Mixture-of-transformers systems denoise video and action through shared attention, so the action reads a representation that changes at every denoising step. These two designs expose the future in almost opposite ways yet report similar success. Our intervention study starts from this observation. Both pay for iterative video generation, which dominates their latency. Prior comparisons report end-task success only, which cannot distinguish a policy that uses the imagined future from one that merely benefits from training alongside it. Implicit-future policies. A third line keeps future prediction as training-time signal while avoiding a visual rollout at inference. Fast-WAM removes the explicit future representation at deployment and preserves much of the aggregate success through a world-aware current representation, which shows how much of the benefit survives co-training alone (45). PFD instead distills the action-side effect of a generated future into a lightweight current-only correction (14; 40; 10). FLARE (49) and DreamVLA (47) learn future-prediction tokens through auxiliary objectives without an explicit test-time rollout, and Being-H0.7 trains latent queries with a future-informed posterior branch that is discarded at inference (31). EvoScene-VLA supervises an action-updated recurrent scene prefix with future scene-token targets and likewise discards its training-time scene predictor at deployment (46). These methods remove a rollout-produced future read, distill its effect, or use future targets to train a compact latent state. In our matched evaluation, Fast-WAM and PFD leave a residual gap to rollout-based policies. Rift takes the opposite decomposition: it keeps an explicit test-time future-position interface and replaces the iterative producer with a single learned pass. Flow-matching uncertainty and runtime monitoring. Concurrent work reads uncertainty from the action flow itself: 35 measures denoising-path acceleration along a single FM action trajectory, validates it against the L2 divergence of resampled action chunks, and accumulates the score with CUSUM for failure detection. Our auxiliary readouts expose a complementary signal in predicted future-latent space. In an optional shadow mode, the stopped-gradient L2 probe supplies a deterministic readout while the conditional-FM head supplies sampled modes; their cross-head discrepancy, normalized by within-FM spread, measures cross-estimator conflict rather than action-flow curvature. Writing μ for the probe readout and x¯ x for the mean of K FM samples xix_i, the score is dratio=∥μ−x¯∥22/(K−1∑i∥xi−x¯∥22+ϵ)d_ratio= μ- x _2^2/(K^-1 _i x_i- x _2^2+ε). Thus their proxy estimates uncertainty internal to one FM action head, whereas ours asks whether two future estimators agree. Both signals can miss confidently wrong predictions. This monitor is not part of the policy-only deployment or latency results. Locating computation by intervention. Our instrument follows causal tracing, which intervenes on selected internal activations to measure their causal effect (32). We apply this idea to a robot policy’s attention cache rather than to an autoregressive language model’s hidden states. This matters because the standard alternatives answer a weaker question: linear probes establish that information is decodable from a representation, not that downstream computation uses it, and attention weights are similarly unreliable as evidence of use (1; 2; 20; 37). Intervening on the future read and measuring the executed trajectory tests use directly. Two properties of the WAM setting make this test unusually clean. The video-to-action attention mask makes the recorded cache action-independent, so we can either mask the read or edit its values under fixed keys. Closed-loop execution then supplies a physical readout in centimeters rather than a distance in an arbitrary latent space. For value edits, the remaining caveat is distributional rather than positional, and we return to it in the next section. 3 What does the action expert read? Figure 2: Future-cache interventions alter executed trajectories and task success. Value-edit rows retain the recorded Original key trajectory and modify only future-position values. Final-clean replay instead substitutes one final-clean K/V cache at every action-denoising step, while masking removes the future read. Bar length gives task-macro E-ADE in centimeters; the label at each bar end gives SR in percent. Each reported intervention result uses 2,0002,000 paired trials over all 4040 LIBERO tasks and is paired with the same model’s Original. Original has zero E-ADE by definition and is marked by colored ticks at the origin. All numbers are direct measurements. Missing bars are not zero: IDM-style models already read one fixed final-clean future K/V cache, while Cosmos-2 exposes only one future timestep and therefore admits no temporal swap. 3.1 The channel we edit With paired closed-loop interventions, we test whether action needs the future read, whether its values must stay at assigned positions, and whether the complete K/V cache must evolve during denoising. Each model–intervention estimate uses 2,0002,000 paired trials across all 4040 LIBERO tasks. Figure 2 reports executed E-ADE and success rate. Fast-WAM-Joint, Fast-WAM-IDM, Cosmos Policy, and LingBot-VA construct futures differently but expose the same per-layer video K/V interface at future positions. We call it the future cache; Original denotes the unmodified checkpoint. We mask the read, edit future values under recorded keys, or replace K and V with one final-clean cache. Because video tokens cannot attend to action tokens, the cache is action-independent given observation (o), language (l), and video-generation randomness. This property supports record and replay. The record pass generates video normally and stores per-layer future K/V at every action-denoising step. Value corruptions retain the recorded key trajectory and edit only the matched future values. The final-clean control instead takes the final clean future from Original’s iterative generation, prefills its complete K/V once, and reuses that fixed cache at every action-denoising step. All non-target inputs remain fixed. Exact replay of the unedited trajectory reproduces Original. Shuffle and noise are location-exact but out of distribution, so we interpret them only beside the structured frozen-present and final-clean K/V controls. 3.2 Scoring each edit Action chunks are not directly comparable across architectures, so we execute paired policies. For episode i, intervention I and Original share the initial state and policy seed. Let i,tIx^I_i,t and i,tOx^O_i,t denote their recorded end-effector positions after environment step t. E-ADE averages their distance over the common executed prefix TiT_i: E-ADEi(I)=1Ti∑t=1Ti‖i,tI−i,tO‖2.E -ADE_i(I)= 1T_i _t=1^T_i ^I_i,t-x^O_i,t _2. (1) We exclude the reset pose and use recorded simulator positions rather than integrating predicted actions. When runs end at different times, only their common prefix contributes. We average episode scores within each task before macro-averaging tasks; appendix B defines the task-cluster bootstrap. E-ADE measures drift from Original; success measures task completion. High drift with similar success indicates another route; high drift with low success associates the intervention with failure during closed-loop task execution. 3.3 Intervention set The interventions map directly to the three questions. Masking removes the read; norm-matched noise replaces its values; and frozen-present supplies plausible, non-predictive values. Spatial shuffle permutes values within frames, temporal swap exchanges value frames, and final-clean replay replaces the evolving complete cache with the same final-clean K/V at every action step. We apply each supported edit to Fast-WAM-Joint, Fast-WAM-IDM, Cosmos Policy, and LingBot-VA, comparing each with its Original. Claims are therefore within-model; cross-model magnitudes remain descriptive because architectures and checkpoints differ. 3.4 Finding 1: WAM action experts use future values at their assigned positions Across all four WAMs, masking yields 11.811.8–20.420.4 cm E-ADE and reduces success from 98.498.4–98.6%98.6\% to 0.00.0–32.0%32.0\%. Under recorded keys, noise yields 10.510.5–17.917.9 cm and at most 40.9%40.9\% success, while frozen-present values yield 18.118.1–21.321.3 cm and at most 6.5%6.5\%. These within-model effects show that action experts use meaningful future values. Position also matters: spatial shuffle yields 5.05.0–19.819.8 cm and 0.00.0–84.5%84.5\% success across all four, while temporal swap yields 15.615.6–16.316.3 cm and 0.00.0–69.0%69.0\% on Joint, IDM, and LingBot-VA. Every supported edit changes execution and lowers own-model success, so future values are not an unordered pool. Severity differs by interface: temporal swap is worse for Joint and IDM, spatial shuffle for LingBot-VA, and Cosmos-2 exposes no temporal swap. Thus, position sensitivity is shared, not a universal severity ordering. 3.5 Finding 2: one final-clean K/V cache nearly preserves execution Finding 1 establishes content and position sensitivity, not whether the complete cache must evolve. Where supported, final-clean replay replaces the entire future-cache trajectory with one final-clean cache, holding both K and V fixed at every action-denoising step. For Joint and Cosmos-2, final-clean K/V replay gives 1.9/1.71.9/1.7 cm E-ADE and 97.9/98.2%97.9/98.2\% success; their Originals reach 98.4/98.4%98.4/98.4\%. These are the smallest nonzero E-ADEs among compatible edits. IDM and LingBot-VA already expose one fixed final-clean K/V cache, so this replay is identical by construction and structurally N/A. This result establishes consumption-side sufficiency: once final-clean K/V is available, Joint and Cosmos-2 nearly preserve execution without the evolving cache trajectory. It does not show that this cache can be produced without rollout, because its clean future came from iterative video generation. The full-cache intervention also does not isolate the separate contributions of keys and values. Section 4 tests the distinct producer-side hypothesis that one learned prefill can construct an effective fixed, complete K/V interface. 4 Rift: One-pass Future-Token Imagination 4.1 From findings to our design The analysis shows that WAM action experts require meaningful, position-bound future values and that Joint and Cosmos-2 can reuse one final-clean, complete K/V cache throughout action denoising with near-Original execution. This establishes a fixed complete cache as a viable consumption interface, but not its rollout-free production: the intervention cache still comes from iterative video generation. Thus, Rift tests whether one learned prefill can produce the complete interface. Figure 3: Rift training and deployment. (a) One VideoStack prefill maps the first-frame latent and anticipation tokens to a fixed per-layer future-position K/V cache, which the action expert reuses throughout denoising. (b) Training pairs native video supervision with a deployment-matched forward for action, conditional-FM, and a stopped-gradient mean-squared probe loss. Action rows use the clean first frame; late perturbation affects only future-supervised rows. At policy-only test time, video co-training, auxiliary heads, and ground-truth futures are removed, leaving the prefill, fixed cache, and action flow. 4.2 Writing the cache in one pass Rift keeps Fast-WAM-Joint’s architecture and future-read interface. For a video stack with hidden width d and L layers, it replaces rolled-out future tokens with learned anticipation tokens E∈ℝm×dE ^m× d. Each token inherits its corresponding future spatiotemporal index. If each latent frame contains n tokens and the clip contains TlatT_lat latent frames, full alignment uses m=n(Tlat−1)m=n(T_lat-1). As fig. 3 shows, we use full alignment with m=196m=196 on LIBERO and m=240m=240 on RoboTwin 2.0. Let f0(o)f_0(o) denote the first-frame tokens extracted from observation o; the remaining observation and instruction l enter through the backbone’s original conditioning path. Let ϕφ collect the shared video expert’s parameters and the learned tokens E; together they form the cache producer. One video-stack prefill writes the full future-position cache: ϕ(o,l)=(KE(ℓ),VE(ℓ))ℓ=1L=CachePrefillϕ([f0(o);E],o,l), splitC_φ(o,l)&= \ (K_E^( ),V_E^( ) ) \_ =1^L\\ &=CachePrefill_φ\! ([f_0(o);E],o,l ), split (2) where KE(ℓ),VE(ℓ)∈ℝm×dK_E^( ),V_E^( ) ^m× d denote the layer-ℓ keys and values of the m anticipation tokens. We retain Fast-WAM-Joint’s mask: first-frame tokens attend only the observed frame, anticipation tokens attend the first frame and one another, and video tokens never attend action tokens. The cache is therefore action-independent. The action expert reads ϕC_φ through the rollout model’s per-layer future-position interface: a^1:H=ActionDenoise(o,l;ϕ(o,l)). a_1:H=ActionDenoise (o,l;C_φ(o,l) ). (3) Here, H is the action-chunk horizon; we use H=32H=32. ϕC_φ has no action-flow index: the same K/V serve every denoising evaluation, matching the fixed-cache consumption pattern tested in Finding 2. The producer-side hypothesis is that one pass from (o,l)(o,l) constructs a usable cache; deployment then needs one prefill per chunk, with no rollout or VAE decoding. 4.3 Training Training uses two forwards through the same video expert per optimization step. Deployment retains only the second forward’s cache prefill. Native video supervision. The first forward applies native video-flow loss ℒvidL_vid (45) to a clean-first-frame clip with noised future latents and no anticipation tokens or action supervision. This preserves dynamics supervision while changing only the deployment interface. Deployment-matched action training. The second forward matches deployment through input [f0;E][f_0;E] and its attention mask. Clean rows train the action expert on this cache with Fast-WAM’s inherited action flow-matching loss ℒactL_act (45). After 70%70\% of the configured curriculum horizon, the probability and scale of first-frame latent noise rise linearly from zero to 0.30.3 and 0.060.06 times the latent standard deviation. Perturbed rows retain ℒFML_FM and ℒprobeL_probe but are masked out before ℒactL_act is reduced. Thus, ℒactL_act averages only clean rows, so all action-loss inputs are clean. Conditional flow-matching supervision. Let Sϕ∈ℝm×dS_φ ^m× d be the final anticipation states from the same deployment-matched forward and Y∈ℝm×dyY ^m× d_y the aligned ground-truth future latent patches, where dyd_y is the dimension of one flattened latent patch. Under a multimodal conditional distribution, direct ℓ2 _2 regression has a conditional-mean optimum and can average distinct valid futures; we instead use conditional flow matching (28) as a distributional auxiliary objective. We sample ϵ∼(0,I)ε (0,I) and σ∈[0,1]σ∈[0,1] with the native video-flow schedule, then define Xσ=(1−σ)Y+σϵ,vσ⋆=ϵ−Y.X_σ=(1-σ)Y+σε, v_σ =ε-Y. (4) Conditioned on SϕS_φ, the training-only FM head vψv_ψ, parameterized by ψ, predicts the velocity from (Xσ,σ)(X_σ,σ). Using the native video-flow timestep weight wvid(σ)w_vid(σ), its loss is ℒFM=Y,ϵ,σ[wvid(σ)‖vψ(Xσ,σ,Sϕ)−vσ⋆‖22],L_FM=E_Y,ε,σ\! [w_vid(σ) \|v_ψ(X_σ,σ;S_φ)-v_σ \|_2^2 ], (5) This auxiliary head shapes the cache producer during training but enters neither action nor policy-only deployment. Stopped-gradient linear probe. The direct-L2 recipe provides a deterministic future-latent readout. In the final conditional-FM recipe, we retain this view as a linear probe, Y^L2=gω(RMS(stopgrad(Sϕ))) Y_L2=g_ω(RMS(stopgrad(S_φ))). We train it with MSE, ℒprobe=mean((Y^L2−Y)2). [rgb]0,0,0L_probe=mean\! (( Y_L2-Y)^2 ). The reduction weights every future-token and latent-channel squared residual uniformly before the batch mean. Detached input confines this loss to the probe; appendix A details the reduction. In optional shadow mode, the probe point prediction and conditional-FM samples define a normalized disagreement score. Both read the same anticipation states, but neither feeds the controller; appendix D defines the score. Objective and gradient routes. Both forwards share the video expert. ℒvidL_vid updates this expert and its video head; ℒactL_act updates the action expert and backpropagates through the cache into the shared video expert and E; ℒFML_FM updates the shared video expert, E, and the FM head; and detached ℒprobeL_probe updates only the probe. ℒ=ℒvid+ℒact+λFMℒFM+λprobeℒprobe,L=L_vid+L_act+ _FML_FM+ _probeL_probe, (6) Both auxiliary weights stay at 11 for the first 70%70\% of the curriculum horizon, then follow a cosine decay to 0.20.2 over the final 30%30\%. Policy-only deployment retains one [f0;E][f_0;E] prefill, fixed ϕC_φ, and the standard action flow; appendix A gives remaining settings. 5 Experiments 5.1 Setup Implementation. For LIBERO, Fast-WAM, Fast-WAM-Joint, Fast-WAM-IDM, and Rift share the same Wan2.2-5B (39) Fast-WAM backbone, training data, and 2020k-step budget, so differences come from the future interface rather than scale or data. We additionally compare with the Fast-WAM-based released PFD checkpoint (14) and the embodied-pretrained LingBot-VA (24). Unless stated otherwise Rift means the full conditional-FM recipe, with Rift-L2 reserved for the base recipe in the ablations. Latency is milliseconds per action chunk on one A800 under each method’s own denoising configuration; Rift’s figure covers its cache prefill and action denoising and excludes optional diagnostic readouts. In the tables, future read denotes explicit future-position attention; rollout denotes iterative future generation at deployment. Benchmarks. LIBERO (29) has 4040 tasks across its Spatial, Object, Goal, and Long suites. We evaluate one checkpoint per method with three evaluation seeds and 2,0002,000 trials per seed (5050 episodes per task), reporting mean and standard deviation across seeds; the error bars therefore measure closed-loop evaluation noise, not variation across training runs. RoboTwin 2.0 (11) has 5050 bimanual tasks under clean and domain-randomized scenes. Following 45, the matched Fast-WAM/Joint/PFD/Rift family trains on 2,5002,500 clean-scene and 25,00025,000 randomized demonstrations for 3030k steps, with externally pretrained LingBot-VA for comparison. Each checkpoint is evaluated for 100100 trials per task in each setting. 5.2 Quantitative results LIBERO. In Table 1, Rift achieves 98.8%98.8\% overall success, close to the 98.4%98.4\% to 98.6%98.6\% achieved by rollout-based Joint, IDM, and LingBot-VA. Unlike these methods, Rift requires only 247.9247.9 ms per action chunk, reducing latency by 68.2%68.2\% to 89.1%89.1\% while remaining close to current-only Fast-WAM at 235.7235.7 ms. Compared with rollout-free Fast-WAM and PFD, Rift improves success by 2.02.0 and 1.51.5 percentage points, respectively, at comparable latency. Table 1: LIBERO success and deployment cost. SR is mean± over three evaluation seeds of one checkpoint (2,0002,000 trials each). Latency is ms per action chunk on one A800 under each method’s original denoising configuration. Emb. PT.: embodied pretraining. † released checkpoint evaluated by us; ∗ our matched reproduction. Method Emb. Fut. Roll- SR Lat. Rel. PT. read out (%) (ms) Fast-WAM† × × × 96.8±0.2796.8 \,± 0.27 235.7 1.0×1.0× PFD† × × × 97.3±0.1297.3 \,± 0.12 257.0 1.1×1.1× Fast-WAM-Joint∗ × ✓ ✓ 98.4±0.2698.4 \,± 0.26 780.2 3.3×3.3× Fast-WAM-IDM∗ × ✓ ✓ 98.6±0.3498.6 \,± 0.34 1081.2 4.6×4.6× LingBot-VA† ✓ ✓ ✓ 98.5±0.0898.5 \,± 0.08 2270.3 9.6×9.6× Rift (ours) × ✓ × 98.8±0.1798.8 \,± 0.17 247.9 1.1×1.1× RoboTwin 2.0. The interface transfers to a second embodiment (table 2; per-task rates in table 5). Rift reaches 92.9/92.692.9/92.6 on clean/randomized scenes, the best observed among the evaluated methods, against 92.5/92.192.5/92.1 for PFD, 92.4/91.492.4/91.4 for rollout-based LingBot-VA, 91.9/91.691.9/91.6 for Fast-WAM, and 91.0/91.191.0/91.1 for rollout-based Fast-WAM-Joint. We observe the same performance recovery on RoboTwin 2.0 while retaining the one-pass deployment path. Table 2: RoboTwin 2.0 closed-loop success (%) on clean and domain-randomized scenes. Values aggregate one checkpoint per method; This report follows single-seed evaluation protocol. Emb. PT.: embodied pretraining. Method Emb. Fut. Roll- Clean Rand. AVG PT. read out LingBot-VA ✓ ✓ ✓ 92.4 91.4 91.9 Fast-WAM-Joint × ✓ ✓ 91.0 91.1 91.0 Fast-WAM × × × 91.9 91.6 91.8 PFD × × × 92.5 92.1 92.3 Rift (ours) × ✓ × 92.9 92.6 92.8 5.3 Ablations Table 3: LIBERO per-suite SR (%). Mean ± std over three seeds (500500 trials/suite/seed; 2,0002,000 overall). Bold: column bests. Method Spatial Object Goal Long Overall Fast-WAM 97.0±0.4397.0 \,± 0.43 99.7±0.0999.7 \,± 0.09 96.7±0.5096.7 \,± 0.50 93.7±0.5793.7 \,± 0.57 96.8±0.2796.8 \,± 0.27 PFD 98.3±0.1998.3 \,± 0.19 99.2±0.3399.2 \,± 0.33 98.1±0.6298.1 \,± 0.62 93.7±0.4193.7 \,± 0.41 97.3±0.1297.3 \,± 0.12 Fast-WAM-Joint 99.3±0.4199.3 \,± 0.41 99.5±0.2599.5 \,± 0.25 98.1±0.6298.1 \,± 0.62 96.6±0.5796.6 \,± 0.57 98.4±0.2698.4 \,± 0.26 Fast-WAM-IDM 99.3±0.2399.3 \,± 0.23 99.3±0.3199.3 \,± 0.31 98.1±1.0398.1 \,± 1.03 97.5±0.9997.5 \,± 0.99 98.6±0.3498.6 \,± 0.34 Cosmos Policy 97.8±0.4397.8 \,± 0.43 100.0±0.00100.0 \,± 0.00 97.8±0.1697.8 \,± 0.16 97.4±0.0097.4 \,± 0.00 98.3±0.1198.3 \,± 0.11 LingBot-VA 98.5±0.1498.5 \,± 0.14 99.6±0.1899.6 \,± 0.18 97.2±0.3897.2 \,± 0.38 98.5±0.1798.5 \,± 0.17 98.5±0.0898.5 \,± 0.08 Rift-L2 98.6±0.4398.6 \,± 0.43 99.7±0.0999.7 \,± 0.09 98.1±0.3498.1 \,± 0.34 97.2±0.5097.2 \,± 0.50 98.4±0.1298.4 \,± 0.12 Rift (full) 98.7±0.2598.7 \,± 0.25 99.7±0.0999.7 \,± 0.09 98.8±0.4398.8 \,± 0.43 98.0±0.2898.0 \,± 0.28 98.8±0.1798.8 \,± 0.17 Anticipation-token supervision. The base recipe Rift-L2 regresses future latents with a direct L2 loss and reaches 98.37±0.1298.37 \,± 0.12; the conditional-FM recipe reaches 98.8±0.1798.8 \,± 0.17. Both use the same one-pass graph and 247.9247.9 ms cost, isolating supervision without deployment overhead. The 0.40.4 point difference approaches the evaluation’s resolution, and the full recipe includes its conditioning curriculum. Table 3 gives per-suite rates; we report FM as the recipe. The number of anticipation tokens. Figure 4 sweeps m=2m=2 to full alignment (m=196m=196); current-only Fast-WAM is the no-cache m=0m=0 reference (96.75%96.75\%). Rift-L2 rises from 97.08%97.08\% to 98.37%98.37\%; conditional FM exceeds it from m=4m=4 and peaks at 98.78%98.78\%. Even small interfaces beat the reference, and full alignment is best for both. Figure 4: Anticipation-interface capacity on LIBERO. Mean success over four suites (three seeds) as token count m grows for L2 and conditional-FM supervision. The dashed line is Fast-WAM without anticipation; full alignment (m=196m=196) gives both recipes their best mean. 5.4 Qualitative results Matched imagined futures. Figure 5 compares Fast-WAM-Joint rollout and Rift one-pass decodes from matched starts on both benchmarks. Both evolve similarly at frames 0, 4, and 8. These visuals diagnose future representations, not cache equivalence; appendix C shows all frames. Figure 5: Matched imagined futures. From matched initial states, Fast-WAM-Joint iterative rollout (top) and Rift one-pass anticipation decodes (bottom) show similar evolution at frames 0, 4, and 8 on LIBERO (upper block) and RoboTwin 2.0 (lower block). Decodes are diagnostic. L2–FM uncertainty warning. The stopped-gradient L2 probe and conditional-FM head yield controller-independent future estimates whose normalized disagreement defines a CUSUM warning (appendix D). Calibrated on 1,9671,967 successful episodes, the mean CUSUM over 3333 failed rollouts crosses η 210 steps before the common t=420t=420 endpoint (fig. 6). Figure 6: L2–FM uncertainty warning over LIBERO. Here η is the CUSUM alarm threshold, conformally calibrated on all successful tasks; failures are excluded from calibration. Across the 33 failed rollouts, the detector raises an alarm an average of 210 steps before failure. 6 Conclusion World action models combine a representation read by the action expert with the iterative rollout that produces it. Intervening on the future K/V interface shows that actions require values bound to token positions, while one fixed final-clean K/V cache nearly reproduces Original execution within 1.91.9 cm E-ADE. Because this cache remains rollout-produced, the intervention establishes consumption-side sufficiency rather than rollout-free production. Rift addresses the remaining production problem with one anticipation-token prefill, preserving the complete future K/V interface while removing video denoising and VAE decoding. It clears the current-only gap on LIBERO at 1.1×1.1× baseline latency and reaches the best observed RoboTwin 2.0 success in both evaluation settings. References Alain and Bengio (2016) G. Alain and Y. Bengio Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644. Cited by: §2. Belinkov (2022) Y. Belinkov Probing classifiers: promises, shortcomings, and advances. Computational Linguistics 48 (1), p. 207–219. Cited by: §2. Bharadhwaj et al. (2024) H. Bharadhwaj, D. Dwibedi, A. Gupta, S. Tulsiani, C. Doersch, T. Xiao, D. Shah, F. Xia, D. Sadigh, and S. Kirmani Gen2Act: human video generation in novel scenarios enables generalizable robot manipulation. arXiv preprint arXiv:2409.16283. Cited by: §2. Bi et al. (2025) H. Bi, H. Tan, S. Xie, Z. Wang, S. Huang, H. Liu, R. Zhao, Y. Feng, C. Xiang, Y. Rong, H. Zhao, H. Liu, Z. Su, L. Ma, H. Su, and J. Zhu Motus: a unified latent action world model. arXiv preprint arXiv:2512.13030. External Links: Link Cited by: §1, §2. Bjorck et al. (2025) J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al. GR00T N1: an open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734. Cited by: §2. Black et al. (2024a) K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. π0 _0: a vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164. Cited by: §2. Black et al. (2024b) K. Black, M. Nakamoto, P. Atreya, H. Walke, C. Finn, A. Kumar, and S. Levine Zero-shot robotic manipulation with pre-trained image-editing diffusion models. In International Conference on Learning Representations (ICLR), External Links: Link Cited by: §2. Cen et al. (2025) J. Cen, S. Huang, Y. Yuan, K. Li, H. Yuan, C. Yu, Y. Jiang, J. Guo, X. Li, H. Luo, F. Wang, F. Wang, and D. Zhao RynnVLA-002: a unified vision-language-action and world model. arXiv preprint arXiv:2511.17502. Cited by: §2. Cheang et al. (2024) C. Cheang, G. Chen, Y. Jing, T. Kong, H. Li, Y. Li, Y. Liu, H. Wu, J. Xu, Y. Yang, H. Zhang, and M. Zhu GR-2: a generative video-language-action model with web-scale knowledge for robot manipulation. arXiv preprint arXiv:2410.06158. External Links: Link Cited by: §2. Chen et al. (2019) D. Chen, B. Zhou, V. Koltun, and P. Krähenbühl Learning by cheating. In Conference on Robot Learning (CoRL), External Links: Link Cited by: §2. Chen et al. (2025) T. Chen, Z. Chen, B. Chen, Z. Cai, Y. Liu, Z. Li, Q. Liang, X. Lin, Y. Ge, Z. Gu, et al. RoboTwin 2.0: a scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation. arXiv preprint arXiv:2506.18088. Cited by: §5.1. Chi et al. (2023) C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song Diffusion policy: visuomotor policy learning via action diffusion. In Robotics: Science and Systems (RSS), Cited by: §2. Du et al. (2023) Y. Du, M. Yang, B. Dai, H. Dai, O. Nachum, J. B. Tenenbaum, D. Schuurmans, and P. Abbeel Learning universal policies via text-guided video generation. In Advances in Neural Information Processing Systems (NeurIPS), External Links: Link Cited by: §2. Fang et al. (2026) P. Fang, H. Chen, and X. Cai Privileged foresight distillation: zero-cost future correction for world action models. arXiv preprint arXiv:2604.25859. Cited by: §1, §2, §5.1. Feng et al. (2025) Y. Feng, H. Tan, X. Mao, C. Xiang, G. Liu, S. Huang, H. Su, and J. Zhu Vidar: embodied video diffusion model for generalist manipulation. arXiv preprint arXiv:2507.12898. External Links: Link Cited by: §2. Gemini Robotics Team et al. (2025) Gemini Robotics Team, S. Abeyruwan, J. Ainslie, J. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, et al. Gemini robotics: bringing AI into the physical world. arXiv preprint arXiv:2503.20020. Cited by: §2. Hafner et al. (2023) D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104. External Links: Link Cited by: §2. Hansen et al. (2024) N. Hansen, H. Su, and X. Wang TD-MPC2: scalable, robust world models for continuous control. In International Conference on Learning Representations (ICLR), External Links: Link Cited by: §2. Hu et al. (2025) Y. Hu, Y. Guo, P. Wang, X. Chen, Y. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen Video prediction policy: a generalist robot policy with predictive visual representations. In International Conference on Machine Learning (ICML), External Links: Link Cited by: §2. Jain and Wallace (2019) S. Jain and B. C. Wallace Attention is not explanation. In Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), Cited by: §2. Jang et al. (2025) J. Jang, S. Ye, Z. Lin, J. Xiang, J. Bjorck, Y. Fang, F. Hu, S. Huang, K. Kundalia, Y. Lin, et al. DreamGen: unlocking generalization in robot learning through video world models. arXiv preprint arXiv:2505.12705. External Links: Link Cited by: §2. Kim et al. (2026) M. J. Kim, Y. Gao, T. Lin, Y. Lin, Y. Ge, G. Lam, P. Liang, S. Song, M. Liu, C. Finn, and J. Gu Cosmos policy: fine-tuning video models for visuomotor control and planning. arXiv preprint arXiv:2601.16163. Cited by: §2. Kim et al. (2024) M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn OpenVLA: an open-source vision-language-action model. In Conference on Robot Learning (CoRL), Cited by: §2. Li et al. (2026) L. Li, Q. Zhang, Y. Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, Y. Shen, and Y. Xu Causal world modeling for robot control. arXiv preprint arXiv:2601.21998. Cited by: §1, §2, §5.1. Li et al. (2025) S. Li, Y. Gao, D. Sadigh, and S. Song Unified video action model. arXiv preprint arXiv:2503.00200. Cited by: §2. Liang et al. (2025) J. Liang, P. Tokmakov, R. Liu, S. Sudhakar, P. Shah, R. Ambrus, and C. Vondrick Video generators are robot policies. arXiv preprint arXiv:2508.00795. Cited by: §2. Liao et al. (2025) Y. Liao, P. Zhou, S. Huang, D. Yang, S. Chen, Y. Jiang, Y. Hu, J. Cai, S. Liu, J. Luo, L. Chen, S. Yan, M. Yao, and G. Ren Genie envisioner: a unified world foundation platform for robotic manipulation. arXiv preprint arXiv:2508.05635. Cited by: §2. Lipman et al. (2023) Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le Flow matching for generative modeling. In International Conference on Learning Representations (ICLR), External Links: Link Cited by: §4.3. Liu et al. (2023) B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone LIBERO: benchmarking knowledge transfer for lifelong robot learning. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §5.1. Liu et al. (2024) S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu RDT-1B: a diffusion foundation model for bimanual manipulation. arXiv preprint arXiv:2410.07864. Cited by: §2. Luo et al. (2026) H. Luo, W. Zhang, Y. Feng, S. Zheng, H. Xu, C. Xu, Z. Xi, Y. Fu, and Z. Lu Being-H0.7: a latent world-action model from egocentric videos. arXiv preprint arXiv:2605.00078. External Links: Link Cited by: §2. Meng et al. (2022) K. Meng, D. Bau, A. Andonian, and Y. Belinkov Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems, Vol. 35, p. 17359–17372. External Links: Document, Link Cited by: §2. Pai et al. (2025) J. Pai, L. Achenbach, V. Montesinos, B. Forrai, O. Mees, and E. Nava Mimic-video: video-action models for generalizable robot control beyond VLAs. arXiv preprint arXiv:2512.15692. Cited by: §2. Physical Intelligence et al. (2025) Physical Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al. π0.5 _0.5: a vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054. Cited by: §2. Rao et al. (2026) Z. Rao, Y. Zhao, W. Guo, B. Fei, Y. Guo, and H. Xiong The geometry of flow-matching uncertainty: a cost-free uncertainty proxy and its application in flow-based VLA failure detection. arXiv preprint arXiv:2607.27933. External Links: Document, Link Cited by: Appendix D, §2. Schrittwieser et al. (2020) J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. Lillicrap, and D. Silver Mastering atari, go, chess and shogi by planning with a learned model. Nature 588 (7839), p. 604–609. External Links: Document Cited by: §2. Serrano and Smith (2019) S. Serrano and N. A. Smith Is attention interpretable?. In Annual Meeting of the Association for Computational Linguistics (ACL), Cited by: §2. Shukor et al. (2025) M. Shukor, D. Aubakirova, F. Capuano, P. Kooijmans, S. Palma, A. Zouitine, M. Aractingi, C. Pascal, M. Russi, A. Marafioti, et al. SmolVLA: a vision-language-action model for affordable and efficient robotics. arXiv preprint arXiv:2506.01844. Cited by: §2. Team Wan et al. (2025) Team Wan, A. Wang, B. Ai, B. Wen, C. Mao, et al. Wan: open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314. External Links: Link Cited by: Appendix A, §5.1. Vapnik and Vashist (2009) V. Vapnik and A. Vashist A new learning paradigm: learning using privileged information. Neural Networks 22 (5). External Links: Document Cited by: §2. Wen et al. (2025) J. Wen, Y. Zhu, J. Li, Z. Tang, C. Shen, and F. Feng DexVLA: vision-language model with plug-in diffusion expert for general robot control. arXiv preprint arXiv:2502.05855. Cited by: §2. Won et al. (2025) J. Won, K. Lee, H. Jang, D. Kim, and J. Shin Dual-stream diffusion for world-model augmented vision-language-action model. arXiv preprint arXiv:2510.27607. External Links: Link Cited by: §2. Wu et al. (2024) H. Wu, Y. Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong Unleashing large-scale video generative pre-training for visual robot manipulation. In International Conference on Learning Representations (ICLR), External Links: Link Cited by: §2. Ye et al. (2026) S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, et al. World action models are zero-shot policies. arXiv preprint arXiv:2602.15922. External Links: Link Cited by: §2. Yuan et al. (2026) T. Yuan, Z. Dong, Y. Liu, and H. Zhao Fast-WAM: do world action models need test-time future imagination?. arXiv preprint arXiv:2603.16666. External Links: Link Cited by: Appendix A, Table 5, §1, §1, §2, §2, §4.3, §4.3, §5.1. Zhang et al. (2026) C. Zhang, R. Lu, J. Tong, X. Li, Y. Wang, and H. Li EvoScene-VLA: evolving scene beliefs inside the action decoder for chunked robot control. arXiv preprint arXiv:2605.21862. External Links: Document, Link Cited by: §2. Zhang et al. (2025) W. Zhang, H. Liu, Z. Qi, Y. Wang, X. Yu, J. Zhang, R. Dong, J. He, H. Wang, Z. Zhang, L. Yi, W. Zeng, and X. Jin DreamVLA: a vision-language-action model dreamed with comprehensive world knowledge. CoRR abs/2507.04447. External Links: Document Cited by: §2. Zhao et al. (2025) Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn, A. Handa, M. Liu, D. Xiang, G. Wetzstein, and T. Lin CoT-VLA: visual chain-of-thought reasoning for vision-language-action models. arXiv preprint arXiv:2503.22020. External Links: Link Cited by: §2. Zheng et al. (2025) R. Zheng, J. Wang, S. Reed, J. Bjorck, Y. Fang, F. Hu, J. Jang, K. Kundalia, Z. Lin, L. Magne, et al. FLARE: robot learning with implicit world modeling. arXiv preprint arXiv:2505.15659. External Links: Link Cited by: §2. Zhou et al. (2024) S. Zhou, Y. Du, J. Chen, Y. Li, D. Yeung, and C. Gan RoboDreamer: learning compositional world models for robot imagination. arXiv preprint arXiv:2404.12377. Cited by: §2. Zhu et al. (2025) C. Zhu, R. Yu, S. Feng, B. Burchfiel, P. Shah, and A. Gupta Unified world models: coupling video and action diffusion for pretraining on large robotic datasets. arXiv preprint arXiv:2504.02792. External Links: Link Cited by: §2. Zitkovich et al. (2023) B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. RT-2: vision-language-action models transfer web knowledge to robotic control. In Conference on Robot Learning (CoRL), Cited by: §2. Appendix A Training implementation details Loss normalization. For ℒvidL_vid, unreduced squared error is averaged over channel and spatial dimensions and then over valid latent frames. For ℒFML_FM, it is averaged over each patch vector and then over valid future patches. For each sample, both losses receive the native video-flow timestep weight before the batch mean. For ℒactL_act, squared error is first averaged over action dimensions, then masked for padding and averaged over the action horizon. The action-timestep weight from the base objective is applied before the batch mean. The coefficients of ℒvidL_vid and ℒactL_act are 11. Both λFM _FM and λprobe _probe remain at 11 for the first 70%70\% of the configured curriculum horizon and follow a cosine decay to 0.20.2 over the remaining 30%30\%. Because the probe receives a stopped-gradient input, its weight cannot affect the deployed representation. Deployment-matched perturbation. After 70%70\% of the configured curriculum horizon, the perturbation probability rises linearly from 00 to 0.30.3 and its standard deviation from 00 to 0.060.06 times the latent standard deviation. Perturbed rows retain ℒFML_FM and ℒprobeL_probe but are excluded from ℒactL_act. Diagnostic probe. An RMS-normalized linear probe gives Y^L2=gω(RMS(stopgrad(Sϕ))) Y_L2=g_ω(RMS(stopgrad(S_φ))). We use the elementwise mean squared error ℒprobe=mean((Y^L2−Y)2)L_probe=mean\! (( Y_L2-Y)^2 ), where the mean is over all prediction elements in the batch. Its gradients reach only gωg_ω. Gradient routes. ℒvidL_vid updates the video expert and its output head, but not the action expert. ℒactL_act updates the action expert and backpropagates through the attended video states into the video expert and anticipation tokens. ℒFML_FM updates the video expert, anticipation tokens, and FM head; detached ℒprobeL_probe updates only the linear probe. Architecture and training configuration. All in-house models share the pretrained Wan2.2-5B backbone (39): its video DiT, text encoder, and video VAE. The action expert reuses the video-branch architecture with reduced hidden dimension da=1024d_a=1024, giving a 11B action expert and a 66B total model. The action horizon is H=32H=32. Images from multiple cameras are concatenated into a single image before the VAE, and video is temporally downsampled by 4×4× to 99 frames per chunk. Both branches use Fast-WAM’s continuous flow-matching formulation. During training, we sample u∼[0,1)u [0,1), set σ=5u/(1+4u)σ=5u/(1+4u), and supply t=1000σt=1000σ to the model. The FM head reuses the video scheduler. At deployment, the action is sampled with 1010 flow-matching steps and classifier-free guidance scale 1.01.0, with no video denoising and no VAE decoding. We train with AdamW at learning rate 1×10−41× 10^-4, weight decay 0.010.01, cosine annealing, mixed precision, and gradient clipping at 1.01.0. LIBERO models train for 2020k steps and RoboTwin 2.0 models for 3030k steps. Relative to the Fast-WAM backbone (45), Rift preserves these settings, adding only anticipation tokens, ℒFML_FM, and a diagnostic probe. Appendix B Uncertainty for the intervention study Figure 2 reports point estimates; table 4 gives E-ADE intervals and the task-level resampling protocol. Table 4: Uncertainty for the E-ADE estimates in fig. 2. The table reports percentile task-cluster bootstrap 95%95\% confidence intervals in centimeters. Each replicate resamples whole tasks with replacement and recomputes the task-macro mean, keeping all trials from a sampled task together. Original is omitted because its E-ADE is zero by definition. Final-clean replay substitutes both final-step keys and values at every action-denoising step, as in fig. 2. All reported intervention estimates use 2,0002,000 paired trials across all 4040 LIBERO tasks. N/A entries are structural rather than missing runs: IDM-style models already consume one fixed final-clean future K/V cache, so final-clean replay coincides with their Original, and Cosmos-2 exposes only one future timestep and therefore admits no temporal swap. Model Intervention E-ADE 95%95\% CI Fast-WAM-Joint Final clean K/V 1.908 [1.579, 2.260][1.579,\ 2.260] Spatial shuffle 14.306 [12.649, 16.057][12.649,\ 16.057] Temporal swap 15.606 [13.973, 17.189][13.973,\ 17.189] Noise 17.221 [16.017, 18.405][16.017,\ 18.405] Frozen present 18.255 [16.876, 19.655][16.876,\ 19.655] Masked future 18.725 [17.499, 19.953][17.499,\ 19.953] Fast-WAM-IDM Final clean K/V N/A N/A Spatial shuffle 15.293 [13.214, 17.418][13.214,\ 17.418] Temporal swap 16.325 [15.068, 17.520][15.068,\ 17.520] Noise 17.939 [17.070, 18.779][17.070,\ 18.779] Frozen present 21.028 [19.901, 22.104][19.901,\ 22.104] Masked future 20.378 [19.389, 21.370][19.389,\ 21.370] Cosmos-2 Final clean K/V 1.690 [0.260, 3.620][0.260,\ 3.620] Spatial shuffle 4.976 [4.342, 5.643][4.342,\ 5.643] Temporal swap N/A N/A Noise 10.494 [8.984, 12.034][8.984,\ 12.034] Frozen present 18.090 [16.800, 19.323][16.800,\ 19.323] Masked future 11.788 [10.516, 13.034][10.516,\ 13.034] LingBot-VA Final clean K/V N/A N/A Spatial shuffle 19.820 [16.171, 22.550][16.171,\ 22.550] Temporal swap 15.706 [12.171, 20.044][12.171,\ 20.044] Noise 17.466 [16.099, 18.831][16.099,\ 18.831] Frozen present 21.323 [17.515, 25.022][17.515,\ 25.022] Masked future 19.176 [17.066, 21.131][17.066,\ 21.131] Appendix C Comparison of imagined futures Matched iterative Fast-WAM-Joint and one-pass Rift decodes on both benchmarks appear in figs. 7 and 8. These are diagnostic; section 3.5 provides the closed-loop evidence for complete-K/V cache sufficiency. Figure 7: Imagined future: Fast-WAM-Joint iterative diffusion versus Rift one pass. Per state, the top strip is Fast-WAM-Joint’s future decoded from its full iterative diffusion rollout; the bottom strip is decoded from Rift’s anticipation tokens after one backbone pass. The first column is the shared current observation at frame 0; the remaining columns show decoded frames 2, 4, 6, and 8. Visual similarity is illustrative; policy behavior is evaluated separately. Figure 8: RoboTwin 2.0 imagined future: Fast-WAM-Joint iterative diffusion versus Rift one pass. Each row pair starts from the same raw observation from three cameras. The first column is the shared observation at frame 0; the remaining columns show decoded frames 2, 4, 6, and 8. The top strip uses Fast-WAM-Joint’s iterative video diffusion with 20 denoising steps. The bottom strip shows Rift’s anticipation states from one backbone pass, decoded by the trained linear diagnostic probe and frozen VAE. The Rift video decode is diagnostic and is not part of its deployed action path; policy behavior is evaluated separately. Appendix D Optional L2–FM uncertainty warning The final recipe retains direct L2’s deterministic future-latent readout as a stopped-gradient diagnostic probe. Both heads read the same anticipation states but do not feed the controller. This controller-independent comparison runs only in optional monitor-only mode and is excluded from policy-only deployment and reported latency. Let μL2=Y^L2∈ℝm×dy _L2= Y_L2 ^m× d_y denote the probe estimate, let xiFM∈ℝm×dyx_i^FM ^m× d_y be the iith of K samples from the conditional-FM head, and let x¯FM=K−1∑ixiFM x_FM=K^-1 _ix_i^FM denote their sample mean at the current environment step. We measure their normalized cross-estimator discrepancy as dratio=∥μL2−x¯FM∥F2K−1∑i=1K∥xiFM−x¯FM∥F2+ϵ,d_ratio= _L2- x_FM _F^2K^-1 _i=1^K x_i^FM- x_FM _F^2+ε, (7) where ϵ>0ε>0 stabilizes the denominator. Large dratiod_ratio means that the probe point estimate departs from the FM mean beyond the dispersion of the FM samples. It is a warning statistic, not a failure probability: both heads may still agree on the same wrong future. Unlike the action-flow acceleration of 35, this score compares two estimators in future-latent space. In fig. 6, a one-sided CUSUM accumulates St=max(0,St−1+dratio(t)−μd−0.25σd),S_t= \! (0,S_t-1+d_ratio(t)- _d-0.25 _d ), (8) where μd _d and σd _d are calibrated from the 1,9671,967 successes in the 2,0002,000-episode ledger, and η is the resulting conformal CUSUM alarm threshold. Figure 6 averages StS_t over all 3333 failed rollouts, which are excluded from calibration. This mean crosses η 210 steps before the common t=420t=420 endpoint; individual crossing times can differ. Appendix E Limitations and future work Evaluation is simulation-only. Physical robots and non-WAM fusion backbones remain future work. Appendix F Per-task RoboTwin 2.0 success rates Table 5 reports per-task RoboTwin 2.0 success rates. Table 5: Per-task RoboTwin 2.0 success rates (%) under clean and randomized settings. ∗ marks results from 45; bold marks each row’s best per setting. π0.5∗ _0.5^\,* Motus∗ LingBot-VA Fast-WAM-Joint Fast-WAM PFD Rift (Ours) Task Clean Rand. Clean Rand. Clean Rand. Clean Rand. Clean Rand. Clean Rand. Clean Rand. Adjust Bottle 100 99 89 93 90 94 99 98 100 100 100 100 100 100 Beat Block Hammer 96 93 95 88 95 98 100 99 100 97 100 97 100 97 Blocks Ranking RGB 92 85 99 97 98 98 100 100 100 98 100 98 100 98 Blocks Ranking Size 49 26 75 63 94 96 87 89 93 98 94 98 94 98 Click Alarmclock 98 89 100 100 98 99 99 100 100 100 100 100 100 100 Click Bell 99 66 100 100 99 99 100 97 100 100 100 100 100 100 Dump Bin Bigbin 92 97 95 91 89 96 96 94 97 96 97 96 97 97 Grab Roller 100 100 100 100 99 99 100 100 100 100 100 100 100 100 Handover Block 66 57 86 73 98 78 92 93 97 81 97 82 97 83 Handover Mic 98 97 78 63 94 96 100 100 99 100 99 100 99 100 Hanging Mug 18 17 38 38 40 28 63 68 60 62 63 65 65 67 Lift Pot 96 85 96 99 99 98 100 100 100 100 100 100 100 100 Move Can Pot 51 55 34 74 94 97 98 98 90 91 91 92 91 92 Move Pillbottle Pad 84 61 93 96 98 99 100 99 100 99 100 99 100 99 Move Playingcard Away 96 84 100 96 99 99 100 100 100 100 100 100 100 100 Move Stapler Pad 56 42 83 85 91 79 83 84 72 70 74 72 75 74 Open Laptop 90 96 95 91 92 94 91 90 99 100 99 100 99 100 Open Microwave 34 77 95 91 82 86 6 25 59 37 62 41 64 45 Pick Diverse Bottles 81 71 90 91 89 82 89 85 81 88 82 89 83 89 Pick Dual Bottles 93 63 96 90 99 99 97 100 100 96 100 96 100 97 Place A2B Left 87 82 88 79 96 93 97 95 94 95 94 95 95 96 Place A2B Right 87 84 91 87 96 95 94 96 94 96 94 96 95 97 Place Bread Basket 77 64 91 94 96 95 91 92 93 93 94 94 94 94 Place Bread Skillet 85 66 86 83 95 90 88 95 95 95 95 95 96 96 Place Burger Fries 94 87 98 98 96 95 100 99 94 98 94 98 95 98 Place Can Basket 62 62 81 76 81 84 53 35 68 63 70 65 72 68 Place Cans Plasticbox 94 84 98 94 99 99 99 97 100 97 100 97 100 97 Place Container Plate 99 95 98 99 98 97 98 100 97 100 97 100 97 100 Place Dual Shoes 75 75 93 87 94 89 95 87 94 89 94 90 95 90 Place Empty Cup 100 99 99 98 99 99 100 100 100 100 100 100 100 100 Place Fan 87 85 91 87 98 93 98 97 95 95 95 95 96 96 Place Mouse Pad 60 39 66 68 93 96 94 92 88 89 89 90 89 90 Place Object Basket 80 76 81 87 91 88 88 80 86 84 87 85 88 86 Place Object Scale 86 80 88 85 96 95 95 100 93 93 94 94 94 94 Place Object Stand 91 85 98 97 98 96 94 96 89 93 90 94 90 94 Place Phone Stand 81 81 87 86 96 97 100 100 99 99 99 99 99 99 Place Shoe 92 93 99 97 97 98 94 98 96 97 96 97 96 97 Press Stapler 87 83 93 98 85 82 52 58 94 96 94 96 95 97 Put Bottles Dustbin 84 79 81 79 87 91 95 93 91 88 92 89 92 89 Put Object Cabinet 80 79 88 71 85 87 93 91 93 90 94 91 94 91 Rotate QRcode 89 87 89 73 96 91 90 94 91 90 92 91 92 91 Scan Object 72 65 67 66 96 91 93 91 90 91 91 92 91 92 Shake Bottle 99 97 100 97 99 97 100 100 100 100 100 100 100 100 Shake Bottle Horizontally 99 99 100 98 99 99 99 100 100 100 100 100 100 100 Stack Blocks Three 91 76 91 95 98 98 99 96 95 96 95 96 96 97 Stack Blocks Two 97 100 100 98 99 98 100 100 100 100 100 100 100 100 Stack Bowls Three 77 71 79 87 86 83 87 84 78 81 80 82 81 83 Stack Bowls Two 95 96 98 98 94 98 96 97 93 94 94 94 94 95 Stamp Seal 79 55 93 92 96 97 97 98 89 97 90 97 90 97 Turn Switch 62 54 84 78 44 45 71 75 60 66 63 68 65 70 Average 82.7 76.8 88.7 87.0 92.4 91.4 91.0 91.1 91.9 91.6 92.5 92.1 92.9 92.6