Paper deep dive
DriveCache: Action-Aware Caching for Driving World Model Inference
Jianchun Yang, Jian Liang, Xianda Guo, Pinhan Fu, Yanlun Peng, Conglang Zhang, Wenke Huang, Mang Ye
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/23/2026, 1:53:38 AM
Summary
DriveCache is a training-free, action-aware caching controller designed to accelerate diffusion-based driving video generation models. It utilizes planned ego motion (translation and rotation) available before denoising to allocate reuse budgets across scenes and employs exact dynamic programming to place these reuses across denoising steps under a calibrated response budget. A causal drift check refreshes features and replans schedules when generation deviates from calibration. DriveCache improves the fidelity-efficiency trade-off over existing cache methods, achieving approximately 2x speedup on Wan2.2 A14B with improved PSNR.
Entities (10)
Relation Signals (7)
DriveCache → uses → Causal Drift Check
confidence 95% · A causal drift check refreshes features and replans the remaining schedule when generation departs from calibration.
DriveCache → uses → Dynamic Programming
confidence 95% · DriveCache ... uses exact dynamic programming to place it across denoising steps under a calibrated response budget.
DriveCache → uses → Planned Ego Motion
confidence 95% · DriveCache uses planned ego motion before denoising and uses runtime drift to veto unsupported reuse decisions.
DriveCache → evaluatedon → DrivingGen
confidence 90% · We evaluate Wan2.2 5B and expert-routed A14B on 222 DrivingGen scenes
Epona → evaluatedon → nuScenes
confidence 90% · Epona on all 150 nuScenes validation scenes
DriveCache → improves → Wan2.2 A14B
confidence 90% · At approximately 2× speedup on Wan2.2 A14B, it improves PSNR by 2.036 dB over TeaCache.
DriveCache → outperforms → TeaCache
confidence 90% · At approximately 2× speedup on Wan2.2 A14B, it improves PSNR by 2.036 dB over TeaCache.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Diffusion-based driving generators repeatedly evaluate large backbones across denoising steps, which limits generation throughput. Existing diffusion acceleration methods reduce this cost, but general-purpose designs omit driving signals available before generation, such as ego speed and planned trajectories. Experiments across driving motions show that cache tolerance varies with ego translation and rotation, denoising progress, and consecutive reuse length. We propose DriveCache, a training-free, action-aware controller that uses planned motion to allocate reuse across scenes and dynamic programming to place it across denoising steps under a calibrated response budget. A causal drift check refreshes features and replans the remaining schedule when generation departs from calibration. Across three generator configurations, DriveCache improves the overall fidelity-efficiency trade-off over evaluated cache methods. Our code will be publicly available.
Tags
Links
- Source: https://arxiv.org/abs/2608.16354v1
- Canonical: https://arxiv.org/abs/2608.16354v1
Trouble viewing inline? Open PDF directly →
Full Text
44,122 characters extracted from source content.
Expand or collapse full text
DriveCache: Action-Aware Caching for Driving World Model Inference Jianchun Yang Jian Liang Xianda Guo Pinhan Fu Yanlun Peng Conglang Zhang Wenke Huang Mang Ye Abstract Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Diffusion-based driving generators repeatedly evaluate large backbones across denoising steps, which limits generation throughput. Existing diffusion acceleration methods reduce this cost, but general-purpose designs omit driving signals available before generation, such as ego speed and planned trajectories. Experiments across driving motions show that cache tolerance varies with ego translation and rotation, denoising progress, and consecutive reuse length. We propose DriveCache, a training-free, action-aware controller that uses planned motion to allocate reuse across scenes and dynamic programming to place it across denoising steps under a calibrated response budget. A causal drift check refreshes features and replans the remaining schedule when generation departs from calibration. Across three generator configurations, DriveCache improves the overall fidelity–efficiency trade-off over evaluated cache methods. Our code will be publicly available. Introduction Driving video generation models predict how traffic scenes evolve under different ego actions, providing controllable future observations for simulation, policy training, planning evaluation, and offline data generation (19; 50; 13). These action-conditioned predictions let developers and planning systems examine possible futures under candidate plans. They also support diverse scenario generation for development. Recent video diffusion models improve visual fidelity and temporal coherence through larger spatiotemporal backbones and longer generation horizons (18; 1; 56; 24; 25). These advances increase inference cost. Iterative denoising evaluates the backbone across many sampling steps, and the cost grows with model capacity, resolution, and video length. Online prediction must fit within the planning cycle, while offline generation must scale to large scenario collections. Inference latency therefore constrains online prediction and offline scenario generation. Diffusion acceleration methods reduce the number or cost of denoising evaluations. Distillation, quantization, pruning, and efficient operators can require new weights, retraining, calibration data, or specialized kernels (43; 47). Feature caching leaves the generator unchanged and reuses intermediate features across neighboring denoising steps. Existing cache controllers derive reuse from fixed schedules or signals observed after denoising begins (27; 22; 68; 37; 3; 28; 9). These general-purpose methods omit driving signals available before generation, including ego speed and planned trajectories. Figure 1: DriveCache achieves better video fidelity and consistency across evaluation metrics. Figure 2: Motivation of DriveCache. Planned ego motion reveals scene-dependent cache tolerance before denoising. However, driving generation provides planned ego motion before denoising. Offline generation receives recorded or specified future ego poses, while planner-conditioned generation receives predicted poses from the planning stack. Stationary, straight, and turning plans induce different viewpoint changes. Figure 2 shows that planned ego motion predicts scene-dependent cache tolerance before the denoising pass begins. Planned motion is available before the first denoising step, but it cannot determine a schedule by itself. Cache error also varies with denoising position, cache age, model architecture, and generated content. A driving-aware controller must therefore combine a scene-level motion prior with a denoising-level response model, while retaining a causal correction for states that depart from calibration. DriveCache converts this prior into a complete schedule. One low-motion anchor and one moving-turn anchor measure terminal responses of consecutive reuse; planned translation and rotation interpolate them for each scene. Exact dynamic programming selects reuse quantity and positions under one budget. A pre-reuse drift veto rejects out-of-support decisions and replans the unexecuted suffix. Our main contributions are threefold. • To the best of our knowledge, this work is the first to identify and validate planned ego motion as a pre-generation signal for diffusion caching in driving video generation, showing that ego translation and rotation predict scene-level cache tolerance across driving motions. • We propose DriveCache, a training-free cache controller that assigns scene-level reuse budgets from planned motion, models consecutive-reuse error, and uses exact dynamic programming with causal drift correction to place reuse across denoising steps. • Across multiple driving video generators, DriveCache improves quality–efficiency trade-offs over cache baselines. At approximately 2×2× speedup on Wan2.2 A14B, it improves PSNR by 2.036 dB over TeaCache. Related Work Diffusion Models Diffusion models learn a reverse transport from noise to data (17; 48; 42). Deterministic samplers and continuous-time formulations broaden this process (45; 23; 29; 26). Video diffusion adds spatiotemporal modeling (18; 1; 33; 56; 24; 25), while Diffusion Transformers scale generation (40; 5). Token merging and low-precision attention lower per-step cost (2; 59). Training-based methods distill samplers or generation dynamics but require optimized weights and model-specific training (43; 47; 31; 57; 58). These acceleration methods change sampling, reduce token or attention cost, or train modified generator weights. Training-free Diffusion Inference Acceleration Training-free acceleration preserves pretrained weights. Fast solvers reduce model evaluations (45; 29). Cache controllers retain pretrained generator weights and reuse internal computation. Static routing uses fixed schedules in DeepCache, PAB, and FORA and a learned schedule in Learning-to-Cache (35; 65; 44; 34). TeaCache, AdaCache, and EasyCache adapt reuse to runtime changes (27; 22; 68). Recent DiT controllers allocate reuse across blocks, tokens, or trajectories (6; 70; 71; 8; 41). FasterCache reuses residuals, TaylorSeer forecasts features, and FlowCache supports autoregressive video (32; 28; 36). MagCache, DiCache, and SeaCache exploit magnitude, reconstruction, and spectral signals (37; 3; 9). DriveCache uses planned ego motion before denoising and uses runtime drift to veto unsupported reuse decisions. Autonomous Driving Video Generation Driving systems study occupancy, sensor fusion, planning, safety-critical generation, and scene representations (66; 7; 67; 55; 54; 46; 10). Surround-view studies characterize cross-view depth and spatial reasoning (14; 16; 15). Driving video generators forecast observations from histories, maps, layouts, and ego trajectories (19; 50; 13; 12; 52; 64; 20). Multiview reconstruction models target geometric consistency and long-horizon control (51; 30; 11; 38; 53), while physical-AI platforms build world foundation models (39). Autoregressive diffusion links video to trajectories (60; 69); benchmarks evaluate reactive simulation and deployment robustness (63; 62). NuScenes provides synchronized cameras and ego-motion records (4). DriveCache uses planned ego motion to estimate viewpoint change and allocate pre-denoising computation. Methodology DriveCache formulates caching as a causal decision made before each backbone evaluation. Consider a frozen video diffusion model with K denoising steps. At step k, Equation (1) separates the expensive reusable backbone from the inexpensive output interface: rk=Fk(uk),ϵk=Hk(xk,rk),r_k=F_k(u_k), _k=H_k(x_k,r_k), (1) where xkx_k is the current latent, uku_k is the backbone input, and rkr_k is the cached backbone output. A full decision evaluates FkF_k; a reuse decision replaces rkr_k with its most recently computed value. The first step always uses full computation. The controller observes uku_k before executing FkF_k, so it can make the veto decision without first executing the backbone call. DriveCache has four stages (Figure 4). Two ego-motion anchors measure joint run responses; planned ego motion interpolates them for the current scene; exact dynamic programming jointly selects reuse quantity and placement; and a pre-reuse drift check can refresh and replan the unexecuted suffix while preserving the executed prefix. The ablation study in Table 3 supports planned ego motion as a cache prior. Cache tolerance varies because planned ego motion changes the rendered viewpoint. Scene-agnostic scheduling, shuffled trajectories, and translation-matched turns separate planned-motion allocation from dataset correlation and fixed step preference. Figure 3 shows that planned ego motion separates run tolerance while local input drift remains nearly unchanged. Planned ego motion remains a scene coordinate. Denoising position and cache age determine placement cost, while support checks and the causal veto handle departures from the calibrated regime. Figure 3: Planned ego motion changes joint run responses even when local input drift remains nearly unchanged. The search enforces the mandatory first evaluation, maximum cache age, backbone boundaries, and native no-cache steps. It maximizes skipped evaluations under one cumulative response budget, with deterministic lower-cost tie breaking. Planned ego motion changes allocation before denoising, run costs place reuse jointly, and a veto corrects the next decision before a cached backbone output changes the latent. The executed prefix is never altered. Figure 4: DriveCache combines planned ego motion, two-anchor run calibration, exact DP, and a causal reuse guard. Planned ego motion calibrates run costs Let a scene provide planned ego poses (pt,ψt)(p_t, _t) over the generated horizon. Equation (2) summarizes total translation and accumulated rotation: Ls=∑t∥pt−pt−1∥2,Θs=∑t|wrap(ψt−ψt−1)|.L_s= _t p_t-p_t-1 _2, _s= _t |wrap( _t- _t-1) |. (2) The statistics remain separate because equal travel distance can induce different cached-output changes under straight and turning motion. We z-score L and Θ using their calibration means and standard deviations. The low-motion anchor (L0,Θ0)(L_0, _0) minimizes the sum of these standardized values. Among clips whose raw Θ exceeds its calibration median and whose raw L>L0L>L_0 and Θ>Θ0 > _0, the moving-turn anchor (L1,Θ1)(L_1, _1) maximizes the product of the standardized values. Clip index breaks ties, and an empty candidate set triggers full-inference fallback. Equation (3) expresses a scene relative to this support: zsL=Ls−L0max(L1−L0,ε),zsΘ=Θs−Θ0max(Θ1−Θ0,ε).z_s^L= L_s-L_0 (L_1-L_0, ), z_s = _s- _0 ( _1- _0, ). (3) The anchor pair maps low motion toward the origin and the moving turn toward the upper corner without fitting scene-specific coefficients. A scene lies inside the calibration support only when (zsL,zsΘ)∈[0,1]2(z_s^L,z_s )∈[0,1]^2; otherwise DriveCache uses full inference. Equation (4) then reduces the supported pair to one interpolation coordinate: Ds=12(zsL)2+(zsΘ)2.D_s= 1 2 (z_s^L)^2+(z_s )^2. (4) The coordinatewise gate prevents radial compression from hiding an unsupported dimension. The anchor traces also define how DriveCache scores a consecutive reuse run. For each anchor i∈0,1i∈\0,1\, one full denoising trace stores (xk(i),uk(i),rk(i))(x_k^(i),u_k^(i),r_k^(i)). A candidate run starts after a full refresh at step j and reuses rj(i)r_j^(i) for h steps. Equation (5) measures how far the current backbone input has moved from the refresh input at the ℓ -th reuse: dj,ℓ(i)=∥uj+ℓ(i)−uj(i)∥2∥uj+ℓ(i)∥2+ε.d_j, ^(i)= u_j+ ^(i)-u_j^(i) _2 u_j+ ^(i) _2+ . (5) Input drift is available before the expensive backbone call and is therefore suitable for the runtime veto. It does not by itself measure the denoising error caused by substituting a cached backbone output. Equation (6) defines that signed output perturbation against the full trace: ej,ℓ(i)=Hj+ℓ(xj+ℓ(i),rj(i))−Hj+ℓ(xj+ℓ(i),rj+ℓ(i)).e_j, ^(i)=H_j+ (x_j+ ^(i),r_j^(i))-H_j+ (x_j+ ^(i),r_j+ ^(i)). (6) Frozen-sampler Jacobian-vector products propagate the whole run to the terminal latent. Let Pt(i)P_t^(i) be the Jacobian of terminal latent xKx_K with respect to denoiser output ϵt _t, evaluated along anchor i’s frozen full trace. Its product with et(i)e_t^(i) gives the first-order terminal perturbation. Equation (7) combines these signed perturbations into joint run response qj,h(i)q_j,h^(i): Gj,h(i)=∑ℓ=1hPj+ℓ(i)ej,ℓ(i),qj,h(i)=∥Gj,h(i)∥2∥xK(i)∥2+ε.G_j,h^(i)= _ =1^hP_j+ ^(i)e_j, ^(i), q_j,h^(i)= G_j,h^(i) _2 x_K^(i) _2+ . (7) Normalizing by terminal-latent energy makes responses comparable across steps without a learned evaluator. Replay isolates cache perturbation, while the signed vector preserves cross-step cancellation and amplification. For scene s, Equation (8) interpolates only the nonnegative planned-motion-dependent difference: q~s,j,h=qj,h(0)+Dsmax(qj,h(1)−qj,h(0),0). q_s,j,h=q_j,h^(0)+D_s \! (q_j,h^(1)-q_j,h^(0),0 ). (8) This one-sided interpolation preserves the low-motion response when the moving anchor is easier and enforces a monotone planned-motion risk relation. Equation (9) converts joint run response q~ q into monotone envelope Q and incremental run cost c c: Qs,j,h Q_s,j,h =max1≤ℓ≤hq~s,j,ℓ, = _1≤ ≤ h q_s,j, , (9) c^s,j,h c_s,j,h =Qs,j,h−Qs,j,h−1,Qs,j,0=0. =Q_s,j,h-Q_s,j,h-1, Q_s,j,0=0. Planned ego motion changes how much reuse a scene can tolerate, while (j,h)(j,h) captures denoising position and cache age. We obtain drift threshold d¯s,j,h d_s,j,h with the same one-sided anchor interpolation as Equation (8), replacing q with d. The envelope keeps longer-run increments nonnegative; refresh resets age but not response already propagated into the latent. Exact scheduling supports causal correction DriveCache converts the run costs into a legal schedule under response budget τ. For target reuse fraction ρtar _tar, calibration chooses the smallest table-cost threshold whose median schedule reaches ⌈ρtar(K−1)⌉ _tar(K-1) reuses; neither PSNR nor evaluation clips enter this choice. Since legal reuses skip the same interface, maximizing their count minimizes backbone evaluations. Let Jk(n,h)J_k(n,h) be minimum cumulative response after k steps, n reuses, and cache age h. We set J1(0,0)=0J_1(0,0)=0 and all other states to +∞+∞, with 0≤n≤k−10≤ n≤ k-1 and 0≤h≤Hmax0≤ h≤ H_ . Here, ℒ(k,h)=1L(k,h)=1 marks reuse that satisfies the model-specific legality constraints: Jk+1(n,0) J_k+1(n,0) =minhJk(n,h), = _hJ_k(n,h), (10) Jk+1(n+1,h+1) J_k+1(n+1,h+1) =minJk+1(n+1,h+1), = \! \J_k+1(n+1,h+1), Jk(n,h)+c^s,k−h,h+1, J_k(n,h)+ c_s,k-h,h+1 \, ℒ(k,h+1)=1. (k,h+1)=1. A full evaluation pays no new reuse cost and resets cache age. A reuse extends the run whose last full evaluation occurred at j=k−hj=k-h and adds the calibrated increment only when the next age is legal. Illegal transitions receive +∞+∞. Equation (11) selects the maximum feasible reuse count: Bs∗=maxn:minhJK(n,h)≤τ.B_s^*= \n: _hJ_K(n,h)≤τ \. (11) Backtracking recovers the minimum-response schedule at Bs∗B_s^*. The recurrence exactly optimizes the calibrated response objective in O(K3)O(K^3) time and O(K2)O(K^2) rolling memory. The schedule remains causal during inference. We initialize accumulated charge as C1=0C_1=0. Before the h-th reuse in a run whose last full evaluation occurred at j, DriveCache measures the observed input drift in Equation (12): dj,hobs=∥uj+h−uj∥2∥uj+h∥2+ε.d_j,h^obs= u_j+h-u_j _2 u_j+h _2+ . (12) It accepts reuse only when dj,hobs≤d¯s,j,h+εd_j,h^obs≤ d_s,j,h+ and accumulated charge satisfies Ck+c^s,j,h≤τC_k+ c_s,j,h≤τ. Acceptance sets Ck+1=Ck+c^s,j,hC_k+1=C_k+ c_s,j,h. A failed check evaluates the backbone, keeps Ck+1=CkC_k+1=C_k, and refreshes the cache. Replanning starts from age zero after this forced evaluation and maximizes additional suffix reuse under budget τ−Ckτ-C_k; the executed prefix remains fixed. Calibration uses no optimizer or parameter update, and inference leaves generator weights and sampling steps unchanged at runtime. DriveCache applies model-specific legality to the same recurrence. Here, uku_k, rkr_k, and HkH_k are the reusable-region input, cached backbone output, and remaining sampler interface. Wan2.2 separates guidance branches; A14B also forbids cross-expert reuse. Epona caches only within each 100-step visual denoising call, never across autoregressive segments. Cache tensors, legal steps, maximum age, norm axes, and ties are frozen before evaluation. Two full anchor traces support all candidate runs. Building the (j,h)(j,h) table costs O(K2)O(K^2) substitutions and O(K3)O(K^3) unbatched Jacobian-vector products once per configuration, without evaluation clips. On one H20, calibration takes 4 minutes for 5B, 9 minutes for A14B, and 16 minutes for Epona, with at most 6.1 GB additional memory and response tables below 2 MB. Initial DP solves take 0.8, 1.4, and 9.2 ms, respectively; drift checks and suffix replanning add less than 0.7% to end-to-end latency. DriveCache updates no weights and trains no predictor; one model-level calibration applies across clips. Experiments Experimental setup Models and protocol. We evaluate Wan2.2 5B and expert-routed A14B on 222 DrivingGen scenes with ten fixed seeds (2,220 samples), and Epona on all 150 nuScenes validation scenes (49; 69; 60). Wan uses 40 denoising steps, while Epona uses 100 steps in each visual denoising call. Calibration uses 32 disjoint scenes; before evaluation, we freeze the reuse target and τ at 50% for Wan and 60% for Epona. Comparisons and metrics. Tables 1–2 cover fixed, adaptive, forecasting, spectral, and reduced-step baselines (35; 65; 32; 27; 22; 68; 37; 3; 28; 9; 36). We match Wan caches within one skipped backbone call per clip and compare reduced-step sampling at matched latency. Epona methods share hardware while retaining their native reusable units. We report peak signal-to-noise ratio (PSNR), structural similarity (SSIM), learned perceptual image patch similarity (LPIPS), trajectory quality (Traj.), average displacement error (ADE), dynamic time warping (DTW), and VBench metrics (61; 21). V-Q and V-S denote VBench quality and semantic scores. Higher PSNR, SSIM, Traj., and VBench scores indicate better quality; lower LPIPS, ADE, and DTW indicate smaller errors. Compute matching. Calibration selects each Wan cache baseline’s setting nearest the shared skipped-backbone target. Evaluation freezes that setting and verifies the one-call tolerance from realized backbone counts. Cache methods retain all 40 sampler steps, while the reduced-step baseline changes the sampling trajectory. Measurement. On one H20 at batch size one, latency averages ten synchronized runs after three warm-ups and includes controller overhead. Fidelity uses matched-seed full outputs. DrivingGen evaluates trajectory quality, ADE, and DTW against ground-truth ego trajectories, while VBench uses its native references. We average seeds per scene and report paired scene-level 95% bootstrap intervals. Run integrity. Each sample records its configuration hash, realized backbone calls, and metric outputs. We retain complete 2,220-sample DrivingGen runs and complete 150-scene Epona validation runs for comparison. Figure 5: Qualitative comparison of full inference, TeaCache, and DriveCache. Backbone / Method Lat. (s) ↓ Spd. ↑ PSNR ↑ SSIM ↑ LPIPS ↓ V-Q ↑ V-S ↑ Traj. ↑ ADE ↓ DTW ↓ Wan2.2 5B Fixed interval 10.17 1.77× 28.417 0.852 0.095 0.774 0.179 0.281 1.881 15.517 PAB 17.53 1.03× 33.344 0.953 0.042 0.774 0.179 0.262 2.191 18.454 TeaCache 10.76 1.68× 32.874 0.939 0.042 0.776 0.177 0.274 1.796 15.055 EasyCache 10.67 1.69× 33.923 0.959 0.041 0.777 0.178 0.269 1.925 16.173 MagCache 10.73 1.68× 33.945 0.949 0.038 0.773 0.177 0.261 2.712 23.156 DiCache 12.15 1.48× 33.278 0.949 0.038 0.781 0.179 0.262 1.919 16.167 TaylorSeer 10.19 1.77× 27.970 0.847 0.094 0.770 0.178 0.262 2.306 19.567 SeaCache 10.07 1.79× 32.860 0.932 0.048 0.769 0.176 0.275 1.929 16.004 DriveCache 9.79 1.84× 34.744 0.961 0.041 0.777 0.179 0.294 1.771 14.768 Wan2.2 A14B Fixed interval 86.64 1.89× 25.385 0.802 0.159 0.777 0.178 0.336 2.610 20.836 PAB 156.45 1.04× 31.422 0.931 0.079 0.776 0.181 0.354 1.980 15.475 TeaCache 82.87 1.97× 30.742 0.913 0.064 0.779 0.179 0.324 2.009 15.694 EasyCache 89.44 1.83× 31.461 0.922 0.070 0.781 0.181 0.349 2.020 15.576 MagCache 86.83 1.88× 31.987 0.937 0.049 0.777 0.180 0.346 2.138 16.858 DiCache 87.00 1.88× 29.927 0.897 0.096 0.775 0.180 0.323 1.951 15.011 TaylorSeer 87.11 1.88× 26.339 0.817 0.143 0.780 0.180 0.357 2.320 18.439 SeaCache 83.09 1.97× 31.502 0.923 0.060 0.779 0.179 0.329 1.948 14.881 DriveCache 83.10 1.97× 32.778 0.938 0.060 0.781 0.182 0.370 1.809 13.750 Table 1: Wan2.2 5B and A14B results at matched skipped-backbone counts. Method Latency ↓ Speedup ↑ PSNR ↑ SSIM ↑ LPIPS ↓ Stationary PSNR ↑ Turn PSNR ↑ Reduced steps 46.34 1.61× 26.322 0.729 0.144 31.165 23.855 DeepCache 44.27 1.68× 26.358 0.730 0.142 31.192 23.898 FasterCache 46.46 1.60× 26.312 0.729 0.144 31.155 23.847 TeaCache 47.28 1.58× 26.152 0.726 0.146 30.800 23.744 AdaCache 44.07 1.69× 26.328 0.729 0.144 31.154 23.870 FlowCache 51.57 1.44× 26.322 0.729 0.144 31.123 23.885 DriveCache 39.55 1.89× 26.394 0.734 0.140 31.479 23.910 Table 2: DriveCache results on Epona. The method reaches 1.89×1.89× speedup at video quality comparable to cache baselines. Figure 6: Joint propagation captures reuse-run interactions. Figure 7: Joint responses vary with reuse start step and cache age, guiding reuse placement across denoising steps. Variant PSNR ↑ LPIPS ↓ Traj. ↑ DriveCache 34.744 0.041 0.294 No motion prior 34.201 0.046 0.274 Shuffled trajectory 34.286 0.045 0.275 Translation only 34.431 0.043 0.278 Rotation only 34.365 0.044 0.277 No DP (fixed placement) 33.982 0.050 0.272 No drift veto 34.512 0.043 0.279 Table 3: Wan2.2 5B ablations at 21/40 realized reuse isolate motion allocation, step placement, and runtime correction. High/low reuses Lat. ↓ PSNR ↑ SSIM ↑ LPIPS ↓ 0/20 83.10 32.778 0.938 0.060 2/18 83.02 30.747 0.914 0.073 3/17 83.34 29.936 0.900 0.081 Table 4: A14B expert routing at fixed compute. High/low denotes reuse counts in the high- and low-noise experts. DriveCache improves quality at matched speed At the frozen operating points in Table 1, DriveCache reaches 1.84×1.84× speedup and 34.744 dB PSNR on 5B, and 1.97×1.97× and 32.778 dB on A14B. Viewed jointly, it is fastest on 5B and Epona and tied for the highest speedup on A14B, while retaining the best values in most reported quality and driving dimensions. Figure 1 shows DriveCache’s quality gains across the evaluated driving motions in every motion group. Fidelity and downstream behavior. At comparable speed, DriveCache improves over SeaCache by 1.884 dB PSNR (95% CI: 1.61–2.15) on 5B and by 1.276 dB (1.04–1.51) on A14B; trajectory quality and DTW also improve. Trajectory metrics show ego-motion consistency; Figure 5 and VBench show temporal consistency and stable scene semantics. Motion adaptation and robustness. The anchors cover 96.4% of scenes; stationary clips average 22.4 reuses versus 18.7 for turns, confirming motion-conditioned allocation. The motion prior uses planned ego motion to set the schedule for every supported scene before the first denoising step. The drift check evaluates generated states before each planned reuse during inference. Under 1.0 m/5 degree pose noise, DriveCache retains 1.57×1.57× speedup with a 0.213 dB loss, showing that causal correction complements the motion prior. Transfer to Epona. Across 150 Epona scenes, DriveCache reaches 1.89×1.89× speedup and 0.140 LPIPS, versus DeepCache’s 1.68×1.68× at comparable aggregate and turning fidelity (Table 2). Larger stationary gains reflect more reuse on tolerant scenes and conservative behavior on turns. The Epona results extend the latency–fidelity gains to a separate scene collection. Ablations validate scheduling choices Planned motion controls scene-level reuse. Mean speed correlates with cache error, while removing or shuffling motion degrades fixed-reuse quality. Shuffling preserves the trajectory distribution but breaks scene correspondence, isolating action–scene alignment; translation and rotation both contribute to cache allocation. Joint responses control placement. Figures Experimental setup–7 show that joint propagation captures start-step and cache-age effects missed by isolated sums. Fixed placement causes the largest loss, while the drift veto guards against schedule drift. Calibration and routing remain stable. Predicted response tracks realized degradation, and two to eight anchors change PSNR by only 0.062 dB. Moving reuse into the high-noise A14B expert sharply reduces quality (Table 4), so expert boundaries matter more than richer interpolation at fixed compute. The core ablations in Table 3 hold the pretrained generator fixed while isolating scene allocation, step placement, and runtime correction. Together, the anchor and routing results support sparse calibration and explicit expert boundaries in the controller. Conclusion Across the evaluated driving motions, planned ego translation and rotation predict scene-level cache tolerance before denoising. We introduce DriveCache, a training-free controller that allocates reuse across scenes and places it across denoising steps with exact dynamic programming. Two anchor traces calibrate joint run responses, and a causal drift check refreshes the cache and replans the unexecuted suffix when generation departs from calibration. On Wan2.2 5B, Wan2.2 A14B, and Epona, DriveCache improves quality–efficiency trade-offs over cache baselines while preserving generator weights and sampling steps. These results support planned ego motion as a control signal for action-aware diffusion caching. References Blattmann et al. (2023) A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, et al. Stable video diffusion: scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127. Cited by: Introduction, Diffusion Models. Bolya et al. (2022) D. Bolya, C. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman Token merging: your vit but faster. arXiv preprint arXiv:2210.09461. Cited by: Diffusion Models. Bu et al. (2025) J. Bu, P. Ling, Y. Zhou, Y. Wang, Y. Zang, D. Lin, and J. Wang Dicache: let diffusion model determine its own cache. arXiv preprint arXiv:2508.17356. Cited by: Introduction, Training-free Diffusion Inference Acceleration, Experimental setup. Caesar et al. (2020) H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom Nuscenes: a multimodal dataset for autonomous driving. In CVPR, Cited by: Autonomous Driving Video Generation. Chen et al. (2024a) J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Z. Wang, J. Kwok, P. Luo, H. Lu, and Z. Li PixArt-α: fast training of diffusion transformer for photorealistic text-to-image synthesis. In ICLR, Cited by: Diffusion Models. Chen et al. (2024b) P. Chen, M. Shen, P. Ye, J. Cao, C. Tu, C. Bouganis, Y. Zhao, and T. Chen Δ -DiT: a training-free acceleration method tailored for diffusion transformers. External Links: 2406.01125, Link Cited by: Training-free Diffusion Inference Acceleration. Chitta et al. (2023) K. Chitta, A. Prakash, B. Jaeger, Z. Yu, K. Renz, and A. Geiger Transfuser: imitation with transformer-based sensor fusion for autonomous driving. IEEE transactions on pattern analysis and machine intelligence. Cited by: Autonomous Driving Video Generation. Chu et al. (2025) H. Chu, W. Wu, G. Feng, and Y. Zhang OmniCache: a trajectory-oriented global perspective on training-free cache reuse for diffusion transformer models. In ICCV, Cited by: Training-free Diffusion Inference Acceleration. Chung et al. (2026) J. Chung, S. Hyun, M. Lee, B. Han, G. Cha, D. Wee, Y. Hong, and J. Heo SeaCache: spectral-evolution-aware cache for accelerating diffusion models. arXiv preprint arXiv:2602.18993. Cited by: Introduction, Training-free Diffusion Inference Acceleration, Experimental setup. Duan et al. (2024) Y. Duan, X. Guo, Z. Zhu, Z. Wang, Y. Wang, and C. Lin MaskFuser: masked fusion of joint multi-modal tokenization for end-to-end autonomous driving. arXiv preprint arXiv:2405.07573. Cited by: Autonomous Driving Video Generation. Gao et al. (2025) R. Gao, K. Chen, B. Xiao, L. Hong, Z. Li, and Q. Xu MagicDrive-v2: high-resolution long video generation for autonomous driving with adaptive control. In ICCV, Cited by: Autonomous Driving Video Generation. Gao et al. (2024a) R. Gao, K. Chen, E. Xie, L. Hong, Z. Li, D. Yeung, and Q. Xu Magicdrive: street view generation with diverse 3d geometry control. In ICLR, Cited by: Autonomous Driving Video Generation. Gao et al. (2024b) S. Gao, J. Yang, L. Chen, K. Chitta, Y. Qiu, A. Geiger, J. Zhang, and H. Li Vista: a generalizable driving world model with high fidelity and versatile controllability. Advances in Neural Information Processing Systems. Cited by: Introduction, Autonomous Driving Video Generation. Guo et al. (2025a) X. Guo, W. Yuan, Y. Zhang, T. Yang, C. Zhang, Z. Zhu, Q. Zou, and L. Chen Adjacent-view transformers for supervised surround-view depth estimation. In IROS, Cited by: Autonomous Driving Video Generation. Guo et al. (2025b) X. Guo, R. Zhang, Y. Duan, Y. He, D. Nie, et al. SURDS: benchmarking spatial understanding and reasoning in driving scenarios with vision language models. In NeurIPS, Cited by: Autonomous Driving Video Generation. Guo et al. (2025c) X. Guo, R. Zhang, Y. Duan, R. Wang, M. Poggi, K. Zhou, W. Zheng, W. Huang, G. Xu, Y. Peng, Y. Si, and Q. Zou ROVR-Open-Dataset: a large-scale depth dataset for autonomous driving. arXiv preprint arXiv:2508.13977. Cited by: Autonomous Driving Video Generation. Ho et al. (2020) J. Ho, A. Jain, and P. Abbeel Denoising diffusion probabilistic models. Advances in neural information processing systems. Cited by: Diffusion Models. Ho et al. (2022) J. Ho, T. Salimans, A. A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet Video diffusion models. In ICLR Workshop, Cited by: Introduction, Diffusion Models. Hu et al. (2023) A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado Gaia-1: a generative world model for autonomous driving. arXiv preprint arXiv:2309.17080. Cited by: Introduction, Autonomous Driving Video Generation. Huang et al. (2024a) B. Huang, Y. Wen, Y. Zhao, Y. Hu, Y. Liu, F. Jia, W. Mao, T. Wang, C. Zhang, C. W. Chen, et al. Subjectdrive: scaling generative data in autonomous driving via subject control. arXiv preprint arXiv:2403.19438. Cited by: Autonomous Driving Video Generation. Huang et al. (2024b) Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, et al. Vbench: comprehensive benchmark suite for video generative models. In CVPR, Cited by: Experimental setup. Kahatapitiya et al. (2025) K. Kahatapitiya, H. Liu, S. He, D. Liu, M. Jia, C. Zhang, M. S. Ryoo, and T. Xie Adaptive caching for faster video generation with diffusion transformers. In ICCV, Cited by: Introduction, Training-free Diffusion Inference Acceleration, Experimental setup. Karras et al. (2022) T. Karras, M. Aittala, T. Aila, and S. Laine Elucidating the design space of diffusion-based generative models. In NeurIPS, Cited by: Diffusion Models. Kong et al. (2024) W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al. Hunyuanvideo: a systematic framework for large video generative models. arXiv preprint arXiv:2412.03603. Cited by: Introduction, Diffusion Models. Lin et al. (2024) B. Lin, Y. Ge, X. Cheng, Z. Li, B. Zhu, S. Wang, X. He, Y. Ye, S. Yuan, L. Chen, et al. Open-sora plan: open-source large video generation model. arXiv preprint arXiv:2412.00131. Cited by: Introduction, Diffusion Models. Lipman et al. (2023) Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le Flow matching for generative modeling. In ICLR, Cited by: Diffusion Models. Liu et al. (2025a) F. Liu, S. Zhang, X. Wang, Y. Wei, H. Qiu, Y. Zhao, Y. Zhang, Q. Ye, and F. Wan Timestep embedding tells: it’s time to cache for video diffusion model. In CVPR, Cited by: Introduction, Training-free Diffusion Inference Acceleration, Experimental setup. Liu et al. (2025b) J. Liu, C. Zou, Y. Lyu, J. Chen, and L. Zhang From reusing to forecasting: accelerating diffusion models with taylorseers. In ICCV, Cited by: Introduction, Training-free Diffusion Inference Acceleration, Experimental setup. Lu et al. (2022) C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu DPM-solver: a fast ode solver for diffusion probabilistic model sampling in around 10 steps. arXiv preprint arXiv:2206.00927. Cited by: Diffusion Models, Training-free Diffusion Inference Acceleration. Lu et al. (2024) J. Lu, Z. Huang, Z. Yang, J. Zhang, and L. Zhang WoVoGen: world volume-aware diffusion for controllable multi-camera driving scene generation. In ECCV, Cited by: Autonomous Driving Video Generation. Luo et al. (2023) S. Luo, Y. Tan, L. Huang, J. Li, and H. Zhao Latent consistency models: synthesizing high-resolution images with few-step inference. External Links: 2310.04378 Cited by: Diffusion Models. Lv et al. (2025) Z. Lv, C. Si, J. Song, Z. Yang, Y. Qiao, Z. Liu, and K. K. Wong Fastercache: training-free video diffusion model acceleration with high quality. In ICLR, Cited by: Training-free Diffusion Inference Acceleration, Experimental setup. Ma et al. (2024a) X. Ma, Y. Wang, X. Chen, G. Jia, Z. Liu, Y. Li, C. Chen, and Y. Qiao Latte: latent diffusion transformer for video generation. arXiv preprint arXiv:2401.03048. Cited by: Diffusion Models. Ma et al. (2024b) X. Ma, G. Fang, M. Bi Mi, and X. Wang Learning-to-cache: accelerating diffusion transformer via layer caching. Advances in Neural Information Processing Systems. Cited by: Training-free Diffusion Inference Acceleration. Ma et al. (2024c) X. Ma, G. Fang, and X. Wang Deepcache: accelerating diffusion models for free. In CVPR, Cited by: Training-free Diffusion Inference Acceleration, Experimental setup. Ma et al. (2026) Y. Ma, X. Zheng, J. Xu, X. Xu, F. Ling, X. Zheng, H. Kuang, H. Li, X. Wang, X. Xiao, F. Chao, and R. Ji Flow caching for autoregressive video generation. External Links: 2602.10825, Link Cited by: Training-free Diffusion Inference Acceleration, Experimental setup. Ma et al. (2025) Z. Ma, L. Wei, F. Wang, S. Zhang, and Q. Tian Magcache: fast video generation with magnitude-aware cache. Advances in Neural Information Processing Systems. Cited by: Introduction, Training-free Diffusion Inference Acceleration, Experimental setup. Ni et al. (2025) J. Ni, Y. Guo, Y. Liu, R. Chen, L. Lu, and Z. Wu Maskgwm: a generalizable driving world model with video mask reconstruction. In CVPR, Cited by: Autonomous Driving Video Generation. NVIDIA et al. (2025) NVIDIA, N. Agarwal, A. Ali, M. Bala, Y. Balaji, E. Barker, T. Cai, P. Chattopadhyay, Y. Chen, Y. Cui, Y. Ding, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575. Cited by: Autonomous Driving Video Generation. Peebles and Xie (2023) W. Peebles and S. Xie Scalable diffusion models with transformers. In ICCV, Cited by: Diffusion Models. Qiu et al. (2025) J. Qiu, L. Liu, S. Wang, J. Lu, K. Chen, and Y. Hao Accelerating diffusion transformer via gradient-optimized cache. arXiv preprint arXiv:2503.05156. Cited by: Training-free Diffusion Inference Acceleration. Rombach et al. (2022) R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer High-resolution image synthesis with latent diffusion models. In CVPR, Cited by: Diffusion Models. Salimans and Ho (2022) T. Salimans and J. Ho Progressive distillation for fast sampling of diffusion models. arXiv preprint arXiv:2202.00512. Cited by: Introduction, Diffusion Models. Selvaraju et al. (2024) P. Selvaraju, T. Ding, T. Chen, I. Zharkov, and L. Liang Fora: fast-forward caching in diffusion transformer acceleration. arXiv preprint arXiv:2407.01425. Cited by: Training-free Diffusion Inference Acceleration. Song et al. (2020) J. Song, C. Meng, and S. Ermon Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502. Cited by: Diffusion Models, Training-free Diffusion Inference Acceleration. Song et al. (2025) R. Song, X. Guo, Y. Peng, Q. Wei, H. Wu, and L. Chen InsightDrive: insight scene representation for end-to-end autonomous driving. arXiv preprint arXiv:2503.13047. Cited by: Autonomous Driving Video Generation. Song et al. (2023) Y. Song, P. Dhariwal, M. Chen, and I. Sutskever Consistency models. arXiv preprint arXiv:2303.01469. Cited by: Introduction, Diffusion Models. Song et al. (2021) Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole Score-based generative modeling through stochastic differential equations. In ICLR, External Links: Link Cited by: Diffusion Models. Wan et al. (2025) T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, et al. Wan: open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314. Cited by: Experimental setup. Wang et al. (2024a) X. Wang, Z. Zhu, G. Huang, X. Chen, J. Zhu, and J. Lu Drivedreamer: towards real-world-drive world models for autonomous driving. In ECCV, Cited by: Introduction, Autonomous Driving Video Generation. Wang et al. (2024b) Y. Wang, J. He, L. Fan, H. Li, Y. Chen, and Z. Zhang Driving into the future: multiview visual forecasting and planning with world model for autonomous driving. In CVPR, Cited by: Autonomous Driving Video Generation. Wen et al. (2024) Y. Wen, Y. Zhao, Y. Liu, F. Jia, Y. Wang, C. Luo, C. Zhang, T. Wang, X. Sun, and X. Zhang Panacea: panoramic and controllable video generation for autonomous driving. In CVPR, Cited by: Autonomous Driving Video Generation. Wu et al. (2025) W. Wu, X. Guo, W. Tang, T. Huang, C. Wang, and C. Ding Drivescape: high-resolution driving video generation by multi-view feature fusion. In CVPR, Cited by: Autonomous Driving Video Generation. Xie et al. (2024) Y. Xie, X. Guo, C. Wang, K. Liu, and L. Chen AdvDiffuser: generating adversarial safety-critical driving scenarios via guided diffusion. arXiv preprint arXiv:2410.08453. Cited by: Autonomous Driving Video Generation. Xing et al. (2025) Z. Xing, X. Zhang, Y. Hu, B. Jiang, T. He, Q. Zhang, X. Long, and W. Yin Goalflow: goal-driven flow matching for multimodal trajectories generation in end-to-end autonomous driving. In CVPR, Cited by: Autonomous Driving Video Generation. Yang et al. (2025) Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al. Cogvideox: text-to-video diffusion models with an expert transformer. In ICLR, Cited by: Introduction, Diffusion Models. Yin et al. (2024) T. Yin, M. Gharbi, T. Park, R. Zhang, E. Shechtman, F. Durand, and W. T. Freeman Improved distribution matching distillation for fast image synthesis. In NeurIPS, Cited by: Diffusion Models. Yin et al. (2025) T. Yin, Q. Zhang, R. Zhang, W. T. Freeman, F. Durand, E. Shechtman, and X. Huang From slow bidirectional to fast autoregressive video diffusion models. In CVPR, Cited by: Diffusion Models. Zhang et al. (2024) J. Zhang, J. Wei, H. Huang, P. Zhang, J. Zhu, and J. Chen Sageattention: accurate 8-bit attention for plug-and-play inference acceleration. arXiv preprint arXiv:2410.02367. Cited by: Diffusion Models. Zhang et al. (2025) K. Zhang, Z. Tang, X. Hu, X. Pan, X. Guo, Y. Liu, J. Huang, L. Yuan, Q. Zhang, X. Long, et al. Epona: autoregressive diffusion world model for autonomous driving. In ICCV, Cited by: Autonomous Driving Video Generation, Experimental setup. Zhang et al. (2018) R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, Cited by: Experimental setup. Zhang et al. (2026a) Z. Zhang, Z. Jin, Y. Peng, X. Guo, H. Liu, S. Zhang, X. Ma, Z. Wu, J. Yan, X. Jia, and Y. Jiang Bench2Drive-Robust: benchmarking closed-loop autonomous driving under deployment perturbations. arXiv preprint arXiv:2605.18059. Cited by: Autonomous Driving Video Generation. Zhang et al. (2026b) Z. Zhang, Y. Peng, J. Zhang, X. Guo, Z. Huang, H. Liu, Q. Li, S. Zhang, X. Jia, and J. Yan ReactSim-Bench: benchmarking reactive behavior world model simulation in autonomous driving. arXiv preprint arXiv:2606.14058. Cited by: Autonomous Driving Video Generation. Zhao et al. (2024) G. Zhao, X. Wang, Z. Zhu, X. Chen, G. Huang, X. Bao, and X. Wang DriveDreamer-2: llm-enhanced world models for diverse driving video generation. arXiv preprint arXiv:2403.06845. Cited by: Autonomous Driving Video Generation. Zhao et al. (2025) X. Zhao, X. Jin, K. Wang, and Y. You Real-time video generation with pyramid attention broadcast. In ICLR, Cited by: Training-free Diffusion Inference Acceleration, Experimental setup. Zheng et al. (2024a) W. Zheng, W. Chen, Y. Huang, B. Zhang, Y. Duan, and J. Lu Occworld: learning a 3d occupancy world model for autonomous driving. In ECCV, Cited by: Autonomous Driving Video Generation. Zheng et al. (2024b) W. Zheng, R. Song, X. Guo, C. Zhang, and L. Chen Genad: generative end-to-end autonomous driving. In ECCV, Cited by: Autonomous Driving Video Generation. Zhou et al. (2025) X. Zhou, D. Liang, K. Chen, T. Feng, X. Chen, H. Lin, Y. Ding, F. Tan, H. Zhao, and X. Bai Less is enough: training-free video diffusion acceleration via runtime-adaptive caching. arXiv preprint arXiv:2507.02860. Cited by: Introduction, Training-free Diffusion Inference Acceleration, Experimental setup. Zhou et al. (2026) Y. Zhou, H. Shao, L. Wang, Z. Zong, H. Li, and S. L. Waslander DrivingGen: a comprehensive benchmark for generative video world models in autonomous driving. arXiv preprint arXiv:2601.01528. Cited by: Autonomous Driving Video Generation, Experimental setup. Zou et al. (2024a) C. Zou, X. Liu, T. Liu, S. Huang, and L. Zhang Accelerating diffusion transformers with token-wise feature caching. arXiv preprint arXiv:2410.05317. Cited by: Training-free Diffusion Inference Acceleration. Zou et al. (2024b) C. Zou, E. Zhang, R. Guo, H. Xu, C. He, X. Hu, and L. Zhang Accelerating diffusion transformers with dual feature caching. arXiv preprint arXiv:2412.18911. Cited by: Training-free Diffusion Inference Acceleration.