Paper deep dive
Disagree to Accelerate: Closing the Loop on Diffusion Feature Forecasts
Yanchao Li, Jiaqing Xie, Ben Gao, Wanhao Liu, Yanbo Wang, T. Y. Tsui, Jinfei Liu, Yuqiang Li, Tianfan Fu
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps. Recent work has mainly focused on designing stronger forecasters. Yet forecast error varies sharply across steps, and open-loop caches trust the forecast in full at every skipped step. This fixed trust is what breaks as acceleration turns aggressive. The missing question is not only how to forecast better, but when and how much to trust a forecast. We show that reliability can be observed from the cache itself. Two forecasts agree where the feature trajectory is smooth, and they diverge where prediction turns hard. Their disagreement is a cheap runtime signal, and it costs no extra denoiser evaluation. Based on this signal, we introduce RACER, a training-free closed-loop controller with two responses. It continuously shrinks uncertain forecasts toward the last computed feature. At the riskiest steps, RACER refreshes the feature and repays the added evaluation by skipping a later scheduled one. We derive a deterministic error bound for the shrinkage and empirically evaluate its validity and tightness across acceleration regimes. At the same number of denoiser evaluations, RACER improves the strongest open-loop baseline across SD3.5-Large, FLUX.1-dev, Wan2.1-14B, and HunyuanVideo on DrawBench, VBench, and COCO. On SD3.5, we further show that RACER samples faster at equal quality. RACER generalizes across forecasting designs as well. For example, it recovers much of the quality lost on a Taylor base. These results show that reliable diffusion acceleration also depends on how forecasts are used. Code is available at this https URL
Tags
Links
- Source: https://arxiv.org/abs/2608.01740v1
- Canonical: https://arxiv.org/abs/2608.01740v1
Trouble viewing inline? Open PDF directly →
Full Text
38,300 characters extracted from source content.
Expand or collapse full text
Disagree to Accelerate: Closing the Loop on Diffusion Feature Forecasts Yanchao Li1,2, Jiaqing Xie2, Ben Gao2, Wanhao Liu2, Yanbo Wang3, T. Y. Tsui4, Jinfei Liu5, Yuqiang Li2 , Tianfan Fu1,2 Abstract Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps. Recent work has mainly focused on designing stronger forecasters. Yet forecast error varies sharply across steps, and open-loop caches trust the forecast in full at every skipped step. This fixed trust is what breaks as acceleration turns aggressive. The missing question is not only how to forecast better, but when and how much to trust a forecast. We show that reliability can be observed from the cache itself. Two forecasts agree where the feature trajectory is smooth, and they diverge where prediction turns hard. Their disagreement is a cheap runtime signal, and it costs no extra denoiser evaluation. Based on this signal, we introduce RACER, a training-free closed-loop controller with two responses. It continuously shrinks uncertain forecasts toward the last computed feature. At the riskiest steps, RACER refreshes the feature and repays the added evaluation by skipping a later scheduled one. We derive a deterministic error bound for the shrinkage and empirically evaluate its validity and tightness across acceleration regimes. At the same number of denoiser evaluations, RACER improves the strongest open-loop baseline across SD3.5-Large, FLUX.1-dev, Wan2.1-14B, and HunyuanVideo on DrawBench, VBench, and COCO. On SD3.5, we further show that RACER samples faster at equal quality. RACER generalizes across forecasting designs as well. For example, it recovers much of the quality lost on a Taylor base. These results show that reliable diffusion acceleration also depends on how forecasts are used. Code is available at https://github.com/LiZaiyuan0619/RACER. Introduction Diffusion models dominate image and video generation (Ho et al. 2020; Rombach et al. 2022; Peebles and Xie 2023), but their sampling is slow. One sample runs a large denoiser tens to hundreds of times. This number of function evaluations (NFE) drives the sampling time. Training-free feature caching is a leading remedy (Ma et al. 2024; Selvaraju et al. 2024; Liu et al. 2025a). It runs the network at only some steps and predicts the skipped features from a cache. Recent work has focused on improving the forecaster. The rules have moved from local Taylor expansions to global fits such as Chebyshev polynomials (Liu et al. 2025b; Han et al. 2026). The remaining blind spot lies in how each forecast is used. Whatever the rule, the prediction is used at full weight, with no check at run time. Forecast error is far from uniform. It stays small on smooth stretches of the feature trajectory and spikes where it bends. Even the strongest forecaster therefore becomes unreliable at high speedups. The challenge is then not the next forecasting rule. It is how far a given forecast should be trusted at each step. The way a forecast fails leaves a trace in the cache. Each cheap forecaster carries its own bias, so each drifts in its own way as prediction gets harder. Two different forecasters therefore agree on the easy stretches and split apart on the hard ones. The split is visible at run time, from the cache alone and with no extra denoiser evaluation. We test this disagreement as a reliability signal on four image and video models. It flags the highest-error steps at a mean AUROC of 0.94 (Figure 1). The input-side signal that prior caches rely on averages 0.77. To fill the gap, we introduce RACER, a Reliability-Aware Closed-loop controller with Exact Repay. RACER keeps the base forecaster unchanged. When the signal grows uncertain, RACER shrinks the forecast continuously toward the last computed feature. At a step too risky to trust at all, RACER spends one denoiser evaluation to recompute the feature, then repays it by skipping a later scheduled step. The per-prompt NFE therefore never moves from the base. The shrinkage is grounded in a deterministic endpoint-error bound and an MSE-optimality analysis. We empirically evaluate the bound and its tightness on real traces. A short offline trace fixes the controller scalars, with no training or denoiser modification. Figure 1: The disagreement r is a cross-model reliability signal. (a) On four image and video models r rises with the true forecast error, and it flags the high-error steps better than prior caches. The regime of this case is α=3. (b) The signal separates the highest-error steps on every model and across all five acceleration regimes. By default, image models are evaluated on DrawBench, and video models on VBench. At a matched NFE, RACER improves the strongest open-loop baseline on all four models and all three benchmarks. The advantage holds in every regime, and it is clearest at the aggressive end where the open-loop base breaks. The margin also buys speed. On SD3.5, for instance, RACER reaches a measured 5.4× over the 50-step sampler. Even against the strongest base, it matches the base’s best quality at 0.76×0.76× the latency and 0.72×0.72× the NFE. Under a fixed refresh-only executor, the disagreement outperforms horizon and random signals by up to 2.42 dB. The mechanism is also not tied to one forecaster pair. It rescues a Taylor base by up to 3.83 dB, and the observer itself can be swapped. The controller form is shared across all four models, and image-side calibration transfers from SD3.5 to FLUX with only a 4.7% change in the trust threshold. We make four major contributions. 1. We recast feature caching as a question of the reliability of a forecast. This reliability is observable from the cache itself, as the disagreement between two cached forecasts. 2. We introduce RACER, a training-free closed-loop controller that turns this signal into two responses, a continuous trust shrinkage and an exact refresh-and-repay. 3. The mechanism is decoupled from the forecaster. It generalizes to a new base forecaster and to a new observer, improving quality in both cases. 4. Across four image and video models and three benchmarks, RACER improves the strongest open-loop baseline at matched NFE. On SD3.5, it also reaches equal quality at lower latency. Related Work Diffusion sampling is slow because it runs the denoiser many times. Two lines of work cut this cost. One line takes fewer steps, with faster solvers, distillation, or straighter flow paths (Song et al. 2021; Lu et al. 2022; Zhao et al. 2023; Salimans and Ho 2022; Luo et al. 2023; Liu et al. 2022; Lipman et al. 2023). The other makes each step cheaper, through quantization, pruning, or token merging, or spreads it across devices (Li et al. 2023; Fang et al. 2023; Bolya and Hoffman 2023; Li et al. 2024). Feature caching belongs to the second line and needs no extra training (Ma et al. 2024). Early caches reuse a stored feature, and recent work turns from reuse to forecasting (Zou et al. 2025). A forecaster reads the cache and extrapolates the skipped feature. Early rules used a local expansion (Liu et al. 2025b). A global basis fit came next, and its error grows more slowly with the skip horizon than a local expansion (Han et al. 2026; Trefethen 2019). Later work varies the basis, with Hermite, learned, or spectral forms (Zheng et al. 2026). RACER is orthogonal to these designs. It treats the forecaster as a proposal and decides how much to trust it. Most of these forecast caches apply the prediction open loop. Once a step is skipped, the forecast is used in full, whatever its reliability. A more recent group adds a runtime check, asking how far a cached prediction can be trusted (Zheng et al. 2025). Forecast-then-verify caches recompute an actual feature and accept or reject the forecast. Others read a shallow probe, an input-side change, or a feature change rate (Liu et al. 2025a; Zhou et al. 2025; Cui et al. 2026). RACER sits in this line and reads its signal on the output side, as the disagreement between two forecasts already in the cache. It needs no extra denoiser evaluation and no input-side probe. Prior signals mostly gate a binary recompute, while RACER turns the signal into continuous trust. Method Preliminaries A diffusion model draws a sample over a set of N steps T=t1,…,tNT=\t_1,…,t_N\. Each step evaluates the denoiser ϵθ _θ once, and a solver then advances the state, xi=Solve(xi−1,ϵθ(xi−1,ti),ti).x_i=Solve (x_i-1,\, _θ(x_i-1,t_i),\,t_i ). (1) One sample therefore runs the denoiser N times, and this NFE dominates the sampling time. Feature caching lowers the NFE by running the denoiser at only some steps and forecasting the rest (Ma et al. 2024; Selvaraju et al. 2024). The computed set U⊆TU T holds the steps that run the denoiser and cache their features. The skip set V=T∖UV=T U holds the rest. At a skipped step a forecaster f predicts the missing feature from the cache, written h^t=f(Ct) h_t=f(C_t). Only |U||U| steps run the denoiser, so this triple (U,V,f)(U,V,f) fixes the NFE. Methods differ only in f, from feature reuse to local or global extrapolation. We adopt a global Chebyshev forecaster as our base f (Han et al. 2026). We call the nearest cached feature the anchor a. Prior caches use this base open loop, trusting h^t h_t in full at every skipped step. But the forecast error is not uniform across steps. A single fixed trust therefore breaks under aggressive skipping. Our starting point is that the reliability of each forecast is readable from the cache alone, at no extra denoiser evaluation. Alongside the base forecast h^t h_t, we compute a second forecast g^t=g(Ct) g_t=g(C_t) from the same cache. We call this second forecaster the observer, and use a local Taylor rule for it. Disagreement between predictors is a standard proxy for predictive uncertainty (Seung et al. 1992; Lakshminarayanan et al. 2017). The two forecasts carry different biases, so their disagreement is our reliability signal r, rt=‖h^t−g^t‖/‖h^t‖.r_t=\| h_t- g_t\|\,/\,\| h_t\|. (2) When the trajectory is smooth the two forecasts nearly agree, and r stays small. Where the trajectory bends they diverge, and r grows, marking a step where the forecast is unreliable. Computing r reuses the cache and adds only O(F)O(F) arithmetic in the feature size F. In the measured SD3.5 setting, this auxiliary computation adds about 2% latency. The effect is not tied to this pair, only to two forecasts with different biases. We close the loop on r with RACER, a training-free controller. It keeps the base (U,V,f)(U,V,f) unchanged and adds the observer g and a control policy π. The policy maps the runtime signal to two responses, π:(rt,kt)↦(κt,γt)π:(r_t,k_t) ( _t, _t), where κt _t is the trust placed on the forecast and γt∈0,1 _t∈\0,1\ flags a refresh. The method is thus the tuple (U,V,f,g,π)(U,V,f,g,π), orthogonal to the choice of f. An open-loop cache is the special case in which neither output depends on rtr_t. Figure 2: Empirical evaluation of Proposition 1. The bound covers every evaluated step, remains tight with an error-to-bound ratio near 0.9–1.0, and tracks measured error across regimes. Closed-Loop Control SD3.5-Large FLUX.1-dev Method PSNR↑ SSIM↑ LPIPS↓ ImageReward↑ CLIP↑ PSNR↑ SSIM↑ LPIPS↓ ImageReward↑ CLIP↑ Reference – – – 1.04 28.47 – – – 1.01 27.38 α=0.25α=0.25 NFE 18 FORA 10.84 0.485 0.591 0.85 28.33 16.17 0.673 0.387 0.99 27.32 TaylorSeer 11.49 0.547 0.505 0.81 27.96 19.87 0.781 0.227 1.01 27.46 TeaCache 10.84 0.487 0.589 0.86 28.21 18.66 0.749 0.280 0.98 27.46 ToCa 12.79 0.586 0.449 0.91 28.46 10.67 0.450 0.665 0.99 27.34 Chebyshev 18.35 0.761 0.232 1.01 28.48 24.48 0.849 0.139 1.00 27.40 Ours 19.68 0.792 0.201 1.01 28.48 24.62 0.850 0.142 1.03 27.46 α=0.75α=0.75 NFE≈14 FORA 10.33 0.425 0.661 0.78 28.47 15.24 0.638 0.442 0.99 27.20 TaylorSeer 10.19 0.480 0.601 0.48 27.12 18.29 0.732 0.286 0.97 27.40 TeaCache 11.85 0.555 0.513 0.91 28.26 18.71 0.751 0.277 1.03 27.34 ToCa 11.87 0.530 0.526 0.74 28.17 18.66 0.726 0.294 0.97 27.35 Chebyshev 17.85 0.737 0.267 0.97 28.38 24.08 0.837 0.146 0.99 27.29 Ours 18.84 0.750 0.245 0.99 28.53 24.27 0.847 0.164 1.03 27.41 α=3.0α=3.0 NFE≈10 FORA 10.02 0.364 0.739 0.59 27.86 14.46 0.603 0.506 0.86 27.01 TaylorSeer 9.10 0.411 0.727 −-0.22 25.23 16.31 0.662 0.384 0.96 27.16 TeaCache 10.21 0.415 0.676 0.73 28.04 16.75 0.681 0.380 1.03 27.32 ToCa 11.44 0.486 0.574 0.55 27.60 17.52 0.671 0.374 0.95 27.57 Chebyshev 15.70 0.613 0.393 0.79 27.73 22.15 0.772 0.234 0.99 27.58 Ours 16.52 0.653 0.355 0.90 28.39 22.19 0.778 0.252 1.00 27.81 α=5.0α=5.0 NFE 9 FORA 10.02 0.364 0.739 0.59 27.86 14.46 0.603 0.506 0.86 27.01 TaylorSeer 9.10 0.411 0.727 −-0.22 25.23 16.31 0.662 0.384 0.96 27.16 TeaCache 10.21 0.415 0.676 0.73 28.04 16.75 0.681 0.380 1.03 27.32 ToCa 11.44 0.486 0.574 0.55 27.60 17.52 0.671 0.374 0.95 27.57 Chebyshev 15.06 0.553 0.465 0.47 26.83 20.78 0.717 0.339 0.95 27.76 Ours 15.73 0.618 0.397 0.83 28.14 21.44 0.742 0.299 1.00 27.81 α=7.0α=7.0 NFE 8 FORA 10.01 0.341 0.762 0.48 27.48 14.22 0.592 0.530 0.79 27.18 TaylorSeer 8.63 0.383 0.784 −-0.84 23.64 14.90 0.610 0.463 0.90 26.95 TeaCache 10.01 0.344 0.761 0.47 27.47 15.54 0.627 0.479 0.84 27.22 ToCa 11.14 0.450 0.632 0.05 26.79 10.93 0.435 0.696 0.91 27.70 Chebyshev 14.39 0.494 0.484 0.34 26.74 20.06 0.662 0.405 0.87 27.79 Ours 14.95 0.566 0.444 0.70 27.70 20.48 0.686 0.369 0.92 27.91 Table 1: Image main results on DrawBench, quality at matched NFE. Best in bold, second best underlined. The four cache baselines land on one schedule for α=3.0α=3.0 and 5.05.0, so their entries repeat there. The first component of π shrinks trust continuously. At a skipped step it outputs a feature between the anchor ata_t and the forecast h^t h_t, y^t=at+κt(h^t−at) y_t=a_t+ _t( h_t-a_t). The trust κt _t falls as the horizon ktk_t grows and as rtr_t rises, κt=exp(−λmax(kt−1,0))⋅σ(β(θκ−rt)). _t= \! (-λ\, (k_t-1,0) )·σ\! (β\,( _κ-r_t) ). (3) Here σ is the sigmoid, and ktk_t is the number of steps since the last computed step. This response spends no denoiser evaluation, and only combines cached features. Its two extremes recover existing methods, pure reuse at κt=0 _t=0 and the full forecast at κt=1 _t=1. RACER lies between them, guided by r. The second component of π is an exact refresh-and-repay. It sets γt=1 _t=1 where the disagreement exceeds a refresh threshold, rt>θrefr_t> _ref, marking a forecast too unreliable to trust even weakly. At that step RACER spends one real evaluation to recompute the true feature, a refresh. It repays this evaluation by dropping a later scheduled step from a repayable subset R⊆UR U. At the dropped step it uses the shrinkage output instead. A debt counter tracks the unpaid refreshes, and RACER refreshes only while a repayable step still lies ahead. Every refresh therefore has a later step to pay it back. The per-prompt count of real evaluations stays exactly at |U||U|. A short offline trace fixes the trust scalars λ, β, θκ _κ and the refresh threshold θref _ref. It needs no training and does not modify the denoiser. A fixed regime rule uses the step coordinate for α≤0.75α≤ 0.75 and the logSNR coordinate for α≥3.0α≥ 3.0. The cross-model calibration protocol is specified in Experimental Setup. Theoretical Analysis Wan2.1-14B HunyuanVideo Method PSNR↑ SSIM↑ LPIPS↓ VBench-Quality↑ PSNR↑ SSIM↑ LPIPS↓ VBench-Quality↑ Reference – – – 83.13 – – – 84.49 α=0.75α=0.75 FORA 14.47 0.426 0.532 80.53 18.24 0.663 0.427 83.34 TaylorSeer 19.46 0.660 0.299 82.66 24.73 0.805 0.243 84.03 TeaCache 19.13 0.628 0.347 82.59 23.88 0.783 0.270 84.08 ToCa 16.88 0.537 0.410 82.18 21.26 0.732 0.357 83.79 Chebyshev 23.01 0.755 0.158 82.41 26.27 0.814 0.125 83.14 Ours 23.80 0.773 0.151 82.74 26.81 0.847 0.147 84.14 α=3.0α=3.0 FORA 13.02 0.353 0.598 80.82 17.29 0.594 0.475 83.35 TaylorSeer 17.24 0.585 0.367 81.38 22.27 0.740 0.303 83.40 TeaCache 16.53 0.557 0.390 81.85 21.42 0.725 0.334 81.71 ToCa 15.10 0.488 0.453 81.14 19.48 0.656 0.438 83.62 Chebyshev 21.85 0.694 0.208 80.90 24.42 0.767 0.173 82.84 Ours 22.77 0.729 0.215 82.14 25.01 0.799 0.202 83.68 α=5.0α=5.0 Chebyshev 21.60 0.666 0.272 79.90 23.78 0.747 0.231 80.77 Ours 22.34 0.707 0.248 81.66 23.81 0.752 0.228 81.31 α=7.0α=7.0 Chebyshev 21.22 0.637 0.316 79.10 21.82 0.635 0.436 80.01 Ours 21.87 0.673 0.292 80.77 22.67 0.685 0.328 80.12 Table 2: Main results of video generation on VBench, quality at matched NFE. Best in bold, second best underlined. We establish three complementary properties of the controller. They provide a deterministic error bound for shrinkage, an MSE-optimal target for trust, and exact per-prompt budget conservation. The shrinkage output interpolates the anchor and the forecast, so its error is controlled by the two endpoint errors. Proposition 1. Error bound. For any κ∈[0,1]κ∈[0,1], ‖y^t(κ)−ht‖≤κBf+(1−κ)Bh,\| y_t(κ)-h_t\|≤κ\,B_f+(1-κ)\,B_h, where hth_t is the true feature, BfB_f bounds the base-forecast error, and BhB_h bounds the drift from the anchor. The result is distribution-free. On our traces the bound holds at every step, and the measured error averages 0.932 of it (Figure 2). It also preserves the ordering of the steps by error, at a Spearman correlation of 0.997. The trust structure follows from an MSE-optimality argument. Proposition 2. MSE-optimal shrinkage. Treat the forecast and the anchor as two estimates of the true feature with finite error second moments. The trust that minimizes the expected squared error of y^t(κ) y_t(κ) over κ∈[0,1]κ∈[0,1] is κ∗=Π[0,1](σh2−ρfhσfσhσf2+σh2−2ρfhσfσh),κ^*= _[0,1]\! ( _h^2- _fh\, _f _h _f^2+ _h^2-2\, _fh\, _f _h ), where Π[0,1] _[0,1] denotes projection onto [0,1][0,1], σf2 _f^2 and σh2 _h^2 are their error second moments, and ρfh _fh is their normalized error correlation. This is the optimal combination of two correlated estimates (Wang et al. 2023), and it reduces to inverse-MSE weighting σh2/(σf2+σh2) _h^2/( _f^2+ _h^2) when the two errors are uncorrelated. As forecast risk grows relative to anchor risk, the optimal trust decreases. Equation (3) implements this structure with the horizon as a proxy for extrapolation risk and rtr_t as a proxy for forecast uncertainty. Its exponential–sigmoid map keeps trust bounded, while quantile calibration absorbs the signal scale. On our traces the true error rises with r within each horizon bucket, at rank correlations that are typically above 0.89. This runtime surrogate requires r to be a faithful proxy, which in turn depends on the observer being a valid check on the base. Write the disagreement as Δt=h^t−g^t _t= h_t- g_t. Two conditions make the check valid. First, a large disagreement forces a large forecast error efe_f, since ef≥‖Δt‖−Bo,e_f≥\| _t\|-B_o, (4) whenever the observer error is bounded by BoB_o. Second, a large forecast error must surface in the disagreement, since ‖Δt‖≥ef1−ρfo2,\| _t\|≥ e_f\, 1- _fo^2, (5) when the forecast and observer errors have correlation ρfo _fo, in the mean-square sense. A good observer therefore has a small error and a low correlation with the forecast. The Chebyshev and Taylor pair meets both. On the worst steps their error correlation runs from 0.17 to 0.44, which keeps 1−ρfo2 1- _fo^2 between 0.90 and 0.98. The refresh-and-repay policy provides a constructive exact-budget guarantee. It spends a real evaluation on a flagged step and later drops a scheduled one to pay it back. Proposition 3. Exact budget. For every prompt, the number of denoiser evaluations equals |U||U|. The debt grows only when a distinct repayable step still lies ahead, so it never exceeds the repayable steps that remain. Every refresh is thus matched to one later skipped step, and the count holds by construction. Across 20 configurations of 200 prompts each, it never deviated. Together these results bound the error, characterize an MSE-optimal target for κ, and fix the budget. Experiments Experimental Setup Tasks. We evaluate on four diffusion models, two for images and two for video. For image generation, we use SD3.5-Large (Esser et al. 2024) and FLUX.1-dev (Labs et al. 2025) at 1024×1024. For video generation, we use Wan2.1-14B (Wan et al. 2025) and HunyuanVideo (Kong et al. 2025). The reference for each model is its own 50-step sampler. We benchmark on DrawBench (Saharia et al. 2022) for image generation and VBench (Huang et al. 2023) for video generation. We also test generalization to a second image dataset, COCO (Lin et al. 2015), in Table 3. Baselines. The strongest open-loop baseline is the Chebyshev forecaster. We also compare against other baselines, FORA, TaylorSeer, TeaCache, and ToCa (Selvaraju et al. 2024; Liu et al. 2025b, a; Zou et al. 2025). Each baseline uses its published configuration. Evaluation metrics. We report reference fidelity using PSNR, SSIM (Wang et al. 2004), and LPIPS (Zhang et al. 2018) against the 50-step reference. We separately report ImageReward (Xu et al. 2023) and CLIP (Radford et al. 2021) for images, and VBench-Quality for videos. NFE measures denoiser computation, while measured latency and its speedup over the reference measure end-to-end speed. Implementation details. All methods are compared at a matched NFE. The per-prompt NFE of RACER exactly equals that of its base, so every paired comparison uses the same number of full denoiser evaluations. We sweep several acceleration regimes per model, from mild to aggressive. Each regime fixes a target NFE, while the controller form is shared across models. Image parameters are transferred from SD3.5 to FLUX, whereas video thresholds are calibrated per model and regime on a short trace. α Model Method 0.25 0.75 3.0 5.0 7.0 SD3.5-L Chebyshev 17.53 17.08 15.02 14.38 13.82 Ours 18.70 17.96 15.79 15.11 14.33 Δ ++1.17 ++0.88 ++0.77 ++0.73 ++0.51 FLUX.1 Chebyshev 23.74 23.34 21.42 20.29 19.49 Ours 23.83 23.53 21.49 20.82 19.86 Δ ++0.09 ++0.19 ++0.07 ++0.53 ++0.37 Table 3: Generalization to COCO, PSNR of the Chebyshev base and RACER. Δ is the gain of Ours over the base. The gain transfers across regimes and models. Figure 3: Quality versus compute on SD3.5. RACER tops the Chebyshev base at every NFE and achieves the highest quality. Main Results At matched NFE, RACER improves PSNR and SSIM over the strongest open-loop baseline on all four models (Table 1 and 2). The Chebyshev forecaster leads the other baselines on most metrics and regimes, and we treat it as the primary reference throughout. The PSNR and SSIM gains hold in every regime, while the task metrics are matched or improved relative to the Chebyshev base. LPIPS is the only metric with occasional losses. The gain also generalizes to a second image distribution (Table 3). We evaluate on COCO to test how well the gain generalizes. We compare against the Chebyshev base alone here. RACER raises the base in every regime. The quality margin converts to speed (Figure 3). For example, RACER runs at a measured 2.6 to 5.4× speedup over the 50-step sampler across the sweep on SD3.5. It tops the base at every point. α Metric Signal 0.25 0.75 3.0 7.0 PSNR↑ Ours 19.38 18.13 15.50 14.35 horizon 17.56 15.71 14.72 13.42 random 17.90 16.30 14.04 13.16 ImageReward↑ Ours 1.01 0.97 0.81 0.28 horizon 0.99 0.78 0.63 0.03 random 0.98 0.85 0.35 −-0.29 Table 4: Signal ablation on SD3.5. The executor is the refresh alone, with the NFE and the trigger budget held fixed, so only the driving signal changes. Ours is the cache disagreement r. α Method 0.25 0.75 3.0 7.0 base 18.35 17.85 15.70 14.39 refresh only ++1.03 ++0.28 −-0.20 −-0.04 shrinkage only ++0.32 ++0.31 ++0.25 ++0.31 Ours ++1.54 ++0.96 ++0.49 ++0.47 Table 5: Component ablation on SD3.5, under one configuration held fixed across regimes. The base row is the PSNR of the Chebyshev base, and the rows below it are the PSNR change over that base. Turning both responses on is best everywhere. Ablation Studies The disagreement outperforms horizon and random signals under a fixed refresh-only executor (Table 4). We swap only the driving signal. Replacing r with the horizon or with a random rule loses quality in every regime. Both responses contribute, and they divide the work by regime (Table 5). On images, turning both on is best in every regime. The refresh carries the mild regimes, and the shrinkage carries the aggressive ones, where the refresh alone no longer helps. The controller also transfers to a Taylor base (Table 6, Figure 4). Here, a Taylor-internal observer drives the same control policy. RACER improves PSNR by up to 3.83 dB in the aggressive regimes. α Model Method 0.25 0.75 3.0 7.0 SD3.5-L Taylor 17.66 16.92 12.98 10.49 refresh only 17.16 15.02 13.50 11.40 Ours 17.95 17.39 15.76 14.32 FLUX.1 Taylor 26.01 24.73 21.87 17.48 refresh only 25.97 24.40 20.84 16.65 Ours 26.16 25.71 22.63 20.92 Table 6: The RACER controller on a Taylor base with PSNR. Figure 4: The RACER controller placed on a Taylor base. The PSNR improvement over the Taylor base increases with acceleration. The Reliability Signal The disagreement identifies high-error forecasts (Figure 1). It flags the top-20% forecast-error steps at a mean AUROC of 0.94, against 0.77 for the input-side signals that prior caches rely on (Figure 1a). The separation holds on every model and across all five regimes (Figure 1b). Refreshes are concentrated at high-error steps (Figure 5). Wherever the budget allows, refreshes occur at high error ranks. The late steps carry no repayable step and stay forecast-only by design. Figure 5: The signal places compute where it is needed. Refreshes are concentrated at high error ranks when later repayment is possible. The Chebyshev and Taylor pair is a principled default (Figure 6). The signal applies when the observer has low self-error and weak error correlation with the base. Taylor and FoCa best balance these criteria, and Taylor-2 against Taylor-1 provides another effective pair. The disagreement is a signal, not a better forecast. Figure 6: Observer choice on SD3.5 at α=0.25α=0.25. Taylor and FoCa yield the largest gains among ten observers. We use Taylor as the standard forecaster. Conclusion This paper asked when and how much a diffusion feature forecast should be trusted. The answer sits in the cache itself. Two cheap forecasts agree on the easy stretches and split where prediction turns hard, and this split flags the risky steps without an extra denoiser evaluation. RACER closes the loop on this signal with a continuous trust shrinkage and an exact refresh-and-repay. At a matched NFE this improves the strongest open-loop base across four image and video generation models. On SD3.5, RACER also reaches equal quality faster. The mechanism continues to work when either the base or the observer is replaced. The signal scores how risky a step is, not which endpoint is better. The gain is also smaller where the base already tracks the reference closely. As forecasters grow stronger, the next lever for training-free acceleration may be reliable control under an exact budget. Acknowledgments We thank Zhehong Ai for helpful discussions. References D. Bolya and J. Hoffman (2023) Token merging for fast stable diffusion. External Links: 2303.17604, Link Cited by: Related Work. H. Cui, Z. Tang, Q. Ma, Z. Yao, and W. Jia (2026) Predict to skip: linear multistep feature forecasting for efficient diffusion transformers. External Links: 2602.18093, Link Cited by: Related Work. P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, D. Podell, T. Dockhorn, Z. English, K. Lacey, A. Goodwin, Y. Marek, and R. Rombach (2024) Scaling rectified flow transformers for high-resolution image synthesis. External Links: 2403.03206, Link Cited by: Experimental Setup. G. Fang, X. Ma, and X. Wang (2023) Structural pruning for diffusion models. In Thirty-seventh Conference on Neural Information Processing Systems, External Links: Link Cited by: Related Work. J. Han, J. Shi, P. Li, H. Ye, Q. Guo, and S. Ermon (2026) Adaptive spectral feature forecasting for diffusion sampling acceleration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 43320–43330. Cited by: Introduction, Related Work, Preliminaries. J. Ho, A. Jain, and P. Abbeel (2020) Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33, p. 6840–6851. External Links: Link Cited by: Introduction. Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, Y. Wang, X. Chen, L. Wang, D. Lin, Y. Qiao, and Z. Liu (2023) VBench: comprehensive benchmark suite for video generative models. External Links: 2311.17982, Link Cited by: Experimental Setup. W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, K. Wu, Q. Lin, J. Yuan, Y. Long, A. Wang, A. Wang, C. Li, D. Huang, F. Yang, H. Tan, H. Wang, J. Song, J. Bai, J. Wu, J. Xue, J. Wang, K. Wang, M. Liu, P. Li, S. Li, W. Wang, W. Yu, X. Deng, Y. Li, Y. Chen, Y. Cui, Y. Peng, Z. Yu, Z. He, Z. Xu, Z. Zhou, Z. Xu, Y. Tao, Q. Lu, S. Liu, D. Zhou, H. Wang, Y. Yang, D. Wang, Y. Liu, J. Jiang, and C. Zhong (2025) HunyuanVideo: a systematic framework for large video generative models. External Links: 2412.03603, Link Cited by: Experimental Setup. B. F. Labs, S. Batifol, A. Blattmann, F. Boesel, S. Consul, C. Diagne, T. Dockhorn, J. English, Z. English, P. Esser, S. Kulal, K. Lacey, Y. Levi, C. Li, D. Lorenz, J. Müller, D. Podell, R. Rombach, H. Saini, A. Sauer, and L. Smith (2025) FLUX.1 kontext: flow matching for in-context image generation and editing in latent space. External Links: 2506.15742, Link Cited by: Experimental Setup. B. Lakshminarayanan, A. Pritzel, and C. Blundell (2017) Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in neural information processing systems 30. Cited by: Preliminaries. M. Li, T. Cai, J. Cao, Q. Zhang, H. Cai, J. Bai, Y. Jia, M. Liu, K. Li, and S. Han (2024) DistriFusion: distributed parallel inference for high-resolution diffusion models. External Links: 2402.19481, Link Cited by: Related Work. X. Li, Y. Liu, L. Lian, H. Yang, Z. Dong, D. Kang, S. Zhang, and K. Keutzer (2023) Q-diffusion: quantizing diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 17535–17545. Cited by: Related Work. T. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Dollár (2015) Microsoft coco: common objects in context. External Links: 1405.0312, Link Cited by: Experimental Setup. Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le (2023) Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, External Links: Link Cited by: Related Work. F. Liu, S. Zhang, X. Wang, Y. Wei, H. Qiu, Y. Zhao, Y. Zhang, Q. Ye, and F. Wan (2025a) Timestep embedding tells: it’s time to cache for video diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 7353–7363. Cited by: Introduction, Related Work, Experimental Setup. J. Liu, C. Zou, Y. Lyu, J. Chen, and L. Zhang (2025b) From reusing to forecasting: accelerating diffusion models with taylorseers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 15853–15863. Cited by: Introduction, Related Work, Experimental Setup. X. Liu, C. Gong, and Q. Liu (2022) Flow straight and fast: learning to generate and transfer data with rectified flow. External Links: 2209.03003, Link Cited by: Related Work. C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu (2022) DPM-solver: a fast ODE solver for diffusion probabilistic model sampling in around 10 steps. In Advances in Neural Information Processing Systems, A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho (Eds.), External Links: Link Cited by: Related Work. S. Luo, Y. Tan, L. Huang, J. Li, and H. Zhao (2023) Latent consistency models: synthesizing high-resolution images with few-step inference. External Links: 2310.04378, Link Cited by: Related Work. X. Ma, G. Fang, and X. Wang (2024) DeepCache: accelerating diffusion models for free. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 15762–15772. Cited by: Introduction, Related Work, Preliminaries. W. Peebles and S. Xie (2023) Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, p. 4195–4205. Cited by: Introduction. A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever (2021) Learning transferable visual models from natural language supervision. External Links: 2103.00020, Link Cited by: Experimental Setup. R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer (2022) High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 10684–10695. Cited by: Introduction. C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al. (2022) Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems 35, p. 36479–36494. Cited by: Experimental Setup. T. Salimans and J. Ho (2022) Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations, External Links: Link Cited by: Related Work. P. Selvaraju, T. Ding, T. Chen, I. Zharkov, and L. Liang (2024) FORA: fast-forward caching in diffusion transformer acceleration. External Links: 2407.01425, Link Cited by: Introduction, Preliminaries, Experimental Setup. H. S. Seung, M. Opper, and H. Sompolinsky (1992) Query by committee. In Proceedings of the fifth annual workshop on Computational learning theory, p. 287–294. Cited by: Preliminaries. J. Song, C. Meng, and S. Ermon (2021) Denoising diffusion implicit models. In International Conference on Learning Representations, External Links: Link Cited by: Related Work. L. N. Trefethen (2019) Approximation theory and approximation practice, extended edition. SIAM. Cited by: Related Work. T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, J. Wang, J. Zhang, J. Zhou, J. Wang, J. Chen, K. Zhu, K. Zhao, K. Yan, L. Huang, M. Feng, N. Zhang, P. Li, P. Wu, R. Chu, R. Feng, S. Zhang, S. Sun, T. Fang, T. Wang, T. Gui, T. Weng, T. Shen, W. Lin, W. Wang, W. Wang, W. Zhou, W. Wang, W. Shen, W. Yu, X. Shi, X. Huang, X. Xu, Y. Kou, Y. Lv, Y. Li, Y. Liu, Y. Wang, Y. Zhang, Y. Huang, Y. Li, Y. Wu, Y. Liu, Y. Pan, Y. Zheng, Y. Hong, Y. Shi, Y. Feng, Z. Jiang, Z. Han, Z. Wu, and Z. Liu (2025) Wan: open and advanced large-scale video generative models. External Links: 2503.20314, Link Cited by: Experimental Setup. X. Wang, R. J. Hyndman, F. Li, and Y. Kang (2023) Forecast combinations: an over 50-year review. International Journal of Forecasting 39 (4), p. 1518–1547. Cited by: Theoretical Analysis. Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli (2004) Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13 (4), p. 600–612. Cited by: Experimental Setup. J. Xu, X. Liu, Y. Wu, Y. Tong, Q. Li, M. Ding, J. Tang, and Y. Dong (2023) ImageReward: learning and evaluating human preferences for text-to-image generation. External Links: 2304.05977, Link Cited by: Experimental Setup. R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang (2018) The unreasonable effectiveness of deep features as a perceptual metric. External Links: 1801.03924, Link Cited by: Experimental Setup. W. Zhao, L. Bai, Y. Rao, J. Zhou, and J. Lu (2023) UniPC: a unified predictor-corrector framework for fast sampling of diffusion models. In Thirty-seventh Conference on Neural Information Processing Systems, External Links: Link Cited by: Related Work. S. Zheng, L. Feng, X. Wang, Q. Zhou, P. Cai, C. Zou, J. Liu, Y. Lin, J. Chen, Y. Ma, et al. (2026) Forecast then calibrate: feature caching as ode for efficient diffusion transformers. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 13449–13457. Cited by: Related Work. Z. Zheng, X. Wang, C. Zou, S. Wang, and L. Zhang (2025) Compute only 16 tokens in one timestep: accelerating diffusion transformers with cluster-driven feature caching. In Proceedings of the 33rd ACM International Conference on Multimedia, p. 10181–10189. Cited by: Related Work. X. Zhou, D. Liang, K. Chen, T. Feng, X. Chen, H. Lin, Y. Ding, F. Tan, H. Zhao, and X. Bai (2025) Less is enough: training-free video diffusion acceleration via runtime-adaptive caching. External Links: 2507.02860, Link Cited by: Related Work. C. Zou, X. Liu, T. Liu, S. Huang, and L. Zhang (2025) Accelerating diffusion transformers with token-wise feature caching. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: Related Work, Experimental Setup.