Paper deep dive
$π\mathbf{R}^2$: Reactive Real-time Flow Policies
Sungjae Park, Shubham Tulsiani
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing \emph{reactivity}. Replanning more often would restore it, but the perception-to-action pipeline (a large backbone plus multiple denoising steps) is too slow: this \emph{latency} forbids frequent replanning and leaves committed actions stale, making such policies ill-suited for dynamic, closed-loop control. We present $\pi\mathbf{R}^2$, which makes these policies reactive and real-time while retaining large backbones, expressive multi-modal policies, and multi-action prediction. Built on the per-position noise schedule of diffusion forcing, $\pi\mathbf{R}^2$ contributes two ideas. First, it splits conditioning into a fast channel (proprioception, fresh every tick) and an asynchronously updated slow channel (vision-language features), so the policy reacts to proprioception within a chunk while tolerating stale vision. Second, a latency-adaptive flow schedule treats in-flight actions as inpainting conditioning and emits actions in one denoising step per call, letting one trained model adapt to varying hardware latency. Requiring minimal modification to existing architectures, $\pi\mathbf{R}^2$ can be finetuned from a pretrained policy: applied to GR00T-N1.7 on a real xArm6+XHand platform, it replans closed-loop roughly $4\times$ faster than the base policy (~$25$Hz on an A5000 GPU), acting on a fresh observation every $40$ms. Across simulation and real-world manipulation tasks, $\pi\mathbf{R}^2$ improves the success rate by up to $23\%$ in simulation and $30\%$ in the real world over the strongest baseline. Project page: this https URL
Tags
Links
- Source: https://arxiv.org/abs/2607.26055v1
- Canonical: https://arxiv.org/abs/2607.26055v1
Trouble viewing inline? Open PDF directly →
Full Text
71,282 characters extracted from source content.
Expand or collapse full text
π2 ^2: Reactive Real-time Flow Policies Sungjae Park, Shubham Tulsiani Carnegie Mellon University sungjae2, stulsian@andrew.cmu.edu Abstract Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing reactivity. Replanning more often would restore it, but the perception-to-action pipeline (a large backbone plus multiple denoising steps) is too slow: this latency forbids frequent replanning and leaves committed actions stale, making such policies ill-suited for dynamic, closed-loop control. We present π2 ^2, which makes these policies reactive and real-time while retaining large backbones, expressive multi-modal policies, and multi-action prediction. Built on the per-position noise schedule of diffusion forcing, π2 ^2 contributes two ideas. First, it splits conditioning into a fast channel (proprioception, fresh every tick) and an asynchronously updated slow channel (vision-language features), so the policy reacts to proprioception within a chunk while tolerating stale vision. Second, a latency-adaptive flow schedule treats in-flight actions as inpainting conditioning and emits actions in one denoising step per call, letting one trained model adapt to varying hardware latency. Requiring minimal modification to existing architectures, π2 ^2 can be finetuned from a pretrained policy: applied to GR00T-N1.7 on a real xArm6+XHand platform, it replans closed-loop roughly 4×4× faster than the base policy (2525 Hz on an A5000 GPU), acting on a fresh observation every 4040 ms. Across simulation and real-world manipulation tasks, π2 ^2 improves the success rate by up to 23%23\% in simulation and 30%30\% in the real world over the strongest baseline. Project page: https://pi-r2-flow.github.io/ Keywords: Reactivity, Real-time Inference, Flow Policies, Robot Foundation Models 1 Introduction We have witnessed remarkable progress in training ‘robotics foundation models’ [5, 15, 4, 8, 42, 20], and three design choices have become ubiquitous across recent efforts. First, large pre-trained backbones such as vision-language models (VLMs) [23, 1, 2, 30, 3, 12] are leveraged to process visual (and language) input and extract rich, generalizable representations for downstream action prediction. Second, expressive policy architectures such as diffusion [10] and flow matching [34, 26] allow effectively capturing the multi-modal action distributions inherent in diverse human demonstrations. Finally, action chunking [39, 38] enables these models to jointly predict multiple actions, providing a richer training signal and improving the temporal consistency of executed actions. Figure 1: Overview of π2 ^2. Top. While standard diffusion/flow matching relies on stale observations to predict actions via iterative denoising, π2 ^2 disentangles the observation into a fast channel (proprioception) and a slow channel (image encoding, VLM embedding, etc.), and uses up-to-date observations for each denoising step. Bottom. To incorporate latency and smooth execution, we adopt an adaptive noise schedule, where the action chunk is divided into three regions: clean actions to be taken during inference, actions with increasing noise level following Diffusion Forcing [9], and pure noise with the same length as clean actions. Compared to Train-Time RTC, this enables faster inference and smoother actions. On the one hand, these choices have been crucial in unlocking generalizable and scalable imitation learning. However, they result in limited reactivity and increased latency, making these foundation models ill-suited for dynamic manipulation tasks that demand fast, smooth, and reactive control. Specifically, the execution of a ‘chunk’ of predicted actions is performed open-loop, without the ability to react to the sensory input streaming in. While the reactivity can in theory be improved by only executing a small ‘sub-chunk’ (in the limit just one action) and re-predicting with updated sensory input, naively doing so is infeasible due to the significant computation required in the ‘perception-to-action’ prediction – a large visual (and language) processing backbone followed by multiple denoising iterations for flow matching policy inference. This computational bottleneck coupled with action chunking flow matching thus results in two (related) fundamental limitations: a) reduced reactivity, as a large (sub-)chunk of actions must be executed open-loop while the next chunk is predicted, and b) increased latency, as the predicted actions are a function of ‘stale’ sensory input. In this work, we develop π2 ^2: a framework that allows reactive real-time policy inference without compromising on the design principles of large pre-trained backbones, expressive multi-modal policies, and multi-action prediction. We note that instead of flow (or diffusion) policy inference, which maps uniformly random noise to action chunks over multiple denoising iterations, leveraging the flexible schedule in ‘diffusion forcing’ [9] offers several benefits: a) the immediate actions to be executed can be at a lower noise level and predicted in fewer iterations, and b) different denoising iterations involved in predicting an action can rely on updated sensory input as available. Although prior work has explored diffusion forcing for policy learning [14], we address the key issues that prevent its widespread adoption for robotics foundation models. First, while the (immediate) action denoising requires less computation, the backbone (e.g., VLM) feature processing becomes the bottleneck for ‘perception-to-action prediction’ latency. Our key insight that allows us to circumvent this bottleneck is that not all input modalities need to be processed at the same rate: proprioceptive signals (joint positions, velocities, torques, and contact forces) can be retrieved and processed orders of magnitude faster than images or text. Moreover, for dynamic tasks, proprioception carries sufficient information for local reactive corrections (e.g., responding to an unexpected contact or compensating for object slip), while vision and language provide the global context needed to guide coarse motion plans. We thus disentangle the conditioning for action denoising into asynchronously updated ‘slow’ and always up-to-date ‘fast’ features, effectively allowing high-frequency proprioceptive control within an action chunk. Second, despite our asynchronous diffusion forcing, practical constraints (communication delays, GPU bandwidth) often prevent zero latency. We develop a delay-adaptive noise schedule for diffusion forcing that enables seamless real-time execution despite latency, allowing the model to continue executing the ‘previous’ actions while ensuring smooth temporally coherent outputs from subsequent denoising iterations. Together, these design choices make π2 ^2 roughly 4×4× faster than the base policy for closed-loop replanning, achieving 25 Hz with a large-scale VLA (GR00T N1.7 [4]) on A5000 GPUs. Crucially, this means the policy predicts each action using fresh observation every 40 ms, rather than committing to a long action chunk and replanning only sparsely (e.g., at ∼ 7 Hz). π2 ^2 thus matches the reactivity of simple feedforward policies while retaining the expressivity of flow matching, the training benefits of chunk-based inference, and the semantic grounding of pretrained backbones. We first validate π2 ^2 in simulation where we highlight the benefits of increased reactivity and reduced latency compared to the typical action chunking flow policies. We also show how, instead of training from scratch, we can finetune pre-trained VLAs with a diffusion forcing schedule and demonstrate the efficacy of π2 ^2 for complex real-world manipulation tasks, where we find that it improves over the strongest baseline by up to 23%23\% in simulation and 30%30\% in the real world. 2 Related Work Action Chunking and Flow Policies. Predicting a chunk of future actions rather than a single action, introduced by ACT [39], improves the temporal coherence of imitation-learned policies and has since become standard, including for mobile and contact-rich manipulation [13, 40]. This design has been widely adopted by expressive multi-modal policies based on diffusion [10] and flow matching [37, 32]. For dynamic tasks, however, committing a (sub-)chunk open-loop is suboptimal, as it cannot react to sensory input arriving mid-execution. While our work also reasons over a chunk of actions, its use of diffusion forcing predicts the immediate actions with lower latency and restores reactivity by conditioning successive denoising iterations on progressively fresher sensory input. Generalist Manipulation Policies. Foundation robotics models also couple chunked flow-matching action heads with large vision-language backbones, yielding vision-language-action (VLA) models that scale imitation learning across diverse tasks and embodiments [11, 17, 42, 19]. Notable examples include π0 _0/π0.5 _0.5 and the dual-system GR00T N1, whose vision-language module feeds a flow-matching action head conditioned on robot proprioception [5, 15, 25]. While this has yielded impressive, generalizable policies, coupling a large backbone with multi-step flow-matching denoising only exacerbates the reactivity problem: the combined per-call latency limits how often the policy can replan. Crucially, this backbone latency would persist even under a diffusion-forcing schedule, since each emitted action still requires a full forward pass through the slow semantic backbone. We instead exploit an asymmetry these models already expose—proprioception is far cheaper to process than vision and language—refreshing proprioceptive conditioning at every control tick while letting the slow vision-language features update asynchronously, making this same class of models reactive. Real-time Execution of Action-Chunking Policies. The high inference latency of modern VLAs has motivated methods for executing action chunks smoothly while the next prediction is computed [31]. RTC maintains continuity across asynchronous chunks via inference-time inpainting on a frozen action prefix, while Training-Time RTC instead folds this prefix conditioning into training, avoiding the extra inference-time cost [6, 7]. Streaming Diffusion Policy and Streaming Flow Policy emit actions incrementally from a diffusion-forcing-style variable-noise buffer maintained across observations [14, 16]. None, however, address the latency of repeatedly conditioning on a large semantic backbone; Streaming Diffusion Policy, for instance, uses compact visuomotor policies and runs synchronously. Concurrent to our work, FASTER [24] similarly reduces reaction latency in flow-based VLAs via a diffusion-forcing-inspired horizon-aware schedule with action-prefix conditioning, but forwards the backbone once per chunk and conditions all actions on a single fixed observation, improving latency and smoothness without improving reactivity to sensory input arriving during execution. π2 ^2 instead refreshes proprioceptive conditioning within the chunk, reacting to incoming feedback as actions execute while vision-language features update asynchronously. 3 π2 ^2 Our goal is to modify existing large-scale flow policies for higher reactivity and lower inference latency. We first review the building blocks – flow matching and diffusion forcing – and then introduce two modifications that together yield a real-time, closed-loop flow policy. The modifications require only a flow matching action head conditioned on a heavy pretrained backbone, so they apply equally to world-action models [33, 18, 21, 41]; we instantiate and evaluate on VLAs, the most common member of this family. 3.1 Preliminary 3.1.1 Flow Matching and Action Chunking Policy We build on flow matching [22, 36] with action chunking [39, 38]. Flow matching learns a velocity field vθ(t,t)v_θ(x_t,t) that transports samples from noise p0=(,)p_0=N(0,I) to data p1=pdatap_1=p_data along the conditional interpolation t=(1−t)ϵ+t1x_t=(1-t)\, ε+t\,x_1, t∈[0,1]t∈[0,1], trained against the conditional velocity ut(t∣1)=1−ϵu_t(x_t _1)=x_1- ε: ℒFM=t,1,ϵ[‖vθ(t,t)−(1−ϵ)‖2].L_FM=E_t,x_1, ε\! [\|v_θ(x_t,t)-(x_1- ε)\|^2 ]. (1) In the action-chunking policy setting, 1=(t,…,t+H)x_1=(a_t,…,a_t+H) is a chunk of H future actions conditioned on observation to_t, generated jointly by K denoising steps; the robot executes the first h≤Hh≤ H actions before re-planning. Crucially, all H positions share a single noise level t and the chunk is committed open-loop: the h executed actions never account for new observations during their execution window, limiting the reactivity that manipulation tasks demand. 3.1.2 Diffusion Forcing and Streaming Diffusion Diffusion forcing [9, 28] generalizes flow matching by assigning each chunk position p∈0,…,H−1p∈\0,…,H-1\ an independent noise level τp∈[0,1] _p∈[0,1]: τ,p=(1−τp)ϵp+τpp,x_τ,p=(1- _p)\, ε_p+ _p\,a_p, (2) and the model vθ(τ,,)v_θ(x_τ, τ,o) predicts a per-position velocity. The flexibility of position-dependent noise levels enables closed-loop action prediction by conditioning successive denoising steps on progressively more recent observations within a single chunk. Streaming diffusion [14] is one such instantiation: it imposes a linearly increasing noise schedule across the chunk so that the leading positions are fully denoised within a few steps, and each denoising step incorporates a fresh observation. In this work, we also adopt diffusion forcing, focusing on an underexplored regime: enabling large-scale flow policies (e.g., VLAs) to be more reactive and closed-loop, where large latency naturally arises from its model size and components. The following two subsections present the architectural and scheduling modifications that make diffusion forcing a practical real-time, reactive policy on top of any flow matching action head. 3.2 Proprioception-Reactive Diffusion Forcing A VLA processes the observation t=(t,t,t)o_t=(s_t,I_t,T_t) – proprioception ts_t (joint positions, end-effector pose, etc.), image tI_t, and language tT_t – to predict an action chunk (t,…,t+H)(a_t,…,a_t+H). Internally, the inputs are preprocessed (dominated by image processing) and passed through the large backbone(e.g., VLM), producing language-aligned features; proprioception is processed by a small MLP into a state embedding; the DiT action head then conditions on the concatenated representation and runs K denoising steps over the chunk. In practice, these components differ greatly in cost. On RTX A6000 with GR00T-N1.7, image preprocessing + VLM (∼60 60 ms) plus DiT denoising (K=4K=4 steps, ∼80 80 ms) sum to ∼140 140 ms per call – ∼7 7 control ticks at 5050 Hz. This motivates disentangling the slow VLM path from the fast action prediction loop. We propose to disentangle the DiT’s conditioning (Fig. 1, left) into a slow channel (vision and language features from the VLM and image/text encoders) and a fast channel (proprioception, fresh every tick). This split mirrors the structure of manipulation itself: vision and language provide coarse spatial and task guidance, while fresh proprioception drives the fine motor refinement that precision demands. Diffusion forcing additionally reduces the per-call denoising work: because the front of the chunk is kept near τ=1τ=1, a single denoising step per call suffices to produce clean actions at the front (one or more, depending on the schedule; see Sec. 3.3), versus K steps in standard chunked denoising. Combining the two, the DiT action head runs one denoising step against fresh proprioception and a cached slow feature; vision and language are processed asynchronously in a background thread (image preprocessing + VLM forward) and refresh the cache when complete, so the action head conditions on a slow feature of age dvlmd_vlm ticks (the vision/text wall-clock latency). Per-call cost of the action head is therefore one NFE only, independent of backbone size. To absorb the resulting slow-channel staleness, we delay the slow channel at training time by dvlm∼Uniform0,…,dvlmmaxd_vlm \0,…,d_vlm \ and supply the delay value through a learned embedding indexed by the integer delay, added to the slow representation; at deploy time the measured dvlmd_vlm uses the same embedding. The resulting policy is proprioception-reactive: it reacts to fresh joint state every denoising step while tolerating bounded staleness on vision and language, recovering closed-loop control even behind a large VLA backbone. We refer to this architecture variant as Proprioception-Reactive Diffusion Forcing. 3.3 Latency-Adaptive Flow Schedule By combining diffusion forcing with asynchronous vision/text processing and single-step DiT inference, the per-call delay d of the action prediction loop (distinct from the vision-language staleness dvlmd_vlm of Sec. 3.2) reduces to one denoising step of the action head. However, in reality, d varies: it depends on each component size and inference GPU, includes network latency when policy and robot run on separate machines (e.g., Ethernet round-trips in remote-inference setups), and fluctuates with computational load and per-call jitter. The policy must therefore handle a range of inference delays d∈[1,dmax]d∈[1,d_max]. A naive linearly increasing noise schedule, as in standard streaming diffusion [14], does not address this regime: it implicitly assumes d=0d=0 (instantaneous inference, no in-flight actions), so when applied at d>0d>0 the returning chunk can jump at the chunk boundary [6, 29]. We address both issues with a per-position noise schedule parameterized by the inference delay d: a clamped-clean front carries the d in-flight actions as inpaint conditioning, a ramped interior emits d clean actions per call at the cost of one denoising step, and randomizing d at training lets a single model adapt to whatever value is measured at each call. 03366991111011pτp _p(a) start of call03366991111011emit d new clean actionsppτp _p(b) after one substep03366991111011fresh noiseppτp _p(c) after slide Figure 2: Inference cycle, one call. (a) Buffer at τ⋆,dτ ,d at the start of the call. (b) One Euler substep (one NFE) shifts the schedule right by d slots: positions [d,2d)[d,2d) reach τ=1τ=1 and are released, and the ramp extends through the back (c) The buffer slides d positions: just-emitted actions become the new front conditioning, d fresh-noise slots (τ=0τ=0) are appended – the schedule reproduces exactly. Staircase schedule (Fig. 1, right). For a target inference delay d, π2 ^2 uses the three-region staircase ⋆,d∈[0,1]H τ ,d∈[0,1]^H τp⋆,d=10≤p<d1−p−dH−2d≤p<H−d0H−d≤p≤H−1τ ,d_p= cases1&0≤ p<d\\[2.0pt] 1- p-dH-2d&d≤ p<H-d\\[6.0pt] 0&H-d≤ p≤ H-1 cases (3) with interior slope s=1/(H−2d)s=1/(H-2d). Three roles by position: front [0,d)[0,d), the in-flight actions, clamped clean as inpaint conditioning; interior [d,H−d)[d,H-d), a linear ramp from clean to noise; tail [H−d,H)[H-d,H), the d pure-noise slots appended at the back per cycle. This staircase can be viewed as a diffusion-forcing generalization of training-time RTC [7]: both clamp a clean front of d in-flight actions for inpaint conditioning, but RTC applies a single shared noise level over the remaining H−dH-d positions, whereas the staircase replaces it with a ramped interior plus a pure-noise tail to enable single-step emission under diffusion forcing. Training. At each batch we sample d∼Uniform1,…,dmaxd \1,…,d_max\ and build ⋆,d τ ,d. The front d slots are filled by the ground-truth actions and excluded from the loss via mp=[p≥d]m_p=1[p≥ d] (mirroring inference-time inpaint conditioning); the interior and tail are noised at the staircase levels and contribute to the per-position MSE. We additionally apply symmetric jitter τp←clip(τp+δp,0,1) _p ( _p+ _p,0,1) with δp∼Uniform[−j,j] _p [-j,j] to absorb small deviations from the central staircase that arise from per-call d variation. With probability p=0.2p=0.2, we instead train on a standard flow schedule (single τ∼Uniform[0,1]τ [0,1] shared across positions, no mask), so the same network can also denoise a full chunk from pure noise, used at inference to warm-start the buffer at episode start. Compatibility with existing flow policy action heads. The procedure above plugs into any DiT-style flow matching action with a one-line architectural change: the AdaLN conditioning (which modulates each DiT block by the noise level τ) becomes per-position, with one (γp,βp)( _p, _p) pair per chunk position rather than one shared across positions. All other components (attention, MLP, pretrained large backbone, slow/fast feature paths) remain unchanged, so existing pretrained policies, such as VLAs, can be fine-tuned with π2 ^2 without touching the backbone. Inference cycle (Fig. 2). At episode start we warm up the buffer with standard-flow inference (denoising all H positions from pure noise) and re-noise to match ⋆,d τ ,d for the initial measured d. From then on, each call applies one denoising step that updates each position by p←p+Δτp⋅vθ(,,)p,τp←τp+Δτp,x_p\;←\;x_p+ _p· v_θ(x, τ,o)_p, _p\;←\; _p+ _p, (4) with region-dependent per-position advances Δτp _p chosen so the schedule shifts right by d slots: positions [d,2d)[d,2d) reach τ=1τ=1 and are emitted, and the rest of the buffer rotates forward (Fig. 2b). The buffer then slides by d positions, with d fresh-noise slots (τ=0τ=0) appended at the back (Fig. 2c). The schedule reproduces exactly when d holds steady; when d changes between calls the Δτp _p pattern adapts to the new measured d, and the buffer is pulled toward the new ⋆,d τ ,d over a few calls. 4 Experiments 4.1 Simulation Experiments Setup. We first test π2 ^2 on the Leap Cube Reorientation task in MuJoCo Playground [35], a common task that requires reactivity to keep the cube from falling and to achieve the target cube pose. Demonstrations are collected at 5050 Hz from 44 RL experts (different seeds), yielding 200200 trajectories total. All methods are trained as state-based flow policies, with prediction chunk length H=16H=16. We run two studies: an execution-horizon h sweep under zero inference delay (Sec. 4.1.1), and a deployment study under a fixed latency budget (Sec. 4.1.2). Figure 3: Simulation results. Left. Without any inference delay, executing a smaller number of actions within the chunk benefits the performance for a flow matching policy. π2 ^2 also achieves the same performance via replanning every timestep, while the only last denoising step is conditioned on up-to-date state. Right. When there is an inference delay for each component, π2 ^2 reduces the effective delay d by reducing the number of denoising steps and asynchronously processing the visual features. For each datapoint, d indicates the effective delay of proprioception, while dvd_v is visual delay. We train 3 policies with different seeds for each method and report the mean and std of the success rate. 4.1.1 Execution-horizon (h) sweep without inference delay We first study the setting where inference delay d=0d=0, to isolate two things: (i) the tradeoff between task performance and execution horizon h (executing more actions open-loop misses the recent state), and (i) whether π2 ^2’s amortized denoising, where each action’s final denoising step conditions on the most recent observation, matches the reactivity of h=1h=1 flow at a fraction of the per-call cost. At evaluation, standard flow matching uses 1616 denoising steps to predict the whole chunk, whereas π2 ^2 uses one denoising step and emits one action at a time. We sweep h∈1,2,4,8h∈\1,2,4,8\ for standard flow; π2 ^2 emits one action per call (h=1h=1, 11 NFE per call). Results. Results are highlighted in Fig. 3 (left). Intuitively, the performance of standard flow degrades as h grows, confirming the h tradeoff for reactive tasks. Meanwhile, π2 ^2 matches the flow configuration of standard flow h∈1,2h∈\1,2\. This supports that leveraging fresh observation more frequently is beneficial for such reactive tasks, and amortizing the denoising budget across calls does not sacrifice quality, even though each call costs only one NFE. We next introduce realistic latency and test each deployment regime under that constraint. 4.1.2 Deployment under VLA-level inference delay To consider a realistic latency, we take into account GR00T-N1.7 computation cost (Sec. 3.2: vision/text VLM processing ≈60≈ 60 ms, K=4K=4 NFE denoising ≈80≈ 80 ms) and define unit delay d0d_0 as the K=4K=4 denoising cost, resulting in vision/text processing ≈0.75d0≈ 0.75\,d_0 and one NFE =d0/4=d_0/4. Baselines. We compare four methods: (i) naive-async runs the policy inference for predicting the new chunk while executing previous chunk. The inference does not take into account the actions being executed during inference. (i) train-time RTC [7] explicitly conditions the in-flight actions while predicting the new chunk, resulting in smooth chunk boundaries. (i) π2 ^2 without asynchronous fast/slow channel processing only applies latency-adaptive flow schedule (Sec. 3.3) without asynchronously processing the fast and slow channel, and (iv) π2 ^2 with asynchronous processing applies all proposed modifications. The effective per-call delays are then 1.75d01.75\,d_0 (end-to-end inference: naive-async and train-time RTC), 1.0d01.0\,d_0 (π2 ^2 w/o async), and 0.25d00.25\,d_0 (π2 ^2 w async). Setting d0∈1,2,3d_0∈\1,2,3\ control ticks and ceil-rounding gives effective proprioception delay d∈2,4,6/1,2,3/1,1,1d∈\2,4,6\/\1,2,3\/\1,1,1\, respectively. For π2 ^2 with async processing, we additionally delay the object pose by dvis∈1,2,3.d_vis∈\1,2,3\. For all methods, we set execution horizon h=dh=d across all methods. Results. Fig. 3 (right) reports the mean over three training seeds (error bars show std); for each cell we pick the best epoch per seed. π2 ^2 w/ async wins at every d0d_0 (0.430.43, 0.420.42, 0.450.45), beating naive-async (0.330.33, 0.290.29, 0.220.22) and Train-time RTC (0.360.36, 0.320.32, 0.190.19) with margins that widen as delay grows. The reason is effective delay: baselines pay d=⌈1.75d0⌉d= 1.75\,d_0 per call while π2 ^2 pays d=d0d=d_0 (w/o async) or d=1d=1 (w/ async). Async processing lets us hold the action delay at 11 while only the visual delay dvisd_vis grows, so performance degrades gracefully with dvisd_vis – supporting the view that vision provides coarse guidance while up-to-date proprioception drives manipulation. 4.2 Real World Experiments Figure 4: Real-world manipulation tasks. Arrows in each image show the desired task outcome. Table 1: Real World Results. (xArm6 + XHand, fine-tuned GR00T N1.7). d is the measured wall-clock inference delay in 2525-Hz control ticks (11 tick ≈40≈ 40 ms), which is the same as execution horizon h for all methods other than synchronous inference. For each task we report success rate (SR) over N trials and a partial-progress score (Prog) capturing the fraction of subgoals completed per episode. Don’t Spill Tidy up Book Insert Box Catch Book Setting SR Prog SR Prog SR Prog SR Flow, Synchronous (h=10h=10) 4/20 16/80 4/20 9/40 11/20 56/80 4/20 Flow, Naive Async, dense, TE [39] 7/20 30/80 7/20 15/40 12/20 61/80 2/20 Flow, Train-Time RTC [7] 9/20 45/80 8/20 18/40 10/20 53/80 5/20 π2 ^2 10/20 55/80 12/20 24/40 16/20 68/80 11/20 We now mirror the deployment study above on a real xArm6 + XHand setup by fine-tuning GR00T-N1.7 model [25] for dexterous manipulation. Demos are collected at 2525 Hz (capped by the teleoperation recording pipeline), so we report d in 2525-Hz ticks (11 tick ≈40≈ 40 ms); We use RTX A5000 for both training and inference, resulting in d=4∼5d=4 5 for baselines and d=1∼2d=1 2 for π2 ^2. We include arm and hand joint angles, along with hand per-joint torque and fingertip forces in the proprioceptive state. Tasks. We evaluate our method on four contact-rich dexterous, reactive manipulation tasks, shown in Fig. 4. Don’t Spill requires placing a ball into a bowl and moving the bowl onto a cutting board without dropping the ball. Tidy Up Book requires extracting a book from a pile, grasping it, and placing it into a basket. Insert Box requires pushing a box against a wall to stand it upright before inserting it between a row of books. Catch Book requires the robot reacting to the falling book (induced by the ball hitting the book) and grasping it without dropping it. All four tasks demand highly reactive behavior. Picking up the ball requires dynamic grasping, moving the bowl without tilting it relies on proprioceptive feedback, and both extracting a book from a pile and pushing a box upright require precise contact-rich interactions and reactive control. Lastly, catching a book requires the robot to quickly react so that the book stays within its palm. We collect 200, 300, 300, and 100 teleoperated demonstrations for each task, respectively. Baselines. We compare π2 ^2 against three baselines, all fine-tuned from GR00T-N1.7 with same training budget. (1) Flow, Synchronous stalls the robot for the full d=4∼5d=4 5 ticks while inference runs, then executes the predicted chunk (h=10h=10). (2) Flow, Naive Async (dense, TE) [40] runs the policy asynchronously without conditioning on in-flight actions: dense re-issues a query every d ticks (whenever the previous call returns), and TE (temporal ensembling) [39] for overlapping action predictions for the same timestep. (3) Flow, Train-Time RTC [7] additionally trains the policy to inpaint a clean d-action front, anchoring the new chunk to the executed history at the chunk boundary. Results. Table 1 summarizes the experimental results. π2 ^2 with asynchronous vision–language inference operates near the per-control-tick limit (d=1d=1 at 2525 Hz, with occasional increases to d=2d=2 under network delays), whereas all flow-based baselines incur the full GR00T pipeline latency of d∈4,5d∈4,5. Synchronous inference produces jittery motion at chunk boundaries because the robot pauses while waiting for the next action chunk, leading to failures such as the ball rolling off the table. Naive asynchronous inference combined with temporal ensembling yields smoother motion but sacrifices execution precision. Train-Time RTC [7] is the strongest baseline; however, π2 ^2 outperforms it across all metrics, with the largest gains on reactivity-critical tasks such as Tidy Up Book, Insert Box, and Catch Book, achieving approximately 20∼3020 30% absolute improvement. We observe that RTC also struggles to recover from failures because previously generated in-flight actions bias the policy toward continuing its recent motion, which reduces its ability to react promptly, particularly in Insert Box. Qualitatively, π2 ^2 produces smoother and more accurate motions, even during failure recovery. These results demonstrate the value of the real-time proprioceptive feedback enabled by π2 ^2. Reactivity analysis. Fig. 5 plots a fingertip contact force against the action each method emits over time on Tidy Up Book. Because π2 ^2 refreshes proprioception at every denoising step, it modulates its grip from live force feedback and stops near ∼ 50 N, gripping just enough to reorient the book. Train-Time RTC instead commits to a stale plan and reacts late, so it keeps pushing down after contact, and its middle-finger force overshoots to ∼ 120 N, crushing the book. This force modulation is the mechanism behind our quantitative gains: reacting to contact as it happens rather than after the current chunk finishes. Tmaihe same pattern holds in all remaining tasks (Appendix A.4). Figure 5: π2 ^2 reacts to proprioception, while baselines run a stale plan (Tidy Up Book). Fingertip force (solid, left axis) and the emitted action (dashed, right axis) over time, for π2 ^2 and Train-Time RTC; numbered markers link plot times to the overhead frames above. π2 ^2 grips just enough (∼ 50 N), while RTC reacts late and over-grips to ∼ 120 N, dropping the book. 5 Discussion We presented π2 ^2, a principled modification to existing VLA architectures that yields real-time, closed-loop policies through two orthogonal contributions: asynchronous vision/text processing that decouples the slow VLM backbone from the action loop, and a latency-adaptive flow schedule that absorbs variable inference delay via a per-position noise schedule parameterized by d. Together, they reduce per-call delay by up to 4×4× while remaining compatible with any VLA action head. This enables π2 ^2 to handle highly reactive behaviors – non-prehensile, contact-rich manipulation – where baselines struggle, both in simulation and on real hardware. Limitations Our approach has several limitations. First, we do not address sources of latency external to the model itself, such as communication delays between the inference server and the robot client. Second, we keep the base policy architecture unchanged; designing it to emphasize proprioceptive features more strongly (e.g., dedicated attention heads for proprioceptive tokens) could further amplify reactivity, and we leave this direction to future work. Acknowledgements We appreciate the helpful discussions with CISCO, Tony Tao, Andrew Wang, Jason Liu, and Yuxuan Kuang. We also thank Yishu Li, Kallol Saha, and Soumojit Bhattacharya for helping real-world experiments. Lastly, we thank Hyeonwoo Kim for the feedback on website design. This work was supported by gift awards from CISCO, Google, and NSF Award IIS-2442282. References [1] J. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al. (2022) Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35, p. 23716–23736. Cited by: §1. [2] S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. (2025) Qwen3-vl technical report. arXiv preprint arXiv:2511.21631. Cited by: §1. [3] L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Bošnjak, X. Chen, M. Minderer, P. Voigtlaender, I. Bica, I. Balazevic, J. Puigcerver, P. Papalampidi, O. Henaff, X. Xiong, R. Soricut, J. Harmsen, and X. Zhai (2024) PaliGemma: a versatile 3b vlm for transfer. External Links: 2407.07726, Link Cited by: §1. [4] J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al. (2025) Gr00t n1: an open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734. Cited by: §1, §1. [5] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. (2025) π 0: A vision-language-action flow model for general robot control. In RSS, Cited by: §1, §2. [6] K. Black, M. Galliker, and S. Levine (2026) Real-time execution of action chunking flow policies. Advances in Neural Information Processing Systems 38, p. 33383–33407. Cited by: §2, §3.3. [7] K. Black, A. Z. Ren, M. Equi, and S. Levine (2025) Training-time action conditioning for efficient real-time chunking. arXiv preprint arXiv:2512.05964. Cited by: 2nd item, §A.1, §2, §3.3, §4.1.2, §4.2, §4.2, Table 1. [8] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al. (2023) Rt-1: robotics transformer for real-world control at scale. In RSS, Cited by: §1. [9] B. Chen, D. Martí Monsó, Y. Du, M. Simchowitz, R. Tedrake, and V. Sitzmann (2024) Diffusion forcing: next-token prediction meets full-sequence diffusion. Advances in Neural Information Processing Systems 37, p. 24081–24125. Cited by: Figure 1, §1, §3.1.2. [10] C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song (2023) Diffusion policy: visuomotor policy learning via action diffusion. In IJRR, Cited by: §1, §2. [11] O. X. Collaboration, A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Gupta, A. Wang, A. Kolobov, A. Singh, A. Garg, A. Kembhavi, A. Xie, A. Brohan, A. Raffin, A. Sharma, A. Yavary, A. Jain, A. Balakrishna, A. Wahid, B. Burgess-Limerick, B. Kim, B. Schölkopf, B. Wulfe, B. Ichter, C. Lu, C. Xu, C. Le, C. Finn, C. Wang, C. Xu, C. Chi, C. Huang, C. Chan, C. Agia, C. Pan, C. Fu, C. Devin, D. Xu, D. Morton, D. Driess, D. Chen, D. Pathak, D. Shah, D. Büchler, D. Jayaraman, D. Kalashnikov, D. Sadigh, E. Johns, E. Foster, F. Liu, F. Ceola, F. Xia, F. Zhao, F. V. Frujeri, F. Stulp, G. Zhou, G. S. Sukhatme, G. Salhotra, G. Yan, G. Feng, G. Schiavi, G. Berseth, G. Kahn, G. Yang, G. Wang, H. Su, H. Fang, H. Shi, H. Bao, H. B. Amor, H. I. Christensen, H. Furuta, H. Bharadhwaj, H. Walke, H. Fang, H. Ha, I. Mordatch, I. Radosavovic, I. Leal, J. Liang, J. Abou-Chakra, J. Kim, J. Drake, J. Peters, J. Schneider, J. Hsu, J. Vakil, J. Bohg, J. Bingham, J. Wu, J. Gao, J. Hu, J. Wu, J. Wu, J. Sun, J. Luo, J. Gu, J. Tan, J. Oh, J. Wu, J. Lu, J. Yang, J. Malik, J. Silvério, J. Hejna, J. Booher, J. Tompson, J. Yang, J. Salvador, J. J. Lim, J. Han, K. Wang, K. Rao, K. Pertsch, K. Hausman, K. Go, K. Gopalakrishnan, K. Goldberg, K. Byrne, K. Oslund, K. Kawaharazuka, K. Black, K. Lin, K. Zhang, K. Ehsani, K. Lekkala, K. Ellis, K. Rana, K. Srinivasan, K. Fang, K. P. Singh, K. Zeng, K. Hatch, K. Hsu, L. Itti, L. Y. Chen, L. Pinto, L. Fei-Fei, L. Tan, L. ”. Fan, L. Ott, L. Lee, L. Weihs, M. Chen, M. Lepert, M. Memmel, M. Tomizuka, M. Itkina, M. G. Castro, M. Spero, M. Du, M. Ahn, M. C. Yip, M. Zhang, M. Ding, M. Heo, M. K. Srirama, M. Sharma, M. J. Kim, M. Z. Irshad, N. Kanazawa, N. Hansen, N. Heess, N. J. Joshi, N. Suenderhauf, N. Liu, N. D. Palo, N. M. M. Shafiullah, O. Mees, O. Kroemer, O. Bastani, P. R. Sanketi, P. ”. Miller, P. Yin, P. Wohlhart, P. Xu, P. D. Fagan, P. Mitrano, P. Sermanet, P. Abbeel, P. Sundaresan, Q. Chen, Q. Vuong, R. Rafailov, R. Tian, R. Doshi, R. Mart’in-Mart’in, R. Baijal, R. Scalise, R. Hendrix, R. Lin, R. Qian, R. Zhang, R. Mendonca, R. Shah, R. Hoque, R. Julian, S. Bustamante, S. Kirmani, S. Levine, S. Lin, S. Moore, S. Bahl, S. Dass, S. Sonawani, S. Tulsiani, S. Song, S. Xu, S. Haldar, S. Karamcheti, S. Adebola, S. Guist, S. Nasiriany, S. Schaal, S. Welker, S. Tian, S. Ramamoorthy, S. Dasari, S. Belkhale, S. Park, S. Nair, S. Mirchandani, T. Osa, T. Gupta, T. Harada, T. Matsushima, T. Xiao, T. Kollar, T. Yu, T. Ding, T. Davchev, T. Z. Zhao, T. Armstrong, T. Darrell, T. Chung, V. Jain, V. Kumar, V. Vanhoucke, V. Guizilini, W. Zhan, W. Zhou, W. Burgard, X. Chen, X. Chen, X. Wang, X. Zhu, X. Geng, X. Liu, X. Liangwei, X. Li, Y. Pang, Y. Lu, Y. J. Ma, Y. Kim, Y. Chebotar, Y. Zhou, Y. Zhu, Y. Wu, Y. Xu, Y. Wang, Y. Bisk, Y. Dou, Y. Cho, Y. Lee, Y. Cui, Y. Cao, Y. Wu, Y. Tang, Y. Zhu, Y. Zhang, Y. Jiang, Y. Li, Y. Li, Y. Iwasawa, Y. Matsuo, Z. Ma, Z. Xu, Z. J. Cui, Z. Zhang, Z. Fu, and Z. Lin (2023) Open X-Embodiment: robotic learning datasets and RT-X models. Note: https://arxiv.org/abs/2310.08864 Cited by: §2. [12] M. Deitke, C. Clark, S. Lee, R. Tripathi, Y. Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, J. Lu, T. Anderson, E. Bransom, K. Ehsani, H. Ngo, Y. Chen, A. Patel, M. Yatskar, C. Callison-Burch, A. Head, R. Hendrix, F. Bastani, E. VanderBilt, N. Lambert, Y. Chou, A. Chheda, J. Sparks, S. Skjonsberg, M. Schmitz, A. Sarnat, B. Bischoff, P. Walsh, C. Newell, P. Wolters, T. Gupta, K. Zeng, J. Borchardt, D. Groeneveld, C. Nam, S. Lebrecht, C. Wittlif, C. Schoenick, O. Michel, R. Krishna, L. Weihs, N. A. Smith, H. Hajishirzi, R. Girshick, A. Farhadi, and A. Kembhavi (2024) Molmo and pixmo: open weights and open data for state-of-the-art vision-language models. External Links: 2409.17146, Link Cited by: §1. [13] Z. Fu, T. Z. Zhao, and C. Finn (2024) Mobile aloha: learning bimanual mobile manipulation with low-cost whole-body teleoperation. External Links: 2401.02117, Link Cited by: §2. [14] S. H. Høeg, Y. Du, and O. Egeland (2024) Streaming diffusion policy: fast policy synthesis with variable noise diffusion models. arXiv preprint arXiv:2406.04806. Cited by: §1, §2, §3.1.2, §3.3. [15] P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky (2025) π0.5 _0.5: A vision-language-action model with open-world generalization. External Links: 2504.16054, Link Cited by: §1, §2. [16] S. Jiang, X. Fang, N. Roy, T. Lozano-Pérez, L. P. Kaelbling, and S. Ancha (2025) Streaming flow policy: simplifying diffusion/flow-matching policies by treating action trajectories as flow trajectories. External Links: 2505.21851, Link Cited by: §2. [17] A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al. (2024) Droid: a large-scale in-the-wild robot manipulation dataset. In RSS, Cited by: §2. [18] M. J. Kim, Y. Gao, T. Lin, Y. Lin, Y. Ge, G. Lam, P. Liang, S. Song, M. Liu, C. Finn, et al. (2026) Cosmos policy: fine-tuning video models for visuomotor control and planning. In ICLR, Cited by: §3. [19] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. (2024) Openvla: an open-source vision-language-action model. arXiv preprint arXiv:2406.09246. Cited by: §2. [20] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. (2024) OpenVLA: an open-source vision-language-action model. In CoRL, Cited by: §1. [21] S. Li, Y. Gao, D. Sadigh, and S. Song (2025) Unified video action model. In RSS, Cited by: §3. [22] Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le (2022) Flow matching for generative modeling. arXiv preprint arXiv:2210.02747. Cited by: §3.1.1. [23] H. Liu, C. Li, Q. Wu, and Y. J. Lee (2023) Visual instruction tuning. Advances in neural information processing systems 36, p. 34892–34916. Cited by: §1. [24] Y. Lu, Z. Liu, X. Fan, Z. Yang, J. Hou, J. Li, K. Ding, and H. Zhao (2026) FASTER: rethinking real-time flow vlas. External Links: 2603.19199, Link Cited by: §2. [25] NVIDIA, :, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. ”. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y. L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y. Xie, Y. Xu, Z. Xu, S. Ye, Z. Yu, A. Zhang, H. Zhang, Y. Zhao, R. Zheng, and Y. Zhu (2025) GR00T n1: an open foundation model for generalist humanoid robots. External Links: 2503.14734, Link Cited by: §A.2, §2, §4.2. [26] C. Pan, G. Anantharaman, N. Huang, C. Jin, D. Pfrommer, C. Yuan, F. Permenter, G. Qu, N. Boffi, G. Shi, et al. (2025) Much ado about noising: dispelling the myths of generative robotic control. arXiv preprint arXiv:2512.01809. Cited by: §1. [27] E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville (2018) Film: visual reasoning with a general conditioning layer. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32. Cited by: §A.1. [28] K. Song, B. Chen, M. Simchowitz, Y. Du, R. Tedrake, and V. Sitzmann (2025) History-guided video diffusion. arXiv preprint arXiv:2502.06764. Cited by: §3.1.2. [29] J. Tang, Y. Sun, Y. Zhao, S. Yang, Y. Lin, Z. Zhang, J. Hou, Y. Lu, Z. Liu, and S. Han (2025) Vlash: real-time vlas via future-state-aware asynchronous inference. arXiv preprint arXiv:2512.01031. Cited by: §3.3. [30] G. Team, R. Anil, S. Borgeaud, J. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. (2023) Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805. Cited by: §1. [31] H. Xie, B. Wen, J. Zheng, Z. Chen, F. Hong, H. Diao, and Z. Liu (2026) DynamicVLA: a vision-language-action model for dynamic object manipulation. arXiv preprint arXiv:2601.22153. Cited by: §2. [32] G. Yan, J. Zhu, Y. Deng, S. Yang, R. Qiu, X. Cheng, M. Memmel, R. Krishna, A. Goyal, X. Wang, and D. Fox (2025) ManiFlow: a general robot manipulation policy via consistency flow training. External Links: 2509.01819, Link Cited by: §2. [33] S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang, et al. (2026) World action models are zero-shot policies. arXiv preprint arXiv:2602.15922. Cited by: §3. [34] C. Yuan, C. Wen, T. Zhang, and Y. Gao (2024) General flow as foundation affordance for scalable robot learning. arXiv preprint arXiv:2401.11439. Cited by: §1. [35] K. Zakka, B. Tabanpour, Q. Liao, M. Haiderbhai, S. Holt, J. Y. Luo, A. Allshire, E. Frey, K. Sreenath, L. A. Kahrs, et al. (2025) Mujoco playground. arXiv preprint arXiv:2502.08844. Cited by: §A.1, §4.1. [36] F. Zhang and M. Gienger (2024) Affordance-based robot manipulation with flow matching. arXiv preprint arXiv:2409.01083. Cited by: §3.1.1. [37] Q. Zhang, Z. Liu, H. Fan, G. Liu, B. Zeng, and S. Liu (2024) FlowPolicy: enabling fast and robust 3d flow-based policy via consistency flow matching for robot manipulation. External Links: 2412.04987, Link Cited by: §2. [38] T. T. Zhang, D. Pfrommer, C. Pan, N. Matni, and M. Simchowitz (2025) Action chunking and exploratory data collection yield exponential improvements in behavior cloning for continuous control. arXiv preprint arXiv:2507.09061. Cited by: §1, §3.1.1. [39] T. Z. Zhao, V. Kumar, S. Levine, and C. Finn (2023) Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705. Cited by: 2nd item, §1, §2, §3.1.1, §4.2, Table 1. [40] T. Z. Zhao, J. Tompson, D. Driess, P. Florence, K. Ghasemipour, C. Finn, and A. Wahid (2024) Aloha unleashed: a simple recipe for robot dexterity. arXiv preprint arXiv:2410.13126. Cited by: §2, §4.2. [41] C. Zhu, R. Yu, S. Feng, B. Burchfiel, P. Shah, and A. Gupta (2025) Unified world models: coupling video and action diffusion for pretraining on large robotic datasets. In RSS, Cited by: §3. [42] B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. (2023) Rt-2: vision-language-action models transfer web knowledge to robotic control. In CoRL, Cited by: §1, §2. Appendix A Implementation Details A.1 Simulation Task. The Leap Cube Reorientation task in MuJoCo Playground [35] requires a 1616-DoF Leap Hand to orient a cube without dropping it from the palm. We modify the original environment in two ways: we raise the control rate from 2020 Hz to 5050 Hz to enable more reactive control, and we restrict the goal distribution to four blue side up yaw poses 0,π/2,π,3π/2\0,π/2,π,3π/2\ rather than uniformly random goals. An episode succeeds if the cube reaches any of the four goals within 0.20.2 rad in 600600 steps; we report the success rate over 100100 episodes. Figure 6: Leap Cube Reorientation goals. The Leap Hand grasps a cube and must rotate it until the blue face points up at one of four target yaw poses (0∘,90∘,180∘,270∘0 ,90 ,180 ,270 ). The bottom caption shows the corresponding goal quaternions (w,x,y,z)(w,x,y,z). Dataset. Four PPO experts (different seeds) generate 5050 trajectories each, 200200 total at 5050 Hz. The simulator injects per-step sensing noise into the observation, sampled independently for each component: uniform noise in [−0.05,0.05][-0.05,0.05] rad on joint angles, uniform noise in [−0.02,0.02][-0.02,0.02] m on cube position, and Gaussian noise with scale 0.10.1 on the cube orientation quaternion (the quaternion is re-normalized after perturbation). The recorded demonstrations preserve this noise, so the training distribution already includes realistic sensing noise on both the proprioceptive and vision-derived components. Observation space. Each observation to_t concatenates: • Proprioception (3232 dim): noisy joint angles (1616) and per-joint tracking error – noisy joint angle minus the most recent motor target (1616). • Vision-derived cube pose (99 dim): palm-to-cube position error (33) and absolute cube orientation in the world frame (66). The 99-dim vision-derived subset stands in for VLM features per Sec. 3.2; asynchronous-vision/text evaluations apply staleness only to this subset. Action space. 1616-dim relative joint commands at 5050 Hz. Each action t∈[−1,1]16a_t∈[-1,1]^16 is scaled by 0.50.5 and added to the current motor target. Architecture. All variants share a Conditional U-Net 1D action head with chunk length H=16H=16 and nobs=2n_obs=2. The U-Net uses FiLM [27] modulation by the noise level τ at every residual block; π2 ^2 replaces the shared FiLM projection with a per-position one – one (γp,βp)( _p, _p) pair per chunk position – analogous to the per-position AdaLN change described in Sec. 3.3 for DiT-based heads. The rest of the architecture is identical across all variants. Training schedule. At each batch we sample d∼Uniform1,…,dmaxd \1,…,d_max\ with dmax=5d_max=5, build ⋆,d τ ,d, fill the front-d slots with ground-truth actions, and mask them out of the loss via mp=[p≥d]m_p=1[p≥ d]. With probability 0.20.2 we instead train on a standard flow schedule (single τ∼Uniform[0,1]τ [0,1] shared across positions, no mask), so the same network can also denoise a full chunk from pure noise. The standard-flow branch supplies the inference pass that initializes the buffer at episode start (Sec. 3.3, warm-up). For the zero-delay sync curve in Sec. 4.1.1, we train a d=0d=0-only variant (dmax=0d_max=0). Slow-subset delay in simulation. The simulation policies are purely state-based: no image or language input. The 99-dim vision-derived subset of the state vector (palm-to-cube position error and absolute cube orientation) plays the slow-channel role from Sec. 3.2. In this setting we find that neither delayed-visual-state sampling at training nor the learned delay embedding is necessary: the env already injects sensor noise on this subset at every step, so the policy is already tolerant to perturbed slow inputs out of the box. For all rows of Sec. 4.1.2, a single π2 ^2 checkpoint is used regardless of dvisd_vis: the synchronous slow-channel rows (dvis=0d_vis=0) and the asynchronous-vision/text rows (dvis>0d_vis>0) share the same network, and evaluation simply feeds a dvisd_vis-step-old version of the 99-dim subset at test time. Evaluation protocol. At inference, π2 ^2 performs one denoising step per policy call (each call outputs nactn_act actions). Train-time RTC and standard flow runs K=15K=15 denoising steps to predict the whole chunk. Each (method,d)(method,d) cell in Sec. 4.1.2 reports the mean success rate over 100100 episodes. Full sim hyperparameters are in Tab. 2. Table 2: Simulation training and inference hyperparameters for the Leap Cube Reorientation Conditional U-Net 1D policies. Values shown are π2 ^2 defaults; the right column lists deviations used by baselines. Empty = baseline same as π2 ^2. Parameter π2 ^2 Baseline deviations Shared across all variants Chunk length H 1616 Observation history nobsn_obs 22 Control rate 5050 Hz Optimizer AdamW (β=[0.95,0.999]β=[0.95,0.999]) Learning rate (peak) 1×10−41×10^-4 LR schedule cosine, 500500-step warm-up Weight decay 1×10−61×10^-6 Gradient-norm clip 1.01.0 Batch size (global) 256256 GPUs 1×1× A5000 Training budget 800800 epochs Method-specific (noise schedule + inference) Train-time dmaxd_max 55 1010 (Train-time RTC); n/a (Flow) Standard-flow warm-up prob. α 0.20.2 n/a Inference budget per call 11 1515 (Train-time RTC, Flow) Baselines. All three variants share the dataset, optimizer, and training budget. π2 ^2 and train-time RTC additionally share the per-batch delay-sampling protocol: each batch draws d∈0,…,dmaxd∈\0,…,d_max\ from an exponentially-decaying distribution p(d=k)∝e−αdkp(d=k) e^- _dk with αd=1 _d=1 (smaller d more likely following [7]), and the front-d positions are clamped to ground-truth actions and masked out of the loss. The three variants then differ only in: • Flow: no d sampling and no front clamp; a single shared noise level τ∼Uniform[0,1]τ [0,1] is applied to the entire chunk. This is standard chunked flow matching, evaluated at NFE=15=15. • Train-time RTC [7]: same per-batch d sampling as π2 ^2, but the noise schedule applies a single shared noise level τ∼Uniform[0,1]τ [0,1] to the back H−dH-d positions (no per-position structure). No standard-flow warm-up mixing, and larger dmax=10d_max=10 as the effective delay is larger. • π2 ^2: same per-batch d sampling as train-time RTC, plus the α=0.2α=0.2 standard-flow warm-up branch (with probability 0.20.2 the batch instead uses a single τ∼Uniform[0,1]τ [0,1] on all positions and no front clamp, so the same network can also initialize a buffer from pure noise at episode start). On the remaining 0.80.8 of batches we apply the three-region staircase ⋆,d τ ,d of Sec. 3.3, and the U-Net uses per-position FiLM heads (γp,βp)p=0H−1\( _p, _p)\_p=0^H-1. A.2 Real-World Hardware and Workspace. Figure 7: Real-world workspace. xArm6 + XHand with a single overhead RGB camera. A 66-DoF xArm6 manipulator carries a 1212-DoF XHand at its wrist; a single overhead 640×480640× 480 RGB camera looks down on the workspace from above (Fig. 7). We fine-tune on 8×8× NVIDIA RTX A5000/A6000 (2424/4848 GB VRAM). Deployment uses a single RTX A5000 for the baselines. π2 ^2 uses 2×2× RTX A5000 so the slow VLM and the fast DiT action head do not share compute or memory bandwidth. Sharing one GPU between the two workers measurably inflates DiT latency and would conflate the measured d. Control rate. Predicted actions run closed-loop at 2525 Hz at deployment (Tctrl≈40T_ctrl≈ 40 ms per tick). Teleoperation captures demonstrations at the same rate, bottlenecked by the camera and the XHand–workstation USB link. Base model. We fine-tune NVIDIA GR00T-N1.7 [25]. The Qwen-VL backbone and Eagle-2 vision encoder remain frozen; only the action head – state/action projectors, the position embedding, and the DiT – is trained. The only architectural change to the action head replaces the DiT’s shared AdaLN modulation with the per-position (γp,βp)( _p, _p) heads of Sec. 3.3. Observation space. Each call to the policy receives: • Proprioception (4545 dim): 66 xArm6 joint angles, 1212 XHand joint angles (the actively-controlled finger DoFs), 1212 corresponding per-joint torques, and 1515 fingertip-force values (33-axis force from each of 55 sensorized fingertips). • Vision and language: one 640×480640× 480 RGB image from the overhead camera, encoded by Eagle-2 into vision tokens, concatenated with the task prompt (Tab. 4) at the VLM backbone input. Action space and low-level drivers. Each policy call emits a chunk of H=50H=50 absolute joint-position targets (66 xArm6 + 1212 XHand) at 2525 Hz, spanning 22 s of motion. We send one target per control tick to each robot’s driver in position-control mode. The arms’ onboard firmware interpolates between consecutive commands at their internal servo rates(Tab. 3). Table 3: Low-level driver settings. Both robots run in position-only mode at 2525 Hz outbound from the policy; the onboard firmware interpolates between targets at its internal rate. xArm6 XHand Driver mode mode=1 (streaming servo) RS485, position mode Command call set_servo_angle_j finger set_target Bus / link TCP RS485 Onboard control API (PID + interp.) per-joint PID, kp=100,ki=0,kd=0k_p=100,\ k_i=0,\ k_d=0 Safety limits collision sensitivity =2=2 torque limit =300=300 Data collection. We teleoperate the robot setup to collect 200200 demonstrations for Don’t Spill, 300300 for Tidy Up Book and Insert Box, and 100100 for Catch Book, randomizing the target object’s initial pose. Table 4: Real-world task prompts. Exact language input passed to GR00T-N1.7 at both training and deployment. Task Language prompt Don’t Spill put the ball in the bowl and put them on the cutting board Tidy Up Book put the books in the basket Insert Box put the box in the basket Catch Book catch the book Training the asynchronous vision/text mechanism. The real-world π2 ^2 training explicitly exercises the asynchronous-vision/text mechanism of Sec. 3.2. Each batch samples a vision delay dvis∼Uniform0,…,dvismaxd_vis \0,…,d_vis \ with dvismax=5d_vis =5 ticks (200200 ms at 2525 Hz); the policy then receives the image recorded dvisd_vis ticks earlier (i.e., the corresponding past frame in the demonstration), while proprioception stays current. For each batch, the integer dvis∈0,…,dvismaxd_vis∈\0,…,d_vis \ indexes a learned embedding e(dvis)e(d_vis) from a (dvismax+1)(d_vis +1)-entry lookup table. The DiT adds e(dvis)e(d_vis) to its action-token features (broadcast across the H chunk positions), so the action head conditions on both the cached visual feature and its age. We zero-initialize the lookup table so that the untrained checkpoint reproduces the no-delay variant exactly. At deployment, the measured wall-clock vision latency passes through the same embedding before each DiT call. Training hyperparameters. We fine-tune the GR00T-N1.7 action head (projectors + DiT) and keep the VLM backbone and vision encoder frozen. Full optimizer settings, batch size, and step budgets are in Tab. 5. The per-task budgets (40,00040,000 steps for Don’t Spill, 28,00028,000 for Tidy Up Book, 49,00049,000 for Insert Box, 4,5504,550 for Catch Book) correspond to approximately 200/100/100/100200/100/100/100 epochs over each dataset. The per-position AdaLN parameters (γp,βp)p=0H−1\( _p, _p)\_p=0^H-1 initialize from the pretrained shared pair, so each head starts from a position-uniform schedule and gradually specializes during fine-tuning. All four rows in Tab. 1 fine-tune from the same GR00T-N1.7 checkpoint with identical data and budget; they differ only in the noise schedule and (for π2 ^2) the AdaLN modification. The train-time RTC baseline samples its per-batch delay from d∈0,…,10d∈\0,…,10\ (i.e., dmax=10d_max=10) rather than the dmax=5d_max=5 used in π2 ^2, as π2 ^2is separating the delay into two parts while RTC treats it as a whole. Full hyperparameters are in Tab. 5. Evaluation protocol. Each (method,task)(method,task) cell runs N=20N=20 trials with randomized initial object placement. Success rate (SR) counts whole-task completion within the episode time limit (30 seconds). The progress score (Prog) decomposes each task into sub-goals and reports the per-trial fraction completed: • Don’t Spill – 44 sub-goals: pick up ball, place ball in bowl, lift bowl off the table, place bowl on the cutting board. • Tidy Up Book – 22 sub-goals: pull a book free from the pile, place it inside the basket. • Insert Box – 44 sub-goals: push the box towards the wall, make it stand, grasp, and insert between books. • Catch Book – 11 sub-goal: grasp the falling book and pick it up. Deployment query modes. The four rows of Tab. 1 differ in how the action-head worker queries the policy: • Flow, Synchronous: the main loop blocks until each call returns; the robot freezes during inference. • Flow, Naive Async, dense, ensemble: the action worker queries back-to-back and combines overlapping chunks with the temporal-ensembling weights of Zhao et al. [39]. • Train-time RTC and π2 ^2: the action worker queries back-to-back; the new chunk replaces the active one at the next swap boundary. The actions executed during policy call is given as input. π2 ^2 additionally spawns a VLM-cache worker on a separate GPU that refreshes the cached visual features in the background. Table 5: Real-world training and inference hyperparameters for the GR00T-N1.7 fine-tunes. Values shown are π2 ^2 defaults; the right column lists deviations used by baselines. Empty = baseline same as π2 ^2. Parameter π2 ^2 Baseline deviations Shared across all variants Chunk length H 5050 Observation history nobsn_obs 11 Control rate 2525 Hz Optimizer fused AdamW Learning rate (peak) 1×10−41×10^-4 LR schedule cosine, 5%5\% warm-up Weight decay 1×10−51×10^-5 Gradient-norm clip 1.01.0 Batch size 512512 Precision bf16 + tf32 GPUs 8×8× A5000/A6000 Training budget 100100 epoch (Tidy Up Book / Insert Box Catch Book) / 200200 epoch (Don’t Spill) Method-specific (noise schedule + delay + inference) Train-time dmaxd_max 55 1010 (Train-time RTC); n/a (Flow) Train-time dvismaxd_vis 55 n/a Delay-embedding lookup size 66 entries → DiT hidden dim n/a Standard-flow warm-up prob. α 0.20.2 0 (Flow, Train-time RTC) NFE per call 11 (one sub-chunk) 44 (Flow, Train-time RTC; full chunk) A.3 Training and Inference Algorithms We give pseudocode for the per-batch training step (Alg. 1) and the deployment-time inference loop (Alg. 2). The inference loop is the loop implemented in our deployment script; the wall-clock measurement that auto-derives d uses a rolling window of the most recent query times. Algorithm 1 π2 ^2 training step. Each batch samples a delay d, builds the staircase ⋆,d τ ,d, masks the front, applies jitter, and (for real-world training) additionally delays the slow channel. 1:Action chunk 0:H−1a_0:H-1, observation (tfast,t:t−Tslow+1slow)(o^fast_t,o^slow_t:t-T_slow+1), model vθv_θ 2:Hyperparams dmax,dvismax,j,αd_max,d_vis ,j,α 3:u∼Uniform[0,1]u [0,1] 4:if u<αu<α then ⊳ warm-up branch: standard flow 5: t∼Uniform[0,1]t [0,1]; τp←t,∀p _p← t,\ ∀ p 6: τ,p←(1−τp)ϵp+τppx_τ,p←(1- _p) ε_p+ _pa_p 7: mp←1,∀pm_p← 1,\ ∀ p ⊳ all positions contribute to loss 8:else⊳ staircase branch 9: d∼Uniform1,…,dmaxd \1,…,d_max\ 10: Build ⋆,d τ ,d as in Sec. 3.3⊳ three-region staircase 11: δp∼Uniform[−j,j] _p [-j,j]; τp←clip(τp⋆,d+δp,0,1) _p (τ ,d_p+ _p,0,1) 12: τ,p←(1−τp)ϵp+τppx_τ,p←(1- _p) ε_p+ _pa_p 13: τ,p←px_τ,p _p if p<dp<d⊳ front d slots are ground-truth (inpaint conditioning) 14: mp←[p≥d]m_p 1[p≥ d] 15:end if 16:dvis∼Uniform0,…,dvismaxd_vis \0,…,d_vis \⊳ slow-channel staleness (real-world only) 17:slow←t−dvisslowo^slow ^slow_t-d_vis; append e(dvis)e(d_vis) to slow representation 18:^p←vθ(τ,,tfast,slow)p v_p← v_θ(x_τ, τ,o^fast_t,o^slow)_p 19:ℒ←∑p=0H−1mp‖^p−(p−ϵp)‖2L← _p=0^H-1m_p\,\| v_p-(a_p- ε_p)\|^2 20:Update θ on ℒL Algorithm 2 π2 ^2 inference loop on the robot. The action head runs continuously with fresh proprioception and a cached slow feature; the VLM runs asynchronously in a background thread; the per-call delay d is auto-derived from a rolling measurement of action-head wall-clock latency. 1:Policy vθv_θ, observation source, robot controller, control period TctrlT_ctrl 2:Rolling window W of recent query times (e.g. W=20W=20) 3:←WarmStart()C← WarmStart()⊳ Sec. 3.3: full standard-flow inference, denoise all H positions 4:←C← re-noise C to match ⋆,d0 τ ,d_0 for initial estimate d0d_0 5:i←0i← 0⊳ chunk index 6:queue Q←∅Q← ⊳ rolling latency window 7:Spawn VLM_Worker thread (continuously updates cached slow feature) 8:Spawn Action_Worker thread (continuously queries action head) 9:loop⊳ one iteration per control tick (2525 Hz) 10: t0←t_0← current wall-clock time 11: Read fresh proprioception tfasto^fast_t from robot sensors 12: Publish tfasto^fast_t to shared state for workers 13: if new chunk newC_new is available from Action_Worker then 14: ←newC _new; i←di← d⊳ swap and skip past in-flight front actions 15: end if 16: Send action iC_i to robot; i←i+1i← i+1 17: Sleep until t0+Tctrlt_0+T_ctrl 18:end loop 19: 20:procedure Action_Worker ⊳ background thread; one denoising step per call 21: loop 22: Snapshot (tfast,tstate)(o^fast_t,t_state) and the cached (cachedslow,timage)(o^slow_cached,t_image) 23: dvis←round((tstate−timage)/Tctrl)d_vis ((t_state-t_image)/T_ctrl ) ⊳ age of cached slow feature in control ticks 24: d←max(1,round(mean(Q)/Tctrl))d← (1,\ round(mean(Q)/T_ctrl) )⊳ auto-derived from rolling window 25: Snapshot front-d inpaint: 0:dinflight←i:i+da^inflight_0:d _i:i+d 26: q0←q_0← wall-clock 27: new←vθ.euler_step(inflight,fast,cachedslow,e(dvis))C_new← v_θ. euler\_step(a^inflight,o^fast,o^slow_cached,e(d_vis)) ⊳ one NFE; per-position Δτp _p as in Eq. (4) 28: Push wall-clock−q0wall -clock-q_0 into Q 29: Hand off newC_new to main loop 30: end loop 31:end procedure 32: 33:procedure VLM_Worker ⊳ background thread; runs vision/text forward asynchronously 34: loop 35: Read fresh image and language; timage←t_image← image capture time 36: cachedslow←o^slow_cached← VLM forward (image, language) 37: Atomically update cache with (cachedslow,timage)(o^slow_cached,t_image) 38: end loop 39:end procedure A.4 Additional Reactivity Analysis The main-text reactivity comparison (Fig. 5) is on Tidy Up Book; the same pattern holds across the other tasks (Fig. 8). On Catch Book, π2 ^2 closes on the force spike as the book lands, whereas Train-Time RTC reacts too late and the book slips. On Insert Box, RTC reacts late to the force feedback and its index force overshoots out of distribution, while π2 ^2 stays controlled. On Don’t Spill, RTC’s thumb never contacts the ball (thumb force ≈0≈ 0) yet it keeps going, while π2 ^2 grasps it with both fingers. Figure 8: Reactivity on the remaining tasks. Fingertip force (solid, left axis) and the emitted action (dashed, right axis) over time for π2 ^2 and Train-Time RTC, with overhead frames at the marked times. Top: Catch Book. Middle: Insert Box. Bottom: Don’t Spill.