Paper deep dive
AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation
Kai Yang, Zedong Chu, Yingnan Guo, Zhengbo Wang, Shichao Xie, Yanfen Shen, Xiaolong Wu, Xing Li, Mu Xu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 6/21/2026, 7:59:28 AM
Summary
AsyncShield is a plug-and-play asynchronous control framework designed for mobile robot navigation using cloud-based Vision-Language-Action (VLA) models. It addresses the issue of spatiotemporal misalignment caused by network jitter and inference latency. The framework replaces traditional black-box time-series prediction with a deterministic 'white-box' spatial mapping using SE(2) kinematic transformations and a temporal pose buffer. To balance VLA intent tracking with physical safety, it employs a Reinforcement Learning (RL) adapter trained via the PPO-Lagrangian algorithm, formulated as a Constrained Markov Decision Process (CMDP). The system demonstrates robust zero-shot generalization across heterogeneous robot chassis through domain randomization and a standardized universal sub-goal interface.
Entities (8)
Relation Signals (5)
VLA Models → require → cloud-based deployment
confidence 100% · their massive parameter sizes typically necessitate cloud-based deployment.
AsyncShield → trainedin → OmniSafe
confidence 100% · We train the AsyncShield in a highly stochastic simulated environment using the OmniSafe framework
AsyncShield → uses → PPO-Lagrangian algorithm
confidence 100% · Solved via the PPO-Lagrangian algorithm, a reinforcement learning adapter dynamically trades off...
AsyncShield → utilizes → SE(2) kinematic transformation
confidence 100% · the system employs an analytical SE(2) kinematic transformation to eliminate the cloud-to-edge latency misalignment
AsyncShield → addresses → spatiotemporal misalignment
confidence 95% · To address this issue, we propose AsyncShield, a plug-and-play asynchronous control framework.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While Vision-Language-Action (VLA) models have been demonstrated possessing strong zero-shot generalization for robot control, their massive parameter sizes typically necessitate cloud-based deployment. However, cloud deployment introduces network jitter and inference latency, which can induce severe spatiotemporal misalignment in mobile navigation under continuous displacement, so that the stale intents expressed in past ego frames may become spatially incorrect in the current frame and lead to collisions. To address this issue, we propose AsyncShield, a plug-and-play asynchronous control framework. AsyncShield discards traditional black-box time-series prediction in favor of a deterministic physical white-box spatial mapping. By maintaining a temporal pose buffer and utilizing kinematic transformations, the system accurately converts temporal lag into spatial pose offsets to restore the VLA's original geometric intent. To balance intent restoration fidelity and physical safety, the edge adaptation is formulated as a constrained Markov decision process (CMDP). Solved via the PPO-Lagrangian algorithm, a reinforcement learning adapter dynamically trades off between tracking the VLA intent and responding to high-frequency LiDAR obstacle avoidance hard constraints. Furthermore, benefiting from a standardized universal sub-goal interface, domain randomization, and perception-level adaptation via Collision Radius Inflation, AsyncShield operates as a lightweight, plug-and-play module. Simulation and real-world experiments demonstrate that, without fine-tuning any cloud-based foundation models, the framework exhibits zero-shot and robust generalization capabilities, effectively improving the success rate and physical safety of asynchronous navigation.
Tags
Links
- Source: https://arxiv.org/abs/2604.24086v1
- Canonical: https://arxiv.org/abs/2604.24086v1
Trouble viewing inline? Open PDF directly →
Full Text
43,780 characters extracted from source content.
Expand or collapse full text
AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation Kai Yang 1 , Zedong Chu 1† , Yingnan Guo 1 , Zhengbo Wang 1 , Shichao Xie 1 , Yanfen Shen 1 , Xiaolong Wu 1 , Xing L ̈ u 2 , and Mu Xu 1 Abstract— While Vision-Language-Action (VLA) models have been demonstrated possessing strong zero-shot gener- alization for robot control, their massive parameter sizes typically necessitate cloud-based deployment. However, cloud deployment introduces network jitter and inference latency, which can induce severe spatiotemporal misalignment in mobile navigation under continuous displacement, so that the stale intents expressed in past ego frames may become spatially incorrect in the current frame and lead to collisions. To address this issue, we propose AsyncShield, a plug-and-play asyn- chronous control framework. AsyncShield discards traditional black-box time-series prediction in favor of a deterministic physical white-box spatial mapping. By maintaining a temporal pose buffer and utilizing kinematic transformations, the system accurately converts temporal lag into spatial pose offsets to restore the VLA’s original geometric intent. To balance intent restoration fidelity and physical safety, the edge adaptation is formulated as a constrained Markov decision process (CMDP). Solved via the PPO-Lagrangian algorithm, a reinforcement learning adapter dynamically trades off between tracking the VLA intent and responding to high-frequency LiDAR obstacle avoidance hard constraints. Furthermore, benefiting from a standardized universal sub-goal interface, domain random- ization, and perception-level adaptation via Collision Radius Inflation, AsyncShield operates as a lightweight, plug-and-play module. Simulation and real-world experiments demonstrate that, without fine-tuning any cloud-based foundation models, the framework exhibits zero-shot and robust generalization capabilities, effectively improving the success rate and physical safety of asynchronous navigation. I. INTRODUCTION While Vision-Language-Action (VLA) models have been demonstrated possessing impressive zero-shot generalization capabilities in robotic manipulation [1]–[4] and robotic nav- igation [5]–[10], their massive parameter sizes usually ne- cessitate cloud-based deployment. This inevitably introduces systemic cloud-to-edge latency. In dynamic environments, inference suffers from significant latency while control must operate in real-time. Consequently, semantic information frequently corresponds to past states but is utilized as current, thereby triggering a systemic temporal misalignment be- tween “thinking” and “control” [11]. How to safely transfer the powerful capabilities of cloud-based large models to highly dynamic mobile navigation tasks remains an urgent challenge. Existing asynchronous control frameworks (e.g., RTC [12] and A2C2 [13]) predominantly focus on temporal align- † Corresponding author. 1 Amap,AlibabaGroup,Beijing,China.Emails: yk496472, chuzedong.czd@alibaba-inc.com 2 Beijing Jiaotong University, Beijing, China. ment, mitigating latency via smooth action chunk splicing or local residual correction. While effective in fixed-base manipulator operations, these methods face a fundamental logic failure when transferred to mobile robots undergoing continuous, large-scale displacements: blindly and smoothly fitting an outdated path prevents the robot from adequately responding to dynamic obstacles. Furthermore, end-to-end delay-injection training methods struggle to cope with long- tail, irregular network jitter in the real world; meanwhile, traditional control frameworks based on black-box time- series prediction also prove extremely fragile under extreme communication latency, easily leading to catastrophic task failures. To address these issues, we propose the AsyncShield, a framework that transforms traditional asynchronous control based on black-box time prediction into a deterministic physical white-box mapping and a safe execution closed loop. We maintain a historical temporal pose buffer at the edge. Upon receiving outdated VLA commands, the system employs an analytical SE(2) kinematic transformation to eliminate the cloud-to-edge latency misalignment, thereby restoring the VLA’s true geometric intent. To balance the restoration fidelity of the VLA intent and physical safety, we formulate the edge adaptation as a CMDP. Through the PPO-Lagrangian algorithm [14], the reinforcement learning adapter can adaptively trade off between tracking the true VLA intent and responding to the high-frequency LiDAR obstacle avoidance hard constraints. Furthermore, AsyncShield is designed as an extremely lightweight, independent edge adaptation architecture, em- phasizing “plug-and-play” capability. Unlike existing solu- tions that require deep modifications or fine-tuning of the cloud-based foundation models, we standardize the cloud VLA output via interpolation resampling into 5 local way- points with a 20 cm spacing as the input to the adapter. Based on this, the policy network solely outputs a Universal Local Sub-goal. This standardized interface design, combined with various randomization training strategies, endows the system with plug-and-play generalization capabilities, seamlessly accommodating various cloud VLA models and generalizing across multiple heterogeneous robot chassis without any fine- tuning. In summary, the main contributions of this paper are as follows: • Cloud-based VLA Navigation Edge Adapter: We propose the AsyncShield framework, which replaces black-box time prediction with deterministic white-box arXiv:2604.24086v1 [cs.RO] 27 Apr 2026 spatial mapping. Concurrently, by formulating a CMDP- based safe execution closed loop, it achieves an adaptive dynamic trade-off between the restoration of the large model’s geometric intent and low-level, high-frequency obstacle avoidance. • Strong Plug-and-Play Capability: We propose a lightweight, independent edge adaptation strategy with a standardized interface. Without requiring any fine- tuning of the cloud-based foundation models, Async- Shield can zero-shot seamlessly accommodate various cloud VLA models and reliably generalize across mul- tiple heterogeneous mobile robot chassis. I. RELATED WORKS A. Vision-Language-Action Models and Hierarchical Adap- tation Inrecentyears,VLAmodels(e.g.,RT-2[2], OmniVLA [15], Green-VLA [16]) have demonstrated formidable zero-shot generalization capabilities, gradually expanding into highly dynamic scenarios such as unmanned aerialvehicles(AutoFly[17]),autonomousdriving (Impromptu VLA [18]), and dynamic object interaction (DynamicVLA [19]). However, their massive parameter sizes lead to extremely high inference latency and network jitter. To alleviate the deployment bottleneck at the edge, existing research has primarily evolved along multiple trajectories. One approach lowers the computational threshold through model lightweighting, context compression, or post-training reinforcement learning fine-tuning (e.g., SmolVLA [20], ContextVLA [21], SimpleVLA-RL [22]). Another constructs a “fast-slow dual-system” hierarchical architecture (e.g., Mobility VLA [23], IROS Dual-Process [24]), or introduces world models with latent state spaces to enhance generalization(e.g.,X-MOBILITY[25]).Particularly in the domain of cross-embodiment navigation, works such as X-Nav [26] explore the end-to-end distillation of massive expert policies, while CE-Nav [27] goes a step further by successfully decoupling high-level geometric reasoning from low-level dynamic execution via a two- stage architecture. Similarly, ABot-Explorer [28] enhances high-level guidance through online scene graph memory construction. Nevertheless, the aforementioned works often implicitly rely on synchronous control assumptions. When faced with the irregular communication delays inherent in real-world cloud deployment, the closed-loop execution between high-level, low-frequency semantics and low-level, high-frequency physical execution is highly susceptible to fracture. As a fully plug-and-play module, our AsyncShield circumvents communication and computational bottlenecks without necessitating any intervention in the internal weights of the cloud-based foundation model. B. Asynchronous Control and Latency-Aware Local Naviga- tion Action Chunking techniques (e.g., ACT [29], Diffusion Policy [30]) have been widely validated in fixed-base ma- nipulator operations, and recent works like Mixture of Hori- zons [31] further explore adaptive fusion strategies for multi- horizon chunks. To address the asynchronous latency of mobile-based large models, existing strategies predominantly focus on “temporal fitting and prediction.” For instance, VLASH [32] relies on forward-state prediction to estimate the robot’s future pose; DuoCore-FS [33] attempts to con- struct a latent representation buffer to bridge the dual-track system, whereas AsyncVLA [34] concentrates on edge-side asynchronous adaptation manipulation via cloud-edge col- laboration. Simultaneously, within low-level local perception modules, traditional frameworks (e.g., NavFormer [35], V- STRONG [36]) and various reinforcement learning schemes (e.g., Hierarchical RL Nav [37], Enhanced PPO [38], De- centralized RL [39], APD for SRL [40]) exhibit excel- lent environmental adaptability. However, when receiving expired commands from cloud-based large models, blind action chunk splicing along the temporal axis still leads to catastrophic spatial misalignment for mobile bases. This paper breaks the traditional paradigm of temporal sequence fitting by proposing “Latency is Geometry.” Our system directly utilizes analytical kinematic transformations within the SE(2) space to precisely map temporal lag into spatial relative offsets within the ego-centric coordinate frame. Com- bined with off-the-shelf local reinforcement learning obstacle avoidance algorithms acting as low-level cost constraints, this physical white-box mapping eliminates the uncertainty brought by black-box predictions, achieving high-frequency, robust control free from manual parameter tuning. I. METHODOLOGY The AsyncShield framework is designed to bridge the spatio-temporal gap between low-frequency, high-latency cloud VLA models and high-frequency, safety-critical edge execution. We formulate this asynchronous control problem as a constrained Markov decision process, where the edge adapter rectifies stale semantic intents through explicit ge- ometric realignment and ensures physical safety via con- strained policy optimization. A. Spatio-Temporal Intent Realignment To address displacement-induced misalignment during continuous motion, we discard implicit temporal fitting and instead employ an analytical SE(2) transformation to map “instruction lag” into a deterministic “spatial offset.” 1) Temporal Pose Buffer: The edge device maintains a circular buffer B = (t k , T O W (t k )), recording the robot’s odometry (from world W to ego frame O) at f edge Hz. When a VLA packet arrives with an anchor timestamp t a , the system retrieves the corresponding historical pose T O W (t a ) via linear interpolation for translation and shortest- path angular interpolation for the heading in the SE(2) space. 2) Geometric Re-projection: Let P A = ̄p A i N i=1 be the set of N local waypoints generated by the VLA model in the Anchor Ego Frame at t a . The realigned waypoints ̄p E i (t) in the Current Ego Frame at time t > t a are computed Fig. 1: Overview of the AsyncShield framework. Top: Behavioral comparison under network degradation. Naive Execution blindly follows stale intents, leading to collisions, whereas AsyncShield safely bypasses obstacles. Bottom: The edge adaptation pipeline. The system first utilizes a temporal pose buffer to perform spatio-temporal intent realignment on delayed cloud waypoints. Subsequently, a policy optimized via PPO-Lagrangian adaptively trades off between intent fidelity and hard safety constraints to output universal local sub-goals. Finally, by incorporating actuator domain randomization during training, the framework achieves plug-and-play generalization across embodiments. analytically: p E i (t) = (T O W (t)) −1 T O W (t a )p A i (1) wherep denotes homogeneous coordinates. The spatial re-projection formula exclusively computes relative pose changes within the minimal delay window ∆t. By establishing a new temporal anchor t a for each incoming VLA packet, this mechanism strictly confines odometry drift (e.g., from wheel slip) to a single communication cycle. Consequently, previous spatial alignment errors are instantly reset to zero upon receiving new waypoints, effectively preventing global divergence over time. We define the edge adaptation as a CMDP tuple (S,A,P,R,C,γ,d) with transition dynamics P and dis- count factor γ, aiming to maximize the intent-tracking reward J R while keeping the expected safety cost J C below a threshold d. 1) State and Action Space: The state s t = [o geo t , o lidar t ] integrates geometric and reactive features. o geo t ∈ R 10 consists of 5 look-ahead points sampled from the realigned path at 0.2 m intervals. o lidar t ∈ R 144 provides 2D LiDAR proximity data. To ensure cross-embodiment compatibility, the action a t = (∆x t , ∆y t ) is defined as a Universal Local Sub-goal in the current ego frame, which is subsequently converted into velocity commands by a low-level controller. 2) Reward Design (Intent Fidelity): The reward function r t focuses on trajectory fidelity and smoothness, independent of obstacle avoidance, r t = w flow (a ⊤ t ˆτ t )− w cte tanh(d cte t )− w smooth ∥a t − a t−1 ∥ 2 (2) where w flow ,w cte ,w smooth are positive weight coefficients, and ˆτ t is the local path unit tangent.To ensure gradient continuity, the cross-track error d cte t is calculated via the point-to-line segment distance: d cte t =∥p 0 + t ∗ (p 1 − p 0 )∥ 2 (3) t ∗ = clip −p ⊤ 0 (p 1 − p 0 ) ∥p 1 − p 0 ∥ 2 , 0, 1 (4) where p 0 , p 1 are the immediate waypoints in the look-ahead window. 3) Safety Constraints: Physical safety is enforced as an independent cost c t based on the minimum LiDAR distance d min . Given a safety radius R safe , the cost is triggered as: c t = I(d min < R safe ) + α max(0,R safe − d min )(5) where I(·) is the indicator function and α is a penalty scaling factor. We employ the PPO-Lagrangian algorithm [14] to update a dual variable λ that balances J R and J C . When stale VLA intents pose collision risks, the surge in λ forces the policy to prioritize safety advantages, resulting in proactive collision avoidance . TABLE I: Quantitative comparison under two network conditions. SR, CTE, and RER are computed over all 600 episodes (CTE is episode-weighted). RER is the percentage of time steps with d min < d risk . TTG is computed on the shared-success set (intersection of success episodes across all compared methods; 120/600 for Ideal and 100/600 for Delayed). Method Ideal (Fast Update)Non-ideal (Mixed Degradation) SR (%)↑CTE (m)↓RER (%)↓TTG (s)↓SR (%)↑CTE (m)↓RER (%)↓TTG (s)↓ Ours80.00.7171.28.8776.70.7251.39.29 A2C256.70.9371.58.4143.31.1461.79.64 RTC40.00.6733.08.1830.01.1784.09.81 Naive20.01.1755.27.9216.71.2725.510.25 B. Training Environment and Kinematic Domain Random- ization To ensure zero-shot transferability, we train the Async- Shield in a highly stochastic simulated environment using the OmniSafe framework with extensive perturbations. 1) Stochastic Environment Configuration: Training is conducted in a 10 m× 10 m workspace featuring both static and dynamic collision risks. Each episode is initialized with: • Geometric Diversity: 6 static and 6 dynamic obstacles with heterogeneous geometries (cylinders, polygons, and irregular shapes). The equivalent radius R obs is sampled from U (0.2, 2.0) m. • Dynamic Interference: Dynamic obstacles follow arandom-walkmodelwithlinearvelocities v obs ∈[0.2, 1.0] (m/s)andangularvelocities ω obs ∈ [0.1, 1.0] (rad/s). • Asynchronous Disturbances: We simulate irregular communication by sampling latencies δ ∼U (0.3, 1.5) s and packet loss probabilities p loss ∈ [0, 0.2] , and intermittent transient outages (blackouts) lasting up to several seconds . 2) Cross-Embodiment Actuator Randomization: We decouple the policy from specific robot dynamics by random- izing the actuator response model. The transition from a sub- goal a t to the actual executed velocity v act is governed by acceleration constraints followed by a first-order lag system: v act (t) = clip τv cmd (t) + (1− τ )v act (t− 1) +η v ,−v max ,v max (6) where v cmd is the intended command derived from a t , and v max is the velocity limit. The randomization parameters are sampled per episode to cover diverse chassis profiles: • System Latency: The first-order lag coefficient τ ∼ U (0.2, 0.9), representing varying motor response times. • AccelerationConstraints: Maximum acceleration a max ∼ U (0.5, 1.5) m/s 2 to simulate different power- to-weight ratios. • Stochastic Noise & Bias: Velocity-dependent Gaussian noise η v with intensity σ ∼ U (0.05, 0.20) and system- atic angular bias b ω ∼U (−0.05, 0.05) rad/s. This extensive kinematic perturbation ensures that the AsyncShield learns a dynamics-agnostic control law capable of seamless deployment on heterogeneous mobile platforms. IV. EXPERIMENTS This section evaluates the performance of the Async- Shield framework under asynchronous cloud-edge control. We design extensive simulation experiments to answer the following questions: (1) Can AsyncShield maintain system robustness and navigational safety under varying network latencies? (2) What is the inherent trade-off between “action smoothness” and “physical safety” in mobile robot naviga- tion? (3) How do the core modules—temporal pose buffer, RL adapter, and constrained optimization—contribute to the overall performance? A. Experimental Setup and Metrics Environment and Task Setup: We construct a 3D navi- gation environment in OmniSafe. The robot navigates to a target point based on VLA-generated local waypoints. To evaluate the system’s obstacle avoidance and error correction capabilities, we randomize the test scenarios: • Spatial Complexity: Each episode is conducted in a randomly generated 10m× 10m map. • Mixed Obstacle Interference: We deploy 7 static and 4 dynamic wandering obstacles to evaluate physical safety responses. • Cross-Embodiment Dynamics Randomization: We apply a ±40% domain randomization to the robot’s kinematic constraints (velocity and acceleration limits) to verify the RL Adapter’s universality. Network Conditions & Degradation Modeling: Assum- ing the cloud VLA generates waypoints at roughly 3 Hz, we evaluate the edge execution under two communication profiles: • Ideal (Fast & Stable Update): A reliable network with a deterministic ∼200 ms delay and zero packet loss. • Non-ideal (Mixed Degradation): A mixed degradation model emulating unreliable wireless communication. It applies three independent perturbations: – Heavy-Tail Latency: To emulate irregular cloud- side queuing and network delays, the one-way cloud-to-edge delay ∆t is sampled from a mixture distribution: ∆t ∼ (1 − q) · U (0.15, 0.25) + q · U (0.5, 1.5) s, where q = 0.1 is the probability of a latency spike. – Stochastic Packet Loss: To simulate signal inter- ference, each transmitted message is subject to an Fig. 2: Qualitative comparison of executed trajectories under the Mixed Degradation network condition. We visualize the actual executed robot trajectories (blue solid lines) and the realigned/stale VLA intents (red thin lines) across four highly challenging dynamic scenarios. RTC generates extremely smooth curves but blindly guides the robot to crash into dynamic obstacles. A2C2 exhibits severe Intent Deviation, causing the robot to oscillate or completely distort the intended path. AsyncShield (Ours) achieves a perfect trade-off: it rigorously restores the VLA’s original intent in free space, and autonomously deviates to ensure safety when the intent becomes dangerous. independent Bernoulli drop with probability p loss = 0.15. – Transient Outages (Black-outs): To emulate phys- ical signal dead zones, we define an outage time ratio r out = 0.05, where during randomly sampled continuous segments D ∼U (1.0, 2.0) s, the packet loss rate is temporarily forced to 100%. Baselines: We adapt two state-of-the-art asynchronous ex- ecution strategies to the embodied navigation domain for comparison: • Naive (Direct Execution): Directly executes the VLA- generated local waypoints as current commands without any spatio-temporal alignment. • RTC [12] (Real-Time Chunking): Focuses on the smooth temporal transition of action chunks. We adapt its core mechanism to smoothly stitch consecutive VLA waypoint chunks to eliminate trajectory jitter. • A2C2 [13] (Asynchronous Action Chunk Correc- tion): Utilizes a lightweight, high-frequency RL correc- tion head to output residual actions based on the latest observation. Evaluation Metrics: • SR (Success Rate): Overall task completion rate. • CTE (Cross-Track Error): Episode-weighted trajec- tory tracking error as defined in [41], measuring the fidelity to the VLA’s original intent. • RER (Risk Exposure Rate): The percentage of time steps where the robot is exposed to high-collision-risk regions (d min < d risk ), reflecting physical safety. • TTG (Time-to-Goal): Navigation efficiency as utilized in [42]. To eliminate survivorship bias, TTG is cal- culated strictly on the shared-success subset across all methods. B. Robustness and Safety Analysis Statistical results across 600 evaluation episodes (summa- rized in Table I) reveal several critical insights regarding asynchronous navigation: 1) Overall Task Completion and Safety: Across both net- work conditions, AsyncShield achieves the highest task com- pletion (SR: 80.0% → 76.7%) and maintains consistently low risk exposure (RER: 1.2% → 1.3%). Baseline methods degrade substantially under injected stochastic latency (e.g., A2C2 drops from 56.7% to 43.3%, and RTC from 40.0% to 30.0%). This discrepancy stems from domain differences: RTC and A2C2 excel in manipulation tasks where intent smoothing and local residual fitting suffice. However, in mobile navigation, large-scale spatial misalignment leads to critical collisions. AsyncShield absorbs this misalignment via SE(2) geometric transformations and an RL adapter, providing robustness to stochastic latency. 2) The CTE Paradox: Tracking Fidelity vs. Safety: A counter-intuitive phenomenon emerges under the Ideal con- TABLE I: Ablation study under the same evaluation protocol as Table I. Note: ‘w/o Safety Constraints’ disables the Lagrangian safety optimization during training/execution. Method Variant Ideal (Fast Update)Non-ideal (Mixed Degradation) SR (%) ↑CTE (m) ↓RER (%) ↓SR (%) ↑CTE (m) ↓RER (%) ↓ AsyncShield (Full)80.00.7171.276.70.7251.3 w/o Temporal Alignment53.30.9151.636.71.1943.6 w/o RL Adapter66.71.2321.353.31.4431.4 w/o Safety Constraints40.00.6814.223.30.6924.7 dition: RTC achieves the lowest trajectory tracking error (CTE = 0.673 m), yet exhibits a low SR (40.0%) and high RER (3.0%). This reveals a critical navigation paradigm: blind smooth tracking does not equate to safety. While RTC seamlessly stitches action chunks, it smoothly executes stale VLA intents, directing the robot toward dynamic obstacles. AsyncShield records a slightly higher CTE (0.717 m) be- cause the PPO-Lagrangian mechanism proactively deviates from unsafe intents for obstacle avoidance. This demon- strates that our “intent-safety decoupling” effectively trades minor tracking fidelity for substantial survival rate gains. 3) Efficiency on Shared-Success Episodes: Under the Ideal condition, the Naive method appears fastest (TTG = 7.92 s). However, this manifests survivorship bias—with an SR of only 20.0%, Naive strictly succeeds in trivial, obstacle- free scenarios. Evaluated strictly on the shared-success subset of complex tasks, latency inflates TTG across all methods. AsyncShield exhibits the smallest efficiency degradation (8.87 s → 9.29 s), outperforming A2C2 (9.64 s), RTC (9.81 s), and Naive (10.25 s), demonstrating robust decision- making under adverse latency conditions. 4) Qualitative Analysis: To intuitively understand behav- ioral differences under the severe Mixed Degradation net- work, we visualize the executed trajectories (blue solid lines) and realigned/stale VLA intents (red thin lines) across four challenging dynamic scenarios in Fig. 2. The visualizations highlight the failure modes of the baselines: • The Pitfall of Blind Smoothness (RTC): In Case 2 and Case 3, RTC successfully eliminates trajectory oscillation, generating a smooth curve. However, lack- ing explicit spatial realignment and obstacle perception, RTC smoothly but blindly steers the robot into dynamic obstacles. This matches its high RER. • The Collapse of Residual Fitting (A2C2): In Case 3 and Case 4, A2C2 exhibits severe Intent Deviation. Without an explicit geometric anchor, a pure RL resid- ual network struggles to implicitly map large time de- lays into spatial coordinate corrections. This cumulative error causes the robot to exhibit severe oscillation or completely distort the intended path. • Adaptive Trade-off between Geometry and Safety (AsyncShield): Our method demonstrates exceptional robustness. Driven by the PPO-Lagrangian optimiza- tion, the system dynamically adjusts the weights between intent-tracking rewards and safety costs. Consequently, AsyncShield rigorously restores the VLA’s original intent in obstacle-free spaces, and au- tonomously prioritizes collision avoidance when stale intents become dangerous, seamlessly blending task fidelity with physical safety. C. Ablation Studies To validate the necessity of AsyncShield’s core architec- tural designs, we conduct targeted ablation experiments (see Table I): • w/o Temporal Alignment (Naive Grafting): Remov- ing the temporal pose buffer forces the agent to rigidly graft stale local waypoints onto the current ego frame. Under Delayed latency, SR plummets from 76.7% to 36.7%, and CTE deteriorates to 1.194 m. This strongly corroborates our core premise, “Latency is Geometry”: under severe stochastic latency, relying on backend RL networks to implicitly fit large-scale coordinate frame discrepancies is highly inefficient without explicit spatial realignment. • w/o RL Adapter: Replacing the RL policy with a classical dynamic window approach (DWA) planner. While DWA effectively avoids obstacles (66.67% SR under Ideal), its success drops to 53.30% under latency, accompanied by severe trajectory deviation (CTE > 1.2 m). This reveals that traditional cost-based planners struggle to balance stale intents with safety, surviving only by aggressively abandoning the VLA’s guidance. Conversely, our CMDP-based RL adapter achieves a tuning-free, optimal trade-off between intent fidelity and physical safety. • w/o Safety Constraints: Removing the Lagrangian- based obstacle penalty relies purely on flow-tracking rewards. Consequently, the CTE decreases as the agent rigidly adheres to the VLA intent, but RER skyrockets, resulting in task failures due to direct collisions. This inverse validation underscores the indispensable role of the PPO-Lag mechanism as a hard safety baseline in asynchronous heterogeneous navigation. D. Cross-Embodiment Validation in Simulation To validate the zero-shot cross-embodiment capability of AsyncShield, we directly deploy the base policy onto two morphologically distinct agents in the OmniSafe simulator: Doggo (a quadruped robot) and Racecar (a vehicle with Ackermann steering constraints), designing significantly dif- ferent kinematic parameters for each category of robot. Specifically, to eliminate the discrepancy in physical volumes across different embodiments without the need for retraining, we introduce a Collision Radius Inflation mechanism. By subtracting a specific inflation compensation value directly from the raw LiDAR scan distances at the perception level (i.e., equivalently pushing obstacles closer in the observation space), the base policy can zero-shot adapt to physical entities with larger collision volumes. As shown in Table I, the differences in Success Rate (SR) and Risk Exposure Rate (RER) across all variants are minimal. This proves that our method can safely and effectively generalize to other embodiments. TABLE I: Zero-Shot Cross-Embodiment Performance Embodiment VariantSR (%) ↑RER (%) ↓ Doggo A78.001.25 Doggo B76.001.31 Racecar A79.001.20 Racecar B76.001.29 E. Real-world Hardware Deployment and Zero-Shot Trans- fer To validate the “plug-and-play” feasibility and zero-shot transferability, we deployed AsyncShield on a Unitree Go2 quadruped robot. The system adopts an edge-cloud archi- tecture: the onboard computer processes 2D LiDAR and odometry at high frequencies, while a cloud GPU runs the VLA models, communicating via real-world Wi-Fi with a baseline round-trip latency of approximately 200ms. Experimental Setup and Multi-Task Validation: We se- lected three SOTA VLA models covering distinct tasks to verify the generalizability of AsyncShield: SocialNav [9] for point-goal navigation, TrackVLA [43] for person-following, and Nav-R 2 [7] for object-goal navigation. For each model, we evaluated performance across four challenging scenarios: Dynamic Crowds, Narrow Corridors, Doorway Entry/Exit, and Extreme Network Jitter. In the jitter scenario, we arti- ficially injected stochastic latencies ranging from 500ms to 1500ms on top of the baseline delay. We conducted 5 trials per scenario, totaling 20 cases for each VLA model. TABLE IV: Real-world Success Rates (20 trials each) demonstrat- ing plug-and-play compatibility across different cloud VLA models. Cloud VLA ModelDirect VLAVLA + AsyncShield SocialNav (Point-goal)6/20 (30%)17/20 (85%) TrackVLA (Following)8/20 (40%)16/20 (80%) Nav-R 2 (Object-goal)5/20 (25%)18/20 (90%) Results and Qualitative Analysis: As shown in Table IV, due to the severe cloud-to-edge latency, naive asynchronous execution (Direct VLA) suffers catastrophic performance degradation (25%–40% SR) in challenging scenarios, fre- quently leading to collisions caused by stale commands. Inte- grating our module immediately restores robust performance (80%–90% SR) without any VLA fine-tuning. In dynamic environments and narrow spaces, AsyncShield allows the robot to autonomously deviate from stale global intents based on real-time local perception. For instance, dur- ing doorway traversal, it recalculates corrective trajectories in real-time to prevent lateral drift or wall collisions induced by command lag. Furthermore, in scenarios with induced high network latency, the robot exhibits no oscillations or freezing; instead, it safely proceeds along the primary intent direction while maintaining active obstacle avoidance, demonstrating remarkable resilience to communication fail- ures. Synthesizing the aforementioned simulation and real- world experiments, the cross-embodiment generalization capability validated in simulation mutually corroborates the hardware deployment performance on the real-world quadruped robot. This demonstrates that AsyncShield can serve as a universal plug-and-play edge module, safely and robustly enabling the real-world physical deployment of cloud-based VLA models without requiring any fine-tuning. V. CONCLUSION This paper proposes AsyncShield, a plug-and-play asyn- chronous control framework designed for cloud-based Vision-Language-Action (VLA) models in mobile navigation tasks. To address the spatio-temporal misalignment caused by cloud-to-edge latency, we discard traditional black-box time- series prediction in favor of a deterministic physical white- box spatial mapping. We restore the original intent of the VLA through geometric transformations and formulate the edge adaptation as a CMDP. Utilizing the PPO-Lagrangian algorithm, we develop a reinforcement learning adapter to achieve an adaptive dynamic trade-off between tracking the true VLA intent and responding to high-frequency LiDAR obstacle avoidance hard constraints. Furthermore, benefiting from a standardized universal sub-goal interface, domain randomization, and perception- level adaptation (i.e., Collision Radius Inflation), Async- Shield demonstrates zero-shot cross-embodiment general- ization capabilities. Simulation and real-world experiments show that, without fine-tuning any cloud-based foundation model weights, the framework exhibits stable generalization capability and effectively improves the success rate and physical safety of asynchronous navigation. Future work will explore extending this analytical geo- metric re-projection mechanism to more complex 3D un- structured environments, and investigate the introduction of lightweight multimodal local perception models at the edge to further enhance the system’s safety redundancy and robustness in dynamic scenarios. REFERENCES [1] J. Zhang, K. Wang, S. Wang, M. Li, H. Liu, S. Wei, Z. Wang, Z. Zhang, and H. Wang, “Uni-navid: A video-based vision-language- action model for unifying embodied navigation tasks,” arXiv preprint arXiv:2412.06224, 2024. [2] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich, “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” 2023. [Online]. Available: https://arxiv.org/abs/2307.15818 [3] O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y. L. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine, “Octo: An open-source generalist robot policy,” 2024. [Online]. Available: https://arxiv.org/abs/2405.12213 [4] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn, “Openvla: An open-source vision-language-action model,” 2024. [Online]. Available: https://arxiv.org/abs/2406.09246 [5] Z. Chu, S. Xie, X. Wu, Y. Shen, M. Luo, Z. Wang, F. Liu, X. Leng, J. Hu, M. Yin, J. Lu, Y. Guo, K. Yang, J. Han, X. Chen, Y. Zhu, Y. Zhao, X. Liu, Y. Yang, Y. He, J. Wang, Y. Cai, T. Zhang, L. Gao, L. Liu, M. Sun, F. Jiang, C. Wang, Z. Liu, H. Pan, H. Han, Z. Gu, K. Yang, J. Zhang, D. Jing, Z. Guan, W. Guo, G. Liu, D. Yang, X. Yang, M. Yang, H. Xing, W. Li, and M. Xu, “Abot-n0: Technical report on the vla foundation model for versatile embodied navigation,” 2026. [Online]. Available: https://arxiv.org/abs/2602.11598 [6] J. Hu, J. Chen, H. Bai, M. Luo, S. Xie, Z. Chen, F. Liu, Z. Chu, X. Xue, B. Ren, X. Wu, M. Xu, and S. Zhang, “Astranav-world: World model for foresight control and consistency,” 2025. [Online]. Available: https://arxiv.org/abs/2512.21714 [7] W. Xiang, H. Zhang, T. Yang, Z. Chu, R. Chu, S. Xie, Y. Yuan, J. Sun, Z. Gu, J. Wang, X. Wu, M. Xu, and Y. Yang, “Nav-r 2 dual-relation reasoning for generalizable open-vocabulary object-goal navigation,” 2025. [Online]. Available: https://arxiv.org/abs/2512.02400 [8] F. Liu, S. Xie, M. Luo, Z. Chu, J. Hu, X. Wu, and M. Xu, “Navforesee: A unified vision-language world model for hierarchical planning and dual-horizon navigation prediction,” 2025. [Online]. Available: https://arxiv.org/abs/2512.01550 [9] Z. Chen, Y. Guo, Z. Chu, M. Luo, Y. Shen, M. Sun, J. Hu, S. Xie, K. Yang, P. Shi, Z. Gu, L. Liu, H. Han, X. Wu, M. Xu, and Y. Zhang, “Socialnav: Training human-inspired foundation model for socially-aware embodied navigation,” 2025. [Online]. Available: https://arxiv.org/abs/2511.21135 [10] X. Xue, J. Hu, M. Luo, S. Xie, J. Chen, Z. Xie, K. Quan, W. Guo, M. Xu, and Z. Chu, “Omninav: A unified framework for prospective exploration and visual-language navigation,” 2026. [Online]. Available: https://arxiv.org/abs/2509.25687 [11] Z. Huang, Y. Zhang, J. Liu, R. Song, C. Tang, and J. Ma, “Tic-vla: A think-in-control vision-language-action model for robot navigation in dynamic environments,” 2026. [Online]. Available: https://arxiv.org/abs/2602.02459 [12] K. Black, M. Y. Galliker, and S. Levine, “Real-time execution of action chunking flow policies,” 2025. [Online]. Available: https://arxiv.org/abs/2506.07339 [13] K. Sendai, M. Alvarez, T. Matsushima, Y. Matsuo, and Y. Iwasawa, “Leave no observation behind: Real-time correction for vla action chunks,” 2025. [Online]. Available: https://arxiv.org/abs/2509.23224 [14] J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in International conference on machine learning. Pmlr, 2017, p. 22–31. [15] N. Hirose, C. Glossop, D. Shah, and S. Levine, “Omnivla: An omni-modal vision-language-action model for robot navigation,” 2025. [Online]. Available: https://arxiv.org/abs/2509.19480 [16] I. Apanasevich, M. Artemyev, R. Babakyan, P. Fedotova, D. Grankin, E. Kupryashin, A. Misailidi, D. Nerus, A. Nutalapati, G. Sidorov, I. Efremov, M. Gerasyov, D. Pikurov, Y. Senchenko, S. Davidenko, D. Kulikov, M. Sultankin, K. Askarbek, O. Shamanin, D. Statovoy, E. Zalyaev, I. Zorin, A. Letkin, E. Rusakov, A. Silchenko, V. Vorobyov, S. Sobolnikov, and A. Postnikov, “Green-vla: Staged vision-language-action model for generalist robots,” 2026. [Online]. Available: https://arxiv.org/abs/2602.00919 [17] X. Sun, W. Si, W. Ni, Y. Li, D. Wu, F. Xie, R. Guan, H.-Y. Xu, H. Ding, Y. Wu, Y. Yue, Y. Huang, and H. Xiong, “Autofly: Vision-language-action model for uav autonomous navigation in the wild,” 2026. [Online]. Available: https://arxiv.org/abs/2602.09657 [18] H. Chi, H.-a. Gao, Z. Liu, J. Liu, C. Liu, J. Li, K. Yang, Y. Yu, Z. Wang, W. Li, et al., “Impromptu vla: Open weights and open data for driving vision-language-action models,” arXiv preprint arXiv:2505.23757, 2025. [19] H. Xie, B. Wen, J. Zheng, Z. Chen, F. Hong, H. Diao, and Z. Liu, “Dynamicvla: A vision-language-action model for dynamic object manipulation,” arXiv preprint arXiv:2601.22153, 2026. [20] M. Shukor, D. Aubakirova, F. Capuano, P. Kooijmans, S. Palma, A. Zouitine, M. Aractingi, C. Pascal, M. Russi, A. Marafioti, et al., “Smolvla: A vision-language-action model for affordable and efficient robotics,” arXiv preprint arXiv:2506.01844, 2025. [21] H. Jang, S. Yu, H. Kwon, H. Jeon, Y. Seo, and J. Shin, “Contextvla: Vision-language-action model with amortized multi-frame context,” arXiv preprint arXiv:2510.04246, 2025. [22] H. Li, Y. Zuo, J. Yu, Y. Zhang, Z. Yang, K. Zhang, X. Zhu, Y. Zhang, T. Chen, G. Cui, et al., “Simplevla-rl: Scaling vla training via reinforcement learning,” arXiv preprint arXiv:2509.09674, 2025. [23] H.-T. L. Chiang, Z. Xu, Z. Fu, M. G. Jacob, T. Zhang, T.-W. E. Lee, W. Yu, C. Schenck, D. Rendleman, D. Shah, F. Xia, J. Hsu, J. Hoech, P. Florence, S. Kirmani, S. Singh, V. Sindhwani, C. Parada, C. Finn, P. Xu, S. Levine, and J. Tan, “Mobility vla: Multimodal instruction navigation with long-context vlms and topological graphs,” 2024. [Online]. Available: https://arxiv.org/abs/2407.07775 [24] J. Lee, H. Shin, and J. Ko, “Iros: A dual-process architecture for real-time vlm-based indoor navigation,” 2026. [Online]. Available: https://arxiv.org/abs/2601.21506 [25] W. Liu, H. Zhao, C. Li, J. Biswas, B. Okal, P. Goyal, Y. Chang, and S. Pouya, “X-mobility: End-to-end generalizable navigation via world modeling,” in 2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, p. 7569–7576. [26] H. Wang, A. H. Tan, A. Fung, and G. Nejat, “X-nav: Learning end-to- end cross-embodiment navigation for mobile robots,” IEEE Robotics and Automation Letters, vol. 11, no. 1, p. 698–705, 2025. [27] K. Yang, T. Zhang, Z. Wang, Z. Chu, X. Wu, Y. Cai, and M. Xu, “Ce-nav: Flow-guided reinforcement refinement for cross-embodiment local navigation,” arXiv preprint arXiv:2509.23203, 2025. [28] X. Chen, S. Xie, Z. Gu, L. Jia, M. Luo, F. Liu, Z. Chu, Y. Shen, X. Wu, and M. Xu, “Explore like humans: Autonomous exploration with online sg-memo construction for embodied agents,” 2026. [Online]. Available: https://arxiv.org/abs/2604.19034 [29] T. Z. Zhao, V. Kumar, S. Levine, and C. Finn, “Learning fine-grained bimanual manipulation with low-cost hardware,” 2023. [Online]. Available: https://arxiv.org/abs/2304.13705 [30] C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” 2024. [Online]. Available: https://arxiv.org/abs/2303.04137 [31] D. Jing, G. Wang, J. Liu, W. Tang, Z. Sun, Y. Yao, Z. Wei, Y. Liu, Z. Lu, and M. Ding, “Mixture of horizons in action chunking,” arXiv preprint arXiv:2511.19433, 2025. [32] J. Tang, Y. Sun, Y. Zhao, S. Yang, Y. Lin, Z. Zhang, J. Hou, Y. Lu, Z. Liu, and S. Han, “Vlash: Real-time vlas via future- state-aware asynchronous inference,” 2025. [Online]. Available: https://arxiv.org/abs/2512.01031 [33] T. Zou, H. Zeng, Y. Nong, Y. Li, K. Liu, H. Yang, X. Ling, X. Li, and L. Ma, “Asynchronous fast-slow vision-language-action policies for whole-body robotic manipulation,” arXiv preprint arXiv:2512.20188, 2025. [34] N. Hirose, C. Glossop, D. Shah, and S. Levine, “Asyncvla: An asynchronous vla for fast and robust navigation on the edge,” arXiv preprint arXiv:2602.13476, 2026. [35] H. Wang, A. H. Tan, and G. Nejat, “Navformer: A transformer ar- chitecture for robot target-driven navigation in unknown and dynamic environments,” IEEE Robotics and Automation Letters, vol. 9, no. 8, p. 6808–6815, 2024. [36] S. Jung, J. Lee, X. Meng, B. Boots, and A. Lambert, “V-strong: Visual self-supervised traversability learning for off-road navigation,” in 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2024, p. 1766–1773. [37] J. Gao, X. Pang, Q. Liu, and Y. Li, “Hierarchical reinforcement learning for safe mapless navigation with congestion estimation,” in 2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, p. 8849–8855. [38] H. Taheri, S. R. Hosseini, and M. A. Nekoui, “Deep reinforcement learning with enhanced ppo for safe mobile robot navigation,” arXiv preprint arXiv:2405.16266, 2024. [39] X. Lin, Y. Huang, F. Chen, and B. Englot, “Decentralized multi- robot navigation for autonomous surface vehicles with distributional reinforcement learning,” in 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2024, p. 8327–8333. [40] W. Chen, J. Onyejizu, L. Vu, L. Hoang, D. Subramanian, K. Kar, S. Mishra, and S. Paternain, “Adaptive primal-dual method for safe reinforcement learning,” arXiv preprint arXiv:2402.00355, 2024. [41] A. Penumarti and J. Shin, “Global uncertainty-aware planning for magnetic anomaly-based navigation,” 2024. [Online]. Available: https://arxiv.org/abs/2409.10366 [42] V. Rajagopal, K. W. K. Mudiyanselage, G. D. Seneviratne, P. A. Sankaralingam, M. Elnoor, J. Liang, R. Chandra, and D. Manocha, “Dr. nav: Semantic-geometric representations for proactive dead-end recovery and navigation,” 2025. [Online]. Available: https://arxiv.org/abs/2511.12778 [43] S. Wang, J. Zhang, M. Li, J. Liu, A. Li, K. Wu, F. Zhong, J. Yu, Z. Zhang, and H. Wang, “Trackvla: Embodied visual tracking in the wild,” arXiv preprint arXiv:2505.23189, 2025.