Paper deep dive
RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation
Boxiong Wang, Hui Kang, Geng Sun, Jiahui Li, Chao Yu, Daxin Tian
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited corrective supervision for interactive closed-loop execution. Reinforcement learning (RL) offers a promising solution, while its effectiveness is constrained by inefficient use of samples, long-tailed scene distributions, and policy distribution shift during optimization. To this end, we propose RecoverFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies. Specifically, RecoverFly adapts token-level RL for stable optimization of grammar-constrained autoregressive UAV actions, revisits unresolved failure cases to strengthen corrective learning and sample utilization, and combines a two-stage long-tail scene curriculum with reference-policy regularization to improve scene adaptation while preserving acquired capabilities. Experiments on the TravelUAV benchmark demonstrate that RecoverFly achieves the best performance on the seen, unseen-map, and unseen-object splits. Moreover, compared to the AerialVLA initialization, RecoverFly improves success rate by 3.12 to 8.37 percentage points under a total rollout budget of about 30\% of the training-set size, validating its effectiveness, robustness, and generalization capabilities.
Tags
Links
- Source: https://arxiv.org/abs/2608.09467v1
- Canonical: https://arxiv.org/abs/2608.09467v1
Trouble viewing inline? Open PDF directly →
Full Text
45,591 characters extracted from source content.
Expand or collapse full text
RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation Boxiong Wang1,4, Hui Kang1, Geng Sun1, Jiahui Li111footnotemark: 1, Chao Yu2,4, Daxin Tian3,4 Corresponding authors: Geng Sun and Jiahui Li. Preprint. Under Review. Abstract Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited corrective supervision for interactive closed-loop execution. Reinforcement learning (RL) offers a promising solution, while its effectiveness is constrained by inefficient use of samples, long-tailed scene distributions, and policy distribution shift during optimization. To this end, we propose RecoverFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies. Specifically, RecoverFly adapts token-level RL for stable optimization of grammar-constrained autoregressive UAV actions, revisits unresolved failure cases to strengthen corrective learning and sample utilization, and combines a two-stage long-tail scene curriculum with reference-policy regularization to improve scene adaptation while preserving acquired capabilities. Experiments on the TravelUAV benchmark demonstrate that RecoverFly achieves the best performance on the seen, unseen-map, and unseen-object splits. Moreover, compared to the AerialVLA initialization, RecoverFly improves success rate by 3.12 to 8.37 percentage points under a total rollout budget of about 30% of the training-set size, validating its effectiveness, robustness, and generalization capabilities. Introduction Vision-language navigation (VLN) requires embodied agents to associate natural-language instructions with visual observations and execute trajectories toward semantic goals (Anderson et al. 2018; Krantz et al. 2020). Unmanned aerial vehicle (UAV) VLN extends this problem to large-scale three-dimensional (3D) environments, where an aerial agent searches for a language-described object while continuously adjusting forward motion, altitude, yaw, and termination decisions (Liu et al. 2023; Fan et al. 2023; Wang et al. 2025). Recent vision-language-action (VLA) models provide a promising end-to-end interface from onboard observations and instructions to executable controls, reducing reliance on separately engineered perception, planning, and control modules (Zitkovich et al. 2023; Kim et al. 2025; Xu et al. 2026). However, closed-loop UAV navigation remains difficult for two reasons. First, rapidly changing viewpoints, occlusion, and long trajectories cause local errors to compound into collisions, navigation drift, premature landing, timeouts, or target misses (Jiang et al. 2025; Xu et al. 2026). Behavior-cloned policies are especially vulnerable because they are trained on expert demonstrations while they may need to act on their own error-induced states at test time (Ross et al. 2011). Second, aerial datasets are costly and highly imbalanced across scenes, causing standard sampling to emphasize frequent scenes while neglecting rare ones, thus weakening transfer to unseen maps and difficult routes (Lin et al. 2025; Jiang et al. 2025). These problems are coupled since rare or difficult cases that most need correction are also least likely to be revisited during standard training. Figure 1: Comparison of training paradigms for UAV-VLN. Behavior cloning lacks corrective learning, and standard RL may discard informative failures during online sampling. RecoverFly revisits unresolved tasks and learns recovery behaviors from interaction feedback. Existing UAV-VLN systems mitigate the challenges of long-horizon planning through waypoint prediction, hierarchical planners, memory mechanisms, or external perception modules (Wang et al. 2025; Zhang et al. 2025a; Ding et al. 2026; Ning et al. 2026; Jiang et al. 2025). Although end-to-end UAV VLA policies such as AerialVLA (Xu et al. 2026) can generate flight actions directly, their optimization relies primarily on behavior cloning, which provides no specific corrective signal for failures encountered during closed-loop execution. While reinforcement learning (RL) can leverage interactive rewards to improve policy behavior in specific states, its direct application to autoregressive action policies introduces four issues. First, a failure event may result from previous actions rather than from the instantaneous actions. Second, unresolved failure cases with substantial learning value may rapidly disappear from subsequent sampling batches. Third, scene imbalance may concentrate policy optimization on frequently observed environments. Finally, unconstrained policy updates may degrade previously acquired navigation capabilities (Ouyang et al. 2022; Lin et al. 2025). As a result, a practical post-training method should address these issues jointly rather than treating RL as a generic second-stage optimizer. To address the aforementioned challenges, we propose RecoverFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies. Figure 1 conceptually compares RecoverFly with behavior cloning and standard RL for end-to-end UAV-VLN. Specifically, RecoverFly adapts RL to grammar-constrained UAV action tokens, enabling stable closed-loop optimization of executable autoregressive actions. Furthermore, RecoverFly incorporates failure-aware replay to retain and revisit unresolved cases with substantial learning value, thereby improving sample utilization and strengthening corrective learning during closed-loop execution. Building on this, a two-stage curriculum progressively adjusts the training distribution toward long-tailed scenes, while reference-policy Kullback-Leibler (KL) regularization limits excessive policy deviation and preserves previously acquired capabilities. Consequently, RecoverFly integrates fine-grained action optimization, failure-oriented learning, long-tail scene adaptation, and stable policy updating into a unified post-training process, thereby improving navigation performance and generalization across diverse environments. The main contributions of this paper are as follows. • RL Framework for End-to-End UAV-VLA. We develop RecoverFly as an integrated closed-loop RL post-training framework. Specifically, RecoverFly preserves the pretrained textual action interface and adapts token-level RL to grammar-constrained UAV commands, enabling autoregressive action generation to be optimized through online navigation feedback. • Failure-Aware Replay Learning. We propose a failure-aware replay mechanism that retains and revisits unresolved navigation cases with substantial learning value. This design improves the utilization of sparse failure feedback without applying policy updates to stale trajectories. • Long-Tail Adaptation with Stable Policy Updates. We introduce a two-stage long-tail scene curriculum together with reference-policy KL regularization. Building on this design, RecoverFly strengthens learning from underrepresented scenes while limiting policy distribution shift and preserving previously acquired navigation capabilities. Related Work VLN for UAVs Early VLN methods established cross-modal instruction grounding and action prediction (Anderson et al. 2018; Krantz et al. 2020). AerialVLN (Liu et al. 2023) and AVDN (Fan et al. 2023) extended this setting to outdoor aerial navigation and dialog-conditioned target search, while UAV-ON (Xiao et al. 2025) and OpenFly (Gao et al. 2026) broadened it toward open-world object goals and larger-scale aerial datasets. To execute long-horizon flights, TravelUAV (Wang et al. 2025) predicts continuous waypoints, whereas CityNavAgent (Zhang et al. 2025b), SkyVLN (Li et al. 2025), TypeFly (Chen et al. 2025), and training-free VLM approaches (Hu et al. 2025) combine language reasoning with memory, trajectory generation, model-based control, or program synthesis. Moreover, NavFoM (Zhang et al. 2025a) and LongFly (Jiang et al. 2025) further improve generalization through large-scale pretraining and spatiotemporal modeling. These methods substantially improve semantic planning, while most retain explicit interfaces between high-level reasoning and low-level flight execution, through which prediction and control errors can accumulate. VLA Policies for UAV-VLN RT-2 (Zitkovich et al. 2023) and OpenVLA (Kim et al. 2025) show that robotic actions can be represented in language-compatible output spaces, allowing pretrained vision-language representations to support embodied control. This paradigm has been extended to fine-grained imitation, racing, and cognitive UAV control (Wang et al. 2026; Serpiva et al. 2025; Lykov et al. 2025). Furthermore, AerialVLA (Xu et al. 2026) maps onboard observations and linguistic prompts directly to continuous 3-degree-of-freedom (DoF) controls and an intrinsic landing decision, reducing reliance on oracle waypoints, external detectors, and separate landing modules. However, these UAV-VLA policies are trained mainly by supervised behavior cloning. Thus, they learn strong expert-action priors while receiving limited corrective learning from collisions, moving-away behavior, stuck actions, early stops, and timeouts induced by their own closed-loop execution (Ross et al. 2011). RL in UAV-VLN RL has also been explored directly in language-conditioned UAV control. SuReAL (Blukis et al. 2019) combines supervised position prediction with RL-based continuous control and evaluates the resulting policy on a physical quadcopter. More recently, HTNav (Fan et al. 2026) integrates staged imitation and RL with tiered decision making for urban aerial VLN. OpenVLN (Lin et al. 2025) employs value-based waypoint rewards and KL-regularized policy updates, while FlightGPT (Cai et al. 2025) combines supervised fine-tuning with group relative policy optimization (GRPO)-style optimization for goal accuracy, reasoning quality, and output compliance. While these studies establish the value of RL for aerial navigation, online post-training of autoregressive end-to-end UAV-VLA policies remains underexplored. Specifically, existing methods fail to comprehensively address sparse closed-loop feedback, adaptation to long-tail scenarios, and the mitigation of policy drift. RecoverFly establishes a unified framework based on the standard proximal policy optimization (PPO) algorithm (Schulman et al. 2017) to address these complex and coupled challenges. Figure 2: Overview of RecoverFly. RecoverFly combines token-level PPO with dynamic failure replay to learn corrective behaviors from closed-loop interaction while preserving on-policy rollouts. Furthermore, a two-stage long-tail scene curriculum and stage-wise reference-policy KL regularization improve rare-scene adaptation and constrain policy distribution shift. Method Overview We propose RecoverFly, a failure-aware online RL post-training framework for end-to-end UAV-VLA policies in UAV-VLN, as illustrated in Fig. 2. Building on the native token-level VLA optimization in RLinf (Yu et al. 2025), RecoverFly preserves the pretrained textual action interface and integrates dynamic failure replay, a two-stage long-tail scene curriculum, and stage-wise reference-policy KL regularization to strengthen corrective learning and rare-scene adaptation while constraining policy drift. Textual UAV Policy and Online RL Objective Textual Action Representation. We consider end-to-end UAV-VLN control, where the agent receives a natural-language instruction x and a visual observation oto_t at time step t. We denote the resulting policy context by ht=(x,ot)h_t=(x,o_t), and the policy directly generates an action-token sequence t=(zt,1,…,zt,Kt)z_t=(z_t,1,…,z_t,K_t), where K is the maximum sequence length and Kt≤K_t≤ K is the realized length. The autoregressive policy factorizes as πθ(t∣ht)=∏k=1Ktπθ(zt,k∣ht,zt,<k). _θ(z_t h_t)= _k=1^K_t _θ(z_t,k h_t,z_t,<k). (1) Following AerialVLA (Xu et al. 2026), the action grammar contains three numerical control tokens followed by an optional LAND token. The decoded 3-DoF control is t=⟨Δxt,Δzt,Δψt⟩a_t= x_t, z_t, _t , corresponding to forward progression, vertical movement, and yaw adjustment, respectively. Each control dimension is uniformly quantized into 99 bins over [0,5][0,5], [−5,5][-5,5], and [−π,π][-π,π], respectively. A deterministic decoder maps valid sequences to continuous controls. Moreover, we use mt,k∈0,1m_t,k∈\0,1\ to select valid action-token positions after padding to length K, and define Mt=max(1,∑k=1Kmt,k)M_t= (1, _k=1^Km_t,k). The same action grammar constrains rollout generation and policy optimization, and a sequence that still fails parsing is mapped to a safe no-op action and receives an invalid-action penalty. Event-Aware Reward and Advantage Estimation. In UAV-VLN, success and failure events provide direct supervision, whereas they occur sparsely. Therefore, RecoverFly combines dense distance progress with event-specific reward, which is as follows: rt=clip(κp(Dt−1−Dt),rminprog,rmaxprog)+∑e∈ℰRe[et=e],r_t=clip\! ( _p(D_t-1-D_t),r_ ^prog,r_ ^prog )+ _e R_eI[e_t=e], (2) where DtD_t is the Euclidean distance to the target and ℰ=success,collision,stuck,away,early-stop,timeoutE=\ success, collision, stuck, away, early-stop, timeout\ denotes the event set. In long-horizon UAV navigation, the success or failure of a trajectory may become observable only several steps after the key actions, making the attribution of delayed rewards to earlier decisions essential. Therefore, we apply PPO with generalized advantage estimation (GAE) (Schulman et al. 2016), which propagate delayed returns to preceding actions while constraining the magnitude of each policy update to support effective credit assignment and stable policy optimization. Moreover, we attach a value head Vϕ(ht)V_φ(h_t) to the base VLA policy and optimize it using the standard clipped PPO value loss ℒVL_V. Consequently, let dt∈0,1d_t∈\0,1\ indicate whether the transition terminates the episode. For a rollout ending at step T, we compute A^t=∑l=0T−t−1(γλ)lδt+l,δt=rt+γ(1−dt)Vϕ(ht+1)−Vϕ(ht). split& A_t= _l=0^T-t-1(γλ)^l _t+l,\\ & _t=r_t+γ(1-d_t)V_φ(h_t+1)-V_φ(h_t). split (3) where γ is the discount factor and λ controls the bias–variance trade-off. For terminal transitions, dt=1d_t=1 removes the bootstrap value, and non-terminal rollout truncations bootstrap from the final value estimate. Token-Level Policy Optimization Backbone. Standard continuous-action PPO is not directly applicable, as the VLA policy parameterizes token probabilities rather than an explicit density over decoded commands. A sequence-level alternative forms a joint ratio from autoregressive token probabilities, while this applies a single importance weight and clipping decision to the entire action sequence. Thus, RecoverFly applies the clipped surrogate separately to each valid action token and assigns all tokens encoding the same action the shared action-level advantage A^t A_t. For the k-th action token, the importance ratio is ρt,k(θ)=πθ(zt,k∣ht,zt,<k)πold(zt,k∣ht,zt,<k) _t,k(θ)= _θ(z_t,k h_t,z_t,<k) _old(z_t,k h_t,z_t,<k), where πold _old denotes the rollout policy used to collect the current batch. Defining ρ¯t,k(θ)=clip(ρt,k(θ),1−ϵ,1+ϵ) ρ_t,k(θ)=clip ( _t,k(θ),1-ε,1+ε ), the token-level objective is JPPOtoken(θ)=t[1Mt∑k=1Kmt,kmin(ρt,k(θ)A^t,ρ¯t,k(θ)A^t)],J_PPO^token(θ)=E_t\! [ 1M_t _k=1^Km_t,k \! ( _t,k(θ) A_t, ρ_t,k(θ) A_t ) ], (4) where ϵε is the PPO clipping threshold. This formulation enables token-wise updates in the native autoregressive action space while preserving the action-step learning signal and avoiding a shared sequence-level clipping decision. Dynamic Failure Replay Although token-level PPO enables optimization of autoregressive VLA actions, ordinary on-policy sampling can rapidly dilute difficult failures, leading to insufficient learning of high-value samples. Prior work improves sample efficiency by relabeling unsuccessful experience in hindsight (Andrychowicz et al. 2017) or by replaying levels with high estimated learning potential (Jiang et al. 2021). In contrast, RecoverFly introduces a dynamic failure replay mechanism, which stores unresolved task initializations rather than old trajectories and regenerates each rollout with the current policy, thereby retaining on-policy PPO while repeatedly learning informative failures. Specifically, RecoverFly maintains a dynamic failure pool ℳ=ξiM=\ _i\, where each entry is represented as ξi=(idi,ci,fi,σi,ni). _i=(id_i,c_i,f_i, _i,n_i). (5) where idiid_i, cic_i, and fif_i denote the task identifier, scene, and failure type, respectively. The state σi∈active,solved,dropped _i∈\active,solved,dropped\ records the active state, and nin_i counts unsuccessful replay attempts. At each online reset, RecoverFly samples fresh and failed task instances according to a replay ratio η. After each rollout, failed fresh instances are added to the pool or reactivated within it, successful replays are marked as solved, and entries that remain unsuccessful after NmaxN_ replay attempts are marked as dropped. As a result, the dynamic replay pool evolves with policy updates and increasingly focuses the training effort on unresolved failure cases. Two-Stage Long-Tail Scene Curriculum Although failure replay improves the use of observed failures, it does not correct scene-level imbalance in sampling. To mitigate this issue, RecoverFly introduces a two-stage long-tail scene curriculum that adjusts scene-sampling weights across training stages to balance the experience obtained from different scenarios. Curriculum learning controls the examples or distributions presented over training to improve optimization and generalization (Bengio et al. 2009). In RecoverFly, the curriculum acts on scene frequency rather than trajectory difficulty. Specifically, stage I samples scenes according to their empirical frequencies as follows: P1(c)=Nc∑c′∈Nc′,P_1(c)= N_c _c N_c , (6) where C denotes the set of training scenes and NcN_c denotes the number of training episodes associated with scene c. This stage expands the basic navigation capability under the original data distribution. Subsequently, Stage I continues from Stage I and replaces proportional scene sampling with an equal quota for each scene. Compared with the empirical distribution used in Stage I, this allocation increases rare-scene training frequency and prevents certain scenes from dominating online sampling. Moreover, the curriculum couples balanced scene exposure with failure correction to improve policy generalization. Due to space limitations, details of the rare-scene partition and stage I sampling strategy are provided in Appendix A.1. Reference-Policy KL Regularization Constraining policy updates is a standard mechanism for stabilizing policy optimization (Schulman et al. 2015). RecoverFly adopts this principle through a stage-wise frozen reference policy that persistently anchors online adaptation. Specifically, Stage I uses the initial VLA policy, while Stage I uses the final Stage I policy. For stage s∈1,2s∈\1,2\, let pθ,t,k(⋅)=πθ(⋅∣ht,zt,<k)p_θ,t,k(·)= _θ(· h_t,z_t,<k) and pref,t,k(s)(⋅)=πref(s)(⋅∣ht,zt,<k)p_ref,t,k^(s)(·)= _ref^(s)(· h_t,z_t,<k) denote the corresponding reference distribution over valid action tokens. The token-level reference loss is as follows: ℒKL(s)=t[1Mt∑k=1Kmt,kDKL(pθ,t,k∥pref,t,k(s))].L_KL^(s)=E_t\! [ 1M_t _k=1^Km_t,kD_KL\! (p_θ,t,k\,\|\,p_ref,t,k^(s) ) ]. (7) The two-stage constraints anchor online updates to the initial VLA policy distribution and the capabilities acquired during Stage I, respectively. This anchoring mechanism discourages excessive cross-stage drift and moderates the cross-stage adaptation–retention trade-off. As a result, RecoverFly aims to minimize the policy, value, and reference losses jointly, and the final training objective is defined as follows: ℒ(θ,ϕ)=−JPPOtoken(θ)+cvℒV(ϕ)+βℒKL(s),L(θ,φ)=-J_PPO^token(θ)+c_vL_V(φ)+ _KL^(s), (8) where cvc_v controls value regression and β controls the reference-policy constraint. Experiments Method Full Easy Hard NE↓ SR↑ OSR↑ SPL↑ NE↓ SR↑ OSR↑ SPL↑ NE↓ SR↑ OSR↑ SPL↑ Human 14.15 94.51 94.51 77.84 11.68 95.44 95.44 76.19 17.16 93.37 93.37 79.85 Random Action 222.20 0.14 0.21 0.07 142.07 0.26 0.39 0.13 320.12 0.00 0.00 0.00 Fixed Action 188.61 2.27 8.16 1.40 121.36 3.48 11.48 2.14 270.69 0.79 4.09 0.49 CMA 135.73 8.37 18.72 7.90 84.89 11.48 24.52 10.68 197.77 4.57 11.65 4.51 TravelUAV-DA 98.66 17.45 48.87 15.76 66.40 20.26 51.23 18.10 138.04 14.02 45.98 12.90 NavFoM 93.05 29.17 49.24 25.03 58.98 32.91 53.16 27.87 143.83 23.58 43.40 20.80 LongFly 60.02 36.39 65.87 31.07 38.10 38.52 71.90 31.24 85.20 33.94 58.94 30.88 AerialVLA 65.88 47.96 57.69 38.54 43.76 49.30 61.30 37.14 93.16 46.30 53.23 40.26 RecoverFly (Ours) 54.96 56.33 65.40 45.98 37.60 56.28 67.39 42.91 76.36 56.38 62.94 49.75 ± 1.19±\,1.19 ± 0.15±\,0.15 ± 0.80±\,0.80 ± 1.13±\,1.13 ± 1.31±\,1.31 ± 0.48±\,0.48 ± 0.98±\,0.98 ± 0.58±\,0.58 ± 1.07±\,1.07 ± 0.88±\,0.88 ± 0.59±\,0.59 ± 1.80±\,1.80 Table 1: Comparison on the Test Seen set. RecoverFly reports mean and standard deviation over three evaluation seeds. NE is in meters, with other metrics as percentages (%). Bold and underline indicate the best and second-best results, respectively. Human performance is provided for reference only. Method Full Easy Hard NE↓ SR↑ OSR↑ SPL↑ NE↓ SR↑ OSR↑ SPL↑ NE↓ SR↑ OSR↑ SPL↑ Random Action 202.98 0.00 0.00 0.00 158.46 0.00 0.00 0.00 265.88 0.00 0.00 0.00 Fixed Action 180.47 0.52 2.61 0.39 132.89 0.89 4.28 0.67 247.72 0.00 0.25 0.00 CMA 141.68 2.30 10.02 2.16 102.29 3.57 14.26 3.33 197.35 0.50 4.03 0.50 TravelUAV 138.80 4.18 20.77 3.84 102.94 4.63 22.82 4.24 189.46 3.53 17.88 3.28 NavFoM 125.10 6.30 18.95 5.68 102.41 6.77 20.07 6.04 170.58 5.36 15.71 4.97 LongFly 108.32 11.27 30.27 9.32 78.56 12.96 34.31 10.32 148.10 9.02 24.88 7.98 AerialVLA 67.42 37.58 52.92 28.22 44.99 41.89 58.47 29.72 99.11 31.49 45.09 26.11 RecoverFly (Ours) 58.88 42.97 60.09 31.83 44.02 46.88 63.04 32.44 79.88 37.45 55.92 30.98 ± 2.09±\,2.09 ± 1.76±\,1.76 ± 1.01±\,1.01 ± 1.25±\,1.25 ± 0.99±\,0.99 ± 1.24±\,1.24 ± 0.37±\,0.37 ± 0.72±\,0.72 ± 4.95±\,4.95 ± 2.65±\,2.65 ± 2.27±\,2.27 ± 2.05±\,2.05 Table 2: Comparison on the Test Unseen Map set. RecoverFly reports mean and standard deviation over three seeds. NE is in meters, and the other metrics are in percentage. Bold and underline indicate the best and second-best results, respectively. Method Full Easy Hard NE↓ SR↑ OSR↑ SPL↑ NE↓ SR↑ OSR↑ SPL↑ NE↓ SR↑ OSR↑ SPL↑ Random Action 260.14 0.16 0.16 0.16 174.10 0.48 0.48 0.48 302.96 0.00 0.00 0.00 Fixed Action 212.84 3.66 9.54 2.16 151.66 6.70 13.88 3.72 243.29 2.14 7.38 1.38 CMA 155.79 9.06 16.06 8.68 102.92 14.83 22.49 13.90 182.09 6.19 12.86 6.08 TravelUAV 118.11 22.42 46.90 20.51 86.12 24.40 49.28 22.03 134.03 21.43 45.71 19.75 NavFoM 108.04 29.83 47.99 27.20 70.51 32.54 50.72 29.54 133.01 28.03 46.18 25.64 LongFly 66.74 43.87 64.56 38.39 54.84 38.01 56.84 31.36 57.07 50.25 74.16 45.27 AerialVLA 61.45 56.60 64.86 46.61 45.72 56.94 64.11 43.76 69.27 56.43 65.24 48.03 RecoverFly (Ours) 53.42 59.72 68.31 51.34 35.49 63.80 72.09 50.97 62.34 57.70 66.43 51.52 ± 0.97±\,0.97 ± 0.81±\,0.81 ± 1.66±\,1.66 ± 0.77±\,0.77 ± 1.94±\,1.94 ± 1.20±\,1.20 ± 2.41±\,2.41 ± 1.58±\,1.58 ± 0.52±\,0.52 ± 1.45±\,1.45 ± 1.95±\,1.95 ± 1.02±\,1.02 Table 3: Comparison on the Test Unseen Object set. RecoverFly reports mean and standard deviation over three seeds. NE is in meters, and the other metrics are in percentage. Bold and underline indicate the best and second-best results, respectively. Framework Components Success Rate (%) ID Failure Replay KL Regularization Two-Stage Curriculum Seen Unseen Map Unseen Object Avg. −- −- −- −- 47.96 37.58 56.60 47.38 1 × × × 48.17 (+0.21↑)(+0.21\, ) 42.90 (+5.32↑)(+5.32\, ) 56.44 (−0.16↓)(-0.16\, ) 49.17 (+1.79↑)(+1.79\, ) 2 ✓ × × 56.21 (+8.25↑)(+8.25\, ) 31.21 (−6.37↓)(-6.37\, ) 59.78 (+3.18↑)(+3.18\, ) 49.07 (+1.69↑)(+1.69\, ) 3 ✓ ✓ × 55.01 (+7.05↑)(+7.05\, ) 37.47 (−0.11↓)(-0.11\, ) 62.32 (+5.72↑)(+5.72\, ) 51.60 (+4.22↑)(+4.22\, ) 4 ✓ ✓ ✓ 56.28 (+8.32↑)(+8.32\, ) 44.89 (+7.31↑)(+7.31\, ) 60.41 (+3.81↑)(+3.81\, ) 53.86 (+6.48↑)(+6.48\, ) Table 4: Incremental ablation of the RecoverFly framework. A checkmark indicates that the corresponding component is enabled. We report full-split success rate (SR, %) and the unweighted mean across the three evaluation splits. Parenthesized values are absolute changes relative to AerialVLA. Best results are in bold, and second-best results are underlined. Experimental Setup Dataset. We evaluate RecoverFly on the TravelUAV dataset (Wang et al. 2025). Following the UAV-Need-Help task adopted by AerialVLA, we use 7922 trajectories for training and evaluate on 1,418 Seen, 958 Unseen Map, and 629 Unseen Object trajectories. Based on the benchmark settings, trajectories that are shorter than 250 meters are categorized as easy, while the remaining trajectories form the hard subset. We report results on the full test set as well as on both difficulty subsets. Furthermore, as mentioned earlier, the TravelUAV training set exhibits a long-tailed scene distribution, as detailed in Appendix A.1. Metrics. We report navigation error (NE), success rate (SR), oracle success rate (OSR), and success weighted by path length (SPL). Specifically, NE is the final Euclidean distance to the destination. SR measures successful termination within the target region, including a correct LAND output or maintaining near-zero movement for 10 consecutive steps within target area. OSR records whether the executed trajectory ever enters that region, and SPL jointly evaluates task completion and path efficiency. Lower NE and higher SR, OSR, and SPL indicate better performance. Implementation Details. RecoverFly initializes from AerialVLA, which combines the OpenVLA-7B (Kim et al. 2025) backbone with a LoRA (Hu et al. 2022) adapter, and adds a value head for actor-critic optimization. RL post-training further optimizes the LoRA adapter. Including failure replays, the total rollout budget is approximately 30% of the training-set size. The training uses 8 × NVIDIA A100 (80GB) GPUs and takes about 21 hours. We implement RecoverFly on RLinf (Yu et al. 2025) by integrating AirSim (Shah et al. 2017), the TravelUAV environment, and the AerialVLA action-token pipeline. In addition, the details about the rest configurations can be found in Appendix A.2. Baselines. We compare RecoverFly with heuristic controls, task-specific UAV-VLN methods, generalist navigation models, and an end-to-end VLA baseline. Unlike RecoverFly, these methods either rely on fixed policies, specialized planning modules, or behavior-cloning-based training. For all baselines, we report results directly from the corresponding publications on the same evaluation splits. • Heuristic Methods. Random Action samples controls without using visual or language inputs, while Fixed Action executes predefined commands. • Task-Specific UAV-VLN Models. CMA (Anderson et al. 2018) is a recurrent cross-modal navigation baseline, while TravelUAV (Wang et al. 2025) combines multimodal features with hierarchical trajectory decoders and an external target detector. TravelUAV-DA extends TravelUAV by aggregating additional corrective trajectories. • Generalist Navigation Models. NavFoM (Zhang et al. 2025a) uses a navigation foundation model with a dedicated trajectory-planning head, and LongFly (Jiang et al. 2025) introduces explicit spatiotemporal modeling for long-horizon UAV navigation. • End-to-End VLA Baseline. AerialVLA (Xu et al. 2026) autoregressively generates action tokens for UAV control, and its checkpoint initializes RecoverFly, making their comparison a direct evaluation of RL post-training. Performance Comparison Performance on Seen Environments. As shown in Table 1, RecoverFly improves AerialVLA by 8.378.37 percentage points in Full-set SR, and the gain increases to 10.0810.08 points on Hard trajectories. The improvement is consistently larger on Hard than Easy trajectories across all four metrics. This pattern indicates that RL post-training is most beneficial when navigation errors accumulate over long horizons, rather than only refining short-range control. The nearly matched gains in SR and SPL further show that the additional successes are achieved without sacrificing path efficiency. Generalization to Unseen Maps. As shown in Table 2, RecoverFly improves AerialVLA by an absolute 5.395.39 percentage points in Full-set SR. Unlike the Easy subset, where NE changes only slightly, Hard-set NE decreases by 19.2319.23 m, indicating that RL post-training is especially effective when navigation errors accumulate over long routes. Moreover, the concurrent gains in OSR and SR show that RecoverFly improves both target-region reachability and successful completion on unseen maps. Furthermore, the stronger performance on long, previously unseen routes suggests that the learned corrective behavior transfers beyond the environment represented in the training set. Generalization to Unseen Objects. As shown in Table 3, RecoverFly achieves the best Full-set results across all four metrics and raises SR to 59.72%59.72\%. On Hard trajectories, LongFly achieves lower NE and higher OSR, whereas RecoverFly obtains higher SR and SPL. This contrast shows that RecoverFly converts target encounters into successful and efficient completion more reliably, rather than only reaching the vicinity of an unseen object. Since RL post-training introduces no additional object annotations or external detector, the gain suggests that it strengthens the mapping from open-vocabulary representations to approach and landing actions for novel objects inherited from AerialVLA. Sampling Strategy Seen Unseen Map Unseen Object Avg. Uniform Sampling 50.71 41.23 52.31 48.08 Original Distribution 56.13 40.92 59.14 52.06 Two-Stage Curriculum 56.28 44.89 60.41 53.86 Table 5: Scene-sampling strategies comparison using full-split SR. Uniform Sampling assigns equal scene quotas for one stage, and Original Distribution preserves empirical proportions for two stage. Best and second-best results are bold and underlined, respectively. Ablation Study We conduct ablation studies across all test splits. All ablations use seed 1. Table 4 incrementally adds components, whereas Table 5 isolates the scene sampling strategy. Appendices B and C analyze replay behavior and compare the performance of token-level and sequence-level PPO. Effect of RL Post-Training and Failure Replay. We compare the AerialVLA baseline with ID 1, which applies a token-level adapted base PPO. As illustrated in Table 4, ID 1 leaves Seen SR nearly unchanged, increases Unseen Map SR by 5.32 percentage points, and decreases Unseen Object SR by only 0.16 points. Accordingly, its average SR improves by 1.79 points, indicating that token-level PPO provides a modest gain with uneven effects across splits. Furthermore, We isolate failure replay by comparing ID 2 with ID 1. Adding failure replay increases Seen and Unseen Object SR by 8.04 and 3.34 percentage points, respectively, whereas Unseen Map SR decreases by 11.69 points. Consequently, average SR decreases slightly by 0.10 points, showing that failure replay alone redistributes performance across splits rather than providing a uniform gain. These results suggest that revisiting unresolved tasks and regenerating trajectories under the current policy can transform corrective signals into additional learning experience. However, since replay is conditioned on the sampled scene, it amplifies failure-focused updates within the same empirically weighted distribution instead of correcting its long tail. Effect of Reference-Policy KL Regularization. To verify the effect of reference-policy KL regularization, we compare ID 2 and ID 3. As shown in Table 4, the KL-regularized policy improves Unseen Map SR by 6.26 percentage points and raises the average by 2.53 points, whereas Seen SR decreases by 1.2 points. This is consistent with the intended role of constraining excessive deviations from the stage-initial policy. Consequently, KL regularization may help retain transferable behaviors from the initial VLA policy on unseen maps while achieving superior results across seen and unseen object splits, which may alleviate excessive policy drift introduced by RL updates and failure replay, while remains the capabilities of the initial VLA model. Effect of the Two-Stage Curriculum and Sampling Strategy Analysis. To evaluate the two-stage curriculum, we compare IDs 3 and 4. As reported in Table 4, the two-stage curriculum improves Unseen Map SR by 7.42 percentage points and average SR by 2.26 points, with only a 1.91-point decrease in Unseen Object SR. Nevertheless, this comparison adds a training stage. Accordingly, Table 5 introduces two controls to isolate the effect of the sampling strategy. As shown in Table 5, RecoverFly outperforms both controls across all three test splits. Importantly, Original Distribution and RecoverFly involve the same number of training stages, indicating that the improvement cannot be explained solely by additional optimization steps. Consequently, the result supports the complementary roles of the stages of the long-tail scene curriculum. Specifically, stage I establishes the policy under the original scene distribution, while the second stage increases the proportion of rare scenes to achieve balanced performance gains without significantly sacrificing previously acquired capabilities. Conclusions This paper presents RecoverFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies. RecoverFly combines token-level PPO, dynamic failure replay, a two-stage long-tail scene curriculum, and reference-policy regularization. Experiments on the TravelUAV benchmark demonstrate RecoverFly outperforms all comparison methods across all three splits. With a total rollout budget of about 30% of the training-set size, RecoverFly improves the SR of the initial VLA policy by 3.12 to 8.37 percentage points. Moreover, ablation studies demonstrate that revisiting unresolved tasks converts sparse failure feedback into reusable corrective experience, while shifting training from the empirical distribution to a distribution that is more balanced towards rare scenes outperforms uniform sampling. Furthermore, reference-policy regularization balances performance across different stages and reduces capability degradation during optimization. We hope this work provides an effective RL solution for advancing UAV-VLN. References P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. van den Hengel (2018) Vision-and-language navigation: interpreting visually-grounded navigation instructions in real environments. In Proc. CVPR, p. 3674–3683. Cited by: Introduction, VLN for UAVs, 2nd item. M. Andrychowicz, D. Crow, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba (2017) Hindsight experience replay. In Proc. NeurIPS, p. 5048–5058. Cited by: Dynamic Failure Replay. Y. Bengio, J. Louradour, R. Collobert, and J. Weston (2009) Curriculum learning. In Proc. ICML, p. 41–48. Cited by: Two-Stage Long-Tail Scene Curriculum. V. Blukis, Y. Terme, E. Niklasson, R. A. Knepper, and Y. Artzi (2019) Learning to map natural language instructions to physical quadcopter control using simulated flight. In Proc. CoRL, Vol. 100, p. 1415–1438. Cited by: RL in UAV-VLN. H. Cai, J. Dong, J. Tan, J. Deng, S. Li, Z. Gao, H. Wang, Z. Su, A. Sumalee, and R. Zhong (2025) Flightgpt: Towards generalizable and interpretable UAV vision-and-language navigation with vision-language models. In Proc. EMNLP, p. 6670–6687. Cited by: RL in UAV-VLN. G. Chen, X. Yu, N. Ling, and L. Zhong (2025) TypeFly: Low-latency drone planning with large language models. IEEE Trans. Mob. Comput., p. 9068–9079. Cited by: VLN for UAVs. X. Ding, J. Gao, C. Pan, W. Wang, and J. Qin (2026) History-enhanced two-stage transformer for aerial vision-and-language navigation. In Proc. AAAI, Vol. 40, p. 18225–18233. Cited by: Introduction. C. Fan, C. Pan, Z. Liu, N. Liu, and J. Qin (2026) HTNav: A hybrid navigation framework with tiered structure for urban aerial vision-and-language navigation. In Proc. CVPR, p. 10976–10985. Cited by: RL in UAV-VLN. Y. Fan, W. Chen, T. Jiang, C. Zhou, Y. Zhang, and X. Wang (2023) Aerial vision-and-dialog navigation. In Proc. ACL, p. 3043–3061. Cited by: Introduction, VLN for UAVs. Y. Gao, C. Li, Z. You, J. Liu, Z. Li, P. Chen, Q. Chen, Z. Tang, L. Wang, P. Yang, et al. (2026) OpenFly: A comprehensive platform for aerial vision-language navigation. In Proc. ICLR, Cited by: VLN for UAVs. C. Y. Hu, Y. Lin, Y. Lee, C. Su, J. Lee, S. Tsai, C. Lin, K. Chen, T. Ke, and Y. Liu (2025) See, point, fly: A learning-free VLM framework for universal unmanned aerial navigation. In Proc. CoRL, p. 4697–4708. Cited by: VLN for UAVs. E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) LoRA: Low-rank adaptation of large language models. In Proc. ICLR, Cited by: Implementation Details.. M. Jiang, E. Grefenstette, and T. Rocktäschel (2021) Prioritized level replay. In Proc. ICML, Vol. 139, p. 4940–4950. Cited by: Dynamic Failure Replay. W. Jiang, L. Wang, K. Huang, W. Fan, J. Liu, S. Liu, H. Duan, B. Xu, and X. Ji (2025) LongFly: Long-horizon UAV vision-and-language navigation with spatiotemporal context integration. External Links: 2512.22010 Cited by: Introduction, Introduction, VLN for UAVs, 3rd item. M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. (2025) OpenVLA: An open-source vision-language-action model. In Proc. CoRL, p. 2679–2713. Cited by: Introduction, VLA Policies for UAV-VLN, Implementation Details.. J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee (2020) Beyond the nav-graph: vision-and-language navigation in continuous environments. In Proc. ECCV, A. Vedaldi, H. Bischof, T. Brox, and J. Frahm (Eds.), p. 104–120. Cited by: Introduction, VLN for UAVs. T. Li, T. Huai, Z. Li, Y. Gao, H. Li, and X. Zheng (2025) SkyVLN: Vision-and-language navigation and NMPC control for UAVs in urban environments. In Proc. IROS, p. 17199–17206. Cited by: VLN for UAVs. P. Lin, G. Sun, C. Liu, F. Li, W. Ren, and Y. Cong (2025) OpenVLN: Open-world aerial vision-language navigation. External Links: 2511.06182 Cited by: Introduction, Introduction, RL in UAV-VLN. S. Liu, H. Zhang, Y. Qi, P. Wang, Y. Zhang, and Q. Wu (2023) AerialVLN: Vision-and-language navigation for UAVs. In Proc. ICCV, p. 15338–15348. Cited by: Introduction, VLN for UAVs. A. Lykov, V. Serpiva, M. H. Khan, O. Sautenkov, A. Myshlyaev, G. Tadevosyan, Y. Yaqoot, and D. Tsetserukou (2025) Cognitivedrone: A VLA model and evaluation benchmark for real-time cognitive task solving and reasoning in UAVs. External Links: 2503.01378 Cited by: VLA Policies for UAV-VLN. Y. Ning, G. Zhao, Y. Qin, S. Liu, Y. Liu, L. Lin, and G. Li (2026) LookasideVLN: direction-aware aerial vision-and-language navigation. In Proc. CVPR, p. 32441–32450. Cited by: Introduction. L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe (2022) Training language models to follow instructions with human feedback. In Proc. NeurIPS, p. 27730–27744. Cited by: Introduction. S. Ross, G. J. Gordon, and D. Bagnell (2011) A reduction of imitation learning and structured prediction to no-regret online learning. In Proc. AISTATS, p. 627–635. Cited by: Introduction, VLA Policies for UAV-VLN. J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz (2015) Trust region policy optimization. In Proc. ICML, Vol. 37, p. 1889–1897. Cited by: Reference-Policy KL Regularization. J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel (2016) High-dimensional continuous control using generalized advantage estimation. In Proc. ICLR, Cited by: Event-Aware Reward and Advantage Estimation.. J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017) Proximal policy optimization algorithms. External Links: 1707.06347 Cited by: RL in UAV-VLN. V. Serpiva, A. Lykov, A. Myshlyaev, M. H. Khan, A. A. Abdulkarim, O. Sautenkov, and D. Tsetserukou (2025) RaceVLA: VLA-based racing drone navigation with human-like behaviour. External Links: 2503.02572 Cited by: VLA Policies for UAV-VLN. S. Shah, D. Dey, C. Lovett, and A. Kapoor (2017) Airsim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and service robotics: Results of the 11th international conference, p. 621–635. Cited by: Implementation Details.. X. Wang, D. Yang, H. Kwan, J. Chen, H. Li, Y. Liao, S. Liu, et al. (2025) Towards realistic UAV vision-language navigation: Platform, benchmark, and methodology. In Proc. ICLR, p. 7292–7310. Cited by: Introduction, Introduction, VLN for UAVs, 2nd item, Dataset.. X. Wang, D. Yang, Y. Liao, W. Zheng, B. Dai, H. Li, S. Liu, et al. (2026) UAV-flow colosseo: A real-world benchmark for flying-on-a-word UAV imitation learning. In Proc. NeurIPS, Vol. 38. Cited by: VLA Policies for UAV-VLN. J. Xiao, Y. Sun, Y. Shao, B. Gan, R. Liu, Y. Wu, W. Guan, and X. Deng (2025) UAV-ON: A benchmark for open-world object goal navigation with aerial agents. In Proc. ACM M, p. 13023–13029. Cited by: VLN for UAVs. P. Xu, Z. Deng, J. Deng, Z. Gu, and S. Wan (2026) AerialVLA: A vision-language-action model for UAV navigation via minimalist end-to-end control. External Links: 2603.14363 Cited by: Introduction, Introduction, Introduction, VLA Policies for UAV-VLN, Textual Action Representation., 4th item. C. Yu, Y. Wang, Z. Guo, H. Lin, S. Xu, H. Zang, Q. Zhang, Y. Wu, C. Zhu, J. Hu, et al. (2025) RLinf: Flexible and efficient large-scale reinforcement learning via macro-to-micro flow transformation. External Links: 2509.15965 Cited by: Overview, Implementation Details.. J. Zhang, A. Li, Y. Qi, M. Li, J. Liu, S. Wang, H. Liu, G. Zhou, Y. Wu, X. Li, Y. Fan, W. Li, Z. Chen, F. Gao, Q. Wu, Z. Zhang, and H. Wang (2025a) Embodied navigation foundation model. External Links: 2509.12129 Cited by: Introduction, VLN for UAVs, 3rd item. W. Zhang, C. Gao, S. Yu, R. Peng, B. Zhao, Q. Zhang, J. Cui, X. Chen, and Y. Li (2025b) Citynavagent: Aerial vision-and-language navigation with hierarchical semantic planning and global memory. In Proc. ACL, p. 31292–31309. Cited by: VLN for UAVs. B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V. Vanhoucke, H. T. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. Sanketi, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y. Lu, S. Levine, L. Lee, T. E. Lee, I. Leal, Y. Kuang, D. Kalashnikov, R. Julian, N. J. Joshi, A. Irpan, B. Ichter, J. Hsu, A. Herzog, K. Hausman, K. Gopalakrishnan, C. Fu, P. Florence, C. Finn, K. A. Dubey, D. Driess, T. Ding, K. M. Choromanski, X. Chen, Y. Chebotar, J. Carbajal, N. Brown, A. Brohan, M. G. Arenas, and K. Han (2023) RT-2: vision-language-action models transfer web knowledge to robotic control. In Proc. CoRL, J. Tan, M. Toussaint, and K. Darvish (Eds.), p. 2165–2183. Cited by: Introduction, VLA Policies for UAV-VLN.