Paper deep dive
HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL
Langzhe Gu, Chengkai Hou, Meng Li, Xinhua Wang, Jiaming Liu, Xinyuan Lv, Bowei Zhang, Shuanghao Bai, Guangrun Li, Jingyang He, Gaole Dai, Ziluo Ding, Zhiyuan Xu, Kuan Cheng, Jian Tang, Zhengping Che, Shanghang Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/23/2026, 2:58:13 AM
Summary
The paper introduces HAF (Humanoid Adaptation Framework), a two-part system designed to adapt generalist Vision-Language-Action (VLA) models for humanoid whole-body loco-manipulation. HAF consists of HAF-VLA, a hierarchical action-flow generator that splits action denoising into three sequential stages (locomotion/head, waist, manipulation) using cross-stage KV caches to maintain kinematic coherence, and HAF-Steer, a latent offline-to-online reinforcement learning pipeline. HAF-Steer uses flow-matching invertibility and DCT-based dimensionality reduction to optimize a regularized SAC policy in a compact noise subspace, avoiding expensive updates to the large VLA backbone. The framework is evaluated on seven real-world humanoid tasks, demonstrating superior coordination and performance compared to single-stage VLA baselines.
Entities (8)
Relation Signals (8)
HAF → consistsof → HAF-Steer
confidence 98% · HAF consists of two complementary components: HAF-VLA... and HAF-Steer
HAF → consistsof → HAF-VLA
confidence 98% · HAF consists of two complementary components: HAF-VLA... and HAF-Steer
HAF → targetsproblem → Humanoid Whole-Body Loco-manipulation
confidence 95% · Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation
HAF-VLA → usestechnique → Flow Matching
confidence 95% · HAF-VLA is a hierarchical action-flow generator built on a pretrained flow-matching VLA.
HAF-Steer → usesalgorithm → SAC
confidence 92% · train a regularized SAC policy.
HAF-VLA → usesmechanism → Cross-stage KV Cache
confidence 92% · It splits full-body action denoising into three sequential stages with stage embeddings and cross-stage KV caches
HAF-Steer → usestechnique → DCT
confidence 92% · HAF-Steer is a latent offline-to-online RL pipeline that leverages flow-matching invertibility and DCT-based dimensionality reduction
HAF → improves → Whole-body coordination
confidence 90% · HAF surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively. Moreover, policies trained through offline behavior cloning can remain suboptimal during real-world deployment. Although online reinforcement learning can refine policies through real-world interaction, directly tuning large VLA backbones demands excessive computation and may introduce safety risks during real-robot exploration. To address these bottlenecks, we introduce HAF (Humanoid Adaptation Framework), a two-part framework consisting of HAF-VLA and HAF-Steer that transfers off-the-shelf generalist VLA foundation models to humanoid whole-body loco-manipulation. HAF-VLA is a hierarchical action-flow generator built on a pretrained flow-matching VLA. It splits full-body action denoising into three sequential stages with stage embeddings and cross-stage KV caches that retain kinematic dependencies, avoiding incoherent whole-body actions from one-shot generation. On top of the frozen HAF-VLA, HAF-Steer is a latent offline-to-online RL pipeline that leverages flow-matching invertibility and DCT-based dimensionality reduction to restrict RL optimization to a compact noise subspace and train a regularized SAC policy. This avoids updating the large VLA backbone and enables efficient real-world policy refinement. Evaluated on seven real-world humanoid loco-manipulation tasks, HAF surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance. Project website: this https URL .
Tags
Links
- Source: https://arxiv.org/abs/2608.16837v1
- Canonical: https://arxiv.org/abs/2608.16837v1
Trouble viewing inline? Open PDF directly →
Full Text
57,061 characters extracted from source content.
Expand or collapse full text
HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL Langzhe Gu Chengkai Hou Meng Li Xinhua Wang Jiaming Liu Xinyuan Lv Bowei Zhang Shuanghao Bai Affiliation: Xi’an Jiaotong University Guangrun Li Jingyang He Gaole Dai Ziluo Ding Zhiyuan Xu Kuan Cheng Jian Tang Zhengping Che Shanghang Zhang [0.5em] State Key Laboratory of Multimedia Information Processing School of Computer Science Peking University Beijing Innovation Center of Humanoid Robotics Nankai University [0.3em] Equal contribution Abstract Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively. Moreover, policies trained through offline behavior cloning can remain suboptimal during real-world deployment. Although online reinforcement learning can refine policies through real-world interaction, directly tuning large VLA backbones demands excessive computation and may introduce safety risks during real-robot exploration. To address these bottlenecks, we introduce HAF (Humanoid Adaptation Framework), a two-part framework consisting of HAF-VLA and HAF-Steer that transfers off-the-shelf generalist VLA foundation models to humanoid whole-body loco-manipulation. HAF-VLA is a hierarchical action-flow generator built on a pretrained flow-matching VLA. It splits full-body action denoising into three sequential stages with stage embeddings and cross-stage KV caches that retain kinematic dependencies, avoiding incoherent whole-body actions from one-shot generation. On top of the frozen HAF-VLA, HAF-Steer is a latent offline-to-online RL pipeline that leverages flow-matching invertibility and DCT-based dimensionality reduction to restrict RL optimization to a compact noise subspace and train a regularized SAC policy. This avoids updating the large VLA backbone and enables efficient real-world policy refinement. Evaluated on seven real-world humanoid loco-manipulation tasks, HAF surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance. Project website: https://grange007.github.io/HAF. Keywords: Humanoid Robot, Vision-Language-Action (VLA) models, Reinforcement Learning, Imitation Learning 1 Introduction Humanoid robots hold the promise of becoming universal agents capable of operating in diverse human-centric environments. Unlike fixed-base or wheeled manipulators, humanoids require the synchronized coordination of locomotion, torso posture, and bimanual manipulation. However, this complexity introduces a critical challenge: unstable base movements often induce erratic upper-body compensatory motions, which severely degrade balance and manipulation accuracy [4, 42, 19, 32, 11, 21, 40, 12]. Figure 1: Overview of HAF. We demonstrate seven humanoid loco-manipulation tasks in real-world settings. HAF consists of two complementary components: HAF-VLA is a hierarchical action-flow vision-language-action model that generates whole-body motions stage by stage according to the kinematic dependencies among locomotion, torso adjustment, and dual-arm manipulation; HAF-Steer performs latent-space reinforcement learning on discrete cosine transform compressed temporal noise from frozen flow-based VLA policies. Together, they improve task performance in real-world settings. Although recent generalist vision-language-action (VLA) models have shown strong generalization on conventional robotic platforms [5, 29, 27, 14], adapting them to humanoid whole-body loco-manipulation remains challenging. Standard flow-matching VLAs typically generate all body-part actions in a single stage, without explicitly modeling the dependencies among locomotion, posture, and manipulation [21, 11, 10, 44]. Existing humanoid VLAs often address this issue through humanoid-specific pretraining or large embodiment-specific datasets [12, 27, 37], which substantially increases the cost of adapting a generalist VLA to a new humanoid platform. Moreover, offline behavior-cloned policies often degrade under deployment-time distribution shifts, yet directly fine-tuning large VLA backbones for high-dimensional humanoid action spaces with online reinforcement learning is computationally expensive and may induce unsafe exploration. Existing latent-noise RL methods avoid backbone updates but either optimize the full high-dimensional temporal noise or repeat a single noise vector over the entire horizon, sacrificing either efficiency or temporal expressiveness [35, 34, 20, 23]. To address these challenges, we introduce HAF, the Humanoid Adaptation Framework for transferring pretrained flow-matching VLAs to humanoid whole-body loco-manipulation. HAF-VLA restructures whole-body action generation according to the dependency among locomotion, body posture, and manipulation, while HAF-Steer performs lightweight policy refinement in a compact latent noise space. The two components operate at complementary levels: HAF-VLA provides a structured whole-body action generator, while HAF-Steer further refines the frozen generator through real-world offline-to-online adaptation. HAF-VLA repurposes a pretrained generalist VLA [29] by orchestrating a progressive, kinematics-aware generation process. Instead of producing full-body actions in a single shot, HAF-VLA sequentially outputs commands for locomotion and head orientation, followed by torso adjustment, and finally bimanual manipulation. This hierarchical ordering prioritizes base stabilization to mitigate the erratic compensatory motions often observed in upper-body control [6, 13]. To ensure global coherence, we employ cross-stage KV-cache conditioning: clean actions from earlier stages are re-encoded into cross-stage KV caches to condition subsequent generation. Since the active action sets are cumulative, later stages can refine previously generated dimensions, and only the final full-body action chunk is executed. Built upon the frozen VLA model, HAF-Steer introduces a low-dimensional spectral action space for offline-to-online reinforcement learning (RL). We numerically reverse the frozen flow matching field to recover the initial noise corresponding to demonstrated action chunks and apply the discrete cosine transform (DCT) along the temporal dimension, retaining only the first 8 coefficients. This design avoids the high computational cost of optimizing in the raw high-dimensional action space or fine-tuning the massive VLA backbone, while preserving smoothly varying temporal modes. An RL actor for spectral action is trained through behavior cloning (BC) with offline expert spectral samples first and subsequently optimized using mixed offline–online Soft Actor-Critic (SAC) [9] with BC regularization. We validate HAF on two physical humanoid robots across seven complex household loco-manipulation tasks. HAF-VLA significantly outperforms imitation learning and standard VLA baselines, particularly in tasks requiring long-distance travel and coordinated full-body motion. With the HAF-Steer module, our framework achieves robust performance gains in both in-distribution and challenging OOD scenarios. Our contributions are summarized as follows: • Humanoid Adaptation Framework: We propose HAF, an effective framework that repurposes pretrained generalist VLAs for humanoid whole-body loco-manipulation, eliminating the need to train humanoid-specific foundation models from scratch. • Kinematics-Aligned Generation: We design HAF-VLA, a hierarchical action-flow module that aligns with humanoid kinematics to suppress unstable compensatory movements and improve action coherence in whole-body loco-manipulation. • Efficient Real-World Refinement: We develop HAF-Steer, a DCT-based latent RL pipeline that enables stable mixed offline–online adaptation for frozen VLA backbones. • Comprehensive Real-World Validation: We validate HAF on two humanoid platforms across seven real-world loco-manipulation tasks, demonstrating improved task performance, whole-body coordination, and robustness under distribution shifts. 2 Related Work Humanoid VLAs. Vision-language-action models (VLAs) aim to map visual observations and language instructions directly to robot actions [22, 15, 46, 36]. Representative works such as RT-2 [47] and OpenVLA [15] show that vision-language pretraining can be effectively transferred to robotic manipulation, while π0 _0 and π0.5 _0.5 further move toward continuous generative action modeling with flow matching [5, 29]. However, most existing VLAs still focus on tabletop tasks or mobile manipulators, leaving high-dimensional humanoid whole-body action generation underexplored. Recent works have begun to extend VLAs to humanoid robots [31, 38, 7, 2]. For instance, GR00T N1.7 [27] studies general-purpose humanoid action generation, WholeBodyVLA [12] explores latent VLA learning for large-space mobile manipulation, and Ψ0 _0 [37] introduces an open foundation model pretrained on egocentric human videos. In contrast to these methods that rely on latent skills, cross-embodiment transfer, or learning from human videos, our method is directly built upon a pretrained generalist VLA. We improve its performance on humanoid robots by generating structured whole-body actions grounded in the forward-kinematic chain, sparing the computational cost of training a new model. RL post-training for VLA models. Pre-trained VLA policies obtained solely through offline behavior cloning can remain suboptimal during real-world deployment, motivating reinforcement learning (RL) for post-training refinement [3, 43, 16, 24, 39, 20]. Relevant approaches can be broadly grouped into three categories. The first directly fine-tunes the entire VLA backbone using algorithms like PPO [16, 26, 30], but this incurs massive computational overhead and risks generating unstable, unsafe motions that erase pre-trained priors in high-dimensional whole-body tasks. The second freezes the VLA backbone and trains lightweight residual networks to revise raw actions (e.g., ResFit [1], EXPO [8], DICE-RL [33]). Despite variance reduction techniques, these methods still optimize in the original action space, leading to low sample efficiency and accumulated temporal errors in long-horizon tasks. The third branch performs RL within the latent noise space, which is most relevant to our work. However, existing methods also have some limitations: DSRL [35] and FRS-based noise policies [34] reduce the search dimension by repeating a single noise vector across the entire action chunk, but this construction departs largely from the independently sampled Gaussian temporal noise used during VLA pretraining. UniSteer [23] directly optimizes full-dimensional temporal noise over the entire action chunk, maintaining a large exploration space. Compared with prior latent RL work, HAF-Steer significantly reduces the search dimension of temporal noise, suppresses unstable high-frequency exploratory movements during real-world interaction, and enables more efficient real-world adaptation on long-horizon humanoid loco-manipulation tasks. Whole-body manipulation. Recent research on humanoid whole-body control lays an essential technical foundation for integrated locomotion and manipulation tasks. Advanced low-level controllers and RL-based motion pipelines enable legged robots to perform agile, dynamic movements [25, 18, 21, 28]. Language-conditioned whole-body controllers such as LangWBC [31] and LeVERB [38] support high-level language-guided navigation but lack fine-grained bimanual dexterous manipulation capabilities. VR teleoperation frameworks including TWIST2 [41], AMO [17] and SONIC [25] deliver effective pipelines to collect full-body humanoid trajectories and optimize low-level motion tracking, yet they merely serve as data acquisition tools rather than generalizable vision-language-driven policies for long-horizon loco-manipulation. None of these whole-body control architectures are integrated with hierarchical generative VLA backbones. 3 HAF-VLA: Hierarchical Action Flow for Humanoid Whole-Body Loco-Manipulation HAF-VLA is a hierarchical action-flow module developed to repurpose pretrained generalist flow-matching VLA foundation models for humanoid whole-body loco-manipulation. Unlike single-stage VLA generators that entangle all body movements, it reflects humanoid kinematic dependencies by producing motions step by step: locomotion and head signals are predicted first, followed by waist adjustment, and finally fine bimanual manipulation trajectories. To link motion outputs across generation phases, we introduce a cross-stage KV-cache conditioning mechanism that feeds encoded action features from earlier stages as context to subsequent generation branches, unifying shared visual-linguistic representations and reducing unstable upper-body compensatory movements triggered by flawed base or torso poses. Figure 2: Overview of HAF-VLA. A shared action expert progressively expands the active action space from locomotion and head control to waist posture and bimanual manipulation. Clean action caches from earlier stages condition subsequent generation, and only the final full-body action chunk is executed. 3.1 Hierarchical Whole-Body Action Generation Conditioned on the current observation ot=(It,qt)o_t=(I_t,q_t), which consists of the egocentric RGB image ItI_t and robot proprioception qtq_t, as well as the language instruction ℓ , the VLA model predicts a structured whole-body action chunk AtA_t. Formally, the action chunk is defined as a sequence of future control steps: At=[at,…,at+H−1]∈ℝH×DA_t=[a_t,…,a_t+H-1] ^H× D, where H denotes the predicted action horizon length and D denotes the dimension of per-step whole-body action. In our implementation, H=100H=100 and the robot executes the first 40 steps before the next inference call. To support hierarchical action generation for whole-body loco-manipulation, each instantaneous action at∈ℝDa_t ^D is explicitly decomposed into four orthogonal kinematic sub-components: at=[atmove,athead,atwaist,atmanip]a_t=[a_t^move,a_t^head,a_t^waist,a_t^manip], which correspond to locomotion and skill-mode commands, head orientation control, waist posture adjustment, and bimanual manipulation targets. Based on this explicit action decomposition, we further partition the full D-dimensional action index space into four mutually disjoint subsets corresponding to the above kinematic components: tmoveA_t^move, theadA_t^head, twaistA_t^waist, and tmanipA_t^manip. Instead of predicting all action dimensions simultaneously in a single forward pass, HAF-VLA achieves progressive, coarse-to-fine whole-body generation via three nested cumulative action index sets constructed from the four subsets: t1=tmove∪thead,t2=t1∪twaist,t3=t2∪tmanip.A_t^1=A_t^move _t^head, _t^2=A_t^1 _t^waist, _t^3=A_t^2 _t^manip. (1) This construction strictly establishes the nested inclusion relation t1⊂t2⊂t3A_t^1 _t^2 _t^3, supporting hierarchical action generation across three progressive stages. In Stage 1, the model generates fundamental locomotion behaviors, task skill modes, and gaze directions over the minimal action set t1A_t^1. Stage 2 expands the active action space to t2A_t^2, incorporating waist actuation signals to refine torso posture and reshape the upper-body workspace. Stage 3 fully activates the complete action space t3A_t^3 to output the final executable whole-body action commands. Different from rigid decoupled generation schemes that fix early-stage outputs, our cumulative nested design allows each subsequent stage to refine and update action dimensions predicted in earlier stages, enabling more accurate and coherent whole-body loco-manipulation. All three hierarchical stages share identical weights for the action-flow expert network. To reduce redundant computation, the vision-language prefix KV cache is computed only once at the start and reused across all stages of hierarchical inference: Pt=Fprefix(ot,ℓ)P_t=F_prefix(o_t, ), where PtP_t is the shared vision-language prefix cache derived from the current observation oto_t and language instruction ℓ . Each stage starts sampling from an independent Gaussian noise vector ϵts∼(,) _t^s (0,I), and stage-specific behavior is modulated via a trainable stage embedding ese_s for stage index s∈1,2,3s∈\1,2,3\. The full multi-stage generation process is formalized below: At1 A_t^1 =Fθ(ϵt1∣Pt,e1), =F_θ( _t^1 P_t,e_1), Ct1 C_t^1 =Cacheθ(At1∣Pt,e1), =Cache_θ(A_t^1 P_t,e_1), (2) At2 A_t^2 =Fθ(ϵt2∣Pt,Ct1,e2), =F_θ( _t^2 P_t,C_t^1,e_2), Ct2 C_t^2 =Cacheθ(At2∣Pt,Ct1,e2), =Cache_θ(A_t^2 P_t,C_t^1,e_2), At3 A_t^3 =Fθ(ϵt3∣Pt,Ct1,Ct2,e3). =F_θ( _t^3 P_t,C_t^1,C_t^2,e_3). Here FθF_θ denotes the shared numerical action-flow map, with the stage-specific action mask applied internally according to the corresponding stage. The Cacheθ(⋅)Cache_θ(·) operator re-encodes the denoised action prediction from the current stage and compresses it into an action KV cache (Ct1C_t^1 for Stage 1, Ct2C_t^2 for Stage 2). These cached motion embeddings inject coarse prior motion context into later generation steps. Critically, only the final output chunk At≡At3A_t≡ A_t^3 is sent to the robot controller for real execution; intermediate predictions At1A_t^1 and At2A_t^2 serve purely as context providers and are never deployed directly. 3.2 Training Strategy and Deployment Let ms∈0,1Dm_s∈\0,1\^D be the binary indicator of the active action set tsA_t^s, broadcast along the temporal dimension. For compactness, we denote the conditioning of the three stages as ht1=(Pt,e1)h_t^1=(P_t,e_1), ht2=(Pt,Ct1,e2)h_t^2=(P_t,C_t^1,e_2), and ht3=(Pt,Ct1,Ct2,e3)h_t^3=(P_t,C_t^1,C_t^2,e_3). Given a clean action chunk At∈ℝH×DA_t ^H× D, each Stage s independently samples Gaussian noise ϵts∼(0,I) _t^s (0,I) and flow time τs∼(0,1)τ^s (0,1). The noisy input and target velocity for Stage s are Xτs,ts X_τ^s,t^s =[(1−τs)ϵts+τsAt]⊙ms, = [(1-τ^s) _t^s+τ^sA_t ] m_s, (3) uts u_t^s =(At−ϵts)⊙ms, = (A_t- _t^s ) m_s, where Xτs,tsX_τ^s,t^s is the masked noisy action input and utsu_t^s is the corresponding masked flow-matching target. All stages operate in the same global action space. The stage mask is applied to both the input and target, such that inactive dimensions are assigned zero target velocity rather than excluded from optimization. The shared action expert is trained with the stage-wise flow-matching objective ℒHAF=At,ϵts,τss=13[∑s=131HD‖vθ(Xτs,ts,τs,hts)−uts‖F2],L_HAF=E_A_t,\ _t^s,τ^s\_s=1^3 [ _s=1^3 1HD \|v_θ (X_τ^s,t^s,τ^s;h_t^s )-u_t^s \|_F^2 ], (4) where vθv_θ denotes the velocity field predicted by the shared action expert. The three stage losses are summed with equal weights. Apart from their independently sampled noise and flow time, the stages differ only in their conditioning htsh_t^s and active action mask msm_s, while sharing the same model parameters. During training, teacher forcing is used to compute cross-stage caches Ct1C_t^1 and Ct2C_t^2 from re-encoded masked ground-truth actions rather than stage outputs, which prevents error accumulation while retaining the hierarchical inference structure. At inference, the three stages run sequentially with independent Gaussian noise samples. Ct1C_t^1 and Ct2C_t^2 are built from the denoised outputs of earlier stages to condition later ones, and only Stage 3’s result is executed via receding-horizon control. Each stage starts from independently sampled noise and performs 10 flow-denoising steps conditioned on the visual-language cache and the available cross-stage action caches. We deploy HAF-VLA with a receding-horizon execution scheme. The robot executes the first 40 steps before the next inference call. In our real-robot deployment, the complete three-stage inference takes approximately 0.12 seconds on a single RTX 5090 GPU, supporting real-time closed-loop humanoid loco-manipulation. 4 HAF-Steer: Spectral Latent RL for Flow-Matching VLAs HAF-Steer is a lightweight reinforcement-learning method for adapting a frozen flow-matching VLA through its initial flow noise, as illustrated in fig. 3. Demonstrated actions are first mapped back to their corresponding noise and compressed into low-dimensional spectral actions. A stochastic policy is then trained in this spectral space and decoded through the frozen flow generator to produce executable action chunks. We first define the spectral latent space and then introduce its offline-to-online optimization. Figure 3: Overview of HAF-Steer. (a) A conventional flow-matching VLA samples full-dimensional Gaussian noise and decodes it into an action trajectory through a frozen flow generator. (b) HAF-Steer instead constructs a low-dimensional spectral action space from expert demonstrations: demonstrated trajectories are mapped back to their initial flow noise through numerical reverse integration, compressed by retaining the first 8 temporal DCT coefficients, and normalized to form expert spectral actions in the offline replay buffer. An RL policy is trained with mixed offline and online experience to predict spectral actions, which are de-normalized, reconstructed into full temporal noise by inverse DCT, and decoded by the frozen VLA to produce executable robot trajectories. 4.1 Spectral Latent Space via Flow Reversal HAF-Steer builds upon a frozen pre-trained flow-matching VLA backbone FθF_θ, which generates whole-body action chunks via a deterministic flow mapping: At=Fθ(ϵt∣ζt),ϵt∈ℝH×D,A_t=F_θ ( _t _t ), _t ^H× D, where H is the action chunk length, D is the dimension of robot action, At∈ℝH×DA_t ^H× D denotes the predicted action chunk, ϵt _t is the initial flow noise, and ζt _t represents the unified vision-language-robot conditioning of the frozen VLA model. The network parameters θ are fully fixed during subsequent reinforcement learning adaptation. Directly optimizing raw temporal noise at all H time steps leads to an excessively high-dimensional exploration space. Prior steering methods [35, 34] repeat a single noise vector over the entire chunk, which we found unstable and frequently inducing severe body jerking or unsafe transient behaviors in whole-body RL training. To enforce temporal smoothness and compress action space complexity, we constrain RL exploration within a low-dimensional spectral subspace spanned by the first K discrete cosine transform (DCT) bases. In our practice, we empirically choose K=8K=8, which greatly reduces the exploration space from H×DH× D to K×DK× D. Our multi-mode spectral parameterization retains diverse smooth temporal variation patterns and yields physically plausible whole-body action corrections. Based on the above spectral dimensionality reduction design, we further build offline supervised spectral targets to guide policy training. Specifically, to construct supervised spectral targets from offline demonstrations, we first invert each expert action chunk back to its corresponding flow noise space using the frozen VLA model. Given a demonstrated action Ai∗A_i^* and its conditioning ζi∗ _i^*, we recover and normalize its spectral representation via DCT transformation: ϵi∗=Fθ−1(Ai∗∣ζi∗),ci∗=DCTK(ϵi∗),zi∗=ci∗−μcσc+δ. _i^*=F_θ^-1 (A_i^* _i^* ), c_i^*=DCT_K( _i^*), z_i^*= c_i^*- _c _c+δ. (5) Here, Fθ−1F_θ^-1 denotes numerical backward integration of the fixed flow field rather than a learned inverse model. ϵi∗ _i^* is the recovered full-scale flow noise, ci∗∈ℝK×Dc_i^* ^K× D truncates the first K DCT temporal coefficients, and zi∗z_i^* is the normalized expert spectral action. The statistics μc,σc _c, _c are precomputed over the entire demonstration set, and δ is a small constant for numerical stability. According to the derived spectral expert representations, we construct an offline spectral replay buffer for policy training. We formulate VLA-level transition tuples consisting of spectral expert targets, sparse task rewards, and episodic termination flags: off=(xi,zi∗,ri,xi+1,di)i=1N−1,D_off= \ (x_i,z_i^*,r_i,x_i+1,d_i ) \_i=1^N-1, (6) Here, xix_i and xi+1x_i+1 denote consecutive observational states, zi∗z_i^* is the precomputed spectral expert target, and did_i indicates episode termination. We adopt a sparse reward scheme for demonstration-guided training: terminal transitions of successful trajectories are assigned ri=1r_i=1, while all intermediate transitions receive ri=0r_i=0. This simple yet effective reward formulation drives the policy to learn task-aligned spectral noise corrections from offline data. 4.2 Offline-to-Online Spectral Policy Learning Given the spectral action space and offline replay buffer constructed above, we first train the actor through behavior cloning and then improve it using mixed offline–online reinforcement learning. After training, the learned spectral actor generates smooth whole-body corrections at inference time through DCT spectral decoding. Behavior-cloning initialization. Let μψ(x) _ψ(x) denote the mean output of the stochastic spectral actor πψ(z∣x) _ψ(z x). We first initialize the actor by regressing toward the recovered expert spectral actions in offD_off: ℒBC=(xi,zi∗)∼off[‖μψ(xi)−zi∗‖22].L_BC=E_(x_i,z_i^*) _off [ \| _ψ(x_i)-z_i^* \|_2^2 ]. (7) Only the spectral actor is optimized during this initialization stage; the VLA backbone and value functions remain unchanged. Mixed Offline–Online Reinforcement Learning. After BC initialization, the policy interacts with the real robot following the spectral generation pipeline in Eq. (9), and newly collected transitions are stored in an online replay buffer onD_on. We adopt the same sparse terminal reward rule used for offline data: only successful terminal steps receive r=1r=1, while all other transitions yield r=0r=0. In each training iteration, we sample mixed mini-batches ℬ=ℬoff∪ℬonB=B_off _on from both buffers. All samples participate in standard SAC policy updates, while BC regularization is only applied to offline expert data to preserve demonstration quality: ℒπ= _π= xi∼ℬ,zi∼πψ(⋅∣xi)[αlogπψ(zi∣xi)−minjQωj(xi,zi)] _x_i ,\,z_i _ψ(· x_i) [α _ψ(z_i x_i)- _jQ_ _j(x_i,z_i) ] (8) +λBC(xi,zi∗)∼ℬoff[‖μψ(xi)−zi∗‖22]. + _BCE_(x_i,z_i^*) _off [ \| _ψ(x_i)-z_i^* \|_2^2 ]. Here, α denotes the SAC entropy temperature, Qωj\Q_ _j\ are twin critic networks, and λBC _BC balances demonstration regularization strength. The critics are updated via standard entropy-regularized SAC Bellman loss using the full mixed mini-batch. This hybrid training scheme enables offline demonstrations to stabilize early policy learning and online interactions to adapt to real-world deployment distributions. We gradually decay the offline sampling ratio throughout training to shift optimization from demonstration behavior toward real robot experience. Consistent with our adaptation protocol, the entire flow-matching VLA backbone remains frozen; only the spectral actor, critics, target critics, and entropy temperature are updated during training. Inference Pipeline. At test time, the trained spectral actor πψ _ψ operates purely in the normalized spectral space. Given the current observation xtx_t, the model samples spectral actions, recovers full temporal noise via inverse DCT, and generates refined whole-body action chunks for robot execution: zt∼πψ(⋅∣xt),ct=μc+(σc+δ)⊙zt,ϵt=IDCTK(ct),At=Fθ(ϵt∣ζt)z_t _ψ(· x_t), c_t= _c+( _c+δ) z_t, _t=IDCT_K(c_t), A_t=F_θ( _t _t) (9) The recovered noise ϵt _t is fed into the frozen VLA flow model conditioned on ζt _t to produce final executable whole-body actions. The IDCTKIDCT_K operation zero-pads truncated spectral coefficients to reconstruct smooth full-horizon temporal noise. By restricting RL optimization to the low-dimensional spectral subspace (K≪HK H), our framework significantly reduces optimization complexity, suppresses high-frequency jitter, and yields robust, physically plausible whole-body loco-manipulation behaviors. 4.3 Instantiation on Flow-Matching VLA Backbones The formulation above is independent of a particular flow-matching VLA architecture. We apply the same spectral parameterization, flow-reversal procedure, and offline-to-online objective to both π0.5 _0.5 and HAF-VLA. For π0.5 _0.5, FθF_θ is its native action-flow map and ζt _t denotes its standard visual–language and robot-state conditioning. No change to the pretrained VLA parameters is required. For HAF-VLA, we instantiate the generic flow map only on its final generation stage. Stages 1 and 2 execute normally and provide their clean-action caches, while HAF-Steer replaces only the initial noise of Stage 3. During demonstration inversion, Ct1C_t^1 and Ct2C_t^2 are constructed through teacher forcing; during deployment, they are obtained from the predicted clean actions of the first two stages. In both π0.5 _0.5 and HAF-VLA, the complete VLA backbone remains frozen. 5 Experiments Figure 4: Robot teleoperation data collection setup. The teleoperation pipeline uses an IMU for head control, kinematically matched master arms for dual-arm manipulation, and a joystick to command leg locomotion and waist posture. In this section, we aim to answer four key research questions: Q1. Does HAF-VLA enable long-horizon humanoid loco–manipulation beyond existing state-of-the-art methods? Q2. Does HAF-VLA’s hierarchical action-flow design improve imitation-learning performance? Q3. Does HAF-VLA generalize to unseen visual and positional disturbances? Q4. Can HAF-Steer improve real-world deployment performance through offline-to-online adaptation? 5.1 Experiment Setup Hardware and Data Collection. We conduct real-robot experiments on the TienKung 2.0 and TienKung 3.0 humanoid platforms. Perception relies on an onboard egocentric RGB camera mounted on the robot head. Training demonstrations are collected via isomorphic teleoperation: a kinematic master arm controls bimanual manipulation, a handheld joystick governs locomotion and waist motion, and an inertial measurement unit (IMU) tracks head orientation commands, as visualized in fig. 4. We collect 120 teleoperated trajectories for each household task. Tasks and Evaluation Metric. As illustrated in fig. 5, we benchmark our framework on seven real-world household loco-manipulation tasks: Laundry Loading, Clothes Retrieval, Table Tidy, Basket Transfer, Toy Storage, Ball Tossing, and Box Transfer. These tasks demand long-horizon coordination across locomotion, whole-body posture adjustment, dual-arm manipulation, and object interaction, including challenging skills such as traversing separated regions, torso bending, squatting, object carrying, and throwing. Since binary pass/fail success rates cannot capture partial progress during long sequential executions, we adopt a normalized task score as the primary evaluation metric. Each task is split into predefined milestone sub-goals with individual scores. For each trial rollout, we compute the normalized score as: Score(%)=sachievedsmax×100.Score(\%)= s_achieveds_max× 100. We report the average normalized score across 10 independent rollouts per method and task, which enables fine-grained quantitative comparison across tasks with different completion criteria. Figure 5: Seven real-world humanoid loco-manipulation tasks. Each row depicts a representative execution sequence requiring navigation, whole-body posture adjustment, and physical object interaction. Baselines. We compare our approach against four representative state-of-the-art baselines: (i) ACT [45], a well-established action-chunking imitation learning algorithm; (i) π0.5 _0.5 [29], a strong generalist flow-matching vision-language-action policy; (i) GR00T N1.7 [27], a large-scale pretrained humanoid foundation model for general robot skill learning; and (iv) Cosmos Policy [14], a recent world-model-based framework for robotic manipulation. We strictly follow the official open-source implementations and recommended hyperparameters for all baselines. 5.2 Q1: Does HAF-VLA enable long-horizon humanoid loco–manipulation beyond existing SOTA methods? We evaluate HAF-VLA across seven long-horizon humanoid loco-manipulation tasks requiring navigation, whole-body coordination, and precise object interaction. As summarized in table 1, HAF-VLA achieves the best or tied-best average normalized score on all seven tasks, improving the overall average performance from 53.3% with the strongest baseline π0.5 _0.5 to 70.5%. Performance gains are particularly pronounced for tasks combining locomotion with subsequent fine manipulation, including Box Transfer, Ball Tossing, and Laundry Loading. Empirically, both π0.5 _0.5 and GR00T N1.7 frequently exhibit locomotion drift, while Cosmos Policy is hindered by relatively high online inference latency. Overall, these results demonstrate the effectiveness of hierarchical action generation for high-dimensional humanoid action spaces in whole-body loco-manipulation. Table 1: Main HAF-VLA results on seven long-horizon humanoid loco–manipulation tasks. We report the average normalized task score (%) over 10 rollout trials. Higher values indicate better performance. Task HAF-VLA 0.5 _0.5 GR00T N1.7 Cosmos Policy ACT Laundry Loading 66.7 53.3 40.0 0.0 10.0 Clothes Retrieval 53.3 53.3 33.3 26.7 23.3 Table Tidy 80.0 70.0 40.0 16.7 23.3 Basket Transfer 63.3 50.0 43.3 33.3 16.7 Toy Storage 80.0 53.3 30.0 40.0 23.3 Ball Tossing 56.7 33.3 36.7 3.3 30.0 Box Transfer 93.3 60.0 43.3 73.3 50.0 Average 70.5 53.3 38.1 27.6 25.2 5.3 Q2: Does HAF-VLA’s hierarchical action-flow design improve imitation-learning performance? We conduct targeted ablation experiments on the representative Laundry Loading task to isolate the contribution of HAF-VLA’s hierarchical action-flow design. Full HAF-VLA progressively expands the active action space across three stages, with 10 flow-matching denoising steps per stage and 30 steps in total. We compare it against three variants: (1) All-Joint Denoising, where every stage denoises all action dimensions simultaneously; (2) Arm-First Hierarchy, which reverses the coarse-to-fine generation order by prioritizing manipulation before locomotion; and (3) vanilla π0.5 _0.5 with 30 denoising steps, which controls for the total denoising budget. Table 2: Ablation study of HAF-VLA’s hierarchical action-flow design. We report the average normalized task score (%) over 10 rollout trials on the Laundry Loading task. Higher scores are better. Method / Variant Stage Design Total Steps Norm. Score (%) HAF-VLA Locomotion/Head → + Waist → + Manip. 30 66.7 All-Joint Denoising All joints at every stage 30 53.3 Arm-First Hierarchy Manip. → + Waist → + Locomotion/Head 30 50.0 π0.5 _0.5, 30 steps Full-body action 30 20.0 As shown in table 2, our full hierarchical design outperforms all ablated variants, indicating that the improvement does not simply result from increasing the number of denoising iterations. The inferior performance of the arm-first variant further highlights the importance of establishing locomotion and body posture before fine manipulation. We also observe that increasing π0.5 _0.5 from 10 to 30 denoising steps decreases performance despite the larger computation budget, which is mainly caused by the increased inference latency (from 0.075s to 0.115s). Since π0.5 _0.5 already exhibits locomotion-prediction errors on humanoid tasks, the additional latency can amplify the temporal mismatch between predicted velocity commands and robot execution. In contrast, HAF-VLA remains effective at a similar latency, suggesting greater robustness to inference delay. 5.4 Q3: Does HAF-VLA generalize to unseen visual and positional disturbances? Figure 6: Generalization experiments. Left: object-distraction test during Laundry Loading. Right: position-disturbance test during Clothes Retrieval. We evaluate out-of-distribution robustness under two controlled distribution shifts visualized in fig. 6. For Laundry Loading, we place an unseen black office chair along the robot navigation path to introduce visual distraction. For Clothes Retrieval, we shift the robot initial position 20 cm backward relative to the demonstration collection setup. We compare HAF-VLA against vanilla π0.5 _0.5 using the same normalized task score metric across 10 rollouts per condition. As reported in table 3, HAF-VLA maintains higher task performance under both visual distraction and positional perturbation, which demonstrates stronger generalization beyond the exact demonstration layout. Table 3: Generalization evaluation under controlled disturbances. We report the average normalized task score (%) over 10 rollout trials. Higher scores indicate stronger robustness. Task Disturbance HAF-VLA 0.5 _0.5 Laundry Loading Unseen chair near the path 40.0 26.7 Clothes Retrieval 20 cm backward start shift 43.3 36.7 5.5 Q4: Can HAF-Steer improve real-world deployment performance through offline-to-online adaptation? We evaluate HAF-Steer on two representative whole-body loco-manipulation tasks, Toy Storage and Basket Transfer, using the TienKung 3.0 humanoid platform. For each task, we consider both in-distribution (ID) settings covered by the offline demonstrations and out-of-distribution (OOD) settings with unseen goal locations. Specifically, for Toy Storage, the destination table is moved 30 cm beyond the spatial range covered during demonstration collection; for Basket Transfer, the destination table is similarly shifted by 30 cm from its demonstrated position. These controlled spatial perturbations require the policy to adapt its locomotion and whole-body manipulation behavior to previously unseen goal configurations. To examine whether the proposed spectral adaptation mechanism generalizes across different flow-matching policies, we conduct all experiments on both the vanilla π0.5 _0.5 backbone and HAF-VLA. The corresponding ID and OOD evaluation setups are visualized in fig. 7. Figure 7: In-distribution and out-of-distribution evaluation setups for HAF-Steer. Blue markers indicate destination-table locations covered during demonstration collection, while red markers indicate the OOD evaluation locations. For Toy Storage, the destination table is placed 30 cm beyond the demonstrated spatial range. For Basket Transfer, the destination table is shifted by 30 cm from its demonstrated position. Baselines and Variants. We compare four policy variants. Base Model directly deploys the frozen VLA without latent adaptation. DSRL [35] optimizes a single D-dimensional noise vector and repeats it across the entire action horizon, thereby reducing the search dimension at the cost of removing temporal variation in the flow noise. Noise BC evaluates our spectral actor immediately after the spectral behavior-cloning initialization, without subsequent SAC training. Finally, HAF-Steer further optimizes the BC-initialized spectral actor using mixed offline–online SAC while keeping the entire VLA backbone frozen. This comparison isolates both the effect of the spectral noise representation and the additional benefit of online reinforcement learning. Results Analysis. As shown in fig. 8, HAF-Steer consistently improves both π0.5 _0.5 and HAF-VLA across ID and OOD settings. Noise BC already improves performance in several settings, while subsequent mixed offline–online RL further increases real-world success rates, demonstrating that HAF-Steer can refine deployment performance beyond the offline-initialized policy in both ID and OOD conditions. In contrast, the repeated-noise parameterization of DSRL removes temporal variation and produces unsafe exploratory motions in several real-robot experiments, leading to early training termination. By retaining multiple low-frequency temporal modes, HAF-Steer enables more stable latent-space exploration and effective online adaptation. Importantly, HAF-VLA and HAF-Steer operate at complementary levels. HAF-VLA first provides a structured whole-body policy through hierarchical action generation, while HAF-Steer further adapts this frozen policy to deployment-time distribution shifts. Applying HAF-Steer to HAF-VLA improves its success rate in all four evaluated ID/OOD settings and achieves the best or tied-best performance in three of them. These results demonstrate the benefit of combining structured action generation and lightweight latent policy adaptation within the unified HAF framework. Figure 8: HAF-Steer performance under in-distribution and out-of-distribution settings. We report successful trials out of 10 rollouts on Toy Storage and Basket Transfer using both π0.5 _0.5 and HAF-VLA backbones. Noise BC denotes the spectral actor after behavior-cloning initialization without SAC post-training. † DSRL failed to complete training because unsafe exploratory motions triggered early termination during real-robot rollouts. 6 Conclusion and Limitations In this paper, we introduce HAF, a unified framework for adapting pretrained VLA models to humanoid whole-body loco-manipulation. It comprises two key components: HAF-VLA generates structured sequential actions using progressive denoising and cross-stage KV-cache conditioning, while HAF-Steer leverages SAC-based latent reinforcement learning to optimize DCT-compressed noise coefficients and improve real-world deployment performance through offline-to-online adaptation. HAF successfully unifies structured generative modeling and adaptive policy optimization for complex humanoid whole-body loco-manipulation. Limitations. The hierarchical pipeline increases denoising computation and introduces deployment latency. The latent RL module is also constrained by the base VLA’s inherent priors and may therefore fail to correct erroneous motions in extreme unseen scenarios. Future work will explore efficient denoising schemes and lightweight RL implementations to enhance real-time performance and task robustness. Acknowledgements This work was supported by the National Natural Science Foundation of China (62476011), the Beijing Natural Science Foundation (L252060), and the Beijing Major Science and Technology Project under Contract no. Z191100010618003. References [1] L. Ankile, Z. Jiang, R. Duan, G. Shi, P. Abbeel, and A. Nagabandi (2025) Residual off-policy rl for finetuning behavior cloning policies. arXiv preprint arXiv:2509.19301. Cited by: §2. [2] S. Bai, M. Li, X. Lv, J. Wang, X. Wang, F. Liao, C. Hou, L. Gu, W. Zhou, K. Wu, Z. Ding, Z. Xu, L. Sun, S. Zhang, Z. Che, J. Tang, and B. Chen (2026) HEX: humanoid-aligned experts for cross-embodiment whole-body manipulation. External Links: 2604.07993, Link Cited by: §2. [3] P. J. Ball, L. Smith, I. Kostrikov, and S. Levine (2023) Efficient online reinforcement learning with offline data. In International Conference on Machine Learning, p. 1577–1594. Cited by: §2. [4] Q. Ben, F. Jia, J. Zeng, J. Dong, D. Lin, and J. Pang (2025) HOMIE: humanoid loco-manipulation with isomorphic exoskeleton cockpit. External Links: 2502.13013, Link Cited by: §1. [5] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. (2024) π 0_ 0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164. Cited by: §1, §2. [6] S. Chen, J. Liu, S. Qian, H. Jiang, L. Li, R. Zhang, Z. Liu, C. Gu, C. Hou, P. Wang, Z. Wang, and S. Zhang (2025) AC-dit: adaptive coordination diffusion transformer for mobile manipulation. External Links: 2507.01961, Link Cited by: §1. [7] P. Ding, J. Ma, X. Tong, B. Zou, X. Luo, Y. Fan, T. Wang, H. Lu, P. Mo, J. Liu, et al. (2025) Humanoid-VLA: towards universal humanoid control with visual integration. arXiv preprint arXiv:2502.14795. Cited by: §2. [8] P. Dong, Q. Li, D. Sadigh, and C. Finn (2025) Expo: stable reinforcement learning with expressive policies. arXiv preprint arXiv:2507.07986. Cited by: §2. [9] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine (2018) Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. External Links: 1801.01290, Link Cited by: §1. [10] T. He, Z. Luo, X. He, W. Xiao, C. Zhang, W. Zhang, K. M. Kitani, C. Liu, and G. Shi (2025) OmniH2O: universal and dexterous human-to-humanoid whole-body teleoperation and learning. In Conference on Robot Learning, Cited by: §1. [11] M. Ji, X. Peng, F. Liu, J. Li, G. Yang, X. Cheng, and X. Wang (2024) ExBody2: advanced expressive humanoid whole-body control. arXiv preprint arXiv:2412.13196. Cited by: §1, §1. [12] H. Jiang, J. Chen, Q. Bu, L. Chen, M. Shi, Y. Zhang, D. Li, C. Suo, C. Wang, Z. Peng, and H. Li (2026) WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1, §1, §2. [13] Y. Jiang, R. Zhang, J. Wong, C. Wang, Y. Ze, H. Yin, C. Gokmen, S. Song, J. Wu, and L. Fei-Fei (2025) BEHAVIOR robot suite: streamlining real-world whole-body manipulation for everyday household activities. External Links: 2503.05652, Link Cited by: §1. [14] M. J. Kim, Y. Gao, T. Lin, Y. Lin, Y. Ge, G. Lam, P. Liang, S. Song, M. Liu, C. Finn, and J. Gu (2026) Cosmos policy: fine-tuning video models for visuomotor control and planning. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1, §5.1. [15] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn (2025) OpenVLA: an open-source vision-language-action model. In Proceedings of The 8th Conference on Robot Learning, P. Agrawal, O. Kroemer, and W. Burgard (Eds.), Proceedings of Machine Learning Research, Vol. 270, p. 2679–2713. External Links: Link Cited by: §2. [16] K. Lei, H. Li, D. Yu, Z. Wei, L. Guo, Z. Jiang, Z. Wang, S. Liang, and H. Xu (2025) Rl-100: performant robotic manipulation with real-world reinforcement learning. arXiv preprint arXiv:2510.14830. Cited by: §2. [17] J. Li, X. Cheng, T. Huang, S. Yang, R. Qiu, and X. Wang (2025) AMO: adaptive motion optimization for hyper-dexterous humanoid whole-body control. External Links: 2505.03738, Link Cited by: §2. [18] Y. Li, Z. Luo, T. Zhang, C. Dai, A. Kanervisto, A. Tirinzoni, H. Weng, K. Kitani, M. Guzek, A. Touati, A. Lazaric, M. Pirotta, and G. Shi (2025) BFM-zero: a promptable behavioral foundation model for humanoid control using unsupervised reinforcement learning. External Links: 2511.04131, Link Cited by: §2. [19] Y. Li, Y. Zhang, W. Xiao, C. Pan, H. Weng, G. He, T. He, and G. Shi (2025) Learning gentle humanoid locomotion and end-effector stabilization control. arXiv preprint arXiv:2505.24198. Cited by: §1. [20] Y. Li, X. Ma, J. Xu, Y. Cui, Z. Cui, Z. Han, L. Huang, T. Kong, Y. Liu, H. Niu, W. Peng, J. Qiao, Z. Ren, H. Shi, Z. Su, J. Tian, Y. Xiao, S. Zhang, L. Zheng, H. Li, and Y. Wu (2025) GR-rl: going dexterous and precise for long-horizon robotic manipulation. External Links: 2512.01801, Link Cited by: §1, §2. [21] Q. Liao, T. E. Truong, X. Huang, Y. Gao, G. Tevet, K. Sreenath, and C. K. Liu (2025) BeyondMimic: from motion tracking to versatile humanoid control via guided diffusion. arXiv preprint arXiv:2508.08241. Cited by: §1, §1, §2. [22] S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu (2025) RDT-1b: a diffusion foundation model for bimanual manipulation. External Links: 2410.07864, Link Cited by: §2. [23] J. Lu, X. Qin, Y. Jiang, K. Wang, C. Zhang, B. Liang, J. Yang, M. Xu, and L. Zhao (2026) UniSteer: unified noise steering for efficient human-guided vla adaptation. External Links: 2605.10821, Link Cited by: §1, §2. [24] J. Luo, Z. Hu, C. Xu, Y. L. Tan, J. Berg, A. Sharma, S. Schaal, C. Finn, A. Gupta, and S. Levine (2024) Serl: a software suite for sample-efficient robotic reinforcement learning. In 2024 IEEE International Conference on Robotics and Automation (ICRA), p. 16961–16969. Cited by: §2. [25] Z. Luo, Y. Yuan, T. Wang, C. Li, F. Castañeda, S. Chen, Z. Cao, J. Li, D. Minor, Q. Ben, J. Park, D. Sami, Z. Wang, X. Da, R. Ding, C. Hogg, L. Song, E. Lim, E. Jeong, T. He, H. Xue, W. Xiao, S. Yuen, J. Kautz, Y. Chang, U. Iqbal, L. ". Fan, and Y. Zhu (2026) SONIC: supersizing motion tracking for natural humanoid whole-body control. External Links: 2511.07820, Link Cited by: §2. [26] D. McAllister, S. Ge, B. Yi, C. M. Kim, E. Weber, H. Choi, H. Feng, and A. Kanazawa (2025) Flow matching policy gradients. arXiv preprint arXiv:2507.21053. Cited by: §2. [27] NVIDIA, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al. (2025) GR00T N1: an open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734. Cited by: §1, §2, §5.1. [28] X. B. Peng, P. Abbeel, S. Levine, and M. van de Panne (2018) DeepMimic: example-guided deep reinforcement learning of physics-based character skills. ACM Trans. Graph. 37 (4), p. 143:1–143:14. External Links: ISSN 0730-0301, Link, Document Cited by: §2. [29] Physical Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al. (2025) π 0.5_0.5: a vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054. Cited by: §1, §1, §2, §5.1. [30] A. Ren, J. Lidard, L. Ankile, A. Simeonov, P. Agrawal, A. Majumdar, B. Burchfiel, H. Dai, and M. Simchowitz (2025) Diffusion policy policy optimization. In International Conference on Learning Representations, Vol. 2025, p. 77288–77329. Cited by: §2. [31] Y. Shao, X. Huang, B. Zhang, Q. Liao, Y. Gao, Y. Chi, Z. Li, S. Shao, and K. Sreenath (2025) LangWBC: language-directed humanoid whole-body control via end-to-end learning. External Links: 2504.21738, Link Cited by: §2, §2. [32] J. Shi, X. Liu, D. Wang, O. Lu, S. Schwertfeger, F. Sun, C. Bai, and X. Li (2025) Adversarial locomotion and motion imitation for humanoid policy learning. arXiv preprint arXiv:2504.14305. Cited by: §1. [33] Z. Sun and S. Song (2026) From prior to pro: efficient skill mastery via distribution contractive rl finetuning. arXiv preprint arXiv:2603.10263. Cited by: §2. [34] A. Tang, W. Chen, A. Wagenmaker, C. Finn, and S. Levine (2026) Improving robotic generalist policies via flow reversal steering. arXiv preprint arXiv:2606.13675. Cited by: §1, §2, §4.1. [35] A. Wagenmaker, M. Nakamoto, Y. Zhang, S. Park, W. Yagoub, A. Nagabandi, A. Gupta, and S. Levine (2025) Steering your diffusion policy with latent space reinforcement learning. arXiv preprint arXiv:2506.15799. Cited by: §1, §2, §4.1, §5.5. [36] Y. Wang, X. Li, W. Wang, J. Zhang, Y. Li, Y. Chen, X. Wang, and Z. Zhang (2025) Unified vision-language-action model. External Links: 2506.19850, Link Cited by: §2. [37] S. Wei, H. Jing, B. Li, Z. Zhao, J. Mao, Z. Ni, S. He, J. Liu, X. Liu, K. Kang, S. Zang, W. Yuan, M. Pavone, D. Huang, and Y. Wang (2026) Ψ0 _0: An open foundation model towards universal humanoid loco-manipulation. External Links: 2603.12263, Link Cited by: §1, §2. [38] H. Xue, X. Huang, D. Niu, Q. Liao, T. Kragerud, J. T. Gravdahl, X. B. Peng, G. Shi, T. Darrell, K. Sreenath, and S. S. Sastry (2025) LeVERB: humanoid whole-body control with latent vision-language instruction. CoRR abs/2506.13751. External Links: Link, Document, 2506.13751 Cited by: §2, §2. [39] J. Yang, M. S. Mark, B. Vu, A. Sharma, J. Bohg, and C. Finn (2024) Robot fine-tuning made easy: pre-training rewards and policies for autonomous real-world reinforcement learning. In 2024 IEEE International Conference on Robotics and Automation (ICRA), p. 4804–4811. Cited by: §2. [40] Y. Ze, Z. Chen, W. Wang, T. Chen, X. He, Y. Yuan, X. B. Peng, and J. Wu (2025) Generalizable humanoid manipulation with 3d diffusion policies. In IROS, Cited by: §1. [41] Y. Ze, S. Zhao, W. Wang, A. Kanazawa, R. Duan, P. Abbeel, G. Shi, J. Wu, and C. K. Liu (2025) TWIST2: scalable, portable, and holistic humanoid data collection system. External Links: 2511.02832, Link Cited by: §2. [42] Y. Zhang, Y. Yuan, P. Gurunath, T. He, S. Omidshafiei, A. Agha-mohammadi, M. Vazquez-Chanlatte, L. Pedersen, and G. Shi (2025) FALCON: learning force-adaptive humanoid loco-manipulation. arXiv preprint arXiv:2505.06776. Cited by: §1. [43] Y. Zhang, L. Ke, A. Deshpande, A. Gupta, and S. Srinivasa (2023) Cherry-picking with reinforcement learning: robust dynamic grasping in unstable conditions. arXiv preprint arXiv:2303.05508. Cited by: §2. [44] Z. Zhang, C. Chen, H. Xue, J. Wang, S. Liang, Y. Liu, Z. Zhang, H. Wang, and L. Yi (2025) Unleashing humanoid reaching potential via real-world-ready skill space. arXiv preprint arXiv:2505.10918. Cited by: §1. [45] T. Z. Zhao, V. Kumar, S. Levine, and C. Finn (2023) Learning fine-grained bimanual manipulation with low-cost hardware. External Links: 2304.13705, Link Cited by: §5.1. [46] J. Zheng, J. Li, Z. Wang, D. Liu, X. Kang, Y. Feng, Y. Zheng, J. Zou, Y. Chen, J. Zeng, T. Wang, Y. Zhang, J. Liu, and X. Zhan (2026) X-VLA: soft-prompted transformer as scalable cross-embodiment vision-language-action model. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §2. [47] B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V. Vanhoucke, H. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. Sanketi, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y. Lu, S. Levine, L. Lee, T. E. Lee, I. Leal, Y. Kuang, D. Kalashnikov, R. Julian, N. J. Joshi, A. Irpan, B. Ichter, J. Hsu, A. Herzog, K. Hausman, K. Gopalakrishnan, C. Fu, P. Florence, C. Finn, K. A. Dubey, D. Driess, T. Ding, K. M. Choromanski, X. Chen, Y. Chebotar, J. Carbajal, N. Brown, A. Brohan, M. G. Arenas, and K. Han (2023) RT-2: vision-language-action models transfer web knowledge to robotic control. In Proceedings of The 7th Conference on Robot Learning, J. Tan, M. Toussaint, and K. Darvish (Eds.), Proceedings of Machine Learning Research, Vol. 229, p. 2165–2183. External Links: Link Cited by: §2.