Paper deep dive
LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action
Qi Lyu, Baicheng Liu, Xudong Wang, Jiahua Dong, Lianqing Liu, Zhi Han
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/10/2026, 4:27:58 AM
Summary
The paper introduces LEEVLA, a Vision-Language-Action (VLA) architecture designed to enhance robot manipulation by explicitly guiding attention to task-critical regions and preserving structured latent environment evolution. It integrates Drift-Guided Dynamic Prioritization (DGDP), which combines Dynamic Position Prioritization (DPP) and Semantic Drift Guidance (SDG) to focus on salient, instruction-relevant features, with Structured Feature Flow Generation (SFFG), which employs Prototype-to-Periphery (P2P) prediction and a Mutual-neighborhood Contrastive (MC) loss to maintain topological consistency in latent space. Extensive experiments demonstrate that LEEVLA consistently outperforms existing VLA methods on standard benchmarks.
Entities (12)
Relation Signals (13)
LEEVLA → isa → VLA
confidence 98% · LEEVLA is a VLA architecture for seeing what matters in Latent Environment Evolution
LEEVLA → incorporates → DGDP
confidence 95% · LEEVLA adopts drift-guided dynamic prioritization (DGDP) to automatically discover salient and instruction-relevant regions
LEEVLA → incorporates → SFFG
confidence 95% · we introduce structured feature flow generation (SFFG) to make the model reason over these regions on how to evolve
DGDP → consistsof → DPP
confidence 92% · DGDP, which combines dynamic position prioritization (DPP) with semantic drift guidance (SDG)
DGDP → consistsof → SDG
confidence 92% · DGDP, which combines dynamic position prioritization (DPP) with semantic drift guidance (SDG)
SFFG → consistsof → MC Loss
confidence 90% · ensuring semantic neighborhood consistency via mutual-neighborhood contrastive (MC) loss
SFFG → consistsof → P2P
confidence 90% · SFFG strategy... through prototype-to-periphery (P2P) prediction mechanism
P2P → enforces → structured latent evolution
confidence 88% · P2P forecasting loss function... to model structured state evolution
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasize task-critical evidence and ignore underlying factors. To address this issue, we propose LEEVLA, a VLA architecture for seeing what matters in Latent Environment Evolution that explicitly guides the model toward informative regions while preserving the structured evolution of latent world representations. To identify salient and instruction-relevant regions, we introduce drift-guided dynamic prioritization (DGDP), which combines dynamic position prioritization (DPP) with semantic drift guidance (SDG) to guide the VLA agent where to attend during training. On top of this, we introduce structured feature flow generation (SFFG), which models how these prioritized features should evolve in latent space via prototype-to-periphery (P2P) prediction, and a mutual-neighborhood contrastive (MC) loss to maintain topological consistency among neighborhoods. Together, DGDP and SFFG form a task-aware "where-how" training framework. Extensive experiments on VLA benchmarks show that LEEVLA consistently outperforms prior methods, confirming that explicit task-evidence guidance and structured latent reasoning are both crucial for scalable VLA. Our code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.08182v1
- Canonical: https://arxiv.org/abs/2607.08182v1
Trouble viewing inline? Open PDF directly →
Full Text
55,071 characters extracted from source content.
Expand or collapse full text
LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action Qi Lyu 1,2,3 , Baicheng Liu 1,2 , Xudong Wang 1,2,3 , Jiahua Dong 4 , Lianqing Liu 1,2 , Zhi Han 1,2 1 State Key Laboratory of Robotics and Intelligent Systems 2 Shenyang Institute of Automation, Chinese Academy of Sciences 3 University of Chinese Academy of Sciences 4 Mohamed bin Zayed University of Artificial Intelligence Abstract Vision-language-action (VLA) models aim to map multi- modal inputs to robot actions. However, most existing ap- proaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to empha- size task-critical evidence and ignore underlying factors. To address this issue, we propose LEEVLA, a VLA ar- chitecture for seeing what matters in Latent Environment Evolution that explicitly guides the model toward infor- mative regions while preserving the structured evolution of latent world representations. To identify salient and instruction-relevant regions, we introduce drift-guided dy- namic prioritization (DGDP), which combines dynamic po- sition prioritization (DPP) with semantic drift guidance (SDG) to guide the VLA agent where to attend during training. On top of this, we introduce structured feature flow generation (SFFG), which models how these priori- tized features should evolve in latent space via prototype- to-periphery (P2P) prediction, and a mutual-neighborhood contrastive (MC) loss to maintain topological consistency among neighborhoods. Together, DGDP and SFFG form a task-aware “where–how” training framework. Exten- sive experiments on VLA benchmarks show that LEEVLA consistently outperforms prior methods, confirming that ex- plicit task-evidence guidance and structured latent reason- ing are both crucial for scalable VLA. Our code is available in the https://github.com/LyuQi127/LEEVLA. 1. Introduction Vision-Language-Action (VLA) [3, 8, 11, 15, 18, 23, 29, 42, 53, 58, 63] models aim to ground perception and lan- guage in action, mapping observation and language in- structions directly to low-level controls in the closed loop. Context Clues Multimodal inputs (b) Prior methods(c) Our method VLA Multimodal inputs Action tokens Action Sample Action LEEVLA Action tokens Sample Instruction: pick up the black bowl 푍 ଵ 푍 ଶ 푍 ଷ 푍 푍 ଵ 푍 ଶ 푍 ଷ 푍 ஶ 푍 Unknown factors 푍 Known factors 푎 ௧ାଵ 푎 ௧ା் Multimodal inputs (a) Causal graph Actions 푍 푍 Evolve Spatial CurrentFuture Task relevance Semantic Context Clues Figure 1. Comparison between our method and prior methods. (a): Causal graph between multimodal inputs and action chunk. (b): Prior methods: VLA reasons on factors of specific context clues. (c): Our method: VLA reasons on abstract context clues. By jointly encoding visual observations, proprioception, and language instruction, VLAs offer a path toward robots that can perceive, reason, and act in open environments. Recent research typically employs large language models (LLMs) or visual-language models (VLMs) to construct VLAs [31, 35, 43, 46], and then trains these VLAs on large- scale robot demonstrations or simulated interaction traces [28, 34]. Notably, the generation of action is governed by both known and unknown factors [47], as illustrated by the causal graph in Fig. 1 (a). Recently, several works have further introduced explicit intermediate reasoning prior to decoding action to strengthen world understanding and task decomposition. As shown in Fig. 1 (b), these methods con- fine the scope of reasoning to specific context clues selected 1 arXiv:2607.08182v1 [cs.CV] 9 Jul 2026 by humans. Such vision-language reasoning is effective in sharpening spatial perception and making downstream poli- cies more interpretable. However, those methods based on reasoning still suffer from a restricted search space due to their dependence on external conditions or priors such as subgoal images, segmentation, or depth [57, 58, 60]. Considering that most available embodied demonstra- tion datasets [20, 21, 28, 32, 56] have restricted modality diversity (e.g., few scene layouts, limited object categories, and limited concepts), methods depending on specific con- text clues require auxiliary models to produce these modal- ities.Beyond incurring additional computational over- head, such methods are contingent on progress in auxil- iary tasks. Consequently, these methods [57, 58] overem- phasize known factors while under-exploring unknown but task-relevant factors [6, 16], leading to fragility when the same object varies across modalities or different objects look similar within a single modality. As a result, such de- signs can inadvertently limit the exploratory capacity of the model in latent space, making it harder for the backbone to discover task-critical but unknown factors [49]. To address these limitations, we advocate reasoning di- rectly in the latent feature space, treating actions as drivers of how environment states evolve in high-dimensional space, as shown in Fig. 1 (c). The success of general vi- sion models [36, 40, 55] on all sorts of tasks [25, 38, 52] suggests that features produced by general vision models embed multimodal information, such as category and depth. Through reasoning over these latent representations, a pol- icy can jointly exploit semantic, appearance, and geomet- ric cues that are entangled for the same object [1, 27]. Specifically, considering only depth makes it difficult to recover the spatial position of the target object under par- tial occlusion, while other modalities can help the robot localize it. Operating in latent space also avoids auxiliary pixel-level reconstruction or externally supplied condition pipelines during training, reducing overhead while enlarg- ing search space of the policy to uncover both known and unknown task-relevant factors [47]. Additionally, naive pre- diction in a high-dimensional space often diverts attention toward static background or instruction-irrelevant objects, undermining reasoning. Thus, deciding where to attend is equally critical: our insight is to prioritize regions that ex- hibit pronounced spatial change and whose semantic evo- lution aligns with the language instruction, so the model concentrates supervision and capacity on scene components that are causally tied to the intended manipulation. To this end, we propose Latent Environment Evolution VLA (LEEVLA), in which drift-guided dynamic prioritiza- tion (DGDP) tells the agent where to attend, while struc- tured feature flow generation (SFFG) models how to evolve latent environment representation. LEEVLA adopts drift- guided dynamic prioritization (DGDP) to automatically dis- cover salient and instruction-relevant regions in the feature space via dynamic position prioritization (DPP) and seman- tic drift guidance (SDG), thereby “focusing” the model on where to attend during training. At the same time, we intro- duce structured feature flow generation (SFFG) to make the model reason over these regions on how to evolve by en- forcing prototype-to-periphery (P2P) prediction. Further- more, we found that the latent space is inevitably con- taminated by spurious neighbors and asymmetric affinity. Similar to the phenomenon observed in passive discrim- inant analysis, where only reciprocal neighbors or high- affinity neighbors can reliably share semantics, we intro- duce a mutual-neighborhood contrastive (MC) loss to fil- ter out these noisy links and maintain semantic neighbor- hood consistency. To sum up, drift-guided dynamic prior- itization (DGDP) explicitly steers the model toward task- critical context clues, while structured feature flow genera- tion (SFFG) preserves the structured evolution of environ- ment representations. Through extensive experimentation, we demonstrate that our approach achieves state-of-the-art performance among existing methods. Our contributions include the following points: • We propose a drift-guided dynamic prioritization (DGDP) mechanism for automatic discovery of where to attend, composed of dynamic position prioritization (DPP) and semantic drift guidance (SDG), which iden- tifies dynamically active and instruction-relevant regions. • We propose a structured feature flow generation (SFFG) strategy incorporating prototype-to-periphery (P2P) pre- diction and mutual-neighborhood contrastive (MC) loss, guiding the model to learn how to evolve in latent space. • Extensive experiments shows consistent performance gains, validating that explicit task-relevance guidance and structured latent reasoning together enhance the general- ization and long-horizon reasoning capability of VLA. 2. Related Works 2.1. Vision–Language–Action Models Vision–language–action (VLA) models [5, 10, 23, 24, 26, 48, 57, 59, 60] aim to map multi-view visual observations and natural-language instructions to robot actions in a sin- gle sequence modeling framework. PaLM-E [10] showed that injecting embodied signals into a pretrained language model enables grounded instruction following, and RT- 2 [63] further demonstrated that internet-scale vision– language pretraining can transfer to real robot manipula- tion and arrangement tasks. More recent open efforts such as OpenVLA [23] and Octo [42] make this recipe acces- sible by standardizing the use of a VLM/LLM backbone combined with an action head over large cross-embodiment datasets, while π 0 [3, 37] improve action expressivity via flow matching on top of a frozen VLM. CogACT [26] 2 ⋮ Vision Encoder Projector 푎 ௧ାଵ Continuous Action Chunk 휃 Semantic Vision ⋮ Compute Semantic Drift Normalize Prototype-to-Periphery Forecasting Loss 푐 ଷ 푐 푐 ଵ 푐 ଶ Drift-Guided Dynamic Prioritization Dynamic Position Prioritization Semantic Drift Guidance 휔 Visual observation at time 푡 Language embedding Nonlinear Function Product Operation Patch-wise feature (time 푡+푇) Patch-wise feature (time 푡) Forward (only training) Instruction pick up the black bowl on the wooden cabinet and place it on the plate. Euclidean distance Mutual-neighborhood Contrastive Loss 456123 145623 14 5 62 3 1 3 2 5 6 4 1 KNN-Map 푎 ௧ା் 푎 ௧ାଶ Parallel Continuous Action Predictor Future Feature Decoder Cluster boundary Positive Negative Nearest neighborhood Structured Feature Flow Generation Tokenizer Action prompt Large Language Model Figure 2. Overview of our LEEVLA. The Structured Feature Flow Generation is composed of Prototype-to-Periphery (P2P) mechanism and a Mutual-neighborhood Contrastive (MC) Loss as shown in the pink panel. A Future Feature Decoder predicts features at time step t+ T within the target latent space. We impose a P2P forecasting loss to model structured state evolution. For each feature at t+ T , we build a second-order k-N mutual-neighborhood graph and optimize the MC loss. The Drift-Guided Dynamic Prioritization (DGDP) module (yellow panel) computes prioritization weight based on the Dynamic Position Prioritization (DPP) and Semantic Drift Guidance (SDG). The Parallel Continuous Action Predictor (action policy) consumes embeddings from LLM to output continuous actions. factorizes cognition and action by attaching a diffusion- transformer-based action module to a pretrained VLM and demonstrates excellent performance. Recent works fur- ther improve VLA perception by selecting or compress- ing visual tokens. OTTER [19] extracts instruction-aligned visual features with pretrained vision-language alignment, and Compressor-VLA [12] compresses instruction-relevant visual tokens for efficient manipulation.In contrast, LEEVLA keeps the inference token stream unchanged and uses future-feature prediction with structured latent topol- ogy constraints during training, teaching action-conditioned environment dynamics without extra inference memory overhead. However, existing VLA models typically treat all tokens or patches uniformly during training, which may dilute supervision on task-relevant regions [22].Based on those prior works, we keep the VLA formulation. Si- multaneously, we introduce drift-guided dynamic prioriti- zation (DGDP) and a structured feature flow generation mechanism, which enable the vision-language-action model to learn where to attend and how these attended features should evolve. 2.2. World Models World models [4, 9, 13, 17, 27, 50, 51, 61, 62] learn to forecast future states.CASCADE [51] learns a world model from collaboratively gathered data across multiple agents under an information-theoretic objective motivated by Bayesian active learning. DreamerV3 [13] showed that scalable latent dynamics can outperform specialized model- free methods across many domains. GAIA-1 [17] per- forms action-conditioned world-model video generation for autonomous driving, and Genie [4] learns a latent action space enabling unsupervised and action-controllable inter- active environments. RoboDreamer [61] brought action- conditioned video/world modeling closer to robot manipu- lation by factorizing objects, goals, and actions. UWM [62] integrates the action diffusion process and the video diffu- sion process within a unified Transformer architecture, ef- fectively combining policy generation with robot dynamics. DreamZero [54] jointly models videos and actions for zero- shot policy learning. Unlike video-generation-based con- trol, LEEVLA uses latent-space future prediction only as auxiliary training supervision for the action policy. 3. Preliminaries Following prior VLA models [23, 24], the robot observes a third-person image I p t ∈R H×W×C , a wrist-camera im- age I w t ∈R H×W×C , proprioception s t , and a language in- struction l at time step t. Here H , W and C denote the height, width, and number of channels of the visual in- puts. We adopt pretrained visual backbones (DinoV2 [36] and SigLIP [55]) to form visual encoder E v (·) that pro- duces N v patch features f p t = E v (I p t ) = v p t,i N v i=1 and f w t = E v (I w t ) =v w t,i N v i=1 , where each patch vector lives in R d 0 and i indexes spatial locations. A projectorP :R d 0 → R d is applied patch-wise to obtain model-dimension tokens x p t = P(f p t ) ∈R N v ×d and x w t = P(f w t ) ∈R N v ×d . The language instructions l are encoded by a tokenizer Ψ(·) into c = Ψ(l) ∈R M×d , and proprioception s t is embedded by E s (·) into x s t = E s (s t ) ∈R d . However, prior VLA mod- 3 els [23, 24] struggle to handle complex tasks due to lack of perception of the future state of the environment. To incor- porate reasoning of future representation, we build a future feature decoderD, which maps the vision embedding to the future feature space: D :R d →R d 0 . The detailed archi- tecture of the future feature decoder is presented in the sup- plementary material §9. The parallel continuous action pre- dictor (action policy) generates the continuous action chunk by a 1:T ∈R T×F , where T denotes the length of the action chunk and F is the action-space degrees of freedom. 4. Methodology In this section, we introduce drift-guided dynamic prioriti- zation (DGDP, §4.1), which consists of dynamic position prioritization (DPP) and semantic drift guidance (SDG), to automatically discover key task-relevant regions during training. Subsequently, we introduce the structured feature flow generation (SFFG, §4.2) strategy, allowing the model to perform structured reasoning in latent space through prototype-to-periphery (P2P) prediction, while ensuring se- mantic neighborhood consistency via mutual-neighborhood contrastive (MC) loss. Finally, §4.3 details our training ob- jective, where an L 1 action regression loss is combined with the P2P and MC losses to jointly optimize continuous ac- tion generation and future feature structure. We visualize an overview of our LEEVLA in Fig. 2. Detailed hyperpa- rameter settings are included in the supplementary material. Notably, DGDP and SFFG are used only during training and do not incur any additional inference cost during test time. 4.1. Drift-Guided Dynamic Prioritization Establishing where to attend is essential. Paying equal attention to all information in the observation space can easily lead to gradients being diluted across static back- grounds and objects irrelevant to the task. To steer optimiza- tion direction toward task-relevant regions, we quantify fea- ture dynamics between adjacent timestamps and estimate semantic-drift direction of each patch, then adaptively mod- ulate its task relevance. Dynamic Position Prioritization. When forecasting fu- ture representations in the feature space, the semantic dy- namics vary markedly across spatial regions: local fea- tures associated with the robot and task-relevant objects typically exhibit strong temporal variability, whereas back- ground or otherwise static regions remain comparatively stable. Therefore, we introduce a dynamic position prioriti- zation mechanism that adaptively modulates prediction-loss weights at the feature level, guiding the model to focus on regions more sensitive to task execution and environment interaction. We quantify the dynamism score at each spatial position via the change in cosine similarity between adja- cent time steps. For patch i, the feature v t,i at time t moves to v t+T,i after the execution of T time steps. The dynamism (a)(b) Figure 3. Example of Weight Visualization. (a): Visualization of dynamic position prioritization (DPP). (b): Visualization of drift-guided dynamic prioritization (DGDP). The white arrow in- dicates the target direction of the robotic arm’s movement. The se- mantic drift guidance factor suppresses the weights of instruction- irrelevant background regions (as indicated by the yellow box) and enhances regions with significant semantic changes along edges (as indicated by the blue box). score θ i is denoted as: θ i = 1− cos v t,i , v t+T,i , i = 1,..., 2N v ,(1) where cos(·) is the cosine similarity function. As shown in Fig. 3 (a), the regions exhibiting salient changes in vi- sual features between time t and t + T are assigned higher dynamism scores, indicating stronger attentional focus. Semantic Drift Guidance. Focusing solely on dynamic regions ignores the directionality of task signals, leading the model to over-attend to patches whose semantics drift from task-relevant to task-irrelevant. To address this, we introduce semantic drift guidance.First, we assign an instruction-relevant score r to each patch of time t and t+T according to the language instructions. Instruction-relevant score assignment is denoted as: r t,i = max 1≤j≤M ⟨ ̃ x t,i , ̃ c j ⟩,(2) where⟨·,·⟩ denotes the Euclidean inner product, ̃ x t,i is the normalized vision token of patch i at time t, ̃ c j denotes the normalized token of language token j, M indicates the number of language tokens. Based on Eq. 2, we then ob- tain r t,i and r t+T,i which represent the instruction-relevant scores of patch i at time t and t + T . We define instruction-relevant semantic drift ∆ i as: ∆ i = clip( r t+T,i − r t,i τ , −δ, δ), (3) where clip(·) is a numerical constraint function with a boundary of δ and τ is temperature. Both τ and δ are pos- itive. To ensure a consistent dynamic range across sam- ples and prevent domination by outliers, we normalize the instruction-relevant score to the interval [-1, 1] and map it to 4 the semantic drift guidance factor. Semantic drift guidance factor of patch i is denoted as: ω i = exp( ̃ ∆ i τ ), (4) where the semantic drift ̃ ∆ i = 2·(∆ i −min(∆)) max(∆)−min(∆) − 1. We can adjust the intensity of the modulation by controlling the temperature. Finally, we couple semantics and dynamics to yield the prioritization weight for each token: β i = σ ω i · θ i ,(5) where σ(·) is a non-linear function. In this work, we employ the sigmoid function. Fig. 3 (b) illustrates a visualization example of DGDP weights. Through the DGDP compo- nent, the model maintains hierarchical focus over visual fea- tures, assigning higher importance to dynamic patch-level features whose semantics evolve toward the instruction- relevant, and lower importance to static features that drift toward background semantics. 4.2. Structured Feature Flow Generation It is crucial to specify how to evolve environment repre- sentation in the latent space. Flattened token prediction, generated from top left to bottom right, corrupts the local structure of the feature space [14, 44], i.e., features from the same semantic unit are split due to sequence order, which impairs the agent’s spatial reasoning ability. This disruption breaks contextual continuity. To address it, we propose structured feature flow generation (SFFG) strat- egy. SFFG alleviates semantic fragmentation caused by flat prediction through prototype-to-periphery (P2P) prediction mechanism. We also leverage mutual-neighborhood con- trastive loss to align semantically similar features, thereby preserving the topology of the visual feature space. Prototype-to-Periphery (P2P) prediction. We perform joint clustering on features of multi-view observations to obtain setC, where each element represents a cluster: C =F [f p t+T ; f w t+T ]) =c l L l=1 , μ ℓ = 1 |c ℓ | X j∈c ℓ f t+T,j , (6) whereF (·) represents clustering operator, [·;·] denotes se- quence concatenation along the token dimension, L is the number of clusters,|·| indicates the set cardinality, and μ ℓ denotes centroid of cluster c ℓ . And then we sort members by Euclidean distance between members and centroid from nearest to farthest (prototype → periphery) within each c ℓ as follows: −→ c ℓ =v ε ℓ (k) |c ℓ | k=1 , s.t. ε ℓ (1)≤·≤ ε ℓ (|c ℓ |),(7) where ε ℓ (k) =∥ v i − μ ℓ ∥ 2 denotes Euclidean distance be- tween visual feature v i ∈ c ℓ and centroid μ l . Based on the ordered sequence constructed by Eq. 7, the P2P forecasting loss function is expressed as: L P2P = 1 |C| |C| X ℓ=1 1 | −→ c ℓ | | −→ c ℓ | X j∈ −→ c ℓ α + β j φ ˆv t+T,j , v t+T,j , (8) where α denotes the global modulation factor to maintain attention to global information, φ(ˆv t+T,j , v t+T,j ) = 1− cos(ˆv t+T,j , v t+T,j ) represents the cosine embedding loss, and ˆv t+T is the visual feature predicted at time t for time t + T . In this work, we set α = 1. Mutual-neighborhood Contrastive (MC) Loss.To achieve more robust contrastive supervision under noisy clustering, we construct contrastive pairs based on high- confidence neighborhood relations in the feature space. Let S ij = cos(v t+T,i ,v t+T,j ) denote the cosine similarity be- tween the future visual feature of samples i and j. For each anchor i, we first form a first-order neighbor set: G (1) i = TopK S iℓ ℓ̸=i which keeps the K most simi- lar tokens to i. To further enlarge the pool of potentially clean positives while still staying in a locally consistent region, we define a second-order neighbor set: G (2) i = S j∈G (1) i TopM S jr r̸=j , i.e. the union of the M near- est neighbors of each first-order neighbor. We set K = 10 and M = 5. We then select only those tokens that are mu- tual neighbors to suppress the asymmetric or spurious links introduced by clustering noise. Concretely, the positive set for anchor i is: G + i =j | j ∈G (1) i ∪G (2) i , i∈G (1) j ∪G (2) j .(9) We adopt the InfoNCE loss [7, 33] over these mutual- neighborhood positives to pull them closer in the feature space. The mutual-neighborhood contrastive loss is cal- culated as a function of the similarity relationships among samples within their respective neighborhoods in the repre- sentation space, and is formally defined as follows: L MC =− 1 |I| X i∈I 1 |G + i | X j∈G + i log exp S ij /τ c P ℓ∈V\i exp S iℓ /τ c , (10) where I = i ∈ V | |G + i | > 0 represents non-empty set of positive samples, V = 1,...,2N v is index set repre- senting multi-view visual features, and τ c is the tempera- ture. Rather than simply enlarging the neighborhood size, the mutual-neighborhood mechanism adaptively identifies high-confidence and symmetric feature relations, provid- ing more stable supervision under noisy and weakly labeled robot demonstration data. 4.3. Training Objective Similar to [24], we adopt an L 1 regression strategy and em- ploy parallel action decoding, which is efficient and tends 5 to produce more accurate actions. The action policy is an MLP head that directly regresses continuous actions from the last-layer hidden states of the large language model. Training minimizes the average L 1 distance to ground truth actions to filter noise from the training demonstrations [24]. The action prediction loss is calculated as : L action = 1 T T X i=1 |a t+i − ˆa t+i | 1 ,(11) where ˆa t+i represents the predicted action of time t + i and |·| is the L 1 norm. The overall training objective is expressed as: L total = λ 1 L action + λ 2 L P2P + λ 3 L MC ,(12) where λ 1 , λ 2 , and λ 3 are hyperparameters that balance the contributions of the action regression loss, prototype-to- periphery (P2P) forecasting loss, and mutual-neighborhood contrastive (MC) loss. 5. Experiment 5.1. Implementation Details LEEVLA-large is initialized from OpenVLA-7B and fur- ther pretrained on a large mixture of datasets from Open X- Embodiment [34], which covers diverse robot and vision– language trajectories. We train LEEVLA-large for 50k to 150k optimization steps, where more challenging tasks typ- ically require longer training schedules. LEEVLA-mini is initialized from miniVLA [2, 23], which is pretrained on LIBERO-90 [28], and we train LEEVLA-mini for 20k to 50k steps. For LEEVLA-large, we use a learning rate of 5× 10 −4 ; for LEEVLA-mini, we use 2× 10 −5 . All models are optimized using an AdamW optimizer [30], with both training and inference performed on a computing infrastruc- ture equipped with 8× A100 (80 GB) GPUs. Detailed hy- perparameters are provided in the supplementary materials §8. For real-world experiments, we adopt Universal Robots UR5 collaborative robotic arm, which has 6 degrees of free- dom. The experiments require the robot to complete three tasks: placing an object, pressing a button, and closing a drawer. Each experimental setup is evaluated over 20 con- secutive trials. 5.2. Benchmark We compare our method against representative VLA sys- tems on the LIBERO suite [28], which groups manipulation tasks into four categories: Spatial, Object, Goal, and Long. We report success rates of each task and the average success rate. We report the success rate for each task and the overall average success rate across the 10 language instructions and 50 episodes under 3 random seeds. Small-scale baselines. Tab. 2 compares our LEEVLA- mini with recent small-scale VLAs whose sizes are less than 1 billion parameters. Octo [42] is an open-source gener- alist policy for robotic manipulation which pretrained on the Open X-Embodiment trajectories. UniACT [59] builds an embodied foundation model in a universal action space. Seer [45] is an end-to-end Predictive Inverse Dynamics Model that jointly performs conditional visual foresight and inverse-dynamics action prediction. DreamVLA [57] in- troduces explicit reasoning by forecasting visual goals be- fore action decoding. FLOWER [39] is a 950M-parameter VLA policy that improves the efficiency of action gener- ation. Our LEEVLA-mini includes an explicit reasoning stage through structured feature flow generation, which is reflected in the Reasoning column. This set isolates the ef- fect of reasoning and token prioritization at similar param- eter budgets. Large-scale baselines. Tab. 2 shows the comparison re- sults between LEEVLA-large and prior large-scale mod- els.OpenVLA [23] is a widely used 7B open-source baseline built on Llama-2-7B [46] with DINOv2 [36] and SigLIP [55] vision features. π 0 [3] adopts a pretrained VLM (PaliGemma) with a flow-matching action expert and ac- tion chunking for continuous control. OpenVLA-OFT [24] instantiates an Optimized Fine-Tuning recipe for Open- VLA [23] with parallel decoding, chunked continuous ac- tions, and an L 1 regressive policy. UniVLA [5] learns cross-embodiment VLA policies by extracting task-centric latent action representations from large-scale, heteroge- neous videos and decoding them into robot-specific actions. MemoryVLA [41] adds a perceptual-cognitive memory to handle long-horizon temporal dependence. We mark Rea- soning according to whether a method introduces an explicit intermediate stage before action output. 5.3. Experimental Results Simulation Environment Results. As shown in Tab. 2, LEEVLA effectively adapts to various task settings of LIBERO, achieving optimal or competitive performance across most task suites. In Fig. 4, we further visualize the correlation between vision features and the instruction by computing the cosine similarity between the visual and instruction embeddings. The top part of Fig. 4 (Baseline) indicates that a model that does not infer future states of the environment fails to leverage visual embeddings to effectively guide action gen- eration. Benefiting from the SFFG and DGDP modules, LEEVLA achieves a much tighter coupling between visual observations and action generation. Real-world Results. We provide a quantitative analysis in real-world settings, as shown in Table 3. Across mul- tiple tasks, our approach consistently outperforms Open- VLA. Additionally, we offer qualitative insights through vi- 6 MethodsReasoningSize CALVIN Results 12345Avg.(%) OpenVLA (CoRL’2025) [23]✗7B91.377.862.052.143.53.80 π 0 (RSS’2025) [3]✗3B94.387.077.968.559.43.87 π 0.5 (RSS’2025) [3]✗3B91.984.679.475.571.04.02 OpenVLA-OFT (RSS’2025) [24]✗7B96.389.182.475.866.54.10 UniVLA (RSS’2025) [5]✗7B95.585.575.466.956.53.80 LEEVLA-large (Ours)✓7B98.894.587.380.672.74.34 Table 1. The success rates of large-scale models on the CALVIN benchmark. We evaluate LEEVLA-large on four CALVIN ABC-D tasks and report the success rate for each task and the average length. ScaleMethodReasoningSize LIBERO Results Spatial(%)Object(%)Goal(%)Long(%)Average(%) Small Octo (RSS’2024) [42]✗0.1B78.985.784.651.175.1 UniACT (CVPR’2025) [59]✗0.5B77.087.077.070.077.8 Seer (ICLR’2025) [45]✗0.3B–87.7– DreamVLA (NeurIPS’2025) [57]✓0.3B97.594.089.589.592.6 FLOWER (CoRL’2025) [39]✗1B97.196.795.693.595.7 LEEVLA-mini (Ours)✓0.5B98.699.097.095.597.5 Large OpenVLA (CoRL’2024) [23]✗7B84.788.479.253.776.5 CoT-VLA (CVPR’2025) [58]✓7B87.591.687.669.081.1 π 0 (RSS’2025) [3]✗3B96.898.895.885.294.1 π 0.5 (CoRL’2025) [3]✗3B97.099.098.096.097.5 OpenVLA-OFT (RSS’2025) [24]✗7B97.698.497.994.597.1 UniVLA (RSS’2025) [5]✗7B96.596.895.692.095.2 LEEVLA-large (Ours)✓7B98.899.098.696.498.2 Table 2. Success rates on the LIBERO benchmark. We evaluate LEEVLA-mini and LEEVLA-large on four LIBERO tasks and report the success rate for each task and the average success rate. Reasoning indicates whether the model performs an explicit reasoning stage before action generation, and Size denotes the parameter scale of the language backbone. Figure 4. Comparison of instruction-relevance between LEEVLA- mini and baseline vision features. We present a visualization of the cosine similarity between vision embeddings and language in- struction embeddings. “Baseline” denotes the miniaturized model variant where both DGDP and SFFG components are ablated. Place the yellow toy duck into the box. Close the drawer. Press the yellow button. Pick up the small block and place the small block on the large block. Figure 5. The qualitative results of LEEVLA in real-world en- vironment. We set up three real-world tasks to demonstrate the generalization performance of our model in real world. sualizations, as illustrated in Fig. 5. 5.4. Ablation study In this section, we conduct a series of ablations on LIBERO using LEEVLA-mini to better understand the contribution of each component in LEEVLA. As shown in Tab. 5, each proposed component brings a consistent improvement 7 MethodsPlacePressDrawerLongAverage OpenVLA [23]3055403540 LEEVLA-large (Ours)7080656078.5 Table 3. Real-world performance comparison on manipulation tasks.We evaluate LEEVLA-large (Ours) and OpenVLA in real-world environments across three representative tasks: Place, Press, Drawer, and their overall average success rate (%). StageComponentPeak Memory (GB)Latency (ms) Train P2P0.2372.01 MC0.6521200.91 DPP0.7269.60 SDG1.05222.68 Inference OpenVLA-OFT15.639124.03 LEEVLA-large15.639124.36 LEEVLA-mini5.08848.61 Table 4. Complexity analysis of different components during training and inference. We report the peak memory usage and la- tency of P2P, MC, DPP, and SDG. over the baseline. The base LEEVLA-mini model without prototype-to-periphery (P2P), mutual-neighborhood con- trastive (MC) loss, dynamic position prioritization (DPP), or semantic-drift guidance (SDG) achieves a success rate of 94.8%. Introducing P2P prediction alone improves per- formance to 95.2% (+0.4), indicating that enforcing an or- dered feature flow is beneficial for policy learning. Adding MC loss further boosts the success rate to 95.6% (+0.8 over baseline), suggesting that preserving local semantic topol- ogy in latent space stabilizes future feature prediction. On top of this structured feature flow generation, enabling DPP yields the largest single gain, reaching 96.3% (+1.5 over baseline), which highlights the importance of concentrating supervision on interaction-centric regions. Finally, incor- porating the SDG leads to the best performance of 96.6%. The experiments demonstrate that SFFG (P2P+MC) and DGDP (DPP+SDG) are complementary, jointly contribut- ing to more accurate and robust action policies. 6. Discussion Why do we encourage models to reason in the structured latent space? Human-selected external conditions capture a narrow and specific concept of the environment, often misaligned with the clues the model actually uses. Pre-trained visual features capture more structural signals. Our SFFG imposes a prototype-to-periphery ordering so tokens from the same semantic unit are predicted together, preserving spatial se- mantic continuity and improving long-horizon prediction. As shown in Fig. 6, LEEVLA considers the structured in- formation between different patches. Why do we need to reorder the visual features? Flat token prediction processes visual tokens in a fixed order, which ignores how features are actually organized in the (a)(b) (c)(d) Figure 6. Wrist-View Image Clustering Visualization Results. We present the actual patches corresponding to different clusters after clustering, where (a), (b), (c), and (d) represent a bowl, a plate, a gripper, and a tabletop, respectively. It can be observed that within the set of encoded image patch features, patches with similar se- mantics are grouped into the same feature cluster. latent space. As a result, the model reasons within a dis- continuous semantic space, compromising generalization. As shown in §4.2, tokens that belong to the same semantic unit can be far apart in the flattened sequence, even though they are close in feature space. This mismatch breaks local contextual continuity and makes it harder for the policy to reason about spatially coherent changes. Tab. 6 shows the effect of feature reordering on future feature prediction. Why do we use the global factor α instead of rely- ing only on prioritization weights β? Intuitively, β am- plifies task-relevant tokens. Without the global factor α and relying solely on β, the model becomes overly selective: contact regions are over-emphasized, while background to- kens are almost discarded. However, background in manip- ulation scenes provides crucial spatial context (e.g., table boundaries, obstacles, robot base) that is important for ge- ometry and long-horizon feasibility. The global factor α ensures that even down-weighted regions retain a weak but non-zero contribution, preserving global layout. As shown in Tab. 7, using both α and β enables the model to focus on task-critical areas without losing overall scene awareness, whereas the β-only variant tends to over-focus and degrades performance. Therefore, we train LEEVLA with both α and β learn to sharply highlight task-critical regions while still maintaining understanding of the whole environment. 8 SFFGDGDP Avg. Suc. (%) Instructions P2PMCDPPSDGI1I2I3I4I5I6I7I8I9I10 ✗94.890.096.098.090.096.098.080.0100.0100.0100.0 ✓✗95.296.0100.092.092.0100.0100.076.098.098.0100.0 ✓✗95.696.094.094.092.0100.0100.084.0100.098.098.0 ✓✗96.398.098.094.787.3100.097.389.3100.0100.098.7 ✓97.0100.0100.098.088.0100.0100.094.0100.098.092.0 Table 5. Results of the ablation study on LIBERO-Goal. We investigate the contribution of each component by progressively ablating the four components: P2P, MC, DPP, and SDG. MethodSuc. (%) SFFG w/o reorder94.7 SFFG w/ reorder95.2 Table 6. Effect of feature reordering on future feature prediction. We evaluate LEEVLA-mini on LIBERO-Goal with and without the feature reordering module, using only the “Baseline+P2P” MethodSuc. (%) LEEVLA-mini w/o α95.8 LEEVLA-mini w/ α97.0 Table 7. Performance comparison between using the global factor α and not using the global factor α. We tested the effect of the global modulation factor α on the action policy. 7. Conclusion We introduce LEEVLA for reasoning in latent feature space. By forecasting structured future features, LEEVLA exploits the relational structure already encoded in the vi- sual backbone and avoids hand-crafted hypothesis spaces. Our structured feature flow generation (SFFG) treats pre- diction as a latent state transition: prototype-to-periphery (P2P) anchors the flow on robust prototypes before refin- ing toward cluster periphery, while mutual-neighborhood contrastive (MC) loss preserves local topology by empha- sizing reciprocal neighbors. Complementing this, drift- guided dynamic prioritization (DGDP) component of dy- namic position prioritization (DPP) and semantic drift guid- ance (SDG) focuses supervision on dynamically active, instruction-relevant patches, reducing the impact of static background. Evaluated at two scales, LEEVLA-mini (0.5B) and LEEVLA-large (7B) achieve the state-of-the-art perfor- mance on the LIBERO and Calvin benchmark. References [1] Randall Balestriero and Yann LeCun. How learning by re- construction produces uninformative features for perception. In Proceedings of the 41st International Conference on Ma- chine Learning. JMLR.org, 2024. 2 [2] Suneel Belkhale and Dorsa Sadigh. Minivla: A better vla with a smaller footprint, 2024. 6 [3] Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael R. Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xi- aoyang Shi, Laura Smith, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky. pi 0 : A vision-language-action flow model for general robot control. In Proceedings of Robotics: Science and Systems (RSS), 2025. 1, 2, 6, 7 [4] Jake Bruce, Stephanie Zhai, Igor Mordatch, et al.Ge- nie: Generative interactive environments. arXiv preprint arXiv:2402.15391, 2024. 3 [5] Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. Univla: Learning to act anywhere with task-centric latent ac- tions. Proceedings of Robotics: Science and Systems (RSS), 2025. 2, 6, 7 [6] Samuel Carton, Surya Kanoria, and Chenhao Tan. What to learn, and how: Toward effective learning from rationales. In Findings of the Association for Computational Linguis- tics: ACL 2022, pages 1075–1088, Dublin, Ireland, 2022. Association for Computational Linguistics. 2 [7] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Ge- offrey Hinton. A simple framework for contrastive learn- ing of visual representations. In Proceedings of the 37th In- ternational Conference on Machine Learning (ICML), pages 1597–1607, 2020. 5 [8] Muhayy Ud Din, Waseem Akram, Lyes Saad Saoud, Jan Rosell, and Irfan Hussain. Vision language action models in robotic manipulation: A systematic review, 2025. 1 [9] Jiahua Dong, Qi Lyu, Baichen Liu, Xudong Wang, Wenqi Liang, Duzhen Zhang, Jiahang Tu, Hongliu Li, Hanbin Zhao, Henghui Ding, Yulun Zhang, Zhi Han, Nicu Sebe, Fahad Shahbaz Khan, Salman Khan, Mubarak Shah, Philip Torr, Ming-Hsuan Yang, and Dacheng Tao. Learning to model the world: A survey of world models in artificial in- telligence. TechRxiv, 2026. 3 [10] Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm- e: an embodied multimodal language model. In Proceedings of the 40th International Conference on Machine Learning, pages 8469–8488, 2023. 2 [11] Danny Driess, Jost Tobias Springenberg, Brian Ichter, Lili Yu, Adrian Li-Bell, Karl Pertsch, Allen Z. Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, and Sergey Levine. Knowledge insulating vision-language-action models: Train fast, run fast, generalize better, 2025. 1 [12] Juntao Gao, Feiyang Ye, Jing Zhang, and Wenjing Qian. Compressor-vla: Instruction-guided visual token compres- sion for efficient robotic manipulation.arXiv preprint arXiv:2511.18950, 2025. 3 [13] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap.Mastering diverse control tasks through world models. Nature, pages 1–7, 2025. 3 9 [14] Jian Han, Jinlai Liu, Yi Jiang, Bin Yan, Yuqi Zhang, Zehuan Yuan, Bingyue Peng, and Xiaobing Liu. Infinity: Scaling bit- wise autoregressive modeling for high-resolution image syn- thesis, 2024. 5 [15] Zebin Han, Xudong Wang, Baichen Liu, Qi Lyu, Zhenduo Shang, Jiahua Dong, Lianqing Liu, and Zhi Han. Seqwalker: sequential-horizon vision-and-language navigation with hi- erarchical planning. In Proceedings of the Fortieth AAAI Conference on Artificial Intelligence and Thirty-Eighth Con- ference on Innovative Applications of Artificial Intelligence and Sixteenth Symposium on Educational Advances in Arti- ficial Intelligence. AAAI Press, 2026. 1 [16] Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal. Leakage-adjusted simulatability: Can models generate non- trivial explanations of their behavior in natural language? In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4351–4367, Online, 2020. Association for Computational Linguistics. 2 [17] Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gian- luca Corrado. Gaia-1: A generative world model for au- tonomous driving. arXiv preprint arXiv:2309.17080, 2023. 3 [18] Yucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, and Jianyu Chen. Video prediction policy: A generalist robot policy with predictive visual representations, 2025. 1 [19] Huang Huang, Fangchen Liu, Letian Fu, Tingfan Wu, Mustafa Mukadam, Jitendra Malik, Ken Goldberg, and Pieter Abbeel.Otter: A vision-language-action model with text-aware visual feature extraction.arXiv preprint arXiv:2503.03734, 2025. 3 [20] Stephen James, Zicong Ma, David Rovick Arrojo, and An- drew J. Davison. Rlbench: The robot learning benchmark & learning environment. IEEE Robotics and Automation Let- ters, 2020. 2 [21] Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandku- mar, Yuke Zhu, and Linxi Fan. Vima: robot manipulation with multimodal prompts. In Proceedings of the 40th Inter- national Conference on Machine Learning. JMLR.org, 2023. 2 [22] Yitong Jiang, Jinwei Gu, Tianfan Xue, Ka Chun Cheung, Pavlo Molchanov, Hongxu Yin, and Sifei Liu.Token- efficient vlm: High-resolution image understanding via dy- namic region proposal. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV), 2025. 3 [23] Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Fos- ter, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kol- lar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn.Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024. 1, 2, 3, 4, 6, 7, 8 [24] Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-tuning vision-language-action models: Optimizing speed and suc- cess. arXiv preprint arXiv:2502.19645, 2025. 2, 3, 4, 5, 6, 7 [25] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ́ ar, and Ross Girshick. Segment anything. arXiv:2304.02643, 2023. 2 [26] Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizhong Zhang, Xiaofan Wang, Bei Liu, Jianlong Fu, Jian- min Bao, Dong Chen, Yuanchun Shi, Jiaolong Yang, and Baining Guo.Cogact: A foundational vision-language- action model for synergizing cognition and action in robotic manipulation, 2024. 2 [27] Yingyan Li, Lue Fan, Jiawei He, Yu-Quan Wang, Yuntao Chen, Zhaoxiang Zhang, and Tieniu Tan. Enhancing end- to-end autonomous driving with latent world model. The Thirteenth International Conference on Learning Represen- tations (ICLR), abs/2406.08481, 2025. 2, 3 [28] Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowl- edge transfer for lifelong robot learning, 2023. 1, 2, 6 [29] Dongxiu Liu, Haoyi Niu, Zhihao Wang, Jinliang Zheng, Yi- nan Zheng, Zhonghong Ou, Jianming Hu, Jianxiong Li, and Xianyuan Zhan. Efficient robotic policy learning via latent space backward planning. In Proceedings of the 42nd In- ternational Conference on Machine Learning (ICML), 2025. 1 [30] Ilya Loshchilov and Frank Hutter. Decoupled weight de- cay regularization. In International Conference on Learning Representations (ICLR), 2017. 6 [31] Qi Lyu, Jiahua Dong, Baichen Liu, Xudong Wang, Mingfei Han, Yulun Zhang, Fahad Shahbaz Khan, Salman Khan, Lianqing Liu, and Zhi Han. Sab-lvlm: Significance-aware binarization for large vision-language models, 2026. 1 [32] Oier Mees, Lukas Hermann, Erick Rosete-Beas, and Wol- fram Burgard.Calvin:A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks. IEEE Robotics and Automation Letters, 7(3): 7327–7334, 2022. 2 [33] Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Repre- sentation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018. 5 [34] Open X-Embodiment Collaboration. Open x-embodiment: Robotic learning datasets and rt-x models. In Proceedings of the IEEE International Conference on Robotics and Au- tomation (ICRA), 2024. 1, 6 [35] OpenAI. Introducing chatgpt. https://openai.com/ index/chatgpt/, 2022. Accessed: 2025-10-30. 1 [36] Maxime Oquab, Timoth ́ e Darcet, Theo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Rus- sell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang- Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nico- las Ballas, Gabriel Synnaeve, Ishan Misra, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bo- janowski. Dinov2: Learning robust visual features without supervision, 2023. 2, 3, 6 10 [37] Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. Fast: Efficient action tokenization for vision- language-action models. arXiv preprint arXiv:2501.09747, 2025. 2 [38] Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R ̈ adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao- Yuan Wu, Ross Girshick, Piotr Doll ́ ar, and Christoph Feicht- enhofer. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024. 2 [39] Moritz Reuss, Hongyi Zhou, Marcel R ̈ uhle, ̈ Omer Erdinc ̧ Ya ̆ gmurlu, Fabian Otto, and Rudolf Lioutikov. Flower: De- mocratizing generalist robot policies with efficient vision- language-action flow policies. In Proceedings of the 9th Conference on Robot Learning (CoRL), pages 3736–3761. PMLR, 2025. 6, 7 [40] Nur Muhammad (Mahi) Shafiullah, Chris Paxton, Ler- rel Pinto, Soumith Chintala, and Arthur Szlam.Clip- fields: Weakly supervised semantic fields for robotic mem- ory. Proceedings of Robotics: Science and Systems (RSS), abs/2210.05663, 2022. 2 [41] Hao Shi, Bin Xie, Yingfei Liu, Lin Sun, Fengrong Liu, Tiancai Wang, Erjin Zhou, Haoqiang Fan, Xiangyu Zhang, and Gao Huang. Memoryvla: Perceptual-cognitive memory in vision-language-action models for robotic manipulation, 2025. 6 [42] Octo Model Team. Octo: An open-source generalist robot policy. In Proceedings of Robotics: Science and Systems (RSS), Delft, Netherlands, 2024. 1, 2, 6, 7 [43] Qwen Team.Qwen2 technical report.arXiv preprint arXiv:2407.10671, 2024.Large language and vision- language models from Alibaba Group. 1 [44] Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang. Visual autoregressive modeling: scalable image gen- eration via next-scale prediction. In Proceedings of the 38th International Conference on Neural Information Processing Systems, Red Hook, NY, USA, 2024. Curran Associates Inc. 5 [45] Yang Tian, Sizhe Yang, Jia Zeng, Ping Wang, Dahua Lin, Hao Dong, and Jiangmiao Pang. Predictive inverse dynamics models are scalable learners for robotic manipulation. arXiv preprint arXiv:2412.15109, 2024. 6, 7 [46] Hugo Touvron, Louis Martin, Kevin Stone, Peter Al- bert, Amjad Almahairi, Yasmine Babaei, Nikolay Bash- lykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fer- nandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Vik- tor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Ko- renev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizen- stein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiao- qing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine- tuned chat models, 2023. 1, 6 [47] Jiaxuan Wang, Sarah Jabbour, Maggie Makar, Michael Sjod- ing, and Jenna Wiens. Learning concept credible models for mitigating shortcuts. In Advances in Neural Information Processing Systems, pages 33343–33356. Curran Associates, Inc., 2022. 1, 2 [48] Yuqi Wang, Xinghang Li, Wenxuan Wang, Junbo Zhang, Yingyan Li, Yuntao Chen, Xinlong Wang, and Zhaoxiang Zhang. Unified vision-language-action model, 2025. 2 [49] Yi Wang, Jiaze Wang, Ziyu Guo, Renrui Zhang, Donghao Zhou, Guangyong Chen, Anfeng Liu, and Pheng-Ann Heng. What we miss matters: Learning from the overlooked in point cloud transformers. In Advances in Neural Informa- tion Processing Systems (NeurIPS), 2025. 2 [50] Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, and Mingsheng Long. ivideogpt: Interactive videogpts are scalable world models. Advances in Neural Information Processing Systems, 37:68082–68119, 2024. 3 [51] Yingchen Xu, Jack Parker-Holder, Aldo Pacchiano, Philip J. Ball, Oleh Rybkin, Stephen J. Roberts, Tim Rockt ̈ aschel, and Edward Grefenstette. Learning general world models in a handful of reward-free deployments. In Advances in Neural Information Processing Systems (NeurIPS), Red Hook, NY, USA, 2022. Curran Associates Inc. 3 [52] Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2 [53] Shuai Yang, Hao Li, Yilun Chen, Bin Wang, Yang Tian, Tai Wang, Hanqing Wang, Feng Zhao, Yiyi Liao, and Jiangmiao Pang. Instructvla: Vision-language-action instruction tuning from understanding to manipulation, 2025. 1 [54] Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, et al. World action models are zero-shot policies.arXiv preprint arXiv:2602.15922, 2026. 3 [55] Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 2, 3, 6 [56] Shiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu, Yuan Li, Qinghui Gao, Zhaoye Fei, Zhangyue Yin, Zuxuan Wu, Yu- Gang Jiang, et al.Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long- horizon reasoning tasks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11142– 11152, 2025. 2 [57] Wenyao Zhang, Hongsi Liu, Zekun Qi, Yunan Wang, Xin- qiang Yu, Jiazhao Zhang, Runpei Dong, Jiawei He, He 11 Wang, Zhizheng Zhang, Li Yi, Wenjun Zeng, and Xin Jin. Dreamvla: A vision-language-action model dreamed with comprehensive world knowledge. CoRR, abs/2507.04447, 2025. 2, 6, 7 [58] Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Qianli Ma, Song Han, Chelsea Finn, Ankur Handa, Tsung-Yi Lin, Gordon Wet- zstein, Ming-Yu Liu, and Donglai Xiang. Cot-vla: Visual chain-of-thought reasoning for vision-language-action mod- els. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1702– 1713, 2025. 1, 2, 7 [59] Jinliang Zheng, Jianxiong Li, Dongxiu Liu, Yinan Zheng, Zhihao Wang, Zhonghong Ou, Yu Liu, Jingjing Liu, Ya-Qin Zhang, and Xianyuan Zhan. Universal actions for enhanced embodied foundation models. In 2025 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR), pages 22508–22519, 2025. 2, 6, 7 [60] Ruijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao, Hal Daum ́ e I, Andrey Kolobov, Furong Huang, and Jianwei Yang. Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies, 2025. 2 [61] Siyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li, Dit-Yan Yeung, and Chuang Gan. Robodreamer: Learning compo- sitional world models for robot imagination. arXiv preprint arXiv:2404.12377, 2024. 3 [62] Chuning Zhu, Raymond Yu, Siyuan Feng, Benjamin Burch- fiel, Paarth Shah, and Abhishek Gupta. Unified world mod- els: Coupling video and action diffusion for pretraining on large robotic datasets. ArXiv, abs/2504.02792, 2025. 3 [63] Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In Conference on Robot Learning, pages 2165–2183. PMLR, 2023. 1, 2 12