Paper deep dive
Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning
Taozhao Chen, Ian Manchester, Huaming Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 7/5/2026, 1:56:22 AM
Summary
This position paper argues that current Vision-Language-Action (VLA) systems, which rely on pretrained Vision-Language Models (VLMs), cannot be verified to perform genuine physical reasoning. The authors identify a 'semantic sufficiency assumption'âthe unverified belief that semantic representations from internet-scale data are sufficient for physical action decisions. They decompose VLA policies into semantic mapping and physical action decision, noting that physical properties (mass, friction, etc.) are not recoverable from static image-text data used in VLM pretraining. The paper identifies three levels of non-identifiability in current evaluation protocols: Attribution, Source, and Representation-level. Consequently, improvements in task success rates cannot distinguish between semantic matching, distributional overlap, and genuine physical generalization. The authors propose a research direction involving evaluation designs with controlled physical variation to enable causal attribution of performance.
Entities (9)
Relation Signals (4)
Vision-Language-Action (VLA) systems â builton â Vision-Language Models (VLMs)
confidence 100% · Vision-Language-Action (VLA) systems, built on pretrained vision-language models (VLMs)...
RT-2 â isatypeof â Vision-Language-Action (VLA) systems
confidence 100% · Since RT-2 demonstrated that pretrained VLMs can be repurposed as robot policies...
Semantic Generalization â isdistinctfrom â Physical Generalization
confidence 95% · These correspond to two distinct components: semantic generalization... and physical execution.
VLA research â suffersfrom â Narrative Drift
confidence 90% · We further argue that this identifiability gap has been reinforced through narrative drift...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision-Language-Action (VLA) systems, built on pretrained vision-language models (VLMs), have shown rapidly improving performance on robot manipulation benchmarks. These gains are commonly interpreted as evidence that semantic representations learned from internet-scale data transfer to physical execution generalization. This position paper argues that the assumption underlying this interpretation -- that semantic generalization is sufficient to support physical action decisions -- has not been independently verified and cannot be tested under current evaluation protocols. We support this claim by decomposing VLA policies into semantic mapping and physical action decision, and showing that task success rate -- the dominant evaluation metric -- cannot distinguish between these two sources of capability. As a result, improvements in benchmark performance are consistent with multiple competing explanations, including semantic matching, distributional overlap, and genuine physical generalization. We further argue that this identifiability gap has been reinforced through narrative drift, whereby successive systems inherit and strengthen prior interpretations of performance gains without isolating the underlying causal mechanism. To address this limitation, we propose a research direction based on evaluation designs that introduce controlled variation to separately measure semantic and physical generalization. Such designs make it possible to causally attribute performance without requiring access to model internals, and to empirically assess the role of VLM backbones as semantic interfaces rather than implicit sources of physical competence. Our goal is not to refute the role of VLMs in robotics, but to clarify the conditions under which claims of physical generalization can be meaningfully evaluated.
Tags
Links
- Source: https://arxiv.org/abs/2606.30686v1
- Canonical: https://arxiv.org/abs/2606.30686v1
Trouble viewing inline? Open PDF directly â
Full Text
56,834 characters extracted from source content.
Expand or collapse full text
Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning Taozhao Chen School of Electrical and Computer Engineering The University of Sydney Sydney, NSW, Australia tche8294@uni.sydney.edu.au &Ian Manchester School of Aerospace, Mechanical and Mechatronic Engineering The University of Sydney Sydney, NSW, Australia ian.manchester@sydney.edu.au &Huaming Chen School of Electrical and Computer Engineering The University of Sydney Sydney, NSW, Australia huaming.chen@sydney.edu.au Abstract Vision-Language-Action (VLA) systems, built on pretrained vision-language models (VLMs), have shown rapidly improving performance on robot manipulation benchmarks. These gains are commonly interpreted as evidence that semantic representations learned from internet-scale data transfer to physical execution generalization. This position paper argues that the assumption underlying this interpretationâthat semantic generalization is sufficient to support physical action decisionsâhas not been independently verified and cannot be tested under current evaluation protocols. We support this claim by decomposing VLA policies into semantic mapping and physical action decision, and showing that task success rateâthe dominant evaluation metricâcannot distinguish between these two sources of capability. As a result, improvements in benchmark performance are consistent with multiple competing explanations, including semantic matching, distributional overlap, and genuine physical generalization. We further argue that this identifiability gap has been reinforced through narrative drift, whereby successive systems inherit and strengthen prior interpretations of performance gains without isolating the underlying causal mechanism. To address this limitation, we propose a research direction based on evaluation designs that introduce controlled variation to separately measure semantic and physical generalization. Such designs make it possible to causally attribute performance without requiring access to model internals, and to empirically assess the role of VLM backbones as semantic interfaces rather than implicit sources of physical competence. Our goal is not to refute the role of VLMs in robotics, but to clarify the conditions under which claims of physical generalization can be meaningfully evaluated. 1 Introduction Teaching robots to perform manipulation tasks has long required extensive task-specific training data and engineering effort, with each new task demanding carefully curated demonstrations Argall et al. (2009); Ravichandar et al. (2020); Zare et al. (2024) or hand-engineered reward functions Kalashnikov et al. (2018); Sutton et al. (1998). The emergence of large vision-language models (VLMs)âpretrained on internet-scale visual and linguistic dataâoffers a compelling alternative: if such models encode broad semantic knowledge about objects, instructions, and goals, they may provide a generalizable foundation for robot control Firoozi et al. (2025); Hu et al. (2023); Kawaharazuka et al. (2025). Since RT-2 demonstrated that pretrained VLMs can be repurposed as robot policies Zitkovich et al. (2023), the field has rapidly converged on the Vision-Language-Action (VLA) paradigmâend-to-end architectures that map visual observations and language instructions directly to action commands. As illustrated in Figure 1, subsequent systems largely follow this formulation while exploring variations in output parameterization and control interfaces, forming several architectural lineages, including discrete token policies, diffusion-based policies, and dual-system designs. Despite these differences, they share a common structural core: a pretrained VLM backbone that mediates perception, language grounding, and action generation, and an implicit assumptionâintroduced by RT-2 and inherited by subsequent workâthat semantic representations learned from large-scale vision-language pretraining are sufficient to guide physical action decisions. Under this premise, improvements in semantic generalization are expected to translate into more robust physical execution, motivating extensions along multiple axes, including diffusion-based action modeling, increased model scale, and additional modalities such as proprioception. This trajectory is reflected in systems such as OpenVLA Kim et al. (2024), OpenVLA-OFT Kim et al. (2025), Ï0 _0 Black et al. (2024), and GR00T N1 Bjorck et al. (2025), which report steady benchmark improvements while maintaining this shared architectural and conceptual core. However, task success measures outcomes, not the mechanism producing them. A central unresolved question is whether semantic generalization learned by VLMs transfers to physical execution generalization in robotic systems. These correspond to two distinct components: semantic generalization, which maps observations and instructions to task-relevant abstractions, and physical execution, which selects actions under constraints of dynamics, contact, embodiment, and temporal interaction. We argue that current VLA research relies on an implicit but unverified assumption: that semantic generalization is sufficient to support physical execution. This assumption has been reinforced through what we term narrative drift, whereby successive systems inherit and extend prior interpretations of performance gains without explicit causal verification. Crucially, existing evaluation protocols do not provide the signals necessary to distinguish semantic generalization from physical execution, and therefore do not support causal attribution of performance improvements. Figure 1: VLM-backbone VLA systems (2023â2026) organized into three architectural lineages: discrete token output, diffusion head, and dual-system. Arrows indicate architectural similarity only, not dependency or inheritance relationships. We develop this argument as follows. Section 2 formalizes the implicit assumption underlying VLM-backbone VLA research. Section 3 shows that this assumption is not identifiable under current evaluation protocols. Section 4 analyzes the resulting systemic consequences. Section 5 proposes directions for evaluation reform. Section 6 concludes. 2 The Implicit Assumption Underlying VLM-Backbone VLA Research We formalize the distinction introduced in Section 1 by decomposing a VLA policyâmapping observations and instructions to actionsâinto two components. Semantic mapping transforms visual observations and language instructions into task-relevant representations (e.g., object categories, spatial relations, and intended goals), whereas physical action decision selects actions conditioned on this representation under constraints of dynamics, contact interactions, embodiment, and temporal execution. A capability is physical if, holding the semantic interpretation fixed, variations in physical context (e.g., object mass, friction, geometry, or embodiment) require different actions for successful task completion. This distinction induces a criterion for attribution: improvements in task success can only be attributed to physical generalization if they remain robust under controlled variation of physical conditions that do not alter task semantics. It also implies distinct failure modes: semantic errors lead to incorrect task interpretation, while physical failures arise when the task is correctly identified but execution fails due to unmodeled physical factorsâfailure modes that are not distinguishable from task success rate alone. We formalize this decomposition as: atâŒÏ(at|zsem,zphys)a_t Ï\! (a_t\; |\;z^sem,\,z^phys ) (1) where zsem=fvlmâ(ot,l)z^sem=f_vlm(o_t,l) denotes the semantic representation extracted by the VLM backbone from current observation oto_t and language instruction l, and zphysz^phys denotes physical contextâproperties such as mass, friction, and contact geometry that are not recoverable from static image-text data. The semantic sufficiency assumption implicit in current VLA architectures is: Ïâ(atâŁzsem,zphys)âÏâ(atâŁzsem)Ï\! (a_t z^sem,\,z^phys )âÏ\! (a_t z^sem ) (2) We decompose zsemz^sem into three constituent representations. Categorical representations zcatz^cat encode object identity, category membership, and visual appearance; they are acquired through large-scale imageâtext contrastive training, in which visual and linguistic signals are aligned via co-occurrence supervision Radford et al. (2021); Li et al. (2023). Relational representations zrelz^rel encode spatial configurations, affordance descriptions, and scene-level structure; they are learned through grounding and captioning objectives that map linguistic relational expressions to visual scene layouts Li et al. (2023); Zeng et al. (2022). Goal representations zgoalz^goal encode task intent, sub-goal decomposition, and success criteria; they are derived from instruction-following and, in recent systems, chain-of-thought supervision over language and imageâtext pairs Zitkovich et al. (2023); Kim et al. (2024). Each component is acquired from static image-text data: visual and linguistic signals are jointly available, but action consequences are absent from the supervision signal throughout Radford et al. (2021); Li et al. (2023). This shared supervision structure defines the information boundary of zsemz^sem. Physical properties such as mass distribution, surface friction, and contact geometry are not recoverable from static visual observation: two objects may be visually indistinguishable while differing substantially in physical behavior Battaglia et al. (2013); Lerer et al. (2016). These properties are only revealed through interactionâthrough the consequences of applied force, the resistance of contact, and the dynamics of motion Agrawal et al. (2016); Pinto and Gupta (2016). Because the training signal for fvlmf_vlm contains no such interaction outcomes, zsemz^sem cannot, in principle, encode zphysz^phys regardless of model scale or data volume Gibson (2014); Lake et al. (2017). The semantic sufficiency assumption in Equation 2 therefore does not follow from the properties of VLM pretraining: it asserts that zsemz^sem is sufficient for physical action decisions, but the supervision structure of fvlmf_vlm excludes precisely the signal that would be required for this sufficiency to hold. Unlike zsemz^sem, which is explicitly computed via VLM encoding, current architectures provide no corresponding mechanism for acquiring zphysz^phys, making 2 both structurally unjustified and empirically untested. Despite substantial variation in output representationsâdiscrete token policies trading action fidelity for semantic expressiveness Zawalski et al. (2024); Duan et al. (2025); Belkhale et al. (2024); Li et al. (2024); Arai et al. (2025); Ding et al. (2024), diffusion-based policies improving action continuity and precision Zhou et al. (2025b); Zhu et al. (2025); Bu et al. (2025); Li et al. (2026a, 2025a); Deng et al. (2025); Lin et al. (2025), and dual-system architectures decoupling reasoning latency from execution speed Chen et al. (2025); Song et al. (2025); Li et al. (2025b)âthe architectural core remains unchanged: all systems rely on a pretrained VLM backbone, with modifications concentrated in output heads and input modalities Kim et al. (2024, 2025); Black et al. (2024). These design choices address execution limitations without isolating whether semantic representations are sufficient to support physical action decisions. Nevertheless, capability claims have escalated alongside this unchanged core. Early work positioned VLMs as transferring web-scale semantic knowledge to robot control Zitkovich et al. (2023). Subsequent systems extended this to cross-task and cross-embodiment generalization Kim et al. (2024, 2025), while later work attributed semantic reasoning and problem-solving capabilities to VLM backbones in support of complex manipulation Black et al. (2024). More recent systems further suggest that scaling VLM backbones improves spatial reasoning and physical adaptation Bjorck et al. (2025). Across these claims, a component that has not fundamentally changed has been credited with progressively stronger capabilities. Underlying this trajectory is a common but unverified condition: that improvements in semantic representations learned through VLM pretraining translate into robustness of physical action decisions. This condition is not necessarily stated explicitly in the cited works, but is required for their capability claims to hold under a causal interpretation. We refer to its progressive reinforcement across systems as narrative drift. Table 1: Narrative drift in VLM-backbone VLA research: the architectural core and verification methodology remain unchanged while capability claims escalate continuously. These systems are selected as representative high-impact examples rather than an exhaustive survey. System Primary Architectural Changes Capability Claim Assumption required for this claim to hold RT-2 Zitkovich et al. (2023) Fine-tune PaLI-X/PaLM-E to output action tokens Web-scale knowledge transfers to robot control Semantic knowledge is sufficient to guide physical action selection OpenVLA Kim et al. (2024) Open-source VLM backbone; adjusted tokenization Robust and generalizable visuomotor policies Semantic generalization entails visuomotor generalization OpenVLA-OFT Kim et al. (2025) Parallel decoding; multi-resolution input Semantic generalization supports physical execution Semantic competence transfers to execution via fine-tuning Ï0 _0 Black et al. (2024) Flow matching action expert Inherits reasoning and problem-solving ability Semantic reasoning transfers to manipulation competence GR00T N1 Bjorck et al. (2025) Dual-system architecture with proprioception Stronger VLM improves adaptation in physical tasks VLM strength reflects physical world modeling capability This pattern extends beyond the systems in Table 1 Wen et al. (2024); Zawalski et al. (2024); Arai et al. (2025); Li et al. (2026b). Across a broader range of VLM-backbone VLA works, improvements in task success rateâmeasured under largely unchanged evaluation protocolsâare interpreted as evidence of increasingly strong physical execution generalization Song et al. (2026); Zhong et al. (2026); Wang (2026); Wang et al. (2026b); Zhang et al. (2026). However, as introduced in Section 1, task success rate does not distinguish between semantic and physical contributions. Consequently, the same observed improvements are consistent with multiple explanations, including genuine physical generalization, improved semantic matching, or increased overlap between training and test distributions. The consequence is not that the assumption is untestable in principle, but that it is not independently identifiable within the current methodological framework. Existing evidence does not permit attribution of performance gains to semantic or physical capabilities separately, and therefore cannot establish whether semantic representations are sufficient to support physical action decisions. 3 The Structural Argument 3.1 Semantic and Physical Representations Are Structurally Distinct The assumption identified in Section 2âthat semantic representations are sufficient to support physical action decisionsâdepends on whether the two representation types impose equivalent learning requirements. We argue that they do not. Semantic representations encode what objects are and what instructions mean, including object identity, category membership, and relational descriptions grounded in language and visual appearance Radford et al. (2021); Li et al. (2023); Zeng et al. (2022). Such representations can be acquired from large-scale image-text co-occurrence, where visual and linguistic signals are jointly available. Physical representations encode how objects behave under interaction, including properties such as mass distribution, friction, contact geometry, and dynamic response to force Billard and Kragic (2019). These properties are not recoverable from static image-text data, as the consequences of action are absent from the training signal Battaglia et al. (2013); Gibson (2014). Increasing data scale does not resolve this limitation: supervision over physical state transitions is not present in the modality. The two representation types therefore differ in their supervision requirements. Semantic representations can be learned from passive observation, whereas physical representations require interaction-based feedback over action outcomes Pinto and Gupta (2016); Agrawal et al. (2016). A model trained exclusively on the former cannot be assumed to have acquired the latter. This distinction implies that semantic generalization and physical execution generalization can dissociate. A system may correctly interpret the task while failing to execute it under novel physical conditions Zitkovich et al. (2023); Kim et al. (2025). The assumption that semantic representations are sufficient for physical action decisions therefore does not follow from the properties of the training data or learning objective. 3.2 Three Levels of Non-Identifiability in Current Evaluation Current benchmarks fail to support causal attribution at three distinct levels. These levels are not independent: each deeper level is revealed only when the shallower one is examined, and together they indicate that the evaluation gap is more fundamental than it first appears. Level 1: Attribution Non-Identifiability. Current robot manipulation benchmarks evaluate performance primarily through task success rate Kim et al. (2024); Liu et al. (2025); Wang et al. (2026a); Bjorck et al. (2025). This metric does not distinguish between semantic and physical sources of success or failure. Successful task completion may result from correct semantic grounding, physically robust execution, or overlap between training and test distributions. Failures may arise from incorrect task interpretation or from inability to execute under novel physical conditions. Staged evaluation protocols improve localization by decomposing tasks into phases Kim et al. (2025); Nasiriany et al. (2024), but do not resolve this ambiguity. A failure at a specific stage remains consistent with both semantic misinterpretation and physical execution error. These protocols identify where failure occurs, but not why. A critical limitation is that current benchmarks do not provide controlled interventions that invalidate semantic matching while preserving task semantics. As a result, policies that rely on distributional alignment between instructionâobservation pairs and action patterns cannot be distinguished from those that model underlying physical dynamics. Attribution of performance to semantic matching or physical generalization is therefore fundamentally underdetermined. However, even if attribution could be resolved, a deeper question remains: where does the observed generalization come from? Level 2: Source Non-Identifiability. Even if attribution were resolved, the source of generalization would remain unidentifiable. Performance gains may arise from at least three sourcesâsemantic priors from VLM pretraining, physical understanding from robot interaction data, or their interactionâyet current evaluation protocols control for none of them. A concrete manifestation appears in the definition of out-of-distribution (OOD). OOD status is typically defined relative to robot demonstration data, not VLM pretraining data Zitkovich et al. (2023); Kim et al. (2024); Black et al. (2024). Objects novel to the robot dataset may be familiar from pretraining, introducing an uncontrolled confound. Combined with task-specific fine-tuning, this creates a mismatch between nominal and actual sources of generalization. As a result, observed improvements remain ambiguous in origin. Gains attributed to physical generalization may instead reflect prior semantic exposure or distributional overlap. Because current evaluation protocols neither control pretraining exposure nor isolate interaction-driven learning, they do not permit disentangling these sources. This leads to source non-identifiability: the origin of generalization cannot be determined from observed outcomes. Level 3: Representation-Level Non-Identifiability. Current benchmarks implicitly assume that all instructions are valid and executable within the agentâs physical context. Under this assumption, evaluation reduces to conditional action generation given a feasible task, and does not test whether the task itself should or can be executed. This obscures the fate of capabilities inherited from pretrained vision-language backbones. Prior work shows that such models exhibit forms of constraint-aware reasoning and feasibility judgment in embodied settingsâincluding scoring action affordances Ahn et al. (2022b), replanning after failures Huang et al. (2022), and recognizing infeasible instructions Zhang et al. (2024). However, it remains unclear whether these capabilities are preserved, degraded, or ignored during fine-tuning into action policies. This concern is supported by evidence that fine-tuning can overwrite pretrained representations Zhou et al. (2025a); Yadav et al. (2025); Fei et al. (2025); Huang et al. (2026). Because current benchmarks exclude infeasible, unsafe, or contradictory instructions, they provide no signal to detect the presence or absence of such capabilities. As a result, representation-level changes during policy learning remain unobservable. This leads to representation-level non-identifiability: the effect of training on pre-existing capabilities cannot be measured, and potential degradation remains undetected. Taken together, these three levels show that non-identifiability is not merely a measurement limitation but a structural property of the current evaluation paradigm. Attribution cannot be resolved at the level of outcomes, the source of generalization cannot be isolated across training signals, and changes in underlying representations cannot be observed during learning. As a result, multiple latent factorsâincluding semantic grounding, physical reasoning, and representation-level capabilitiesâare collapsed into a single observable metric. This prevents causal attribution and renders the central assumption of this paperâthat semantic generalization is sufficient for physical executionâempirically untestable under current evaluation protocols. 3.3 Empirical Signatures of Underdetermined Attribution The three levels of non-identifiability identified above are not merely theoretical. If evaluation systematically cannot distinguish attribution, source, and representation-level effects, the literature should exhibit a predictable empirical signature: physical execution failures are observed but attributed to engineering factors, while the alternative explanationâfailure of the semantic sufficiency assumptionâis not evaluated. Table 2 illustrates this pattern across representative systems. In each case, observed limitations are consistent with failure of physical execution under novel conditions, yet are attributed to data coverage, model scale, or training procedure Zitkovich et al. (2023); Kim et al. (2024, 2025); Black et al. (2024). These attributions are plausible given the available evidence, but not uniquely supported by it. Table 2: Physical execution limitations across representative VLM-backbone VLA systems, and their attributed causes. In each case, an engineering attribution was adopted; the alternativeâthat the semantic sufficiency assumption has failedâwas not considered. System Observed limitation Authorâs attribution RT-2 Zitkovich et al. (2023) Model correctly identifies and approaches target objects but fails to control physical dynamics: pen rolling, banana center of mass displacement [App. G]. Manually identified through qualitative case analysis outside the primary evaluation protocol Physical skills are bounded by the robot training data distribution; more robot demonstration data needed RT-2 Zitkovich et al. (2023) Larger VLM backbone parameters do not produce higher task success rates across evaluation conditions Model scale is not the bottleneck; data quality matters more than parameter count OpenVLA-OFT Kim et al. (2025) Policy relies on spurious visual correlations in multi-camera deployment; language grounding degrades despite unchanged task semantics [§IV.C]. Distractor objects cause substantial success rate collapse while target localization remains intact Spurious correlations in training data; insufficient data coverage Ï0 _0 Black et al. (2024) Task completion is unreliable across conditions; recovery from errors is fragile [§VII]. Success rate varies substantially across tasks with no uniform failure pattern Data recipe problem; higher quality or more diverse data needed This pattern admits two explanations. The first is that independent research groups consistently overlook the same alternative interpretation. The second is structural: if evaluation does not provide signals to distinguish semantic and physical failure, alternative explanations cannot be formulated or tested. The observed consistency of engineering attribution across systems supports the latter. Under this interpretation, attribution in the literature reflects the limits of the evaluation framework rather than definitive evidence about underlying causes. 4 Systemic Consequences The semantic sufficiency assumption, when combined with an evaluation framework that does not provide identifiable signals for detecting its failure, produces three systemic consequences that compound over time. Misallocation of Research Resources. When evaluation cannot attribute performance gains to their source, research investment decisions are made without reliable diagnostic signals Sculley et al. (2015); Kiela et al. (2021). The field cannot determine whether observed improvements reflect advances in semantic generalization, physical execution generalization, or distributional overlapâyet resource allocation depends precisely on this distinction DâAmour et al. (2022); Geirhos et al. (2020). Investment in larger VLM backbones is justified if semantic capability transfers to physical execution; it is unjustified if the transfer does not occur and gains instead reflect distributional overlap. Likewise, investment in data collection depends on whether the bottleneck lies in semantic coverage or physical interaction diversityâtwo distinct requirements that existing benchmarks do not separately measure Kalashnikov et al. (2018); Dasari et al. (2019); Nasiriany et al. (2024); Mees et al. (2022). This issue connects to a broader principle in the study of scaling. It is often argued that increasing scale alone is sufficient to achieve higher capability, with different methods primarily affecting efficiency rather than qualitative behavior Hestness et al. (2017); Kaplan et al. (2020); Huh et al. (2024). Under this view, evaluation plays a critical role: it determines whether scaling is applied to the correct underlying mechanism. If evaluation cannot distinguish between competing explanations of performance, scaling may amplify improvements along directions that do not correspond to the intended capability. False Confidence in Deployment Readiness. Benchmark performance is the primary basis on which VLA systems are assessed for deployment readiness Zitkovich et al. (2023); Liu et al. (2023). If benchmark success rates reflect distributional overlap rather than physical execution generalization, this confidence may not be warranted by the available evidence. OpenVLA-OFT reports near-perfect success rates on constrained tabletop tasks while the same system fails substantially when task-irrelevant distractor objects are introduced Kim et al. (2025): a perturbation that does not alter task semantics produces significant performance degradation. This gap between benchmark performance and robustness to minimal perturbations is not reliably captured by current evaluation frameworks. A system can accumulate a strong benchmark record while its physical execution generalization remains unverified. In deployment contexts where physical configurations cannot be controlled, failure modes that benchmarks cannot detect become consequential: systems that succeed under training distribution overlap may fail when that overlap disappears, without prior evaluation signals that predict or characterize such failures. Persistent Misdiagnosis of Architectural Failures. The attribution pattern identified in Section 3.3âphysical execution limitations consistently attributed to engineering factorsâis not merely a historical observation. It reflects a structural limitation of a field that lacks reliable tools to distinguish architectural causes from data or optimization effects DâAmour et al. (2022). When failures are attributed to insufficient data coverage, the proposed solution is more data. When failures are attributed to model capacity, the solution is larger models. When failures are attributed to action parameterization, the solution is alternative output representations. Each of these responses is reasonable under an engineering diagnosis. However, if the underlying cause is an unverified architectural assumption, these responses may leave the root issue unaddressed while continuing to consume resources. The cost of persistent misdiagnosis is not limited to inefficiency. It produces a compounding divergence between the problems the field believes it is solving and the problems that must be solved for robust physical deployment. Systems may exhibit consistent improvements in benchmark performance while relying on mechanisms that do not generalize under physically novel conditions Geirhos et al. (2020). Without evaluation signals that support causal attribution, progress becomes difficult to distinguish from the accumulation of distributional alignment. 5 Forward Path The consequences identified in Section 4 arise from a single diagnostic limitation: the absence of identifiable signals for separately measuring semantic generalization and physical execution generalization. Addressing this limitation does not primarily require new model architectures; it requires evaluation designs that treat these two capabilities as distinct variables. Evaluation Reform. Separating semantic and physical generalization requires experimental designs that independently manipulate semantic content and physical configuration. Two complementary interventions enable this separation. To evaluate semantic generalization, physical conditions are held constant while semantic content varies. This can be achieved by altering linguistic descriptions of the same scene, or by introducing objects with visually similar appearance but different task-relevant properties. To evaluate physical execution generalization, semantic content is held constant while physical configuration varies. The same instruction and object category are presented under novel poses, surface conditions, or dynamics, requiring different action strategies for successful execution. The key requirement is controlled variation: one factor must be held fixed while the other is systematically perturbed. Increasing dataset scale or benchmark diversity alone does not resolve the attribution problem if these factors remain entangled. A particularly tractable instantiation of this design uses object properties that are linguistically expressible but not visually recoverableâsuch as mass, friction, or material composition Zellers et al. (2019); Gibson (1950); Lerer et al. (2016). A model that correctly grounds expressions such as âthe heavier cupâ or âthe slippery objectâ provides evidence of semantic generalization that cannot be reduced to visual pattern matching Geirhos et al. (2020); Brown et al. (2020). Evaluating whether this grounded representation leads to physically correct manipulation establishes a conditional test: semantic grounding is verified independently, and physical execution is evaluated given correct grounding Agrawal et al. (2016). This design introduces two observable outcomesâ semantic correctness and physical successârather than collapsing both into a single task success signal. Importantly, this protocol does not require access to model internals; it relies only on observable behavior under controlled interventions Bender et al. (2021). In addition, out-of-distribution status should be defined relative to all training data sources, including VLM pretraining corpora, rather than only robot demonstration datasets Zitkovich et al. (2023); Kim et al. (2024); Liu et al. (2023). Without this distinction, nominally novel objects may remain familiar to the model, preventing attribution of performance gains to physical generalization. Reorienting the Role of VLM Backbones. Recent work has highlighted a gap between high-level semantic reasoning and the requirements of physical execution, emphasizing the importance of three-dimensional structure, contact, and spatial relations in manipulation Chen et al. (2026); Wang et al. (2023); Cai et al. (2025). However, under current evaluation protocols, it remains unclear whether observed limitations arise from insufficient physical modeling or from confounding factors such as distributional overlap. Evaluation designs that enable attribution have a direct implication for the role of VLM backbones. Once semantic and physical generalization can be measured separately, the contribution of VLM pretraining to each can be empirically assessed rather than assumed. This may reveal that VLM backbones provide strong semantic interfacesâgrounding open-vocabulary instructions to visual targets Li et al. (2023); Radford et al. (2021); Ahn et al. (2022a)âwhile physical execution generalization depends on mechanisms not captured by passive image-text training LeCun and others (2022); Lake et al. (2017). This perspective suggests a shift in architectural framing. Rather than treating VLM backbones as the foundation from which physical capability is expected to emerge, systems can be designed with explicitly separated components: a semantic interface for grounding language and perception, and a physical module for modeling action-conditioned dynamics Wu et al. (2015); Tamkin et al. (2021); Raghu et al. (2019). On such an architecture, the semantic sufficiency assumption need not be adopted from the outset, because the two components are designed to serve distinct functions independently Bender et al. (2021). This separation has precedent in developmental cognition. Infants exhibit sensitivity to physical properties such as object permanence and causal interaction prior to acquiring language Lake et al. (2017); Leslie and Keeble (1987). Physical reasoning, on this account, develops as a distinct representational capacity rather than as a derivative of semantic knowledge. This reframing does not prescribe a specific architecture. It shifts the design objective: from adapting semantic models to perform physical execution, to constructing systems in which physical execution generalization is measurable, attributable, and improvable independently of semantic grounding. Under such conditions, the contribution of VLM-based representations to physical performance becomes an empirical question rather than an assumption. 6 Conclusion This paper examined a foundational but unverified assumption underlying recent VLM-backbone VLA research: that semantic generalization learned from internet-scale data transfers to physical execution generalization in embodied systems. We showed that, under current evaluation protocols, this assumption is not independently identifiable. Task success rateâdespite being the dominant metricâdoes not distinguish between semantic grounding and physical execution, and therefore does not support causal attribution of performance gains. Reinterpreted through this lens, the steady improvement of benchmark results across VLA systems does not, by itself, establish progress in physical reasoning. Because semantic and physical factors remain entangled in both training and evaluation, the same empirical outcomes are consistent with multiple competing explanations, including improved semantic matching and distributional overlap. What appears as evidence of increasingly general physical capability may instead reflect the limits of what current benchmarks are able to measure. The implication is methodological rather than architectural. Progress on physical execution generalization requires evaluation designs that introduce controlled variation, allowing semantic and physical contributions to be measured separately. This shift also reframes system design: instead of assuming that physical competence emerges from semantic representations, architectures can be structured to treat physical state modeling and semantic grounding as distinct components from the outset, enabling their roles to be empirically assessed rather than inferred. More broadly, this analysis highlights a general principle for embodied AI: when evaluation does not provide identifiable signals for key capabilities, implicit assumptions can become entrenched as de facto premises of a field. This limitation has implications beyond attribution: it also constrains which capabilities are expressed and measured. Pretrained VLMs exhibit abilitiesâfeasibility judgment, constraint awareness, context-sensitive refusalâthat fine-tuning into action policies may degrade, yet task completion benchmarks provide no signal to detect such loss. Progress along the task completion axis alone does not guarantee movement toward general embodied intelligence. In this context, our contribution is not to refute the semantic sufficiency assumption, but to show that it has not yet been tested in a way that would allow it to be confirmed or rejected. Establishing such tests is a necessary step toward grounding claims of generalization in mechanisms that can be measured, attributed, and ultimately understood. References [1] P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine (2016) Learning to poke by poking: experiential learning of intuitive physics. Advances in neural information processing systems 29. Cited by: §2, §3.1, §5. [2] M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, M. Yan, and A. Zeng (2022) Do as i can and not as i say: grounding language in robotic affordances. In arXiv preprint arXiv:2204.01691, Cited by: §5. [3] M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al. (2022) Do as i can, not as i say: grounding language in robotic affordances. arXiv preprint arXiv:2204.01691. Cited by: §3.2. [4] H. Arai, K. Miwa, K. Sasaki, K. Watanabe, Y. Yamaguchi, S. Aoki, and I. Yamamoto (2025) Covla: comprehensive vision-language-action dataset for autonomous driving. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), p. 1933â1943. Cited by: §2, §2. [5] B. D. Argall, S. Chernova, M. Veloso, and B. Browning (2009) A survey of robot learning from demonstration. Robotics and autonomous systems 57 (5), p. 469â483. Cited by: §1. [6] P. W. Battaglia, J. B. Hamrick, and J. B. Tenenbaum (2013) Simulation as an engine of physical scene understanding. Proceedings of the national academy of sciences 110 (45), p. 18327â18332. Cited by: §2, §3.1. [7] S. Belkhale, T. Ding, T. Xiao, P. Sermanet, Q. Vuong, J. Tompson, Y. Chebotar, D. Dwibedi, and D. Sadigh (2024) Rt-h: action hierarchies using language. arXiv preprint arXiv:2403.01823. Cited by: §2. [8] E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell (2021) On the dangers of stochastic parrots: can language models be too big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT â21, New York, NY, USA, p. 610â623. External Links: ISBN 9781450383097, Link, Document Cited by: §5, §5. [9] A. Billard and D. Kragic (2019) Trends and challenges in robot manipulation. Science 364 (6446), p. eaat8414. Cited by: §3.1. [10] J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al. (2025) Gr00t n1: an open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734. Cited by: §1, Table 1, §2, §3.2. [11] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. (2024) pâiâ_â0pi\_0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164. Cited by: §1, Table 1, §2, §2, §3.2, §3.3, Table 2. [12] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. (2020) Language models are few-shot learners. Advances in neural information processing systems 33, p. 1877â1901. Cited by: §5. [13] Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, S. Gao, X. He, X. Hu, X. Huang, et al. (2025) Agibot world colosseo: a large-scale manipulation platform for scalable and intelligent embodied systems. arXiv preprint arXiv:2503.06669. Cited by: §2. [14] W. Cai, I. Ponomarenko, J. Yuan, X. Li, W. Yang, H. Dong, and B. Zhao (2025) Spatialbot: precise spatial understanding with vision language models. In 2025 IEEE International Conference on Robotics and Automation (ICRA), p. 9490â9498. Cited by: §5. [15] H. Chen, J. Liu, C. Gu, Z. Liu, R. Zhang, X. Li, X. He, Y. Guo, C. Fu, S. Zhang, et al. (2025) Fast-in-slow: a dual-system foundation model unifying fast manipulation within slow reasoning. arXiv preprint arXiv:2506.01953. Cited by: §2. [16] K. Chen, C. Li, C. Tu, J. Pan, Y. Ma, W. Chen, Z. Zhou, X. Xu, S. James, C. Fu, R. Xiong, P. Abbeel, Y. Liu, and Q. Dou (2026) A retrieval-augmented framework enabling vlm spatial awareness for object-centric robot manipulation. Science Robotics 11 (113), p. eaea2092. External Links: Document, Link Cited by: §5. [17] A. DâAmour, K. Heller, D. Moldovan, B. Adlam, B. Alipanahi, A. Beutel, C. Chen, J. Deaton, J. Eisenstein, M. D. Hoffman, et al. (2022) Underspecification presents challenges for credibility in modern machine learning. Journal of Machine Learning Research 23 (226), p. 1â61. Cited by: §4, §4. [18] S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn (2019) Robonet: large-scale multi-robot learning. arXiv preprint arXiv:1910.11215. Cited by: §4. [19] S. Deng, M. Yan, S. Wei, H. Ma, Y. Yang, J. Chen, Z. Zhang, T. Yang, X. Zhang, W. Zhang, et al. (2025) Graspvla: a grasping foundation model pre-trained on billion-scale synthetic action data. arXiv preprint arXiv:2505.03233. Cited by: §2. [20] P. Ding, H. Zhao, W. Zhang, W. Song, M. Zhang, S. Huang, N. Yang, and D. Wang (2024) Quar-vla: vision-language-action model for quadruped robots. In European Conference on Computer Vision, p. 352â367. Cited by: §2. [21] Z. Duan, Y. Zhang, S. Geng, G. Liu, J. Boedecker, and C. X. Lu (2025) Fast ecot: efficient embodied chain-of-thought via thoughts reuse. arXiv preprint arXiv:2506.07639. Cited by: §2. [22] S. Fei, S. Wang, J. Shi, Z. Dai, J. Cai, P. Qian, L. Ji, X. He, S. Zhang, Z. Fei, et al. (2025) Libero-plus: in-depth robustness analysis of vision-language-action models. arXiv preprint arXiv:2510.13626. Cited by: §3.2. [23] R. Firoozi, J. Tucker, S. Tian, A. Majumdar, J. Sun, W. Liu, Y. Zhu, S. Song, A. Kapoor, K. Hausman, et al. (2025) Foundation models in robotics: applications, challenges, and the future. The International Journal of Robotics Research 44 (5), p. 701â739. Cited by: §1. [24] R. Geirhos, J. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann (2020) Shortcut learning in deep neural networks. Nature Machine Intelligence 2 (11), p. 665â673. Cited by: §4, §4, §5. [25] J. J. Gibson (1950) The perception of the visual world.. Cited by: §5. [26] J. J. Gibson (2014) The ecological approach to visual perception: classic edition. Psychology press. Cited by: §2, §3.1. [27] J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. M. A. Patwary, Y. Yang, and Y. Zhou (2017) Deep learning scaling is predictable, empirically. arXiv preprint arXiv:1712.00409. Cited by: §4. [28] Y. Hu, Q. Xie, V. Jain, J. Francis, J. Patrikar, N. Keetha, S. Kim, Y. Xie, T. Zhang, H. Fang, et al. (2023) Toward general-purpose robots via foundation models: a survey and meta-analysis. arXiv preprint arXiv:2312.08782. Cited by: §1. [29] S. Huang, J. Shao, K. Wang, Q. Chen, J. Sun, Y. Guo, M. Schwager, and J. Bohg (2026) Breaking lock-in: preserving steerability under low-data vla post-training. arXiv preprint arXiv:2604.23121. Cited by: §3.2. [30] W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, et al. (2022) Inner monologue: embodied reasoning through planning with language models. arXiv preprint arXiv:2207.05608. Cited by: §3.2. [31] M. Huh, B. Cheung, T. Wang, and P. Isola (2024) The platonic representation hypothesis. arXiv preprint arXiv:2405.07987. Cited by: §4. [32] D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al. (2018) Scalable deep reinforcement learning for vision-based robotic manipulation. In Conference on robot learning, p. 651â673. Cited by: §1, §4. [33] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei (2020) Scaling laws for neural language models. arXiv preprint arXiv:2001.08361. Cited by: §4. [34] K. Kawaharazuka, J. Oh, J. Yamada, I. Posner, and Y. Zhu (2025) Vision-language-action models for robotics: a review towards real-world applications. IEEE Access. Cited by: §1. [35] D. Kiela, M. Bartolo, Y. Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia, et al. (2021) Dynabench: rethinking benchmarking in nlp. In Proceedings of the 2021 conference of the North American chapter of the Association for Computational Linguistics: human language technologies, p. 4110â4124. Cited by: §4. [36] M. J. Kim, C. Finn, and P. Liang (2025) Fine-tuning vision-language-action models: optimizing speed and success. arXiv preprint arXiv:2502.19645. Cited by: §1, Table 1, §2, §2, §3.1, §3.2, §3.3, Table 2, §4. [37] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. (2024) Openvla: an open-source vision-language-action model. arXiv preprint arXiv:2406.09246. Cited by: §1, Table 1, §2, §2, §2, §3.2, §3.2, §3.3, §5. [38] B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman (2017) Building machines that learn and think like people. Behavioral and brain sciences 40, p. e253. Cited by: §2, §5, §5. [39] Y. LeCun et al. (2022) A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27. Open Review 62 (1), p. 1â62. Cited by: §5. [40] A. Lerer, S. Gross, and R. Fergus (2016) Learning physical intuition of block towers by example. In International conference on machine learning, p. 430â438. Cited by: §2, §5. [41] A. M. Leslie and S. Keeble (1987) Do six-month-old infants perceive causality?. Cognition 25 (3), p. 265â288. Cited by: §5. [42] C. Li, J. Wen, Y. Peng, Y. Peng, and Y. Zhu (2026) Pointvla: injecting the 3d world into vision-language-action models. IEEE Robotics and Automation Letters 11 (3), p. 2506â2513. Cited by: §2. [43] H. Li, S. Yang, Y. Chen, X. Chen, X. Yang, Y. Tian, H. Wang, T. Wang, D. Lin, F. Zhao, et al. (2025) CronusVLA: towards efficient and robust manipulation via multi-frame vision-language-action modeling. arXiv preprint arXiv:2506.19816. Cited by: §2. [44] J. Li, D. Li, S. Savarese, and S. Hoi (2023) Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, p. 19730â19742. Cited by: §2, §3.1, §5. [45] M. Li, Z. Zhao, Z. Che, F. Liao, K. Wu, Z. Xu, P. Ren, Z. Jin, N. Liu, and J. Tang (2025) Switchvla: execution-aware task switching for vision-language-action models. arXiv preprint arXiv:2506.03574. Cited by: §2. [46] W. Li, R. Zhang, R. Shao, Z. Fang, K. Zhou, Z. Tian, and L. Nie (2026) Semanticvla: semantic-aligned sparsification and enhancement for efficient robotic manipulation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 18397â18405. Cited by: §2. [47] X. Li, C. Mata, J. Park, K. Kahatapitiya, Y. S. Jang, J. Shang, K. Ranasinghe, R. Burgert, M. Cai, Y. J. Lee, et al. (2024) Llara: supercharging robot learning data for vision-language policy. arXiv preprint arXiv:2406.20095. Cited by: §2. [48] F. Lin, R. Nai, Y. Hu, J. You, J. Zhao, and Y. Gao (2025) Onetwovla: a unified vision-language-action model with adaptive reasoning. arXiv preprint arXiv:2505.11917. Cited by: §2. [49] B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone (2023) Libero: benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems 36, p. 44776â44791. Cited by: §4, §5. [50] H. Liu, S. Ruan, J. Long, J. Wu, J. Hou, H. Tang, T. Jiang, W. Zhou, and W. Yao (2025) Eva-vla: evaluating vision-language-action modelsâ robustness under real-world physical variations. arXiv preprint arXiv:2509.18953. Cited by: §3.2. [51] O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard (2022) Calvin: a benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robotics and Automation Letters 7 (3), p. 7327â7334. Cited by: §4. [52] S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y. Zhu (2024) Robocasa: large-scale simulation of everyday tasks for generalist robots. arXiv preprint arXiv:2406.02523. Cited by: §3.2, §4. [53] L. Pinto and A. Gupta (2016) Supersizing self-supervision: learning to grasp from 50k tries and 700 robot hours. In 2016 IEEE international conference on robotics and automation (ICRA), p. 3406â3413. Cited by: §2, §3.1. [54] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748â8763. Cited by: §2, §3.1, §5. [55] M. Raghu, C. Zhang, J. Kleinberg, and S. Bengio (2019) Transfusion: understanding transfer learning for medical imaging. Advances in neural information processing systems 32. Cited by: §5. [56] H. Ravichandar, A. S. Polydoros, S. Chernova, and A. Billard (2020) Recent advances in robot learning from demonstration. Annual review of control, robotics, and autonomous systems 3 (1), p. 297â330. Cited by: §1. [57] D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, M. Young, J. Crespo, and D. Dennison (2015) Hidden technical debt in machine learning systems. Advances in neural information processing systems 28. Cited by: §4. [58] H. Song, D. Qu, Y. Yao, Q. Chen, Q. Lv, Y. Tang, M. Shi, G. Ren, M. Yao, B. Zhao, et al. (2025) Hume: introducing system-2 thinking in visual-language-action model. arXiv preprint arXiv:2505.21432. Cited by: §2. [59] W. Song, Z. Zhou, H. Zhao, J. Chen, P. Ding, H. Yan, Y. Huang, F. Tang, D. Wang, and H. Li (2026) Reconvla: reconstructive vision-language-action model as effective robot perceiver. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 18549â18557. Cited by: §2. [60] R. S. Sutton, A. G. Barto, et al. (1998) Reinforcement learning: an introduction. Vol. 1, MIT press Cambridge. Cited by: §1. [61] A. Tamkin, M. Brundage, J. Clark, and D. Ganguli (2021) Understanding the capabilities, limitations, and societal impact of large language models. arXiv preprint arXiv:2102.02503. Cited by: §5. [62] H. Wang, G. Zhang, Y. Yan, R. R. Kompella, and G. Liu (2026) VLA knows its limits. arXiv preprint arXiv:2602.21445. Cited by: §3.2. [63] J. Wang (2026-Mar.) LatentVLA: taming latent space for generalizable and long-horizon bimanual manipulation. Proceedings of the AAAI Conference on Artificial Intelligence 40 (22), p. 18593â18601. External Links: Document, Link Cited by: §2. [64] R. Wang, J. Mao, J. Hsu, H. Zhao, J. Wu, and Y. Gao (2023) Programmatically grounded, compositionally generalizable robotic manipulation. arXiv preprint arXiv:2304.13826. Cited by: §5. [65] Y. Wang, P. Ding, L. Li, C. Cui, Z. Ge, X. Tong, W. Song, H. Zhao, W. Zhao, P. Hou, et al. (2026) Vla-adapter: an effective paradigm for tiny-scale vision-language-action model. In Proceedings of the AAAI conference on artificial intelligence, Vol. 40, p. 18638â18646. Cited by: §2. [66] J. Wen, M. Zhu, Y. Zhu, Z. Tang, J. Li, Z. Zhou, C. Li, X. Liu, Y. Peng, C. Shen, et al. (2024) Diffusion-vla: generalizable and interpretable robot foundation model via self-generated reasoning. arXiv preprint arXiv:2412.03293. Cited by: §2. [67] J. Wu, I. Yildirim, J. J. Lim, B. Freeman, and J. Tenenbaum (2015) Galileo: perceiving physical object properties by integrating a physics engine with deep learning. Advances in neural information processing systems 28. Cited by: §5. [68] Y. Yadav, Z. Zhou, A. Wagenmaker, K. Pertsch, and S. Levine (2025) Robust finetuning of vision-language-action robot policies via parameter merging. arXiv preprint arXiv:2512.08333. Cited by: §3.2. [69] M. Zare, P. M. Kebria, A. Khosravi, and S. Nahavandi (2024) A survey of imitation learning: algorithms, recent developments, and challenges. IEEE Transactions on Cybernetics 54 (12), p. 7173â7186. Cited by: §1. [70] M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine (2024) Robotic control via embodied chain-of-thought reasoning. arXiv preprint arXiv:2407.08693. Cited by: §2, §2. [71] R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi (2019) From recognition to cognition: visual commonsense reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 6720â6731. Cited by: §5. [72] A. Zeng, M. Attarian, B. Ichter, K. Choromanski, A. Wong, S. Welker, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, et al. (2022) Socratic models: composing zero-shot multimodal reasoning with language. arXiv preprint arXiv:2204.00598. Cited by: §2, §3.1. [73] R. Zhang, M. Dong, Y. Zhang, L. Heng, X. Chi, G. Dai, L. Du, D. Wang, Y. Du, and S. Zhang (2026) Mole-vla: dynamic layer-skipping vision language action model via mixture-of-layers for efficient robot manipulation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 18764â18772. Cited by: §2. [74] W. Zhang, Z. Xu, and H. Cai (2024) Recognizing limits: investigating infeasibility in large language models. arXiv preprint arXiv:2408.05873. Cited by: §3.2. [75] Y. Zhong, X. Huang, R. Li, C. Zhang, Z. Chen, T. Guan, F. Zeng, K. N. Lui, Y. Ye, Y. Liang, et al. (2026) Dexgraspvla: a vision-language-action framework towards general dexterous grasping. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 18836â18844. Cited by: §2. [76] X. Zhou, Y. Xu, G. Tie, Y. Chen, G. Zhang, D. Chu, P. Zhou, and L. Sun (2025) LIBERO-pro: towards robust and fair evaluation of vision-language-action models beyond memorization. arXiv preprint arXiv:2510.03827. Cited by: §3.2. [77] Z. Zhou, Y. Zhu, M. Zhu, J. Wen, N. Liu, Z. Xu, W. Meng, Y. Peng, C. Shen, F. Feng, et al. (2025) Chatvla: unified multimodal understanding and robot control with vision-language-action model. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 5377â5395. Cited by: §2. [78] M. Zhu, Y. Zhu, J. Li, Z. Zhou, J. Wen, X. Liu, C. Shen, Y. Peng, and F. Feng (2025) Objectvla: end-to-end open-world object manipulation without demonstration. arXiv preprint arXiv:2502.19250. Cited by: §2. [79] B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. (2023) Rt-2: vision-language-action models transfer web knowledge to robotic control. In Conference on Robot Learning, p. 2165â2183. Cited by: §1, Table 1, §2, §2, §3.1, §3.2, §3.3, Table 2, Table 2, §4, §5.