Paper deep dive
GeoProp: Grounding Robot State in Vision for Generalist Manipulation
Guoyang Zhao, Quanhao Qian, Gongjie Zhang, Wenhao Li, Jiuniu Wang, Xiaowei Lu, Deli Zhao, Ran Xu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/9/2026, 6:58:48 AM
Summary
The paper introduces GeoProp, a lightweight, plug-and-play adapter that aligns robot proprioception with visual tokens through explicit geometric grounding. By projecting 3D end-effector kinematics onto the 2D image plane, GeoProp samples localized visual features and injects state-derived priors via FiLM modulation. It also uses predictive kinematic sampling to capture motion intent. Evaluated across 67 simulation and real-world tasks, GeoProp significantly improves performance over baselines like Diffusion Policy and Ļ0 with minimal parameter overhead, demonstrating that explicit geometric grounding is a powerful inductive bias for generalist manipulation policies.
Entities (10)
Relation Signals (10)
GeoProp ā alignsmodality ā Proprioception
confidence 95% Ā· aligns proprioception with vision through explicit geometric grounding
GeoProp ā improvesperformance ā Diffusion Policy
confidence 95% Ā· GeoProp improves Diffusion Policy by 8.7% on 63 simulation tasks
GeoProp ā improvesperformance ā Ļ0
confidence 95% Ā· improves ... Ļ0 by 4.0% on the RoboTwin subset
GeoProp ā evaluatedon ā MetaWorld
confidence 92% Ā· MetaWorld contains 50 tabletop manipulation tasks
GeoProp ā evaluatedon ā RoboTwin
confidence 92% Ā· evaluates on ... RoboTwin subset
GeoProp ā evaluatedon ā RLBench
confidence 92% Ā· RLBench provides visually rich manipulation tasks
GeoProp ā utilizestechnique ā FiLM Modulation
confidence 90% Ā· injects state-derived spatial priors into the corresponding visual features via FiLM modulation
GeoProp ā utilizestechnique ā Predictive Kinematic Sampling
confidence 90% Ā· introduce Predictive Kinematic Sampling to capture motion intent
Ļ0 ā usesfusionmethod ā
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to ground the robot's state within the scene, frequently underperforming even vision-only baselines. To address this, we introduce GeoProp, a lightweight, plug-and-play adapter that aligns proprioception with vision through explicit geometric grounding and spatial feature sampling. GeoProp projects the robot state onto the image plane to sample localized visual features, constructing a grounded state token. It then injects state-derived spatial priors into the corresponding visual features via FiLM modulation. To capture motion intent, GeoProp further samples features at a short-horizon predicted coordinate derived from recent kinematics, providing look-ahead visual context. Across 67 tasks, GeoProp improves Diffusion Policy by 8.7% on 63 simulation tasks and pi_0 by 4.0% on the RoboTwin subset, and yields a 10.6% average gain across both policy families in the real world, while adding only 2-3% to the parameter count. These results demonstrate that GeoProp is a simple yet high-impact inductive bias for generalist embodied policies. Project page: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.07101v1
- Canonical: https://arxiv.org/abs/2607.07101v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
63,939 characters extracted from source content.
Expand or collapse full text
GeoProp: Grounding Robot State in Vision for Generalist Manipulation Guoyang Zhao1,ā, Quanhao Qian2,3,ā, Gongjie Zhang4, Wenhao Li5, Jiuniu Wang2,3, Xiaowei Lu2,3, Deli Zhao2,3, Ran Xu2,3, š 1Tongji University 2DAMO Academy, Alibaba Group 3HuPan Lab 4Alibaba Group 5Nanyang Technological University āEqual contribution. š Corresponding author Abstract Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to ground the robotās state within the scene, frequently underperforming even vision-only baselines. To address this, we introduce GeoProp, a lightweight, plug-and-play adapter that aligns proprioception with vision through explicit geometric grounding and spatial feature sampling. GeoProp projects the robot state onto the image plane to sample localized visual features, constructing a grounded state token. It then injects state-derived spatial priors into the corresponding visual features via FiLM modulation. To capture motion intent, GeoProp further samples features at a short-horizon predicted coordinate derived from recent kinematics, providing look-ahead visual context. Across 67 tasks, GeoProp improves Diffusion Policy by 8.7% on 63 simulation tasks and Ļ0 _0 by 4.0% on the RoboTwin subset, and yields a 10.6% average gain across both policy families in the real world, while adding only 2ā3% to the parameter count. These results demonstrate that GeoProp is a simple yet high-impact inductive bias for generalist embodied policies. Project page: https://alibaba-damo-academy.github.io/GeoProp/. Keywords: Robot Perception, Proprioception, Visual Grounding 1 Introduction Figure 1: Proprioceptive-to-image attention in Ļ0 _0: GeoProp concentrates attention on the gripper and manipulated objects, while vanilla attention is diffuse. Despite the diversification of robot learning architecturesāspanning diffusion-based controllers [3, 39, 5], transformer-based action predictors [44, 10], and large-scale visionālanguageāaction (VLA) systems [2, 46, 17, 1]āa persistent representational limitation remains: the structural decoupling between high-dimensional visual observations and low-dimensional proprioceptive feedback. Current mainstream frameworks encode proprioception as a global, ungrounded state embedding and fuse it with vision either through simple concatenation [14, 44, 16, 26] or through more expressive cross-attention [1, 30]. Both designs typically lack an explicit state-vision correspondence and require the model to learn 3D kinematics-to-2D feature alignment implicitly, overlooking the robotās intrinsic geometric relationship with the visual scene. Our empirical analysis shows that ungrounded proprioception can be counterproductive: without explicit alignment, the state vector may introduce spurious correlations that cause the model to underperform vision-only alternatives. We propose GeoProp, a lightweight adapter that bridges the gap between 3D kinematics and 2D vision by transforming the robotic state into a localized geometric modulation within the 2D visual feature map. Instead of treating proprioception as an ungrounded auxiliary input, GeoProp projects the end-effector pose onto the image plane and uses the corresponding localized visual features as the state-aligned visual feature of the robot. This enables the robotās configuration to inherit scene semantics within the same visual latent space, promoting modality-consistent fusion, as visualized by the attention heatmaps in Fig. 1 and Appendix B.3. In the vanilla Ļ0 _0 [1] model, state embeddings often fail to produce spatially coherent attention and yield diffuse, background-driven patterns. In contrast, GeoProp tightly couples proprioceptive cues to the relevant visual regions by establishing correspondences analytically. GeoProp anchors the robotic state within the visual manifold across both static configuration and motion dynamics. First, we employ Spatially-Aligned Modulationāinspired by FiLM [31]āto inject state-derived geometric cues into tokens near the projected end-effector location. Second, we introduce Predictive Kinematic Sampling to capture motion intent by sampling visual features at a short-horizon predicted coordinate estimated from recent kinematics, providing a look-ahead visual context. As GeoProp operates at the interface between perception and state representation, it is inherently framework-agnostic. It addresses a shared limitation of prevailing fusion paradigms: Diffusion Policy [3] typically performs shallow state fusion via concatenation, whereas VLA models such as Ļ0 _0 [1] employ deep cross-modality attention; nevertheless, both commonly represent proprioception as a standalone state vector that is not spatially aligned with visual tokens. GeoProp improves both frameworks by supplying a spatially grounded, geometry-aware signal that aligns robot kinematics with visual semantics, without requiring changes to the backbone architecture. We validate GeoProp across 67 tasks, including 63 tasks in simulation [38, 13, 28] and 4 real-world tasks on a Mobile ALOHA system [9]. Our results show average absolute improvements of 8.7% on Diffusion Policy across 63 simulation tasks and 4.0% on Ļ0 _0 on the RoboTwin subset, together with a 10.6% real-world gain averaged over both policy families, with negligible parameter overhead of 2ā3%. Our contributions: (1) We empirically show that treating proprioception as an ungrounded global vector can lead to modality misalignment and degraded manipulation performance compared to vision-only baselines, motivating the need for an explicit, geometry-based stateāvision correspondence. (2) We propose GeoProp, a plug-and-play adapter that transforms 3D proprioception into image-grounded visual tokens, bridging the representational gap between robot kinematics and scene semantics through geometric projection, localized feature modulation, and predictive sampling of motion intent. (3) We demonstrate that GeoProp consistently enhances performance across distinct policy architectures in both simulation and the real world, validating explicit geometric grounding as an efficient and powerful inductive bias for embodied AI. 2 Related Work 2.1 Generalist Robot Policies and Foundation Models Robotic manipulation has shifted from state-based reinforcement learning [22, 34], which often assumes compact low-dimensional states, to policies learned from high-dimensional visual observations [3, 44, 25, 35]. More recently, Vision-Language-Action (VLA) models [46, 17, 1] leverage large-scale pre-training for semantic generalization and open-vocabulary task following. However, precise manipulation still requires aligning what the robot sees with where the robot physically is. Most frameworks fuse proprioception through concatenation, MLP state encoders, or transformer state/action tokens, treating robot state as a collapsed global context vector. This leaves the correspondence between visual semantics and kinematic state implicit, forcing policies to learn stateāvision alignment from data alone. 2.2 ProprioceptionāVision Alignment Recent work has examined the limitations of ungrounded proprioceptive fusion. Some studies [24, 43] suggest that isolated state vectors can induce embodiment- or trajectory-specific overfitting; relative-action formulations [43] improve spatial generalization but reduce access to absolute kinematic state. Other methods align proprioception through learned objectives or alternative state parameterizations, including contrastive losses [18], discrete codebooks [7], and verbalized state tokens [37]. Most closely related, robot-centric positional encoding (RC-PE) [19] projects the end-effector to the image plane and injects a dense relative-coordinate embedding into every visual token. RC-PE shares GeoPropās premise that projecting the robot state into image space exposes a useful geometric cue, but spreads this signal across the full feature grid rather than co-locating proprioception with the visual evidence. A complementary line targets the input or planning level: LLARVA [29], HAMSTER [21], and PEEK [41] predict visual traces, paths, or affordance masks from a VLM as intermediate task plans, while AimBot [4] overlays analytically-projected robot reticles on the raw RGB input. GeoProp instead operates inside the policyās feature stack, FiLM-modulating intermediate visual features at projected end-effector locations without changing the input image or requiring a VLM planner, and is thus complementary to input/planning-level methods. 2.3 Geometric Grounding in Robotic Perception Explicit geometric priors improve sample efficiency, spatial reasoning, and robustness in manipulation [45, 27, 40]. Many methods construct 3D scene representations, including voxel, point-cloud, or 3D-aware policy backbones [36, 11, 39, 42], giving vision and robot state a shared coordinate substrate but often requiring depth, multi-view observations, 3D reconstruction, or dedicated 3D architectures. Recent VLA-oriented methods use 3D Gaussian Splatting or feature distillation for scene-centric spatial reasoning [20, 32, 33]; for instance, GeoPredict [32] uses predictive 3D Gaussian geometry and robot keypoint trajectories as auxiliary supervision. GeoProp instead takes a lightweight, robot-state-centric route: it anchors end-effector kinematics directly onto 2D visual feature tokens via projective grounding, preserving explicit stateāvision correspondence without full 3D reconstruction. 3 Methodology Figure 2: GeoProp projects end-effector and look-ahead waypoints to image features, producing spatially aligned state tokens for downstream policies. As illustrated in Fig. 2, GeoProp aligns 3D proprioception with 2D visual tokens through three components: i) geometric projection and feature sampling, i) spatially aligned feature modulation, and i) predictive kinematic sampling. The method assumes access to camera intrinsics K and extrinsics (,)(R,t), intermediate spatial feature maps from the vision backbone, and the end-effector 3D position. The visual token sampled at the projected end-effector location is referred to as the grounded state token, as it carries local scene evidence at the robotās current interaction point. GeoProp uses FPN [23] for multi-scale aggregation and applies FiLM modulation [31] only at the aligned feature cell. The resulting grounded and predictive tokens are concatenated with global visual tokens, and language tokens when available, and passed to the downstream policy without modifying the backbone. 3.1 Geometric Projection and Feature Grounding We denote the proprioceptive state at time t as t=[xt,yt,zt,qw,t,qx,t,qy,t,qz,t,gt]ā¤āā8p_t=[x_t,y_t,z_t,q_w,t,q_x,t,q_y,t,q_z,t,g_t] ^8, including the end-effector position, orientation, and gripper state. Standard fusion methods inject tp_t as an ungrounded global vector, requiring the policy to infer the correspondence between 3D kinematics and 2D visual tokens. GeoProp instead makes this correspondence explicit by projecting the end-effector into the image plane and sampling the visual feature at the projected location. Let t=[xt,yt,zt]ā¤r_t=[x_t,y_t,z_t] be the end-effector position in the robot/world frame. Given camera intrinsics K and extrinsics (,)(R,t), we obtain the image-plane coordinate t=Ī ā(t+),t=(ut,vt)ā¤,q_t= _K\! (Rr_t+t ), _t=(u_t,v_t) , (1) where Ī ā(ā ) _K(Ā·) denotes perspective projection with intrinsics K. The image coordinate tq_t is then mapped to a continuous coordinate ĀÆt=Ļā(t) q_t=Ļ(q_t) on the feature grid, where Ļā(ā )Ļ(Ā·) maps image-plane coordinates to the corresponding feature-grid coordinates induced by the visual backbone. Given a spatial feature map F, the feature at a continuous coordinate is obtained via bilinear sampling: ā(,ĀÆ)=āāā(ĀÆ)wā(ĀÆ)ā[],S(F, q)= _c ( q)w_c( q)\,F[c], (2) where ā(ĀÆ)N( q) denotes the four neighboring grid cells and w_c are bilinear interpolation weights. After the modulation and aggregation stage in Sec. 3.2, GeoProp forms the grounded state token as t=ā(mod,ĀÆt). Ļ_t=S(F_mod, q_t). (3) 3.2 Spatially-Aligned Feature Modulation The grounded token t Ļ_t provides an image-aligned state representation, but it is extracted only after visual encoding. To inject the same geometric prior into the visual hierarchy, GeoProp applies state-conditioned FiLM modulation locally at the projected end-effector location before FPN aggregation. For each intermediate feature map encāF _enc, let tāc _t denote the aligned grid cell induced by ĀÆt q_t at level ā . Given the proprioceptive state tp_t, GeoProp predicts channel-wise FiLM parameters tā γ _t and tā β _t for each feature level. The modulated feature map is defined as modāā[]=(1+tā)āencāā[]+tā,=tā,encāā[],otherwise.F _mod[c]= cases(1+ γ _t) _enc[c]+ β _t,&c=c _t,\\ F _enc[c],&otherwise. cases (4) This operation conditions the visual representation only at the feature cell geometrically aligned with the end-effector. Thus, proprioceptive conditioning remains spatially localized prior to multi-scale aggregation. The modulated multi-scale features are then aggregated with an FPN: mod=FPNā(modā=1L).F_mod=FPN(\F _mod\_ =1^L). (5) The grounded state token is obtained by sampling modF_mod at the continuous feature-grid coordinate defined in Sec. 3.1. 3.3 Predictive Kinematic Sampling Instantaneous proprioception localizes the current end-effector state, but provides limited information about imminent motion trends. GeoProp therefore augments grounded state representations with a predictive spatial token constructed from a short-horizon look-ahead waypoint. Given a temporal window of recent 3D positions tāk,ā¦,t\r_t-k,ā¦,r_t\, we use polynomial extrapolation to estimate a future waypoint ^t+1 r_t+1. This waypoint is projected to the image plane using Eq. (1) and mapped to the feature grid, yielding a continuous look-ahead coordinate ĀÆt+1pre q^pre_t+1. Let rawF_raw denote the FPN-aggregated feature map computed from the unmodulated encoder features. We construct the predictive token by sampling this feature map at the look-ahead coordinate: pre=ā(raw,ĀÆt+1pre). Ļ_pre=S(F_raw, q^pre_t+1). (6) Sampling from unmodulated features decouples the look-ahead token from state-conditioned modulation, making it a spatial preview of the anticipated motion rather than an additional modulation site. Let globalh_global denote the set of global visual tokens produced by the vision backbone. The policy input is formed by concatenating the grounded state token, the predictive token, and the global visual tokens: t=[t;pre;global].z_t=[ Ļ_t; Ļ_pre;h_global]. (7) This representation jointly encodes grounded current state, anticipated motion context, and global scene semantics within a unified geometry-aware feature space. For bimanual platforms (e.g., the ALOHA setups in our RoboTwin and Mobile ALOHA experiments), GeoProp is applied independently to each arm, producing one grounded and one predictive token per end-effector that are concatenated with the global visual tokens. 4 Simulation Experiments We evaluate GeoProp on 63 simulated tasks across three benchmarks spanning single-arm and bimanual manipulation, relative and absolute pose control, and both task-specific and multi-task training. 4.1 Experimental Setup Benchmarks. We evaluate on three established manipulation benchmarks. MetaWorld [38] contains 50 tabletop manipulation tasks with a Sawyer robot in MuJoCo; we train one policy per task using 25 scripted demonstrations and a corner-view camera, with actions parameterized as end-effector translation deltas (Īāx,Īāy,Īāz)( x, y, z). We follow Qian et al. [33] to group the 50 tasks into the Easy / Medium / Hard / Very Hard difficulty splits used in Table 1. RLBench [13] provides visually rich manipulation tasks in CoppeliaSim with a Franka Panda; following prior work [15, 33], we evaluate on six tasks using 100 OMPL-generated demonstrations per task, front-view images, and absolute 7-DoF end-effector pose actions. RoboTwin [28] evaluates bimanual manipulation with a simulated ALOHA setup; following [8], we use seven tasks with 50 demonstrations per task, head-camera observations, and absolute 14-DoF actions for the two arms. MetaWorld uses per-task training, while RLBench and RoboTwin use multi-task training. Baselines. We compare GeoProp with two policy families that represent different mechanisms for proprioceptionāvision integration. Diffusion Policy (DP) [3] uses shallow fusion, where proprioception is encoded by an MLP and concatenated with visual tokens for policy conditioning. We instantiate DP with both ResNet-18 [12] and ViT-B/16 [6] vision backbones to test backbone generality. We also evaluate Ļ0 _0 [1], a 3B-parameter vision-language-action model in which state, image, and language tokens interact through deep transformer attention. For each policy family, we compare three variants: i) No-Proprio, which removes proprioceptive input; i) Vanilla, which uses the original proprioceptive conditioning mechanism; and i) GeoProp, which replaces the original proprioceptive conditioning mechanism with our geometry-grounded adapter. All variants use the same data, backbone, action space, and training schedule within each policy family. This setup isolates the effect of proprioceptive conditioning while controlling for backbone capacity, action parameterization, and optimization settings. Implementation details. All policies process RGB observations resized to 224Ć224224Ć224. For success-rate reporting, MetaWorld results are averaged over 22 seeds with 2525 evaluation rollouts each (5050 rollouts per task), while RLBench and RoboTwin use 2525 evaluation rollouts per task. For DP, ResNet-18 is trained from scratch, while ViT-B/16 is initialized from ImageNet-1K pretrained weights using timm. DP is trained for 100 epochs with Adam (β1,β2)=(0.95,0.999)( _1, _2)=(0.95,0.999), a cosine learning-rate schedule, 10% linear warm-up, and initial learning rate 1Ć10ā41Ć10^-4. For Ļ0 _0, we fine-tune for 30k steps with AdamW (β1,β2)=(0.9,0.95)( _1, _2)=(0.9,0.95), a cosine schedule, 3% warm-up, and peak learning rate 5Ć10ā55Ć10^-5. GeoProp uses two-layer MLPs to generate FiLM parameters, and predictive kinematic sampling extrapolates one future waypoint from a 4-frame history window. Full architectural and hyperparameter details are provided in Appendix A. Table 1: Diffusion Policy success rates (%) across simulation benchmarks. Results are reported with ResNet-18 and ViT-Base backbones. Per-task numbers are in Appendix B.2. Parentheses denote absolute gains over Vanilla, and Ī Params reports the parameter overhead of GeoProp. Backbone Method MetaWorld (50) RLBench RoboTwin Overall Ī Params Easy Medium Hard V. Hard Mean Mean Avg. ResNet-18 Vanilla 81.8 49.1 46.0 66.4 66.0 53.7 66.8 ā No-Proprio 80.9 51.1 51.0 74.4 68.0 53.7 68.1 ā GeoProp 82.9 67.5 59.0 82.4 79.3 62.9 75.3 (+8.5) +3.0% ViT-Base Vanilla 78.3 46.7 34.0 63.6 65.3 41.1 62.0 ā No-Proprio 78.5 50.5 30.7 52.8 70.7 42.9 62.3 ā GeoProp 81.5 62.5 52.7 69.2 79.3 52.0 71.0 (+9.0) +2.3% 4.2 Quantitative Results Table 1 reports the main results on Diffusion Policy (DP). Across 63 tasks and two vision backbones, GeoProp achieves the best overall performance. With ResNet-18, GeoProp reaches 75.3% average success, improving over Vanilla by 8.5 points and over No-Proprio by 7.2 points. With ViT-Base, GeoProp achieves 71.0%, yielding 9.0 and 8.7 point gains over Vanilla and No-Proprio, respectively. The gains are especially pronounced on the harder MetaWorld subsets, suggesting that spatially grounded proprioception is more beneficial in manipulation settings that require precise visual-state alignment. These improvements come with limited overhead, adding only 2.3ā3.0% parameters to DP. Table 2: Success rates (%) on Ļ0 _0 across RoboTwin tasks. Method Beat Block Hammer Click Alarm Clock Move Can to Pot Move Playing Card Place Shoe Scan Object Stack Two Blocks Avg. No-Proprio 80 88 80 76 40 24 44 61.7 Ļ0 _0 (Vanilla) 76 100 80 72 48 24 48 64.0 GeoProp 84 88 92 72 52 28 60 68.0 We further evaluate GeoProp on the VLA model Ļ0 _0 in Table 2. Since Ļ0 _0 uses deep cross-modality attention between state, image, and language tokens, it provides a strong baseline for implicit proprioceptionāvision alignment. GeoProp improves the average success rate from 64.0% to 68.0% and outperforms the vanilla Ļ0 _0 baseline on 5 out of 7 tasks. This suggests that explicit geometric grounding remains complementary even when the backbone has substantial capacity for cross-modal interaction. The parameter overhead on Ļ0 _0 is also small (+1.98%). The gain on Ļ0 _0 further holds across data scales (+2.3 / +4.0 / +2.8 p at 25 / 50 / 100 demos per task) and on an expanded 15-task RoboTwin suite (+4.5p); full results are reported in Appendix B.1. Across DP backbones in simulation, Vanilla proprioception fusion is often close to or below No-Proprio. This indicates that ungrounded state vectors are not always used effectively by the policy. By anchoring state information to the corresponding visual feature location, GeoProp makes proprioception more consistently useful at the aggregate level. 4.3 Ablations and Analysis We conduct ablations and diagnostic analyses to identify which design choices drive GeoPropās gains and when the geometric grounding assumption is most beneficial. Specifically, we study i) component contributions, i) whether simpler spatial-coordinate priors can explain the improvement, and i) task-type patterns and robustness to cameraārobot calibration errors. Table 3: Component ablation on 15 MetaWorld tasks. Spa.: geometric grounding; Mot.: predictive sampling; FiLM: localized modulation. Component Vanilla Base + Mot. + FiLM Full Spa. grounding ā ā ā ā ā Mot. sampling ā ā ā ā ā FiLM modulation ā ā ā ā ā Mean (%) 61.2 64.8 66.4 66.0 68.5 Component Effectiveness. We ablate three components of GeoProp: geometric spatial grounding (Spa.), predictive kinematic sampling (Mot.), and spatially-aligned FiLM modulation. Following [33], we evaluate all variants on 15 diverse MetaWorld tasks. As shown in Table 3, replacing vector-based proprioception fusion with geometric grounding improves mean success from 61.2% to 64.8%. Adding predictive sampling and localized FiLM further improves performance to 66.4% and 66.0%, respectively, while the full model reaches 68.5%. These results indicate that geometric grounding is the primary source of improvement, with motion look-ahead and localized modulation providing complementary gains. Notably, motion sampling and FiLM modulation are super-additive (+3.7p jointly vs. +1.6 / +1.2 individually), suggesting they address complementary error modes. Split-level results are provided in Appendix B.1. Alternative Spatial-Grounding Baselines. To test whether GeoPropās gains come merely from providing the projected end-effector location, we compare against Heatmap injection and RC-PE [19]. On the same 15 MetaWorld tasks, Heatmap and RC-PE reach 61.7% and 60.8% mean success, respectively, compared with 59.5% for No-Proprio and 64.8% for GeoProp-Base. Thus, simple coordinate priors help, but do not match co-located visual feature sampling. Full split-level results are provided in Appendix B.1. Task-Type Patterns. GeoProp provides the largest gains on precision-oriented tasks involving small objects: on the four representative MetaWorld tasks Basketball, Hand Insert, Pick Place, and Sweep Into, the mean improvement reaches +22.3p averaged over both backbones (Appendix Table B.3). GeoProp underperforms when the manipulated object occludes the projected end-effector, e.g., on Box Close (RN: 48ā 40; ViT: 54ā 54), where the lid covers the gripper during closing. Parameter-Matched Control. A parameter-matched control that enlarges the proprioceptive encoder to GeoPropās parameter count recovers only +0.5p for DP and +0.1p for Ļ0 _0, indicating the gains are driven primarily by geometric alignment rather than added capacity. Calibration Robustness. GeoProp relies on cameraārobot calibration for projection, so we stress-test it under synthetic extrinsic perturbations at test time. As detailed in Appendix B.4, GeoProp degrades gradually under moderate calibration drift and remains above Vanilla and No-Proprio for translation errors up to 2.5 cm and rotation errors up to 1.5ā1.5 . 5 Real-World Experiments 5.1 Experimental Setup We evaluate GeoProp on the Mobile ALOHA bimanual platform [9], equipped with two 6-DoF arms and a front-facing Intel RealSense camera. The setup reflects a common single-view deployment scenario where proprioception must be aligned with egocentric visual observations under sensor noise, mild calibration drift, and cluttered household scenes. We consider four household manipulation tasks that cover distinct interaction patterns: i) Paper Toss, where the robot grasps a paper ball and throws it into a bin; i) Coffee Retrieval, which requires reaching, grasping, and placing a cup; i) Desk Clearing, where the robot pushes multiple objects into a target region; and iv) Table Cleaning, which requires wiping a designated area with a cloth. Each task uses 50 successful teleoperated demonstrations, and success rates are averaged over 20 evaluation trials per task. We compare No-Proprio, Vanilla, and GeoProp using Diffusion Policy with a ResNet-18 backbone, while keeping the demonstrations, camera setup, action representation, and training schedule fixed. We additionally evaluate Ļ0 _0 under the same real-world protocol to test whether the benefit of geometric grounding transfers beyond diffusion policies. Qualitative execution sequences are provided in Appendix C.1. 5.2 Performance Analysis Table 4: Real-world success rates (%) on Mobile ALOHA. All methods use the same demonstrations, camera setup, and evaluation protocol. Policy Method Paper Coffee Desk Table Avg. DP No-Proprio 25 40 35 15 28.8 Vanilla 35 35 55 30 38.8 GeoProp 50 55 60 35 50.0 Ļ0 _0 No-Proprio 45 40 60 25 42.5 Vanilla 40 45 55 35 43.8 GeoProp 55 55 65 40 53.8 Table 4 reports real-world results for both policy families. GeoProp achieves the highest average success rate, improving from 38.8% with Vanilla to 50.0% (+11.2p) and from 28.8% with No-Proprio to 50.0% (+21.2p). The gains are largest on tasks that require accurate local contact and gripperāobject alignment. For example, GeoProp improves Coffee Retrieval from 35% to 55%, corresponding to four additional successful trials out of 20, and improves Paper Toss from 35% to 50% under varied object placements. The real-world results also reveal the brittleness of ungrounded state fusion. On Coffee Retrieval, Vanilla underperforms the vision-only baseline (35% vs. 40%), suggesting that concatenated proprioceptive vectors can interfere with visual policy learning when they are not spatially aligned. By projecting the end-effector state into the image and sampling co-located visual evidence, GeoProp provides a geometric inductive bias that stabilizes hand-eye coordination under real-world noise factors such as sensor jitter, clutter, and mild calibration drift. GeoProp shows a similar trend with Ļ0 _0 under the same demonstration and evaluation protocol. It improves average success from 43.8% to 53.8% over Vanilla Ļ0 _0 (+10.0p), with per-task gains of +15/+10/+10/+5p on Paper Toss, Coffee Retrieval, Desk Clearing, and Table Cleaning, respectively. These results suggest that GeoPropās real-world benefit is not tied to a single policy family, but transfers to both diffusion-based policies and VLA models. 6 Conclusion and Limitations We introduce GeoProp, a lightweight adapter that anchors proprioception to co-located visual semantics by analytically projecting 3D end-effector states into 2D visual feature maps. Across 67 tasks, GeoProp improves Diffusion Policy by +8.7% and Ļ0 _0 by +4.0% in simulation, and yields a +10.6% real-world gain, with only 2ā3% parameter overhead. Limitations. GeoProp requires known camera intrinsics and extrinsics: it degrades gradually under mild drift (Appendix B.4) but settings without reliable calibration require re-estimating the camera pose. It grounds only the 3D end-effector position, leaving richer kinematic structures (links, joints, fingertips, contact patches) unrepresented, which limits applicability to dexterous or whole-body manipulation. Predictive sampling further assumes locally smooth end-effector motion; abrupt direction changes or contact transitions can produce off-task look-ahead points. Gains shrink or invert when the manipulated object occludes the projected pixel or when the baseline is already near ceiling; all evaluations also use a single fixed primary camera per scene, leaving multi-view fusion and moving-camera setups open. Finally, our evaluation covers three embodiments and two imitation-learning policy families; transfer to humanoids, dexterous hands, or RL, along with calibration-aware alignment and multi-point kinematic grounding, is left to future work. Acknowledgments Acknowledgments omitted for anonymous review. References [1] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. (2024) Ļ 0: A vision-language-action flow model for general robot control. corr, abs/2410.24164, 2024. doi: 10.48550. arXiv preprint ARXIV.2410.24164. Cited by: §1, §1, §1, §2.1, §4.1. [2] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al. (2022) Rt-1: robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817. Cited by: §1. [3] C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song (2023) Diffusion policy: visuomotor policy learning via action diffusion. In Robotics: Science and Systems, External Links: Link Cited by: §1, §1, §2.1, §4.1. [4] Y. Dai, J. Lee, Y. Zhang, Z. Ma, J. Yang, A. Zadeh, C. Li, N. Fazeli, and J. Chai (2025) AimBot: a simple auxiliary visual cue to enhance spatial awareness of visuomotor policies. arXiv preprint arXiv:2508.08113. Cited by: §2.2. [5] S. Dasari, O. Mees, S. Zhao, M. K. Srirama, and S. Levine (2025) The ingredients for robotic diffusion transformers. In 2025 IEEE International Conference on Robotics and Automation (ICRA), p. 15617ā15625. Cited by: §1. [6] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby (2021) An image is worth 16x16 words: transformers for image recognition at scale. In International Conference on Learning Representations, External Links: Link Cited by: §4.1. [7] S. Fan, K. Wu, Z. Che, X. Wang, D. Wu, F. Liao, N. Liu, Y. Zhang, Z. Zhao, Z. Xu, et al. (2025) XR-1: towards versatile vision-language-action models via learning unified vision-motion representations. arXiv preprint arXiv:2511.02776. Cited by: §2.2. [8] Y. Fang, K. Ranasinghe, L. Xue, H. Zhou, J. Tan, R. Xu, S. Heinecke, C. Xiong, S. Savarese, D. Szafir, et al. (2025) Robotic vla benefits from joint learning with motion image diffusion. arXiv preprint arXiv:2512.18007. Cited by: §4.1. [9] Z. Fu, T. Z. Zhao, and C. Finn (2024) Mobile aloha: learning bimanual mobile manipulation using low-cost whole-body teleoperation. In 8th Annual Conference on Robot Learning, Cited by: §1, §5.1. [10] T. Gervet, Z. Xian, N. Gkanatsios, and K. Fragkiadaki (2023) Act3D: 3d feature field transformers for multi-task robotic manipulation. In CoRL, Cited by: §1. [11] A. Goyal, V. Blukis, J. Xu, Y. Guo, Y. Chao, and D. Fox (2024) RVT-2: learning precise manipulation from few demonstrations. In RSS 2024 Workshop: Data Generation for Robotics, Cited by: §2.3. [12] K. He, X. Zhang, S. Ren, and J. Sun (2016) Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 770ā778. Cited by: §4.1. [13] S. James, Z. Ma, D. R. Arrojo, and A. J. Davison (2020) Rlbench: the robot learning benchmark & learning environment. IEEE Robotics and Automation Letters 5 (2), p. 3019ā3026. Cited by: §1, §4.1. [14] E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn (2022) Bc-z: zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning, p. 991ā1002. Cited by: §1. [15] Y. Jia, J. Liu, S. Chen, C. Gu, Z. Wang, L. Luo, X. Li, P. Wang, Z. Wang, R. Zhang, et al. (2025) Lift3D policy: lifting 2d foundation models for robust 3d robotic manipulation. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 17347ā17358. Cited by: §4.1. [16] M. J. Kim, C. Finn, and P. Liang (2025) Fine-tuning vision-language-action models: optimizing speed and success. arXiv preprint arXiv:2502.19645. Cited by: §1. [17] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, et al. (2025) OpenVLA: an open-source vision-language-action model. In Conference on Robot Learning, p. 2679ā2713. Cited by: §1, §2.1. [18] T. Kim, J. Lee, M. Koo, D. Kim, K. Lee, C. Kim, Y. Seo, and J. Shin (2025) Contrastive representation regularization for vision-language-action models. arXiv preprint arXiv:2510.01711. Cited by: §2.2. [19] B. Li, S. He, H. Xu, H. Yuan, X. Xu, Y. Zang, L. Hu, J. Yue, Z. Jiang, P. Hu, et al. (2025) Towards proprioception-aware embodied planning for dual-arm humanoid robots. arXiv preprint arXiv:2510.07882. Cited by: §2.2, §4.3. [20] F. Li, W. Song, H. Zhao, J. Wang, P. Ding, D. Wang, L. Zeng, and H. Li (2025) Spatial forcing: implicit spatial representation alignment for vision-language-action model. arXiv preprint arXiv:2510.12276. Cited by: §2.3. [21] Y. Li, Y. Deng, J. Zhang, J. Jang, M. Memmel, R. Yu, C. R. Garrett, F. Ramos, D. Fox, A. Li, A. Gupta, and A. Goyal (2025) HAMSTER: hierarchical action models for open-world robot manipulation. arXiv preprint arXiv:2502.05485. Cited by: §2.2. [22] T. P. Lillicrap (2015) Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971. Cited by: §2.1. [23] T. Lin, P. DollĆ”r, R. Girshick, K. He, B. Hariharan, and S. Belongie (2017) Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 2117ā2125. Cited by: §3. [24] J. Lu, W. Xia, Y. Wu, Z. Lu, and D. Hu (2025) When would vision-proprioception policy fail in robotic manipulation?. arXiv preprint. Cited by: §2.2. [25] A. Majumdar, K. Yadav, S. Arnaud, J. Ma, C. Chen, S. Silwal, A. Jain, V. Berges, T. Wu, J. Vakil, et al. (2023) Where are we in the search for an artificial visual cortex for embodied intelligence?. Advances in Neural Information Processing Systems 36, p. 655ā677. Cited by: §2.1. [26] O. Mees, D. Ghosh, K. Pertsch, K. Black, H. R. Walke, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, et al. (2024) Octo: an open-source generalist robot policy. In First Workshop on Vision-Language Models for Navigation and Manipulation at ICRA 2024, Cited by: §1. [27] X. Miao, H. Duan, Q. Qian, J. Wang, Y. Long, L. Shao, D. Zhao, R. Xu, and G. Zhang (2025) Towards scalable spatial intelligence via 2d-to-3d data lifting. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Cited by: §2.3. [28] Y. Mu, T. Chen, Z. Chen, S. Peng, Z. Lan, Z. Gao, Z. Liang, Q. Yu, Y. Zou, M. Xu, et al. (2025) Robotwin: dual-arm robot benchmark with generative digital twins. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 27649ā27660. Cited by: §1, §4.1. [29] D. Niu, Y. Sharma, G. Biamby, J. Quenum, Y. Bai, B. Shi, T. Darrell, and R. Herzig (2024) LLARVA: vision-action instruction tuning enhances robot learning. In Conference on Robot Learning, Cited by: §2.2. [30] Nvidia, J. Bjorck, F. Castaneda, N. Cherniadev, X. Da, R. Ding, LinxiJimFan, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y. L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y. Xie, Y. Xu, Z. Xu, S. Ye, Z. Yu, A. Zhang, H. Zhang, Y. Zhao, R. Zheng, and Y. Zhu (2025) GR00T n1: an open foundation model for generalist humanoid robots. ArXiv abs/2503.14734. External Links: Link Cited by: §1. [31] E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville (2018) FiLM: visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, Cited by: §1, §3. [32] J. Qian, B. Han, C. Shi, L. Xiao, L. Yang, S. Shi, and L. Jiang (2025) GeoPredict: leveraging predictive kinematics and 3d gaussian geometry for precise vla manipulation. arXiv preprint arXiv:2512.16811. Cited by: §2.3. [33] Q. Qian, G. Zhao, G. Zhang, J. Wang, R. Xu, J. Gao, and D. Zhao (2025) GP3: a 3d geometry-aware policy with multi-view images for robotic manipulation. arXiv preprint arXiv:2509.15733. Cited by: §2.3, §4.1, §4.3. [34] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017) Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347. Cited by: §2.1. [35] J. Shang, K. Schmeckpeper, B. B. May, M. V. Minniti, T. Kelestemur, D. Watkins, and L. Herlant (2024) Theia: distilling diverse vision foundation models for robot learning. In 8th Annual Conference on Robot Learning, External Links: Link Cited by: §2.1. [36] M. Shridhar, L. Manuelli, and D. Fox (2023) Perceiver-actor: a multi-task transformer for robotic manipulation. In Conference on Robot Learning, p. 785ā799. Cited by: §2.3. [37] K. Suzuki, S. Shimizu, and T. Ogata (2025) Proprioception enhances vision language model in generating captions and subtask segmentations for robot task. arXiv preprint arXiv:2512.20876. Cited by: §2.2. [38] T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine (2020) Meta-world: a benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on robot learning, p. 1094ā1100. Cited by: §1, §4.1. [39] Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu (2024) 3D diffusion policy: generalizable visuomotor policy learning via simple 3d representations. In ICRA 2024 Workshop on 3D Visual Representations for Robot Manipulation, Cited by: §1, §2.3. [40] G. Zhang, W. Li, Q. Qian, J. Wang, D. Zhao, S. Lu, and R. Xu (2026) On the generalization capacities of MLLMs for spatial intelligence. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §2.3. [41] J. Zhang, M. Memmel, K. Kim, D. Fox, J. Thomason, F. Ramos, E. Bıyık, A. Gupta, and A. Li (2025) PEEK: guiding and minimal image representations for zero-shot generalization of robot manipulation policies. In IEEE International Conference on Robotics and Automation, Cited by: §2.2. [42] Q. Zhang, Z. Liu, H. Fan, G. Liu, B. Zeng, and S. Liu (2025) Flowpolicy: enabling fast and robust 3d flow-based policy via consistency flow matching for robot manipulation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 14754ā14762. Cited by: §2.3. [43] J. Zhao, W. Lu, D. Zhang, Y. Liu, Y. Liang, T. Zhang, Y. Cao, J. Xie, Y. Hu, S. Wang, et al. (2025) Do you need proprioceptive states in visuomotor policies?. arXiv preprint arXiv:2509.18644. Cited by: §2.2. [44] T. Zhao, V. Kumar, S. Levine, and C. Finn (2023) Learning fine-grained bimanual manipulation with low-cost hardware. Robotics: Science and Systems XIX. Cited by: §1, §2.1. [45] D. Zheng, S. Huang, Y. Li, and L. Wang (2025) Learning from videos for 3d world: enhancing mllms with 3d vision geometry priors. arXiv preprint arXiv:2505.24625. Cited by: §2.3. [46] B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. (2023) Rt-2: vision-language-action models transfer web knowledge to robotic control. In Conference on Robot Learning, p. 2165ā2183. Cited by: §1, §2.1. Appendix A Implementation Details A.1 Training Setup and Hyperparameters Visual encoders and feature aggregation. GeoProp is backbone-agnostic and only requires spatial feature maps that preserve image-grid correspondence. To make projection-based sampling comparable across backbones, we use an FPN to produce a dense sampling map with a unified spatial resolution. For ResNet-18, we aggregate features from the four residual stages and use a stride-16 FPN output as the sampling map. For ViT-based backbones, including ViT-Base and the SigLIP encoder in Ļ0 _0, we extract intermediate transformer features, reshape patch tokens to spatial grids, and aggregate them with the same FPN interface. This produces a 14Ć1414Ć14 dense feature map for 224Ć224224Ć224 inputs. Localized FiLM modulation. The FiLM generator maps the proprioceptive state to channel-wise modulation parameters through a two-layer MLP. Following Sec. 3.2, FiLM is applied before FPN aggregation and only at the feature cell aligned with the projected end-effector location. We use the residual FiLM form modāā[]=(1+tā)āencāā[]+tā,=tā,F _mod[c]=(1+ γ _t) _enc[c]+ β _t, =c _t, while all non-aligned cells remain unchanged. The last layer of the FiLM generator is zero-initialized, so the initial modulation is an identity transform. The modulated multi-scale features are then aggregated by the FPN to form modF_mod, from which the grounded state token is extracted by bilinear sampling. Predictive kinematic extrapolation. To estimate short-horizon motion intent, we fit a quadratic polynomial to the recent end-effector positions along each Cartesian axis and extrapolate one future waypoint. The predicted waypoint is projected to the image plane and mapped to the feature grid, yielding ĀÆt+1pre q^pre_t+1. The predictive token is sampled from an FPN-aggregated but non-modulated feature map: raw=FPNā(encā=1L),pre=ā(raw,ĀÆt+1pre).F_raw=FPN(\F _enc\_ =1^L), Ļ_pre=S(F_raw, q^pre_t+1). Thus, both tokens are extracted from FPN outputs with the same spatial resolution: t Ļ_t from the modulated map modF_mod, and pre Ļ_pre from the unmodulated map rawF_raw. Implementation and training protocols. All policies are trained using standard visuomotor imitation learning setups. Diffusion Policy is optimized with Adam and a cosine schedule, while Ļ0 _0 follows its native fine-tuning protocol with AdamW and linear warm-up. Detailed architectural and optimization settings are summarized in Table A.1. Table A.1: Detailed training hyperparameters. Hyperparameter Diffusion Policy (DP) Ļ0 _0 VLA Input Resolution 224Ć224224Ć 224 224Ć224224Ć 224 Image Backbone ResNet-18 / ViT-B SigLIP-So400M Observation Horizon 2 1 Action Chunk Steps 32 50 ViT Selected Layers [2,5,8,11][2,5,8,11] [5,12,19,26][5,12,19,26] FPN Target Output cā4c4 (ResNet) / cā3c3 (ViT) cā3c3 Feature Map Size 14Ć1414Ć 14 14Ć1414Ć 14 Feature Dim 512 (ResNet) / 768 (ViT) 2048 Kinematic History 4 4 Extrapolation Order Quadratic Quadratic Training Duration 100 epochs 30k steps Optimizer Adam AdamW Base Learning Rate 1Ć10ā41Ć 10^-4 5Ć10ā55Ć 10^-5 Batch Size 64 32 Weight Decay 1Ć10ā61Ć 10^-6 0.050.05 Learning Rate Schedule Cosine annealing Cosine with warm-up A.2 Baseline Configurations and Fairness Fair comparison protocol. For all comparisons, we keep the dataset split, image preprocessing, visual/language backbone, policy architecture, action representation, rollout horizon, optimizer, and training budget fixed. The only changed factor is the mechanism used to incorporate proprioception. Vanilla proprioception fusion. For Diffusion Policy, Vanilla encodes the proprioceptive vector with a two-layer MLP and concatenates the resulting state embedding with the visual feature before policy conditioning. For Ļ0 _0, proprioception is projected into state tokens in the action expert, while image and language tokens are produced by the VLM backbone. These tokens interact through the standard transformer attention layers, following the original Ļ0 _0 proprioception-conditioning design. Unlike GeoProp, Vanilla does not impose an explicit correspondence between the robot state and image-space visual tokens. No-Proprio. No-Proprio removes the state encoder and trains the policy using only image inputs, and language inputs when available. This baseline measures whether conventional proprioception fusion provides benefit beyond visual observation alone under the same training budget. Figure A.1: Heatmap conditioning baseline. The projected end-effector coordinate is rendered as a Gaussian heatmap and concatenated with the RGB observation as additional input channels. The overlay is shown only for visualization. Heatmap conditioning. Heatmap conditioning provides a spatially explicit but input-level proprioceptive prior. We project the 3D end-effector position to the image plane and render a Gaussian heatmap centered at the projected pixel. The heatmap is concatenated with the RGB observation as additional input channels: Itā²=Concatā(It,Ht).I _t=Concat(I_t,H_t). All downstream architecture and training settings are unchanged. This baseline tests whether exposing the projected state location in image space is sufficient, without co-located feature sampling or localized modulation. RC-PE. RC-PE injects a dense robot-centric positional encoding into visual tokens. For each feature-grid location (x,y)(x,y), we compute its normalized 2D offset to the projected end-effector coordinate (uĀÆ,vĀÆ)( u, v): Īāx=xāuĀÆw/2,Īāy=yāvĀÆh/2, x= x- uw/2, y= y- vh/2, (8) where w and h are the spatial dimensions of the feature grid. The offset (Īāx,Īāy)( x, y) is embedded with a positional encoding function and added to the corresponding visual token. This baseline tests whether dense relative-coordinate cues can replace GeoPropās localized feature grounding. Appendix B Simulation Experiments Details B.1 Additional Quantitative Results This section provides additional quantitative results supporting the ablation and scalability analyses in the main paper. We include split-level component ablations, comparisons with alternative spatial-grounding baselines, and data-scaling results on RoboTwin. Component ablation. Table B.1 reports split-level ablations on 15 MetaWorld tasks. Replacing vector-based proprioception fusion with geometric spatial grounding improves the mean success rate from 61.2% to 64.8%. Adding predictive kinematic sampling and localized FiLM modulation further improves mean success to 66.4% and 66.0%, respectively. The full GeoProp model achieves the best mean success rate of 68.5%, indicating that geometric grounding is the primary source of improvement, while motion look-ahead and localized modulation provide complementary benefits. Table B.1: Split-level component ablation on 15 MetaWorld tasks. Components S.R. (%) Method Spa. Mot. FiLM Easy Med. Hard Mean Vanilla ā ā ā 58.3±3.258.3_± 3.2 64.4±10.764.4_± 10.7 62.0±4.762.0_± 4.7 61.2 GeoProp-Base ā ā ā 57.1±0.857.1_± 0.8 75.6±5.175.6_± 5.1 64.7±0.964.7_± 0.9 64.8 + Motion ā ā ā 57.2±1.657.2_± 1.6 80.0±4.580.0_± 4.5 66.1±2.466.1_± 2.4 66.4 + FiLM ā ā ā 56.9±1.256.9_± 1.2 76.8±2.376.8_± 2.3 69.3±3.869.3_± 3.8 66.0 GeoProp-Full ā ā ā 60.3±1.260.3_± 1.2 78.4±0.078.4_± 0.0 71.3±1.071.3_± 1.0 68.5 Alternative spatial-grounding baselines. To test whether GeoPropās gains come merely from exposing the projected end-effector location, we compare against Heatmap conditioning and RC-PE. Heatmap renders the projected end-effector coordinate as an image-space Gaussian prior, while RC-PE injects relative coordinate encodings into visual tokens. As shown in Table B.2, both baselines improve over No-Proprio, confirming that robot-centric spatial cues are useful. However, GeoProp-Base reaches 64.8% mean success, outperforming Heatmap by 3.1 points and RC-PE by 4.0 points. This suggests that GeoPropās gains come not only from providing location cues, but from grounding proprioception through co-located visual feature sampling. Data scaling. Beyond the 50-demo setting reported in the main text, we evaluate whether GeoPropās benefit holds across data scales on the original 7-task RoboTwin setup with Ļ0 _0 (Table B.4). GeoProp improves over Vanilla by +2.3 / +4.0 / +2.8 points at 25 / 50 / 100 demonstrations per task. The absolute gap peaks in the mid-data regime and narrows as data grows, consistent with the role of geometric grounding as an inductive bias that is most useful when the policy cannot easily infer stateāvision alignment from data alone. Table B.2: Alternative spatial-grounding baselines on 15 MetaWorld tasks. Method Easy Med. Hard Mean No-Proprio 54.6±1.254.6_± 1.2 66.0±1.766.0_± 1.7 58.0±2.858.0_± 2.8 59.5 RC-PE 56.1±1.456.1_± 1.4 67.4±0.967.4_± 0.9 61.0±1.461.0_± 1.4 60.8 Heatmap 58.3±0.058.3_± 0.0 66.9±0.766.9_± 0.7 61.0±1.461.0_± 1.4 61.7 GeoProp-Base 57.1±0.857.1_± 0.8 75.6±5.175.6_± 5.1 64.7±0.964.7_± 0.9 64.8 Small-object precision tasks. To support the claim in Sec. 4.3 that GeoPropās gains are largest on precision-oriented tasks involving small objects, Table B.3 reports per-task results on four representative MetaWorld tasks under both Diffusion Policy backbones. GeoProp improves Basketball, Hand Insert, Pick Place, and Sweep Into by +24.0p on average with ResNet-18 and +20.5p with ViT-Base, giving an overall mean gain of +22.3p. This indicates that explicit geometric grounding is especially helpful when the policy must reason about small visual targets, where vector-based proprioception often fails to localize the relevant interaction region. Table B.3: Per-task success rates (%) on four representative small-object precision MetaWorld tasks (Diffusion Policy). Ī is the absolute gain of GeoProp over Vanilla. Task ResNet-18 ViT-Base Avg Ī Vanilla GeoProp Ī Vanilla GeoProp Ī Basketball 8 38 +30 16 38 +22 +26 Hand Insert 44 52 +8 38 56 +18 +13 Pick Place 12 36 +24 2 30 +28 +26 Sweep Into 64 98 +34 76 90 +14 +24 Mean 32.0 56.0 +24.0 33.0 53.5 +20.5 +22.3 Expanded task suite. On the expanded 15-task RoboTwin setup (Table B.5), GeoProp improves Ļ0 _0 from 66.4% to 70.9% (+4.5p), comparable to the +4.0p gain on the original 7-task setup, indicating that the benefit is not specific to the original task selection. Failure and low-gain cases. The expanded RoboTwin results also show that GeoProp is not uniformly beneficial on every task. Gains are small on saturated tasks such as Adjust Bottle, Click Alarmclock, and Click Bell, where Vanilla already reaches 100% success. GeoProp can also underperform on tasks such as Pick Diverse Bottles, Place Shoe, Scan Object, and Stack Blocks Two, suggesting that end-effector-only grounding is less reliable when the projected point is visually ambiguous, occluded, or not the dominant cue for task progress. These cases are consistent with the limitation discussed in the main paper: projection provides a useful spatial prior only when the projected end-effector region remains informative for control. Parameter-matched control. A parameter-matched control that enlarges the proprioceptive encoder to match GeoPropās parameter count recovers only +0.5p for Diffusion Policy and +0.1p for Ļ0 _0, indicating that the gains are driven primarily by geometric alignment rather than additional model capacity. Table B.4: Data-scaling results on the 7-task RoboTwin setup with Ļ0 _0. Demos / Task Vanilla Ļ0 _0 GeoProp Gain 25 46.3 48.6 +2.3 50 64.0 68.0 +4.0 100 77.7 80.5 +2.8 Table B.5: Per-task success rates (%) on the expanded 15-task RoboTwin setup with Ļ0 _0 (50 demonstrations per task, 25 evaluation rollouts per task). Ī denotes the absolute gain of GeoProp over Vanilla Ļ0 _0. Task Vanilla Ļ0 _0 GeoProp Ī adjust_bottle 100 100 0 beat_block_hammer 68 92 +24 click_alarmclock 100 100 0 click_bell 100 100 0 dump_bin_bigbin 92 88 -4 move_can_pot 76 80 +4 move_playingcard_away 76 80 +4 open_laptop 72 80 +8 pick_diverse_bottles 36 32 -4 pick_dual_bottles 68 76 +8 place_dual_shoes 16 44 +28 place_shoe 64 60 -4 scan_object 28 24 -4 stack_blocks_two 68 64 -4 turn_switch 32 44 +12 Mean 66.4 70.9 +4.5 B.2 Per-Task Success Rates on Simulation Benchmarks To complement the aggregated numbers in Table 1, we report per-task success rates for all 63 Diffusion Policy simulation tasks, covering RLBench (Table B.6), RoboTwin (Table B.7), and MetaWorld-50 (Table B.8). Each entry follows the protocol in Sec. 4; per-row bests are highlighted. These tables make explicit that GeoPropās average-level gains are not concentrated on a small subset of tasks: across both ResNet-18 and ViT-Base backbones, GeoProp matches or improves over Vanilla on the large majority of individual tasks, and the largest gains are concentrated on visually demanding settings (e.g., Water Plants and Put Rubbish In Bin on RLBench; Move Can Pot, Move Playingcard Away, and Scan Object on RoboTwin; Basketball, Sweep Into, Pick Place, and Hand Insert on MetaWorld), consistent with the small-object precision analysis in Sec. 4.3. Table B.6: Per-task success rates (%) on RLBench with Diffusion Policy. Per-row bests are highlighted. Backbone Method RLBench Mean Close Box Put Rubbish In Bin Close Laptop Lid Water Plants Unplug Charger Toilet Seat Down ResNet-18 Vanilla 100 32 88 36 44 96 66.0 No-Proprio 100 44 92 28 48 96 68.0 GeoProp 100 48 100 72 56 100 79.3 ViT-Base Vanilla 100 20 92 36 44 100 65.3 No-Proprio 100 64 84 44 32 100 70.7 GeoProp 100 68 100 52 56 100 79.3 Table B.7: Per-task success rates (%) on the RoboTwin 7-task suite with Diffusion Policy. Per-row bests are highlighted. Backbone Method RoboTwin Mean Beat Block Hammer Click Alarmclock Move Can Pot Move Playingcard Away Place Shoe Scan Object Stack Blocks Two ResNet-18 Vanilla 84 84 84 36 36 20 32 53.7 No-Proprio 84 84 72 48 24 40 24 53.7 GeoProp 84 76 92 72 40 44 32 62.9 ViT-Base Vanilla 68 72 68 24 32 8 16 41.1 No-Proprio 64 60 76 36 24 24 16 42.9 GeoProp 88 76 68 52 32 36 12 52.0 Table B.8: Per-task success rates (%) on the MetaWorld-50 suite with Diffusion Policy. Per-row bests are highlighted within each backbone. Task ResNet-18 ViT-Base Vanilla No-Proprio GeoProp Vanilla No-Proprio GeoProp Button Press 100 100 100 100 100 100 Button Press Topdown 100 100 100 100 100 100 Button Press Topdown Wall 100 100 100 100 100 100 Button Press Wall 100 100 100 100 100 100 Coffee Button 100 100 100 100 100 100 Dial Turn 96 98 88 58 62 62 Door Close 100 100 100 100 100 100 Door Lock 64 68 74 68 70 54 Door Open 100 100 100 100 100 100 Door Unlock 100 98 100 98 96 100 Drawer Close 100 100 100 100 100 100 Drawer Open 74 78 82 76 72 74 Faucet Close 100 100 100 100 100 100 Faucet Open 84 82 96 78 78 94 Handle Press 100 100 100 100 100 100 Handle Pull 74 72 78 54 56 74 Handle Press Side 100 100 100 86 96 100 Handle Pull Side 70 72 74 70 56 74 Lever Pull 6 8 10 14 8 8 Plate Slide 100 100 100 100 98 98 Plate Slide Back 100 100 100 100 100 100 Plate Slide Back Side 100 100 100 100 100 100 Plate Slide Side 98 96 94 98 100 100 Reach 12 8 8 8 10 8 Reach Wall 34 38 42 46 40 46 Window Close 92 84 100 84 90 98 Window Open 46 34 46 32 26 52 Peg Unplug Side 40 28 30 22 40 40 Basketball 8 6 38 16 14 38 Bin Picking 66 70 78 66 72 64 Box Close 48 36 40 54 54 54 Coffee Pull 74 66 84 70 74 86 Coffee Push 82 84 90 52 46 74 Hammer 94 94 96 92 90 92 Peg Insert Side 26 36 38 8 20 24 Push Wall 44 44 88 48 64 92 Soccer 14 12 16 12 12 6 Sweep 20 32 76 20 30 68 Sweep Into 64 82 98 76 80 90 Assembly 96 92 98 50 48 82 Hand Insert 44 40 52 38 24 56 Pick Out Of Hole 42 56 50 32 30 42 Pick Place 12 16 36 2 6 30 Push 34 50 66 34 30 58 Push Back 48 52 52 48 46 48 Pick Place Wall 40 68 86 70 36 76 Stick Pull 74 80 86 74 58 72 Stick Push 100 100 100 100 100 100 Shelf Place 46 46 60 14 14 30 Disassemble 72 78 80 60 56 68 Mean (50) 68.8 70.1 76.6 64.6 64.0 72.6 B.3 Qualitative Grounding Visualization We provide qualitative visualizations to illustrate how GeoProp grounds proprioceptive state in visual feature space. Fig. B.1 shows representative simulation rollouts with projected end-effector locations and their corresponding image-space grounding regions. The red dot denotes the 2D projection of the 3D end-effector position, and the orange box visualizes the image-space region associated with the feature-grid cell used by GeoProp. These examples show that the projected robot state remains spatially aligned with task-relevant interaction regions across different simulated embodiments and manipulation tasks. Fig. B.2 further compares attention maps between Vanilla proprioception fusion and GeoProp across multiple RoboTwin tasks and transformer layers. Vanilla fusion often produces diffuse activations over background regions, whereas GeoProp yields more localized responses around the end-effector and task-relevant objects. This supports our central hypothesis that explicit geometric grounding helps the policy associate proprioceptive state with co-located visual evidence, rather than treating state as a disjoint global vector. Figure B.1: Simulation grounding visualization. Rows show representative rollouts from simulation tasks. Red dots indicate projected 2D end-effector positions, and orange boxes show the corresponding image-space regions associated with GeoPropās feature-grid grounding. Figure B.2: Additional attention-map visualizations. We compare Vanilla proprioception fusion and GeoProp across multiple RoboTwin tasks and transformer layers. GeoProp produces more localized activations around task-relevant end-effector/object regions, while Vanilla fusion often attends diffusely to background regions. B.4 Calibration Robustness GeoProp relies on cameraārobot calibration to project 3D end-effector states into the image plane. To characterize this assumption, we evaluate GeoProp under synthetic test-time perturbations to the camera extrinsics. All policies are trained under the same setting, and perturbations are applied only during evaluation. We compare GeoProp with Vanilla proprioception fusion and No-Proprio using the same 15-task MetaWorld protocol as in the ablation study. For translation drift, we perturb the cameraārobot translation along each Cartesian axis independently. For rotation drift, we perturb the camera orientation around roll, pitch, and yaw. The dashed horizontal lines denote the mean performance of Vanilla and No-Proprio baselines, which are unaffected by projection noise. As shown in Fig. B.3, GeoProp degrades gradually under increasing calibration error. It remains above both baselines under mild-to-moderate translation drift up to 2.5 cm and rotation drift up to 1.5ā1.5 . The degradation is axis-dependent: translation errors along the depth-related axis and rotation errors around pitch/yaw cause larger drops, while roll drift is substantially less harmful. These results show that GeoProp benefits from accurate geometric grounding, but remains usable under moderate calibration noise. (a) Translation drift. (b) Rotation drift. Figure B.3: Calibration robustness under synthetic extrinsic drift. GeoProp remains above Vanilla and No-Proprio under mild-to-moderate perturbations, while performance gradually decreases as projection error grows. To visualize how calibration drift affects GeoPropās spatial grounding, we further show projected locations and sampled grounding regions under perturbed extrinsics. In Figs. B.4 and B.5, the red dot indicates the end-effector projection under the original extrinsics, while the blue dot indicates the projection under perturbed extrinsics. The orange box denotes the grounding region from the original projection, the cyan box denotes the grounding region from the perturbed projection, and the purple region shows their overlap. As calibration error increases, the perturbed grounding region gradually shifts away from the original one, explaining the performance degradation observed in Fig. B.3. Figure B.4: Qualitative visualization under translation drift. Red and blue dots show unperturbed and perturbed end-effector projections, respectively. Orange and cyan boxes show the corresponding grounding regions, with overlap highlighted in purple. Figure B.5: Qualitative visualization under rotation drift. Red and blue dots show unperturbed and perturbed end-effector projections, respectively. Orange and cyan boxes show the corresponding grounding regions, with overlap highlighted in purple. Appendix C Real-World Experiments Details C.1 Qualitative Execution Sequences Fig. C.1 shows representative real-world execution sequences on the Mobile ALOHA platform. Each row corresponds to one household manipulation task and visualizes task progress from the initial state to the final state. The red dot denotes the projected 2D end-effector position, and the orange box indicates the image-space region corresponding to GeoPropās feature-grid grounding. These examples illustrate that GeoProp can maintain spatial correspondence between proprioceptive state and task-relevant visual evidence across grasping, object clearing, and wiping behaviors. Figure C.1: Real-world execution sequences on Mobile ALOHA. Rows show task progress from the initial state to the final state for Paper Toss, Coffee Retrieval, Desk Clearing, and Table Cleaning. Red dots indicate projected end-effector positions, and orange boxes show the corresponding grounded visual regions used by GeoProp.