Paper deep dive
Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation
Xiwen Chen, Zelin Li, Zhiruo Zhou, Huiming Chen, Chenwei Wang, Xiaojun Zhu
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text. For the evaluated Bridge/WidowX and physical pick-place deployments, an explicit execution contract assigns PICK to source-conditioned OFG and PLACE to Pointing, providing direct stage-aligned spatial targets. Pointing-VLA achieves SOTA performance on Bridge/WidowX, averaging 72.9\% across the evaluated four-task set without Bridge-specific finetuning under collision-enabled CuRobo execution. Pointing and OFG show complementary strengths across native and cross-dataset evaluations. The OFG/contact readout transfers to NORA-1.5, preserving or improving success while reducing recorded controller time by more than 20$\times$; typed heads are also 6.68--6.90$\times$ faster than Embodied-R1 text decoding on a shared external suite. When integrated as spatial guidance for a $\pi_{0.5}$ action policy, Pointing-VLA raises autonomous real-robot success from 52.7\% to 80.7\% across three visual contexts. These results establish typed spatial readouts as an efficient, inspectable interface between embodied reasoning and robot execution.
Tags
Links
- Source: https://arxiv.org/abs/2608.23138v1
- Canonical: https://arxiv.org/abs/2608.23138v1
Trouble viewing inline? Open PDF directly →
Full Text
37,328 characters extracted from source content.
Expand or collapse full text
Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation Xiwen Chen Zelin Li Zhiruo Zhou Huiming Chen Chenwei Wang Xiaojun Zhu Abstract Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text. For the evaluated Bridge/WidowX and physical pick-place deployments, an explicit execution contract assigns PICK to source-conditioned OFG and PLACE to Pointing, providing direct stage-aligned spatial targets. Pointing-VLA achieves SOTA performance on Bridge/WidowX, averaging 72.9% across the evaluated four-task set without Bridge-specific finetuning under collision-enabled CuRobo execution. Pointing and OFG show complementary strengths across native and cross-dataset evaluations. The OFG/contact readout transfers to NORA-1.5, preserving or improving success while reducing recorded controller time by more than 20×; typed heads are also 6.68–6.90× faster than Embodied-R1 text decoding on a shared external suite. When integrated as spatial guidance for a π0.5 _0.5 action policy, Pointing-VLA raises autonomous real-robot success from 52.7% to 80.7% across three visual contexts. These results establish typed spatial readouts as an efficient, inspectable interface between embodied reasoning and robot execution. Introduction Large vision-language-action (VLA) models can connect observations, language instructions, and robot behavior, but their output interface remains a bottleneck. Many manipulation subtasks first require an explicit spatial commitment: the point to touch, the functional part to use, or the short visual trace an executor should follow. Autoregressive text coordinates are convenient for prompting, yet they make geometry an afterthought: the model must spend tokens spelling out numbers, the evaluator must parse strings back into pixels, and dense part-level evidence is discarded before execution. Direct action tokens avoid parsing, but they also hide the spatial rationale that a robot stack or human auditor may need to inspect. Figure 1 separates two interface failures observed in representative simulator cases: text generation can fail before producing valid geometry, and a syntactically valid coordinate can still be spatially wrong. Parsing is therefore an additional execution dependency rather than a guarantee of correct grounding. Figure 1: Representative text-coordinate failures. Left: generation yields no parseable coordinate. Right: [222, 172] is parseable but lies outside the target region. The central claim of this paper is that embodied grounding should be treated as typed interface prediction rather than text-coordinate generation. Referring points, functional regions, and visual trajectories are related but not identical targets: a point is sparse, an affordance is often a dense part-level region, and a trajectory is temporally structured. Collapsing them into one output distribution can create negative interference even when they share a visual-language backbone, while serializing all of them as text makes downstream execution slower and less stable. This formulation is grounded in the vision-for-action and population-readout view of computation: downstream readouts decode spatial variables from distributed internal activity instead of serializing them through the system that represents them (7; 1; 5). Pointing-VLA realizes this principle with separate point, affordance-heatmap, and trajectory heads that emit each target in the geometry consumed by robot-side execution. We propose Pointing-VLA, a lightweight typed grounding framework built on an embodied vision-language backbone. Rather than asking the backbone to say coordinates, Pointing-VLA reads spatial intent from multimodal hidden states. A compact adapter and learned queries feed point, object-functional grounding (OFG), and visual trajectory generation (VTG) decoders. Deployment contracts bind each structured execution slot to the geometry it requires; the primary pick-place scaffold uses source-conditioned OFG for PICK and Pointing for PLACE. The resulting spatial targets connect embodied reasoning to downstream wrappers, planners, and executors through deterministic, stage-aligned contracts. Experiments evaluate the interface from native spatial grounding through head specialization, cross-backbone transfer, and robot deployment. On Bridge/WidowX, the fixed OFG PICK and Pointing PLACE composition achieves SOTA performance, averaging 72.9% across the evaluated four-task set without Bridge-specific finetuning under collision-enabled CuRobo execution. Cross-dataset results reveal complementary Pointing and OFG geometries, while NORA-1.5 transfer preserves or improves success with more than 20× shorter recorded controller time. Real-robot deployment completes the evidence chain from hidden-state readout to physical execution. Our contributions are: • We formulate embodied grounding as typed interface prediction, replacing serialized text coordinates with geometry-specific robot-facing readouts. • We introduce Pointing, OFG, and VTG hidden-state readouts with explicit execution contracts, including a fixed source-conditioned OFG-PICK/Pointing-PLACE deployment that aligns each stage with its required geometry. • We validate the interface through native and cross-dataset grounding, runtime and cross-backbone transfer, Bridge/WidowX evaluation, and autonomous real-robot deployment. Figure 2: Pointing-VLA architecture and typed interfaces. Embodied-R1 hidden states feed point, OFG/contact-heatmap, and VTG-waypoint heads; fixed PICK/PLACE slots use source-conditioned OFG/Pointing, and a wrapper maps image geometry to robot coordinates. Figure 3: Stage-wise training. Geometry warm-up trains the adapter, eight queries, and geometry head; LoRA specialization updates only low-rank backbone parameters; task continuation updates only the required subset. Snowflakes and flames mark frozen and trainable modules. Related Work Generalist robot policies and VLA systems adapt large pretrained models to robot action spaces, from embodied multimodal models and Robotics Transformers to cross-embodiment datasets and open robot policies (6; 3; 21; 12; 11; 9; 2). This line of work has improved policy learning, action parameterization, and inference throughput, but it mainly asks how perception and language should become actions. Pointing-VLA asks a complementary interface question: what spatial representation should be exposed before final execution? Spatial structure has long been useful for manipulation. Neuroscience distinguishes object recognition from action-facing visuomotor representations, including posterior-parietal intention maps and competing affordance-like action candidates (7; 1; 5). Robot learning makes a compatible engineering move: Transporter Networks preserve spatial feature maps for pick-and-place displacements, CLIPort combines semantic “what” and spatial “where” pathways, and PerAct discretizes 3D voxel actions for language-conditioned manipulation (19; 14; 15). Recent grounding-centric VLA work extends this idea to embodied reasoning: RoboPoint predicts spatial affordance keypoints from language instructions (16), FSD studies seeing-to-doing affordance grounding (17), and Embodied-R1 formalizes embodied pointing across REG, relational-region grounding, OFG, and VTG (18). Affordance datasets such as AGD20K further emphasize that actionable regions are often object parts rather than whole-object centers (10). These systems show that spatial outputs can bridge VLM reasoning and robot behavior, but they leave open which interface an executor should consume. Pointing-VLA studies this interface choice directly: sparse referring points, dense functional heatmaps, and visual waypoint traces are decoded in their native geometries rather than serialized as text coordinates. This interface complements action-sequence policies such as Diffusion Policy and ACT (4; 20) by exposing typed spatial targets before low-level action generation. In the evaluated pick-place systems, the execution phase itself supplies a stable contract: functional contact is decoded for PICK and a compact target point is decoded for PLACE. Method Hidden-State Spatial Readout Pointing-VLA operationalizes typed interface prediction by decoding geometry-specific robot targets directly from shared multimodal hidden states. Figure 2 summarizes this hidden-state readout and its execution-facing spatial outputs. Let e∈pt,ofg,vtge∈\pt,ofg,vtg\ index a spatial expert and let Hmme=hiei=1NH_m^e=\h_i^e\_i=1^N denote its Embodied-R1 multimodal states. Pointing-VLA reuses one backbone architecture, base weights, and Grounding Adapter body, while each specialist retains its own LoRA, learned queries, and decoder state. For the point and OFG experts, the shared adapter body projects the sequence into a grounding space and reads it with expert-specific task-conditioned queries: Xe=WheHmme+Pe,Ae=softmax((QeWQe)(XeWKe)⊤d),Ge=FFNe(AeXeWVe),ge=Pool(Ge), splitX_e&=W_h^eH_m^e+P_e,\\ A_e&=softmax\! ( (Q_eW_Q^e)(X_eW_K^e) d ),\\ G_e&=FFN_e\! (A_eX_eW_V^e ), g_e=Pool(G_e), split where PeP_e supplies sequence position, GeG_e preserves a fixed set of grounded tokens, and geg_e is its global summary. The point and OFG experts use eight grounding queries. The strongest VTG expert instead lets learned temporal queries attend directly to the full multimodal sequence, avoiding a fixed grounding-token bottleneck. Spatial Interfaces and Deployment Contracts Learned Spatial Heads. The Pointing decoder uses a learned pointer query qptq_pt to select the grounded evidence needed for one image location: Apt=softmax((qptWQ)(GptWK)⊤d),zpt=Apt(GptWV),(x^,y^)=σ(WpLN(zpt)+bp). splitA_pt&=softmax\! ( (q_ptW_Q)(G_ptW_K) d ),\\ z_pt&=A_pt(G_ptW_V), ( x, y)=σ\! (W_pLN(z_pt)+b_p ). split Thus (x^,y^)∈[0,1]2( x, y)∈[0,1]^2 is emitted directly by the learned head, without generating or parsing coordinate text. For an image of width W and height H, the corresponding pixel is (u^,v^)=((W−1)x^,(H−1)y^).( u, v)= ((W-1) x,(H-1) y ). The prediction is evaluated with point-in-bbox (PIB) or point-in-mask (PIM). OFG preserves spatial extent instead of collapsing the source to one token. Let FvisF_vis be the visual tower’s 2D feature map. The global grounded summary modulates this map through FiLM before a convolutional heatmap decoder (10): Fofg′=γ(gofg)⊙Fvis+β(gofg),H^ofg=fhm(Fofg′),(i∗,j∗)=argmaxi,jH^ofg[i,j],p^ofg=(j∗Wo−1,i∗Ho−1). splitF _ofg&=γ(g_ofg) F_vis+β(g_ofg),\\ H_ofg&=f_hm(F _ofg),\\ (i^*,j^*)&= _i,j H_ofg[i,j],\\ p_ofg&= ( j^*W_o-1, i^*H_o-1 ). split Here H^ofg H_ofg contains heatmap logits, and its normalized peak is the execution-facing functional contact. For VTG, learned waypoint queries qtτq_t^τ with temporal encodings attend to the multimodal sequence in order: ztτ=Attn(qtτ,Hmmvtg),τ^t=σ(Wτztτ+bτ),t=1,…,T. splitz_t^τ&=Attn(q_t^τ,H_m^vtg),\\ τ_t&=σ(W_τz_t^τ+b_τ), t=1,…,T. split We use T=8T=8 normalized image-space waypoints and report RMSE, average displacement error (ADE), and final displacement error (FDE). These waypoints describe a visual trace rather than low-level robot actions. Structured PICK/PLACE Contracts. Training follows target geometry: point, region, and visual-trace samples supervise the corresponding learned heads. For pick-place deployment, the scaffold constructs slot-specific queries from instruction ℓ , source description csrcc_src, and goal description cgoalc_goal: qpick=pick(ℓ,csrc),qplace=place(ℓ,cgoal),y^pick=Fofg(Hmm,G,qpick),y^place=Fpt(Hmm,G,qplace). array[]rclq_pick&=&T_pick( ,c_src),\\ q_place&=&T_place( ,c_goal),\\ y_pick&=&F_ofg(H_m,G,q_pick),\\ y_place&=&F_pt(H_m,G,q_place). array This assignment is deterministic: source-conditioned OFG emits the functional PICK contact, while Pointing emits the PLACE target. VTG serves annotated waypoint requests; the reported pick-place deployments use the fixed OFG-PICK/Pointing-PLACE contract. An external wrapper converts image-space outputs to robot-frame targets. With calibrated depth D, camera intrinsics K, and camera-to-robot transform Tr←cT_r← c, this boundary can be written as p~=[u^,v^,1]⊤,Xc=D(u^,v^)K−1p~,X¯r=Tr←cX¯c, split p&=[ u, v,1] ,\\ X_c&=D( u, v)K^-1 p,\\ X_r&=T_r← c X_c, split where X¯ X denotes homogeneous coordinates. The existing executor then consumes XrX_r; the learned heads remain image-space predictors. Training Objective The experts use a weighted spatial objective: ℒ=λptℒpt+λboxℒbox+λofgℒofg+λvtgℒvtg+λauxℒaux.L= _ptL_pt+ _boxL_box+ _ofgL_ofg+ _vtgL_vtg+ _auxL_aux. Point and trajectory coordinates use SmoothL1, while OFG mask or keypoint supervision is converted to a dense target H∗H^*: ℒpt=1K∑kSmoothL1(p^k,pk),ℒofg=1|Ω|∑u∈Ωℓhm(H^ofg(u),H∗(u)),ℒvtg=1T∑tSmoothL1(τ^t,τt). array[]rclL_pt&=& 1K _kSmoothL1( p_k,p_k),\\ L_ofg&=& 1| | _u∈ _hm( H_ofg(u),H^*(u)),\\ L_vtg&=& 1T _tSmoothL1( τ_t, _t). array Here ℓhm _hm is mean-squared error for Gaussian keypoint targets and class-balanced binary cross-entropy with logits for dense affordance masks. Language-modeling and legacy reconstruction terms are inactive; optimization focuses on adapters, heads, and LoRA specialization (8). For source-conditioned OFG specialization, we keep the Embodied-R1 base, Pointing/VTG query and decoder modules, and executor frozen while jointly updating the OFG LoRA, shared Grounding Adapter body, OFG learned query, and heatmap decoder. Dense source-mask reconstruction is combined with source-over-distractor ranking and distractor suppression. This adaptation teaches OFG to preserve functional contact while resolving the source instance directly in the PICK readout. The final pointing model additionally applies GRPO-style group-normalized refinement. With policy center μ=p^μ= p and learned σ=exp(s)σ= (s), samples are ai=clip(μ+σϵi,0,1),ϵi∼(0,I).a_i=clip(μ+σ _i,0,1), _i (0,I). For target box B∗B^*, reward and group-normalized advantage are ri=1,ai∈B∗,0.5max(0,1−di/ρ),ai∉B∗,Ai=ri−meanj(rj)stdj(rj)+ϵ. array[]rclr_i&=& \ array[]l1,&a_i∈ B^*,\\ 0.5 (0,1-d_i/ρ),&a_i∉ B^*, array .\\[2.0pt] A_i&=& r_i-mean_j(r_j)std_j(r_j)+ε. array With πi=π(ai∣μ,σ) _i=π(a_i μ,σ) and Gaussian reference KL DrefD_ref, the objective is ℒGRPO=−1M∑i=1Msg(Ai)logπi+λsftSmoothL1(μ,p∗)+λklDref+λσℒσ. array[]rclL_GRPO&=&- 1M _i=1^Msg(A_i) _i\\ &&+ _sftSmoothL1(μ,p^*)\\ &&+ _klD_ref+ _σL_σ. array Here sgsg stops advantage gradients and p∗p^* is the supervised point target; refinement acts directly on the continuous spatial head. Experiments Evaluation Questions and Protocol Four questions structure the evaluation. Q1: Geometry. Do typed heads preserve task-appropriate geometry? Q2: Phase specialization. Does source-conditioned OFG resolve PICK source selection while preserving the benefits of a fixed OFG-PICK/Pointing-PLACE contract? Q3: Transfer and efficiency. Does the readout transfer across datasets and VLA backbones while reducing inference cost? Q4: Deployment. Do the gains persist in simulation and real-robot deployment? Native and external evaluations answer Q1; source-conditioning and interface-composition ablations answer Q2. These controlled comparisons isolate component-level grounding quality, while the deployment studies measure complete robot-task success. Cross-dataset and cross-backbone transfer plus known-head runtime address Q3; controlled simulation and real-robot studies on shared task sets and executor families address Q4. Evaluation Protocol and Metrics. We evaluate each output at the level of the interface it is designed to serve. Pointing and OFG predictions are converted to image coordinates and scored by point-in-mask (PIM) or point-in-box (PIB) sample success; VTG is measured by normalized trajectory root mean squared error (RMSE), average displacement error (ADE), and final displacement error (FDE). Robot studies require full task completion. Within each experiment, all compared methods use the same samples or episode identities, ensuring that measured differences arise from the spatial interface or adaptation strategy. Experiments used 40-GB NVIDIA A100 and dual 48-GB RTX A6000 GPUs; complete hardware, software, and seed records appear in the Supplement. Native Spatial-Head Results Table 1 answers the first question. Pointing-VLA matches Embodied-R1 on sparse referring-point grounding at 64.3%, while OFG raises full Part-Affordance-2K accuracy from 40.9% to 57.3%. This 16.4-point improvement demonstrates geometry specialization: Pointing preserves sparse target accuracy, whereas OFG substantially strengthens part-level functional grounding by retaining the spatial extent required for contact prediction. Model VABench-P acc. ↑ Part-Afford.2K acc. ↑ GPT-4o 9.3 10.2 ASMv2 10.1 13.8 RoboBrain 7.0 25.3 Qwen2.5-VL 9.9 23.4 RoboPoint (16) 19.1 27.6 FSD (17) 61.8 9.6 Embodied-SFT (18) 50.5 39.4 Embodied-R1 (18) 63.7 40.9 Pointing-VLA (ours) 64.3 57.3 Table 1: Native point-in-region grounding accuracy (%). Pointing-VLA uses its Pointing head for VABench-P and its OFG head for the full Part-Affordance-2K run. Sequential geometry. VTG completes the native evaluation for ordered spatial outputs. On 300 VABench-V examples, the trajectory readout records 0.1042 RMSE, 0.1368 ADE, and 0.1493 FDE in normalized image coordinates. These measures evaluate the waypoint sequence and its terminal target rather than collapsing trajectory quality into a single point-in-region score. Mechanism and Fixed-Contract Ablations PICK requires both source-instance disambiguation and a contact location within the selected object. Source-conditioned OFG preserves this dense functional extent and exposes its peak as the grasp cue. PLACE instead requires a compact target in the destination region, for which Pointing provides a direct normalized coordinate. The fixed OFG-PICK/Pointing-PLACE contract therefore aligns each execution stage with the geometry it requires, without introducing an inference-time selector. (a) Spatial-head configuration Component Settings PIB Query tokens 4 / 8 / 16 44.7 / 50.3 / 45.7 LoRA rank none / 16 / 32 / 64 45.3 / 47.3 / 53.3 / 50.3 Refinement SFT / GRPO 61.3 / 64.3 (b) Interface-composition ablation Configuration Spoon on towel Carrot on plate Stack cube Eggplant basket Avg. Attention → Attention 16.7 41.7 100.0 62.5 55.2 Attention → Pointing 12.5 50.0 100.0 70.8 58.3 OFG → Attention 16.7 54.2 16.7 58.3 36.5 OFG → Pointing 50.0 50.0 100.0 91.7 72.9 Table 2: Mechanism and fixed-contract results. (a) VABench-P PIB (%) across spatial-head configurations. (b) Bridge/WidowX success (%) for fixed PICK/PLACE interfaces (24 episodes/task); the final row is the deployed OFG-PICK/Pointing-PLACE contract with collision-enabled CuRobo. Mechanism analysis. Panel (a) identifies the effective capacity regime of the spatial readout. Eight queries and rank-32 LoRA produce the strongest PIB within their respective sweeps, while GRPO refinement raises PIB from 61.3% to 64.3%. The non-monotonic query and rank trends demonstrate that performance is governed by the spatial-readout design rather than parameter count alone. Fixed-contract validation. Panel (b) validates the phase-aligned assignment: the deployed contract completes 70/96 episodes (72.9%) under collision-enabled CuRobo execution, outperforming the strongest alternative composition by 14.6 percentage points while completing all Stack episodes and 22/24 Eggplant episodes. The controlled frozen multicolor Stack evaluation further verifies this mechanism. Source-conditioned OFG selects the instructed source in all 48 online cases and completes 43/48 episodes, versus 40/48 for Attention PICK under identical Pointing PLACE targets. Perfect source selection establishes OFG as a direct PICK readout that resolves the instructed source while preserving functional contact grounding. Figure 4: Representative successful real-robot rollouts across three visual contexts, each ordered from initialization to stable release. External Diagnostics and Cross-Backbone Transfer Cross-Dataset Geometry. Table 3 asks whether performance follows the geometry exposed by each interface. These diagnostics standardize outputs as points and score PIM/PIB, isolating spatial grounding from downstream control. Pointing is strongest on RefCOCO and RoboAff, whereas OFG is strongest on AGD20K and raises full Part-Affordance-2K PIM from 40.9% with Embodied-R1 text generation to 57.3%. The crossover supports expert-specific output geometries rather than one serialized coordinate interface. The Supplement provides qualitative OFG heatmaps. Pointing, OFG, and Embodied-R1 text generation are evaluated on the same cross-dataset samples. Each trained head emits one point, while text generation can emit multiple points; sample-level success therefore keeps the comparison independent of candidate count. Dataset composition and deterministic split construction are detailed in the Supplement and released evaluation code. This shared point protocol enables a controlled cross-dataset comparison of spatial interfaces under identical output and evaluation rules. Suite Metric Pointing-VLA Point Pointing-VLA OFG Embodied-R1 Text Generation RefCOCO PIM 40.7% 22.2% 10.6% RefCOCO PIB 56.1% 36.0% 25.9% AGD20K PIM 53.9% 64.4% 39.9% Part-Aff.2K PIM – 57.3% 40.9% RoboAff. PIM 11.5% 9.8% 2.4% Table 3: Cross-dataset PIM/PIB accuracy (%). System Total (s) Per sample (s) Speedup OFG head 1,274.61 0.434 6.90× Pointing head 1,316.29 0.448 6.68× Embodied-R1 text decoder 8,798.06 2.993 1.00× Method Pose Success Time (s) Speedup NORA base Horizontal 100.0 107.11 1.0× + OFG/contact Horizontal 100.0 4.85 22.1× NORA base Laid vert. 89.0 111.57 1.0× + OFG/contact Laid vert. 95.0 5.44 20.5× Table 4: Known-head runtime (top) and NORA-1.5 transfer (bottom). Interface-Level Runtime. Table 4 measures inference from multimodal hidden states to typed spatial outputs under the shared external protocol. The geometric readouts are 6.68–6.90× faster than autoregressive text decoding, establishing the interface-level efficiency gained by eliminating coordinate serialization, token-by-token generation, and post-hoc parsing. NORA-1.5 Transfer. Table 4 shows that the frozen NORA-1.5 backbone with the transferred OFG/contact readout preserves horizontal-pose success, improves laid-vertical success from 89.0% to 95.0%, and cuts recorded controller time by more than 20× under the shared wrapper. System Spoon Carrot Stack Eggpl. Avg. Octo-S 47.2 9.7 4.2 56.9 30.0 SpatialVLA (FT) 16.7 25.0 29.2 100.0 42.7 SoFar 55.5 56.9 62.5 40.2 53.8 MemoryVLA 75.0 75.0 37.5 100.0 71.9 Embodied-R1 + CuRobo 20.8 45.8 62.5 79.2 52.1 Pointing-VLA + CuRobo 50.0 50.0 100.0 91.7 72.9 Table 5: Bridge/WidowX success (%; 24 episodes/task for Embodied-R1 and Pointing-VLA). Pointing-VLA uses fixed OFG-PICK/Pointing-PLACE with collision-enabled CuRobo. Deployment Studies Table 5 evaluates four Bridge/WidowX manipulation structures. The Pointing-VLA condition contains 24 episodes per task, and success requires complete task execution. Its fixed OFG PICK and Pointing PLACE composition reaches 72.9% with collision checking active in every rollout; CuRobo completes all planned motion segments without planner failure or fallback. The tasks stress distinct spatial bottlenecks: thin-object contact for Spoon, surface targeting for Carrot, source-instance disambiguation for Stack, and container placement for Eggplant. The source-conditioned OFG expert preserves 100.0% Stack success as the direct PICK readout, while Pointing supplies an explicit placement target. Together, these results demonstrate reliable four-task deployment through a collision-aware motion-planning backend. Real-Robot Deployment Scene success π0.5 _0.5 π0.5 _0.5 + P-VLA Δ No distractor 20/50 (40.0) 36/50 (72.0) +32.0 Yellow cylinder 26/50 (52.0) 42/50 (84.0) +32.0 Red cylinder 33/50 (66.0) 43/50 (86.0) +20.0 All three scenes 79/150 (52.7) 121/150 (80.7) +28.0 Failure stage Grasp Transfer Tray Uncertain Count (/150) →1647\!→\!16 →98\!→\!9 →413\!→\!4 →03\!→\!0 Δ (p) -20.7 +0.7 -6.0 -2.0 Table 6: Autonomous PiPER outcomes. Success is count/trials (%); failures show baseline→ -VLA for pre-lift grasp, post-grasp transfer/drop, tray arrival without upright release, and uncertain cases. We deploy Pointing-VLA on an AgileX PiPER and compare it with a π0.5 _0.5 baseline (13) through the same dual-camera perception-to-control stack. Pointing-VLA retains π0.5 _0.5 as the action-generating policy and inserts typed spatial guidance before execution. During PICK, the OFG head predicts an affordance heatmap over the lower rectangular part of the white composite object, and the external wrapper converts its peak into a robot-frame grasp cue. During PLACE, the Pointing head predicts a normalized image-space placement coordinate inside the green tray, which the wrapper converts into the placement target. The same deterministic OFG-PICK/Pointing-PLACE contract is used throughout all three visual contexts: no distractor, a yellow-cylinder distractor, and a red-cylinder distractor. We evaluate both systems over 50 trials per context. The reported endpoint is full task completion: the designated part must be grasped and lifted, positioned upright in the tray, and stably released. Across all three visual contexts, typed spatial guidance consistently improves autonomous completion over the action-policy baseline. Failure-stage analysis confirms that the gains concentrate at the intended control boundaries: pre-lift grasp failures fall from 47 to 16, and tray-arrival failures fall from 13 to 4. Post-grasp transfer remains the primary residual bottleneck. These reductions verify the stage-aligned contribution of OFG-PICK and Pointing-PLACE guidance; Figure 4 traces representative successes from approach through stable release. Discussion and Conclusion Pointing-VLA treats referring points, functional regions, and waypoint traces as distinct robot-facing geometries. Embodied-R1 provides shared multimodal states, while lightweight geometry heads expose normalized points, heatmaps, and trajectories directly. Explicit deployment contracts align these outputs with execution stages and preserve the geometry required by each robot-facing slot. Together, the native, cross-dataset, runtime, transfer, and deployment results establish geometry-appropriate typed readouts as an efficient and inspectable interface between embodied VLM reasoning and robot execution. The fixed pick/place scaffold provides a controlled deployment setting that verifies the effectiveness of typed spatial guidance. Task-dependent variation, most visible on spoon-on-towel, identifies closed-loop geometry-aware execution as the next leverage point. Future work will extend this interface with visual correction, broader structured task decompositions, and cross-robot transfer. References Andersen and Buneo (2002) R. A. Andersen and C. A. Buneo Intentional maps in posterior parietal cortex. Annual Review of Neuroscience 25, p. 189–220. External Links: Document, Link Cited by: Introduction, Related Work. Black et al. (2025) K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky π0 _0: a vision-language-action flow model for general robot control. In Proceedings of Robotics: Science and Systems, External Links: Document, Link Cited by: Related Work. Brohan et al. (2023) A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich RT-1: robotics transformer for real-world control at scale. In Proceedings of Robotics: Science and Systems, External Links: Document, Link Cited by: Related Work. Chi et al. (2023) C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song Diffusion policy: visuomotor policy learning via action diffusion. In Robotics: Science and Systems, External Links: Document, Link Cited by: Related Work. Cisek (2007) P. Cisek Cortical mechanisms of action selection: the affordance competition hypothesis. Philosophical Transactions of the Royal Society B: Biological Sciences 362 (1485), p. 1585–1599. External Links: Document, Link Cited by: Introduction, Related Work. Driess et al. (2023) D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence PaLM-E: an embodied multimodal language model. In Proceedings of the 40th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 202, p. 8469–8488. External Links: Link Cited by: Related Work. Goodale and Milner (1992) M. A. Goodale and A. D. Milner Separate visual pathways for perception and action. Trends in Neurosciences 15 (1), p. 20–25. External Links: Document, Link Cited by: Introduction, Related Work. Hu et al. (2021) E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen LoRA: low-rank adaptation of large language models. External Links: 2106.09685, Link Cited by: Training Objective. Kim et al. (2025) M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn OpenVLA: an open-source vision-language-action model. In Proceedings of The 8th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 270, p. 2679–2713. External Links: Link Cited by: Related Work. Luo et al. (2022) H. Luo, W. Zhai, J. Zhang, Y. Cao, and D. Tao Learning affordance grounding from exocentric images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 2252–2261. External Links: Document, Link Cited by: Related Work, Learned Spatial Heads.. Octo Model Team et al. (2024) Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y. L. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine Octo: an open-source generalist robot policy. In Robotics: Science and Systems, External Links: Document, Link Cited by: Related Work. Open X-Embodiment Collaboration et al. (2024) Open X-Embodiment Collaboration, A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, et al. Open X-embodiment: robotic learning datasets and RT-X models. In Proceedings of the IEEE International Conference on Robotics and Automation, External Links: Document, Link Cited by: Related Work. Physical Intelligence et al. (2025) Physical Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky π0.5 _0.5: A vision-language-action model with open-world generalization. External Links: 2504.16054, Link Cited by: Real-Robot Deployment. Shridhar et al. (2022) M. Shridhar, L. Manuelli, and D. Fox CLIPort: what and where pathways for robotic manipulation. In Proceedings of the 5th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 164, p. 894–906. External Links: Link Cited by: Related Work. Shridhar et al. (2023) M. Shridhar, L. Manuelli, and D. Fox Perceiver-actor: a multi-task transformer for robotic manipulation. In Proceedings of the 6th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 205, p. 785–799. External Links: Link Cited by: Related Work. Yuan et al. (2025a) W. Yuan, J. Duan, V. Blukis, W. Pumacay, R. Krishna, A. Murali, A. Mousavian, and D. Fox RoboPoint: a vision-language model for spatial affordance prediction in robotics. In Proceedings of The 8th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 270, p. 4005–4020. External Links: Link Cited by: Related Work, Table 1. Yuan et al. (2026) Y. Yuan, H. Cui, Y. Chen, Z. Dong, F. Ni, L. Kou, J. Liu, P. Li, Y. Zheng, and J. Hao From seeing to doing: bridging reasoning and decision for robotic manipulation. External Links: 2505.08548, Link Cited by: Related Work, Table 1. Yuan et al. (2025b) Y. Yuan, H. Cui, Y. Huang, Y. Chen, F. Ni, Z. Dong, P. Li, Y. Zheng, H. Tang, and J. Hao Embodied-R1: reinforced embodied reasoning for general robotic manipulation. External Links: 2508.13998, Link Cited by: Related Work, Table 1, Table 1. Zeng et al. (2021) A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, and J. Lee Transporter networks: rearranging the visual world for robotic manipulation. In Proceedings of the 4th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 155, p. 726–747. External Links: Link Cited by: Related Work. Zhao et al. (2023) T. Z. Zhao, V. Kumar, S. Levine, and C. Finn Learning fine-grained bimanual manipulation with low-cost hardware. In Proceedings of Robotics: Science and Systems, External Links: Document, Link Cited by: Related Work. Zitkovich et al. (2023) B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V. Vanhoucke, H. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. Sanketi, G. Salazar, M. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y. Lu, S. Levine, T. E. Lee, L. Lee, I. Leal, Y. Kuang, D. Kalashnikov, R. Julian, N. Joshi, A. Irpan, B. Ichter, J. Hsu, A. Herzog, K. Hausman, K. Gopalakrishnan, C. Fu, P. Florence, C. Finn, A. Dubey, D. Driess, T. Ding, K. Choromanski, X. Chen, Y. Chebotar, J. Carbajal, N. Brown, A. Brohan, M. G. Arenas, and K. Han RT-2: vision-language-action models transfer web knowledge to robotic control. In Proceedings of the 7th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 229, p. 2165–2183. External Links: Link Cited by: Related Work.