Paper deep dive
Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation
Zhiruo Zhou, Zelin Li, Xiwen Chen, Jiazhuo Li, Chenwei Wang, Huiming Chen, Xiaojun Zhu
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text becomes a control intervention once it enters the executed prompt. In a matched audit, raw appended text reduces mean success from 92.47\% to 3.00\%, while meaningful and length-matched meaningless appends both fail on all 500 states. This result identifies \emph{prompt-form collapse}: changing the instruction form, rather than adding useful semantics, can dominate execution. We introduce TOWN-VLA (Think Only When Needed), a prompt-authority interface that separates candidate generation from permission to alter the policy input. A fixed compatibility rule authorizes a canonical compact instruction; otherwise, the interface restores the original Base prompt exactly. Across 900 audited routes, every route follows this contract: 525 routes recover Base with matching hashes, and all 375 authorized prompts preserve the task signature. On a matched $4\times7$ LIBERO-Plus evaluation with 10{,}030 episodes per method, success rises from 69.5\% to 73.1\% ($+362$ episodes; 95\% CI 1.89--5.45 points), improving on six perturbation axes and all four suites. On a physical PiPER arm with a frozen \pizerofive{} checkpoint, success rises from 52.7\% to 78.7\% over 150 trials per method ($p=3.16\times10^{-6}$). Prompt authority is enforceable for a frozen controller; oracle-free admission calibration is the next deployment target.
Tags
Links
- Source: https://arxiv.org/abs/2608.23224v1
- Canonical: https://arxiv.org/abs/2608.23224v1
Trouble viewing inline? Open PDF directly →
Full Text
49,661 characters extracted from source content.
Expand or collapse full text
Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation Zhiruo Zhou Zelin Li Xiwen Chen Jiazhuo Li Chenwei Wang Huiming Chen Xiaojun Zhu Abstract Retrieval can efficiently and effectively augment a frozen vision–language–action (VLA) policy without retraining, yet retrieved text becomes a control intervention once it enters the executed prompt. In a matched audit, raw appended text reduces mean success from 92.47% to 3.00%, while meaningful and length-matched meaningless appends both fail on all 500 states. This result identifies prompt-form collapse: changing the instruction form, rather than adding useful semantics, can dominate execution. We introduce TOWN-VLA (Think Only When Needed), a prompt-authority interface that separates candidate generation from permission to alter the policy input. A fixed compatibility rule authorizes a canonical compact instruction; otherwise, the interface restores the original Base prompt exactly. Across 900 audited routes, every route follows this contract: 525 routes recover Base with matching hashes, and all 375 authorized prompts preserve the task signature. On a matched 4×74× 7 LIBERO-Plus evaluation with 10,030 episodes per method, success rises from 69.5% to 73.1% (+362+362 episodes; 95% CI 1.89–5.45 points), improving on six perturbation axes and all four suites. On a physical PiPER arm with a frozen π0.5 _0.5 checkpoint, success rises from 52.7% to 78.7% over 150 trials per method (p=3.16×10−6p=3.16× 10^-6). Prompt authority is enforceable for a frozen controller; oracle-free admission calibration is the next deployment target. Introduction Vision-language-action (VLA) policies map visual observations and language instructions to closed-loop robot actions, connecting the semantic breadth of vision-language pretraining with continuous control (9; 52; 18). Cross-embodiment data, generalist pretraining, and continuous-action models have broadened their capabilities (28; 27; 5; 25; 29). Practical systems now wrap frozen policies with planners, memories, verifiers, or steering modules that supply test-time guidance (1; 14; 48; 21). Once such text enters the executed instruction, it becomes a control intervention. We study this interface rather than policy retraining, asking which retrieved content may modify the frozen policy input (Figures 1 and 2). Figure 1: Prompt authority at the frozen-policy boundary. Direct injection passes retrieved text to the VLA; TOWN-VLA (Think Only When Needed) authorizes a canonical capsule or restores Base while retaining the same frozen action generator. Holding the policy, initial states, and execution protocol fixed, raw appended text reduces mean success from 92.47% to 3.00% (Figure 3(a)). In a paired 500-state control, meaningful and length-matched meaningless appends both yield 0/500, called prompt-form collapse. We hypothesize that OFT (Optimized Fine-Tuning) adaptation exposes the policy to a stable instruction template but no appended retrieval tokens; appending text shifts sequence length, token positions, and the instruction boundary. The identical failure of semantically different appends, while canonical prompts remain operative, supports template departure over semantic quality, demonstrating task-preserving rewordings and multimodal perturbations (40; 45) by isolating the interface around an otherwise frozen policy. Existing planners ground proposals through affordances, feedback, programs, or spatial value maps (1; 14; 22; 13); newer systems request assistance, route experts, retrieve memory, or steer frozen policies (31; 49; 32; 21; 51; 15). They make the proposal-to-action interface consequential, but do not generally separate candidate quality from an explicit, logged authorization decision. This distinction parallels value-of-computation reasoning and selective prediction with a reject option (33; 12): a slow path may be worth querying without its output being worth obeying.We therefore formulate retrieval augmentation as prompt-authority control in TOWN-VLA. Figure 1 contrasts direct injection with rejection followed by Base-prompt restoration. Memory is one replaceable source of proposals; the validated method concerns the contract governing whether and how any candidate may alter the frozen policy input, rather than attributing gains to retrieved strategy content. TOWN-VLA separates candidate preparation from prompt authorization (Figure 2). A Compatibility-Reranked Capsule retrieves structured candidates and orders them by a fixed text-level compatibility rule. A Top-2 Fail-Closed Cascade inspects at most two candidates: the first eligible candidate becomes a canonical compact instruction; if neither passes, the interface retains the exact Base Policy prompt. The frozen Base Policy remains the sole action generator, with its parameters and control loop unchanged. The construction supplies two verifiable properties: (P1) every unauthorized route resolves to the bit-identical Base prompt, and (P2) each inspected episode uses at most one retrieval, five compatibility scores, and two checks. Task-Prior Admission estimates an oracle-side upper bound on removable slow-path computation; oracle-free admission is evaluated in Q4. Key Contributions: • We formulate prompt-authority control as a previously underexplored test-time interface problem for frozen VLAs and identify prompt-form collapse as its motivating failure mode. Raw appended text reduces mean success from 92.47% to 3.00%; task-aligned and length-matched meaningless suffixes both yield 0/500, whereas exact Base restoration reaches 499/500. • We operationalize proposal ≠ authority at the frozen-policy boundary: TOWN-VLA admits only canonical guidance, restores rejected routes to the exact Base prompt, and limits inspection to Top-2. Across 900 routes, 525 recover hash-identical Base prompts and all 375 authorized prompts preserve the task signature. • In matched evaluation, TOWN-VLA raises LIBERO-Plus success from 69.46% to 73.07% (+362+362 successes; paired-cell bootstrap 95% CI, 1.89–5.45 points), with gains on six of seven axes and all four suites; physical PiPER success rises from 52.7% to 78.7%. • A preregistered boundary quantifies admission timing. Oracle routing preserves 2,826 successes, halving slow-path calls; on the fixed 24/36 development–held-out split, oracle-free gates peak at the 91.81% Base Policy. Admission calibration is the next challenge. Figure 2: TOWN-VLA separates invocation, candidate ranking, prompt authorization, and frozen control. Ranked contexts are checked and either rendered as a canonical capsule or rejected for exact Base restoration. Logged scores, flags, and prompt hashes make routes auditable. Task-Prior is used only for the oracle control in Equation (9). Related Work Frozen policies and inference-time interfaces. PaLM-E, RT-1, and RT-2 established scalable language-conditioned robot control (9; 6; 52), while Open X-Embodiment, Octo, and OpenVLA scaled cross-dataset adaptation (28; 27; 18). Continuous and latent-action models extend bimanual, spatial, and open-world control (25; 5; 4), with deployability explored by SmolVLA (38). These lines redesign the policy, training data, or action representation. We instead freeze the generator and regulate external text through the interface it already understands. External reasoning and fast–slow authority. SayCan and VoxPoser ground language through affordances and spatial value maps (1; 13). ECoT, InstructVLA, Acting While Understanding, and VLS move reasoning closer to action (48; 47; 46; 24). Fast–slow systems separate semantic and reactive computation (50; 8; 54). Our slow path remains external: it may compute a proposal, but its text cannot condition the action generator without authorization. The separation therefore concerns control rights, not only execution rate. Inspectable memory at the policy boundary. Unlike latent memory, an external capsule can be inspected and hashed before it crosses the policy boundary. MemoryVLA and MemoryVLA++ integrate perceptual and episodic history (37; 36). WorldVLA, VLA-JEPA, WAMs, and latent-action world models instead use predicted visual or latent dynamics (7; 41; 43; 11). External memories retrieve experience or action priors (39; 26; 23; 21; 51). We instead ask when an inspectable textual proposal may replace the frozen policy instruction. Prompt brittleness beyond language generation. Language models can be insensitive to prompt meaning yet highly sensitive to formatting or optimized suffixes (44; 34; 53). These studies concern output quality in open-loop generation. In a frozen closed-loop VLA, out-of-template text conditions actions, turning prompt brittleness into control failure. We therefore regulate boundary crossing rather than optimize a prompt. Selective intervention and runtime assurance. LIBERO-Plus, Q-DIG, and STRONG-VLA expose or train against visual and linguistic shift (10; 40; 45). Assistance and routing methods calibrate uncertainty or select experts (31; 49; 32); Mostly Harmless VLA Steering learns when feedback may help (15). Runtime-assurance systems supervise learned controllers through Simplex or action shielding (35; 2). Our action generator and fallback remain fixed, so metareasoning or rejection (33; 12) becomes exact Base-prompt restoration rather than controller switching. Method Problem Setup and Interface Contract Symbol Meaning ℓ,u⋆,p⋆ ,\ u ,\ p Base instruction, resolved instruction, and resolved prompt (u⋆)P(u ) h,ℋKh,\ H_K One memory candidate and the retrieved top-K set g,j⋆g,\ j Slow-path invocation bit and authorized candidate rank sigsig Normalized object–relation–target task signature TOWN-VLA wraps a frozen VLA policy πθ _θ, with θ fixed throughout evaluation. It implements a deterministic interface contract that constrains candidate-generated policy inputs while retaining πθ _θ as the sole action generator. At step t, the policy receives ot=(It,qt,τ<t)o_t=(I_t,q_t, _<t), comprising the image, robot state, and action history. A fixed renderer P maps the original instruction to the Base prompt: pbase=(ℓ),atbase=πθ(ot,pbase),θfixed.p_ base=P( ), a_t base= _θ(o_t,p_ base), θ\ fixed. (1) The interface initializes u⋆=ℓu = ; only the authorization rule in Equation (8) may replace it with a canonical instruction and set p⋆=(u⋆)p =P(u ). Otherwise p⋆=pbasep =p_ base byte for byte. One resolved prompt is reused throughout rollout, leaving policy parameters, action space, and control frequency fixed. Retrieval logs candidate identifiers, reranking logs scores, the checker logs its flag and reason, and the interface logs the resolved-prompt hash. Together they separate proposal order, authorization, and controller input. Compatibility-Reranked Capsule The frozen memory ℳ=hii=148M=\h_i\_i=1^48 contains demonstration trajectories but no evaluation rollouts. CLIP text features and a nearest-neighbor index retrieve K=5K=5 candidates (30; 16): ℋK(ℓ)=RetrieveK(ℓ;ℳ).H_K( )=Retrieve_K( ;M). (2) Each memory entry stores a task description, structured context, raw plan, and compact hint. Retrieval indexes the description; parsing and rendering consume only the context. Raw plans and hints remain provenance fields, never executed text, and the memory is read-only during evaluation. A frozen parser extracts object–target pairs from the task and each stored context. Both parsers read text only—never images, robot state, action history, or benchmark labels. Fixed pick/place templates and deterministic fallbacks cover unmatched strings. With J denoting token-set Jaccard overlap, xℓ x_ =ParseTask(ℓ)=(xobj,xtgt), =ParseTask( )=(x_ obj,x_ tgt), (3) yh y_h =ParseContext(h)=(yobj,ytgt), =ParseContext(h)=(y_ obj,y_ tgt), mobj m_ obj =J(xobj,yobj),mtgt=J(xtgt,ytgt), =J(x_ obj,y_ obj), m_ tgt=J(x_ tgt,y_ tgt), mctx m_ ctx =J(xobj->xtgt,context(h)). =J(x_ obj ->x_ tgt,context(h)). Here xobj->xtgtx_ obj ->x_ tgt is the literal ordered string formed by concatenating the normalized object phrase, the delimiter ->, and the normalized target phrase; it is not a functional map. For route identity, the same frozen parser supplies normalized object, relation, and target fields: sig(u) (u) =(cobj(u),crel(u),ctgt(u)), = (c_ obj(u),c_ rel(u),c_ tgt(u) ), SigEq(u,v) (u,v) =[sig(u)=sig(v)]. =1\! [sig(u)=sig(v) ]. Signature equality is componentwise after the parser’s deterministic text normalization and fixed relation-alias mapping. The mismatch indicators are robj=[yobj≠∅∧mobj=0]r_ obj=1[y_ obj≠ m_ obj=0] and rtgt=[ytgt≠∅∧mtgt=0]r_ tgt=1[y_ tgt≠ m_ tgt=0]. The frozen compatibility score is s(h,xℓ)= s(h,x_ )= αclipsclip(h,ℓ)+αobjmobj+αtgtmtgt _ clips_ clip(h, )+ _ objm_ obj+ _ tgtm_ tgt (4) +αctxmctx−λobjrobj−λtgtrtgt + _ ctxm_ ctx- _ objr_ obj- _ tgtr_ tgt −ηrankclip(h). -η\,rank_ clip(h). We fix (αclip,αobj,αtgt,αctx,λobj,λtgt,η)=(1,2,1.5,0.8,0.6,0.4,0.01)( _ clip, _ obj, _ tgt, _ ctx, _ obj, _ tgt,η)=(1,2,1.5,0.8,0.6,0.4,0.01) before evaluation. These coefficients rerank textual agreement and are not fitted to outcomes. The resulting order is (h(1),…,h(K))=argsorth∈ℋK(ℓ)↓s(h,xℓ).(h_(1),…,h_(K))= *argsort _h _K( )s(h,x_ ). (5) The score combines CLIP proximity with structural overlap, explicit mismatch penalties, and deterministic rank-based tie resolution. Reranking changes only inspection order; the hard checker alone grants authorization. Top-2 Fail-Closed Cascade The hard checker records deterministic conflict reasons ℛ(h,ℓ)R(h, ). Candidate eligibility is Gcomp(h,ℓ)=[ℛ(h,ℓ)=∅].G_ comp(h, )=1\! [R(h, )= ]. (6) The checker flags a nonempty candidate object or target with zero task overlap. Context affects only soft ranking, so this remains a text-level rule. Rank 2 is inspected only after rank 1 is rejected: j⋆=minj∈1,2:Gcomp(h(j),ℓ)=1,j = \! \j∈\1,2\:G_ comp(h_(j), )=1 \, (7) with j⋆=∅j = if neither passes. Let g∈0,1g∈\0,1\ denote whether the slow path is invoked; always-on evaluation sets g=1g=1, while the oracle ablation sets g=gprior(m)(z)g=g_ prior^(m)(z). The deterministic renderer and final authority rule are ucap u_ cap =RenderCapsule(ℓ,context(h(j⋆))), =RenderCapsule\! ( ,context(h_(j )) ), (8) u⋆ u =ucap,g=1∧j⋆≠∅,ℓ,otherwise, = casesu_ cap,&g=1 j ≠ ,\\ ,&otherwise, cases p⋆ p =(u⋆). =P(u ). The cascade short-circuits after the first eligible candidate and never merges contexts. An accepted context is rendered with the fixed template put <object> <relation> <target>; empty or unrenderable context returns ℓ . For example, if Base is Place the black bowl on the plate., rejection returns that byte-identical string, whereas authorization resolves to put the black bowl on the plate. Fallback adds no retrieved text or rejection message. In the 1,200-state audit, all authorized routes resolve at rank 1, leaving the second slot unexercised. Top-2 therefore serves as a fixed safety budget that preserves one recovery opportunity under future memory or ranking changes. Task-Prior Admission as an Oracle Control To quantify the oracle-side upper bound on removable slow-path computation, a benchmark label z supplies a manifest-indexed routing bit: gprior(m)(z)=[z∈allow(m)],g_ prior^(m)(z)=1\! [z _ allow^(m) ], (9) where m∈five-suite,paired-3000m∈\five-suite,paired-3000\ indexes the frozen routing manifest. For the five-suite manifest, allow(five-suite)Z_ allow^(five-suite) contains red-stick, mug, and yellow-book spatial suites; standard-spatial and object-mug execute Base. Its 3/5 allow–2/5 skip lookup saves 40% of calls. Table 4b instead uses a separately frozen paired-3000 manifest with 1,500 invoked and 1,500 bypassed states, hence 50% savings. The two denominators are reported separately and never pooled. For N episodes, slow-path coverage and calls saved are ρslow=1N∑i=1Ngi,Scall=1−ρslow. _ slow= 1N _i=1^Ng_i, S_ call=1- _ slow. (10) Routing precedes retrieval and changes only whether slow-path work runs; Equations (6)–(8) separately control prompt intervention. Per episode, the interface adds at most one retrieval, five scores, two checks, one rendering, and one prompt resolution. All components and routing tables are frozen before evaluation. Runtime States and Interface Guarantees TOWN-VLA exposes three auditable runtime states: bypass executes (ℓ)P( ) without retrieval; inspected-but-unauthorized also executes Base after checking; authorized renders ucapu_ cap. Thus spending slow-path computation and granting prompt authority remain separate decisions. The construction enforces (I1) prompt-boundary non-interference on rejection, (I2) πθ _θ as the sole action generator, and (I3) a memory-size-independent post-retrieval inspection budget. Candidate ID, conflict reason, route state, and resolved-prompt hash provide the corresponding execution trace. Execution initializes u⋆=ℓu = , evaluates routing, and, if invoked, retrieves, scores, and checks at most two candidates. One prompt hash is fixed before the first action; the slow path is not queried again during rollout. Unless labeled oracle-side, all results use g=1g=1; the main evaluation therefore measures prompt authority independently of benchmark labels and Task-Prior routing. Method Setting Camera Robot Language Light Background Noise Layout Weighted Mean OpenVLA Direct 0.8 3.5 23.0 8.1 34.8 15.2 28.5 15.6 NORA Direct 2.2 37.0 65.1 45.7 58.6 12.8 62.1 39.0 WorldVLA Direct 0.1 27.9 41.6 43.7 17.1 10.9 38.0 25.0 UniVLA Direct 1.8 46.2 69.6 69.0 81.0 21.2 31.9 42.9 π0 _0 Flow policy 13.8 6.0 58.8 85.0 81.4 79.0 68.9 53.6 π0 _0-FAST Autoregressive 65.1 21.6 61.0 73.2 73.2 74.4 68.8 61.6 OpenVLA-OFTw Third-view only 10.4 38.7 70.5 76.8 93.6 49.9 69.9 55.8 OpenVLA-OFTm Mixed supervised tuning 55.6 21.7 81.0 92.7 91.0 78.6 68.7 67.9 OpenVLA-OFT Optimized fine-tuning 56.4 31.5 78.7 88.6 93.1 75.6 74.9 69.5 TOWN-VLA (Always-On) Frozen-backbone interface 61.5 37.5 85.2 92.7 95.6 73.6 78.0 73.1 Table 1: Seven-axis LIBERO-Plus SR (%). Published rows are from 10; the local OpenVLA-OFT and TOWN-VLA rows use the same frozen backbone, 28 cells, and 10,030 episodes per method. Experiments We organize the evaluation around five questions: end-to-end robustness (Q1), prompt-form failure (Q2), authorization and exact restoration (Q3), selective computation (Q4), and transfer across backbones and embodiments (Q5). Setup and Protocol Map OpenVLA with the Optimized Fine-Tuning recipe (OpenVLA-OFT) is the Base Policy (17). TOWN-VLA and its ablations modify only the prompt interface while holding the backbone, action space, policy parameters, and control frequency fixed. Table 1 and all Q1–Q3/Q5 end-to-end comparisons use the always-on setting (g=1g=1), independently of benchmark labels and Task-Prior routing. Experiments run on NVIDIA RTX A6000 hardware with CUDA 12.1 and headless MuJoCo rendering through EGL (42). Success rate (SR) is the primary metric. Each frozen protocol includes its own matched Base Policy, and every reported difference uses the Base evaluated on the same states. Simulation contrasts use paired initial states whenever episode-level outcomes are available. The LIBERO-Plus interval resamples its 28 matched suite–axis cells, the prompt-form audit uses paired bootstrap and exact McNemar tests, and physical task success uses a two-sided Fisher exact test. Published comparison rows provide context only; claims about the interface use the locally matched backbone and execution stack. The LIBERO-Plus interval summarizes variation across matched cells, rather than independent run-to-run variation. Perturbation Breadth (Q1) Under the matched local protocol, TOWN-VLA raises weighted LIBERO-Plus SR by 3.61 points, adding 362 successes over 10,030 episodes against the same OpenVLA-OFT backbone (Table 1). The paired-cell 95% interval is 1.89–5.45 points. The axis view improves in six of seven conditions, led by Language (+6.5+6.5) and Robot (+6.0+6.0); image-only Noise is the sole lower estimate. Aggregating the same episodes by task suite gives a positive difference for Spatial, Object, Goal, and Long-Horizon (Table 2). The two partitions show that the overall gain is distributed across perturbation types and task families. Method Spatial Object Goal Long Overall OpenVLA 19.4 14.0 15.1 14.3 15.6 WorldVLA 32.5 28.6 31.8 8.2 25.0 UniVLA 55.5 36.7 40.7 39.9 42.9 π0 _0 60.7 61.4 44.9 48.4 53.6 π0 _0-FAST 74.4 72.7 57.6 43.4 61.6 OpenVLA-OFT 84.0 65.8 62.9 65.9 69.5 TOWN-VLA 86.3 69.5 68.1 69.1 73.1 Table 2: Four-suite LIBERO-Plus SR (%). Published rows follow 10; local rows aggregate the matched episodes in Table 1. Prompt-Form Collapse (Q2) Across three fixed-schedule executions, mean SR falls from 92.47% under Base to 3.00% under raw appended text. Compatibility reranking and Top-2 fail-closed both recover to 91.08% (Figure 3a). The executions reuse all 1,200 initial-state hashes and measure implementation reproducibility, not environmental-seed variance. In the paired diagnostic, the first raw candidate changes SR by −88.67-88.67 points (95% CI −94.00-94.00 to −82.33-82.33). Increasing K beyond one neither worsens nor mitigates the collapse, so retrieval depth is not the operative variable. A matched 500-state factorial control separates prompt form from semantics (Table 3). Base and exact restoration each reach 499/500, and the correct canonical instruction reaches 497/500. Correct and meaningless appends both yield 0/500, whereas canonical wrong-object and wrong-target controls reach 497/500 and 496/500. Task-only and retrieved canonical arms share all prompt hashes and outcomes, indicating that prompt form, rather than retrieved strategy content, explains performance in this audit. Prompt form Correct Wrong obj. Wrong tgt. Meaningless Canonical 497/500 497/500 496/500 – Appended 0/500 – – 0/500 Table 3: Prompt form × semantics on 500 matched states. Canonical controls alter the object or target; the meaningless append is length matched. Base and exact restoration are 499/500. Figure 3: Prompt-form collapse and retrieval-depth diagnostics. (a) Mean SR over three fixed-schedule repeats on one manifest (1,200 episodes per method and repeat); the Raw annotation is its change from matched Base. Compat. R and Top-2 denote the reranked capsule and fail-closed cascade. (b) Paired Top-K results on 300 states with Wilson 95% intervals. (a) Interface configuration, five-suite manifest Variant Task-prior Capsule Rerank Fail-closed SR (%) Base Policy – – – – 89.48 Strategy Capsule (Top-1) – ✓ – – 90.64 Compatibility-Reranked Capsule – ✓ ✓ – 91.60 Top-2 Fail-Closed Cascade – ✓ ✓ ✓ 91.32 Complete Interface ✓ ✓ ✓ ✓ 91.32 (b) Oracle-side selective invocation Metric Always-On Routed Change Paired SR (%) 94.20 94.20 0.000.00 Slow-path calls 3000/3000 1500/3000 −50%-50\% Decision latency (ms) 212.50 190.75 −10.24%-10.24\% Table 4: Interface and invocation controls. Complete Interface is TOWN-VLA. (a) Five-suite SR over 2,500 episodes per variant; checks mark enabled components. (b) Oracle routing on a separate paired-3,000 manifest, split equally between invoke and bypass. Authorization and Exact Restoration (Q3) Route-level identity. A 1,200-state always-on audit separates authorization from restoration. Reranking admits all states and preserves the task signature on 1,100/1,200. The cascade admits 1,000 states, preserves the signature on 940/1,000 authorized routes, and restores exact Base on the other 200. Thus its 94.00% denominator is the authorized subset, not all states. A separate 900-route audit includes routing. The slow path runs on 450 routes: 375 are authorized with SigEq=1SigEq=1, and 75 are inspected then rejected. The other 450 bypass retrieval. Consequently, the 525 exact-Base routes comprise 450 bypasses and 75 rejections. The identity and restoration rates differ because the two audits use different manifests and denominators. The high-Base always-on audit yields +1.00/+1.67 points for signature-changing prompts (McNemar p=1p=1), whereas Q1 localizes gains to perturbed inputs, especially the lower-Base Language and Robot cells. This contrast is consistent with task-equivalent canonicalization rather than semantic departure. Restricted-OOD replication. The same ordering appears under spatial/object shift (Table 5): Raw Top-5 drops to 74.20/79.40, reranking recovers to 88.20/96.67, and fail-closed authorization reaches 89.33/97.53, above Base in both reported shift aggregates. We additionally evaluate five released checkpoints under the same local protocol. The citations identify the model sources rather than supply these measurements (41; 20; 38; 19; 3); their locally measured spatial means span 55.20%–88.13%, and object means span 45.60%–95.07%. On the five-suite manifest (Table 4a), reranking scores 91.60% and the bounded cascade 91.32%. The 0.28-point difference is seven outcomes over 2,500 episodes; the cascade is retained for exact restoration rather than claimed as an incremental SR mechanism. Method Spatial shift Object shift Base Policy 85.80 93.93 Raw Top-5 Plan 74.20 79.40 Compatibility-Reranked Capsule 88.20 96.67 Top-2 Fail-Closed Cascade 89.33 97.53 Table 5: Restricted-OOD success rate (%) under spatial and object shifts. All rows are local evaluations sharing one frozen Base Policy and execution protocol. Selective Computation and Gate Controls (Q4) Oracle-side headroom. On 3,000 paired states, benchmark-side routing preserves the same 2,826 successes while halving slow-path calls (Table 4b). Decision latency falls 10.24%, from 212.50 to 190.75 ms, over three rounds of 300 decisions after 50 warm-ups. This episode-level decision includes routing, retrieval, reranking, tokenization, and one VLA decode, but excludes simulator and startup time. The five-suite manifest separately saves 40% of calls. Oracle-free admission remains a calibration challenge. On the preregistered 24-development/36-held-out cell split, the learned selector authorizes 2/36 cells and preserves the paired outcomes across those 40 episodes, matching the 91.81% Base Policy; CLIP authorizes 0/36 cells. Always-on intervention is 1.53 points below Base, and the form-aware and matched-budget controls do not recover this gap. Thus, the oracle analysis quantifies removable computation, while the evaluated gates do not yet identify which calls to remove. A separate compositional-language probe identifies the relational fields needed to extend canonical rendering. Base reaches 20/30, whereas canonical prompts without drawer, cabinet, and relative-location clauses reach 0/30. Together with the 500-state object–target control, this result bounds the current renderer and motivates preserving richer relational structure for compositional instructions. Arm Authorized cells SR (%) Exact Base (reference) 0/36 91.81 Learned, no oracle 2/36 91.81 Fixed/random matched budget 32/36 90.14–90.97 Form-aware 32/36 90.28 Table 6: Oracle-free gates on the fixed 24-development/36-held-out split, reporting held-out authorization and SR. Figure 4: Representative π0.5 _0.5/PiPER rollouts in the three evaluated scenes. Columns are time ordered; rows compare Base and TOWN-VLA, and red boxes mark interaction regions. Prompts resolve once; results use all 150 trials per method. Transfer to a Second Backbone and a Physical Robot (Q5) The instruction is “Grasp the white object and place it upright in the green tray.” We test no-distractor, yellow-cup, and red-cylinder scenes. Object poses are randomized around a teleoperated reference, and Base/TOWN-VLA trials are randomly interleaved. A human records success after the target remains upright for two seconds (50 trials per scene–method cell). Both methods share the PiPER arm, dual RealSense D405 cameras, frozen π0.5 _0.5 checkpoint, and success criterion. Controller settings are fixed to a 10-action-step replanning horizon, speed ratio 5, and a 60-second limit. Drops, collisions, human intervention, timeout, or failure to maintain the upright placement count as failures. Formal trials are excluded from training, memory, and calibration, leaving the prompt interface as the method-level change. TOWN-VLA raises success from 79/150 to 118/150 (+26.00+26.00 points; p=3.16×10−6p=3.16× 10^-6, Fisher exact test; Table 7), with 22–32-point gains in all three scenes. Because prompts resolve before rollout, corrective re-grasp is closed-loop policy behavior (Figure 4). Post-miss recovery rises from 24.4% to 52.4%, but remains descriptive: method-dependent misses create unequal post-treatment denominators (41 versus 21). Because pose distributions differ by scene, the controlled comparison is Base versus TOWN-VLA (Complete Interface) within each row. All three favor the Complete Interface, so the pooled gain is not carried by one configuration. Task success Post-miss recovery Scene Base Policy TOWN-VLA Base Policy TOWN-VLA No distractor 21/50 (42.0) 37/50 (74.0) 3/16 (18.8) 2/6 (33.3) Yellow distractor 27/50 (54.0) 39/50 (78.0) 3/13 (23.1) 3/7 (42.9) Red distractor 31/50 (62.0) 42/50 (84.0) 4/12 (33.3) 6/8 (75.0) Overall 79/150 (52.7) 118/150 (78.7) 10/41 (24.4) 11/21 (52.4) Table 7: π0.5 _0.5/PiPER task success and descriptive recovery. Cells show successes/trials (%); recovery conditions on an initial miss and uses method-dependent denominators. These trials evaluate interface transfer rather than policy adaptation: backbone, cameras, feedback loop, and criterion remain fixed; only the resolved instruction changes. Discussion TOWN-VLA results support separating retrieval relevance from prompt authority: canonical rendering constrains accepted intervention, while exact restoration makes rejection reversible. Across protocols, the contrast between Q1 and Q3 suggests that intervention is most valuable when perturbations have degraded the executed instruction and is largely neutral when Base remains strong. Route-level attribution of the Q1 gain is an important next step. Simulation isolates this interface behavior under matched manifests, while PiPER tests the same behavior on a second backbone and embodiment. The present evaluation deliberately isolates the interface using a 48-entry same-domain memory, text-only compatibility, and a controlled single-task, single-operator physical study with human-scored outcomes. Natural extensions include larger cross-domain memories, visually conditioned admission, and broader blinded robot trials. Conclusion We studied how retrieved text should cross the input boundary of a frozen VLA. Prompt-form controls identify severe prompt-form collapse under raw appends, while TOWN-VLA (Think Only When Needed) separates candidate generation from authorization and restores Base exactly on rejected routes. Under matched evaluation, the interface improves LIBERO-Plus by 3.61 points and the PiPER real-robot platform by 26.00 points without retraining the action generator. These results establish prompt authority as an enforceable control primitive for frozen VLAs and identify reliable oracle-free admission as the next frontier for selective slow-path control. References Ahn et al. (2022) M. Ahn et al. Do as i can, not as i say: grounding language in robotic affordances. External Links: 2204.01691, Document, Link Cited by: Introduction, Introduction, External reasoning and fast–slow authority.. Alshiekh et al. (2018) M. Alshiekh, R. Bloem, R. Ehlers, B. Könighofer, S. Niekum, and U. Topcu Safe reinforcement learning via shielding. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32. External Links: Document, Link Cited by: Selective intervention and runtime assurance.. Beeface (2026) Beeface SmolVLA fine-tuned on LIBERO-Spatial: model card. Note: https://huggingface.co/Beeface/smolvla-libero-spatialAccessed: 2026-07-28 Cited by: Restricted-OOD replication.. Black et al. (2025a) K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. R. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky π0.5 _0.5: a vision-language-action model with open-world generalization. In Proceedings of The 9th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 305, p. 17–40. External Links: Link Cited by: Frozen policies and inference-time interfaces.. Black et al. (2025b) K. Black, N. Brown, D. Driess, A. Esmail, M. R. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, L. Smith, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky π0 _0: a vision-language-action flow model for general robot control. In Proceedings of Robotics: Science and Systems, Los Angeles, CA, USA. External Links: Document, Link Cited by: Introduction, Frozen policies and inference-time interfaces.. Brohan et al. (2022) A. Brohan et al. RT-1: robotics transformer for real-world control at scale. External Links: 2212.06817, Document, Link Cited by: Frozen policies and inference-time interfaces.. Cen et al. (2025) J. Cen, C. Yu, H. Yuan, Y. Jiang, S. Huang, J. Guo, X. Li, Y. Song, H. Luo, F. Wang, D. Zhao, and H. Chen WorldVLA: towards autoregressive action world model. External Links: 2506.21539, Document, Link Cited by: Inspectable memory at the policy boundary.. Chen et al. (2025) H. Chen, J. Liu, C. Gu, Z. Liu, R. Zhang, X. Li, X. He, Y. Guo, C. Fu, S. Zhang, and P. Heng Fast-in-slow: a dual-system foundation model unifying fast manipulation within slow reasoning. External Links: 2506.01953, Document, Link Cited by: External reasoning and fast–slow authority.. Driess et al. (2023) D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence PaLM-E: an embodied multimodal language model. External Links: 2303.03378, Document, Link Cited by: Introduction, Frozen policies and inference-time interfaces.. Fei et al. (2026) S. Fei, S. Wang, J. Shi, Z. Dai, J. Cai, P. Qian, L. Ji, X. He, S. Zhang, Z. Fei, J. Fu, J. Gong, and X. Qiu LIBERO-Plus: a progressive robustness benchmark for visual-language-action models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 38574–38583. External Links: Link Cited by: Selective intervention and runtime assurance., Table 1, Table 2. Garrido et al. (2026) Q. Garrido, T. Nagarajan, B. Terver, N. Ballas, Y. LeCun, and M. Rabbat Learning latent action world models in the wild. External Links: 2601.05230, Document, Link Cited by: Inspectable memory at the policy boundary.. Geifman and El-Yaniv (2019) Y. Geifman and R. El-Yaniv SelectiveNet: a deep neural network with an integrated reject option. In Proceedings of the 36th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 97, p. 2151–2159. External Links: Link Cited by: Introduction, Selective intervention and runtime assurance.. Huang et al. (2023) W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei VoxPoser: composable 3d value maps for robotic manipulation with language models. External Links: 2307.05973, Document, Link Cited by: Introduction, External reasoning and fast–slow authority.. Huang et al. (2022) W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, P. Sermanet, N. Brown, T. Jackson, L. Luu, S. Levine, K. Hausman, and B. Ichter Inner monologue: embodied reasoning through planning with language models. External Links: 2207.05608, Document, Link Cited by: Introduction, Introduction. Jeong et al. (2026) H. J. Jeong, G. Swamy, and A. Bajcsy Learning what to say to your VLA: mostly harmless vision language action model steering. External Links: 2606.12299, Document, Link Cited by: Introduction, Selective intervention and runtime assurance.. Johnson et al. (2021) J. Johnson, M. Douze, and H. Jégou Billion-scale similarity search with GPUs. IEEE Transactions on Big Data 7 (3), p. 535–547. External Links: Document Cited by: Compatibility-Reranked Capsule. Kim et al. (2025a) M. J. Kim, C. Finn, and P. Liang Fine-tuning vision-language-action models: optimizing speed and success. In Proceedings of Robotics: Science and Systems, Los Angeles, CA, USA. External Links: Document, Link Cited by: Setup and Protocol Map. Kim et al. (2025b) M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn OpenVLA: an open-source vision-language-action model. In Proceedings of The 8th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 270, p. 2679–2713. External Links: Link Cited by: Introduction, Frozen policies and inference-time interfaces.. LeRobot Team (2025) LeRobot Team SmolVLA-LIBERO model card. Note: https://huggingface.co/lerobot/smolvla_liberoAccessed: 2026-07-28 Cited by: Restricted-OOD replication.. LeRobot Team (2026) LeRobot Team VLA-JEPA-LIBERO model repository. Note: https://huggingface.co/lerobot/VLA-JEPA-LIBEROAccessed: 2026-07-28 Cited by: Restricted-OOD replication.. Li et al. (2025) R. Li, W. Guo, Z. Wu, C. Wang, H. Deng, Z. Weng, Y. Tan, and Z. Wang MAP-VLA: memory-augmented prompting for vision-language-action model in robotic manipulation. External Links: 2511.09516, Document, Link Cited by: Introduction, Introduction, Inspectable memory at the policy boundary.. Liang et al. (2022) J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng Code as policies: language model programs for embodied control. External Links: 2209.07753, Document, Link Cited by: Introduction. Lin et al. (2026) Z. Lin, R. Cui, J. Xu, X. Jin, W. Li, L. Fan, and Z. Zhang World pilot: steering vision-language-action models with world-action priors. External Links: 2606.12403, Document, Link Cited by: Inspectable memory at the policy boundary.. Liu et al. (2026a) S. Liu, I. S. Singh, Y. Xu, J. Duan, and R. Krishna VLS: steering pretrained robot policies via vision-language models. External Links: 2602.03973, Document, Link Cited by: External reasoning and fast–slow authority.. Liu et al. (2024) S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu RDT-1B: a diffusion foundation model for bimanual manipulation. External Links: 2410.07864, Document, Link Cited by: Introduction, Frozen policies and inference-time interfaces.. Liu et al. (2026b) Z. Liu, X. Ning, Z. Hu, X. Xie, W. Li, Z. Tang, C. Wang, Z. Yang, H. Wang, Y. Liu, and Z. Pu Goal2Skill: long-horizon manipulation with adaptive planning and reflection. External Links: 2604.13942, Document, Link Cited by: Inspectable memory at the policy boundary.. Octo Model Team et al. (2024) Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y. L. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine Octo: an open-source generalist robot policy. External Links: 2405.12213, Document, Link Cited by: Introduction, Frozen policies and inference-time interfaces.. Open X-Embodiment Collaboration (2023) Open X-Embodiment Collaboration Open X-embodiment: robotic learning datasets and RT-X models. External Links: 2310.08864, Document, Link Cited by: Introduction, Frozen policies and inference-time interfaces.. Pertsch et al. (2025) K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine FAST: efficient action tokenization for vision-language-action models. In Proceedings of Robotics: Science and Systems, Los Angeles, CA, USA. External Links: Document, Link Cited by: Introduction. Radford et al. (2021) A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 139, p. 8748–8763. External Links: Link Cited by: Compatibility-Reranked Capsule. Ren et al. (2023) A. Z. Ren, A. Dixit, A. Bodrova, S. Singh, S. Tu, N. Brown, P. Xu, L. Takayama, F. Xia, J. Varley, Z. Xu, D. Sadigh, A. Zeng, and A. Majumdar Robots that ask for help: uncertainty alignment for large language model planners. In Proceedings of The 7th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 229, p. 661–682. External Links: Link Cited by: Introduction, Selective intervention and runtime assurance.. Ren et al. (2026) X. Ren, C. Yi, and Y. Sun RouterVLA: budgeted commissioning and expert onboarding for growing VLA pools. External Links: 2606.27355, Document, Link Cited by: Introduction, Selective intervention and runtime assurance.. Russell and Wefald (1991) S. Russell and E. Wefald Principles of metareasoning. Artificial Intelligence 49 (1–3), p. 361–395. External Links: Document Cited by: Introduction, Selective intervention and runtime assurance.. Sclar et al. (2024) M. Sclar, Y. Choi, Y. Tsvetkov, and A. Suhr Quantifying language models’ sensitivity to spurious features in prompt design or: how i learned to start worrying about prompt formatting. In International Conference on Learning Representations, External Links: Link Cited by: Prompt brittleness beyond language generation.. Seto et al. (1998) D. Seto, B. H. Krogh, L. Sha, and A. Chutinan Dynamic control system upgrade using the simplex architecture. IEEE Control Systems Magazine 18 (4), p. 72–80. External Links: Document Cited by: Selective intervention and runtime assurance.. Shi et al. (2026a) H. Shi, W. Li, B. Xie, Y. Wang, R. Zhou, T. Wang, X. Zhang, P. Luo, and G. Huang MemoryVLA++: temporal modeling via memory and imagination in vision-language-action models. External Links: 2606.09827, Document, Link Cited by: Inspectable memory at the policy boundary.. Shi et al. (2026b) H. Shi, B. Xie, Y. Liu, L. Sun, F. Liu, T. Wang, E. Zhou, H. Fan, X. Zhang, and G. Huang MemoryVLA: perceptual-cognitive memory in vision-language-action models for robotic manipulation. In International Conference on Learning Representations, External Links: 2508.19236, Document, Link Cited by: Inspectable memory at the policy boundary.. Shukor et al. (2025) M. Shukor, D. Aubakirova, F. Capuano, P. Kooijmans, S. Palma, A. Zouitine, M. Aractingi, C. Pascal, M. Russi, A. Marafioti, S. Alibert, M. Cord, T. Wolf, and R. Cadene SmolVLA: a vision-language-action model for affordable and efficient robotics. External Links: 2506.01844, Document, Link Cited by: Frozen policies and inference-time interfaces., Restricted-OOD replication.. Sridhar et al. (2025) A. Sridhar, J. Pan, S. Sharma, and C. Finn MemER: scaling up memory for robot control via experience retrieval. External Links: 2510.20328, Document, Link Cited by: Inspectable memory at the policy boundary.. Srikanth et al. (2026) S. Srikanth, F. Liang, Y. Hsu, V. Bhatt, S. Zhao, H. Chen, B. Tjanaka, M. Hwang, A. Saran, D. Seita, A. Tabrez, and S. Nikolaidis Red-teaming vision-language-action models via quality diversity prompt generation for robust robot policies. External Links: 2603.12510, Document, Link Cited by: Introduction, Selective intervention and runtime assurance.. Sun et al. (2026) J. Sun, W. Zhang, Z. Qi, S. Ren, Z. Liu, H. Zhu, G. Sun, X. Jin, and Z. Chen VLA-JEPA: enhancing vision-language-action model with latent world model. External Links: 2602.10098, Document, Link Cited by: Inspectable memory at the policy boundary., Restricted-OOD replication.. Todorov et al. (2012) E. Todorov, T. Erez, and Y. Tassa MuJoCo: a physics engine for model-based control. In IEEE/RSJ International Conference on Intelligent Robots and Systems, p. 5026–5033. External Links: Document Cited by: Setup and Protocol Map. Wang et al. (2026) S. Wang, J. Shi, Z. Fu, X. He, F. Liu, C. Yang, Y. Zhou, Z. Fei, J. Gong, J. Fu, M. Z. Shou, X. Huang, X. Qiu, and Y. Jiang World action models: the next frontier in embodied AI. External Links: 2605.12090, Document, Link Cited by: Inspectable memory at the policy boundary.. Webson and Pavlick (2022) A. Webson and E. Pavlick Do prompt-based models really understand the meaning of their prompts?. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p. 2300–2344. External Links: Document, Link Cited by: Prompt brittleness beyond language generation.. Xie et al. (2026) Y. Xie, Y. Yan, Y. Zhao, H. Wang, and Y. Jin STRONG-VLA: decoupled robustness learning for vision-language-action models under multimodal perturbations. External Links: 2604.10055, Document, Link Cited by: Introduction, Selective intervention and runtime assurance.. Yan et al. (2026) S. Yan, G. Wang, Q. Liu, W. Meng, J. Yang, C. Yao, F. Feng, X. Ma, Y. Zhao, and Y. Han Acting while understanding: asynchronous semantic-action decoupling for real-time vision-language-action models. External Links: 2606.15285, Document, Link Cited by: External reasoning and fast–slow authority.. Yang et al. (2025) S. Yang, H. Li, B. Wang, Y. Chen, Y. Tian, T. Wang, H. Wang, F. Zhao, Y. Liao, and J. Pang InstructVLA: vision-language-action instruction tuning from understanding to manipulation. External Links: 2507.17520, Document, Link Cited by: External reasoning and fast–slow authority.. Zawalski et al. (2025) M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine Robotic control via embodied chain-of-thought reasoning. In Proceedings of The 8th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 270, p. 3157–3181. External Links: Link Cited by: Introduction, External reasoning and fast–slow authority.. Zhang et al. (2026) H. Zhang, S. Li, Y. Zhang, Z. Huai, H. Chen, C. Shen, J. Gong, and X. Qiu CoRE-VLA: towards scalable and robust vision-language-action modeling via conditional routing of experts. External Links: 2607.03693, Document, Link Cited by: Introduction, Selective intervention and runtime assurance.. Zhang et al. (2025) J. Zhang, Y. Guo, X. Chen, Y. Wang, Y. Hu, C. Shi, and J. Chen HiRT: enhancing robotic control with hierarchical robot transformers. In Proceedings of The 8th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 270, p. 933–946. External Links: Link Cited by: External reasoning and fast–slow authority.. Zhang et al. (2026) Y. Zhang et al. Harness VLA: steering frozen VLAs into reliable manipulation primitives via memory-guided agents. External Links: 2607.08448, Document, Link Cited by: Introduction, Inspectable memory at the policy boundary.. Zitkovich et al. (2023) B. Zitkovich et al. RT-2: vision-language-action models transfer web knowledge to robotic control. In Proceedings of The 7th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 229, p. 2165–2183. External Links: Link Cited by: Introduction, Frozen policies and inference-time interfaces.. Zou et al. (2023) A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043. External Links: Link Cited by: Prompt brittleness beyond language generation.. Zou et al. (2025) T. Zou, H. Zeng, Y. Nong, Y. Li, K. Liu, H. Yang, X. Ling, X. Li, and L. Ma Asynchronous fast-slow vision-language-action policies for whole-body robotic manipulation. External Links: 2512.20188, Document, Link Cited by: External reasoning and fast–slow authority..