Paper deep dive
Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction
Yi Liu
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent. We develop a decision-resource view of terminal symmetry: process evidence supplies directionality, terminal correspondence transports that structure across equivalent outcomes, realized-state evidence refines its current decision relevance after transitions, and a fixed verifier certifies execution. This decomposition yields transport--refine--certify. \method{} instantiates the principle with an episode-fixed transported process structure, its state-restricted process rank, a state-dependent residual rank refreshed after accepted transitions, and an ordinal rank meet whose top-$k$ set is exactly the union of the two proposal prefixes. The meet provides a completion guarantee under prefix coverage and attains the tight worst-case verifier-query bound under the corresponding prefix information model; a two-state construction predicts a strict post-transition dynamic--static separation. Across CAD assembly, Mini-Programs, and exact-fill packing, statewise refresh improves anytime AUC by up to $6.77$, $21.75$, and $8.68$ points, respectively. On 1,135 target-removal episodes from the official GRN OOD scenes, \method{} attains the lowest mean capped verifier cost at all three scales among the compared GRN and CDGS-style planners. The statewise signal also transfers across aggregation and scheduler organizations. Terminal symmetry thereby becomes a reusable decision resource for directed construction.
Tags
Links
- Source: https://arxiv.org/abs/2608.11318v1
- Canonical: https://arxiv.org/abs/2608.11318v1
Trouble viewing inline? Open PDF directly →
Full Text
102,518 characters extracted from source content.
Expand or collapse full text
Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction Yi Liu Affiliation: MS Student Affiliation: School of Astronomy and Space Science Affiliation: University of Science and Technology of China Email: scnuliuyi@mail.ustc.edu.cn Abstract Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent. We develop a decision-resource view of terminal symmetry: process evidence supplies directionality, terminal correspondence transports that structure across equivalent outcomes, realized-state evidence refines its current decision relevance after transitions, and a fixed verifier certifies execution. This decomposition yields transport–refine–certify. SymBuild instantiates the principle with an episode-fixed transported process structure, its state-restricted process rank, a state-dependent residual rank refreshed after accepted transitions, and an ordinal rank meet whose top-k set is exactly the union of the two proposal prefixes. The meet provides a completion guarantee under prefix coverage and attains the tight worst-case verifier-query bound under the corresponding prefix information model; a two-state construction predicts a strict post-transition dynamic–static separation. Across CAD assembly, Mini-Programs, and exact-fill packing, statewise refresh improves anytime AUC by up to 6.776.77, 21.7521.75, and 8.688.68 points, respectively. On 1,135 target-removal episodes from the official GRN OOD scenes, SymBuild attains the lowest mean capped verifier cost at all three scales among the compared GRN and CDGS-style planners. The statewise signal also transfers across aggregation and scheduler organizations. Terminal symmetry thereby becomes a reusable decision resource for directed construction. 1 Introduction Many sequential tasks end in objects with exact symmetries even when the processes that realize them are directed and history-dependent. A chair may be unchanged when two legs are exchanged, a program output may be invariant to renaming, and an exact-fill packing is unchanged when bins are relabeled; yet blocker order, dependency order, and remaining capacity can make only some next actions useful. This raises a basic question: what decision information can exact terminal symmetry provide when the realizing process itself is asymmetric? The key distinction is informational. Process evidence is the source of directionality; terminal correspondence carries that structure into an equivalent target frame; realized-state evidence refines its current decision relevance after transitions; and a fixed verifier certifies execution. This source–carrier–state decomposition turns terminal symmetry into a reusable decision resource for directed construction. Symmetry has long been used to reduce or guide planning search and to simplify assembly reasoning (8; 7; 19); modern methods encode symmetry in state–action models, planning operators, trajectories, or canonical representations (31; 36; 29; 12). We focus on the regime where exact symmetry is available only at terminal outcomes: directed process structure is grounded independently in demonstrations with appropriate coverage, transition queries, geometry, or execution feedback, while terminal correspondence supplies the reusable transport map. Theorem 1 formalizes the identification–transport separation: the same symmetric terminal object and observed demonstration can coexist with opposite commutativity and precedence relations. This separation yields statewise refinement: transport what the outcome preserves, and refresh what realized history changes. A transported process prior records which asymmetric order was useful in a symmetric reference frame; a local residual records what the current realized state makes newly accessible, blocked, satisfied, or capacity-limited. Statewise refinement refreshes this transient state evidence after accepted transitions while keeping the transported prior semantically fixed. SymBuild uses the resulting proposal order to allocate queries to the fixed verifier. Figure 1 shows the predicted signature: dynamic and static policies take the same first action, then refresh redirects verification after the realized state changes and reaches a verified completion within the same budget. Figure 1: Terminal correspondence transports directed process knowledge; realized state refines its decision value after accepted transitions. (A) Process evidence grounds an asymmetric prior and a terminal automorphism transports it into the target frame. (B) After the shared first action changes feasibility, a cached composition becomes stale and the verifier rejects its next proposal. (C) SymBuild refreshes the state residual and reaches a verified plan under the same budget and verifier. Our contributions are: • Decision-resource formulation. We formulate terminal-symmetric, process-asymmetric construction as an information-flow problem: process evidence identifies directionality, terminal correspondence transports it, and realized-state evidence refines its decision relevance. • Statewise refinement and theory. This decomposition yields transport–refine–certify. SymBuild realizes the resulting union-prefix semantics with an ordinal rank meet, a completion guarantee under prefix coverage, a tight worst-case verifier-query bound under the prefix information model, and a strict post-transition dynamic–static separation. • Cross-domain evidence and transfer. CAD, program construction, packing, and GRN OOD experiments show improved verified efficiency, the predicted temporal signature, distinct transport/state contributions, and transfer across aggregation and scheduler organizations; SymBuild has the lowest mean capped verifier cost at all three GRN OOD scales among the compared direct and CDGS-style planners. 2 Related Work Symmetry as shared process structure. Equivariant RL and differentiable planning encode symmetry in state–action values, policies, planning operators, or action-sampling procedures (31; 32; 36; 24; 35). Symmetry-aware sequential, trajectory, and generative methods extend related structure to solution trajectories, structured generation, or policy spaces (3; 29; 37; 13; 9). In these lines, symmetry is part of the structure shared by transformed states, actions, computations, or trajectories. Symmetry in planning and assembly. Classical planning exploited exact problem symmetries to reduce equivalent search (8; 25) and abstraction-derived almost symmetry to proactively guide action ordering (7). Assembly planning likewise used component symmetries and group theory to simplify kinematic/spatial constraint inference and geometric assembly reasoning (19; 26). This lineage establishes symmetry as active computational guidance. Building on it, the decision-resource formulation centers the symmetry locus on terminal outcomes and separates the source of directionality from the correspondence that transports it: process evidence identifies the directed relation, terminal correspondence carries it, and realized state refines its current relevance. Representative selection and imperfect symmetry. Canonicalization chooses representatives of symmetry orbits (12; 17; 38); soft and approximate equivariance relax hard symmetry constraints under mismatch (6; 33), while probabilistic symmetry breaking and partial equivariance control how symmetry is broken or selectively applied (14; 4). These works adapt the representation, strength, or locality of equivariance; in our setting the terminal correspondence remains exact, while statewise refinement updates the decision relevance of transported asymmetric process information. Appendix B gives a systematic operational comparison, including symmetry-based trajectory reuse (18; 15). Search under expensive evaluation. Partial-order planning and partial-order reduction exploit ordering structure (22; 34), while anytime A*, LAMA, LazySP, and Multi-Heuristic A* organize search or defer expensive evaluation once scores, constraints, or evaluators are available (16; 27; 10; 5; 1). The decision-resource view supplies an information layer before those choices: transported process structure and realized-state evidence form the proposal signals consumed by a fixed verifier. Section 4.5 instantiates the same signals in canonical, partial-order, multi-heuristic, and LazySP-style interfaces; Appendix C.1 gives their definitions. Public SayCanPay, GRN, and CDGS-style comparisons test aggregation and wider guided search; Appendix Table 16 summarizes those interfaces. 3 Terminal Symmetry as a Decision Resource 3.1 Terminal Correspondence as Transport The decision-resource view begins with a transport operation: terminal symmetry provides a correspondence between equivalent outcomes, while directed process structure is supplied independently by process evidence. 3.1.1 Construction systems and terminal lifts A discrete construction system is =(,,F,s0,Ψ)C=(S,A,F,s_0, ), where F:×⇀F:S×A is a deterministic partial transition and Ψ maps a final state to a semantic object. A complete executable trajectory is τ=(a1,…,aT)∈τ=(a_1,…,a_T) _C and its terminal object is O(τ)=Ψ(sT)O_C(τ)= (s_T). Let a group G act on terminal objects. A potentially partial and set-valued lift g(τ) L_g(τ) satisfies O(τ′)=g⋅O(τ)∀τ′∈g(τ).O_C(τ )=g· O_C(τ) ∀τ ∈ L_g(τ). (1) Equation 1 specifies covariance of the terminal object. When a selected lift preserves length and supplies an event bijection PgP_g, a directed relation tensor R(τ)R(τ) transforms by conjugation, R(τ′)=PgR(τ)Pg⊤.R(τ )=P_gR(τ)P_g . (2) Conjugation transports precedence, commutativity, or action correspondence while preserving directionality. Equation 2 gives the transport condition across terminal frames; equality R=PgRPg⊤R=P_gRP_g within one realized trajectory corresponds to the stronger stabilizer-invariance property. 3.1.2 Identification–Transport Separation Identification and transport are distinct information operations. Equation 2 specifies how an identified directed relation moves across equivalent outcomes; identifying that relation requires process evidence beyond terminal observations. Theorem 1 (Identification–transport separation). Over finite deterministic partial-transition construction systems, no estimator that observes only a terminal object and its stabilizer can universally recover true precedence or commutativity. The result still holds when the estimator observes one demonstration whose sampling policy need not cover counterfactual action orders. On a two-system witness class, every randomized estimator has worst-case error at least 1/21/2. A two-system witness shares the same terminal object, stabilizer, and observed demonstration, yet one system admits a commuting diamond while the other forces a chain; Appendix A gives the construction and proof. The theorem establishes two distinct operations: process evidence identifies directed relations, terminal correspondence transports identified relations between equivalent terminal frames according to equation 2, and statewise refinement updates their decision value after transitions. 3.2 Statewise Refinement 3.2.1 Three information roles, two timescales, one verifier The architecture contains three information roles and one certification role. Process evidence supplies directionality and terminal correspondence carries it into the target frame, yielding an episode-fixed transported process structure RTR_T. At state s with candidates AsA_s, restricting RTR_T to the current candidate set induces the state-restricted process rank rTs:As→[1,|As|]r_T^s:A_s→[1,|A_s|]. Realized-state evidence supplies the state-dependent residual rank rXsr_X^s, refreshed after accepted actions as geometry, capacity, dependencies, or extension viability change. A fixed verifier Vs(a)V_s(a) certifies execution. The ranks are algorithmically comparable but informationally distinct. A proposal should remain early whenever either channel ranks it early, because transport and state evidence resolve different uncertainties. The ordinal operator with this union-prefix semantics is ms(a)=minrTs(a),rXs(a).m_s(a)= \r_T^s(a),r_X^s(a)\. (3) Lemma 1 gives the exact prefix identity. The resulting planner, SymBuild, tests actions in nondecreasing msm_s, accepts the first action approved by VsV_s, re-restricts the fixed RTR_T to current candidates after a transition, and refreshes rXsr_X^s from new realized-state evidence; the initial-static control caches the initial composition. Search-interface view. At the algorithmic interface, rTsr_T^s is the decision projection of transported process structure and rXsr_X^s is the decision projection of realized-state evidence. Search methods may consume them as proposal heuristics, while their semantic roles remain distinct: multi-heuristic scheduling keeps two queues, POCL-style scheduling constrains an agenda, LazySP-style selection searches verifier-free prefixes, and rank meet compiles the two prefixes into one order. Section 4.5 tests these mappings across planning interfaces. Figure 2: Mechanism of SymBuild. Process evidence grounds the directed prior; terminal correspondence transports it once; realized-state evidence refreshes decision relevance after accepted transitions. Their ordinal meet orders proposals to a fixed verifier. Algorithm 1 gives the executable interface behind Figure 2. The transported structure and model parameters remain fixed; accepted transitions restrict that structure to the new candidate set, refresh the residual rank, and rebuild the candidate queue. Algorithm 1 SymBuild: statewise residual rank meet 1: target X, correspondence PgP_g, reference RrefR_ ref, state s0s_0, verifier V, resource cap B 2: RT←PgRrefPg⊤R_T← P_gR_ refP_g ; s←s0s← s_0; q←0q← 0; τ←()τ←() 3: while s is incomplete and q<Bq<B do 4: As←A_s← remaining actions; rT←RestrictRank(RT,As)r_T← RestrictRank(R_T,A_s) 5: rX←ResidualRank(s,As)r_X← ResidualRank(s,A_s) 6: Q←Sort(As,min(rT,rX),rT+rX,fixed tie-break)Q← Sort(A_s, (r_T,r_X),r_T+r_X,fixed tie-break) 7: accepted←falseaccepted 8: while Q≠∅Q≠ and q<Bq<B do 9: a←PopFirst(Q)a← PopFirst(Q); (y,c)←Vs(a)(y,c)← V_s(a); q←q+cq← q+c 10: if y=1y=1 then 11: s←F(s,a)s← F(s,a); τ←τ∘aτ←τ a; accepted←trueaccepted ; break 12: end if 13: end while 14: if ¬accepted then 15: return interrupted / infeasible 16: end if 17: end while 18: return τ and q Instantiations. CAD transports blocker order and refreshes geometric accessibility; Mini-Programs transport a frozen relation checkpoint and refresh satisfied/exposed dependencies; packing transports a reference assignment and refreshes item–bin viability from remaining capacities. Rank meet gives a transparent ordinal realization: it preserves candidates ranked highly by either source and exposes an exact union prefix. Section 4.5 confirms the same refresh signature with independently selected aggregation. 3.2.2 Prefix coverage and query optimality For weak rank r, define Topk(r)=a:r(a)≤kTop_k(r)=\a:r(a)≤ k\ and Br(k)=|Topk(r)|B_r(k)=|Top_k(r)|. Lemma 1 (Union coverage). For the rank meet in equation 3, Topk(ms)=Topk(rTs)∪Topk(rXs).Top_k(m_s)=Top_k(r_T^s) _k(r_X^s). (4) If either expert has a verifier-feasible action in its top-k set, meet finds a feasible action in at most BTs(k)+BXs(k)B_T^s(k)+B_X^s(k) verifier calls, or at most 2k2k for strict ranks. Under the prefix information model, if feasibility is known only to lie within the union of two expert prefixes, that union is the minimax query set, and rank meet reaches it exactly. Theorem 2 (Prefix-information optimality and completion guarantee). Let Us(k)=Topk(rTs)∪Topk(rXs)U_s(k)=Top_k(r_T^s) _k(r_X^s). Suppose every candidate set is finite, Vs(a)=1V_s(a)=1 only if F(s,a)F(s,a) admits a complete suffix, and every incomplete reachable state contains an action a∈Us(ks)a∈ U_s(k_s) with Vs(a)=1V_s(a)=1. Then SymBuild completes within ∑t=0H−1|Ust(kst)|≤∑t=0H−1(BTst(kst)+BXst(kst)) _t=0^H-1|U_s_t(k_s_t)|≤ _t=0^H-1 (B_T^s_t(k_s_t)+B_X^s_t(k_s_t) ) (5) queries; with ks=|As|k_s=|A_s|, it is complete relative to the finite candidate graph and extension-preserving verifier. At one state, if the only promise is that Us(k)U_s(k) contains a feasible action, every deterministic query policy requires |Us(k)||U_s(k)| calls in the worst case, and meet attains this bound. Lemma 1 and Theorem 2 show that rank meet fully exploits the union-prefix information exposed by the two channels, while the fixed verifier preserves acceptance semantics. Appendix A gives the proofs and covariance details. 3.3 Anytime Value Under Expensive Verification Statewise refinement reallocates expensive verification, so performance is naturally measured by completion before an uncertain resource limit. Let QπQ_π be verifier resource to a complete verified solution and Fπ(b)=Pr(Qπ≤b)F_π(b)= (Q_π≤ b). For a predeclared budget distribution μ, AUCμ(π)=∑b∈ℬμ(b)Fπ(b)=Prω,B(Qπ(ω)≤B),B∼μ.AUC_μ(π)= _b μ(b)F_π(b)= _ω,B(Q_π(ω)≤ B), B μ. (6) Thus anytime AUC is expected success under random interruption; Appendix A gives the formal proposition, proof, and capped-cost identity. Proposition 1 (Strict dynamic–static separation). There exists a two-state problem in which dynamic and initial-static policies use the same terminal prior, fixed verifier, and first action, yet dynamic completes in two calls while static is interrupted at budget two. After the shared first action, the cached order tests infeasible a, while the refreshed residual promotes feasible b; under budget two, dynamic completes and static is interrupted. Figure 1 gives the conceptual demonstration, Appendix A the formal construction, and Figure 4 the empirical signature. 4 Experiments Experiments test four consequences of transport–refine–certify: verified efficiency across domains and learned planning, the predicted post-transition signature, distinct transport/state contributions, and transfer across planning organizations. 4.1 Domains and protocol We use CAD assembly (72 call objects plus an independent 48-object work panel), 480 Mini-Programs, 480 symmetric exact-fill packings, and 1,135 solvable target-removal episodes from official GRN OOD scenes at 10/15/20 components. All paired comparisons share instances, candidate sets, observable information, deterministic tie-breaking, verifier semantics, and budget definitions; temporal comparisons additionally share the first accepted action by construction or pairing. Primary CAD/program/packing intervals use 10,00010,000 paired bootstrap draws, and GRN targets use a common 5N5N budget. Domain construction, verifier definitions, budgets, seeds, and inference are specified in Appendix C and Table 13; planning-interface controls are specified in Appendix C.1. 4.2 Verified efficiency across domains and GRN Table 1: Matched comparisons across three construction domains. Primary pairs share channels, candidates, first actions, and verifier; initial-static caches the composition. The public-product column repeats refresh–static with the SayCanPay rule. Entries are AUC percentages with paired 95% CIs. Primary rank meet Public product Domain Shift/resource N Initial-static SymBuild Δ [95%CI][95\%\,CI] Cost ↓ Δ [95%CI][95\%\,CI] CAD ID / calls 72 78.73 81.60 +2.86+2.86 [.12,5.76][.12,5.76] .0253 +3.30+3.30 [1.01,5.79][1.01,5.79] CAD ID / work 48 91.93 98.70 +6.77+6.77 [1.30,14.06][1.30,14.06] .0523 +4.43+4.43 [.52,9.90][.52,9.90] Program ID / calls 240 72.67 85.08 +12.42+12.42 [9.50,15.33][9.50,15.33] .0737 +12.42+12.42 [9.50,15.33][9.50,15.33] Program length OOD 240 42.92 64.67 +21.75+21.75 [18.33,24.92][18.33,24.92] .1351 +22.58+22.58 [19.33,25.83][19.33,25.83] Packing ID / calls 240 49.38 57.92 +8.54+8.54 [7.01,10.14][7.01,10.14] .0539 +6.46+6.46 [4.51,8.47][4.51,8.47] Packing size OOD 240 25.35 34.03 +8.68+8.68 [7.01,10.42][7.01,10.42] .0981 +6.53+6.53 [4.58,8.47][4.58,8.47] Table 2: Official GRN 2025 integration and 2026 CDGS-style comparisons. Mean capped verifier queries (↓ ) on the official checkpoint and OOD data. Panel (a) compares cached/statewise scoring; panel (b) widens the same interface with population guided search. (a) Official GRN 2025 model/data integration Split Targets GRN static GRN refresh Combined static SymBuild SymBuild–static gap [95%CI][95\%\,CI] OOD10 129 5.457 5.132 3.349 2.597 −.752-.752 [−1.279,−.270][-1.279,-.270] OOD15 432 8.213 6.734 4.523 3.347 −1.175-1.175 [−1.502,−.867][-1.502,-.867] OOD20 574 11.027 8.028 6.332 3.873 −2.459-2.459 [−3.225,−1.774][-3.225,-1.774] (b) CDGS-style population guided search Split Targets CDGS–GRN CDGS–static CDGS–refresh SymBuild Refresh–static gap [95%CI][95\%\,CI] OOD10 129 5.163 4.202 3.194 2.597 −1.008-1.008 [−2.031,−.203][-2.031,-.203] OOD15 432 8.515 5.844 3.828 3.347 −2.016-2.016 [−2.612,−1.463][-2.612,-1.463] OOD20 574 8.482 7.489 4.158 3.873 −3.331-3.331 [−4.224,−2.493][-4.224,-2.493] Statewise refresh improves verified efficiency in all six pre-specified cross-domain aggregates (Table 1). The same principle remains effective inside a contemporary learned/search stack: on 1,135 target-removal episodes from official GRN OOD scenes, SymBuild gives the lowest mean capped verifier-query count at 10, 15, and 20 components among the compared direct and CDGS-style planners, while evaluating only 2.002.00–3.143.14 GRN states per target versus 13.1413.14–38.3238.32 for CDGS–refresh (Table 2; Appendix Table 17). Figure 3: Statewise refinement shifts the verified-success frontier across the full predeclared budget range. Strong-expert and permuted controls appear where defined; the flat CAD-work curves isolate rescue rate. The gain persists across the full predeclared budget region (Figure 3). Exact frontier values are in Appendix E; full subgroup replications are in Appendix Table 19. 4.3 Predicted post-transition signature Proposition 1 predicts when statewise refinement should matter: the first accepted decision matches, and divergence begins after an accepted transition changes residual relevance. All 384384 CAD-work and 480480 packing pairs share the first accepted decision; later sequences diverge in 93.75%93.75\% of CAD-work and 67.50/93.75%67.50/93.75\% of packing ID/OOD pairs, yielding 26/0 and 29/1 dynamic/static-only completions on CAD work and packing OOD. The gain appears exactly where the state residual is refreshed. Product scoring and an independent CAD panel also reproduce the shared-first-action/post-transition-divergence pattern (Tables 18 and 3(b)). Figure 4: Statewise refresh acts after the shared first decision. Accepted sequences diverge after the transition, and dynamic-only verified completions dominate. 4.4 Distinct roles of transport and state evidence Information roles are domain-dependent: CAD and packing are transport-sensitive, whereas Mini-Programs are state-dominant; statewise refinement improves all six primary matched domain comparisons. Timing diagnoses refinement and alignment diagnoses transport (Table 3); graded correspondence mismatch and complete comparisons are in Appendix Figure 5, Table 6, and Appendix D. Table 3: Information-role evidence. (a) Timing/single-source ablations; (b) independent CAD and observed-process tests. (a) Temporal and information-source ablations Domain Panel N Full Static-once Process-only State-only CAD calls 72 81.60 78.73 79.57 73.12 CAD work 48 98.70 91.93 87.50 85.63 Program ID 240 85.08 72.67 84.83 85.17 Program length OOD 240 64.67 42.92 63.58 65.83 Packing ID 240 57.92 49.38 50.21 38.26 Packing size OOD 240 34.03 25.35 28.40 21.67 (b) Independent mechanism and observed-process evidence Evidence Panel Result CAD algorithm 120 objects Rank meet 84.58%84.58\%; +5.00/+7.71+5.00/+7.71 points, −.143/−.168-.143/-.168 calls vs. PCS-ICP/outsideness. CAD timing 120 objects Dynamic/static 84.38/80.00%84.38/80.00\%, 1.569/1.7211.569/1.721 calls; reduction .152.152 [.064,.246][.064,.246]; disagreement 96.46%96.46\%. CAD alignment 60 objects × 4 lifts Aligned/permuted 81.25/65.00%81.25/65.00\%; +16.25+16.25 [8.33,25.00][8.33,25.00] points; calls reduced .637.637 [.501,.782][.501,.782]. IKEA topology 36 families, 97 videos, 1,022 states Process structure +14.14+14.14 points [6.41,20.18][6.41,20.18]; action core +19.20+19.20 [10.26,25.56][10.26,25.56]; hybrid +2.64+2.64 [.77,4.38][.77,4.38]. 4.5 Transfer Across Planning Organizations The same signals transfer across canonical, replanning, partial-order, and multi-heuristic interfaces. On ID/length OOD, SymBuild gains 12.58/21.7512.58/21.75 AUC points over canonical-slice static planning and 4.42/4.254.42/4.25 over weighted-A*/LazySP-style replanning; within POCL-style and multi-heuristic schedulers, refresh adds 3.833.83 and 5.005.00 length-OOD points. Rank meet retains Theorem 2’s prefix guarantee. Table 4: Transfer across planning organizations. AUC gaps with paired 95% CIs; full results are in Appendix C.1. Comparison ID Length OOD Interface feature SymBuild −- canonical-slice static +12.58+12.58 [9.67,15.50][9.67,15.50] +21.75+21.75 [18.42,25.17][18.42,25.17] canonical representative SymBuild −- weighted-A*/LazySP-style +4.42+4.42 [2.58,6.33][2.58,6.33] +4.25+4.25 [1.83,6.67][1.83,6.67] lazy replanning POCL-style refresh −- fixed Appendix +3.83+3.83 [2.58,5.08][2.58,5.08] partial-order agenda Multi-heuristic refresh −- fixed Appendix +5.00+5.00 [2.83,7.17][2.83,7.17] two-queue scheduling Aggregation and deployment transfer. Product aggregation is positive across all six panels; validation-selected weighted sum adds +13.50/+24.17+13.50/+24.17 ID/length-OOD AUC. Appendix C.2 reports approximate-verifier robustness and verifier-latency break-even. 5 Conclusion Terminal symmetry is a decision resource for directed construction: process evidence supplies directionality, terminal correspondence transports it, realized state refines relevance, and the verifier certifies execution. SymBuild realizes this flow with union-prefix rank meet and a tight verifier-query bound under the prefix information model. The predicted post-transition signature persists across domains, while the same statewise signal transfers across planning organizations; SymBuild also gives the lowest mean capped verifier cost at all three GRN OOD scales. Transport what the outcome preserves; refine what history changes. References Aine et al. (2016) S. Aine, S. Swaminathan, V. Narayanan, V. Hwang, and M. Likhachev Multi-heuristic A*. The International Journal of Robotics Research. Cited by: §2. Ait Bouhsain et al. (2025) S. Ait Bouhsain, R. Alami, and T. Simeon Learning geometric reasoning networks for robot task and motion planning. In International Conference on Learning Representations, Cited by: Appendix C, §C.2, §C.2, Table 16. Brehmer et al. (2023) J. Brehmer, J. Bose, P. de Haan, and T. S. Cohen EDGI: equivariant diffusion for planning with embodied agents. In Advances in Neural Information Processing Systems, Cited by: §2. Chang et al. (2026) J. Chang, M. Park, J. Seo, R. Horowitz, J. Lee, and J. Choi Partially equivariant reinforcement learning in symmetry-breaking environments. In International Conference on Learning Representations, Cited by: Table 5, §2. Dellin and Srinivasa (2016) C. M. Dellin and S. S. Srinivasa A unifying formalism for shortest path problems with expensive edge evaluations via lazy best-first search over paths with edge selectors. In Proceedings of the International Conference on Automated Planning and Scheduling, Cited by: §2. Finzi et al. (2021) M. Finzi, G. Benton, and A. G. Wilson Residual pathway priors for soft equivariance constraints. In Advances in Neural Information Processing Systems, Vol. 34. Cited by: Appendix B, §2. Fox et al. (2005) M. Fox, D. Long, and J. Porteous Abstraction-based action ordering in planning. In Proceedings of the Nineteenth International Joint Conference on Artificial Intelligence, p. 1214–1219. Cited by: Table 5, Appendix B, §1, §2. Fox and Long (1999) M. Fox and D. Long The detection and exploitation of symmetry in planning problems. In Proceedings of the Sixteenth International Joint Conference on Artificial Intelligence, p. 956–961. Cited by: Table 5, §1, §2. Hadjiloizou et al. (2026) L. Hadjiloizou, R. Pérez-Dattari, and N. Jaquier Symmetries here and there, combined everywhere: cross-space symmetry compositions in robotics. arXiv preprint arXiv:2605.22639. Cited by: Table 5, §2. Hansen and Zhou (2007) E. A. Hansen and R. Zhou Anytime heuristic search. Journal of Artificial Intelligence Research. Cited by: §2. Hazra et al. (2024) R. Hazra, P. Zuidberg Dos Martires, and L. De Raedt SayCanPay: heuristic planning with large language models using learnable domain knowledge. In Proceedings of the AAAI Conference on Artificial Intelligence, Cited by: Appendix C, Table 15, Table 16. Kaba et al. (2023) S. Kaba, A. K. Mondal, Y. Zhang, Y. Bengio, and S. Ravanbakhsh Equivariance with learned canonicalization functions. In Proceedings of the 40th International Conference on Machine Learning, Cited by: §1, §2. Kim et al. (2025) H. Kim, S. Lee, and M. Oh Symmetry-aware GFlowNets. In Proceedings of the 42nd International Conference on Machine Learning, Cited by: Appendix B, §2. Lawrence et al. (2025) H. Lawrence, V. Portilheiro, Y. Zhang, and S. Kaba Improving equivariant networks with probabilistic symmetry breaking. In International Conference on Learning Representations, Cited by: §2. Lee et al. (2023) B. Lee, Y. Lee, S. Kim, M. Son, and F. C. Park Equivariant motion manifold primitives. In Proceedings of the 7th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 229, p. 1199–1221. Cited by: Table 5, §2. Likhachev et al. (2003) M. Likhachev, G. J. Gordon, and S. Thrun ARA*: anytime A* with provable bounds on sub-optimality. In Advances in Neural Information Processing Systems, Cited by: §2. Lin and Levie (2026) Y. E. Lin and R. Levie Adaptive canonicalization with application to invariant anisotropic geometric networks. In International Conference on Learning Representations, Cited by: §2. Lin et al. (2020) Y. Lin, J. Huang, M. Zimmer, Y. Guan, J. Rojas, and P. Weng Invariant transform experience replay: data augmentation for deep reinforcement learning. IEEE Robotics and Automation Letters 5 (4), p. 6615–6622. External Links: Document Cited by: Table 5, §2. Liu and Popplestone (1990) Y. Liu and R. J. Popplestone Symmetry constraint inference in assembly planning: automatic assembly configuration specification. In Proceedings of the Eighth National Conference on Artificial Intelligence, p. 1038–1044. Cited by: Table 5, §1, §2. Liu et al. (2024) Y. Liu, C. Eyzaguirre, M. Li, S. Khanna, J. C. Niebles, V. Ravi, S. Mishra, W. Liu, and J. Wu IKEA Manuals at Work: 4D grounding of assembly instructions on internet videos. In Advances in Neural Information Processing Systems, Vol. 37. Cited by: Appendix C. Matada et al. (2025) S. Matada, L. Bhan, Y. Shi, and N. Atanasov Generalizable motion planning via operator learning. In International Conference on Learning Representations, Cited by: Table 16. McAllester and Rosenblitt (1991) D. McAllester and D. Rosenblitt Systematic nonlinear planning. In Proceedings of the Ninth National Conference on Artificial Intelligence, Cited by: §2. Mishra et al. (2026) U. A. Mishra, D. He, Y. Chen, and D. Xu Compositional diffusion with guided search for long-horizon planning. In International Conference on Learning Representations, Cited by: Appendix C, Appendix D. Nguyen et al. (2023) H. H. Nguyen, A. Baisero, D. Klee, D. Wang, R. Platt, and C. Amato Equivariant reinforcement learning under partial observability. In Proceedings of the 7th Conference on Robot Learning, Cited by: §2. Pochter et al. (2011) N. Pochter, A. Zohar, and J. S. Rosenschein Exploiting problem symmetries in state-based planners. In Proceedings of the Twenty-Fifth AAAI Conference on Artificial Intelligence, Cited by: §2. Popplestone et al. (1990) R. J. Popplestone, Y. Liu, and R. Weiss A group theoretic approach to assembly planning. AI Magazine 11 (1), p. 82–97. Cited by: §2. Richter and Westphal (2010) S. Richter and M. Westphal The LAMA planner: guiding cost-based anytime planning with landmarks. Journal of Artificial Intelligence Research. Cited by: §2. Tian et al. (2024) Y. Tian, K. D. D. Willis, B. Al Omari, J. Luo, P. Ma, Y. Li, F. Javid, E. Gu, J. Jacob, S. Sueda, H. Li, S. Chitta, and W. Matusik ASAP: automated sequence planning for complex robotic assembly with physical feasibility. In IEEE International Conference on Robotics and Automation, Cited by: Appendix C, Appendix C, Appendix C, Appendix C. Tie et al. (2025) C. Tie, Y. Chen, R. Wu, B. Dong, Z. Li, C. Gao, and H. Dong ET-SEED: efficient trajectory-level SE(3) equivariant diffusion policy. In International Conference on Learning Representations, Cited by: §1, §2. Todorov et al. (2012) E. Todorov, T. Erez, and Y. Tassa MuJoCo: a physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, p. 5026–5033. External Links: Document Cited by: Appendix C, Appendix C. van der Pol et al. (2020) E. van der Pol, D. E. Worrall, H. van Hoof, F. A. Oliehoek, and M. Welling MDP homomorphic networks: group symmetries in reinforcement learning. In Advances in Neural Information Processing Systems, Cited by: Appendix B, §1, §2. Wang et al. (2022a) D. Wang, R. Walters, X. Zhu, and R. Platt Equivariant Q learning in spatial action spaces. In Proceedings of the 5th Conference on Robot Learning, Proceedings of Machine Learning Research, Vol. 164, p. 1713–1723. Cited by: Appendix B, §2. Wang et al. (2022b) R. Wang, R. Walters, and R. Yu Approximately equivariant networks for imperfectly symmetric dynamics. In Proceedings of the 39th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 162, p. 23078–23091. Cited by: Appendix B, §2. Xu et al. (2011) Y. Xu, Y. Chen, Q. Lu, and R. Huang Theory and algorithms for partial order based reduction in planning. arXiv preprint arXiv:1106.5427. Cited by: §2. Zhao et al. (2024a) L. Zhao, O. Howell, X. Zhu, J. Y. Park, Z. Zhang, R. Walters, and L. L. S. Wong Equivariant action sampling for reinforcement learning and planning. arXiv preprint arXiv:2412.12237. Cited by: §2. Zhao et al. (2023) L. Zhao, X. Zhu, L. Kong, R. Walters, and L. L. S. Wong Integrating symmetry into differentiable planning with steerable convolutions. In International Conference on Learning Representations, Cited by: Table 5, §1, §2. Zhao et al. (2024b) L. Zhao, X. Ding, and L. Akoglu PARD: permutation-invariant autoregressive diffusion for graph generation. In Advances in Neural Information Processing Systems, Cited by: Appendix B, §2. Zhou et al. (2026) C. Zhou, Z. Chen, Z. Li, J. Wang, K. Jiang, P. Li, R. Yu, M. Zhang, S. Bates, and T. Jaakkola Rethinking diffusion models with symmetries through canonicalization with applications to molecular graph generation. arXiv preprint arXiv:2602.15022. Cited by: Table 5, §2. Appendix A Proofs A.1 Proof of Theorem 1 Let both systems use action set a,b\a,b\ and the same terminal object o with a nonidentity automorphism g that exchanges two indistinguishable terminal roles, so g⋅o=og· o=o and the terminal stabilizers coincide. In ∥C_ , define the diamond transitions s0→sa→sTs_0 as_a bs_T and s0→sb→sTs_0 bs_b as_T; then a and b commute at s0s_0. In →C_→, define only q0→qa→qTq_0 aq_a bq_T, so a≺ba b is forced. Thus the same nontrivial terminal symmetry is compatible with different process relations. Let the behavior policy emit abab with probability one in both systems. The terminal-only and one-demonstration observation distributions are therefore identical. Fix the shared observation z. If a randomized estimator outputs the first relation with probability p and the second with probability q, then p+q≤1p+q≤ 1. Its errors on the two systems are at least 1−p1-p and 1−q1-q, hence max1−p,1−q≥1−p+q2≥12. \1-p,1-q\≥ 1- p+q2≥ 12. (7) Thus no universally consistent estimator exists. Repeated terminal-only samples preserve the identical point-mass observation distribution. A.2 Proof of Lemma 1 For every action a, ms(a)≤k⇔minrTs(a),rXs(a)≤k⇔rTs(a)≤korrXs(a)≤k.m_s(a)≤ k \r_T^s(a),r_X^s(a)\≤ k r_T^s(a)≤ k\ or\ r_X^s(a)≤ k. (8) This proves the set identity. Its cardinality is at most BTs(k)+BXs(k)B_T^s(k)+B_X^s(k). If a feasible action is in the union, an ordering that exhausts meet ranks through k must encounter one within that many queries. Strict ranks give BT=BX=kB_T=B_X=k. A.3 Proof of Theorem 2 At each incomplete state, Lemma 1 places every member of Us(ks)U_s(k_s) before every action outside it. Coverage therefore supplies an accepted member within |Us(ks)||U_s(k_s)| calls. Extension preservation leaves a complete suffix after the transition, so induction over the finite horizon reaches a terminal state; summing the prefix sizes proves the query bound. When ks=|As|k_s=|A_s|, Us(ks)=AsU_s(k_s)=A_s, giving relative completeness. For optimality, fix a state and any deterministic query policy. An adversary labels the first |Us(k)|−1|U_s(k)|-1 queried members of the union infeasible and the last member feasible. This labeling satisfies the sole coverage promise and forces |Us(k)||U_s(k)| calls. Meet exhausts exactly this union before leaving the prefix and hence matches the lower bound. Verifier safety and covariance. Proposal order does not alter the acceptance predicate, so the planner remains verifier-safe. If a group element induces a bijection Pg:As→AgsP_g:A_s→ A_gs and both ranks are covariant, rTgs(Pga)=rTs(a)r_T^gs(P_ga)=r_T^s(a) and rXgs(Pga)=rXs(a)r_X^gs(P_ga)=r_X^s(a), then the meet is covariant as well: mgs(Pga)=ms(a)m_gs(P_ga)=m_s(a). Equality among orbit actions within one realized state additionally requires stabilizer invariance; covariance alone concerns transformed states. A.4 Random-interruption identity Conditioning on the independently drawn budget gives Pr(Qπ≤B)=∑b∈ℬPr(B=b)Pr(Qπ≤b)=∑b∈ℬμ(b)Fπ(b). (Q_π≤ B)= _b (B=b) (Q_π≤ b)= _b μ(b)F_π(b). (9) Pointwise CDF dominance implies dominance for every nonnegative weighting. Conversely, choosing a point mass at each b recovers each pointwise inequality. A.5 Proof of Proposition 1 Consider an initial state with a shared first action p. After p is accepted, two final actions a,ba,b remain. Let the transported prior tie them and let the fixed static tie-break order a before b. In the realized successor state set V(a)=0V(a)=0 and V(b)=1V(b)=1, while the refreshed residual orders b before a. Under resource budget two, both policies spend the first call on p; dynamic spends the second on b and completes, whereas static spends the second on a and is interrupted. The prior, verifier, and first action are identical, so the separation is caused solely by post-transition refresh. A.6 Capped cost identity For dense integer budgets, anytime completion is also linked to capped time-to-solution. For nonnegative integer q and cap M, minq,M=∑b=0M−1q>b \q,M\= _b=0^M-1I\q>b\. Taking expectations and using Pr(Q>b)=1−F(b) (Q>b)=1-F(b) gives 1−[minQπ,M]M=1M∑b=0M−1Fπ(b).1- E[ \Q_π,M\]M= 1M _b=0^M-1F_π(b). (10) Appendix B Operational Relation to Prior Symmetry Uses Table 5 compares the symmetry locus and information role of the closest historical and modern precedents. The work of 7 is particularly important historically: they use abstraction-derived almost symmetry as proactive action-ordering guidance. Our formulation assigns symmetry a more specific carrier role: the directed relation is identified independently from process evidence, terminal correspondence transports that relation across equivalent outcomes, and realized-state evidence refines its decision relevance after transitions. Table 5: Operational comparison of symmetry uses. The key distinction is where symmetry is assumed and what information it contributes to decision making. Work Symmetry locus Primary use Process assumption / information source Relation to our setting 8 Planning problem objects/actions Detect and exploit symmetric search structure Exact problem symmetry exposes equivalent choices Classical symmetry reduction over equivalent planning choices; SymBuild uses terminal correspondence to carry independently grounded directed process structure 7 Abstract planning problem; almost-symmetric objects/actions Proactive action ordering in forward search Guidance is induced from abstraction-revealed symmetry and prior plan actions Historical precedent for symmetry-guided action choice; SymBuild separates directionality identification from terminal-correspondence transport 19 Assembly component geometry Kinematic/spatial constraint inference Component feature symmetries reduce CSP combinatorics Establishes symmetry as an assembly reasoning resource; our transported object is directed sequential process structure SymPlan (36) Path-planning MDP / value-iteration operator Equivariant differentiable planning Relevant state/action transformations are symmetries of the planning problem Structures planner computation through equivariance; SymBuild centers information flow among process evidence, terminal correspondence, and realized-state refinement Invariant Transform Experience Replay (18) Feasible RL trajectory space Symmetry-based experience augmentation Selected transforms preserve feasible trajectories Reuses symmetry-transformed trajectories for experience augmentation; SymBuild transports directed process structure and re-certifies actions in the realized state Equivariant Motion Manifold Primitives (15) Task / solution-trajectory manifold Equivariant trajectory-family learning Task symmetry supports equivariant motion families Reuses symmetry across solution trajectories; we couple terminal correspondence to independent process evidence and statewise residuals Partial equivariant RL (4) State–action regions of an MDP Selective invariant/standard Bellman backups Symmetry validity varies across local process regions Adapts where process-level equivariance is used; our terminal correspondence stays fixed while decision relevance changes Cross-space symmetry composition (9) Configuration and task spaces Lift/descend/compose symmetries for jointly equivariant policies Multiple policy symmetries are composed across spaces Transports/composes symmetries themselves; we transport asymmetric process knowledge through a terminal correspondence Canonical diffusion (38) Orbit representatives of invariant targets Canonical representative selection Invariant target distribution is modeled on a canonical slice Resolves representation ambiguity; our statewise residual addresses transition-dependent decision relevance SymBuild Terminal outcomes plus external process evidence Transport–refine–certify for verified construction Directionality is independently grounded; exact terminal correspondence supplies transport Episode-fixed transported structure is refined by realized state under an unchanged verifier The broader landscape reinforces the same separation. MDP-homomorphic and equivariant-Q methods encode joint state–action symmetry directly in value/policy structure (31; 32); soft and approximate equivariance relax hard constraints under approximate or misspecified symmetries (6; 33); PARD and symmetry-aware GFlowNets use symmetry to structure graph generation or correct symmetry-induced transition bias (37; 13). These precedents show that symmetry can constrain, quotient, augment, or guide a process. The contribution here is the identification–transport–refinement decomposition when exact symmetry is available at terminal outcomes: symmetry carries directed information supplied elsewhere, while realized state updates its current decision relevance. Appendix C Additional Protocol Details Data provenance and domain construction. CAD objects come from the public official ASAP training partition (28); each CAD confirmation panel excludes assembly identities and mesh-content duplicates used by earlier CAD experiments before deterministic selection. Mini-Programs are generated deterministically from the data seed and latent program specification, and every confirmation split excludes identities used in earlier experiments. Exact-fill packing instances are generated deterministically from the data seed, split, weights, true groups, group-to-bin transport, and visible anchor assignments; confirmation IDs are pairwise disjoint and disjoint from development. The domain-specific generation and ranking rules are stated below. ASAP CAD correspondence and geometric-control protocol. The external CAD correspondence controls use the official ASAP assembly assets (28). Assemblies are parsed either from archive meshes already expressed in assembled terminal coordinates or from public repository examples with explicit final poses. The correspondence-development audit admits finite, parseable assemblies with 5–50 parts; a fixed cryptographic hash of the assembly identifier, reduced modulo five, gives the deterministic 80/20 train/validation partition, and augmented views inherit the assembly split. Terminal equivalence is the four-yaw tabletop C4C_4 orbit of the completed assembly. Each correspondence view independently permutes part identities, applies a random sensor-frame SE(3)SE(3) transform and a C4C_4 lift, and renders directional partial observations with 55%55\% retained visibility, at most 256 points per part, and Gaussian coordinate noise 0.0020.002. The gravity/tool directions expressed in the sensor frame provide the process coordinate system used by the geometric controls. PCS-ICP denotes the process-frame ICP/Hungarian transport control. It first expresses the partial observations in the tool/gravity frame, evaluates all four C4C_4 representatives, and runs up to four translation-only ICP refinements per representative with an 80%80\% trimmed correspondence set. Trimmed one-way partial-to-full Chamfer costs define the observed-part/canonical-part matrix; Hungarian assignment gives a one-to-one part correspondence, and the representative with minimum mean assignment cost is selected. The recovered lift and assignment transport the canonical CAD geometry into the target frame. At each remaining-part state and each fixed world direction, the transported process score recomputes a directional blocker relation on the predicted parts; a candidate receives its number of predicted blockers plus a −0.01-0.01 normalized outward-projection tie-break, with lower scores queried first. The outsideness control adapts ASAP’s outside-in part heuristic to the shared part–direction action set: using the same process-aligned centers, candidate actions, and verifier, it scores part i along removal direction d by −ci⊤d-c_i d. Both controls share verifier acceptance and action semantics, isolating the proposal information. Main 72-object CAD call-anytime confirmation. The primary CAD-call panel uses 72 previously unused assemblies from the official ASAP training partition; the official test partition is not used. Eligible assemblies have 5–20 parseable parts and a recoverable assembled configuration. After excluding every assembly identity and mesh-content duplicate used by earlier CAD experiments, candidates are placed in a fixed cryptographic order determined only by assembly identity. The first 24 assemblies with 5–8 parts and the first 48 with 9–20 parts form the frozen panel; all 72 load successfully with zero prohibited content overlap. Sensor seeds 701 and 709 are each combined deterministically with assembly identity to define one base partial view, and every assembly–sensor pair is evaluated under all four forced C4C_4 lifts, yielding 72×2×4=57672× 2× 4=576 conditions per method. Each view permutes part identities, applies a random sensor-frame rigid transform, retains 55%55\% of the directional partial observation with at most 256 points per part, and adds Gaussian coordinate noise with standard deviation .002.002. Process-axis noise applies one shared calibration rotation to the tool and gravity vectors, with a uniformly oriented axis and angle drawn from a zero-mean Gaussian with 5∘5 standard deviation. Actions pair each remaining part with one of the six fixed world directions ±x,±y,±z\± x,± y,± z\. For a remaining-part set R, PCS-ICP supplies the recovered lift and observed-to-canonical assignment used to transport the reference CAD. For action (i,d)(i,d), the transported process score is the number of predicted directional blockers plus the outward tie term sT(i,d,R)=#blockers(i,d,R)−0.01⟨ci,d⟩maxrangej∈R⟨cj,d⟩,10−8,s_T(i,d;R)=\#blockers(i,d;R)-0.01\, c_i,d \range_j∈ R c_j,d ,10^-8\, where cic_i is the transported predicted center. The local geometric score is sX(i,d,R)=−⟨xi,d⟩s_X(i,d;R)=- x_i,d , where xix_i is obtained by centering the observed part centers, expressing them in the tool/gravity-aligned process frame, and dividing all centers by their maximum radius (with a 10−810^-8 floor). With m=6|R|m=6|R| current actions, each score channel is independently converted to average weak ranks and normalized to [0,1][0,1] by subtracting one and dividing by max(m−1,1) (m-1,1). Dynamic rank meet orders actions by min(rT,rX)+10−6rT+10−9rX (r_T,r_X)+10^-6r_T+10^-9r_X, followed by part and direction index as the final deterministic tie-break. After an accepted removal, the blocker relation and both current ranks are recomputed on the reduced action set; a rejected proposal leaves the state unchanged and is skipped for that state. The initial-static control computes the complete fused ordering once at the initial state and subsequently only removes actions whose parts have already been accepted. All methods share the dense 96-step MuJoCo straight-sweep verifier, escape factor 1.5, penetration tolerance max(10−5D,10−7) (10^-5D,10^-7) for assembly diameter D, and a maximum of ⌈4N⌉ 4N verifier calls. The predeclared call budgets are 1.25,1.5,2,2.5,3,4N\1.25,1.5,2,2.5,3,4\N; AUC is their unweighted success mean, capped normalized call cost is calls/4N/4N on success and 1 on failure, and 10,000 assembly-paired bootstrap draws (seed 22101) average sensor/lift conditions within assembly before resampling. Independent 120-object CAD mechanism panels. The two 120-object rows in Table 3(b) use exactly the same 120 ASAP training assemblies and never use the official ASAP test partition. After excluding 335 assemblies that duplicate earlier CAD evaluation content, eligible assemblies are placed in a deterministic pseudorandom order based only on assembly identity and fixed before evaluation; the panel takes 24 assemblies with 5–8 parts, 48 with 9–20 parts, and 48 with 21–50 parts. The algorithm panel uses sensor seed 389 and all four C4C_4 lifts. It evaluates the physics oracle, PCS-ICP transport, outsideness, and rank meet under the dense 96-step verifier, escape factor 1.5, and shared cap max(3N,24) (3N,24), with 10,000 paired bootstrap draws (seed 14801). The timing panel reuses the identical 120 assemblies, changes the sensor seed to 503, and compares dynamic and initial-static rank meet under the same four lifts, verifier fidelity, escape factor, and cap; its paired bootstrap uses 10,000 draws (seed 14901). Thus the two rows change the intended algorithm/timing factor while holding the external CAD panel fixed. IKEA observed-process topology protocol. We use IKEA Video Manuals from 20, which provides aligned furniture parts, manuals, and real assembly-video substeps. A frozen terminal audit yields 36 terminal families and 97 parsed real videos; all 267 referenced OBJ parts pass the identity and geometry-integrity checks used to establish the terminal representation. Terminal correspondence is computed from the exact color-preserving part orbits of an unordered manual-incidence graph, with witness automorphisms supplying the transport maps used for candidate closure. For each OBJ part, vertices are centered and divided by their RMS radius. The part color is determined by the normalized covariance eigenvalues in descending order, radial-distance quantiles at 0,.1,.25,.5,.75,.9,10,.1,.25,.5,.75,.9,1, and the base-two logarithm of the vertex count; eigenvalues and radial quantiles are rounded to three decimals and the vertex-count statistic to two decimals. A zero-radius mesh receives a separate degenerate color indexed by its vertex count. Thus the color depends on rigid-invariant shape statistics rather than pose or process order. Each official manual connection is represented by one connection node joined to two same-colored side nodes that connect to the corresponding endpoint parts. Manual step order, page/step IDs, video order, correct actions, and downstream labels do not enter this terminal graph. Thirteen families contain nontrivial terminal orbits. The video parser admits only nonempty video IDs and substeps with a recorded start time and at least one nonempty part component. Within each video, usable events are sorted by start time and source order; consecutive exact repeats with the same action identity and start time are deduplicated. A predictive state is formed between successive usable events only when their start times differ. These deterministic parsing rules produce the reported 97 videos and 1,022 evaluable states; the latter span 35 families because one audited terminal family contributes no evaluable transition. Each state is the realized video prefix up to the current substep. Its candidate set is the deduplicated union of all remaining observed actions and their terminal transport-map variants, and the positive label is the actually observed next action. Of the 1,022 states, 191 states in 13 families are terminal-ambiguous, meaning that the next action admits a nontrivial transported alternative. All compared predictors use the same candidates and five family-held-out out-of-fold folds, 4,096-dimensional feature hashing, logistic regression, seed 20260802, and C∈.03,.1,.3,1,3C∈\.03,.1,.3,1,3\ selected by the frozen validation macro-MRR rule. The raw-prefix model uses candidate and realized-prefix features; the static-orbit model adds fixed terminal-orbit statistics; the section model refines orbit members using their first-seen prefix cohort and age; and the shuffled-section control preserves terminal structure and prefix length while permuting the realized process-member assignment. Top-1 and MRR rank the true next action among candidates, with 10,000-draw paired family bootstrap intervals. IKEA quantities in Table 3. “Process structure” is the section-minus-raw-prefix Top-1 gap on all 191 terminal-ambiguous states: +14.14+14.14 points [6.41,20.18][6.41,20.18]. “Action core” is the predeclared theory-core subset (125 states, 9 families) in which terminal ambiguity is present, the realized prefix breaks every relevant transport, and the terminal has 8–20 part nodes; the corresponding section-minus-raw/static-orbit Top-1 gap is +19.20+19.20 points [10.26,25.56][10.26,25.56]. “Hybrid” applies the section model on terminal-ambiguous states and the raw-prefix model elsewhere, giving +2.64+2.64 Top-1 points [.77,4.38][.77,4.38] over raw on the full panel. The shuffled-section control gives a further process-membership check by preserving terminal structure while changing which symmetry-related member the realized prefix selects. Mini-Program generation, relation checkpoint, and confirmation population. Each proposed Mini-Program first draws its leaf count uniformly from the requested inclusive range. A full binary expression tree is then generated recursively: for a subtree with more than one leaf, the number of leaves assigned to the left child is uniform on 1,…,L−1\1,…,L-1\, and every internal operator is independently uniform over XOR, AND, and OR. The two symmetric modules receive independent length-L literal vectors of fair binary bits, and the final root operator is uniform over XOR and equality. If the two module outputs are v0,v1∈0,1v_0,v_1∈\0,1\ and the root choice is indexed by h∈0,1h∈\0,1\, the six-way generation label is 3h+v0+v13h+v_0+v_1. For a requested population of C objects, the generator fills these six labels as evenly as possible; C=120C=120 therefore gives exactly 20 programs per label. Duplicate canonical programs are rejected with the two module orders identified under terminal swap, candidates whose label quota is already full are rejected, and the accepted population is finally shuffled. Balance is imposed only on the six output labels, not on leaf counts, tree shapes, internal operators, or literal patterns. The frozen checkpoint uses 240 training and 60 validation objects with 3–5 leaves, disjoint from all development and confirmation objects. A width-64, three-message-layer relation network is trained for 800 AdamW steps (seed 211, learning rate .002.002, batch size 32, gradient clipping at 1); validation macro-F1 measured every 100 steps selects the checkpoint and validation alone selects relation thresholds. Formal anytime confirmation uses the selected frozen checkpoint without retraining, threshold fitting, or adaptation. Two frozen data seeds, 20270117 and 20270129, each generate 120 ID programs with 3–5 leaves using generator seed data seed+901data seed+901 and 120 length-OOD programs with 6–8 leaves using generator seed data seed+902data seed+902, giving 240 objects per split. Object identities are pairwise disjoint across seed/split cells and exclude every object identity used by earlier Mini-Program experiments. The deployment interface retains each direct reference independently with probability .35.35 in one randomly chosen anchor module and exposes no demonstration order or hidden reference. Mini-Program proposal ranks and dynamic/static intervention. Each program presents the same unordered event pool to every method. Let O be the sparse observed-reference graph and S the binary terminal module-swap orbit. The observed expert forms the aligned transported graph by adjoining the orbit-propagated references (SOS)>0(SOS)>0, adding the fixed non-root-to-root constraints, removing self edges, and taking transitive closure. The frozen relation model supplies precedence probabilities; the process graph takes their elementwise maximum with this aligned observed graph so directly observed precedence is preserved. At state s with remaining-event mask RsR_s, either graph G assigns candidate a the residual predecessor mass zG(s,a)=∑u∈RsGua,z_G(s,a)= _u∈ R_sG_ua, with lower values preferred; equal scores receive their average weak rank. The process rank rTr_T uses the learned-plus-observed graph and the state rank rXr_X uses the transported observed graph. Dynamic meet recomputes both residual ranks after every accepted event. Initial-static computes the same two ranks once at the empty prefix and subsequently only removes completed events. The exact proposal key for meet is (min(rT,rX),rT+rX,max(rT,rX),π(a))( (r_T,r_X),\,r_T+r_X,\, (r_T,r_X),\,π(a)), where π is a deterministic seed-fixed presentation order; the learned-only and observed-only controls use rTr_T and rXr_X respectively. The permuted control applies a fixed random permutation to all non-root event identities, conjugates only the learned precedence matrix, and keeps the observed expert aligned. Every proposal is accepted only when the private true-precedence graph contains no unfinished predecessor of that event. Mini-Program anytime metric. The maximum budget is 2n2n verifier calls for n events, with the predeclared grid q/n∈1,1.25,1.5,1.75,2q/n∈\1,1.25,1.5,1.75,2\. Program anytime AUC is the unweighted mean of the five completion indicators on this grid, i.e., the Section 3.3 AUC definition with uniform mass on these five budgets. Capped normalized cost is q/(2n)q/(2n) for a successful plan and 11 for an unsuccessful plan; the Program “Cost ↓ ” entries in Table 1 report the paired static-minus-dynamic reduction in this quantity. Intervals use 10,000 object-paired bootstrap draws (seed 24101), and the two frozen data seeds are also reported separately as replication groups. Exact-fill packing generation, ranks, and size shift. A packing action is a=(i,b)a=(i,b), assigning an unassigned item i of weight wiw_i to bin b of capacity 100. Each latent group contains exactly three items whose weights sum to 100: an integer “large” weight is sampled uniformly from 45–60, a second weight from 12–30, and the third is the residual, required to lie in 12–40; sampling repeats until at least two triplet weights differ. Groups and items are permuted, and a random permutation maps latent groups to terminal bin identities. One visible anchor per group–the heaviest item, with item index breaking weight ties–is preassigned to its transported bin, so the anchors expose a complete group-to-bin bijection τ. Among the n remaining items, max(2,round(.25n)) (2,round(.25n)) reference-group labels are selected without replacement and cyclically shifted, implementing the requested 25% prior-corruption condition. With data seeds 20271103 and 20271117, each seed contributes 120 ID instances with 4–6 bins and 120 size-OOD instances with 7–8 bins. Hence ID instances contain 12–18 total items and 8–12 initially unassigned decision items, while size-OOD instances contain 21–24 total items and 14–16 initially unassigned decision items. At state s, the direct candidate set AsA_s contains exactly the capacity-feasible assignments. Let Lb(s)L_b(s) be current load, gig_i the observed reference-group label, and ci(s)=|b:(i,b)∈As|c_i(s)=|\b:(i,b)∈ A_s\|. For a lexicographic key κ, ordnormAs(κ)ordnorm_A_s(κ) sorts by (κ(a),a)(κ(a),a) and maps the zero-based position to [0,1][0,1] by division by max(|As|−1,1) (|A_s|-1,1). The transported and residual ranks are rT(i,b) r_T(i,b) =ordnormAs([b≠τ(gi)],gi,−wi,b), =ordnorm_A_s\! (1[b≠τ(g_i)],\,g_i,\,-w_i,\,b ), rX(i,b,s) r_X(i,b;s) =ordnormAs(ci(s),−wi, 100−Lb(s)−wi,b). =ordnorm_A_s\! (c_i(s),\,-w_i,\,100-L_b(s)-w_i,\,b ). Rank meet queries actions by (min(rT,rX),rT,rX,a)( (r_T,r_X),r_T,r_X,a). Dynamic recomputes rXr_X after every accepted assignment; initial-static retains the initial residual rank for surviving actions. The process-only control uses rTr_T. The best-fit control orders by (−wi,100−Lb(s)−wi,b)(-w_i,100-L_b(s)-w_i,b), and the permuted-transport control cyclically shifts the recovered group-to-bin transport. The exact extension verifier accepts an action only if memoized branch-and-bound proves that the remaining weights can exactly fill the residual bin capacities. The call grid is q/n∈1,1.125,1.25,1.5,2,3q/n∈\1,1.125,1.25,1.5,2,3\ for the n initially unassigned items. Packing AUC is the uniform mean over these six budgets; capped call cost is q/(3n)q/(3n) on success and 1 on failure, and Table 1 reports the paired static-minus-dynamic reduction. Public-product and multi-heuristic aggregation. The parameter-free SayCanPay-style replication uses the same domain-specific rTr_T and rXr_X as the matched primary comparison and sorts ascending by Kprod(a)=((1+rT(a))(1+rX(a)),rT(a)+rX(a),max(rT(a),rX(a)),tie(a)).K_ prod(a)= ((1+r_T(a))(1+r_X(a)),\;r_T(a)+r_X(a),\; (r_T(a),r_X(a)),\;tie(a) ). For CAD, the two raw channel scores are independently converted to SciPy average ranks shifted to start at zero; the first three key fields are encoded as product +10−9+10^-9 sum +10−12+10^-12 max, followed by the stable action order. Mini-Programs use the average weak ranks and deterministic presentation order defined above. Packing uses the normalized ranks above and the action tuple as its final tie-break. The “MH” columns in Table 14 use the separate parameter-free same-interface key KMH(a)=(min(2rT(a),2rX(a)+1),rT(a)+rX(a),max(rT(a),rX(a)),tie(a)),K_ MH(a)= ( (2r_T(a),2r_X(a)+1),\;r_T(a)+r_X(a),\; (r_T(a),r_X(a)),\;tie(a) ), which is distinct from the validation-selected multi-queue 1:21:2 experiment in Appendix C.2. For CAD, the static versions cache the initial combined score; for Mini-Programs they cache both initial weak ranks and then remove completed events; for packing they retain the initial residual rank on surviving actions while restricting the fixed transported prior to the current candidate set. Dynamic versions recompute the state-dependent rank after each accepted transition. Candidate sets, verifier semantics, object splits, and interruption budgets are unchanged from the primary rank-meet panels. Verifier definitions. CAD proposals are checked by a dense 96-step MuJoCo (30) convex-hull straight sweep; its Boolean decision and actual integration-step count produce the call and work panels. The Mini-Program precedence verifier accepts event a exactly when no unfinished event is a true predecessor of a; because the hidden relation is a DAG, every acceptance preserves a complete topological suffix. The packing verifier first rejects capacity overflow and sum mismatch, then runs memoized exact branch-and-bound over remaining weights and sorted residual capacities until it proves or refutes a complete exact-fill suffix; there is no node cap or unknown return. Visible anchor assignments recover the observable group-to-bin correspondence before ranking, while corrupted non-anchor reference labels affect only the transported prior. CAD learned-residual protocol. The aligned–permuted CAD intervention uses mutually disjoint 60/20/60 train/validation/evaluation assemblies from the same public ASAP partition. The dense 96-step verifier labels a frozen probe union at states visited by analytic rank meet, yielding 5,022 training and 1,460 validation state–action rows. A two-hidden-layer width-64 MLP is trained for 2,000 AdamW steps with seed 307; validation BCE selects the step-100 checkpoint. The 60 evaluation objects use four lifts, sensor seed 613, no adaptation, and the common max(3N,24) (3N,24) call cap. Fresh CAD work-anytime protocol and budget distributions. The CAD work panel uses 48 fresh ASAP training assemblies after recursively excluding 769 previously used assembly identities or content duplicates. Eligible objects are placed in a deterministic pseudorandom order based only on assembly identity and fixed before evaluation; the panel takes 16 assemblies with 5–8 parts and 32 with 9–20 parts. Each assembly is evaluated with sensor seeds 811 and 823, all four C4C_4 lifts, five-degree process-axis noise, escape factor 1.5, and the dense 96-step verifier, under a common maximum of 4N4N proposal calls. Let W be the total number of MuJoCo integration steps actually consumed by proposal verification, with each call contributing 1–96 steps and final replay excluded. At work-budget factor b, a valid completion succeeds iff W≤⌈96bN⌉W≤ 96bN . The work-budget grid is B=(1.25,1.5,2,2.5,3,4)B=(1.25,1.5,2,2.5,3,4) full-sweep equivalents per part. The three predeclared interruption distributions on this ordered grid are μtight _ tight =(.35,.25,.18,.12,.07,.03), =(.35,.25,.18,.12,.07,.03), μuniform _ uniform =(1/6,1/6,1/6,1/6,1/6,1/6), =(1/6,1/6,1/6,1/6,1/6,1/6), μrelaxed _ relaxed =(.03,.07,.12,.18,.25,.35). =(.03,.07,.12,.18,.25,.35). For each distribution, AUCμ=∑b∈Bμ(b)F(b)AUC_μ= _b∈ Bμ(b)F(b). The unqualified CAD work AUC reported in Table 1 is the uniform case; tight and relaxed distributions are robustness checks. Capped normalized work cost is W/(96⋅4N)W/(96· 4N) for a valid completion and 1 otherwise, and the CAD-work “Cost ↓ ” entry is the paired static-minus-dynamic reduction. Intervals use 10,000 assembly-paired bootstrap draws (seed 23101), averaging the two sensor seeds and four lifts within assembly before resampling. Matched controls, freshness, and statistics. Every dynamic–static pair shares the instance, information channels, deterministic tie-break, proposal cap, and verifier. CAD single-channel controls are PCS-ICP and outsideness; the latter follows the outside-in geometric part-selection principle used in ASAP (28); Mini-Program controls are transported learned and observed relations; packing controls are the process prior and best fit. Program and packing formal panels use permuted transport; the object-disjoint CAD learned-residual panel supplies the CAD alignment intervention. Confirmation objects are disjoint from development and one another by deterministic identity. The six primary CAD/Program/Packing panels use 10,00010,000 paired bootstrap draws at the object or assembly level; GRN and CDGS-style intervals use 20,00020,000 scene-clustered resamples after averaging search seeds within scene–target. Oracle and qualification runs are used only as diagnostics. Budgets and integrity. CAD call budgets are 1.25,1.5,2,2.5,3,4\1.25,1.5,2,2.5,3,4\ calls per part; the CAD work budgets and their probability masses are defined above. Program budgets are 1,1.25,1.5,1.75,2\1,1.25,1.5,1.75,2\ calls per event, and packing budgets are 1,1.125,1.25,1.5,2,3\1,1.125,1.25,1.5,2,3\ calls per initially unassigned item. Every formal evaluation checks the expected row count, unique identities, zero prohibited cross-split overlap, deterministic replay where applicable, consistent verifier traces, and valid resource sums before aggregation. External implementation provenance. The CAD source assemblies and outside-in geometric control are grounded in ASAP (28); CAD verification uses MuJoCo (30). The public-product replication adapts the multiplicative scoring rule of SayCanPay (11) using the exact joint key defined above. GRN experiments use the authors’ official checkpoint and Panda OOD data (2), and the wider population-guided stress test follows the CDGS-style search template of 23. Table 16 summarizes the official implementations, licenses, native interfaces, and their role in these comparisons; the repository revision for each implementation was fixed for the audit. Software and hardware. The reported experiments were executed on 64-bit Windows 11 with an AMD Ryzen 7 8845H CPU (8 cores, 16 threads), 31.3 GiB RAM, and an NVIDIA GeForce RTX 4060 Laptop GPU with 8,188 MiB memory (driver 572.83). The software environment uses Python 3.8.20, PyTorch 2.4.1 with CUDA 12.4, NumPy 1.24.4, SciPy 1.10.1, and Matplotlib 3.7.2. Learned relation and GRN scoring use CUDA; discrete verifiers, bootstrap/permutation statistics, and table aggregation run on CPU. Aggregation-operator ablation. On the disjoint 120-object CAD development panel used to select the fixed aggregator, rank meet achieves 86.25%86.25\% success and 1.5471.547 normalized calls per part, compared with 80.00%/1.75280.00\%/1.752 for Borda and 79.17%/1.82079.17\%/1.820 for max-rank join. This intervention holds both ranks and the verifier fixed, supporting the union-coverage mechanism used in Lemma 1. Mini-Program correspondence-noise construction. The graded mismatch diagnostic reuses exactly the frozen 480-object confirmation population, checkpoint, observed expert, verifier, and budget grid. For requested fraction f∈0,.1,.25,.5,1f∈\0,.1,.25,.5,1\, let M be the non-root event identities and set m=round(f|M|)m=round(f|M|), except that a singleton move is promoted to two identities when possible. A seed-fixed sample of m identities is cyclically shifted by one position and used to conjugate the learned precedence matrix; all unsampled identities remain fixed. Thus the realized moved fraction is measured directly as m/|M|m/|M|. The perturbation seed is a deterministic function of data seed, split, object index, and noise level, so every level is reproducible without consulting outcomes. Clean-minus-noisy AUC intervals use 10,000 paired bootstrap draws on the shared objects. Table 6: Graded correspondence-noise sensitivity on the frozen 480-program panel. Noise deterministically moves the indicated fraction of non-root action identities before conjugating the learned precedence channel; the observed channel, verifier, budget grid, and checkpoint remain fixed. Cells are anytime AUC / final success in percent. Requested noise 0%0\% 10%10\% 25%25\% 50%50\% 100%100\% ID 85.08/100.00 75.33/98.33 63.42/95.42 40.58/77.92 13.58/36.67 Length OOD 64.67/97.50 49.58/88.75 32.50/70.00 11.42/31.67 0.75/2.92 The realized moved fractions are 12.78/10.21%12.78/10.21\% for requested 10%10\% on ID/OOD because small graphs require at least a two-node swap; all other realized means are within .74.74 points of the requested rate. Clean-minus-noisy AUC intervals are already positive at the first nonzero level: 9.759.75 [7.75,11.83][7.75,11.83] points on ID and 15.0815.08 [12.67,17.50][12.67,17.50] on length OOD. The monotone frontier makes correspondence quality an observable deployment variable rather than a binary aligned/permuted intervention. Figure 5: Correspondence quality produces a graded performance frontier. Exact values appear in Table 6; both anytime AUC and final verified success decrease as transported action identities are increasingly mismatched. C.1 Planning-interface controls Common interface. These controls use the same 480 Mini-Programs as the primary program panel: two independently generated data groups, each containing 120 ID objects with 3–5 leaves and 120 length-OOD objects with 6–8 leaves. Every method receives the same observable graph, frozen learned relation model, candidate actions, deterministic tie protocol, and Boolean precedence verifier. The hidden precedence graph is unavailable to proposal mechanisms and is used only by the shared verifier and post-evaluation validity checks. The cap is 2n2n verifier calls for a program with n events, and anytime AUC averages success at q/n∈1,1.25,1.5,1.75,2q/n∈\1,1.25,1.5,1.75,2\. All paired intervals use 10,000 object-level bootstrap draws. Each adaptation instantiates the corresponding search principle on this common candidate/verifier interface, enabling matched comparison while preserving the candidate and verification semantics used throughout the paper. Canonical-slice control. The exact terminal correspondence defines an identity/swap orbit. For each program, the observable serialization of both representatives is formed from event types, numeric attributes, and observed edges; the lexicographically smaller representative is selected. Its canonical event positions provide the final tie resolution for the initial-static rank-meet order. Thus the control uses exact terminal canonicalization to reduce orbit-label ambiguity while retaining the same observable information and verifier access as the primary method. State evidence is held at its initial value, isolating representative selection from post-transition refresh. POCL-style agenda. An observable acyclic partial order is constructed from transported observed edges together with learned edges of probability at least .5.5, dropping an added edge whenever it would create a cycle. At each state, candidates with fewer remaining predecessors are served first; state and process ranks, followed by the shared deterministic tie order, resolve agenda ties. The refreshed variant recomputes those ranks after every accepted event. Its fixed counterpart updates which predecessors remain but retains the initial state/process ranks for all tie resolution. Multi-heuristic scheduling. The transported process rank and state residual rank define two independent queues. A deterministic 1:21:2 process:state service ratio, selected on a disjoint development panel and frozen before confirmation, interleaves previously un-emitted candidates from the two queues. The refreshed variant rebuilds both queues after every accepted event; the fixed counterpart filters the two initial queues as events are completed without recomputing their ranks. Weighted-A*/LazySP-style replanning. A verifier-free bounded best-first search proposes short candidate paths from the observable projected partial order and state rank. Search width is 32, horizon is 4, and the weighted best-first coefficient is 2. Only the first action of the selected path is sent to the shared verifier. Rejection invalidates that state–action choice and triggers replanning; acceptance advances the state and restarts planning from the successor. Table 7: Full interface-matched search and canonicalization comparison. AUC and final success are percentages. Selector time is mean Python proposal time per object; expansions count verifier-free prefix expansions. All rows share the same 480-object population, observable information, verifier, and 2n2n cap. ID Length OOD Method AUC Success@2n2n Selector ms Expansions AUC Success@2n2n Selector ms Expansions SymBuild 85.08 100.00 5.5 0 64.67 97.50 11.2 0 Initial-static meet 72.67 100.00 2.0 0 42.92 83.33 3.2 0 POCL-style + refresh 85.67 100.00 9.1 0 66.33 97.92 21.8 0 Canonical slice + static meet 72.50 100.00 8.3 0 42.92 82.92 19.9 0 MHA-style 1:21:2 + refresh 85.92 100.00 8.4 0 65.83 97.08 20.0 0 Weighted-A*/LazySP-style 80.67 99.58 94.2 391.7 60.42 96.25 193.3 609.6 Across the common verifier interface, SymBuild gains 12.5812.58/21.7521.75 AUC points on ID/length OOD over canonical-slice static planning and 4.424.42/4.254.25 points over weighted-A*/LazySP-style replanning; paired intervals are reported in Table 8. On length OOD, refreshed POCL- and MHA-style rows show the same statewise signal in partial-order and multi-queue scheduler organizations. Table 8: Paired contrasts for search-family and canonicalization controls. Gaps are AUC percentage points with object-paired 95% bootstrap intervals. For the POCL- and MHA-style rows, the comparison holds the scheduler family fixed and changes only statewise rank refresh. Comparison ID Δ [95%CI][95\%\,CI] Length-OOD Δ [95%CI][95\%\,CI] SymBuild −- canonical-slice static +12.58+12.58 [9.67,15.50][9.67,15.50] +21.75+21.75 [18.42,25.17][18.42,25.17] SymBuild −- weighted-A*/LazySP-style +4.42+4.42 [2.58,6.33][2.58,6.33] +4.25+4.25 [1.83,6.67][1.83,6.67] POCL-style refresh −- fixed +1.08+1.08 [.25,2.00][.25,2.00] +3.83+3.83 [2.58,5.08][2.58,5.08] MHA-style refresh −- fixed −.42-.42 [−1.75,.83][-1.75,.83] +5.00+5.00 [2.83,7.17][2.83,7.17] On length OOD, refresh also reduces capped normalized verifier cost by 2.352.35 points for the POCL-style scheduler (95%95\% CI [1.81,2.91][1.81,2.91]) and 3.263.26 for the MHA-style scheduler ([2.11,4.42][2.11,4.42]). Every reported completion was replayed against the exact precedence relation after evaluation and was valid, and every method respected the common verifier-call cap. C.2 Validation-tuned and approximate-verifier protocols Scheduler development and confirmation. The confirmation population exactly regenerates the 480 frozen Mini-Program objects used in the primary panel. A separate 160-object development set uses seeds 20270203 and 20270211, with 40 objects per seed and split, and has zero identity overlap with confirmation or any object set used in earlier experiments. Development evaluates deterministic process:state queue ratios 1:1,1:2,2:1,1:3,3:11:1,1:2,2:1,1:3,3:1 and weighted rank sums wrT+(1−w)rXwr_T+(1-w)r_X for w∈.25,.50,.75w∈\.25,.50,.75\. For a queue ratio p:sp:s, each service cycle emits the next p previously un-emitted candidates from the process-rank queue and then the next s previously un-emitted candidates from the state-rank queue; a duplicate encountered in the second queue is skipped without consuming a service slot, and cycling continues until all candidates are emitted. Under refresh, both ranks and both queues are rebuilt after every accepted event; within an unchanged state, rejected candidates are simply passed over in the already constructed order. The static version keeps the initial process/state ranks and only filters completed events. Mean ID/OOD development AUC selects one configuration per family; the selected 1:21:2 queue and w=.25w=.25 sum are then frozen. Confirmation shares checkpoint, candidates, verifier, budgets, tie-breaking, and object-level bootstrap protocol across dynamic and initial-static versions. Approximate-verifier pressure test. On the same 480 confirmation objects, a deterministic cryptographic mapping of object identity, completed set, candidate action, regime, and noise seed defines a fixed verifier realization, so the two compared planners receive the same noisy answer whenever they query the same state–action pair. False-positive runs leave truly valid actions unchanged and accept an otherwise invalid action with probability p; such an acceptance marks the event completed in the planner’s reported state and planning continues from that state, while a hidden validity flag records that the true execution prefix is no longer valid. False-negative runs leave invalid actions rejected and reject a truly valid action with probability p; the state then remains unchanged and evaluation advances to the next candidate. Probabilities are .01,.025,.05,.10\.01,.025,.05,.10\ with 20 noise seeds per object and condition. Noise seeds are averaged within object before 10,000-draw object-paired bootstrap. Reported completion requires all events to be marked completed, whereas true success additionally requires every accepted action to have been valid under the hidden exact verifier; false certification is reported completion without true success. GRN score integration. On target-removal episodes derived from the official Panda OOD10/15/20 scenes of 2, the frozen GRN supplies per-object/per-grasp feasibility scores and the target-conditioned process channel supplies obstruction-depth utility on the same remaining scene. GRN static caches the initial GRN scores, whereas GRN refresh reruns the frozen model after every accepted removal. Combined static caches the initial z-normalized GRN and process utilities and sums them with equal weight; SymBuild recomputes both utilities after each accepted removal. Candidate identities, accept/reject state transition, tie-breaking, annotation verifier, and the common 5N5N query cap are shared across these rows; the exact target construction, verifier predicate, score normalization, search seeds, and population-search adaptation are specified in Appendix C. GRN wall-clock protocol. The official GRN checkpoint and Panda OOD10/15/20 data (2) are timed on 10 scenes, using all available solvable episodes up to 100 (61/82/100), three repeats, one warm-up, CUDA synchronization, and rotated method order. Compute time excludes external verifier latency. For static compute tst_s, refresh compute trt_r, and mean query counts qs>qrq_s>q_r, break-even latency is λ⋆=(tr−ts)/(qs−qr)λ =(t_r-t_s)/(q_s-q_r). Projected end-to-end time adds λqλ q at each latency. Table 9: Complete validation-tuned scheduler confirmation. AUC, final success, and capped normalized cost are percentages except cost. Weighted sum and multi-queue parameters are selected on 160 disjoint development objects, then locked for this 480-object confirmation. Best and second-best dynamic/static values in each row are bold and underlined; ties are unmarked. Split Aggregator Refresh AUC Static AUC Δ [95%CI][95\%\,CI] Final: refresh/static Cost: refresh/static Selection ID Rank meet 85.08 72.67 +12.42+12.42 [9.50,15.33][9.50,15.33] 100/100 .5719/.6456 primary ID Multi-queue 1:21:2 85.92 86.33 −.42-.42 [−1.75,.83][-1.75,.83] 100/100 .5652/.5641 dev-tuned ID Weighted sum w=.25w=.25 85.17 71.67 +13.50+13.50 [10.50,16.42][10.50,16.42] 100/100 .5718/.6530 dev-tuned Length OOD Rank meet 64.67 42.92 +21.75+21.75 [18.33,24.92][18.33,24.92] 97.50/83.33 .6706/.8057 primary Length OOD Multi-queue 1:21:2 65.83 60.83 +5.00+5.00 [2.83,7.17][2.83,7.17] 97.08/95.00 .6635/.6962 dev-tuned Length OOD Weighted sum w=.25w=.25 65.83 41.67 +24.17+24.17 [20.75,27.50][20.75,27.50] 97.50/82.08 .6636/.8132 dev-tuned Table 10: Complete stochastic-verifier pressure test. True AUC is evaluated against the hidden exact verifier. FC is false-certified completion. Values are percentages; every row averages 20 deterministically fixed noise seeds within each of 240 objects before paired bootstrap. Regime Split p Refresh true AUC Static true AUC Δ [95%CI][95\%\,CI] FC: refresh/static Seeds/objects FP ID .01 83.43 69.81 +13.62+13.62 [10.53,16.76][10.53,16.76] 2.60/5.38 20/240 FP ID .025 81.36 66.08 +15.28+15.28 [11.92,18.73][11.92,18.73] 5.79/12.21 20/240 FP ID .05 78.27 60.80 +17.47+17.47 [13.70,21.25][13.70,21.25] 10.67/21.56 20/240 FP ID .10 72.72 52.09 +20.62+20.62 [16.03,25.20][16.03,25.20] 18.92/36.46 20/240 FP Length OOD .01 60.75 38.81 +21.94+21.94 [18.68,25.26][18.68,25.26] 8.10/12.42 20/240 FP Length OOD .025 54.63 33.24 +21.39+21.39 [18.10,24.63][18.10,24.63] 19.88/28.17 20/240 FP Length OOD .05 46.78 26.30 +20.48+20.48 [17.18,23.76][17.18,23.76] 34.12/48.12 20/240 FP Length OOD .10 35.62 17.72 +17.90+17.90 [14.61,21.23][14.61,21.23] 52.85/70.48 20/240 FN ID .01 81.04 69.64 +11.40+11.40 [8.59,14.16][8.59,14.16] 0/0 20/240 FN ID .025 74.96 65.01 +9.95+9.95 [7.39,12.57][7.39,12.57] 0/0 20/240 FN ID .05 66.97 58.33 +8.65+8.65 [6.25,11.13][6.25,11.13] 0/0 20/240 FN ID .10 52.88 45.97 +6.92+6.92 [4.84,9.06][4.84,9.06] 0/0 20/240 FN Length OOD .01 61.28 40.74 +20.54+20.54 [17.35,23.73][17.35,23.73] 0/0 20/240 FN Length OOD .025 55.97 38.09 +17.88+17.88 [14.94,20.85][14.94,20.85] 0/0 20/240 FN Length OOD .05 48.57 33.56 +15.01+15.01 [12.36,17.75][12.36,17.75] 0/0 20/240 FN Length OOD .10 36.38 24.98 +11.40+11.40 [9.16,13.60][9.16,13.60] 0/0 20/240 Table 11: Full scorer-level wall-clock decomposition. Compute is milliseconds per episode; queries are capped verifier calls. Process refresh is inexpensive, while neural GRN recomputation dominates the learned and combined rows. Scale Scorer Compute static/refresh Queries static/refresh Saved queries Break-even (ms/query) OOD10 Process .834/1.133.834/1.133 8.05/6.878.05/6.87 1.18 .253 OOD10 GRN .466/83.693.466/83.693 5.13/4.825.13/4.82 .31 267.20 OOD10 Combined .354/40.062.354/40.062 2.95/2.072.95/2.07 .89 44.85 OOD15 Process 1.811/2.6061.811/2.606 15.44/12.0215.44/12.02 3.41 .233 OOD15 GRN 1.019/142.0731.019/142.073 7.87/6.547.87/6.54 1.33 106.11 OOD15 Combined .638/76.434.638/76.434 4.17/3.054.17/3.05 1.12 67.56 OOD20 Process 2.086/2.6922.086/2.692 15.72/11.2215.72/11.22 4.50 .135 OOD20 GRN 1.118/198.9091.118/198.909 9.26/7.619.26/7.61 1.65 119.87 OOD20 Combined .719/95.812.719/95.812 4.81/3.304.81/3.30 1.51 62.98 Table 12: Projected combined refresh-minus-static end-to-end latency. Entries are milliseconds per episode; negative values favor refresh. Projection adds the stated fixed latency to every measured verifier query. Scale 1 ms/query 10 ms/query 50 ms/query 100 ms/query 500 ms/query 1000 ms/query OOD10 +38.82+38.82 +30.86+30.86 −4.55-4.55 −48.82-48.82 −402.92-402.92 −845.54-845.54 OOD15 +74.67+74.67 +64.58+64.58 +19.70+19.70 −36.40-36.40 −485.18-485.18 −1046.15-1046.15 OOD20 +93.58+93.58 +79.99+79.99 +19.59+19.59 −55.91-55.91 −659.91-659.91 −1414.91-1414.91 The scheduler table shows that statewise updates can be consumed by distinct ranking and queueing mechanisms; the validation-selected weighted sum gives the strongest confirmation gap across the two program splits. The noise table shows that the empirical advantage persists under approximate verifier answers, and the timing tables identify the verifier-latency range in which query savings translate into lower wall-clock solution time. Appendix D Complete Baselines and Mechanism Tables Table 13: Formal evaluation coverage. Counts refer to fresh confirmation objects; every listed seed is evaluated on every applicable method and budget. n denotes parts, events, or remaining items. Domain Objects Seeds Shifts Budgets Verifier resource CAD calls 72 2 sensor ID, 5–20 parts 6: 1.25n6:\ 1.25n–4n4n proposal calls CAD work 48 2 sensor ID, 5–20 parts 6: 1.256:\ 1.25–44 dense sweep steps Mini-Program 480 2 data ID, length OOD 5: 1n5:\ 1n–2n2n precedence queries Exact-fill packing 480 2 data ID, size OOD 6: 1n6:\ 1n–3n3n exact-suffix queries Table 14 groups matched temporal comparisons, component ablations, correspondence mismatch, and alternative aggregators. Oracle qualification runs use hidden construction labels to verify carrier presence and are reported as diagnostics. CAD correspondence evidence comes from the object-disjoint aligned–permuted intervention in Table 3(b); the formal CAD matrix marks those cells with an em dash. Table 14: Complete comparison and ablation matrix. Cells are anytime AUC / final verified success in percent. Process/state are PCS-ICP/outsideness, learned/observed relation, and process-prior/best-fit in the three domains. Best and second-best distinct AUCs per row are bold and underlined; ties share formatting. Temporal comparison Component ablation Mismatch Alternative aggregators Panel N Meet dyn. Meet stat. Process State Permuted Product dyn. Product stat. MH dyn. MH stat. CAD calls 72 81.60/94.62 78.73/89.93 79.57/89.58 73.12/91.49 – 82.38/95.49 79.08/90.10 81.63/94.44 78.76/89.76 CAD work 48 98.70/98.70 91.93/91.93 87.50/87.50 85.63/85.94 – 98.18/98.18 93.75/93.75 98.70/98.70 91.93/91.93 Program ID 240 85.08/100 72.67/100 84.83/100 85.17/100 15.83/37.50 85.08/100 72.67/100 85.08/100 72.83/100 Program OOD 240 64.67/97.50 42.92/83.33 63.58/97.08 65.83/97.50 1.25/4.58 65.50/97.50 42.92/83.75 64.33/97.50 42.83/82.50 Packing ID 240 57.92/100 49.38/99.58 50.21/100 38.26/100 13.33/68.75 72.29/100 65.83/100 50.21/100 50.21/100 Packing OOD 240 34.03/95.00 25.35/83.33 28.40/93.75 21.67/99.17 3.68/20.42 47.92/98.33 41.39/99.17 28.40/93.75 28.40/93.75 Table 15: Full public-product scores underlying Table 1. Product uses the parameter-free multiplicative joint score adapted from SayCanPay (11); every refreshed/static pair shares the same product rule, ranks, verifier, and first action. Rank meet is the primary method pre-specified before formal evaluation. AUC is in percent; best and second-best values among the three shown methods are bold and underlined. Domain Resource/shift N Product + refresh Static product Primary meet Product refresh gap [95%CI][95\%\,CI] CAD calls 72 82.38 79.08 81.60 +3.30+3.30 [1.01,5.79][1.01,5.79] CAD measured work 48 98.18 93.75 98.70 +4.43+4.43 [.52,9.90][.52,9.90] Program ID 240 85.08 72.67 85.08 +12.42+12.42 [9.50,15.33][9.50,15.33] Program length OOD 240 65.50 42.92 64.67 +22.58+22.58 [19.33,25.83][19.33,25.83] Packing ID 240 72.29 65.83 57.92 +6.46+6.46 [4.51,8.47][4.51,8.47] Packing size OOD 240 47.92 41.39 34.03 +6.53+6.53 [4.58,8.47][4.58,8.47] Table 16: Recent public-code planning baselines and interface audit. Each referenced public implementation was fixed to one repository revision for the audit. Native models are compared at the task/interface level; numerical rows use only transformations that preserve our candidates, verifier, and budgets. Method Venue License Native interface Use in this paper SayCanPay (11) AAAI 2024 not stated in official repository language actions + Can/Pay scores product-rule dynamic/static numerical audit GRN (2) ICLR 2025 MIT 3D TAMP graph, grasp/IK/GO labels official OOD scorer and matched-search comparison PNO (21) ICLR 2025 C BY-NC-SA 4.0 continuous map/configuration value field task/interface contextual comparison Table 17: Logical search cost and completion for the GRN-native population-search stress test. GRN states count distinct scored remaining-component states per target; sequence evaluations count sampled action steps. Success is measured at the common 5N5N verifier budget. Split SymBuild GRN states CDGS–refresh GRN states CDGS–GRN states CDGS–refresh seq. evals CDGS–GRN seq. evals Success: SymBuild/CDGS–refresh OOD10 2.00 13.14 42.42 154.18 335.62 100/100 OOD15 2.64 23.64 71.26 214.67 624.21 100/100 OOD20 3.14 38.32 109.06 258.60 669.05 100/100 GRN target-removal interface and CDGS-style search. The official Panda OOD scenes are converted into target-removal episodes using the same deterministic annotation verifier for every method. Starting from the empty removed set, breadth-first search enumerates monotone states reachable only through accepted removals; a movable object is included as a target exactly when it is removed on at least one reachable transition. This produces 129 targets from 20 OOD10 scenes, 432 from 50 OOD15 scenes, and 574 from 50 OOD20 scenes, for 1,135 targets total. Every remaining movable object contributes five candidates in the fixed grasp order Top, Front, Rear, Right, Left. Candidate (o,g)(o,g) is accepted iff the object remains annotated, no remaining object names o as its support-frame parent, the official per-grasp IK annotation is feasible, and the sum of official geometric-obstruction failure ratios from remaining blockers is below 1−10−121-10^-12. Acceptance removes the object and clears the rejected-candidate set; rejection leaves the scene unchanged and suppresses only that object–grasp pair until the next accepted removal. For each remaining object, the five GRN grasp scores are taken from the second through sixth outputs of the official frozen model in the order above and z-normalized over the current object–grasp candidate set. The target-conditioned process utility is built from the current directed obstruction/support graph: reverse transitive depth from the target assigns larger utility to remaining blocker ancestors, while objects outside the target ancestry receive −1-1; this object utility is copied to its five grasps and independently z-normalized. GRN-only ranking uses the GRN z-score, process-only ranking uses the process z-score, and the combined methods use their unweighted sum. Static variants cache the corresponding initial scores, whereas refreshed variants recompute the induced-scene GRN score and/or target-conditioned obstruction depth after every accepted removal. Candidates are ordered by descending score, then a shared seed-fixed tie value, original movable-object index, and grasp index. All methods use the common 5N5N verifier budget. OOD10 uses one tie seed (0); OOD15 and OOD20 use seeds 0–4, which are averaged within scene–target before the 20,000-draw scene-clustered bootstrap and sign-permutation analysis (seed 20260809). CDGS-style search follows the population-guided mechanism of 23 on exactly this candidate, score, transition, and verifier interface. Each search round samples 12 imagined sequences of horizon at most four from a softmax with temperature .8.8; imagined steps remove sampled objects without querying the verifier. A sequence is scored by its length-normalized log sampling probability plus a bonus of 2 if it reaches the target. The top 25%25\% form the elite set. After the first round, inverse-rank-weighted elite first-action counts define a guide distribution whose log probability enters the next round’s first-step score with strength .5.5. After two rounds, the first action receiving the largest exponentiated elite-sequence vote is passed to the shared verifier; subsequent accept/reject handling is identical to the direct ranking methods. CDGS–GRN uses refreshed GRN scores alone, CDGS–static uses the cached combined score, and CDGS–refresh recomputes both GRN and process scores. Population-search seeds are 0 on OOD10 and 0–4 on OOD15/OOD20, with the same within-target averaging and scene-clustered inference described above. Table 18: Numeric timing and mismatch interventions. Every pair agrees on its first accepted action. Sequence disagreement and rescue counts come from formal condition traces; gaps compare refresh with the identical static aggregation. Aggregator Panel Pairs First same Sequence differs Dynamic-only / static-only AUC gap [95%CI][95\%\,CI] Rank meet CAD calls 576 100% 96.35% 32 / 5 +2.86+2.86 [.12,5.76][.12,5.76] Rank meet CAD work 384 100% 93.75% 26 / 0 +6.77+6.77 [1.30,14.06][1.30,14.06] Rank meet Packing ID 240 100% 67.50% 1 / 0 +8.54+8.54 [7.01,10.14][7.01,10.14] Rank meet Packing OOD 240 100% 93.75% 29 / 1 +8.68+8.68 [7.01,10.42][7.01,10.42] Product CAD calls 576 100% 97.40% 36 / 5 +3.30+3.30 [1.01,5.79][1.01,5.79] Product CAD work 384 100% 94.79% 17 / 0 +4.43+4.43 [.52,9.90][.52,9.90] Product Program ID 240 100% 98.75% 0 / 0 +12.42+12.42 [9.50,15.33][9.50,15.33] Product Program OOD 240 100% 100% 36 / 3 +22.58+22.58 [19.33,25.83][19.33,25.83] Product Packing ID 240 100% 97.08% 0 / 0 +6.46+6.46 [4.51,8.47][4.51,8.47] Product Packing OOD 240 100% 100% 0 / 2 +6.53+6.53 [4.58,8.47][4.58,8.47] Table 19: Dynamic–static AUC replication. Values are percentage-point gaps; brackets are paired 95%95\% intervals. Domain/resource Group Gap 95%95\% CI CAD calls sensor 701 / 709 +2.55/+3.18+2.55/+3.18 [−.29,5.44]/[.12,6.31][-.29,5.44]/[.12,6.31] CAD calls 5–8 / 9–20 parts +.61/+3.99+.61/+3.99 [−5.12,6.34]/[1.00,7.29][-5.12,6.34]/[1.00,7.29] CAD work sensor 811 / 823 +6.77/+6.77+6.77/+6.77 [1.04,14.06]/[1.04,14.06][1.04,14.06]/[1.04,14.06] CAD work 5–8 / 9–20 parts +3.91/+8.20+3.91/+8.20 [0,11.72]/[.39,17.97][0,11.72]/[.39,17.97] Program ID seed 20270117 / 20270129 +10.50/+14.33+10.50/+14.33 [6.50,14.50]/[10.17,18.67][6.50,14.50]/[10.17,18.67] Program OOD seed 20270117 / 20270129 +24.50/+19.00+24.50/+19.00 [19.50,29.33]/[14.67,23.33][19.50,29.33]/[14.67,23.33] Packing ID seed 20271103 / 20271117 +8.06/+9.03+8.06/+9.03 [5.83,10.42]/[6.81,11.25][5.83,10.42]/[6.81,11.25] Packing OOD seed 20271103 / 20271117 +7.36/+10.00+7.36/+10.00 [5.28,9.58]/[7.64,12.64][5.28,9.58]/[7.64,12.64] Appendix E Complete Numerical Budget Frontiers Table 20 gives the exact dynamic/static success percentages underlying Figure 3. Single-channel controls are reported in Table 3, and alignment interventions are summarized in Table 3(b). Table 20: All dynamic/static budget points used in the paper. Budgets were pre-specified before formal evaluation and are normalized by parts, events, remaining items, or one dense sweep per part. At each budget, the better value is bold and the second value is underlined; ties are unmarked. Domain/split Pre-specified budgets Dynamic / initial-static success at successive budgets CAD calls 1.25/1.5/2/2.5/3/41.25/1.5/2/2.5/3/4 68.92/74.13/78.65/83.16/90.10/94.6268.92/74.13/78.65/83.16/90.10/94.62 / 64.93¯/71.35¯/78.65/82.47¯/85.07¯/89.93¯ 64.93/ 71.35/78.65/ 82.47/ 85.07/ 89.93 CAD work 1.25/1.5/2/2.5/3/41.25/1.5/2/2.5/3/4 98.70/98.70/98.70/98.70/98.70/98.7098.70/98.70/98.70/98.70/98.70/98.70 / 91.93¯/91.93¯/91.93¯/91.93¯/91.93¯/91.93¯ 91.93/ 91.93/ 91.93/ 91.93/ 91.93/ 91.93 Program ID 1/1.25/1.5/1.75/21/1.25/1.5/1.75/2 48.33/82.92/94.58/99.58/100.0048.33/82.92/94.58/99.58/100.00 / 27.92¯/55.42¯/82.92¯/97.08¯/100.00 27.92/ 55.42/ 82.92/ 97.08/100.00 Program OOD 1/1.25/1.5/1.75/21/1.25/1.5/1.75/2 8.33/47.92/78.75/90.83/97.508.33/47.92/78.75/90.83/97.50 / 5.42¯/18.33¯/40.83¯/66.67¯/83.33¯ 5.42/ 18.33/ 40.83/ 66.67/ 83.33 Packing ID 1/1.125/1.25/1.5/2/31/1.125/1.25/1.5/2/3 9.58/31.25/44.17/72.92/89.58/100.009.58/31.25/44.17/72.92/89.58/100.00 / 9.58/22.50¯/32.92¯/50.42¯/81.25¯/99.58¯9.58/ 22.50/ 32.92/ 50.42/ 81.25/ 99.58 Packing OOD 1/1.125/1.25/1.5/2/31/1.125/1.25/1.5/2/3 0.00/4.58/16.25/30.42/57.92/95.000.00/4.58/16.25/30.42/57.92/95.00 / 0.00/2.50¯/5.42¯/16.67¯/44.17¯/83.33¯0.00/ 2.50/ 5.42/ 16.67/ 44.17/ 83.33