Paper deep dive
Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians
Timothy Heightman, Elena Orlova, Philip Mantrov, Aleksei Ustimenko
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:A central promise of useful quantum advantage is the ability to compute ground states of Hamiltonian systems beyond the reach of classical simulation methods. Here we demonstrate that this problem can be effectively amortized across an arbitrary and universal set of Hamiltonians by a foundation model with $\sim0.5$B variational parameters, trained with contemporary techniques from large language models and deep reinforcement learning. To do this, we formulate $\text{spin-}1/2$ quantum ground-state learning as manifold variational optimisation over centrally odd scalar functions on $\mathrm{SU}(2)^N$. This replaces explicit Hilbert-space vector amplitudes with manifold functions on which the Hamiltonian acts through Lie derivatives, evaluated by custom automatic differentiation primitives. We prove that the resulting variational principle on this manifold preserves the $\text{spin-}1/2$ sector's ground-state upper bound using the Peter-Weyl theorem, then pre-train our foundation model on a dataset of hundreds of thousands of different Hamiltonian systems, varying the connection topology, system size, interaction types and strengths, bringing together a century of many-body literature. Using a novel $\mathrm{SU}(2)$ replica-exchange Langevin sampler and sharded natural-gradient optimisation, we train our model with our own extension of the Kronecker-Factored Approximate Curvature (KFAC) optimiser on system sizes up to 64 qubits. On a held-out generalisation dataset, we fine-tune our model on system sizes of up to 1024 qubits, and evaluate on systems up to 8100 qubits.
Tags
Links
- Source: https://arxiv.org/abs/2608.11911v1
- Canonical: https://arxiv.org/abs/2608.11911v1
Trouble viewing inline? Open PDF directly →
Full Text
396,926 characters extracted from source content.
Expand or collapse full text
Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians Timothy Heightman Affiliation: Simulacra Research Inc., London, UK & Chicago, USA[0.15em] timothyheightman@simulacra-ai.com aleksei@simulacra-ai.com Elena Orlova Affiliation: Simulacra Research Inc., London, UK & Chicago, USA[0.15em] timothyheightman@simulacra-ai.com aleksei@simulacra-ai.com Philip Mantrov Affiliation: Simulacra Research Inc., London, UK & Chicago, USA[0.15em] timothyheightman@simulacra-ai.com aleksei@simulacra-ai.com Aleksei Ustimenko Affiliation: Simulacra Research Inc., London, UK & Chicago, USA[0.15em] timothyheightman@simulacra-ai.com aleksei@simulacra-ai.com [0.35em] ICFO – Institut de Ciències Fotòniques, The Barcelona Institute of Science and Technology,08860 Castelldefels, Barcelona, Spain Abstract A central promise of useful quantum advantage is the ability to compute ground states of Hamiltonian systems beyond the reach of classical simulation methods. Here we demonstrate that this problem can be effectively amortized across an arbitrary and universal set of Hamiltonians by a foundation model with ∼0.5 0.5B variational parameters, trained with contemporary techniques from large language models and deep reinforcement learning. To do this, we formulate spin-1/2spin-1/2 quantum ground-state learning as manifold variational optimisation over centrally odd scalar functions on SU(2)NSU(2)^N. This replaces explicit Hilbert-space vector amplitudes with manifold functions on which the Hamiltonian acts through Lie derivatives, evaluated by custom automatic differentiation primitives. We prove that the resulting variational principle on this manifold preserves the spin-1/2spin-1/2 sector’s ground-state upper bound using the Peter-Weyl theorem, then pre-train our foundation model on a dataset of hundreds of thousands of different Hamiltonian systems, varying the connection topology, system size, interaction types and strengths, bringing together a century of many-body literature. Using a novel SU(2)SU(2) replica-exchange Langevin sampler and sharded natural-gradient optimisation, we train our model with our own extension of the Kronecker-Factored Approximate Curvature (KFAC) optimiser on system sizes up to 64 qubits. On a held-out generalisation dataset, we fine-tune our model on system sizes of up to 1024 qubits, and evaluate on systems up to 8100 qubits. 11footnotetext: These two authors contributed equally. Code: https://github.com/simulacra-research/HamiltonZero Contents 1 Introduction 2 Theory 3 Hamilton-Zero’s Architecture 4 Pretraining 5 Results 6 Discussion and outlook S1 Variational principles on Spinor Manifolds S2 Foundation Ansatz and Entanglement Structure via Reinforcement Learning S3 Datasets and further detail on training S4 Optimization and Engineering A The Casimir regulariser for nonlinear ansätze B Proofs from Supp. Mat. S1 References 1 Introduction A central problem in quantum many-body systems is computing the ground state of a given spin-1/21/2 Hamiltonian. Through a fermion-to-qubit mapping such as Jordan-Wigner, it encodes the electronic-structure Hamiltonians of quantum chemistry, and with them this ground state problem is relevant from materials design through to drug discovery [13, 92, 26]. Ground states across a Hamiltonian family also trace out its phase diagram, as well as encode the solutions of combinatorial optimisation problems relevant to supply-chain [137], portfolio optimisation [54], electric vehical charging [145, 63] or flight routing [41] to name some examples [160, 31]. For a large many-body system, however, computing such ground states has proven famously difficult due to the curse of dimensionality. This curse refers to the fact that a spin-12 12 Hamiltonian on N sites has a ground state in 2N2^N complex dimensions. Decades of research have uncovered many complementary methods to solve this problem as generally as possible. Of note are tensor-network methods [136, 97, 98], variational quantum algorithms (VQAs) [111, 24], and neural quantum states (NQS) [23, 78]. The first is the natural tool in one dimension, where the ground states of gapped local Hamiltonians obey an area law and a fixed bond dimension captures them. In two dimensions, that same area law turns against the method, since the entanglement across a region’s boundary, and with it the bond dimension required, grows with that boundary, confining efficient tensor networks to a low-entanglement corner of Hilbert space. The second, VQAs, has proved initially successful for smaller many-body systems, but has trainability and expressivity tradeoffs due to the so-called barren plateau phenomenon [79, 117, 93]. As such, the third method, NQS, has taken off in popularity over the last years, solving for variational upper-bounds to ground states on large systems , with the largest systems approaching 1500−20001500-2000 qubits [152]. This accuracy on single systems alone however has not, until recently, proven transferable across different Hamiltonians [118]. Indeed, recent works aiming for this transferrability constitute early foundation neural quantum states (FNQS), in which a single neural quantum state is trained on many different hamiltonians [118]. This amortises the cost of training a new network from scratch at the per-system level. Thus far however, the community has gone in a direction where they fix the Hamiltonian’s interaction graph (i.e. the topology of the connections between spin sites), varying only the coefficients of a given Hamiltonian structure, and keeping both said structure fixed as well as the number of spins [162, 118], see Figure 1. This lineage is short but fast moving and highly active area of investigation. The transformer quantum state of Zhang and Di Ventra [165] was the first to carry a single network across many instances of one model class. Rende and collaborators [119] then pretrained at one point of a phase diagram and fine-tuned outward from it. The same group made the idea systematic in their foundation neural quantum states [121], a single network conditioned on the couplings of one lattice and swept across its phase diagram and disorder. The same line later settled a two-dimensional spin glass [151], and the construction carried over to lattice fermions [162]. The cost these models amortise is that of deep variational Monte-Carlo. This per-system paradigm exemplified in the electronic setting by FermiNet, PauliNet and DeepErwin [113, 58, 42], each of which solves a single system from scratch to high accuracy, at a price paid anew for every Hamiltonian. Several of the second-order optimisation primitives we build on below originate there. Within the spin-FNQS lineage reviewed above however, prior models condition primarily on the coefficients of a fixed interaction graph, with the only reconfigurable such variational models being the variational architectures on quantum computers, which are themselves also trained from scratch for every new instance of a problem. Hamilton-Zero (this work)spin-12 12 interactionsmolecular VMCOrbformer, Gaoelectron gasLEM, QERNELspin NQSFNQS, Attn-FMedge weightssystem sizetopologyinteraction typethis workprior foundation modelspartial / model-dependent Figure 1: The generalisation surface of foundation wavefunction models. Each bar is one pretrained-model family. Its length is how the model varies its coupling across its training and evaluation corpus, with the axes ordered by structural depth. Solid bars mark demonstrated generalisation, hatched bars partial or model-dependent reach, with each reviewed in the main text. Two neighbouring programmes do vary the structure, though outside the qubit paradigm considered here. The Large Electron Model [161] and its successor QERNEL [101] are foundation models for interacting electrons which generalise across particle number. QERNEL further conditions on the depth of the moiré superlattice potential, the single parameter that tunes a semiconductor heterobilayer between its liquid and crystalline electronic states. The molecular wavefunction programme reaches furthest of all on structure. It runs from weight sharing across geometries [133, 132] to the joint-molecule ansatz of Gao and Günnemann [39], and on to Orbformer [36], pretrained across 22,00022,000 molecular structures. These genuinely change the structure they condition on. Yet the type interaction remains fixed even as the geometric structure varies owing to the nature of the systems modelled. That is, in these models, we see only the Coulomb for the electrons, and a single symmetry class for the spins. Indeed, this obstruction is mechanical. A Hamiltonian read in as a fixed-length vector of couplings has fixed its interaction graph in the length and the meanings of the encoding of that vector. It cannot then be conditioned on a different type of interaction, nor a different number of spins. To cross from the coefficients to the structure, the Hamiltonian must enter as a variable and typed interaction graph. To-date, no such model has proven successful with such an architectural choice. Two final threads sit just off this map, neither a transfer model in itself. The first is the per-system transformer wavefunction of Roca-Jerat and collaborators [127]. The second is the neural scaling laws of Jiang et al. [65], our empirical reason to expect pretraining at the scale of hundreds of thousands of different Hamiltonians to amortize, also more recently reported in [120]. In this contribution, we introduce Hamilton-Zero, a novel foundation model that generalises across both multiple system sizes, different interaction topologies, and different interaction types, trained on a corpus of Hamiltonian systems spanning a century of quantum many-body literature. This pretraining corpus has two levels of structure. At the top level sit thousands of distinct quadratic interaction graphs. Within each graph family, we then vary both the interaction parameters, so that a single family’s data is comparable in scope to the entire corpus of an existing FNQS model, and the number of spins, up to N=64N=64. Combined with a hot-swapped perturbation scheme in pretraining, these expansions yield the hundreds of thousands of distinct training Hamiltonian topologies, system sizes and interaction types. For larger systems up to 80008000 qubits, in many instances our results offer energy bounds that are not accessible by tensor network methods, or DMRG due to the nature of the non-local interaction topologies made accessible by this model. Indeed, the only other architecture that could currently find these bounds is neural quantum states, trained from scratch on each instance of Hamiltonian interaction topology. However, this is amortised by the open-source release of this model’s weights. To do so, we work in a different configuration space for many-body spin systems, akin to the many-body phase space formalism of [55] with two notable improvements. First, we represent each factor SU(2)≅S3SU(2) S^3 by four global embedding coordinates subject to one unit-norm constraint. Thus the configuration manifold has intrinsic dimension 3N3N and is supplied to the network through 4N4N constrained coordinates. This representation results in a smooth, site-local differential structure, on which the Hamiltonian acts through Lie derivatives. These, in turn, enable the use of automatic differentiation primitives, scalable to 8000+ qubits, well beyond the early proposals of [55]. Second, our model genuinely respects the variational principle in the Hilbert space of spin-1/21/2 systems, something which was acknowledged to be an open issue of phase-space-based quantum machine learning [55]. Our model’s wavefunction hence constitutes a centrally odd scalar function on SU(2)NSU(2)^N, a function class properly containing all possible neural quantum states, as we prove in Theorem S1.2. To do so, when optimising, we constrain the model to the physical sector identified by the Peter-Weyl theorem. At 547,521,152547,521,152 variational parameters, Hamilton-Zero is, to our knowledge, the largest wavefunction ansatz yet trained for spin systems. At this scale, pretraining practices developed for large language models and reinforcement learning become directly relevant. On the fixed pretraining workload defined in SM § S4, an off-the-shelf JAX ecosystem [17] with folx [40] required more than one year of wall time. The forked versions we release with this work reduce that estimate to approximately four days. The rest of this work is structured as follows. In Sec. 2 we describe briefly the key mathematical steps that underpin our automatic differentiation primitives and our choice of manifold function. In Sec. 3 we briefly describe Hamilton-Zero’s structure and similarities to Large Language Models, including a brief overview of our deep reinforcement learning pipeline for inferring entanglement structure. Next in Sec. 4 we describe the key aspects of pretraining and data augmentation. Sec. 5 then presents the results, in which we show (i) a single shared optimisation trajectory improves the energy of a heterogeneous distribution of hundreds of thousands of Hamiltonians, (i) that it can solve for energies of unseen systems when evaluated on frozen weights, and fine-tune to within 1%1\% of ED energies on such systems with relatively little compute, (i) that Hamilton-Zero can be taken far outside the system-size distribution on which it was trained on three case-studies, (iv) that it can witness a phase transition from a single checkpoint of its weights, (v) that it becomes approximately equivariant with respect to gauge groups in spinor-physics and (vi) that the entanglement structure of both pretraining and larger out-of-distribution systems are effectively learned by the router. Of note here is our scaling laws, which empirically suggest our model has 5×5× better scaling than today’s LLMs. The supplementary material develops each component of Hamilton-Zero in full. In SM § S1, we derive our variational principle for the function class SU(2)N→ℂSU(2)^N , describing the necessary theory to facilitate Hamiltonian (and other observable) action via automatic differentiation. Here we also resolve an open problem around harmonic leakage for phase space quantum machine learning that enables future generations of the model to gain even more expressiveness. In SM § S2 we describe the foundation ansatz, which contains a featuriser for the Hamiltonian dataset’s differing interaction topologies, weights and system sizes, an LLM style trunk of attention and FFN layers, and a readout mechanism trained with deep reinforcement learning to enable generalisation across unseen system topologies and system sizes. In SM § S3, we lay out the training routines we employed for pretraining in full, as well as detail for the training, validation and evaluation datasets we created for pre-training, fine-tuning and zero-shot evaluation. SM § S4 details our research into optimisation and engineering that made this model’s training feasible. Of note here is our novel sampling scheme for MCMC on Riemannian manifolds which facilitates high-quality, decorrelated samples from our model’s quaternionic density at unprecedented speeds. Here we also detail an extension of the KFAC optimiser of Martens and Grosse [91], which allows it to work with higher-order Fisher-information tensors with a provably faster convergence rate than its rank-two counterpart. Everything above can be found in the open-source github repository for this paper, along with usage instructions for installing and using the model. It is also available on hugging-face. 2 Theory We consider the ground-state problem for the universal [102] spin-1/21/2 family of Hamiltonians, H^=14∑i<j∑a,bJijabσ^iaσ^jb+12∑i∑ahiaσ^ia, H= 14 _i<j _a,bJ^ab_ij\, σ^a_i σ^b_j+ 12 _i _ah^a_i\, σ^a_i, (2.1) where JijabJ^ab_ij encodes arbitrary quadratic Pauli interactions and hiah^a_i encodes local fields. It is useful to think of J as the data of a weighted graph on N sites, where each unordered pair (i,j)(i,j) with i≠ji≠ j carries a 3×33× 3 Pauli-pair coupling matrix Jij∈ℝ3×3J_ij ^3× 3, which we read as nine flavours of edge, one per Pauli-pair (a,b)∈x,y,z2(a,b)∈\x,y,z\^2 giving the operator σiaσjbσ^a_iσ^b_j, see Fig. 2. 1234 (a) (a) Heisenberg X model on a 1D spin-chain. Each nearest neighbour pair of sites carries X, Y and Z channels. 1234 (b) Cross-channel coupling (Kitaev-Γ ) on a 1D chain. Each bond carries a distinct off-diagonal Pauli pair, 1–2 is YZ, 2–3 is ZX, 3–4 is XY. Source and destination Pauli differ, so each edge changes colour at its midpoint. (c) Compass model on a 2D 2×32× 3 grid, where horizontal bonds carry X; vertical bonds carry Y. Bond orientation determines the Pauli channel. 12345 (d) A J1J1 Heisenberg chain with nearest-neighbour (J1J_1, straight) and next-nearest-neighbour (J2J_2, arcs) couplings, both Z. Figure 2: Four interaction-graph examples drawn from the training data. The colour map X, Y, Z shows different types of edges. For example, the diagonal entries a=ba=b (X, Y, Z) are the Heisenberg exchanges and the off-diagonal entries (XY, XZ, YX, YZ, ZX, ZY) encode Kitaev-Γ [144, 164], Dzyaloshinskii–Moriya [159], compass, and other cross-channel couplings. We will informally refer to the N×N× N edges as bonds. Meanwhile the diagonal (i,i)(i,i) self-bonds do not correspond to any quadratic interaction in this graph. Indeed, standard Neural Quantum States (NQS) [23, 165, 127] learn amplitudes for these interaction graphs on the discrete spin basis ↑,↓N\ , \^N. This makes each model tied to one system size and one Hamiltonian instance, so transfer across Hamiltonian structure is difficult [119, 121, 151]. Only recently researchers have tried to start generalising within a given Hamiltonian interaction topology [165, 119, 121, 151, 162]. Hamilton-Zero instead represents the state as a smooth function on the product Lie group ψθ:SU(2)N→ℂ _θ:SU(2)^N . Each spin is parametrised by a unit quaternion qi∈SU(2)≅S3↪ℝ4q_i∈ SU(2) S^3 ^4, so the network consumes 4N4N continuous coordinates rather than the discrete basis of NQS [23]. Thus spin operators can act on this manifold through left-invariant vector fields LiaL_i^a, σ^ia⟷−iLia σ_i^a -iL_i^a for first order and for second, σ^iaσ^jb⟷−LiaLjb σ_i^a σ_j^b -L_i^aL_j^b. The Hamiltonian becomes a second-order differential operator on SU(2)NSU(2)^N, which we can compute with automatic differentiation of the model with respect to its input coordinates, H^⟷−14∑i<j,a,bJabijLiaLjb−i2∑i,ahiaLia, H - 14 _i<j,a,bJ^ab_ijL_i^aL_j^b- i2 _i,ah_i^aL_i^a, (2.2) extending recent works on moment generating functions [55] to the manifold function directly via custom JAX-primitives of these Lie derivatives. Hence, the Hamiltonian’s interaction data is simply a bare coupling tensor (J,h)∈ℝN×N×9×ℝ3N(J,h) ^N× N× 9×R^3N, while its action on states is evaluated by autograd of the manifold wavefunction’s inputs. The quaternionic frame realising each LiaL_i^a and the log-derivative identities behind the local energy are Supplimentary Material (SM) §S1; the kernels evaluating these operators at scale are SM §S4.2. By moving from amplitudes to functions on SU(2)NSU(2)^N, we have a space L2(SU(2)N)L^2(SU(2)^N) which is larger than the physical spin-1/21/2 Hilbert space. However, the Peter–Weyl theorem [52, 112] gives us a spectrum for this function space, L2(SU(2))=⨁j=0,12,1,…Vj∗⊗Vj,L^2(SU(2))= _j=0, 12,1,…V_j^* V_j^, (2.3) decomposing it site-wise into irreducible spin sectors, VjV_j. The eigenfunctions in the physical subspace satisfy −Δiψ=34ψ- _iψ= 34ψ for every site i, where Δi _i is the Laplace–Beltrami operator on each local set of SU(2)SU(2) coordinates and −Δi- _i realises the per-site Casimir S^i2 S_i^2 [35, 74]. All physical spin-1/21/2 functions on this manifold satisfy this Casimir relation, however an unconstrained neural function on SU(2)NSU(2)^N may leak into higher j sectors giving unphysical energies below the ground state. Our ansatz avoids the leak by construction, as we impose per-site oddness ψ(…,−qi,…)=−ψ(…,qi,…)ψ(…,-q_i,…)=-ψ(…,q_i,…), and we keep the wavefunction linear in each site’s quaternion. This is because oddness removes all integer-j sectors, and linear functions of qiq_i are exactly the span of the four spin-1/21/2 Wigner D-matrix elements Dmn1/2(qi)D^1/2_mn(q_i), so a centrally odd wavefunction that is linear in every quaternion lies analytically in the spin-1/21/2 sector, ψθ(q1,…,qN)=∑,T,(θ)∏k=1NDmknk1/2(qk), _θ(q_1,…,q_N)= _ m, nT_ m, n(θ) _k=1^ND^1/2_m_kn_k(q_k), (2.4) On this sector, the Lie-derivative operator is unitarily equivalent to the physical Hamiltonian, so we may therefore define a variational principle, ⟨ψθ|H^|ψθ⟩⟨ψθ|ψθ⟩≥E0, _θ| H| _θ _θ| _θ \;≥\;E_0, (2.5) and every energy our optimiser reports is a rigorous upper bound on the physical ground-state energy. See SM § S1 for a full exposition and derivation of this variational principle. We emphasise that despite this linearity, nothing prevents a non-linear conditioning on the Hamiltonian’s interaction data, which is precisely where Hamilton-Zero’s expressivity lies. We defer details of its architecture to Sec. 3 and SM § S2. Had the ansatz been nonlinear in the quaternions however, oddness alone would leave the higher half-integer sectors open, meaning leakage could occur in principle. However, we can resolve this open problem of leakage to higher-spin sectors recorded in [55] with a regulariser that pushes the model towards the physical sector, and an energy penalty that analytically restores the variational principle. For a full discussion on these cases, we refer to Sup. Mat. S1 and Appendix A for the proofs. For this generation of model however, the regular variational principle is sufficient when we have centrally odd functions that are multi-linear in the manifolds quaternionic coordinates. We finish this section with a remark on the expressivity of our model’s ambient function class in comparison to NQS. In brief, NQS span a variational sub-manifold of their full Hilbert space unless they have control over (2N)O(2^N) independent amplitudes through variational amplitudes, which is dimensionally cursed (and the reason why practical NQS architectures span a smaller variational sub-manifold). Hamilton-Zero does not inherit this restriction. Since the model is centrally odd and multilinear in the quaternion coordinates, it has the capacity to act in ℱlin≅(V1/2∗⊗V1/2)⊗N,dimℱlin=4N,F_lin\; \; (V_1/2^* V_1/2 ) N, _lin=4^N, (2.6) which strictly contains the canonical lift of the entire NQS Hilbert space, and therefore also contains every practical NQS variational manifold. The Hamiltonian acts on the one Peter–Weyl index and leaves the additional index untouched, so this larger function class still respects the physical variational principle. If we allow the model to depend non-linearly on each quaternion (but remain centrally odd), its capacity grows again to include all higher half-integer sectors. This however requires an energy penalty to remain a variational principle on the original Hilbert space, as well as regularisation. We discuss these aspects further in SM § S1 and prove an expressivity hierarchy in Theorem S1.2. The theorem compares ambient function classes; the expressivity of a finite parameterisation remains architecture-dependent. 3 Hamilton-Zero’s Architecture In this section, we give a brief account of the architecture of our model, deferring its technicalities to SM § S2. The architecture has two types of inputs. First is a sampled spin configuration q∈SU(2)Nq (2)^N, and second is a Hamiltonian’s interaction tensor (J,h)(J,h). The two are routed through separate pathways and meet at a single, controlled point, so that the central oddness and per-site linearity of Sec. 2 are preserved exactly by construction (see Fig. 3). We give a brief account of each component here, deferring technicalities to SM § S2. (J,h)(J,h) Spectral normalise Featuriser Pre-norm trunk ×L× L Leaf contextualiser Odd leaf builder Merge-tree readout logψθ(q|p) _θ(q\,|\,p) Route policy deep RL Spin config q MCMC on SU(2)NSU(2)^N global stream g Figure 3: Hamilton-Zero information flow. The Hamiltonian (J,h)(J,h) is spectrally normalised, featurised in polar form, and propagated through the trunk; a route policy conditioned on the trunk’s outputs permutes sites and padding slots before the leaf contextualiser; the spin configuration q enters only at the odd leaf builder, enforcing per-site oddness before the shared rank-four merge tree returns logψθ(q|p) _θ(q\,|\,p). The rail beneath the pipeline is the global stream g, refreshed at every stage on the unit-RMS sphere. Component diagrams are SM §S2. The Hamiltonian data enter the model through a fixed cache followed by a learnable featuriser. The cache bond-symmetrises J and forms a temporary Hermitian matrix from J and h solely to determine one shared spectral scale. Once that scale has been computed, the diagonal field block is discarded and the normalised bond tensor and local field are supplied to the model as separate channels. This places every system in a common numerical regime without needing the model to relearn an overall energy scale. Next, a featuriser then reads the spectrally normalised Hamiltonian data and produces embeddings for it for the downstream blocks of our model. The featuriser normalises in learned groups by root-mean-square (RMS) maps, and aggregates by column and row attention to produce a per-bond embedding, a per-site embedding, and a global context vector. Bonds and fields absent from a system substitute learned absent tokens rather than zeros. This is so that the absence of a connection, which changes the system’s topology, is a representable direction of the feature basis. Meanwhile padding, which carries no structural meaning to the system’s interaction topology, is done with masks. Once the Hamiltonian’s data is embedded by the featuriser, it is passed through L pre-norm residual blocks. Each block updates the bonds first through an edge map, biases the per-site attention logits with the fresh bonds so heads can attend along the interaction graph, and closes with a feed-forward update. Every sublayer normalised by the same root-mean-square rule and is damped by a 1/L1/ L residual gain. A global data-stream runs alongside this so that downstream components can also act directly on the feature basis. This global stream lives on the sphere of unit root-mean-square norm, and is refreshed once per block by descriptor attention over the states the block has just written. It moves by the normalised-interpolation update of nGPT, so near-identical readings from symmetric systems saturate instead of accumulating with depth. We emphasise that the spin configuration never enters this trunk, and that everything up to the leaves is a function of (J,h)(J,h) alone. This means the bulk of the model’s parameters do not need any symmetry constraints that would limit its expressivity. The symmetry constraints are enforced at the single point where quaternionic coordinates enter. For the coordinate dependence, a bias-free map first lifts each qiq_i into a linear odd data-stream. Hamiltonian context supplies only the coefficients of the map applied to that stream. The resulting raw leaf carrier is therefore linear in qiq_i. We normalise the carrier, u, for numerical stability, record the divided-out leaf magnitude in a log-scale, s, and restore it at the root together with all later merge scales. The normalised carrier alone is odd but nonlinear, meaning only the represented pair (u,s)(u,s) reconstructs the linearity property needed from Sec. 2 carrier exactly. Bilinear tree merges then make the final amplitude multilinear in the site coordinates, while the Hamiltonian-only stream retains the model’s nonlinear capacity. Finally, a readout mechanism contracts the N carriers through a balanced binary tree of merges, in the spirit of tree tensor networks [136, 97, 98], with leaves in a bit-reversed canonical frame and one merge shared across every level and position. The shared merge is a rank-four tensor contracting blockwise, bilinear in its two children so a sign flip at any leaf passes to the root and thus preserves our symmetry constraints from Sec. 2. The leaves and every internal node renormalise their carriers to unit root-mean-square while storing the divided-out magnitudes in an additive log-scale, all of which is again restored at the root readout. A Hamiltonian context stream is computed up the tree beside the carriers by the same nGPT-style normalised interpolation as the global context data-stream. This context stream reads the coarse-grained bond field between the two subtrees being merged, and a level attention with a tree-distance bias lets subtrees exchange information laterally, deleting the depth penalty on long-range correlations. Thus our merge tree is conditioned by a coarse-graining of the even data-stream, which we think of as a neural augmentation of tree tensor networks. Note that padding slots carry learned contexts of their own, so one tree serves many system sizes. A learned site-to-slot map adapts the contraction tree to different interaction topologies and system sizes. To learn entanglement structures, we make a map from physical sites to the leaf slots of the merge contraction itself learnable. This map is a discrete latent variable between physical site indices and leaf slots, which we learn with deep reinforcement learning [141, 37]. A pointer-network policy conditioned on the trunk’s outputs decodes a permutation slot by slot. At every row it composes every remaining candidate afresh from its site state, the fixed Hamiltonian global, the decode clock, bidirectional prefix and suffix bond summaries, an order-aware prefix, and hole-placement statistics. The selected row states become the leaves of a partial dyadic merge tree. Causal same-level refinement builds its complete subtrees, their mixed-scale dyadic cover produces the next query, and the remaining candidates attend the cover, earlier queries, and one another before the pointer chooses. A conditional quotient then removes indistinguishable choices before sampling. The sampled route relabels sites, bonds, and walkers before the merge mechanism and its context construction runs, so that the corresponding context around each site and its bonds is carried through the learned reordering. From a physical perspective, the model represents the incoherent mixture of states over sampled routes, a relaxation that costs nothing because the energy is linear in the density matrix and is minimised at a pure state. The policy trains by a score-function gradient with a beam approximation to the policy mode and a self-normalised importance-sampling baseline evaluated on the routed walkers. The walker channel is centred at its route-conditional mean. Here π is the routing policy, m is the slot mask, and Aut(J,h,m)Aut(J,h,m) is the group of slot permutations preserving the bond tensor J, site fields h, and mask m. For a route p, its orbit and stabiliser are (p):=g⋅p:g∈Aut(J,h,m),Stab(p):=g∈Aut(J,h,m):g⋅p=p.O(p):=\g\!·\!p:g (J,h,m)\, (p):=\g (J,h,m):g\!·\!p=p\. (3.1) For every g∈Aut(J,h,m)g (J,h,m), a route p and g⋅pg\!·\!p have equal energy. A symmetrised policy is uniform within each occupied orbit. Writing H[π]:=−∑pπ(p)logπ(p)H[π]:=- _pπ(p) π(p) for the policy’s Shannon entropy, a policy supported on one orbit (p)O(p) has H[π]=log|(p)|=log|Aut(J,h,m)||Stab(p)|.H[π]= |O(p)|= |Aut(J,h,m)||Stab(p)|. (3.2) Thus the measured entropy of a learned policy and its analytic correspondence to automorphisms of the interaction graph provide us with a heuristic tool to measure whether the router’s training has converge to the true optimum, see SM § §S2.5.4. Indeed in SM §S2 we give a full exposition of this component along with all the others, including a proof of optimality when the router learns automorphisms. We now proceed to Hamilton-Zero’s training data and pretraining. 4 Pretraining To train Hamilton-Zero, we collated a curated dataset of 50005000 different interaction topologies, spanning a century of quantum many-body and quantum computing literature [5, 83, 59, 87, 62, 140, 122, 72, 25, 129, 130, 134, 95, 19, 57, 80, 73, 6, 11, 131, 12, 100, 1, 107, 84, 48, 154, 75, 76, 66, 18, 146, 46, 99, 9, 147, 15, 2, 158, 128, 71, 88, 32, 96, 77, 103]. Because the corpus also includes Hamiltonians native to superconducting, trapped-ion, Rydberg, spin-qubit, and molecular platforms, Hamilton-Zero can be evaluated as a classical foundation surrogate for hardware-relevant ground-state problems. The dataset is stratified into physics cells and coverage tiers, from canonical anchors through frustrated, disordered, gauge and volume-law regimes, each system built deterministically, see SM §S3 for a full account of the contents of the training dataset. By the standards of large language models, however, such a dataset is small, and trained on it as-is, the model would memorise couplings, fields, and frames alike. As such, here we briefly discuss the important aspects of the augmentations of our training datasets that prevent overfitting to specific edge-values and field-values, or overfitting to the topologies seen in pre-training. To that end, our training routine involves a series of perturbations of the training data. The first tier perturbs the physics itself, new bonds, new fields, new reference energies to compute; the second applies exact symmetry transformations, which generate unlimited new coupling tensors whose physics, and whose references, carry over unchanged. Together the two tiers expand the dataset into the hundreds of thousands of distinct training Hamiltonians of our corpus, which is again a relatively small dataset by today’s standard for large-language models. During training, every new epoch receives a perturbed version of the pretraining data, which is made in parallel to the training loop’s live progress. This perturbed data is then hot-swapped at every epoch end. To keep the training loop running continuously, this perturbation routine stays 2 rounds ahead of the loop, precomputing exact-diagonalisation references for every perturbed system to N≤22N≤ 22. Each structured system receives one of five changes at equal probability. The changes can be one of the following: (i) leave the system untouched, (i) add one new bond, (i) add noise to one existing bond, (iv) remove one bond, or (v) adds one random orbit-symmetric field, with sites, bonds, and orbits sampled stratified over the Weisfeiler–Leman orbit partition of (J,h)(J,h). Note that for (v) the magnitudes are set relative to the system’s own coupling scale. The orbit-symmetric field draws one random three-vector of strength η∼(0.25,1)η (0.25,1) times the coupling scale and adds it to every site of one orbit. For small random systems N≤8N≤ 8 , we additionally densify towards all-to-all coupling, with a field-free such system gaining a substantial random dense field half the time, and occasional bond drops. On top of each round’s perturbation the worker applies an exact symmetry change. It draws one Haar-random u∈SU(2)u (2) per Weisfeiler–Leman orbit of the interaction graph, shared by the sites of that orbit. Write uiu_i for the draw assigned to the orbit containing site i. Let Rker(ui)R_ker(u_i) denote the adjoint rotation expressed in the kernel’s ordered (x,y,z)=(k^,ȷ^,ı^)(x,y,z)=( k, , ) frame. The exact gauge augmentation is Jij↦Rker(ui)JijRker(uj)⊤,hi↦Rker(ui)hi,qi↦qiu¯i,J_ij R_ker(u_i)J_ijR_ker(u_j) , h_i R_ker(u_i)h_i, q_i q_i u_i, (4.1) where u¯i=ui−1 u_i=u_i^-1. Under this right multiplication the left-invariant fields transform by the adjoint action, so the simultaneous rewrite of (J,h,q)(J,h,q) leaves the local energy and spectrum unchanged. Sharing u within a class preserves the interaction graph’s symmetry structure and leaves the routing algorithm unchanged. Physically, this means the model cannot overfit to a choice of frame, and the physics of the interaction graph is identical under this transformation. Hence, training across gauged rounds penalises whatever gauge violation the model carries. What this encourages is approximate SU(2)SU(2) covariance resolved per orbit. Single-site fields transform by one rotation per orbit; a bond between two different orbits sees independent left and right factors, hence approximate SU(2)×SU(2)SU(2)×SU(2) covariance across inter-orbit bonds, and in the random small-N band the rotations are fully per-site, so every bond there trains under independent factors. The walkers transform on SU(2)SU(2) itself, the double cover of the SO(3)SO(3) rotations acting on the couplings, and the prediction delta across a gauge transformation is a direct probe of learned covariance (see SM §S3.2). 5 Results In this section, we show that one shared optimisation trajectory improves the energy of a heterogeneous distribution of Hamiltonians in a quantitatively regular way in 5.1. We opt for the V-score [157] as our measure of quality here, as it provides a fair comparison across system types and sizes. In 5.2, we show the results for zero-shot inference and fine-tuning. Here, zero-shot refers to taking Hamilton-Zero’s frozen model weights and evaluating them on the three datasets above, and hence zero training steps were taken for those shots. We applied both to three different datasets detailed in SM § S3.3. In brief, these datasets cover (A) combinatorial optimisation, (B) unseen interaction and channel topologies, and (C) systems we observed to be hardest when initially developing Hamilton-Zero that we subsequently held out for the model weights released in this work. We next show how the same pretrained wavefunction can be taken far outside the system-size distribution on which it was trained in 5.3. We study three rather different problems: weighted MaxCut instances [44, 45], the Pariser–Parr–Pople (P) model of a conjugated carbon chain [108, 109, 114], and frustrated Heisenberg antiferromagnets on square and triangular lattices [14]. Together, these cases test non-local weighted graphs, long-range fermionic interactions, and geometrically local frustration. This allows us to see the maximum system sizes Hamilton-Zero can compile and sample from, and whether the resulting zero-shot and fine-tuned energies remain close to reference values when available at such scales. Next in 5.4 we show that Hamilton-Zero can witness a phase transitions via the fidelity susceptibility method [49]. We will apply this to the Ising model and the J1J2 Heisenberg model. We then show that Hamilton-Zero recovers approximate equivariance with respect to the gauge transformations in the perturbation scheme of its pretraining data in 5.5. Finally, in 5.6 we show how the router learned the entanglement structure of our pretraining dataset, how its commitment to a certain entangling contraction path is indicative of the quality of the energy bounds, and that it generalises to unseen interaction topologies and interaction types on the scale of the large-system case studies. Here we will also show the overall model has learned approximate equivariance with respect to gauge groups actions on the interaction graph (J,h)(J,h). 5.1 Pretrain and scaling laws After the initial transient, Fig. 4 shows the V-score falls approximately as C−0.53C^-0.53, while the ED-relative gap falls as C−0.61C^-0.61. Their similar compute dependence shows that local-energy fluctuations remain informative about energy accuracy across the training distribution, including where ED is unavailable. Moreover, the convergence of V1/2V^1/2 across the system-size bands shows that its normalisation removes most of the extensive dependence on N. These exponents describe this pretraining run and its evolving Hamiltonian mixture. We note that C is measured in cumulative eight-H200 hours rather than exact FLOPs, so this result is hardware dependent. By the standards of neural scaling laws in other domains like LLMs, this exponent shows highly favourable scaling with compute, sitting around 5×5× better than the values reported for LLMs [69]. Figure 4: Compute scaling of variance and variational energy error through step 280 000. (a) V-score decreases approximately as V∝C−0.53V C^-0.53 after the initial transient. (b) The observed relative gap for ED-covered systems decreases approximately as C−0.61C^-0.61. (c) V1/2V^1/2 approaches a common late-time scale across the six system-size bands, showing that the normalization removes most of the extensive size dependence. (d) V1/2V^1/2 and the predicted relative gap decrease together over the full 5,000-system dataset, connecting variance reduction to improved inferred accuracy. Compute is cumulative execution time across eight NVIDIA H200 GPUs. The V-score is V=NVar(EL)/(E−E∞)2V=N\,Var(E_L)/(E-E_∞)^2, with E∞=0E_∞=0 here, and is therefore a normalized variational-accuracy metric rather than raw variance. The relative gap denotes r=(E−EED)/|EED|r=(E-E_ED)/|E_ED|. Solid curves show medians and shaded regions show interquartile ranges. Next we turn our attention to the zero-variance principle [14]. As shown in Fig. 5, the bounded zero-variance model follows the calibration data without a systematic residual drift with N, and a calibration restricted to N≤18N≤ 18 predicts the unseen N=19N=19–22 systems with relative ease. This is supporting evidence for transfer across size within the ED-accessible regime. Figure 5: Zero-variance calibration and held-out size validation. (a) Relative gap against V-score, both defined as in Fig. 4. Here we show the median and IQR as a blue line and shaded blue region, with a bounded zero-variance principle in dashed orange, obtained from q=(E−EED)/|E|=κVq=(E-E_ED)/|E|=κ V with κ=0.2825κ=0.2825 and r=q/(1+q)r=q/(1+q); a free robust fit gives r∝V0.89r V^0.89. (b) The median calibration residual shows no monotonic drift with N, supporting transfer across system size. (c) Fit of gap prediction showing sizes N≥18N≥ 18 follow the same relationship as 19≤N≤2219≤ N≤ 22. Blue points belong to the fitted N≤18N≤ 18 population and orange points to the unseen N=19N=19–22 population. Their overlap around the one-to-one line demonstrates held-out size generalization within the range where ED is still applicable. See also Fig. 6 for further size extrapolation. Subject to that qualification within the ED regime, Fig. 6 shows continuity at the N=22N=22 ED boundary (beyond which we no longer computed ED reference values). For N>22N>22, the predicted median relative gap is 3.45%3.45\%, with an interquartile range of 1.86%1.86\%–10.21%10.21\%. The median therefore remains at the few-percent scale, while the widening distribution shows that larger systems are increasingly heterogeneous rather than uniformly harder. Finally, the exchange discrepancy tracks the total-energy error, whereas the field discrepancy is smaller. Since the component energies are not separately variational and may cancel, this identifies where the remaining energy discrepancy lies. Figure 6: Energy gap extrapolation across full (a) Orange points show observed median relative gaps where ED is available; the blue curve and band show the median and interquartile range of the zero-variance predictions of Fig. 5. The dashed line marks the N=22N=22 ED boundary. The extrapolation remains continuous across this boundary indicating the pretraining run’s relative gap is consistently non-monotonic with N, rather than some cuttoff past the ED regime in accuracy. (b) Total, exchange, and field discrepancies are normalized by |EED||E_ED| of the total Hamiltonian. Only total energy obeys the variational bound; component discrepancies are displayed in absolute value. Exchange tracks the total error, while the field contribution is smaller. (c) Predicted relative-gap distributions for N>22N>22 show modest median variation but broader system-to-system tails with size. Across all N>22N>22 systems, the median prediction is 3.45%, with interquartile range 1.86%–10.21%. 5.2 Zero-shot Inference and Fine-tuning In order to evaluate the model, we let the pretrained router select the tree ordering, and compile the wavefunction on that ordering, and sample the fixed model for 512 measurement steps with 256 walkers across eight tempered replicas and 24 adaptive Langevin moves per step (see SM § S4 for full details on our sampling scheme), retaining 512×256×8≃1.05×106512× 256× 8 1.05× 10^6 samples per system. Energies are reported as E=Ekernel/2E=E_kernel/2 with the pooled standard deviation of per-walker means as the uncertainty. Zero shot, the pretrained model’s energies against ED values (where available) are reported in the top row of Fig. 7, with the cumulative frequency of systems below a given percentage threshold on the bottom row. We see from the top row an overall agreement of the systems energies against ED values, with median signed gaps of 4.02%4.02\%, 4.21%4.21\% and 12.63%12.63\% on set (A) , (B) and (C). There are 28%28\%, 21%21\% and 3%3\% of systems within 1%1\% of ED for these three sets, and the degradation across them is structured. Indeed, the gap’s rank correlation with system size grows from 0.140.14 (A) through 0.250.25 (B) to 0.470.47 (C). Within each evaluation set, the hardest systems are on a distinct family: the maximum-independent-set instances on (a) (gaps up to 31%31\%), the t-fractals on B (3535–39%39\%), and the Schwinger chains at nonzero θ on C (4040–50%50\%), consistent with Axis C sitting farthest from the pretraining distribution (see also SM § S3). Figure 7: Held-out energy accuracy for the 512-step evaluation. Panels (a)–(c) show the magnitudes of the variational Monte Carlo and exact-diagonalization energies per site for Axes A, B, and C, respectively, on logarithmic axes. Energy magnitudes are plotted because the reported energies are negative. Every point carries its corresponding walker-mean tail standard-deviation bar; where no bar extends beyond its marker, the uncertainty is smaller than the marker diameter. Panels (d)–(f) show the empirical cumulative percentage of systems whose relative energy gap is at most the threshold for Axes A, B, and C, respectively, with a logarithmic gap-threshold axis. Colors identify the same axis in both rows. All energies are additionally divided by N for per-site values. The empirical-CDF denominators are the systems with exact-diagonalization references: 89 for Axis A, 66 for Axis B, and 79 for Axis C. Next, we turn our attention to fine-tuning. For fine-tuning, we restart from the same checkpoint as evaluation, let the pretrained router pick the tree ordering, and compile the wavefunction on that ordering. We then only train the merge-tree on that fixed ordering, leaving a compiled ansatz with ∼ 4.7M parameters, against 547M for the full model as the object being trained. Our rationale for this is that the trunk, leaf contextualiser and other upstream parts of the Hamiltonian context need not be retrained on a single system since they are amortizing over many different Hamiltonian families, so it is counter-productive to train them to fit the key features of a single system. This is also justified by the fact that once compiled, we can then fine-tune on a single A100 GPU at a considerably faster step-weight owing to the fact that the number of variational parameters is <1%<1\% of the overall model. Each system is optimised independently on a single A100 for 10410^4 KFAC steps with 256 walkers, four MCMC moves per step and a 128-step burn-in, with learning rate 0.01/(1+t/104)0.01/(1+t/10^4) (ending at ≈0.005≈ 0.005), damping 10−310^-3, and curvature EMA 0.90.9. After training, every system is re-evaluated with its new tree weights. The final checkpoint is resumed at an exact zero learning rate for an additional 512 measurement steps with the sampler state carried over. ED references are never supplied to evaluation or fine-tuning and enter only in the analysis below. Fig. 8 shows the results for fine-tuning against ED values (where available). We see the remaining errors from zero-shot decrease by two to three orders of magnitude on every set of the (A)-(C) evaluation datasets. The median signed gap falls to 3.1×10−4%3.1× 10^-4\%, 1.2×10−3%1.2× 10^-3\% and 6.0×10−3%6.0× 10^-3\% respectively for (A)-(C), with 90%90\%, 94%94\% and 100%100\% of systems landing within 1%1\% of ED, and 75%75\%, 49%49\% and 35%35\% already at the 10−3%10^-3\% error. Because every fine-tuned value comes from the frozen-weight re-evaluation rather than from the training tail, and the two agree within combined uncertainties on all 256 systems (|z|≤2.7|z|≤ 2.7), the comparison against zero shot is like for like, and we see the failure modes do not persist through fine-tuning. Hamilton-Zero was thus able to find these ground states too. The cost of this accuracy is modest, with convergence heavily front-loaded in the number of training steps as shown in Fig. 9. The mean gap drops below 1%1\% within a few hundred steps on B and C, and Axis A plateaus at 3.1%3.1\% by roughly 20002000 steps. In terms of compute cost, we took 234234 GPU h to fine-tune all 256 systems of sets (A)-(C), plus 2424 GPU h to re-evaluate them, against 4747 GPU h for the zero-shot evaluation alone, roughly 5.5×5.5× the zero-shot cost for the two to three orders of magnitude in accuracy. Wall time is set by the padded tree width rather than by physical N, so cost is determined by the power of two a system falls into, with the largest instances (N=40N=40–5454) taking about 3.53.5 h each. Figure 8: Held-out energy accuracy for the 512-step evaluation. Panels (a)–(c) show the magnitudes of the variational Monte Carlo and exact-diagonalization energies per site for Axes A, B, and C, respectively, on logarithmic axes. Energy magnitudes are plotted because the reported energies are negative. Every point carries its corresponding walker-mean tail standard-deviation bar; where no bar extends beyond its marker, the uncertainty is smaller than the marker diameter. Panels (d)–(f) show the empirical cumulative percentage of systems whose relative energy gap is at most the threshold for Axes A, B, and C, respectively, with a logarithmic gap-threshold axis. Colors identify the same axis in both rows. All energies are additionally divided by N for per-site values. The empirical-CDF denominators are the systems with exact-diagonalization references: 89 for Axis A, 66 for Axis B, and 79 for Axis C. Figure 9: Convergence of the fine-tuning runs on the ED-referenced systems of datasets A, B, and C, respectively. Lines show the per-axis median of the absolute signed relative energy gap to exact diagonalization as a function of fine-tuning step, and shaded bands span the interquartile range across systems. Gaps are computed from the per-step walker-mean energies recorded every 50 steps, so the late-time level reflects single-step Monte Carlo noise of order 10−2%10^-2\%. The converged medians quoted in the text are lower because the frozen-weight evaluation averages 512 such steps. 5.3 Large System Case Studies We begin by recalling some important properties of the three types of systems we study here: MaxCut, a P carbon chain, and frustrated 2D magnets. We then show the results for zero-shot and fine-tuning for each system. For a weighted graph G=(V,ℰ,w)G=(V,E,w), MaxCut asks for a bipartition of the vertices which maximises the total weight of edges crossing the partition. If zi∈−1,+1z_i∈\-1,+1\ labels the side containing vertex i, its objective is C(z)=12∑(i,j)∈ℰwij(1−zizj).C(z)= 12 _(i,j) w_ij(1-z_iz_j). (5.1) We drop the additive constant and give Hamilton-Zero the diagonal Ising Hamiltonian H=12∑(i,j)∈ℰwijZiZj,⟨C⟩=12∑(i,j)∈ℰwij−⟨H⟩.H= 12 _(i,j) w_ijZ_iZ_j, C = 12 _(i,j) w_ij- H . (5.2) The ground state is consequently a computational-basis state encoding an optimal cut. The variational energy gives an expected cut weight, whilst sampling the wavefunction also returns ordinary feasible cuts. We report both quantities below, where the first tests the full distribution represented by the wavefunction, whereas the second is the relevant output if the model is used as a combinatorial optimiser. Next, the P Hamiltonian is an interacting π-electron model for conjugated carbon systems. For a chain of L carbon sites, at half filling, we use the particle–hole-centred form HPPP= H_P= −∑⟨i,j⟩,σtij(ciσ†cjσ+cjσ†ciσ)+U∑i(ni↑−12)(ni↓−12) - _ i,j ,σt_ij (c _iσc_jσ+c _jσc_iσ )+U _i (n_i - 12 ) (n_i - 12 ) +∑i<jVij(ni−1)(nj−1),Vij=U1+(Urij/e2)2,e2=14.397eVÅ. + _i<jV_ij(n_i-1)(n_j-1), V_ij= U 1+ (Ur_ij/e^2 )^2, e^2=14.397\,eV\, A. (5.3) Here tijt_ij is the nearest-neighbour hopping, U the on-site repulsion and VijV_ij the Ohno interpolation for the long-range Coulomb interaction [104]. Its Jordan–Wigner image contains N=2LN=2L qubits, one for each spin orbital. We report the centred energy per carbon, Ecent/NCE_cent/N_C; any comparison must use this convention. In particular, the infinite-chain value quoted below is a literature scale for the corresponding P convention rather than a rigorous finite-chain bound, since Coulomb geometry and end corrections differ between calculations [138]. Finally, the square and triangular spin systems can be written without committing to a drawing of the lattice. Let ℰ1E_1 and ℰ2E_2 denote the sets of nearest- and next-nearest-neighbour bonds. We study HJ1J2=J1∑(i,j)∈ℰ1i⋅j+J2∑(i,j)∈ℰ2i⋅j,i=12(Xi,Yi,Zi).H_J_1J_2=J_1 _(i,j) _1S_i\!·\!S_j+J_2 _(i,j) _2S_i\!·\!S_j, _i= 12(X_i,Y_i,Z_i). (5.4) On the square lattice, ℰ1E_1 contains the horizontal and vertical bonds and ℰ2E_2 the plaquette diagonals. The case studied here has J2/J1=1/2J_2/J_1=1/2, near the maximally frustrated region of the model [47]. The triangular calculation instead sets J2=0J_2=0 and places the J1J_1 bonds on a triangular lattice, for which every bulk spin has six neighbours and the antiferromagnetic interactions cannot all be minimised simultaneously. The four constructions are shown in Fig. 10. Figure 10: Large-system case studies. (a) A weighted graph for MaxCut. Blue and white vertices show a candidate bipartition; orange edges cross the cut, while the labels illustrate that the instances are weighted. (b) The P model on a conjugate carbon backbone, drawn in standard skeletal notation. Alternating bonds support the π system, with nearest-neighbour hopping ti,i+1t_i,i+1 and the long-range Ohno interaction VijV_ij; two spin orbitals per carbon map to N=2LN=2L qubits. (c) The square J1J_1–J2J_2 model, with solid nearest-neighbour and dashed diagonal bonds. (d) The nearest-neighbour triangular Heisenberg antiferromagnet. The drawings illustrate the interaction graphs, and the evaluated systems are larger than those drawn. On these systems, Tab 1 shows how Hamilton-Zero’s performance scales. We see relative agreement up to the 2000 qubit scale, with the model degrading at 4000 and deteriorating at 8k. This shows us where its current size-extrapolation limits lie in the zero-shot setting, when trained on system sizes up to 64 qubits. Table 1: Zero-shot and fine-tuned large-system evaluation. Energies are per qubit in the model’s coupling units, except the P rows, which are particle–hole-centred energies in eV per carbon with the number of carbons N/2N/2. The square lattices use open boundary conditions and the triangular lattice periodic boundary conditions. MaxCut comparisons are Efeas/N=(W/2−Cref)/NE_feas/N=(W/2-C_ref)/N from the archived feasible cuts CrefC_ref, so they are feasible reference energies not true ground states. Reference energy values are also in the table. P and lattice comparisons are thermodynamic-limit literature scales e∞e_∞ rather than exact finite-size energies [108, 109, 114]. N/A marks systems with no fine-tuning run. Recovery is the fraction of the comparison value reached; marked with an asterisk ∗ are where no fine-tuning happened. System N Zero-shot E/NE/N Fine-tuned E/NE/N Comparison Recovery MaxCut–12 256 −0.321068±0.000594-0.321068± 0.000594 −1.636595-1.636595 (20k) −1.943359-1.943359 (Cref=2456)(C_ref=2456) 96.8%96.8\% MaxCut–12 512 +0.010246±0.001845+0.010246± 0.001845 −2.250606-2.250606 (20k) −2.785156-2.785156 (Cref=9275)(C_ref=9275) 97.1%97.1\% MaxCut–12 1024 +0.539926±0.003347+0.539926± 0.003347 −2.516722-2.516722 (20k) −3.947754-3.947754 (Cref=35469)(C_ref=35469) 95.9%95.9\% P–Ohno 256 −3.187483±0.004127-3.187483± 0.004127 −4.951437-4.951437 (10k) −5.015(20)-5.015(20) 98.7%98.7\% P–Ohno 512 −3.013334±0.002460-3.013334± 0.002460 −4.935446-4.935446 (10k) −5.015(20)-5.015(20) 98.4%98.4\% P–Ohno 1024 −2.725309±0.003992-2.725309± 0.003992 −4.657041-4.657041 (10k) −5.015(20)-5.015(20) 92.9%92.9\% Square J1J_1–J2J_2 484 −0.446661-0.446661 −0.462856-0.462856 (100k) ≃−0.4968 -0.4968 93.2%93.2\% Square J1J_1–J2J_2 1024 −0.445674-0.445674 −0.451456-0.451456 (30k) ≃−0.4968 -0.4968 90.9%90.9\% Square J1J_1–J2J_2 2025 −0.421763±0.000046-0.421763± 0.000046 N/A ≃−0.4968 -0.4968 84.9%∗84.9\%^* Square J1J_1–J2J_2 4096 −0.350009±0.000048-0.350009± 0.000048 N/A ≃−0.4968 -0.4968 70.5%∗70.5\%^* Square J1J_1–J2J_2 8100 +0.128196±0.000101+0.128196± 0.000101 N/A ≃−0.4968 -0.4968 −25.8%∗-25.8\%^* Triangular J1J_1 1764 −0.398369±0.000050-0.398369± 0.000050 N/A −0.5503(8)-0.5503(8) 72.4%∗72.4\%^* We next restart from the corresponding zero-shot checkpoint, retain the route chosen by the pretrained router, and train only the compiled merge tree. As before, this reduces the optimised object from the 547M-parameter foundation model to approximately 4.7M parameters and permits each system to be trained on a single A100 GPU. The MaxCut and P runs use 256 walkers, eight replicas, two MCMC moves per KFAC step, and a burn-in of N steps (see SM § S4). Their learning rate is 0.002/(1+t/104)0.002/(1+t/10^4), with damping 10−310^-3 and curvature exponential moving average 0.990.99. The ×2222\!×\!22 and ×3232\!×\!32 square-lattice runs use the same walker, replica and update counts, a 128-step burn-in and 0.01/(1+t/104)0.01/(1+t/10^4) with curvature moving average 0.90.9. Figure 11 shows the energy intensiveness in the site count of each problem: qubits for MaxCut, carbons for P and physical spins for J1J_1–J2J_2. The optimisation traces are strongly front-loaded, as we saw before on smaller systems. At the final logged steps, the N=256N=256 and 512512 MaxCut energies correspond to expected cut weights 2377.52377.5 and 8917.18917.1, or 96.8%96.8\% and 96.1%96.1\% of the feasible references. The P chains reach −4.9514-4.9514 and −4.9359-4.9359 eV/C after 10410^4 steps, within 1.3%1.3\% and 1.6%1.6\% of the infinite-chain scale under the centered convention. For square-lattice J1J_1–J2J_2, the ×2222\!×\!22 case has 484 physical spins in the padded N=512N=512 routing bucket and improves from E/N=−0.44666E/N=-0.44666 at initialisation to −0.46316-0.46316 after 10510^5 steps. The 32×3232× 32 case has N=1024N=1024 physical spins and reaches E/N=−0.451456E/N=-0.451456 after 3×1043× 10^4 steps, essentially its initial value of −0.44567-0.44567 after an early transient. Figure 11: Fine-tuning on the large-system case studies. The faint curves are decimated per-step training estimates and the solid curves are non-overlapping block means (250 steps for MaxCut and P; 500 for J1J_1–J2J_2). Open markers show representative block means. Diamonds at step zero are the independent Axis-A zero-shot evaluations in (a,b) and the logged training initialisations in (c). The dashed grey lines in (b,c) are literature scales, not optimisation targets. (a) MaxCut energy per qubit; more negative energy corresponds to a larger expected cut through Eq. (5.2). (b) Particle–hole-centred P energy per carbon. (c) Textbook J1J_1–J2J_2 energy per physical spin: ×2222\!×\!22 means 484 physical spins in the padded N=512N=512 routing bucket, while ×3232\!×\!32 means N=1024N=1024 physical spins. 5.4 Witnessing a Phase Transition with Zero-Shot Weights Consider the periodic transverse-field Ising Hamiltonian H(g)=−J∑i=1Nσizσi+1z−gJ∑i=1Nσix,σN+1≡σ1.H(g)=-J _i=1^N _i^z _i+1^z-gJ _i=1^N _i^x, _N+1≡ _1. (5.5) Here, g is the dimensionless transverse-field strength relative to the nearest-neighbour Ising coupling J. Varying this parameter drives the model through its thermodynamic quantum critical point at gc=1g_c=1. For the normalized ground state |ψ(g)⟩ ψ(g) of H(g)H(g), the fidelity between two nearby Hamiltonians is F(g,g+δg)=|⟨ψ(g)∣ψ(g+δg)⟩|.F(g,g+δ g)= | ψ(g) ψ(g+δ g) |. (5.6) The leading response defines the fidelity susceptibility, χF(g+δg2)≃−2lnF(g,g+δg)(δg)2. _F\! (g+ δ g2 ) - 2 F(g,g+δ g)(δ g)^2. (5.7) Near the transition, the ground-state wavefunction changes rapidly with g, decreasing the fidelity and producing a peak in χF _F. The finite-size quantity χF/N _F/N therefore provides a practical witness of the transition. In Fig. 12, we see Hamilton-Zero witnesses this phase-transition with zero-shot inference. This size-consistent maximum near gc=1g_c=1 provides a proof of principle that foundation models could in future serve as search engines for phase transitions, see Sec. 6 for further discussion. Figure 12: Fidelity susceptibility per site, χF/N _F/N, of the periodic transverse-field Ising chain is shown for N=8N=8, 1616, and 3232 spins. Every curve is evaluated directly from the same pretrained Hamilton-Zero checkpoint without fine-tuning; markers denote adjacent-field overlap estimates at midpoint g+Δg/2g+ g/2 with Δg=0.05 g=0.05, and shaded regions give 95% confidence intervals. The vertical dashed line marks the thermodynamic phase-transition point, gc=1g_c=1. The zero-shot response develops a finite-size maximum close to gcg_c for all three sizes, providing a phase-transition witness from the fixed pretrained wavefunction. 5.5 Approximate Equivariance We tested the invariance of the predicted mean energy under a group action on the complete physical triplet (J,h,q)(J,h,q) corresponding to the interaction graph (J,h)(J,h) and quaternionic coordinates q. Recall for site-dependent rotations ui∈SU(2)u_i (2), the action is Jij↦R(ui)JijR(uj),hi↦R(ui)hi,qi↦qiu¯i,J_ij R(u_i)J_ijR(u_j) T, h_i R(u_i)h_i, q_i q_i u_i, (5.8) where R(ui)∈SO(3)R(u_i) (3) is the adjoint rotation and u¯i=ui−1 u_i=u_i^-1. The global test is the diagonal subgroup obtained by setting ui=u_i=u at every site; it is therefore not a rotation of the fields alone. The site-local test instead draws the uiu_i independently, inducing independent left and right SU(2)SU(2) actions at the two ends of each bond. For each transformed system, we measured the relative mean-energy residual rE=|E¯g−E¯||E¯|,r_E= | E_g- E | | E |, (5.9) using paired baseline and group-transformed samples. To distinguish symmetry violation from Monte Carlo noise, we also estimated a paired-sample standard error. For walker w, let ΔE¯w=1T∑t=1T(Ew,t(g)−Ew,t). E_w= 1T _t=1^T (E^(g)_w,t-E_w,t ). (5.10) The relative standard error plotted in Fig. 13 is SErel=sdw(ΔE¯w)W|E¯|,SE_rel= sd_w\! ( E_w ) W\,| E|, (5.11) with W=256W=256 walkers and T=128T=128 steps. This quantity estimates the Monte Carlo uncertainty of each paired residual. At initialization, the median residual was on the order of unity for both actions: 1.301.30 for the global action and 1.391.39 for the site-local action. By step 70 00070\,000, these values had fallen to 1.41×10−21.41× 10^-2 and 3.28×10−23.28× 10^-2, respectively. Thus, most of the equivariant behaviour was acquired during the first quarter of training. Subsequent checkpoints show a non-monotonic plateau rather than a steady improvement: over steps 70 00070\,000–280 000280\,000, the global median lies between 1.14×10−21.14× 10^-2 and 2.01×10−22.01× 10^-2, whereas the site-local median lies between 2.57×10−22.57× 10^-2 and 3.95×10−23.95× 10^-2. At step 280 000280\,000, the median residuals are 1.14×10−21.14× 10^-2 for the global action and 2.91×10−22.91× 10^-2 for the site-local action. The corresponding median paired-sample standard errors are 2.12×10−32.12× 10^-3 and 6.53×10−36.53× 10^-3. The reduction from order-one residuals at initialization to percent-level residuals after training demonstrates that the model has learned approximate equivariance. The consistently larger site-local residual shows that the full SU(2)NSU(2)^N action is learned less accurately than its global diagonal subgroup. Figure 13: Learning of mean-energy equivariance under joint transformations of (J,h,q)(J,h,q) on 20 held-out systems from the (B) evaluation set (see SM § S3). Blue denotes the global diagonal SU(2)SU(2) action and orange the site-local SU(2)NSU(2)^N action. Solid curves and shaded bands show the median and interquartile range of the relative energy residual, respectively; dashed curves show the median paired-sample Monte Carlo standard error estimated from 256 walkerwise averages of 128 steps. 5.6 Router Commitment and Generalisation Recall from Sec. 3 (see also SM § S2) that the router converts the Hamiltonian-conditioned trunk embeddings into a permutation of the physical sites and padded leaves, thereby selecting a binary contraction tree used by the wavefunction. In this section, we visualise the map that the router picks, showing that it recovers the entanglement structure of the case studies of 5.3. For visualisation, we map the saved route from its padded, bit-reversed compiler labels back to canonical physical coordinates. The resulting diagrams therefore show the contraction actually selected by the model. At zero-shot evaluation, the router decodes eight candidate trees. Each candidate is subjected to a short burn-in and measurement contest, after which the winning ordering is compiled and used for the full evaluation. Owing to the fact that our algorithm selects from a contest, we can quantify the concentration of the route policy by its commitment, Croute:=1−HpolicyHmax,C_route:=1- H_policyH_ , (5.12) where Croute≃0C_route 0 denotes a diffuse policy and Croute≃1C_route 1 denotes concentration on a small set of preferred, potentially symmetry-related routes. Figure 14(a) shows the median commitment across the 13 sufficiently populated system categories of the pretraining dataset (SM § S3), with the grey band giving the interquartile range, while Fig. 14(b) resolves the individual category trajectories. Commitment rises from approximately 0.50.5 to at least 0.960.96 during the first eight augmentation rounds and remains close to saturation thereafter. The agreement across categories is indicative of the router learning a stable structural preference early in pretraining, rather than acquiring a different arbitrary ordering for each family late in the run. It is thus able to train effectively over a heterogeneous set of Hamiltonians in the pretraining set. Figure 14: Entropy-realized commitment, 1−Hpolicy/Hmax1-H_policy/H_ , by augmentation round for the 13 system categories with S≥80S≥ 80. (a) Round-wise median across categories, with the grey band denoting the interquartile range. (b) Stratification by system-category, see Fig. S13 of SM § S3. All categories commit rapidly during the first ∼8 8 rounds and subsequently remain near saturation. Indeed, figure 15 shows there is a direct correlation between commitment and the energy quality of Hamilton-Zero’s predictions; a higher level of commitment generally signals a lower gap to ED. Panel (a) follows the round-wise median in the plane of router commitment and relative gap to exact diagonalisation for systems in the pretraining set where ED data is available. Panel (b) shows the corresponding trajectories for the individual system categories, showing this trend is consistent between different types of system. Commitment is therefore a correlational diagnostic rather than an energy certificate: a concentrated policy may still select a poor ordering. Figure 15: Router commitment versus relative gap to ED over augmentation rounds 0→550→ 55 for the 13 system categories of our pretraining dataset (SM § S3). (a) The round-wise median trajectory; grey shading denotes the interquartile range in the commitment and relative-gap coordinates. (b) System-category trajectories, with the enlarged marker indicating the final round. The router’s selected routes also generalise to system sizes and topologies beyond those seen in training. On the 10×1010× 10 J1J_1–J2J_2 torus (Fig. 16), the first merge pairs nearest-neighbour spins and the second produces exact 2×22× 2 plaquettes for every fully occupied four-site cell. Larger cells then remain spatially local as they are combined towards the root. No lattice geometry is supplied to the contraction mechanism, this plaquette hierarchy is inferred from the Hamiltonian-conditioned representations and selected because it supports a lower variational energy. This is supporting evidence that the router is able to recover the locally entangling geometry of the J1J2 model, at a scale above its pretraining. More strikingly, the same behaviour survives far beyond the pretraining sizes. The 45×4545× 45 square lattice in Fig. 17 contains 2,025 physical sites, more than thirty times the largest physical size used in pretraining. Nevertheless, the zero-shot route first forms local cells and subsequently organises them into extended diagonal bands consistent with the frustrated J1J_1–J2J_2 interaction structure. At Step 6, 75.0%75.0\% of nearest-neighbour bonds lie within the same 64-leaf cell, compared with 3.1%3.1\% under random ordering; at the final root split, 91.4%91.4\% lie within one branch, compared with the random baseline of 50.0%50.0\%. The router has therefore preserved interaction locality across eleven merge levels without having encountered comparable system sizes during training. Figure 16: Learned coarse-graining of the 10×1010× 10 J1J_1–J2J_2 torus. The panels are read along the arrows, with successive rows traversed in opposite directions. Step 0 shows the 100 physical lattice sites. At Step k, the black boundaries enclose the effective cells remaining after k levels of binary merging; a boundary disappears when the two cells on either side are merged. Colours are reassigned independently at each step solely to distinguish adjacent cells and therefore do not track cell identity between panels. For J2/J1=0.3J_2/J_1=0.3, the first merge pairs only nearest-neighbour sites. At Step 2, all 24 fully occupied four-site cells are exact 2×22× 2 plaquettes. The few incompletely occupied cells arise because the 100 physical sites are embedded in a 128-leaf binary tree, with 28 virtual leaves. Subsequent levels combine the plaquettes into progressively larger spatially local regions, culminating in two root branches at Step 6 and a single scalar at Step 7 of the merge-tree. Figure 17: Zero-shot coarse-graining of the 45×4545× 45 square J1J_1–J2J_2 lattice. The panels follow the arrows from the 2,025 physical sites at Step 0 to the scalar output at Step 11. At each step, black boundaries delimit the current effective cells and disappear when those cells are merged. Colours are reassigned independently within each panel only to separate adjacent cells visually. At J2/J1=0.5J_2/J_1=0.5, the early merges are predominantly local. The larger effective cells subsequently organize into extended diagonal bands, showing that the learned hierarchy preserves spatial structure far beyond the training sizes. At Step 6, 75.0%75.0\% of nearest-neighbour bonds lie within the same 64-leaf cell, compared with a random-order mean of 3.1%3.1\%. At the root split in Step 10, 91.4%91.4\% lie within one branch, compared with 50.0%50.0\% for a random ordering. Some cells contain fewer than 2k2^k physical sites because the lattice is embedded in a 2,048-leaf tree containing 23 virtual leaves. The P–Ohno case of 5.3 provides a contrasting route. Fig. 18 shows the router first pairs the two spin orbitals belonging to each carbon, then combines neighbouring carbons into four-orbital cells, and thereafter doubles the length of the contiguous carbon blocks at each level. It recovers the chemically natural nested factorisation of a long-range interacting fermionic chain despite operating on qubit-labelled Hamiltonian data [108, 109, 114]. This shows the router demonstrates transfer across interaction type as well as scale. Together, these examples show that the router has not merely memorised a site ordering or learned a generic preference for balanced trees. It extracts a topology-dependent, multiscale contraction hierarchy from the Hamiltonian representations and transfers that hierarchy across periodic and open lattices, long-range chains, interaction families, and system sizes. The result is evidence that pretraining has produced a reusable structural prior for organising entanglement, one that can be converted directly into the compiled contraction path used for inference. Figure 18: Learned coarse-graining of the 128-carbon P–Ohno chain. The 256 spin orbitals are folded into a serpentine arrangement for visualization only, so that consecutive carbons remain adjacent when the chain passes between rows. The two spin orbitals belonging to each carbon are placed vertically. Panels are read along the arrows. Black boundaries enclose the effective cells present at each merge level and disappear when cells are combined; colours are reassigned independently at every step and serve only to distinguish neighbouring cells. At Step 1, every effective cell contains the two spin orbitals of one carbon. Step 2 combines adjacent carbons into four-orbital cells, and every subsequent level doubles the length of the contiguous carbon block. The hierarchy therefore proceeds through blocks of 1,2,4,8,…,641,2,4,8,…,64 carbons before the final scalar merge at Step 8. This exact nested structure provides direct evidence that the router has learned the interaction-local factorisation of the P Hamiltonian. 6 Discussion and outlook Quantum computing has attracted great interest in both academic and industrial settings on the premise that classical cost will continue rising until useful quantum many-body calculations have to move onto quantum hardware. Hamilton-Zero takes the other side of this premise for ground state problems specifically. As a foundation model, pay once for learning ground states across a distribution of Hamiltonians and store what the model learns in its weights. As a result, we have a classical model that improves with pretraining, transfers across interaction graphs and system sizes, and returns an explicit, differentiable wavefunction on hardware that exists today. We believe this changes the economics of quantum ground-state computation. Most classical ground-state solvers restart, or substantially re-optimise, for each new Hamiltonian. Their cost scales with the number of problems solved. Hamilton-Zero replaces repeated training from scratch with one pretraining run whose cost is shared across every subsequent use of the checkpoint. If we let M denote the number of target Hamiltonians evaluated using the same pretrained checkpoint, and let R denote the number of additional fine-tuning runs performed across those targets. So R=0R=0 for a completely zero-shot workload, R=MR=M if every target is fine-tuned separately, and 0<R<M0<R<M if one fine-tuned checkpoint is reused across several related Hamiltonians. Then we denote CpreC_pre for the dollar cost of producing the pretrained checkpoint, Ceval(m)C_eval^(m) for the dollar cost of applying a fixed checkpoint to target Hamiltonian m, including inference, MCMC sampling and estimation of the reported observables, and Cft(r)C_ft^(r) for the dollar cost of fine-tuning run r. The average cost per target Hamiltonian is then CHZ(M)M=CpreM+1M∑m=1MCeval(m)+1M∑r=1RCft(r). C_HZ(M)M= C_preM+ 1M _m=1^MC_eval^(m)+ 1M _r=1^RC_ft^(r). (6.1) The first term falls as 1/M1/M, while the evaluation term remains the marginal cost of using the model. Meanwhile the fine-tuning term depends on how often fine-tuning is required and how widely each fine-tuned checkpoint can subsequently be reused. In the zero-shot setting the final sum vanishes. At the measured throughput of our inference stack, drawing one sample from Hamilton-Zero costs below 10−610^-6 USD. Public quantum hardware is already priced in the same unit, with services like Amazon Braket listing on-demand QPU prices from 4.25×10−44.25× 10^-4 to 8.0×10−28.0× 10^-2 USD per shot in addition to a 0.300.30 USD charge for every submitted job [4]. Even if we discard the job charge entirely, one quantum shot therefore costs between 425425 and 80,00080,000 times one Hamilton-Zero sample at current public prices. Amazon’s own error-mitigation example prices an IonQ task of 2,5002,500 shots at 200.30200.30 USD [4]; the same number of Hamilton-Zero samples costs approximately 0.00250.0025 USD. This is still the comparison most favourable to quantum hardware, one shot returns one measurement outcome, whereas estimating a ground-state energy requires repeated state preparation and measurements across many operator groups. Indeed, when we compare this more broadly with the larger research and engineering costs or the continued development and improvement of quantum hardware, the difference is also stark. Improving Hamilton-Zero requires more pretraining compute, which distributes immediately with the release of new model weights, whereas hardware expenditure depreciates. Furthermore, our scaling laws show Hamilton-Zero’s compute costs scale around 5 times more favourably than today’s LLMs [69], making this a relatively resource-efficient compute demand by today’s standards. For reference, the checkpoint we release in this work, including its previous iterations, cost around $30k\$30k to pretrain from scratch. Evaluation and fine-tuning cost another $30k\$30k using on-demand resources. Since a foundation model can improve, absorb new problem classes, and serve more users from the same weights, the economic question of useful quantum advantage is now whether a quantum device can produce a ground-state estimate at lower total cost than a pretrained classical model whose marginal cost will continue to fall with every subsequent release of such models. Hence, with Hamilton-Zero, we believe the baseline of useful quantum advantage in ground state computation has moved. Comparisons against a classical model trained from scratch establish only an advantage over cold-start optimization, but should, in fact, compare against amortized computation now that a model like Hamilton-Zero is known to be workable. This is especially crucial in the knowledge that to date, no quantum computer has solved quantum chemistry problems beyond a few dozens of qubits at a time for example, yet Hamilton-Zero is able to compile wave-functions to the 8000 qubits scale. From this point, a ground-state quantum-advantage claim should report the same Hamiltonians, the same observables, the same target error and the same confidence level, while counting state preparation, sampling, mitigation, classical outer-loop optimisation and verification. It should then compare total wall time, dollar-cost and energy-cost against both zero-shot Hamilton-Zero and Hamilton-Zero fine-tuned at the same budget if such a claim is to hold up against this model. Until a quantum device wins that comparison, it beats only a now-obsolete classical baseline. We also believe in future Hamilton-Zero could change how we can search for new physics through phase diagrams. The fact that the fixed weights could witness a phase transition in the Ising model is promising evidence that this is a realistic goal for such foundation models. Indeed, a conventional phase-diagram study solves a sparse collection of coupling points in the Hamiltonian’s parameter-space and spends most of its compute rediscovering nearby ground states. Whereas Hamilton-Zero swept the coupling space of the Ising model in zero-shot inference. This avenue of identifying sharp changes in fidelity, correlations, or order parameters, can let us concentrate fine-tuning where the phase structure changes in future. It could also be used to supply synthetic samples for moment-matrices that underpin recent advances in semi-definite programming techniques for mapping phase diagrams [64]. As such in this role the model nominates regions for exact methods or experiment. Hamilton-Zero is also not without limitations however. First, the cost of compute is poly(2ceil(log2N))poly(2^ceil(log_2N)) due to padding that aims to construct a perfectly balanced tree. So, a 129-qubit system costs the same as a 256-qubit one, as they share the next power of 2 available. Second, the automatic differentiation primitives to evaluate energies have an order of differentiation that scales with the weight of the Pauli string present in the Hamiltonian. As such, we restricted ourselves to quadratic Hamiltonians. However, we do note that such a set is universal in the sense of quantum computation [102, 3] as it corresponds to two-qubit gates. Hence, higher-order terms can be compiled onto a system of physical and ancilla qubits at quadratic order. Finally, even though QMC is variational by design, finite-sample Monte-Carlo estimate may violate the variational inequality by either of two mechanisms: estimator noise and/or mixing bias, which is dependent on the spectral gap [81] and how long the burn-in is [14] , which can always be alleviated with larger effective sample size and longer burn-in. We report confidence intervals, which also do not guarantee that its upper bound would be above the ground state energy, as confidence intervals only have a guarantee to contain true value within repeated sampling. To empirically minimize bias effects, we conducted 3 different stationarity tests on energy values to detect distribution drifts during estimation, and chose burn-in duration empirically based on this; see SM § S4. Three extensions follow naturally from this work. First, we can create versions of Hamilton-zero that were fine-tuned on larger datasets to boost its performance in sector-specific problem families such as quantum chemistry or combinatorial optimisation. This would likely further amortize subsequent fine-tuning of individual systems and lead to even stronger zero-shot inference. Second, we can extend Hamilton-Zero’s capabilities from static ground state problems to dynamics via a time encoding. For a pretraining dataset of time-dependent Hamiltonians H(t)H(t), we could train against the Schrödinger residual i∂tψθ−H(t)ψθi _t _θ-H(t) _θ together with an initial-condition loss, as was done in deep stochastic mechanics of bosonic systems [106]. Finally, a per-site-odd but nonlinear successor could use higher half-integer Peter–Weyl sectors. Unlike the multilinear architecture used here, it would require the certified Casimir penalty of Appendix A to retain a physical upper bound. We leave both constructions and their validation to future work. Acknowledgements We thank the NVIDIA Inception Program for hardware support and technical engagement, the Google for Startups Program for cloud infrastructure support, and Nebius for providing the GPU compute that supported this work’s training runs. We would also like to thank Jose Ramon Martinez for his help proof-reading the manuscript. Supplementary Material In the supplementary material, we describe the theory, architecture, datasets, optimisation, and engineering that underpin Hamilton-Zero. S1 covers variational principles on spinor manifolds, with a full exposition on the theory behind our automatic-differentiation primitives of Lie derivatives, the Peter-Weyl theorem that shows why our model has support on a larger manifold than NQS, as well as how our variational principle manages the case of non-linear quaternionic functions through a Casimir-operator-based regulariser and energy penalty that preserves the spin-1/21/2 variational principle. Next S2 describes Hamilton-Zero’s architecture, drawing parallels and design choices from large-scale pretraining runs of large-language models. Here we also describe the router we use to learn entanglement structure through deep reinforcement learning, characterise its automorphism-orbit entropy, as well as the tree tensor-network-style contraction mechanism that is conditioned on the model’s embeddings of interaction data. Next, we describe in full detail the pretraining routines and datasets we used in S3. We then describe in S4 our state-of-the-art sampling scheme for MCMC on the configuration manifold SU(2)NSU(2)^N, our energy kernels, and our extension of the KFAC optimiser, all of which we made shard-mappable and GEMM-optimised for training on Nvidia H200 GPUs. S1 Variational principles on Spinor Manifolds S1.1 From basis amplitudes to wavefunctions on SU(2)NSU(2)^N The standard spin-12 12 wavefunction is a complex-valued function on +,−N\+,-\^N, equivalently a vector Ψ∈ℂ2N ^2^N with one amplitude per computational-basis configuration. Two features of this picture constrain every method built on it. The dimension is exponential in N, and the domain is discrete. Thus a Hamiltonian acts through explicit sums over connected configurations, enumerated anew for each interaction structure. Tensor networks and NQS address the first constraint by replacing the 2N2^N-dimensional space with a lower-dimensional variational sub-manifold. The second constraint they inherit unchanged however, and it is what ties each trained model to one Hamiltonian’s connectivity. The central objective of this section is to remove both at once, by replacing the discrete domain with the continuous group manifold SU(2)NSU(2)^N, on which Hamiltonians act by differentiation. Tensor networks and neural quantum states replace the exponentially large Hilbert space by a lower-dimensional variational family [14, 22], but the resulting optimisation can still be limited by stability and trainability. On the quantum-circuit side, barren plateaus provide a distinct failure mode in which gradient variance can vanish exponentially with system size [93]. The central objective of this section is to expose the methods that allow us to sidestep both the curse of dimensionality and the curse of linearity, by replacing the discrete domain of the spin-1/21/2 basis vectors with a continuous Lie group domain. The wavefunctions we employ, ψ:(SU(2))N→ℂ,ψ: (SU(2) )^N , (S1.1) are smooth functions on the Lie group manifold. Here, each spin is parametrised by a unit quaternion qi∈S3q_i∈ S^3 via the standard identification SU(2)≅S3SU(2) S^3 embedded in ℝ4R^4. Hence any neural network model in our framework thus consumes 4N4N continuous inputs for N bodies which is linear in N. As usual with representation theory, there is no free lunch. In the following, we will see the signature of the Hilbert-space curse of dimensionality in this picture, and explain how to side-step it by relaxing into a larger domain, L2(SU(2)N)L^2(SU(2)^N), than the spin-12 12 sector, while maintaining the variational principle, so that we can find ground-state energies of many-body systems. First, though, let us build some intuition for the quaternionic representation, so that examples later in this part can be written directly in (q0,q1,q2,q3)(q_0,q_1,q_2,q_3). Recall a unit quaternion q=(q0,q1,q2,q3)∈S3q=(q_0,q_1,q_2,q_3)∈ S^3 (q02+q12+q22+q32=1q_0^2+q_1^2+q_2^2+q_3^2=1) is the standard coordinate on SU(2)SU(2) via g(q)=q0I+qa(iσa)=(q0+iq3q2+iq1iq1−q2q0−iq3),g(q)∈SU(2),g(q)\;=\;q_0\,I+q_a\,(iσ^a)\;=\; pmatrixq_0+i\,q_3&q_2+i\,q_1\\[4.0pt] i\,q_1-q_2&q_0-i\,q_3 pmatrix, g(q) (2), (S1.2) where detg(q)=q02+q12+q22+q32=1 g(q)=q_0^2+q_1^2+q_2^2+q_3^2=1, and g(q)g(q)†=Ig(q)\,g(q) =I. The four entries of g(q)g(q) are the spin-12 12 Wigner D-matrix elements. That is, for any spin j, Dmnj(q)D^j_mn(q) is the matrix element of the rotation representation on the spin-j irrep. For j=12j= 12 with row/column indices m,n∈0,1m,n∈\0,1\, Dmn1/2(q)=gmn(q),D001/2(q)=q0+iq3,D011/2(q)=q2+iq1,D101/2(q)=iq1−q2,D111/2(q)=q0−iq3.D^1/2_mn(q)\;=\;g_mn(q), array[]lD^1/2_00(q)=q_0+i\,q_3,\\[2.0pt] D^1/2_01(q)=q_2+i\,q_1,\\[2.0pt] D^1/2_10(q)=i\,q_1-q_2,\\[2.0pt] D^1/2_11(q)=q_0-i\,q_3. array (S1.3) To identify physical single-site spin states with SU(2)SU(2)-functions, we fix the row m=0m=0, because the column index carries the physical spin. Thus |↑⟩⟷D001/2(q)=q0+iq3,|↓⟩⟷D011/2(q)=q2+iq1.| D^1/2_00(q)=q_0+iq_3, | D^1/2_01(q)=q_2+iq_1. (S1.4) A general single-site state α|↑⟩+β|↓⟩α|\! +β|\! becomes ψ(q)=α(q0+iq3)+β(q2+iq1).ψ(q)=α(q_0+iq_3)+β(q_2+iq_1). (S1.5) For N sites, let sk∈0,1s_k∈\0,1\ with 0≡↑0≡ and 1≡↓1≡ . Then |s1⋯sN⟩⟷∏k=1ND0,sk1/2(qk),|s_1·s s_N _k=1^ND^1/2_0,s_k(q_k), (S1.6) and a general spin-12 12 state is a linear combination of these products; see SM § S1.3. Recall a variational ansatz Ψθ _θ over the physical Hilbert space (ℂ2)⊗N(C^2) N gives an upper bound ⟨Ψθ|H|Ψθ⟩⟨Ψθ|Ψθ⟩≥E0 _θ|\,H\,| _θ _θ|\, _θ \;\;≥\;\;E_0 (S1.7) for any normalised Ψθ _θ. Minimising this over θ converges to (or above) E0E_0. This is the variational principle in its standard form, the rest of SM § S1 is about preserving this inequality under our non-standard quaternionic parameterisation. We instantiate the manifold amplitude as a neural network returning logψθ:SU(2)N⟶ℂ. _θ:SU(2)^N . (S1.8) For the multilinear architecture, its Peter–Weyl coefficients carry a row-index multiplicity and a column-index physical-spin factor. Fixing a row multiplicity vector recovers the state-vector lift described above; a general model state is interpreted through the decomposition of Sec. S1.3. We defer the details of the architecture to SM § S2, but the signature is a smooth non-linear function on the manifold, to be optimised by gradient descent on a variational loss. As is standard practise in Quantum Variational Monte Carlo, we use logψθ _θ to prevent underflow, and to allow for stable natural gradient flow [14]. Having now fixed our model to be a smooth 4N4N-input, 11-output complex-valued network, predicting logψθ _θ on the Lie-group manifold, we can now turn to the question of how the Hamiltonian acts on this object, and how to project this object into the physical spin-12 12 sector of this function family. S1.2 Operators with automatic differentiation In this subsection we explain how the shift away from Hilbert space lets us realise the action of spin operators as Lie derivatives, executable directly through automatic differentiation. We restrict attention to quadratic, that is two-body, Hamiltonians throughout, since two-body interactions are universal for quantum [28], and higher-order interactions can be built from them [105, 70]. Recall from the standard theory that at each site i, the spin is carried by three Hermitian operators S^ix,S^iy,S^iz S_i^x, S_i^y, S_i^z, closing the algebra [S^ia,S^ib]=iεabcS^ic[ S_i^a, S_i^b]=i\, _abc\, S_i^c. Their squared magnitude is the Casimir S^i2:=(S^ix)2+(S^iy)2+(S^iz)2 S_i^2:=( S_i^x)^2+( S_i^y)^2+( S_i^z)^2, which on a single spin-12 12 is the scalar 34 34, and we write σ^ia:=2S^ia σ_i^a:=2\, S_i^a for the associated Pauli matrices. The most general quadratic Hamiltonian we target couples these Paulis pairwise, together with an on-site field, H^=14∑i<j∑a,bJijabσ^iaσ^jb+12∑i,ahiaσ^ia, H\;=\; 14\! _i<j _a,bJ_ij^ab\, σ_i^a\, σ_j^b\;+\; 12\! _i,ah_i^a\, σ_i^a, (S1.9) with exchange tensor J∈ℝN×N×3×3J ^N× N× 3× 3 and field h∈ℝN×3h ^N× 3. As stated in the main text, it is useful to think of J as the data of a weighted interaction graph on N sites, where each unordered pair (i,j)(i,j) with i≠ji≠ j carries a 3×33× 3 Pauli-pair coupling matrix Jij∈ℝ3×3J_ij ^3× 3. To move this operator onto the manifold, we use on the iith copy of SU(2)SU(2) the left-invariant vector fields (Liaf)(qi):=dtf(qiexp(tXa))|t=0,(L_i^af)(q_i):= . ddtf\! (q_i (tX_a) ) |_t=0, (S1.10) the anti-Hermitian infinitesimal generators of right translations. For Dmnj(q)=⟨j,m|Dj(q)|j,n⟩D^j_mn(q)= j,m|D^j(q)|j,n , they act on the column index: LiaDmnj(qi)=∑rDmrj(qi)(dDj(Xa))rn.L_i^aD^j_mn(q_i)= _rD^j_mr(q_i)\, (dD^j(X_a) )_rn. (S1.11) We therefore use the column index n as the physical spin index and the row index m as the Peter–Weyl multiplicity. The fields close [Lia,Lib]=κεabcLic[L_i^a,L_i^b]=κ\, _abcL_i^c, with κ=−2κ=-2 in the unit-length (half-Lie) normalisation we adopt. Setting S^ia=cLia S_i^a=cL_i^a and matching the two algebras forces cκ=icκ=i, hence c=i/κ=−i2c=i/κ=- i2 and σ^ia=2cLia=−iLia σ_i^a=2cL_i^a=-iL_i^a. Under this identification each Pauli pair becomes σ^iaσ^jb=−LiaLjb σ_i^a σ_j^b=-L_i^aL_j^b and each field term becomes σ^ia=−iLia σ_i^a=-i\,L_i^a, so the Hamiltonian (S1.9) turns into the second-order differential operator H^↔−14∑i<j,a,bJijabLiaLjb−i2∑i,ahiaLia H\; \;-\, 14\! _i<j,a,bJ_ij^ab\,L_i^aL_j^b\;-\; i2\! _i,ah_i^a\,L_i^a (S1.12) on L2(SU(2)N)L^2(SU(2)^N). The same identification sends the spin magnitude to a Laplacian: S^i2=c2∑a(Lia)2=−Δi S_i^2=c^2 _a(L_i^a)^2=- _i, where Δi:=14∑a(Lia)2 _i:= 14 _a(L_i^a)^2 is the Laplace–Beltrami operator on the i-th sphere. What makes (S1.12) executable is that each LiaL_i^a is a known linear combination of the four partial derivatives ∂/∂qi,μ∂/∂ q_i,μ at site i, Lia=∑μ=03Aμa(qi)∂qi,μ,L_i^a\;=\; _μ=0^3A^a_μ(q_i)\; ∂ q_i,μ, (S1.13) with coefficients AμaA^a_μ the left-invariant frame on S3≅SU(2)S^3 (2), obtained by right-multiplying qiq_i by the three imaginary quaternion units ı^,ȷ^,k , , k, [Aμx(q)]μ=0,1,2,3 [A^x_μ(q) ]_μ=0,1,2,3 =(−q3,q2,−q1,q0), =(-q_3,\;q_2,\;-q_1,\;q_0), (S1.14) [Aμy(q)] [A^y_μ(q) ] =(−q2,−q3,q0,q1), =(-q_2,\;-q_3,\;q_0,\;q_1), [Aμz(q)] [A^z_μ(q) ] =(−q1,q0,q3,−q2). =(-q_1,\;q_0,\;q_3,\;-q_2). A direct expansion using |q|2=1|q|^2=1 gives the identity we will lean on, ∑a∈x,y,zAμa(q)Aνa(q)=δμν−qμqν=[PTqS3]μν, _a∈\x,y,z\A^a_μ(q)\,A^a_ν(q)\;=\; _μν-q_μq_ν\;=\; [P_T_qS^3 ]_μν, (S1.15) so the three tangents Aa(q)A^a(q) are orthonormal and their outer product is the ambient projector onto TqS3T_qS^3. This is what makes La\L^a\ a genuine orthonormal frame of the round metric, and hence Δi _i the Laplace–Beltrami operator named above. When two Lie derivatives act on the same site the channel sum (S1.15) applies; across two distinct sites the two frames carry different arguments and are simply transported along. Equation (S1.13) makes the differential action computationally tractable. The custom forward-Laplacian propagates the value, the required first derivatives, and their contracted second derivatives together through one structured program traversal, never needing to materialise a dense Hessian. We can also avoid running an independent second-order pass for every pair. For a Hamiltonian with edge set ℰE, the pair contraction has leading combinatorial cost (|ℰ|)O(|E|), which is (N2)O(N^2) in the dense case. Here ℰE is the Hamiltonian’s interacting edge set. In the balanced readout, whose padded width is N^:=2⌈log2N⌉ N:=2 _2N , the cross-derivative work summed over levels is likewise (N^2)O( N^2) with model-width factors suppressed. . Applying the frame (S1.13) to the Lie form (S1.12) now gives the action of H H on ψ in quaternionic coordinates. In a pair term the outer derivative LiaL_i^a meets the frame Ab(qj)A^b(q_j) of the other site as a constant, so it contributes a pure second derivative with no first-order remainder, while the field term stays first order. Hence, for any JijabJ_ij^ab, H^ψ(q)=−14∑i<j,a,bJijab∑μ,ν=03Aμa(qi)Aνb(qj)∂2ψ∂qi,μ∂qj,ν(q)−i2∑i,ahia∑μAμa(qi)∂ψ∂qi,μ(q). H\,ψ(q)\;=\;-\, 14\! _i<j,a,b\!J_ij^ab\! _μ,ν=0^3\!A^a_μ(q_i)\,A^b_ν(q_j)\, ∂^2ψ∂ q_i,μ\,∂ q_j,ν(q)\;-\; i2\! _i,a\!h^a_i\! _μ\!A^a_μ(q_i)\, ∂ψ∂ q_i,μ(q). (S1.16) One step remains to make this usable. The network returns logψθ _θ, not ψθ _θ, whereas the local energy needs ratios LiaLjbψθ/ψθL_i^aL_j^b\, _θ/ _θ. Writing ψ=elogψ=e ψ and applying the two first-order operators in turn bridges the two: for any twice-differentiable ψ≠0ψ≠ 0, we have that, LiaLjbψ=LiaLjblogψ+(Lialogψ)(Ljblogψ). L_i^aL_j^b\,ψ\;=\;L_i^aL_j^b\, ψ\;+\; (L_i^a\, ψ )\, (L_j^b\, ψ ). (S1.17) On a single site, summing the diagonal channels means the spin magnitude S^l2=−Δl S_l^2=- _l reads, −Δlψ=−Δllogψ−14∑a(Llalogψ)2. - _l\,ψ\;=\;-\, _l\, ψ\;-\; 14 _a (L_l^a\, ψ )^2. (S1.18) Both (S1.17) and (S1.18) depend only on logψθ _θ and its first two derivatives at the sampled point. These same quantities determine local estimators for observables containing at most two spin operators, including all two-point correlation functions. In (S1.16), the Hamiltonian enters through the coefficients JijabJ_ij^ab and hiah_i^a that contract these derivatives. In particular, the nonzero pattern of J carries the interaction graph. The graph is therefore supplied as data to the same differential expression, rather than encoded in a lattice-specific rule. Changing the topology or system size changes the tensor entries and index ranges, and the form of the calculation remains identical aside from that. S1.2.1 Three examples from the many-body literature To make (S1.16) concrete, we read off three couplings that recur across the spin-model literature. The isotropic Heisenberg model H^XXX=J∑⟨ij⟩∑aS^iaS^ja H_X=J _ ij _a S_i^a S_j^a has Jijab=JδabJ_ij^ab=J\,δ^ab on each bond and no field; its Lie form is −J4∑⟨ij⟩,aLiaLja- J4 _ ij ,aL_i^aL_j^a, since S^iaS^ja=c2LiaLja=−14LiaLja S_i^a S_j^a=c^2L_i^aL_j^a=- 14L_i^aL_j^a, and (S1.16) collapses to a diagonal channel sum, H^XXXψ(q)=−J4∑⟨ij⟩∑a∈x,y,z∑μ,ν=03Aμa(qi)Aνa(qj)∂2ψ∂qi,μ∂qj,ν(q), H_X\,ψ(q)\;=\;-\, J4\! _ ij \! _a∈\x,y,z\ _μ,ν=0^3A^a_μ(q_i)\,A^a_ν(q_j)\, ∂^2ψ∂ q_i,μ\,∂ q_j,ν(q), (S1.19) three mixed-site Hessian sandwiches per bond, one per Pauli channel. The Kitaev-Γ coupling H^Γ=∑⟨ij⟩Γij(S^iyS^jz+S^izS^jy) H_ = _ ij _ij\,( S_i^y S_j^z+ S_i^z S_j^y) is symmetric off-diagonal, Jijyz=Jijzy=ΓijJ_ij^yz=J_ij^zy= _ij, and pairs the two cross-channels, H^Γψ(q)=−14∑⟨ij⟩Γij∑μ,ν=03[Aμy(qi)Aνz(qj)+Aμz(qi)Aνy(qj)]∂2ψ∂qi,μ∂qj,ν(q). H_ \,ψ(q)\;=\;-\, 14 _ ij _ij _μ,ν=0^3 [\,A^y_μ(q_i)\,A^z_ν(q_j)+A^z_μ(q_i)\,A^y_ν(q_j)\, ] ∂^2ψ∂ q_i,μ\,∂ q_j,ν(q). (S1.20) The Dzyaloshinskii–Moriya coupling H^DM=∑⟨ij⟩D→ij⋅(S→^i×S→^j) H_DM= _ ij D_ij·( S_i× S_j) is antisymmetric, Jijbc=∑aDijaεabcJ_ij^bc= _aD_ij^a\, ^abc, and collecting its three channels into a cross product (well defined because LiaL_i^a and LjbL_j^b commute for i≠ji≠ j) gives H^DMψ(q)=−14∑⟨ij⟩∑μ,ν=03D→ij⋅[A→μ(qi)×A→ν(qj)]∂2ψ∂qi,μ∂qj,ν(q), H_DM\,ψ(q)\;=\;-\, 14\! _ ij \! _μ,ν=0^3 D_ij· [\, A_μ(q_i)× A_ν(q_j)\, ]\, ∂^2ψ∂ q_i,μ\,∂ q_j,ν(q), (S1.21) with A→μ(q):=(Aμx,Aμy,Aμz)(q) A_μ(q):=(A^x_μ,A^y_μ,A^z_μ)(q) the frame sliced at fixed μ. Here the bracket is antisymmetric under the joint swap (i,μ)↔(j,ν)(i,μ) (j,ν) while the mixed Hessian is symmetric, so a Dzyaloshinskii–Moriya term contributes nothing when both legs sit on one site: there is no on-site DMI, exactly as the antisymmetry Jbc=−JcbJ^bc=-J^cb demands. S1.3 Peter–Weyl decomposition and the spin-1/2 sector Here we use the Peter–Weyl decomposition to identify the physical spin-12 12 sector inside L2(SU(2)N)L^2(SU(2)^N). We will show that the multilinear manifold ansatz occupies a strictly larger ambient function class than a fixed lift of any NQS family, while the physical Hilbert space itself remains 2N2^N-dimensional. The comparison concerns ambient function classes; finite-parameter expressivity depends on the chosen architecture. We then establish the variational principle on the target spin-12 12 sector. The Peter–Weyl theorem decomposes the single-site space into the matrix coefficients of the irreducible representations, L2(SU(2))≅⨁j=0,12,1,32,…Vj∗⊗Vj.L^2 (SU(2) )\; \; _j=0, 12,1, 32,…V_j^* V_j. (S1.22) Here Dmnj(q)=⟨j,m|Dj(q)|j,n⟩D^j_mn(q)= j,m|D^j(q)|j,n carries the row index m in the first, multiplicity factor Vj∗V_j^* and the column index n in the second, physical factor VjV_j. The coefficients Dmnjm,n=−j,…,j\D^j_mn\_m,n=-j,…,j form a complete orthogonal basis under the Haar measure, with ‖Dmnj‖L22=1/(2j+1)\|D^j_mn\|_L^2^2=1/(2j+1). The spin-12 12 coefficients Dmn1/2D^1/2_mn are the four entries of g(q)g(q) from § S1.1; the higher-j coefficients are the degree-2j2j monomials in those entries. On N sites the basis is the product Φ,,(q1,…,qN)=∏k=1NDmk,nkjk(qk), _ J, m, n(q_1,…,q_N)= _k=1^ND^j_k_m_k,n_k(q_k), (S1.23) labelled by =(j1,…,jN) J=(j_1,…,j_N), and the full space splits as L2(SU(2)N)≅⨁⨂k=1N(Vjk∗⊗Vjk).L^2(SU(2)^N)\; \; _ J _k=1^N (V_j_k^* V_j_k ). (S1.24) with Jk∈0,12,1,…J_k∈\0, 12,1,…\, and any ψ has a unique expansion, ψ=∑,,c,,Φ,,.ψ= _ J, m, nc_ J, m, n\, _ J, m, n. (S1.25) The per-site Casimir of § S1.2 acts diagonally on this basis. Each DmnjD^j_mn restricts to a degree-2j2j harmonic polynomial in the quaternion coordinates, on which −Δi- _i carries the eigenvalue j(j+1)j(j+1), −ΔiDmi,niji(qi)=ji(ji+1)Dmi,niji(qi),-\, _i\,D^j_i_m_i,n_i(q_i)\;=\;j_i(j_i+1)\,D^j_i_m_i,n_i(q_i), (S1.26) and in particular −ΔiDmn1/2=34Dmn1/2- _i\,D^1/2_mn= 34\,D^1/2_mn, the spin-12 12 value S^i2=34 S_i^2= 34 recorded in § S1.2. The consequence for the network is immediate. An unconstrained ψθ∈L2(SU(2)N) _θ∈ L^2(SU(2)^N) generically has c,,≠0c_ J, m, n≠ 0 for every multi-index J in the expansion, integer and half-integer alike. It is not a spin-12 12 wavefunction, nor even a single-spin-j wavefunction; its image lives across the full direct sum. Indeed, the physical content occupies a single slice of that sum. A wavefunction ψ describes N genuine spin-12 12 degrees of freedom, one per site, precisely when it is an eigenfunction of every per-site Casimir at the spin-12 12 value, −Δkψ=34ψfor every site k=1,…,N.-\, _k\,ψ\;=\; 34\,ψ every site k=1,…,N. (∗ ) Condition ( ∗ ‣ S1.3) selects the Peter–Weyl sector =(12,…,12) J=( 12,…, 12), and projecting a generic manifold function onto it is what turns it into a physical state. This leads us to the following proposition. Proposition S1.1 (The spin-12 12 sector). A function ψ∈L2(SU(2)N)ψ∈ L^2(SU(2)^N) satisfies ( ∗ ‣ S1.3) if and only if ψ(q1,…,qN)=∑,T,∏k=1NDmk,nk1/2(qk),T,∈ℂ.ψ(q_1,…,q_N)\;=\; _ m, nT_ m, n\, _k=1^ND^1/2_m_k,n_k(q_k), T_ m, n . (S1.27) The coefficient tensor T∈ℂ4NT ^4^N parameterises the entire sector. Fixing the fiducial row mk=0m_k=0 at every site fixes the Peter–Weyl multiplicity and leaves the 2N2^N coefficients indexed by n; these are the amplitudes of Ψ∈(ℂ2)⊗N ∈(C^2) N under the identification of § S1.1. The proof, a term-by-term application of the Casimir eigen-equation (S1.26), is given in Appendix B.1. To see how Prop. S1.1 works, it is instructive to write down some entangled states of interest to the quantum information community directly in the quaternionic coordinates of § S1.1. The Bell state |Φ+⟩=12(|↑⟩+|↓⟩)| ^+ = 1 2(|\! +|\! ) and singlet |Ψ−⟩=12(|↑↓⟩−|↓↑⟩)| ^- = 1 2(|\! -|\! ) read ψΦ+(q1,q2) _ ^+(q_1,q_2) =12[(q1,0+iq1,3)(q2,0+iq2,3)+(q1,2+iq1,1)(q2,2+iq2,1)], = 1 2 [(q_1,0+iq_1,3)(q_2,0+iq_2,3)+(q_1,2+iq_1,1)(q_2,2+iq_2,1) ], ψΨ−(q1,q2) _ ^-(q_1,q_2) =12[(q1,0+iq1,3)(q2,2+iq2,1)−(q1,2+iq1,1)(q2,0+iq2,3)]. = 1 2 [(q_1,0+iq_1,3)(q_2,2+iq_2,1)-(q_1,2+iq_1,1)(q_2,0+iq_2,3) ]. The three-site GHZ and W states, which read, |GHZ⟩ |GHZ =12(|↑⟩+|↓⟩), = 1 2(|\! +|\! ), |W⟩ |W =13(|↑↓⟩+|↑↓↑⟩+|↓↑⟩), = 1 3(|\! +|\! +|\! ), have manifold representatives, ψGHZ(q1,q2,q3) _GHZ(q_1,q_2,q_3) =12[∏k=13(qk,0+iqk,3)+∏k=13(qk,2+iqk,1)], = 1 2 [ _k=1^3(q_k,0+iq_k,3)+ _k=1^3(q_k,2+iq_k,1) ], ψW(q1,q2,q3) _W(q_1,q_2,q_3) =13[(q1,0+iq1,3)(q2,0+iq2,3)(q3,2+iq3,1)+(q1,0+iq1,3)(q2,2+iq2,1)(q3,0+iq3,3) = 1 3 [(q_1,0+iq_1,3)(q_2,0+iq_2,3)(q_3,2+iq_3,1)+(q_1,0+iq_1,3)(q_2,2+iq_2,1)(q_3,0+iq_3,3) +(q1,2+iq1,1)(q2,0+iq2,3)(q3,0+iq3,3)]. +(q_1,2+iq_1,1)(q_2,0+iq_2,3)(q_3,0+iq_3,3) ]. Each state is a linear combination of products with the single fiducial row mk=0m_k=0, exactly the physical lift described above and the form allowed by Proposition S1.1. We can now state the comparison with neural quantum states precisely. Let G=SU(2)G=SU(2) and define N:=(V1/2∗)⊗N,ℋN:=V1/2⊗N,K_N:=(V_1/2^*) N, _N:=V_1/2 N, (S1.28) where NK_N carries the row/multiplicity indices m and ℋNH_N carries the column/physical indices n. The multilinear and per-site-odd function spaces are ℱlin _lin :=span∏i=1NDmini1/2(qi)≅N⊗ℋN, :=span\! \ _i=1^ND^1/2_m_in_i(q_i) \ _N _N, (S1.29) ℱodd _odd :=f∈L2(GN):f(…,−qi,…)=−f(…,qi,…)∀i := \f∈ L^2(G^N):f(…,-q_i,…)=-f(…,q_i,…)\ ∀ i \ =⨁j1,…,jN∈12,32,…⨂i=1N(Vji∗⊗Vji). = _j_1,…,j_N∈\ 12, 32,…\ _i=1^N (V_j_i^* V_j_i ). (S1.30) Fix any unit vector χ∈Nχ _N. Its canonical lift is ιχ(Ψ)(q1,…,qN):=2N/2∑,χΨ∏i=1NDmini1/2(qi). _χ( )(q_1,…,q_N):=2^N/2\! _ m, n _ m _ n _i=1^ND^1/2_m_in_i(q_i). (S1.31) Theorem S1.2 (Ambient expressivity hierarchy). For N≥1N≥ 1, the lift (S1.31) is an isometry and ιχ(ℋN)⊊ℱlin⊊ℱodd. _χ(H_N)\; \;F_lin\; \;F_odd. (S1.32) Consequently, for every NQS variational family ℳNQS⊆ℋNM_NQS _N, ιχ(ℳNQS)⊆ιχ(ℋN)⊊ℱlin⊊ℱodd. _χ(M_NQS)\; \; _χ(H_N)\; \;F_lin\; \;F_odd. (S1.33) Let H~ H be the right-regular differential representation of the physical Hamiltonian H, generated by the left-invariant fields. On the linear class, H~|ℱlin≅IN⊗HℋN, H |_F_lin I_K_N H_H_N, (S1.34) where HℋNH_H_N denotes H acting on ℋNH_N. Consequently, every nonzero f∈ℱlinf _lin obeys ⟨f,H~f⟩⟨f,f⟩≥E0,inff∈ℱlin∖0⟨f,H~f⟩⟨f,f⟩=E0. f, Hf f,f ≥ E_0, _f _lin \0\ f, Hf f,f =E_0. (S1.35) On the nonlinear odd class, define C^:=∑i=1N(S^i2−34I). C:= _i=1^N ( S_i^2- 34I ). (S1.36) Then C^⪰0 C 0 on ℱoddF_odd and kerC^=ℱlin C=F_lin. For every certified λ≥μ⋆(J,h)λ≥ _ (J,h) from Theorem A.3, proved below, ⟨f,(H~+λC^)f⟩⟨f,f⟩≥E0for all f∈ℱodd∖0, f,( H+λ C)f f,f ≥ E_0 all f _odd \0\, (S1.37) and the infimum of the regularised quotient over ℱodd∖0F_odd \0\ is exactly E0E_0. The strict containments in Theorem S1.2 concern functions on the group manifold. The additional row Peter–Weyl index is a multiplicity, or purification, degree of freedom on which the physical Hamiltonian acts trivially; the column index carries the physical spin. This does not enlarge the physical 2N2^N-dimensional Hilbert space. Likewise, the nonlinear class adds higher-half-integer sectors that become variationally consistent only after the Casimir penalty. The theorem proves the exact ambient hierarchy and the variational principle for both the linear and nonlinear cases. S1.4 Central oddness and variational principles Notice first that the Peter–Weyl expansion splits into two parity classes under the per-site reflection qi→−qiq_i→-q_i. The point −q-q represents the central element −I∈SU(2)-I (2), which acts on the spin-j irrep as the scalar (−1)2j(-1)^2j, so the Wigner coefficients carry that same sign, Dmnj(−q)=(−1)2jDmnj(q),D^j_mn(-q)\;=\;(-1)^2j\,D^j_mn(q), (S1.38) even (+1+1) on the integer sectors and odd (−1-1) on the half-integer ones. The integer-j sectors are therefore exactly the even subspace of L2(SU(2))L^2(SU(2)), and the half-integer sectors the odd subspace. Hence, per-site oddness, ψ(…,−qi,…)=−ψ(…,qi,…)for every site i,ψ(…,-q_i,…)\;=\;-\,ψ(…,q_i,…) every site i, (S1.39) annihilates by (S1.38) every component with integer jij_i, and the support collapses onto the half-integer sectors ji∈12,32,52,…j_i∈\ 12, 32, 52,…\, with the physical target ji=12j_i= 12 now the lowest of them. The architecture of SM § S2 enforces (S1.39) exactly, so we take it as given here. Although each normalised carrier is a nonlinear function of its quaternion, the leaf and merge log-scales cancel those normalisations exactly in the reconstructed amplitude. Lemma S2.1 therefore shows that the final ψθ _θ is degree one in every qiq_i. A degree-one function on SU(2)SU(2) is a linear combination of the four coefficients Dmn1/2D^1/2_mn, so Proposition S1.1 places the ansatz in ℱlinF_lin identically. Two distinct cases follow as a consequence of Theorem S1.2. First, the foundation architecture we describe next is linear in each quaternion, so lies in ℱlinF_lin and directly obeys the variational principle E[ψθ]≥E0E[ _θ]≥ E_0. Second, any future generation of model which is per-site odd but non-linear in the quaternioncs lies in the larger ℱoddF_odd. Such a model’s higher-half-integeter sectors invalidate the variational principle as it stands, however it can be restored with an energy penalty restores the upper bound while retaining that larger ambient function class, see Appendix A. S2 Foundation Ansatz and Entanglement Structure via Reinforcement Learning In this section, we describe the foundation model’s architecture. The model has two different types of input. First is the Hamiltonian’s interaction tensor (J,h)(J,h), and second is a spin configuration q∈SU(2)Nq∈ SU(2)^N sampled from the model’s own density. In this section, we will assume efficient access to such samples, deferring the details of our novel sampling scheme to SM § S4.1. These two sources of data naturally divide the architecture into two components that together allow us to simultaneously condition the foundation model wavefunction on the Hamiltonian (allowing for generalisation) as well as preserving the necessary exchange anti-symmetry of our quaternionic representation from SM § S1. The first component, which we call the trunk, is a function of the Hamiltonian’s interaction tensor. In (§ S2.1) we describe how a given input (J,h)(J,h) is cached into two tensors J′J and h′h . In (§ S2.2) we then describe our learnable featurizer, which converts the cached inputs J′J and h′h into embeddings of the Hamiltonian’s system context for the trunk. We then describe the architecture of the trunk itself in (§ S2.3), which consists of L residual attention blocks nested between feed-forward networks, as well as a per-site leaf builder where the spin configuration q enters (§ S2.3). The second component is a readout contraction mechanism of the trunk that incorporates spin coordinates. This component is akin to a tree-tensor network, augmented by three routines that enable generalisation to arbitrary quadratic Hamiltonians and system sizes. First, we use edge-level attention in a tree to enable it to condition on the Hamiltonian’s interaction tensor. Second, we use a learnable routing policy to allow the tree to entangle arbitrary sites in the system. Finally, we use a contextualizer to allow the tree to condition on Hamiltonian data while preserving the necessary exchange anti-symmetry of our quaternionic representation. The contraction itself is detailed in (§ S2.4), where we describe a rank-4 quadrilinear merge with level attention on each contraction layer, as well as the contextualizer. Section (§ S2.5) then describes the learnable routing policy to optimise over contraction paths with deep reinforcement learning. Together, the merged contraction with a learnable path enables arbitrary entanglement between sites, with the entanglement conditioned on the Hamiltonian’s interaction tensor (J,h)(J,h) enabling a Hamiltonian-dependent contraction pathway. Figure S1 shows the information flow of our ansatz over a forward pass. Recall for the spin-configuration input q∈(SU(2))Nq∈(SU(2))^N, we use the quaternion representation embedded in ℝ4R^4, so qi=(qi,0,qi,1,qi,2,qi,3)q_i=(q_i,0,q_i,1,q_i,2,q_i,3) with qi,02+qi,12+qi,22+qi,32=1q_i,0^2+q_i,1^2+q_i,2^2+q_i,3^2=1. This makes the input space a product of spheres SU(2)≅S3↪ℝ4SU(2) S^3 ^4. Hamiltonian(J,h)(J,\,h)Build J′,h′J ,\,h SystemFeaturizerPure-even trunk, ×L× L blocksLeaf contextualizer ×2× 2refines (ℓ(L),e(L))( ^(L),e^(L)) over the slot framePer-site leaf buildertakes (ℓ^i,g,qi)( _i,\,g,\,q_i), emits the odd carrier ui(0)u_i^(0)Merge-tree readoutbit-rev leaves; rank-4 quadrilinear merge; level attentionlogψθ(q∣p)∈ℂ _θ(q p)\;∈\;CRoute policyp∼πθ(p|h,ℓ(L),e(L),m)p _θ(p\,|\,h, ^(L),e^(L),m)Spin config q∈(SU(2))N∈(SU(2))^Nfrom MCMCJ,hJ,hJ′∈ℝN×N×10,h′∈ℝN×3J ^N\!×\!N\!×\!10,\;\,h ^N\!×\!3e(0),ℓ(0),g(0)e^(0),\, ^(0),\,g^(0)ℓ(L),e(L),g ^(L),\,e^(L),\,gℓ^∈ℝN^×dℓ,g N\!×\!d_ ,\;\,gu(0)∈ℝN^×duu^(0) N\!×\!d_uurootu_rootℓ(L),e(L),g ^(L),\,e^(L),\,groute ppqq Figure S1: Information flow through the foundation ansatz. The Hamiltonian (J,h)(J,h) is cached into the two tensors J′J and h′h (§ S2.1) and consumed by the SystemFeaturizer (§ S2.2), which emits the per-bond embedding e(0)e^(0), the per-site local embedding ℓ(0) ^(0) and the seed g(0)g^(0) of the global stream. The trunk (§ S2.3) then propagates all three streams through L residual attention blocks, refreshing g once per block (Figure S3 traces the stream on its own). Notice the spin configuration q does not enter the trunk, it is a function of the Hamiltonian interaction alone. After the trunk, the route policy (§ S2.5) samples a permutation p∈SN^p∈ S_ N from πθ(p∣h,ℓ(L),e(L),m) _θ(p h, ^(L),e^(L),m), conditioned by its own fork of the global stream, and that permutation relabels both the trunk outputs (ℓ(L),e(L))( ^(L),e^(L)) and the spin configuration q before downstream consumption. In particular before the contextualizer, whose slot clocks and tree-distance decays encode position and must therefore act in the routed frame. A two-block leaf contextualizer (§ S2.4.3) then refines the routed trunk outputs over the padded slot frame, manufacturing contexts for the holes, and emits the per-slot states ℓ together with the rewritten edges that seed the tree’s level-00 edge field E(0)E^(0) (the latter not drawn). The per-site leaf builder (§ S2.3.1) is the unique entry point for q and consumes (ℓ^i,g,qi)( _i,g,q_i) per site, with g the stream’s refreshed post-context value (it does not see the per-bond edges). Its output ui(0)u_i^(0) feeds the tree readout (§ S2.4), which contracts the N N leaf carriers down to a single root carrier; the per-bond e(L)e^(L) re-enters the tree internally at the level-edge attention in (§ S2.4.4) (not drawn here). The root readout is the routed log-amplitude logψθ(q∣p)∈ℂ _θ(q p) . S2.1 Hamiltonian encoding and spectral normalisation Recall from §S1.2 that the Hamiltonian family we target is H^=14∑i<j∑a,kJijakσiaσjk+12∑i∑ahiaσia, H\;=\; 14 _i<j _a,kJ^ak_ij\,σ^a_i\,σ^k_j\;+\; 12 _i _ah^a_i\,σ^a_i, (S2.1) with exchange tensor J∈ℝN×N×3×3J ^N× N× 3× 3 (cross-channel Pauli-pair couplings on off-diagonal (i,j)(i,j)) and transverse field h∈ℝN×3h ^N× 3 (per-site three-vector). As stated in the main text, it is useful to think of J as the data of a weighted graph on N sites, where each unordered pair (i,j)(i,j) with i≠ji≠ j carries a 3×33× 3 Pauli-pair coupling matrix Jij∈ℝ3×3J_ij ^3× 3, which we read as nine flavours of edge, one per Pauli-pair (a,b)∈x,y,z2(a,b)∈\x,y,z\^2 giving the operator σiaσjbσ^a_iσ^b_j, see Fig. 2. We keep exchange and field as separate input channels for the model. Before feeding them into the model, we normalize them in the following way. Step 1 (off-diagonal — bond-symmetrise). On Hermitian-conjugate pairs (i,j)↔(j,i)(i,j) (j,i), average to enforce the swap symmetry: Jij,absym=12(Jij,ab+Jji,ba∗)(i≠j).J^sym_ij,ab\;=\; 12\, (J_ij,ab+J^*_ji,ba ) (i≠ j). (S2.2) We do this because Pauli interactions are hermitian, so at the level of the interaction tensor and its relation to an interaction graph, it must be bi-directionally symmetric. Step 2 (form the common Hermitian scale matrix). For the scale computation only, the cache builder forms Mij,ab=Jij,absym+δijiεabchic,M_ij,ab\;=\;J^sym_ij,ab+ _ij\,i\, _abc\,h_i^c, (S2.3) and reshapes it as M~:=reshape(M, 3N×3N). M:=reshape(M,\,3N× 3N). (S2.4) For real hih_i, the block Ai,ab:=εabchicA_i,ab:= _abch_i^c is real and antisymmetric, so (iAi)†=iAi.(iA_i) =iA_i. (S2.5) Together with the Hermitian swap symmetrisation of Step 1, this makes M~ M Hermitian. Its eigenvalues are therefore real, and the shared scale is s:=‖M~‖2=maxk|λk(M~)|,s:=\| M\|_2= _k | _k( M) |, (S2.6) computed once per system with jnp.linalg.eigvalsh. The spectral radius maxk|λk| _k| _k| is defined for every square matrix. The role of the factor i is instead to make M~ M Hermitian, so that the Hermitian eigensolver is appropriate and the spectral radius equals the spectral norm. The diagonal field block participates only in this scale computation and is discarded afterwards. Step 3 (normalise and flatten). We divide the separate exchange and field channels by the same scale s. The exchange keeps its off-diagonal bonds and an identity-marker channel, with the diagonal Pauli block left empty since the field no longer lives there, Jij,c′=(Jsym/s)ij,ab(c)c=0,…,8andi≠j,0c=0,…,8andi=j,δijc=9,J _ij,c\;=\; cases(J^sym/s)_ij,ab(c)&c=0,…,8\ and\ i≠ j,\\ 0&c=0,…,8\ and\ i=j,\\ _ij&c=9, cases (S2.7) where c=0,…,8c=0,…,8 enumerates the nine (a,b)(a,b) pairs in row-major order, and the field is cached as its own per-site tensor, hi′=hi/s.h _i\;=\;h_i/s. (S2.8) Building these inputs costs one call to eigvalsh per system, paid once per system when a panel or an augmentation round is loaded, and once per system at evaluation or fine-tuning. We emphasise here that the size of the interaction tensor, and the graph that follows, is polynomial in the number of bodies, since our foundation model handles quadratic interactions which are a small sub-group of the overall Pauli group on N spins. By construction, this normalisation makes both cached inputs invariant under simultaneous rescaling (J,h)→α(J,h)(J,h)→α(J,h) for α>0α>0, since the spectral radius s scales linearly with α. This avoids the attention layers needing to learn the invariance from our datasets. Under α<0α<0 the inputs are sign-flip-equivariant: a global sign passes through to the nine Pauli channels of J′J and to h′h , while the identity channel is unaffected. Finally they are smooth at h=0h=0, since h′h is linear in h and the scale s stays bounded away from zero. These three properties are exactly the ones our featurizer (§ S2.2) relies on for equivariance under rescaling and for the h→0h→ 0 limit. This encoding is a bijection between the original (J,h)(J,h) and the cached pair (J′,h′)(J ,h ), and both representations are in the datasets we have open-sourced. The diagonal field block in M is solely a common-normalisation device. It is neither a cached model input nor a replacement of the physical first-order field operator. Including both exchange and field in M~ M lets one scale normalise the two channels together. Once s has been computed, the diagonal block is discarded, and the model receives the bond tensor J′J and the per-site field h′h separately, while the local-energy expression retains the field as the first-order term in Eq. (S1.16). The per-system rescaling-invariance is then the standard ML input-normalisation trick applied at the level of the Hamiltonian itself. Dividing through by the spectral radius lands every system in a unit-norm regime regardless of physical units, so the training dataset’s overall energy scale stops being a feature the network has to re-learn for every batch. Indeed this has direct analogy to dividing pixel values by 255255 or attention keys by d d, removing a known-uninformative scale so the network spends capacity on what actually distinguishes systems. We can simply rescale the output by the renormalisation factor s to get back to physical units, and optimise in the normalised space where the network is more stable and generalises better. S2.2 Learnable featurizer The featurizer is a single fusion block that consumes the two cached inputs and emits three downstream-ready tensors: SystemFeaturizer:(J′∈ℝN×N×10,h′∈ℝN×3)⟼(e∈ℝN×N×96,ℓ∈ℝN×128,g∈ℝ1024). SystemFeaturizer:\; (\,J ^N× N× 10,\;h ^N× 3\, )\; \; (\;e ^N× N× 96,\;\; ^N× 128,\;\;g ^1024\; ). (S2.9) These three objects correspond to a per-bond embedding e for the trunk’s attention layers, a per-site local embedding ℓ for the trunk’s attention layers, and a global context vector g that seeds the global stream the trunk maintains (§ S2.3). The featurizer is a single-pass architecture with no iteration or recurrence; it is a single function of the Hamiltonian that produces these three tensors in one go. Figure S2 shows the data-flow and we detail each step below. J′∈ℝN×N×10J ^N×N×10Polar splitPer-bond embeddinggrouped RMSColumn attentionRow attention (bonds)h′∈ℝN×3h ^N×3Polar splitZeeman basegrouped RMSRow attention (field)Global fusionLocal cross-conditioningEdge cross-conditioningglobal FiLM(e,ℓ,g)(e,\; ,\;g)x^J∈ℝN×N×24 x^J\!∈\!R^N×N×24Bℓraw ^rawx^h∈ℝN×7 x^h\!∈\!R^N×7ℓz ^zgrawg^rawgzg^zgℓ ℓraw ^rawBℓz ^zg (FiLM)ϕspec,ϕJh _spec,\, _Jh Figure S2: Data flow inside the system featurizer (§S2.2). The two cached inputs enter on separate spines, each first mapped to its polar form (raw, radial, and angular coordinates): the exchange J′J is embedded per bond (B) by a grouped-RMS map and aggregated by column attention into the bond descriptor ℓraw ^raw, while the field h′h is lifted into the Zeeman base ℓz ^z by the same grouped-RMS pattern. Two row-attention layers pool the two views (grawg^raw, gzg^z), which a fusion map, fed also by the spectrum block ϕspec _spec and the relative-scale channels, combines into the global vector g, the seed of the global stream. Solid arrows are the main pipeline; dashed magenta arrows are skip-connections: the Zeeman base ℓz ^z is the residual backbone of the local feature ℓ , conditioned by ℓraw ^raw and g, and the edges are conditioned by B and, through a global FiLM, by g. The outputs (e,ℓ,g)(e, ,g) are the three tensors consumed by the trunk. Both cached tensors enter in polar form. We read the nine Pauli-pair entries of each bond block as a vector xijJ=vec(Jij,1:9′)∈ℝ9x^J_ij=vec(J _ij,1:9) ^9 and split it into a radius and a softened direction, rijJ=∥xijJ∥2,uijJ=xijJ∥xijJ∥22+τ2,r^J_ij= x^J_ij _2, u^J_ij= x^J_ij x^J_ij _2^2+τ^2, (S2.10) where τ=10−3τ=10^-3 is a fixed feature-scale floor, and ||⋅||2||·||_2 is the Euclidean norm. The radius rJr^J shows whether a bond is present, whilst the direction uJu^J reports its angular type and stays bounded as the bond vanishes. Notice τ is a fixed feature-scale floor, not a numerical ε . The bond input stacks the raw, radial, angular, and self-bond coordinates, and the field is lifted the same way, x^ijJ=[xijJ,rijJ,uijJ, 1i=j]∈ℝ20,x^ih=[hi′,∥hi′∥2,hi′∥hi′∥22+τ2]∈ℝ7. x^J_ij= [\,x^J_ij,\;r^J_ij,\;u^J_ij,\;1_i=j\, ] ^20, x^h_i= [\,h _i,\; h _i _2,\; h _i h _i _2^2+τ^2\, ] ^7. (S2.11) They are functions of the cached inputs alone and are computed once per system alongside them (§ S2.1). We hand the model the coordinate, its magnitude, and its direction separately rather than make a dense layer recover all three: the loci that matter physically, vanishing bonds or fields, isotropic blocks, vanishing skew or symmetric-traceless parts, rank-deficient couplings, are exactly where a plain projection smears presence into direction. The two embedding paths normalise in groups, and the normalisation is the root-mean-square kind throughout: for a vector z with learned per-channel scale γ, RMS(z)=γ⊙z2¯+ε,z2¯:=1d∑aza2,RMS(z)\;=\;γ z z^2+ , z^2\;:=\; 1d _az_a^2, (S2.12) with no mean subtraction and no shift. The mean subtraction of LayerNorm is left out on purpose: it couples every channel through a shared offset and worsens the conditioning of the curvature the optimiser preconditions against, and dropping it is the same trade that moved the large-language-model community from LayerNorm to RMSNorm [163]. Every normalisation layer on the even path, here and in every later stage, is of this form. Reshaping a vector into K groups of width dgrpd_grp, the grouped variant normalises each group on its own with its own scale, GRMS(z)k=γkzkzk2¯+ε,γk=at init,GRMS(z)_k= _k\, z_k z_k^2+ , _k=1\;at init, (S2.13) with zk2¯ z_k^2 the mean square within group k. Each group is then free to become a learned filter over the raw, radial, and angular coordinates, a data-driven stand-in for the named decomposition J=tI+Q+AJ=tI+Q+A into isotropic, symmetric-traceless, and skew parts, without hard-coding it. We use K=16K=16 groups of width dgrp=16d_grp=16 on both paths. The per-bond embedding is then a two-layer map, a grouped norm, and two more layers, zijJ=GRMS(SiLU(x^ijJW1J)W2J),Bij=mimj⋅SiLU(SiLU(zijJW3J)W4J),z^J_ij=GRMS\! (SiLU( x^J_ij\,W^J_1)\,W^J_2 ), B_ij=m_im_j·SiLU\! (SiLU(z^J_ij\,W^J_3)\,W^J_4 ), (S2.14) with dbond=128d_bond=128, every bias zero at initialisation, and mimjm_im_j the site-padding pair mask: it zeroes pairs with a padded endpoint and carries no bond information. Bonds absent from the system’s adjacency pattern are not masked to zero. Wherever a bond feature is consumed, absence substitutes a learned token in its place, zij⟵zij,presentijt,otherwisepresentij=[∃c≤8:Jij,c′≠0]∨[i=j],z_ij\; \; casesz_ij,&present_ij\\ t,&otherwise cases _ij\;=\; [\,∃\,c≤ 8:\;J _ij,c≠ 0\, ]\; \;[\,i=j\,], (S2.15) with one trained token t per consumption site. Six sites carry a token: the bond embedding, the bond attention’s key and value input, the Zeeman base, the field’s row-attention view, the field branch of the global fusion, and the field’s entry in the local cross-conditioning. On the field path a zero field substitutes its tokens the same way, per site everywhere except the fusion branch, which gates per system. The physical reading is that “no bond” becomes a representable direction of the basis rather than the numeric zero, while the substitution still cuts the degenerate-input gradients that a hard zero would poison. The site-padding masks mim_i and MjkeyM^key_j are a different object: they mark sites absent from the padded system, not bonds absent from a present system’s graph. Nothing downstream carries a bond mask; bond presence reaches the trunk only through these tokens. Column attention then attends over the j index for each (i,)(i, h) using per-row learnable queries Qcol∈ℝdQ^col_ h ^d_ h, Kij,=WKRMS(Bij),Vij,=WVRMS(Bij),αij,col=softmaxj(Qcol⋅Kij,d+Mjkey),K_ij, h=W_K\,RMS(B_ij), V_ij, h=W_V\,RMS(B_ij), α^col_ij, h=softmax_j\! ( Q^col_ h· K_ij, h d_ h+M^key_j ), (S2.16) where ∈1,…,n h∈\1,…,n_ h\ is the attention-head index (we use sans-serif to disinguish from the transverse field h), d_ h is the per-head width with n⋅d=dbondn_ h· d_ h=d_bond, and WK,WV∈ℝdbond×dbondW_K,W_V ^d_bond× d_bond are linear projectors whose outputs split into n_ h head blocks of width d_ h indexed by h. Here RMSRMS is the map of Eq. (S2.12), applied along the bond-feature axis, softmaxjsoftmax_j normalises over the second site index, and MjkeyM^key_j is the key-padding mask which is 00 when site j is present in the system, and −∞-∞ when it is absent. The kernel computes twice the output head count and gates the heads in bisected pairs, oi,=∑jαij,colVij,,ℓiraw=mi⋅[sigmoid(oi(1:n))⊙oi(n+1:2n)],o_i, h= _jα^col_ij, h\,V_ij, h, ^raw_i=m_i· [\,sigmoid (o_i^(1:n_ h) ) o_i^(n_ h+1:2n_ h)\, ], (S2.17) a GLU-style multiplicative gate at the head level [29, 135]. Every attention layer of SM § S2 shares this micro-structure. The descriptor ℓiraw ^raw_i carries the exchange structure seen from site i. The field enters on its own per-site channel. Because each site carries its own hi′h _i, the field is a genuinely local quantity, and we lift its polar input x^ih x^h_i into a Zeeman base through the same two-layer / grouped-norm / two-layer pattern, zih=GRMS(SiLU(x^ihW1h)W2h),ℓiz=mi⋅SiLU(zihW3h)W4h,z^h_i=GRMS\! (SiLU( x^h_i\,W^h_1)\,W^h_2 ), ^z_i=m_i·SiLU(z^h_i\,W^h_3)\,W^h_4, (S2.18) with dℓ=128d_ =128 and, again, every bias zero at initialisation. The field groups learn their own thresholded detectors — field nearly zero, field isotropic, a single dominant axis — in the same way the bond groups do. The global context pools both views of the system. A first row-attention layer with nqn_q learnable global queries Qq,row∈ℝnq×n×dQ^row_q, h ^n_q× n_ h× d_ h attends over the bond descriptors ℓraw ^raw to give grawg^raw, and a second, structurally identical layer attends over the Zeeman bases ℓz ^z to give gzg^z. Meanwhile the relative weight is already implicit in (J′,h′)(J ,h ), and ϕJh _Jh surfaces it as an explicit bounded feature: ϕJh=[log(1+rJ),log(1+rh)]. _Jh\;=\; [\, (1+r_J),\; (1+r_h) ]. (S2.19) The branch is guarded so its gradient signal remains clean on the h=0h=0 majority of the panel. The fusion normalises each pooled view on its own, takes the two derived blocks raw (both are bounded by construction), and closes with an output normalisation, g=RMS(FFNg([RMS(graw),RMS(gz),ϕJh])).g\;=\;RMS (FFN_g\! ( [\,RMS(g^raw),\,RMS(g^z),\, _Jh\, ] ) ). (S2.20) On a system that carries no field at all, RMS(gz)RMS(g^z) is replaced by the fusion’s own absent-field token. The per-site local feature then takes the Zeeman base as its backbone and folds in the bond descriptor and the global context through a pre-norm residual map, each input normalised on its own, ℓi=mi⋅RMS(ℓiz+FFNℓ([RMS(ℓiz),RMS(ℓiraw),RMS(g)])), _i\;=\;m_i·RMS (\, ^z_i+FFN_ \! ( [\,RMS( ^z_i),\,RMS( ^raw_i),\,RMS(g)\, ] ) ), (S2.21) and at a zero-field site the first entry is this map’s own absent-field token. The field is the residual base because it is the dominant per-site signal, while the exchange reaches a site only by aggregation over its bonds. Edge cross-conditioning then refines each bond with the new local features, and the global vector enters only as a gentle modulation. The core mixes the bond with its two endpoints, while a FiLM scale and shift read off the global context, cij=FFNe(RMS[Bij,ℓi,ℓj]),[γ,β]=WFRMS(g)+bF,γ←0.1tanhγ,β←0.1β,c_ij=FFN_e\! (RMS [\,B_ij,\, _i,\, _j\, ] ), [γ,β]=W_F\,RMS(g)+b_F, γ← 0.1 γ,\;\;β← 0.1β, (S2.22) eij=mimj⋅RMS(PBij+(1+γ)cij+β).e_ij\;=\;m_im_j·RMS\! (P\,B_ij+(1+γ)\,c_ij+β ). (S2.23) The FiLM factors start near zero, P is a bias-free residual projector applied when dbond≠dedged_bond≠ d_edge, and the pair mask is, again, padding alone. Given the graph structure of the Hamiltonian input, we note here that the featurizer is a GNN in the sense that it performs neighbor aggregation and global pooling. There is, however, no separate message-passing module since attention subsumes the GNN role, framed in a form compatible with our GPU kernel pipeline and natural-gradient optimisers see SM § S4 for details of our engineering. S2.3 Trunk and per-site leaf builder Recall that the featurizer of § S2.2 emits three tensors: the per-bond embedding e∈ℝN×N×dee ^N× N× d_e, the per-site local feature ℓ∈ℝN×dℓ ^N× d_ , and the global context g∈ℝdgg ^d_g. Both e and ℓ depend on the Hamiltonian (J,h)(J,h) only through the cached inputs J′J and h′h of § S2.1, and every system-context quantity downstream of this point is, by construction, a function of (J′,h′)(J ,h ) and therefore a function of (J,h)(J,h). The trunk’s job is to propagate (e,ℓ,g)(e, ,g) through L residual blocks: it maintains all three streams, mixing information across sites and bonds and keeping the global summary current as it does so. We emphasise that this is completely independent functions of spin configuration q, which enter only once, at a per-site leaf builder that sits between the trunk and the merge tree. As such the per-site antisymmetry constraint of § S1.4, ψ(…,−qi,…)=−ψ(…,qi,…)for every site i,ψ(…,-q_i,…)\;=\;-\,ψ(…,q_i,…) every site i, (S2.24) becomes a property of the leaf builder alone, and the expressivity of the trunk is not limited by any symmetry requirement and thus we may use any standard transformer recipe [148]. The trunk is structured by blocks, indexed by β∈0,1,…,L−1β∈\0,1,…,L-1\. Let ℓ(β)∈ℝN×dℓ ^(β) ^N× d_ and e(β)∈ℝN×N×dee^(β) ^N× N× d_e be the per-site and per-bond states at block β, with ℓ(0)=ℓ ^(0)= and e(0)=e^(0)=e from the featurizer. We write ℓi∈ℝdℓ _i ^d_ for the row of ℓ(β) ^(β) at site i and eij∈ℝdee_ij ^d_e for the bond slot at (i,j)(i,j), dropping the (β)(β) superscript when the block is unambiguous. The global context is the trunk’s third stream, and it is a maintained one. The natural alternative is to broadcast the featurizer’s g unchanged, as a fixed conditioning signal for every block. But a fixed global summarises the couplings as the featurizer saw them, and that summary grows stale exactly as the trunk’s refinement makes the site and bond states worth summarising. The trunk therefore refreshes g once per block, from the states the block has just updated. The stream lives on the sphere of unit root-mean-square norm: writing Norm(x):=xmax(x2¯,ε),x2¯:=1d∑axa2,Norm(x)\;:=\; x \! ( x^2,\, ), x^2\;:=\; 1d _ax_a^2, (S2.25) for the parameter-free projection ℝd→ℝdR^d ^d onto that sphere (the floor ε=10−4 =10^-4 replaces the usual added constant, so the map is exact whenever x2¯≥ε x^2≥ and sends the origin to itself), the trunk receives the featurizer’s global as g(0)=Norm(g)g^(0)=Norm(g) and returns every update to the same sphere. NormNorm carries no parameters; the RMS layers of § S2.2 normalise the same way but carry a learned per-channel scale, and the two should not be conflated. A block applies three residual sublayers in sequence: an edge-update sublayer on the per-bond stream, then an edge-biased multi-head self-attention (MHA) sublayer and a feed-forward network (FFN) sublayer on the per-site stream; a refresh of the global closes the block. The edge stream moves first, e(β+1)=e(β)+1LEdgeUpdate(RMS(e(β)),ℓ(β)),e^(β+1)\;=\;e^(β)\;+\; 1 L\,EdgeUpdate\! (RMS(e^(β)),\; ^(β) ), (S2.26) so that the attention which follows reads fresh bonds: the just-updated e(β+1)e^(β+1) enters the attention logits as a per-head additive bias, defined below. The per-site stream then takes two residual steps, ℓ′ \; =ℓ(β)+1LMHA(RMS(ℓ(β)),e(β+1)), =\; ^(β)\;+\; 1 L\,MHA\! (RMS( ^(β));\,e^(β+1) ), (S2.27) ℓ(β+1) ^(β+1)\; =ℓ′+1LFFN(RMS(ℓ′)+Wgg(β)), =\; \;+\; 1 L\,FFN\! (RMS( )+W_g\,g^(β) ), (S2.28) and the global closes the block by reading the sites the block has just written, g(β+1)=Update(g(β),pool(g(β),ℓ(β+1))),g^(β+1)\;=\;Update\! (g^(β),\;pool (g^(β),\, ^(β+1) ) ), (S2.29) with poolpool and UpdateUpdate defined below. In Eq. (S2.27) the semicolon indicates that the map is conditioned on the argument after it; RMSRMS is the normalisation of § S2.2, applied along the feature axis; and Wg∈ℝdℓ×dgW_g ^d_ × d_g is a learned projection whose output is broadcast to every site. The global enters the per-site path exactly once, additively, at the FFN’s input. We emphasise that the two per-site steps are sequential residuals (the FFN acts on the attention’s output ℓ′ , not on ℓ(β) ^(β)) and that all three sublayers are pre-norm: each reads a normalised copy of its stream and adds its result to the raw stream. The factor 1/L1/ L applies inverse-square-root depth scaling to each residual branch, matching the depth dependence induced by batch normalisation at initialisation [30]. Under the usual independence approximation at initialisation, each additive update contributes variance (1/L)O(1/L) to the running stream, so the cumulative variance across L blocks stays (1)O(1). The two maps that refresh the global recur at every later stage of the architecture; the trunk’s instance is the first of many. The pool is a descriptor attention: its query is produced from the current global state and determines the attention weights over the input set. For a set xii=1n\x_i\_i=1^n of states xi∈ℝdxx_i ^d_x and the current global g, q=WQg,k,i=WKRMS(xi),v,i=WVRMS(xi),q_ h=W^Q_ h\,g, k_ h,i=W^K_ h\,RMS(x_i), v_ h,i=W^V_ h\,RMS(x_i), (S2.30) a,i=softmaxi(q⋅k,idk+Mi),pool(g,xi)=[∑ia,iv,i]=1n∈ℝndv,a_ h,i=softmax_i\! ( q_ h· k_ h,i d_k+M_i ), (g,\x_i\ )= [\, _ia_ h,i\,v_ h,i\, ]_ h=1^n_ h\;∈\;R^n_ hd_v, (S2.31) where h is the sans-serif head index of § S2.2, WQ∈ℝdk×dgW^Q_ h ^d_k× d_g, WK∈ℝdk×dxW^K_ h ^d_k× d_x and WV∈ℝdv×dxW^V_ h ^d_v× d_x are per-head projectors, [⋅][·]_ h concatenates the head outputs, and MiM_i is the key-padding mask of § S2.2 (00 where the element is present, −∞-∞ where it is absent). This is the featurizer’s row attention with two changes: the learnable queries are replaced by projections of g, and the head outputs concatenate plainly, without the gate of Eq. (S2.17). We use n=4n_ h=4 heads of width dk=dv=64d_k=d_v=64, so the pooled reading has width ndv=256n_ hd_v=256, a quarter of the stream’s own width dg=1024d_g=1024. The Hamiltonian’s data streams use the normalised interpolation of nGPT [86]. For a current state x and candidate Δx _x, let axa_x be a learned per-channel gate, write 1 for the all-ones vector and ⊙ for elementwise multiplication, and define logit(r):=log(r/(1−r))logit(r):= (r/(1-r)). Then α~x:=0.8sigmoid(ax),ax|init=logit(0.20.8),x(x,Δx):=Norm(x+α~x⊙(Norm(Δx)−x)). α_x:=0.8\,sigmoid(a_x), a_x |_init=logit\! ( 0.20.8 )1, _x(x, _x):=Norm\! (x+ α_x (Norm( _x)-x ) ). (S2.32) Thus every channel starts with interpolation weight 0.20.2 and remains in the open interval (0,0.8)(0,0.8). Throughout the architecture, ∙N_ denotes this same normalised-interpolation map with an independently learned gate for the named stream. Given a pooled reading p, the global update reads, Δg=FFNg(RMS[Wtapg,p]),Update(g,p)=g(g,Δg). _g=FFN_g\! (RMS[W_tapg,p] ), (g,p)=N_g(g, _g). (S2.33) The bottleneck WtapW_tap balances the global and pooled entries before the joint RMS normalisation. Figure S3 traces the stream’s stations through the architecture. ggtrunk block ×L× Lpost-trunkg(0)g^(0)ctx block ×2× 2ctx block ×2× 2post-ctxmerge level ×K× Kroot readoutpost-ctxposition ttpointerggg(t)g^(t) Figure S3: The global stream. From the featurizer’s seed onward, g lives on the unit-RMS sphere, and every station applies the same two maps: a descriptor pool (Eq. (S2.31)), then the normalised interpolation of Eq. (S2.33). The trunk refreshes g once per block from the updated site states and closes with the factored row-and-column edge reading (Eq. (S2.37)). The stream then forks: the physics leg and the router leg each refresh their own copy through a leaf-contextualizer stack and its closing edge reading (§ S2.4, § S2.5), and the copies never exchange information again. The physics copy is refreshed once per merge level and conditions the root readout; the router copy advances per decode position by reading the dyadic cover, and conditions the pointer. The stream is wide, dg=1024d_g=1024, and stays wide along the whole ladder; what varies is how the rest of the network reads it. The attention pools and the additive taps (WgW_g above, and their analogues downstream) are learned projections by construction. The remaining consumers, the leaf gate, the merge context leg and the root readout hypernet, each read the stream through their own RMS-normalised linear projection to a narrower interface of width 256256; where a later display writes g as an input to one of these three, it denotes that consumer’s projected copy. The attention itself is not the vanilla recipe, and both of its deviations recur across the architecture. First, the logits carry the bond stream: logit,ij=q,i⋅k,jdMLP(RMS(eijβ+1))logit_ h,\,ij\;=\; q_ h,i· k_ h,j d_ h\;MLP(RMS(e^β+1_ij)) (S2.34) with β:ℝde→ℝ _ h:R^d_e a small per-head scalar projector, so a head can attend along the interaction graph rather than by content alone. Second, the heads are gated in bisected pairs, the GLU micro-structure the featurizer’s attention layers already carry (Eq. (S2.17)). The level attention of § S2.4.4 and the contextualizer’s slot attention [85] reuse it in turn. Figure S4 draws one block. e(β)e^(β)ℓ(β) ^(β)g(β)g^(β)pre-RMS edge update++e(β+1)e^(β+1)1L 1 Lℓi,ℓj _i,\, _jpre-RMS edge-biased MHA++pre-RMS FFN++ℓ(β+1) ^(β+1)1L 1 L1L 1 Lfresh bondspool ++ updateg(β+1)g^(β+1)WgW_gℓ(β+1) ^(β+1) Figure S4: One block of the trunk (§S2.3); sublayers run in the order drawn. The per-bond stream e(β)∈ℝN×N×dee^(β) ^N× N× d_e (deeper teal) moves first through the pre-norm edge update of Eqs. (S2.35)–(S2.36), consuming the endpoint sites ℓi,ℓj _i, _j. The freshly updated bonds e(β+1)e^(β+1) then bias the per-site attention’s logits (Eq. (S2.34)), whose output is head-gated, and a pre-norm FFN closes the per-site path. Every sublayer normalises with RMS and carries the residual gain 1/L1/ L. The global stream g(β)g^(β) (magenta) enters the node FFN through th additive projection Wgg(β)W_g\,g^(β) and is refreshed once per block from the updated site states (Eq. (S2.29)). L identical blocks are stacked. The edge-update map itself combines a direct-pair contribution with a three-part aggregate that pools over every site k acting as a third vertex. Concretely, its increment at bond (i,j)(i,j) is EdgeUpdate(RMS(e),ℓ)ij=FFN([pairij,Pij]),EdgeUpdate (RMS(e),\, )_ij\;=\;FFN\! ( [pair_ij,\,P_ij ] ), (S2.35) where pairij:=[eij,ℓi,ℓj]pair_ij:=[\,e_ij, _i, _j\,] is the direct-pair feature, i.e. the bond together with its two endpoints, and Pij∈ℝdcP_ij ^d_c is the three-part aggregate, Pij,c=1N∑k=1NΨcL(pairik)ΨcR(pairkj),c∈1,…,dc.P_ij,\,c\;=\; 1 N _k=1^N ^L_c(pair_ik)\, ^R_c(pair_kj), c∈\1,…,d_c\. (S2.36) The two embedders ΨL,ΨR:ℝde+2dℓ→ℝdc ^L, ^R:R^d_e+2d_ ^d_c are independent SiLU MLPs that act on the left leg i→ki→ k and the right leg k→jk→ j of the triangle (i,k,j)(i,k,j). The 1/N1/ N prefactor keeps the variance of PijP_ij at (1)O(1) as N grows. Sites k that do not complete a physical triangle with (i,j)(i,j) enter through legs carrying the absent-bond tokens of Eq. (S2.15), which the embedders learn to discount; absence is a representable input here, not a hard zero. The update costs (N2dedc)O(N^2d_ed_c) flops per layer, which is the same asymptotic order as the pair-feature attention already in the block. This is the triangular update pattern of AlphaFold’s Evoformer [67], adapted from residue-pair representations to coupling graphs. We opt for this three-part update because it is well suited to handle frustrated systems. On a triangular antiferromagnet, a kagome lattice, or any J1J_1–J2J_2 chain whose nearest- and next-nearest-neighbour bonds close into triangles, two bonds can share endpoint features and still sit in physically distinct environments. Thus the bond’s effective behaviour is set by the triangle of competing couplings it belongs to, not by its endpoints alone. Without PijP_ij, Eq. (S2.35) reduces to the pair-only form standard in message-passing GNNs, which collapses such bonds into a single representation. With PijP_ij, the FFN sees the full multiset (pairik,pairkj):k\(pair_ik,pair_kj):k\ of triangle completions and separates them. In graph-theoretic language, the pair-only limit of Eq. (S2.35) is 1-Weisfeiler–Lehman in distinguishing power on the line graph, which is the level at which most off-the-shelf message-passing GNNs operate. The triangular aggregate PijP_ij lifts the update strictly to the 2-WL refinement on pairs [21, 89] which distinguishes any two bonds whose triangle structures differ. After the L blocks the global context stream takes one further update, this time reading the bonds rather than the sites. Placing all N2N^2 bond states under a single softmax would be wasteful, so the reading is factored through row and column descriptors, ri=pool(g,ei⋅(L)),cj=pool(g,e⋅j(L)),p=pool(g,r1,…,rN,c1,…,cN),r_i=pool (g,\,e^(L)_i\,· ), c_j=pool (g,\,e^(L)_·\,j ), p=pool (g,\,\r_1,…,r_N,\,c_1,…,c_N\ ), (S2.37) followed by the update g↦Update(g,p)g (g,p) of Eq. (S2.33). The row and column descriptors compress the N2N^2 bonds to 2N2N summaries, and the second-stage pool reads those; every pool here runs under the physical-site mask. The same factored reading recurs whenever the global reads an edge field downstream. The trunk thus emits the triple (e(L),ℓ(L),g)(e^(L), ^(L),g), where g now denotes the stream’s post-trunk value. Each of these is, by construction, a deterministic function of the featurizer’s outputs (e(0),ℓ(0),g(0))(e^(0), ^(0),g^(0)), which are themselves deterministic functions of the cached inputs J′J and h′h of the raw Hamiltonian (J,h)(J,h). Hence the whole pipeline up to here is a deterministic function of (J,h)(J,h) alone, and the spin configuration q coordinates have not yet entered the model. The per-site ℓ(L) ^(L) feeds the leaf builder defined immediately below, refined on the way by a leaf contextualizer that the tree readout introduces (§ S2.4); the per-bond e(L)e^(L) re-enters at the level-edge attention variant of the merge tree (§ S2.4); and the global stream continues through every later stage, refreshed at each in the same pool-and-update pattern. S2.3.1 The per-site leaf builder The leaf builder is the unique entry point for the spin configuration q: everything upstream of it is a function of the Hamiltonian alone, and everything downstream of it is a function of q as well. It acts one site at a time. Given the contextualized per-slot state ℓ^i∈ℝdℓ _i ^d_ — the trunk’s per-site outputs, refined over the readout’s slot frame by the leaf contextualizer whose construction we give in § S2.4.3, the global context g∈ℝ256g ^256, the stream’s refreshed post-context value read through the leaf’s own projection (§ S2.3), and the quaternion-embedded spin qi∈SU(2)↪ℝ4q_i (2) ^4 with ∑a(qia)2=1 _a(q_i^a)^2=1 (§ S1.1), it produces a per-site carrier ui(0)∈ℝduu_i^(0) ^d_u that encodes everything site i contributes to the wavefunction amplitude. The N carriers ui(0)i=1N\u_i^(0)\_i=1^N are the leaves of the tree readout of § S2.4, which contracts them into the single scalar logψθ(q)∈ℂ _θ(q) ; the width dud_u is set by that readout. Recall from the opening of § S2.3 that, because q enters nowhere else, the per-site antisymmetry constraint (S2.24) becomes a property of this one map. The design problem is therefore: build ui(0)u_i^(0) as a function of (ℓ^i,g,qi)( _i,g,q_i) that is expressive in all three arguments yet exactly odd under qi→−qiq_i→-q_i. Our solution keeps the two parities on separate legs and couples them one way only, which keeps the necessary properties of the sector argument we made in § S1.4, namely that the model is odd and linear in each quaternion qiq_i. Indeed, the q-leg is linear in qiq_i because it is bias-free (since a constant would be even), and the other leg is an unconstrained function of (ℓ^i,g)( _i,g). The two legs meet in a product, never in a sum, thus preserving the central oddness of the model. Figure S5 draws one site’s unit. ℓ^i∈ℝdℓ _i\!∈\!R^d_ g∈ℝdgg\!∈\!R^d_gqi∈ℝ4q_i\!∈\!R^4even gatemi=Wh[ℓ^i,g]m_i=W_h\,[ _i,\,g]mi∈ℝRm_i ^R (no qiq_i)zi=Wq→zqiz_i=W_q→ z\,q_i Vzi∈ℝRVz_i ^R ⊙ (mi⊙Vzi)U\,(m_i Vz_i)scale-norm (parity-odd)ui(0)∈ℝduu_i^(0) ^d_u, odd in qiq_i Figure S5: The per-site leaf builder (§ S2.3.1). Two input streams, one output carrier. The Hamiltonian context of per-slot ℓ^i _i and global g (teal) sets the gate mi=Wh[ℓ^i,g]m_i=W_h[ _i,g] of Eq. (S2.39), which is q-independent by construction. The spin configuration qiq_i (from MCMC) enters through the bias-free lift zi=Wq→zqiz_i=W_q→ z\,q_i (orange) and meets the gate in the rank-R Hadamard product of Eq. (S2.40) (violet), followed by the bias-free projection U and the parity-odd scale normalisation of Eq. (S2.43). The RMS normalisation is nonlinear in qiq_i, but its divided-out scale is stored as si(0)=logri(0)s_i^(0)= r_i^(0) and restored exactly at the root. Consequently ui(0)u_i^(0) is odd, while the reconstructed carrier esi(0)ui(0)=Biqie^s_i^(0)u_i^(0)=B_iq_i is exactly linear in qiq_i. Let BiasFreeLinear(din,dout)BiasFreeLinear(d_in,d_out) denote a linear map with no additive bias. The q-leg is a single bias-free lift of the quaternion into what we call the odd stream, zi:=Wq→zqi∈ℝdo,Wq→z∈BiasFreeLinear(4,do),z_i\;:=\;W_q→ z\,q_i\;∈\;R^d_o, W_q→ z (4,\,d_o), (S2.38) with dod_o the odd-stream width. The absence of a bias is not a stylistic choice: an additive constant is even under qi→−qiq_i→-q_i, and (S2.38) is the first link in a chain whose oddness must be exact — this is the bias-free constraint of § S1.4 in its simplest form. The even leg maps the per-slot and global context to a gate vector, mi:=Wh[ℓ^i,g]∈ℝR,m_i\;:=\;W_h\, [\, _i,\,g\, ]\;∈\;R^R, (S2.39) where here [⋅,⋅][·\,,·] denotes concatenation, Wh∈ℝR×(dℓ+256)W_h ^R×(d_ +256), and R is a rank hyperparameter we return to in a moment. Since ℓ^i _i and g are functions of the Hamiltonian alone (§ S2.3), carrying no q-dependence whatsoever, the gate mim_i is q-independent by construction — in particular invariant under qi→−qiq_i→-q_i. The two legs meet in a Hadamard (elementwise) product inside an R-dimensional bottleneck, followed by a bias-free output projection, ui(0)=U(mi⊙Vzi)∈ℝdu,V∈ℝR×do,U∈ℝdu×R,\;u_i^(0)\;=\;U (\,m_i Vz_i\, )\;∈\;R^d_u,\; V ^R× d_o, U ^d_u× R, (S2.40) with V and U bias-free for the same parity reason as Wq→zW_q→ z. Collecting the linear maps shows what (S2.40) really is: ui(0)=M(ℓ^i,g)zi,M(ℓ^i,g):=Udiag(mi)V∈ℝdu×do.u_i^(0)\;=\;M( _i,\,g)\;\,z_i, M( _i,\,g)\;:=\;U\,diag(m_i)\,V\;∈\;R^d_u× d_o. (S2.41) The even leg emits the weights of a linear map of rank at most R, and the spin is read out through that map. In ML terms this is a small hypernetwork [51], one network outputs the parameters of the map applied by another, with the Hadamard gate playing the role of FiLM-style multiplicative conditioning [110] in a rank-R bottleneck. In physics terms, the even sector dresses the odd sector without ever mixing into it. The same three-matrix shape U(⋅⊙V⋅)U(\,·\, V·\,) reappears as the residual hypernet at every merge node of the tree (§ S2.4.3); the leaf builder is its first and simplest instance. Finally, we normalise the carrier to unit root-mean-square while banking the divided-out magnitude in the log-scale stream. To that end, we may write Bi:=M(ℓ^i,g)Wq→z,ui,raw(0):=Biqi,B_i\;:=\;M( _i,g)\,W_q→ z, u_i,raw^(0)\;:=\;B_iq_i, (S2.42) to define ri(0):=(ui,raw(0))2¯+ε,ui(0):=ui,raw(0)ri(0),si(0):=logri(0),r_i^(0)\;:=\; (u_i,raw^(0) )^2+ , u_i^(0)\;:=\; u_i,raw^(0)r_i^(0), s_i^(0)\;:=\; r_i^(0), (S2.43) where, (ri(0))2=1duqi⊤Bi⊤Biqi+ε,x2¯:=1du∑axa2. (r_i^(0) )^2\;=\; 1d_u\,q_i B_i B_iq_i+ , x^2:= 1d_u _ax_a^2. (S2.44) The normalised carrier ui(0)u_i^(0) is odd but, by itself, nonlinear in qiq_i because its denominator is the square root of the quadratic form in Eq. (S2.44). The pair (ui(0),si(0))(u_i^(0),s_i^(0)), however, preserves the raw linear carrier exactly, esi(0)ui(0)=ui,raw(0)=Biqi.e^s_i^(0)u_i^(0)\;=\;u_i,raw^(0)\;=\;B_iq_i. (S2.45) Thus no magnitude degree of freedom is deleted at the leaf. The normalisation is a numerical gauge choice whose scale is banked at the beginning of the merge, propagated additively through the tree, and restored at the root readout. This means the non-linearity introduced by normalising cancels with the global scale added through each step of the merge, and the overall model remains multi-linear in each of the manifold’s quaternion coordinates. Next we highlight the the reason for using a product here and not a standard architectural choice of a FFN. A generic FFN on concat(ℓ^i,g,zi)concat( _i,g,z_i) would mix even and odd inputs nonlinearly, generate terms of even degree in qiq_i, and (S2.24) would hold only to the extent that training happened to learn it. The gated form (S2.41) is the simplest combiner that makes oddness structural. The q-dependence stays linear in ziz_i, with coefficients the Hamiltonian context may set freely. Hence, the even side of the model is completely unconstrained when it comes to representational capacity. Indeed, every map on the q-leg is linear and bias-free, so zi(−qi)=−zi(qi)z_i(-q_i)=-z_i(q_i), the matrix M(ℓ^i,g)M( _i,g) does not move when qiq_i flips, and the normalisation (S2.43) has an odd numerator and an even denominator. The composition is therefore odd, ui(0)(ℓ^i,g,−qi)=−ui(0)(ℓ^i,g,qi).u_i^(0)( _i,\,g,\,-q_i)\;=\;-\,u_i^(0)( _i,\,g,\,q_i). (S2.46) This parity requirement also propagates through the rest of the network. Each merge node of § S2.4 is linear in each of its two children separately, and each scale normalisation along the way flips its output whenever its input flips, so a sign flip at leaf i propagates to a sign flip of the root carrier and changes nothing else. The complex readout at the root converts that flip into logψθ↦logψθ+iπ _θ _θ+iπ. That is, ψθ↦−ψθ _θ - _θ with |ψθ|| _θ| untouched in line with (S2.24). The distinction between the normalised carrier and the represented amplitude is essential. The normalised carrier ui(0)u_i^(0) contains the even denominator of Eq. (S2.44) and is therefore not a degree-one function of qiq_i. The tree never discards that denominator however. Instead it carries si(0)=logri(0)s_i^(0)= r_i^(0), adds every later merge scale, and restores the complete scale at the root. Hence the represented leaf amplitude is esi(0)ui(0)=Biqie^s_i^(0)u_i^(0)=B_iq_i, a pure spin-12 12 object. Thus the scale normalisations do not introduce higher per-site Peter–Weyl harmonics into the final wavefunction, since their apparent nonlinear dependence cancels algebraically in the reconstructed amplitude. S2.4 Neural tensor network readout from leaves The leaf builder from § S2.3.1 produces, at every site i∈1,…,Ni∈\1,…,N\, a carrier ui(0)∈ℝduu_i^(0) ^d_u that is linear in its own site’s qiq_i and independent of every other site’s. To evaluate a wavefunction, we need to collapse these N carriers into a single complex scalar logψθ(q)∈ℂ _θ(q) , and we need the collapse to stay linear in each site separately, as multilinearity is the function class the spin-12 12 sector demands, see also (§ S2.3.1). A natural way to do so with multilinear pieces is a balanced binary tree of merges, which pairs the carriers up, merges each pair into a parent carrier through a bilinear map, then pairs the parents up, merges them in turn, and continue for K:=log2N^K:= _2 N levels until one root carrier survives, where N N is the closest power of 2 larger than N. A linear readout on the root produces the scalar. Indeed, the tensor-network community has used this concept in the contraction schedule of a tree tensor network [136]. The comparison is worth dwelling on for a moment because it reveals limitations of a binary-tree merge. Indeed, a contraction geometry is an entanglement prior, since the tree structure encodes in advance which sites can correlate cheaply (through a shallow common ancestor) and which only expensively (through many intermediate tensors), so choosing the geometry well classically requires knowing the correlation structure of the ground state one is trying to find. For a single Hamiltonian, we are able to make such a prior contraction geometry by hand, and both the Tensor Network and the Neural Quantum State literatures [23, 136] have found that a carefully chosen architecture as a training prior can outperform a generic one. However, a foundation model spanning chains, frustrated lattices, disordered ensembles, and long-range couplings cannot commit to one geometry. In this section, we detail how to circumvent the contraction flexibility issue with a single, shared tree architecture that can be cheaply reconditioned to fit the correlation structure of any Hamiltonian that is quadratic weight in the Pauli group. One merge module is shared across the entire tree, and it is modulated by Hamiltonian context; the levels exchange information laterally through attention, and the assignment of physical sites to leaves is promoted in § S2.5 to a learned, per-Hamiltonian contraction. Every node of the tree carries a state of three objects, and a fourth object lives between the nodes of each level: c∈ℝdc⏟context (even),u∈ℝdu⏟carrier (odd),s∈ℝ⏟log-scale,E(κ)∈ℝNκ×Nκ×dedge⏟level edge field, c ^d_c_context (even), u ^d_u_carrier (odd), s _log-scale, E^(κ) ^N_κ× N_κ× d_edge_level edge field, (S2.47) where Nκ:=N^/2κN_κ:= N/2^κ is the number of nodes at level κ∈0,…,Kκ∈\0,…,K\. The context c is the q-independent stream, and it tells a merge node what physical Hamiltonian it is contracting through its encoding in the trunk, see § S2.4.3. The normalised carrier u is the odd stream of § S2.3.1, together with the log-scale s it represents the reconstructed carrier a:=esu.a\;:=\;e^su. (S2.48) At an active leaf, ai(0)=Biqia_i^(0)=B_iq_i is exactly linear in its own quaternion. The scalar s starts at the leaf with si(0)=logri(0)s_i^(0)= r_i^(0) and then accumulates every magnitude divided out by the K levels of merge normalisation. The individual u stream is kept at unit RMS for numerical stability, and multilinearity is an invariant of the reconstructed stream a=esua=e^su. The edge field E(κ)E^(κ) is a coarse-grained descendant of the trunk’s per-bond data e(L)e^(L), with one feature vector per pair of level-κ subtrees. It plays the role of an effective coupling between blocks, in the spirit of Wilson renormalisation [156]. From these objects in the tree, the readout happens in five steps, one per subsection below, with the routing of §S2.5 having to its own dedicated section. We start by defining the balanced frame of slots and padding holes, then state the bilinear merge and identify the two bottlenecks it imposes (§ S2.4.1). The first bottleneck, a hidden locality prior on which leaves entangle cheaply, is loosened by the bit-reversed canonical frame (§ S2.4.2) and then we fully remove any locality priors by the learned routes of § S2.5. The second bottleneck, limited per-merge expressivity, is fixed by widening the merge to a blockwise rank-4 contraction tensor with a context-conditioned residual gate, after which the leaf contextualizer prepares the per-slot contexts and the level-00 edge field that the tree consumes at its leaves (§ S2.4.3). Level attention with a tree-aware position bias then deletes the log2N _2 N depth penalty on long-range pair correlations (§ S2.4.4), and the root readout closes the construction and states the structural invariant that the energy kernels of SM § S4 exploit (§ S2.4.5), see Figure S6. 15372648mergemergemergemergelevel attention (edge bias ++ LCA-ALiBi)mergemergelevel attention (edge bias ++ LCA-ALiBi)mergeroot readout W([RMS(eroot),croot,g])W ([\,RMS(e_root),\,c_root,\,g\,] )logψθ(q)∈ℂ _θ(q) (0)E^(0): 8×88× 8E(1)E^(1): 4×44× 4E(2)E^(2): 2×22× 2eroot=E00(3)e_root=E^(3)_00coarsen ++ 2-FWLcoarsen ++ 2-FWLcoarsen ++ 2-FWLggpool ++ updatepool ++ updatepool ++ updateκ=0κ=0 (leaves)κ=1κ=1κ=2κ=2κ=Kκ=K (root) Figure S6: The tree readout at N^=8 N=8 (§ S2.4). Orange slots carry the per-site carriers ui(0)∈ℝduu_i^(0) ^d_u in the bit-reversed canonical frame; per-slot contexts and the level-00 edge field come from the leaf contextualizer (§ S2.4.3). Violet merge nodes apply the shared three-stream merge of Figure S8 to the node states (c,u,s)(c,u,s). Lavender bands between merge levels are the level-attention insert of Eq. (S2.81) with the edge-biased LCA-ALiBi logits of Eq. (S2.84). The teal lane on the right is the edge field, coarsened level by level (Eq. (S2.77), with 2-FWL refinement); it biases the bands and, at the top, conditions the root readout of Eq. (S2.86). The magenta lane on the left is the global stream, refreshed once per level from the freshly merged contexts; its final value enters the root readout’s hypernet alongside eroote_root and crootc_root (Eq. (S2.85)). S2.4.1 The balanced frame: slots, holes, and the bilinear merge Let K be the smallest non-negative integer with 2K≥N2^K≥ N, and let N^:=2K N:=2^K, which we refer to as the N N slots for leaf positions. The physical mask m∈0,1N^m∈\0,1\ N has mt=1m_t=1 on the N slots occupied by physical sites and mt=0m_t=0 on the N^−N N-N padded slots, which we call holes. Holes carry zero couplings (Jtu=0J_tu=0 whenever mtmu=0m_tm_u=0, ht=0h_t=0) and the physical mask gates every carrier: ut(0)↦mtut(0),u_t^(0)\; \;m_t\,u_t^(0), (S2.49) so a hole can never inject amplitude into the wavefunction. The carrier mask does not send a zero hole through an ordinary bilinear contraction however. For a parent node P with child subtrees A and B, let nXn_X be the physical-leaf count of subtree X. The odd carrier merge is the masked pass-through ℳP(uA,uB):=MPT[uA,uB],nA>0,nB>0,uB,nA=0,nB>0,uA,nA>0,nB=0,0,nA=nB=0,M_P(u_A,u_B):= casesM_PT[u_A,u_B],&n_A>0,\ n_B>0,\\ u_B,&n_A=0,\ n_B>0,\\ u_A,&n_A>0,\ n_B=0,\\ 0,&n_A=n_B=0, cases (S2.50) Here T is the shared bilinear carrier contraction introduced in Eq. (S2.53) and implemented blockwise in Eq. (S2.61), and MPM_P is the parent-specific, context-dependent linear refinement defined in Eq. (S2.64). The log-scale is passed through with the live child in the one-live-child cases. Hole contexts and edge features still follow the even merge, so a hole can condition later gates without injecting an odd amplitude. A hole therefore does not annihilate a nonempty sibling subtree. Holes are not erased from the even side either. The tree treats them as structural sites in their own right. A hole slot carries a context like any other slot, with two holes still merging their contexts, and the tree is always the perfect binary tree over N N slots. This buys two things. A system’s tree is the same regardless of how far the batch later pads it, so evaluation is independent of batch composition, and a single foundation model can therefore serve many sizes of physical system with this mechanism. And since the network learns that a subtree contains holes only through the contexts prepared for it, where the holes sit becomes a structural decision the routing of § S2.5 can learn, not a convention fixed in advance. The implementation accordingly keeps two masks: the physical mask above for everything that touches amplitude, and a structural typing (a flag to mark the context features as coming from a hole) for everything that does not. These design choices are what allows the model to condition properly on arbitrary interaction topologies, including any topological defects in an otherwise idealised lattice tensor with block-diagonal J tensor. Before the trunk’s outputs can be loaded into this frame, it is worth noting that the frame already comes equipped with a natural notion of distance between slots. Every piece of machinery in this section measures position with it. For two distinct slots a,b∈0,…,N^−1a,b∈\0,…, N-1\, let lca(a,b)lca(a,b) denote the depth of their lowest common ancestor, counted from the root at depth 00. Slots a and b lie in the same level-κ subtree exactly when their binary expansions agree on all bits above position κ, since the subtrees of the balanced tree are the dyadic blocks of the slot index. The level at which two slots first meet is therefore set by their highest differing bit, which gives the common-ancestor depth a closed form via bitwise XOR, lca(a,b)=K−1−⌊log2(a⊕b)⌋,a≠b.lca(a,b)\;=\;K-1- _2(a b) , a≠ b. (S2.51) For example, at K=3K=3, slots 00 and 11 have a⊕b=1a b=1, so lca=2lca=2 and they are siblings, while slots 33 and 44 have a⊕b=7a b=7, so lca=0lca=0, and they only meet at the root, despite being index-adjacent. The example is worth pausing on: the slot metric is dyadic rather than translation-invariant. What matters is therefore the largest aligned block two slots share, counter to our intuition of how far apart their indices sit. From the common-ancestor depth we define the tree distance, dtree(a,b):= 2(K−lca(a,b)),d_tree(a,b)\;:=\;2 (K-lca(a,b) ), (S2.52) which counts edges from a up to the common ancestor and back down to b. Computationally, lca(a,b)lca(a,b) is one XOR plus one bit-scan per pair, cheaper than any learned position table and trivially fused into whatever kernel consumes it. And (S2.51) is a statement about slots, not about sites, so nothing that reassigns sites to slots, including the learned routes of § S2.5, touches it. A merge node at level κ takes two children, (cA,uA,sA)(c_A,u_A,s_A) and (cB,uB,sB)(c_B,u_B,s_B), and produces a parent, on all three per-node objects, (cP,uP,sP)(c_P,u_P,s_P), of (S2.47). The carrier half is the part that must stay multilinear, and its simplest form is one bilinear contraction through a learnable merge tensor, uPα=T[uA,uB]α=∑β,γ=1duTαβγuAβuBγ,α∈1,…,du,u_P^α\;=\;T [u_A,u_B ]^α\;=\; _β,γ=1^d_uT_αβγ\,u_A^β\,u_B^γ, α∈\1,…,d_u\, (S2.53) which is linear in each child separately. Hence a sign flip can pass through (§ S2.3.1), and it is also bilinear in the pair. The context meanwhile has no parity constraint and hence we merge bottom-up, computing a candidate update ΔP=FFNc([cA,cB,EAB(κ),EBA(κ),τP,ΦP;g]) _P\;=\;FFN_c\! ( [\,c_A,\,c_B,\,E^(κ)_AB,\,E^(κ)_BA,\, _P,\, _P;\,g\, ] ) (S2.54) and folding it into the averaged children by the normalised interpolation of Eq. (S2.33), c¯=Norm(12(cA+cB)),cP=c(c¯,ΔP),α~c=0.8sigmoid(ac),ac|init=logit(0.20.8). c=Norm\! ( 12(c_A+c_B) ), c_P=N_c( c, _P), α_c=0.8\,sigmoid(a_c), a_c |_init=logit\! ( 0.20.8 )1. (S2.55) Here [⋅][\,·\,] is concatenation as in § S2.3.1. EAB(κ)E^(κ)_AB and EBA(κ)E^(κ)_BA are the two directed edge-field entries between exactly the two subtrees being merged via the renormalised coupling of (S2.47), and encode the interaction this contraction is about to absorb, and τP _P is a fixed sinusoidal encoding [148] of the node’s level and of its dyadic position counted from the root, which we can think of as a root-centred clock. Centering the clock at the root rather than the leaves means trees of different depth agree near the top, which is one of the small choices that lets a model trained at one N N evaluate at a larger one. The retraction guarantees a context of bounded magnitude at every level, and α~c α_c sets the bounded per-channel interpolation from the averaged children toward the update. The same rule, with its own gate, reappears at the router’s tree-prefix merge (§ S2.5). The input ΦP∈ℝ32 _P ^32 is a sinusoidal encoding of the merge’s counts and of its level. With nAn_A and nBn_B the physical-leaf counts of the two children, nremn_rem the physical sites the subtree under P has not yet absorbed, κ the level and K the tree depth, ΦP=[ _P= [ sin(ωfλ),cos(ωfλ)| ( _f\,λ),\; ( _f\,λ)\; |\; (S2.56) λ∈log2(1+nA),log2(1+nB),log2(1+nrem), λ∈ \ _2(1+n_A),\, _2(1+n_B),\, _2(1+n_rem) \, ωf=π/2f,f=0,…,3]. _f=π/2^f,\;f=0,…,3\, ]. together with the same sine–cosine encoding of the pair (κ,K−1−κ)(κ,\,K-1-κ) at f=0,1f=0,1, for 24+8=3224+8=32 channels. The counts are deliberately ordered: nAn_A enters before nBn_B because the reduction path breaks the child-swap symmetry. They are also absolute rather than fractional, so identical physical content encodes identically across the sampled tree topologies of § S2.5. The level pair reports the node’s distance from the root and from the leaves. The global g in Eq. (S2.54) is the stream’s value at the level being merged, read through the tree’s projection (§ S2.3): after each level’s merges, the tree refreshes the stream once over the contexts they produced, g↦Update(g,pool(g,cP(κ)))g (g,pool(g,\c_P^(κ)\) ), pooling under the structural mask, so virtual siblings participate, and re-projects it for the next level. The final value feeds the root readout of § S2.4.5. We defer the log-scale half sPs_P to § S2.4.3, where its purpose becomes visible. Using (S2.54) allows us to separate top-down conditioning pathway from the trunk of §S2.3 into the tree itself. The contexts are computed up the tree alongside the carriers, rooted in per-slot seeds prepared from the trunk’s outputs. This is the subject of § S2.4.3, which we visit before understanding two necessary constraints that a bilinear cascade imposes in terms of expressivity of the model and a locality prior. Before that, we clarify two limitations of the bare bilinear cascade, since the next three subsections and the design choices that led to them aim to address them. First, the slot-to-site map is a hidden locality prior. Any two leaves at slots a and b can only exchange information through their lowest common ancestor in the tree, and the depth of that ancestor depends entirely on the chosen map. To couple two physical sites through a short contraction path, the map must place them in slots with a shallow common ancestor. Second, the bilinear contraction (S2.53) is either too expensive or too narrow in practise. Materialising T at the full carrier width costs du3d_u^3 parameters and flops per merge, giving over 3×1093× 10^9 at our reference width du=1536d_u=1536, which is infeasible on every merge node of the tree. On the other hand, shrinking dud_u until the merge is cheap enough introduces a bottleneck that limits the model’s capacity to pass information. And at any width, the merge at the common ancestor remains the only channel between two leaves, lca(a,b)lca(a,b) levels up. § S2.4.2 loosens the locality prior by re-indexing the frame; § S2.4.3 fixes the width; § S2.4.4 removes the depth penalty. S2.4.2 The bit-reversed canonical frame The slot metric determines the path lengths induced by the tree, while the site-to-slot assignment determines which physical pairs receive those path lengths. The naive assignment of site i at slot i−1i-1 to start at zero, inherits the dyadic metric of (S2.51) directly. This is because site pairs meet early when their zero-based indices agree on high bits, so chain neighbours that straddle an aligned block boundary (e.g. sites 44 and 55, sitting at the slots 33 and 44 of the worked example in § S2.4.1) are maximally separated, while neighbours inside a block are siblings. The prior that this encodes is therefore biasing towards aligned dyadic blocks of the site index, which is at best a rough proxy for chain locality, and meaningless for frustrated, disordered, or long-range systems. To get around this, we opt for a different canonical frame, which instead places site i at slot brevK(i−1)brev_K(i-1), where brevKbrev_K is the K-bit reversal. To that end, let br(t)∈0,1b_r(t)∈\0,1\ denote the bit of t at position r∈0,…,K−1r∈\0,…,K-1\ (least-significant at r=0r=0), so that t=∑r=0K−1br(t)⋅2rt= _r=0^K-1b_r(t)· 2^r; the reversal writes those bits in the opposite order: brevK(t):=∑r=0K−1br(t)⋅2K−1−r.brev_K(t)\;:=\; _r=0^K-1b_r(t)· 2^K-1-r. (S2.57) At K=3K=3 for example, this gives the map, 0↦0, 1↦4, 2↦2, 3↦6, 4↦1, 5↦5, 6↦3, 7↦7.0 0,\;\;1 4,\;\;2 2,\;\;3 6,\;\;4 1,\;\;5 5,\;\;6 3,\;\;7 7. (S2.58) Figure S7 compares the two placements at K=3K=3 on the same example pair. Identity placementsite i at slot i−1i-11234567815lca(site 1, site 5)=0lca(site 1, site 5)=0 (root)Bit-reversed (canonical frame)site i at slot brev3(i−1)brev_3(i-1)1537264815lca(site 1, site 5)=K−1=2lca(site 1, site 5)=K-1=2 (siblings) Figure S7: Identity placement (left) versus the bit-reversed canonical frame (right) at K=3K=3 (N^=8 N=8). Orange squares are slots carrying physical site indices; violet squares are merge nodes. The bold path traces the common-ancestor route for one representative site pair. Under the identity placement, sites 11 and 55 (index distance 44) only meet at the root; in the bit-reversed canonical frame, the same sites are siblings. The slot metric (S2.51) is identical in both panels; only the occupancy changes. Because brevK(x)⊕brevK(y)=brevK(x⊕y)brev_K(x) _K(y)=brev_K(x y), reversal swaps the roles of high and low bits in (S2.51). Hence, under the bit-reversed placement, two sites meet at the level set by the lowest differing bit of their indices. Two sites whose indices differ by 11 (say sites 11 and 22) land at slots 00 and 44, on opposite halves of the tree; two sites whose indices differ by N^/2=4 N/2=4 (say sites 11 and 55) land at slots 00 and 11, which are siblings. The canonical frame is therefore deliberately anti-local in the site index, in the sense that long aligned strides merge first, and index neighbours meet last. Two further considerations make the anti-local frame the right canonical choice. First, a foundation model’s canonical frame should not bake the nearest-neighbor-chain prior into every system. the bit-reversed frame commits the architecture to no particular correlation structure a priori; whatever locality a system needs must come through the learned routes, rather than some locality prior. The learned routes of § S2.5 condition on (J,h)(J,h) and can recover a contiguous layout when favoured by the couplings, up to automorphisms of the interaction graph. The second is bookkeeping. The packing is arranged so that a system’s N sites plus its N^−N N-N holes occupy the contiguous slot block [0,N^)[0, N), bit-reversed within that block, however far the batch later pads beyond N N. This is what makes a system’s tree the same object at every batch shape, so that evaluation across mixed sizes and size extrapolations beyond the pre-training set are well-defined. S2.4.3 Rank-4 quadrilinear merge and the log-scale chain We now fix the width problem that comes with a naive bi-linear merge (see § S2.4.1). The full-width merge tensor costs du3d_u^3 per merge, and we seek the dud_u-wide carrier. The resolution is to split the carrier into sub-blocks, contiguous slices of the carrier indexed by a counter i∈1,…,Bi∈\1,…,B\. Let B:=du/drB\;:=\;d_u\,/\,d_r (S2.59) denote the sub-block count (config-time constraint: drd_r divides dud_u), and for each carrier u∈ℝduu ^d_u define the 2-D reshape u(2D)∈ℝB×dr,ui,k(2D):=u(i−1)dr+k,u^(2D) ^B× d_r, u^(2D)_i,k\;:=\;u_(i-1)\,d_r+k, (S2.60) with first index i∈1,…,Bi∈\1,…,B\ enumerating sub-blocks and second index k∈1,…,drk∈\1,…,d_r\ enumerating channels within a sub-block. The merge tensor can then keep a narrow contraction width drd_r but increases in expressivity thanks to the sub-block axis, T∈ℝB×dr×dr×drT ^B× d_r× d_r× d_r. Contracting it with the two reshaped children produces the bilinear output, which we write u~ u; the rank-4 quadrilinear contraction then reads, u~i,j=∑k=1dr∑l=1drTi,j,k,luA(2D)uB(2D)i,k,i,li∈1,…,B,j∈1,…,dr, u_i,j\;=\; _k=1^d_r _l=1^d_rT_i,j,k,l\,u_A^(2D)_i,k\,u_B^(2D)_i,l, i∈\1,…,B\,\;j∈\1,…,d_r\, (S2.61) flattened back to u~∈ℝdu u ^d_u for the next contraction level. Each sub-block i runs an independent rank-3 trilinear sub-merge on its own slice of the children via its own slab Ti,:,:,:T_i,:,:,: of the merge tensor, and the parent carrier is the concatenation of the B sub-block outputs. The cost arithmetic is what makes this work. With B⋅dr3B· d_r^3 parameters instead of du3d_u^3, at the reference configuration (dr=32d_r=32, B=48B=48, du=1536d_u=1536) that is ∼106 \!10^6 instead of ∼3×109 \!3× 10^9, cheap enough that T is stored as a learnable tensor during pre-training. We emphasise this because it marks a deliberate design choice; the merge tensor itself is not produced by a hypernet and does not depend on the context. Instead one T, shared across every level and every position of the tree, contracts pairs of leaves to the final pair below the root. Hamiltonian dependence enters the carrier path only through the multiplicative gate we add next. Sharing one merge across the tree is the TTN analogue of weight tying [136], and it is a key step for generalisation. The parameter count is independent of N N, and a deeper tree at evaluation time reuses the same merge it trained at every level. The sub-block axis i is also not a batch dimension to sum over. 11 1 This is because every sub-block keeps its own bilinear sub-merge with its own output slice, and the natural-gradient geometry of our optimizer (see § S4.3) [91] treats the sub-block axis as part of a genuine Kronecker factorisation of the Fisher on T (a balanced two-factor matricisation over the four axes (i,j,k,l)(i,j,k,l)) rather than folding it into the batch. A context-conditioned residual gate augments the quadrilinear merge without breaking parity or multilinearity. The bilinear output u~ u of Eq. (S2.61) can be refined by a context-conditioned residual hypernet, without projecting out of the carrier width. Let R≤duR≤ d_u be a hypernet bottleneck, and let V∈ℝR×duV ^R× d_u, U∈ℝdu×RU ^d_u× R, Wh∈BiasFreeLinear(dc,R)W_h (d_c,R) be three learned linear maps. The hypernet output reads H(u~,cP):=U(Wh(cP)⊙Vu~),H( u,c_P)\;:=\;U (\,W_h(c_P) V u\, ), (S2.62) with ⊙ the elementwise (Hadamard) product. This is the same three-matrix gate as the leaf builder’s Eq. (S2.40), meaning it still preserves parity and multilinearity because the even sector modulates the odd sector multiplicatively. The map is dimension-preserving (u~↦H(u~,cP)∈ℝdu u H( u,c_P) ^d_u) and linear in u~ u with q-independent coefficients, so the parent carrier below stays bilinear in (uA,uB)(u_A,u_B). Hence the gate refines the quadrilinear merge with a parity-preserving signal about the context. For numerical stability, the output is renormalised to unit root-mean-square, with the divided-out magnitude banked in the log-scale stream , uP=uouts~,s~:=uout2¯+ε,sP=sA+sB+logs~,u_P\;=\; u_out s, s\;:=\; u_out^2+ , s_P\;=\;s_A+s_B+ s, (S2.63) where uout:=u~+H(u~,cP)u_out:= u+H( u,c_P), x2¯:=1du∑axa2 x^2:= 1d_u _ax_a^2 as in Eq. (S2.43), and ε>0 >0 is a small numerical-stability constant, see Fig. S8. With the level-zero initialisation of Eq. (S2.43), define the reconstructed carrier aX:=esXuXa_X:=e^s_Xu_X and the context-dependent linear map MP:=Idu+Udiag(Wh(cP))V.M_P:=I_d_u+U\,diag\! (W_h(c_P) )V. (S2.64) Since H(u~,cP)=Udiag(Wh(cP))Vu~H( u,c_P)=Udiag(W_h(c_P))V u, one merge satisfies aP=MPT[aA,aB].a_P=M_P\,T[a_A,a_B]. (S2.65) The coefficients of MPM_P are independent of q, so the map preserves multilinearity while the recorded log-scale cancels the numerical normalisation exactly. The factor s~ s and every sXs_X are even under a single-site sign flip, so the normalised carriers remain odd, and the reconstructed carriers remain multilinear. The accumulated scale is reattached at the root readout (§ S2.4.5). Figure S8 shows one merge node, with all three streams. cAc_AcBc_BuA∈ℝduu_A ^d_uuB∈ℝduu_B ^d_usAs_AsBs_BEAB(κ),EBA(κ)E^(κ)_AB,\,E^(κ)_BAclock τP _Pcounts ΦP _Pg (per level)context FFNnormalised interpolationgate αc _ccP∈ℝdcc_P ^d_cΔP _P12(cA+cB) 12(c_A+c_B)blockwise rank-4 merge TTliteral, sharedu~∈ℝdu u ^d_u++ residual hypernetH(u~,cP)H( u,c_P)scale-norm ÷s~ \, sunit RMSuP∈ℝduu_P ^d_ureshape B×drB\!×\!d_rreshape B×drB\!×\!d_rlog-scale updatebanks logs~ ssPs_Plogs~ s Figure S8: One merge node, blown up (§ S2.4.3); the same module is applied at every level and position of the tree. Left (magenta): the even context leg, Eqs. (S2.54)–(S2.55): an FFN over both child contexts, the two directed sibling entries of the level edge field (teal), the root-centred dyadic clock τP _P, the count features ΦP _P and the per-level-refreshed global g, folded into the averaged children by the normalised interpolation with gate α~c α_c. Centre (violet): the odd carrier leg, children reshaped to B×drB× d_r, contracted by the literal shared merge tensor T (Eq. (S2.61)), gated by the residual hypernet H(u~,cP)H( u,c_P) (Eq. (S2.62); the only place context touches the carrier), and renormalised to unit RMS. Right (grey): the log-scale leg banks logs~ s, Eq. (S2.63), so magnitudes accumulate additively instead of multiplicatively. We now have a mechanism to merge a system of N spin-1/21/2 coordinates that formulates our manifold function’s map to a single complex scalar. Our map uses one merge tensor T shared across the tree, and the same three-stream module at every level and position of the binary tree. It also has a conditioning pathway through the contexts, which are computed up the tree and can therefore carry information about the Hamiltonian and the system’s structure. With the merge machinery in hand, what remains is to feed it context from the trunk that carries the (J,h)(J,h) signal. This corresponds to formulating the bottom row of Figure S6 from the trunk’s outputs in Figure S4. Notice the frames of the trunk and readout mechanism are different. The trunk’s outputs (ℓ(L),e(L),g)( ^(L),e^(L),g) are indexed by physical sites and know nothing about slots. Meanwhile the holes have no trunk data at all, yet we know from § S2.4.1 that holes must carry contexts of their own in order for the model to carry meaningful information across different sized systems. So before the first merge runs, every slot, hole or not, needs a structural description written for it that is a function of the trunk’s output data. We construct these descriptions with a leaf-contextualizer module. As a module, the contextualizer is a single-pass bridge between the two frames, LeafContextualizer:(ℓ(L),e(L),g)⟼(ℓ^∈ℝN^×dℓ,E(0)∈ℝN^×N^×dedge,g), LeafContextualizer:\; (\, ^(L),\;e^(L),\;g\, )\; \; (\, N× d_ ,\;\;E^(0) N× N× d_edge,\;\;g\, ), (S2.66) realised as two blocks in sequence. Two copies of this stack exist, the physics leg described here and the router leg of § S2.5; each forks its own global from the same post-trunk value, and the two copies never exchange information afterwards. The per-slot output does double duty: ℓ is the state that gates the leaf builder of § S2.3.1, and, projected once where dc≠dℓd_c≠ d_ , it seeds the level-00 context array c(0)c^(0) of the merge recursion. Each block makes four updates on the running triple (c,E,g)(c,E,g): a context refresh, then a refresh of the global that reads the new contexts, then an edge rewrite that uses both, then a slot attention that reads the rewritten edges. The trunk seeds the state directly in the slot frame, with ct=ℓt(L)c_t= ^(L)_t and Ett′=ett′(L)E_t =e^(L)_t . A hole inherits whatever the masked trunk left in its rows, and the first update discards exactly this. Two further inputs appear in every step. The physical mask mtm_t of § S2.4.1 decides who keeps their seed. A fixed sinusoidal clock τt∈ℝdℓ _t ^d_ on the slot index gives the module its only notion of slot identity, every learned map here is indifferent to slot order, so position enters only through the clock and the tree-distance weights and biases below. Step 1 (context refresh — tree-weighted bond summary). Every slot first summarises its incident bonds, assigning larger weights to partners that it merges with earlier. Define the normalised tree-decay weights ωtt′=wtt′∑t′≠twtt′,wtt′=e−12(K−lca(t,t′)), _t \;=\; w_t _t ≠ tw_t , w_t \;=\;e^- 12 (K-lca(t,t ) ), (S2.67) recalling that K−lca(t,t′)K-lca(t,t ) is the level at which t and t′t merge, so siblings (K−lca=1K-lca=1) dominate while opposite halves of the tree (K−lca=K-lca=K) barely register. The bond summary and the context refresh are then e¯t e_t\; =∑t′≠tωtt′Ψs(Ett′,Et′t,sgn(t′−t)), =\; _t ≠ t _t \; ^s\! (E_t ,\,E_t t,\,sgn(t -t) ), (S2.68) Δtctx ^ctx_t\; =FFNctx(RMS[ct+τt,e¯t,1−mt]), =FFN_ctx\! (RMS[c_t+ _t, e_t,1-m_t] ), (S2.69) ct c_t\; ⟵ctx(ct,Δtctx) \,N_ctx(c_t, ^ctx_t) (S2.70) where Ψs:ℝ2dedge+1→ℝdℓ ^s:R^2d_edge+1 ^d_ is a small MLP on the two directed edges of the pair and their index order, and FFNctx:ℝ2dℓ+1→ℝdℓFFN_ctx:R^2d_ +1 ^d_ proposes the update. Step 2 (global refresh). The global then takes one update of Eq. (S2.33), pooling the just-refreshed contexts over the full slot set, g⟵Update(g,pool(g,ct))g (g,\,pool(g,\c_t\) ), holes included: after Step 1 a hole’s context carries the tree geometry it was rebuilt from, and that structural information is exactly what the global should read. Step 3 (edge rewrite — synthesise hole-incident bonds). The refreshed contexts and global then rewrite the edge tensor, Δtt′E=FFNctx−edge([mtmt′RMS(Ett′),c~t,c~t′,WEg]),Ett′⟵ctx−edge(Ett′,Δtt′E). ^E_t =FFN_ctx-edge\! ([m_tm_t RMS(E_t ), c_t, c_t ,W_Eg] ), E_t _ctx-edge(E_t , ^E_t ). (S2.71) where c~t c_t is a bias-free projection of the refreshed context’s RMS copy, WE∈ℝ64×dgW_E ^64× d_g reads the refreshed global into 64 further input channels, broadcast to every pair, and FFNctx−edgeFFN_ctx-edge maps the concatenation back to ℝdedgeR^d_edge. Unlike the trunk’s additive tap of Eq. (S2.28), the global enters here as extra channels of the input. The pair mask mtmt′m_tm_t zeroes the edge input whenever either endpoint is a hole, so a hole-incident edge is synthesised purely from its two endpoint contexts. By the time the tree runs, an edge into a hole is as meaningful as an edge into a site. Step 4 (slot attention). Steps 1 and 3 are local, in the sense that they are fixed pairwise maps with a fixed decay profile, so each block closes with the one operation that moves information across the whole slot set in a content-dependent way: a pre-norm self-attention layer. Queries, keys and values are projected from RMS(ct+τt)RMS(c_t+ _t) and split into n_ h heads of width d_ h, with the sans-serif head index of § S2.2, and the logits carry two additive biases, logit,tt′=q,t⋅k,t′d−α(K−lca(t,t′))+β(Ett′),logit_ h,\,t \;=\; q_ h,t· k_ h,t d_ h\;-\; _ h\, (K-lca(t,t ) )\;+\; _ h (E_t ), (S2.72) Δtslot ^slot_t =Attnslot(ct′,E)t, =Attn_slot(\c_t \;E)_t, ct c_t ⟵slot−attn(ct,Δtslot), _slot-attn(c_t, ^slot_t), (S2.73) Δtslot−ffn ^slot-ffn_t =FFNslot(RMS(ct)), =FFN_slot(RMS(c_t)), ct c_t ⟵slot−ffn(ct,Δtslot−ffn). _slot-ffn(c_t, ^slot-ffn_t). (S2.74) where α≥0 _ h≥ 0 is a fixed per-head slope (a scalar per head, not to be confused with the learned per-channel gates α~g α_g and α~c α_c of the global stream and the context leg), so that the heads see the tree’s geometry, and β:ℝdedge→ℝ _ h:R^d_edge is a per-head scalar projector of the freshly rewritten edge, so that the heads see the renormalised physics, again akin to Wilson’s renormalisation group. Both biases reappear in full in the level attention of § S2.4.4. The attention and FFN candidates are incorporated through the bounded updates above. After the second block, the global takes one closing update: the factored row-and-column edge reading of Eq. (S2.37), now over the rewritten edges and under the full slot mask. Holes participate here precisely because the contextualizer has just written their states, where the post-trunk reading of § S2.3 saw physical sites alone. The outputs after the second block are therefore the per-slot contexts ct(0)c_t^(0) that seed the level-00 context array, the rewritten edge tensor, which becomes E(0)E^(0), and the refreshed global. As such, this contextualizer manufactures the structural information for the holes that balance the binary tree, which is what makes hole placement a meaningful learned decision rather than dead padding and allows the model to generalise over different sized systems. With the frame loaded, the merge recursion has everything it needs at κ=0κ=0. For κ≥1κ≥ 1, the recursion additionally requires a coarse-grained edge field E(κ)E^(κ) and the attention that consumes it; these are defined in the next subsection. S2.4.4 The edge field and level attention Even with bit-reversed leaves and a quadrilinear merge, the cascade still imposes a log2N _2 N depth penalty on long-range pair correlations. This is because information between two leaves at common-ancestor depth K−lca(a,b)K-lca(a,b) has to travel through as many merge nodes as there are between it and its lca before they meet. We remove this penalty by interweaving lateral self-attention between the levels’ edge data during the tree contraction. Recall from (S2.47) that level κ carries an edge field E(κ)∈ℝNκ×Nκ×dedgeE^(κ) ^N_κ× N_κ× d_edge, one feature vector per ordered pair of level-κ subtrees, seeded at κ=0κ=0 by the leaf contextualizer’s rewritten edge tensor (§ S2.4.3). As the tree contracts, the size of the edge field must be coarse-grained so that it has the correct shape for a given level’s attention layer. Thus, when a level’s nodes are merged pairwise, the edge field is coarse-grained alongside them. The four child entries connecting the two children of P to the two children of Q, together with the four child contexts, are combined into the single parent entry, E¯PQ(κ+1) E^(κ+1)_PQ :=Norm(14∑a,b∈0,1E2P+a, 2Q+b(κ)), :=Norm\! ( 14 _a,b∈\0,1\E^(κ)_2P+a,\,2Q+b ), (S2.75) ΔPQcoarse ^coarse_PQ :=FFNcoarse([E2P,2Q(κ),E2P,2Q+1(κ),E2P+1,2Q(κ),E2P+1,2Q+1(κ),c2P(κ),c2P+1(κ),c2Q(κ),c2Q+1(κ)]), :=FFN_coarse\! ([\,E^(κ)_2P,2Q,E^(κ)_2P,2Q+1,E^(κ)_2P+1,2Q,E^(κ)_2P+1,2Q+1,c^(κ)_2P,c^(κ)_2P+1,c^(κ)_2Q,c^(κ)_2Q+1\,] ), (S2.76) EPQ(κ+1) E^(κ+1)_PQ :=coarse(E¯PQ(κ+1),ΔPQcoarse). :=N_coarse\! ( E^(κ+1)_PQ, ^coarse_PQ ). (S2.77) This update is applied at every parent pair (P,Q)(P,Q), with one FFNcoarseFFN_coarse shared across all levels, like the merge tensor itself. This again has a direct analogy to Wilson’s real space renormalisation group at the level of a block-spin step [68, 155]. Each level halves the system and rewrites the effective couplings between blocks. Despite needing no spectral rank truncation, the four child edge cells and their contexts are nevertheless compressed into one fixed-width parent cell, and the information retained is exactly what the learned FFNcoarseFFN_coarse and nGPT interpolation preserve. The result is a learned, rather than spectral, coarse-graining of the edge field across levels. The diagonal is included deliberately, since the self-edge EPP(κ)E^(κ)_P, built from the on-diagonal 2×22× 2 child block, carries the internal structure of block P (such as the frustration its subtree encloses) while the off-diagonal entries carry the renormalised couplings between blocks. After this coarsening step, the edge field is refined by the same 2-FWL path-composition feature already used in the trunk’s edge update, applying Eq. S2.36 to the merge-node indices of level κ. This is so triangle-closure information survives the blocking. Figure S9 draws one coarsening step. E(κ)E^(κ): Nκ×NκN_κ× N_κ cells(diagonal = self-edges)c2P,c2P+1c_2P,c_2P+1c2Q,c2Q+1c_2Q,c_2Q+1FFNEFFN_E2×22×2 block ++ 4 contextsE(κ+1)E^(κ+1): one cell persubtree pair (P,Q)(P,Q)2-FWL refinelevel attention Figure S9: One edge-coarsening step (§ S2.4.4). A 2×22× 2 block of child edge cells (teal, highlighted) and the four child contexts (magenta) produce one parent edge cell via Eq. (S2.77) — a learned block-spin step. Diagonal self-edges (shaded) summarise the internal structure of their own subtree. The coarsened field is refined by the tree-level 2-FWL block and then biases the level-attention logits of Eq. (S2.84). The attention layer between contraction levels is defined as follows. Let c(κ)∈ℝNκ×dcc^(κ) ^N_κ× d_c be the context array at level κ∈1,…,Kκ∈\1,…,K\ with one row per merge node, and recall Nκ=N^/2κN_κ= N/2^κ. Let m(κ)∈0,1Nκm^(κ)∈\0,1\^N_κ be the level mask that marks live merge nodes. The level-attention insert applies a pre-norm multi-head self-attention to block to c(κ)c^(κ) and applies two bounded updates, Δattn,(κ) ^attn,(κ) :=Attn(RMS(c(κ)),E(κ)), :=Attn\! (RMS(c^(κ));E^(κ) ), (S2.78) c(κ) c^(κ) ⟵m(κ)⊙level−attn(c(κ),Δattn,(κ))+(1−m(κ))⊙c(κ), m^(κ) _level-attn\! (c^(κ), ^attn,(κ) )+(1-m^(κ)) c^(κ), (S2.79) Δffn,(κ) ^ffn,(κ) :=FFNlevel(RMS(c(κ))), :=FFN_level\! (RMS(c^(κ)) ), (S2.80) c(κ) c^(κ) ⟵m(κ)⊙level−ffn(c(κ),Δffn,(κ))+(1−m(κ))⊙c(κ). m^(κ) _level-ffn\! (c^(κ), ^ffn,(κ) )+(1-m^(κ)) c^(κ). (S2.81) As per the rank-4 contraction tensor and the FFN coarse graining network, the block’s parameters are shared across all K levels. This is again to allow us to act on arbitrary system sizes, since the tree metric is self-similar in that siblings are at distance one at every level. Hence one attention layer can serve every level, and at evaluation time, levels deeper than any seen in training can be constructed. This block is our functional stand-in for the disentanglers of the multiscale entanglement renormalisation ansatz (MERA) [149]. Like them, it communicates across block boundaries at every scale, but it acts on the q-independent conditioning stream rather than on the state itself. The carrier path receives that signal only through the context-conditioned gates that modulate its multilinear coefficients. Consequently the block can alter cross-subtree correlations represented by the wavefunction without introducing nonlinear dependence on any quaternion. Whilst we want the initial attention layers here to act as identity functions, we still need to encode the tree geometry with a positional encoding. Otherwise the attention layer will treat the merge nodes at each level as an unordered set. ALiBi [116] is the simplest construction that does this for a one-dimensional sequence. For a sequence of length L and an attention head h∈1,…,Hh∈\1,…,H\, the sequence-ALiBi logit at query position i and key position j is logith,i,jseq=qh,i⋅kh,jd−αh|i−j|,logit^seq_h,i,j\;=\; q_h,i· k_h,j d\;-\; _h\, i-j , (S2.82) with αh≥0 _h≥ 0 a per-head slope drawn from a geometric schedule (116, 115 use αh=2−8h/H _h=2^-8h/H) and d the per-head dimension. Here the content term qh,i⋅kh,j/dq_h,i· k_h,j/ d scores how compatible the query at position i is with the key at position j in content space, with no positional information mixed into Q or K. Positions enter the attention map only through the additive bias term, and the projection matrices are free to encode content purely. Notice also that the geometric slope schedule spans orders of magnitude, so the heads partition naturally by attention range; steep-slope heads see a short neighbourhood, while shallow-slope heads attend nearly globally, without any learned alignment. Finally, the bias is a function of the distance |i−j| i-j , so longer sequences at inference slot into the same linear penalty, and no out-of-distribution positional embeddings are needed when generalising to larger systems. We adapt the ALiBi encoding [116] to a tree by replacing the sequence distance |i−j| i-j with the tree distance dtree=2(K−lca)d_tree=2(K-lca) of Eq. (S2.52) from § S2.4.1. This means the LCA-ALiBi logit absorbs the factor of 22 into the head slope, logith,a,btree=qh,a⋅kh,bd−αh(K−lca(a,b)),αh=32⋅2−4h/(H−1),\ logit^tree_h,a,b\;=\; q_h,a· k_h,b d\;-\; _h\, (K-lca(a,b) ), _h\;=\; 32· 2^-4h/(H-1), (S2.83) where h∈0,…,H−1h∈\0,…,H-1\. We set the slopes to fixed constants spanning the geometric ladder from 1.51.5 down to ≈0.09≈ 0.09 (tuned empirically), so the heads partition by tree-distance band exactly as in the sequence case, with zero learnable positional parameters. The bias depends on the node pair only through a⊕ba b (via lcalca, Eq. S2.51). This means we have only one XOR plus one bit-scan per pair to compute, and is the same function at every level, meaning the computation is self-similar across the tree and we can share the block across different layers. Hence the same position encoding drives the route policy of § S2.5, and we will frequently refer back to it in § S2.5. Finally, the coarsened edge field enters the logits as a learned additive bias alongside the tree-distance term, logith,a,btree, edge=qh,a⋅kh,bd−αh(K−lca(a,b))+βh(Eab(κ)),logit^tree, edge_h,a,b\;=\; q_h,a· k_h,b d\;-\; _h\,(K-lca(a,b))\;+\; _h\! (E^(κ)_ab ), (S2.84) with βh:ℝdedge→ℝ _h:R^d_edge a small per-head scalar projector, scaled by the same 1/d1/ d as the content term so that changing the head width does not reweight geometry against content. Thus the heads see where two blocks sit in the tree and how strongly the renormalised Hamiltonian couples them. S2.4.5 Root readout and the structural invariant After K levels of merges and level attention, one node remains carrying the root state (croot,uroot,sroot)(c_root,\,u_root,\,s_root) together with the top of the edge stream eroot:=E00(K)∈ℝdedgee_root:=E^(K)_00 ^d_edge as a fully renormalised state of the whole system’s coupling structure. The readout maps the root carrier to a pair of real numbers through a hypernet-generated readout matrix W∈ℝ2×duW ^2× d_u, produced from the root’s full even state, W=W([RMS(eroot),croot,g]),W\;=\;W\! ( [\,RMS(e_root),\;c_root,\;g\, ] ), (S2.85) by the same low-rank gate22 2 its third and final appearance as Eqs. (S2.40) and (S2.62): the renormalised coupling structure eroote_root, the root context crootc_root, and the global stream’s final value, refreshed through every level of the tree and read through the root’s own projection (§ S2.4.1). This allows us to assemble a complex log-amplitude from that pair along with the banked log-scale, (ψ1,ψ2)=Wuroot∈ℝ2,logψθ(q)=12log(ψ12+ψ22)+sroot⏟log|ψθ|+iatan2(ψ2,ψ1).( _1, _2)\;=\;W\,u_root\;∈\;R^2, _θ(q)\;=\; 12\, ( _1^2+ _2^2 )+s_root_ _θ \;+\;i\,atan2( _2, _1). (S2.86) Here atan2atan2 is the two-argument arctangent, giving the usual phase of the complex scalar ψ1+iψ2 _1+i _2 in (−π,π](-π,π]. Hence, decomposition is exactly ψθ=esroot(ψ1+iψ2) _θ=e^s_root( _1+i _2), and the pair is the amplitude in Cartesian form at unit scale, with sroots_root re-attaching every magnitude the scale normalisations banked on the way up (§ S2.4.3). We can formalise this by the following lemma. Lemma S2.1 (Restored-scale multilinearity). For every active site i, let ai(0):=esi(0)ui(0)a_i^(0):=e^s_i^(0)u_i^(0). Under the leaf initialisation (S2.43), the bilinear merge (S2.65), and the root readout (S2.86), the represented amplitude ψθ(q1,…,qN) _θ(q_1,…,q_N) is degree one in each quaternion qiq_i separately. Proof. At a leaf, Eq. (S2.45) gives ai(0)=Biqia_i^(0)=B_iq_i, with BiB_i independent of every quaternion. Suppose inductively that aAa_A and aBa_B are multilinear in the active leaves of two disjoint subtrees A and B. When both children contain physical leaves, the map MPTM_PT is bilinear; when exactly one is empty, Eq. (S2.50) passes the other child through linearly. Induction over the nonempty leaves therefore preserves degree one in every physical quaternion. Induction up the balanced tree gives a multilinear aroota_root. Finally, if W1,W2W_1,W_2 are the two rows of the root readout, then ψθ=esroot(ψ1+iψ2)=(W1+iW2)aroot, _θ\;=\;e^s_root( _1+i _2)\;=\;(W_1+iW_2)a_root, (S2.87) which is a linear functional of a multilinear carrier. Therefore ψθ _θ is degree one in each qiq_i. ∎ One might ask why the readout does not simply emit logψθ _θ as a linear function of urootu_root, which looks more direct. However we emphasise that this would violate the central oddness of our phase space since (S2.24) demands that a single-site flip send ψθ→−ψθ _θ→- _θ, i.e. logψθ→logψθ+iπ _θ→ _θ+iπ. Notice that this is an affine shift at the level of the logarithm, not a sign flip. Eq. (S2.86) is built so that the linearity sits exactly there, and thus, qi→−qi⟹ui(0)→−ui(0)⟹uroot→−uroot⟹(ψ1,ψ2)→−(ψ1,ψ2)⟹ψθ→−ψθ.q_i→-q_i\;\; \;\;u_i^(0)→-u_i^(0)\;\; \;\;u_root→-u_root\;\; \;\;( _1, _2)→-( _1, _2)\;\; \;\; _θ→- _θ. (S2.88) The first arrow is the leaf oddness (S2.46). The second holds because each merge is linear in each child and its normalisation is parity-odd, so the sign flip propagates up the unique root-ward path through exactly one node per level. Every log-scale s~ s is even and stays put, and nothing else can change. The third arrow holds because W is a function of (eroot,croot,g)(e_root,c_root,g), all even, q-independent quantities. And the fourth is (S2.86): ψ12+ψ22 _1^2+ _2^2 is invariant, so log|ψθ| _θ does not move, while the atan2atan2 picks up exactly π. Per-site oddness therefore holds at every parameter value of our model, and the symmetry is a property of the architecture itself. Figure S6 draws the full readout of this section, from the initial bit-reversed canonical frame to the root readout on top. Notice that many different shapes and tensor ranks flow through it — dud_u-wide carriers, scalar log-scales, the complex pair at the root — yet every q-dependent intermediate, irrespective of its overall dimension, shares one structure on its two leading axes: at level κ it splits into N^/2κ N/2^κ independent chunks, one per node, and the chunk at node P depends on q only through the 3⋅2κ3· 2^κ Lie coordinates (§ S1.2) of P’s own subtree. The merge is the only operation that grows this dependence set, uniting two subtrees into one; everything else in the readout leaves it untouched. This is what will allow the energy kernel of § S4.2 to propagate each chunk’s derivatives along its own 3⋅2κ3· 2^κ coordinates and no others, bringing the forward-mode Laplacian to (N^2)O( N^2) instead of the (N^3)O( N^3) a generic implementation would pay. S2.5 Deep reinforcement learning of leaf permutations In this section we detail our deep reinforcement learning strategy to learn the leaf-to-site map, which we call the router. We will not introduce deep reinforcement learning here from first principles; see [141] for a comprehensive introduction. The bit-reversed leaf assignment of § S2.4.2 fixes the address geometry, but the map from physical sites to leaf positions is currently the same irrespective of the Hamiltonian interaction data. Indeed, site i goes to leaf brevK(i−1)brev_K(i-1), regardless of (J,h)(J,h). On systems whose correlation structure aligns with the bit-reversed prior (mainly the short-range and translation-invariant ones), a static assignment is fine. However, for systems whose correlation structure aligns with no fixed assignment, such as frustrated couplings, quenched disorder, topologically non-trivial geometry, or long-range tails, we want the leaf-to-site map to be learnable and to condition on the (J,h)(J,h) tensors of the Hamiltonian. We let the leaf-to-site map be a discrete latent variable, and define route as a permutation of the whole slot block, p∈SN^,p\;∈\;S_ N, (S2.89) where SN^S_ N is the symmetric group on the N N slots, containing physical sites and holes on an equal footing. We emphasise that the holes move, per our earlier discussion on scaling to unseen system sizes. The natural alternative is to constrain routes to be mask-preserving, mpt=mtm_p_t=m_t, so that padding stays frozen at its canonical slots. But § S2.4.1 made hole placement informative, since a hole carries a context, and where the holes sit shapes every context merge above them to the root readout. The router therefore frees the holes and decodes their positions with the same care as the sites’. At decode step t=0t=0 we restrict the candidate set to physical sites. This is a policy-support convention, it anchors the autoregressive prefix in a physical token before the hole-placement statistics are updated. Freed holes do cost an (N^−N)!( N-N)!-fold exchange redundancy, as routes that differ only by which hole went where produce the identical routed system. However, our sampler detailed below removes this redundancy exactly at every decode step, by collapsing the interchangeable holes into a single candidate class. Given a route p, every per-site and per-bond tensor is relabelled into the route’s frame (including masks for the holes), qt′:=qpt,ht′:=hpt,Jtu′:=Jpt,pu,ℓt′:=ℓpt(L),etu′:=ept,pu(L),mt′:=mpt.q _t:=q_p_t, h _t:=h_p_t, J _tu:=J_p_t,p_u, _t:= ^(L)_p_t, e _tu:=e^(L)_p_t,p_u, m _t:=m_p_t. (S2.90) Indeed, under mask-preserving routes the mask was a constant of the system, but once real and virtual positions mix, which slots are real is itself route-dependent, and the energy kernel, the sampler, and the carrier gating of Eq. (S2.49) all read the routed mask m′m . Inside the routed frame everything downstream (the leaf contextualizer, leaf builder, merge tree, and root readout) operates on an ordinary already-routed system. Thus each kernel sees (q′,h′,ℓ′,e′,m′)(q ,h , ,e ,m ) exactly as if the permutation had been applied at the input, and the route is invisible to the conditional ansatz beyond this relabelling. Let πθ(p∣h,ℓ(L),e(L),m) _θ(p h, ^(L),e^(L),m) denote a Hamiltonian-conditioned route policy (parametrised below), and let ψθ(⋅∣p) _θ(· p) be the routed ansatz of § S2.3 and § S2.4 evaluated in the p-frame. The policy returns probabilities, not amplitudes, and a route is drawn afresh at each evaluation, so the object the model represents is not a single wavefunction but an incoherent convex mixture of the routed states, that is, the density matrix ρθ(q,q′)=∑p∈SN^πθ(p∣h,ℓ(L),e(L),m)ψθ(q∣p)ψθ∗(q′∣p)Zp, _θ(q,q )\;=\; _p\,∈\,S_ N _θ(p h, ^(L),e^(L),m)\, _θ(q p)\, _θ^*(q p)Z_p, (S2.91) with Zp=∫|ψθ(q∣p)|2qZ_p= | _θ(q p)|^2\,dq the squared norm of the routed state. The parameters θ are shared between the policy and the conditional ansatz. Its energy is the matching convex combination of the routed states’ Rayleigh quotients, a variational principle over mixed states, but a benign one: the energy is linear in ρθ _θ, so its minimum over the convex set of density matrices is attained at an extreme point, a pure state, namely the ground state. The relaxation therefore costs nothing, and at the optimum ρθ _θ is pure even where the policy keeps its weight spread over routes that realise one and the same physical state. Conditioning only on the trunk’s outputs (h,ℓ(L),e(L))(h, ^(L),e^(L)) and the slot mask m keeps the learned raw scoring map permutation-equivariant under relabelling of the slot block. The minimum-index representative retained after quotienting fixes a bookkeeping gauge inside each WL class, so the labelled representative route itself need not transform equivariantly; class probabilities and the routed physical state remain invariant when the classes are genuine symmetry orbits. Since the policy never sees q, each term of the mixture preserves the per-site oddness of § S2.3.1. The rest of this section sets up the architecture of the policy, and the machinery to train it with policy gradients via a discrete sampling step, a score-function estimator, and some variance-control analysis. We conclude by characterising optimal symmetrised policies through the automorphism orbits of the routed problem. A policy supported uniformly on one energy-minimising orbit O has entropy log|| |O|, which is generally far below log(N^!) ( N!). This orbit construction, establishes that a non-uniform optimum exists beyond the guarantees of having a lower bound below the maximum. The sum in Eq. (S2.91) contains N^! N! routes and is therefore estimated by sampling. Tractable ancestral sampling and exact log-probabilities come from factorising the policy autoregressively into N N masked slot picks. Differentiation through the discrete draw is handled separately by the score-function estimator of § S2.5.3. To that end, let p<t:=(p0,p1,…,pt−1)p_<t:=(p_0,p_1,…,p_t-1) denote the prefix of picks taken before step t, and let p≤t:=(p0,p1,…,pt)p_≤ t:=(p_0,p_1,…,p_t). Then πθ(p∣h,ℓ(L),e(L),m)=πθ(p0∣h,ℓ(L),e(L),m)∏t=1N^−1πθ(pt∣p<t,h,ℓ(L),e(L),m), _θ(p h, ^(L),e^(L),m)\;=\; _θ(p_0 h, ^(L),e^(L),m)\, _t=1 N-1 _θ(p_t p_<t,\,h, ^(L),e^(L),m), (S2.92) where this runs over the full slot block, and holes are placed by the same factors that place sites. The first pick anchors the tree, with a physical site at slot p0p_0. If 0realQ_0 real denotes the WL classes after restricting the empty-prefix candidate set to physical sites and r(C)r(C) is the retained representative of class C, then πθ(p0=r(C)∣h,ℓ(L),e(L),m)=exp(η~0,C)∑D∈0realexp(η~0,D), _θ(p_0=r(C) h, ^(L),e^(L),m)\;=\; ( η_0,C ) _D _0 real ( η_0,D ), (S2.93) where η~ η is the class logit of Eq. (S2.109). We emphasise that the anchor is learned, not uniform. This is because which site the contraction is built around is itself part of the structure the policy should discover. The raw logits at all N N positions are Hamiltonian-conditioned; the final factor has one available class and therefore contributes the identically zero log-probability and gradient. S2.5.1 Hamiltonian-conditioned scoring row Before quotienting, each factor in Eq. (S2.92) has raw pointer logits ηt∈ℝN _t N over the candidate slots, with the empty prefix at t=0t=0; the actual factor is the softmax over the resulting WL-class logits η~t η_t. The decoder is a pointer network in the sense of 150, with one structural difference: a standard pointer decoder represents the placed prefix as a flat sequence, whereas ours materialises the partial merge tree implied by the picks made so far. The logits are produced by a candidate composer, a tree-prefix encoder, a tied dot-product pointer, and a conditional symmetry quotient. We describe them in turn. (a) Candidate composer. Let Pℓ:ℝdℓ→ℝdhP_ :R^d_ ^d_h be a small bias-free linear projector on per-site features from the trunk, as in § S2.2, and let ϕpre,ϕsuf:ℝ2de→ℝdh _pre, _suf:R^2d_e ^d_h be independent per-channel SiLU MLPs. Both read the two directed bonds of a pair through Eiu:=[ei,u(L),eu,i(L)]E_iu:=[e^(L)_i,u,e^(L)_u,i]. Define the projected node embedding at slot i, nodei:=Pℓ(ℓi(L)).node_i\;:=\;P_ ( ^(L)_i). (S2.94) Let Ut:=0,…,N^−1∖p<tU_t:=\0,…, N-1\ p_<t be the available candidate set, and write g¯=Wglobalg+bglobal g=W_globalg+b_global for the projection of the fixed forked router global. Define Pt,i:=1max(t,1)∑u∈p<tϕpre(Eiu),St,i:=1max(|Ut|−1,1)∑v∈Ut∖iϕsuf(Eiv).P_t,i:= 1 (t,1) _u∈ p_<t _pre(E_iu), S_t,i:= 1 (|U_t|-1,1) _v∈ U_t \i\ _suf(E_iv). (S2.95) The implemented composer does not first make a static vocabulary token. Let Tg:ℝdg→ℝ256T_g:R^d_g ^256 denote the active bottleneck on the raw-global channel, and write the router’s stabilized RMS map as RMSr(z):=γrms⊙zmaxz2¯,10−2.RMS_ r(z):= _ rms z \ z^2,10^-2\. (S2.96) Here every occurrence has its own learned scale γrms _ rms. The composer RMS-normalises each learned vector stream, leaves the three hole ratios raw, and adds learned affine projections into one row–candidate accumulator, ai(t)= a_i^(t)= b+WℓRMSr(ℓi(L))+WgRMSr(Tgg)+Wg¯RMSr(g¯+γ(t)) b+W_ RMS_ r( _i^(L))+W_gRMS_ r(T_gg)+W_ gRMS_ r( g+γ(t)) (S2.97) +WPRMSr(Pt,i)+WORMSr(Ot,i)+WSRMSr(St,i) +W_PRMS_ r(P_t,i)+W_ORMS_ r(O_t,i)+W_SRMS_ r(S_t,i) +WHRMSr(Ht,i)+Wρρt,i. +W_HRMS_ r(H_t,i)+W_ρ _t,i. where Ot,iO_t,i is the order-aware prefix stream described below, Ht,iH_t,i is the ordered hole-prefix stream, and ρt,i _t,i collects the three remaining-hole ratios. In the reported model one residual candidate FFN acts on ai(t)a_i^(t); an output projection is then added to nodeinode_i to give xi(t)∈ℝdhx_i^(t) ^d_h. Thus g and g¯ g are fixed across decode rows, while the root-centred clock γ(t)γ(t) tells the composer how deep into the sequence it is. The prefix summary Pt,iP_t,i runs over already-picked slots and the candidate-specific suffix St,iS_t,i runs over every other available slot. Both can be maintained as cumulative scans, so each append at step t costs (N^dh)O( Nd_h) work to update the prefix and suffix summaries at every candidate. Across a complete route this bookkeeping costs (N^2dh)O( N^2d_h). The dense candidate refinement below contributes a cubic core; sampled tree recomputation and the conditional quotient are accounted for separately. The order-aware stream Ot,iO_t,i is a second read of the prefix: the candidate-specific edge message from position s<ts<t is weighted by a learned Gaussian filterbank over the dyadic LCA distance between decode positions t and s, so that when a neighbour was placed matters as well as whether. The hole streams track both how many holes remain and where they have been going. Every one of these quantities depends on the current row t; there is no candidate embedding that is reused unchanged throughout a route. (b) Tree-prefix encoder. The structural heart of the decoder is that the placed prefix is not summarised as a flat sequence; it is materialised as the thing it really is, a partial merge tree. A learned dyadic merge, whose inputs and combination rule (children, directed sibling edges, dyadic clocks, and the normalised interpolation of Eq. (S2.55) with its own gate αr _r) deliberately mirror the physical merge of § S2.4.3, scans the placed tokens bottom-up, with causal per-level 2-FWL refinement and causal level attention, the machinery of § S2.4.4 restricted to the prefix. The prefix [0,t)[0,t) is then represented by its dyadic cover: the at most log2N _2 N complete subtrees given by the binary decomposition of t. Cover segments self-attend, and the candidate tokens cross-attend to the cover through their own pooled candidate-to-segment edges. Four read-only cover-query blocks produce the decode state yt∈ℝdhy_t ^d_h; two candidate-to-cover blocks and four post-prefix blocks refine the candidate states before the pointer is applied. The decode state is, in other words, a learned surrogate of the partial wavefunction contraction that the next choice is about to extend. Positional information enters only through structure-aware signals, decode-position clocks, dyadic segment clocks, and the tree-distance attention biases of Eq. (S2.83), and never through absolute slot identities, which is what keeps the raw scoring map equivariant under relabelling of the slot block before the representative gauge is fixed. The router’s global has two distinct roles. It owns the second copy of the contextualizer stack of § S2.4.3: this copy forks from the same post-trunk value as the physics leg, refreshes through that stack’s per-block updates, and closes with the factored edge reading of Eq. (S2.37) over the route-refined edges. The resulting fixed fork g supplies the two composer channels in Eq. (S2.97). A projection of the same fixed g enters every selected-leaf merge and every cover-query FFN; the same-level 2-FWL and attention refinements receive no additional global tap. Only after the prefix cover has been formed does the stream acquire a row-specific value, g(t):=g,t=0,Update(g,pool(g,t)),t>0,g^(t)\;:=\; casesg,&t=0,\\ Update\! (g,\;pool (g,C_t ) ),&t>0, cases (S2.98) one update of Eq. (S2.33) reading the complete-subtree roots in tC_t; at an empty cover the learned update is masked out and the global remains unchanged. This refreshed value enters only the FFNs of the two candidate-to-cover blocks. It is not fed back into the composer, the query seed, the tree merge, the cover-query blocks, or the four post-prefix blocks. Each position computes Eq. (S2.98) independently from the same fixed g, so teacher scoring of a stored route and sequential sampling of a fresh one evaluate the same row map (up to floating-point evaluation order). The attention pattern is defined explicitly as follows. For a query sequence A=(ai)A=(a_i), a key–value sequence B=(bj)B=(b_j), write a^i=RMSr(ai) a_i=RMS_ r(a_i) and b^j=RMSr(bj) b_j=RMS_ r(b_j) for the block’s tokenwise pre-norms. For an additive structural bias βh _h, a per-head key dimension dkd_k, and a mask M∈0,−∞|A|×|B|M∈\0,-∞\^|A|×|B|, write Attnh(A,B,βh,M):=softmax((A^WQh)(B^WKh)dk+βh+M)B^WVh.Attn_h(A,B; _h,M):=softmax\! ( ( AW_Q^h)( BW_K^h) T d_k+ _h+M ) BW_V^h. (S2.99) Equation (S2.99) displays one head; in the residual equations below, AttnAttn denotes the implemented gated multihead collapse followed by the learned bias-free output projection WOW_O. All projections are learned separately in each block. The important object retained at decode position s is not a static embedding of the selected site, but the raw row-conditioned candidate state v~s v_s. The active tree normalisation maps it through the floored RMS sphere transform before the first merge, v~s:=xps(s),vs:=Sphere(v~s):=v~smaxv~s2¯,10−4. v_s:=x_p_s^(s), v_s:=Sphere( v_s):= v_s \ v_s^2,10^-4\. (S2.100) Thus a tree leaf already records the Hamiltonian, the prefix and suffix seen when it was selected, the hole statistics, and the decode clock. The states v0,…,vt−1v_0,…,v_t-1 are merged bottom-up over aligned dyadic intervals. After every merge level, ordered subtree-pair features undergo a row-causal 2-FWL update, and roots of the same dyadic size attend one another using Eq. (S2.99) with the refined pair features as β and the mask Mab=0M_ab=0 for b≤ab≤ a and −∞-∞ otherwise. This is a same-scale causal refinement, since information can move between complete subtrees without flattening their internal contractions. Let the binary expansion of t decompose the prefix into its unique set of maximal aligned intervals, [0,t)=⨆a=1ctIt,a,t:=z(It,a)a=1ct,ct=popcount(t)≤⌈log2N^⌉,[0,t)= _a=1^c_tI_t,a, _t:=\z(I_t,a)\_a=1^c_t, c_t=popcount(t)≤ _2 N , (S2.101) where z(I)z(I) is the root state of the complete subtree on I. The fresh query seed rt(0)=g¯+γ(t)r_t^(0)= g+γ(t), with γ(t)γ(t) the root-centred decode clock, is appended to this cover. Four edge-biased cover blocks then apply Eq. (S2.99) with a deliberately asymmetric mask: every live cover root may read all live cover roots, and the appended query may read the roots and itself, but no root may read the appended query. Padding is masked in both directions. The cover roots generally have different dyadic sizes; their all-to-all interaction is causal because every member of tC_t lies strictly inside the already selected prefix. The resulting query yt=CoverReadθ(rt(0),t,Etcover)y_t=CoverRead_θ (r_t^(0),C_t,E_t^cover ) (S2.102) where EtcoverE_t^cover is the directed root-pair message tensor formed by pooling the physical pair features over the leaves of each ordered pair of cover segments. Thus yty_t is a read-only observation of every complete subtree that exactly tiles the current prefix. At t=0t=0 the cover is empty and y0y_0 is the learned four-block transform of the global–clock seed, which attends itself. In particular, yty_t is rebuilt for every t from the complete-subtree roots in that row’s cover. First, two blocks let each xi(t)x_i^(t), i∈Uti∈ U_t, cross-attend to tC_t; their biases pool the directed candidate–subtree edges over the leaves of each cover segment, and their FFNs alone receive the refreshed g(t)g^(t) of Eq. (S2.98). Four post-prefix blocks then repeat the ordered update. Let Xt(0)X_t^(0) be the output of the second candidate-to-cover block, set Lpost=4L_post=4, and write λpost:=Lpost−1/2 _post:=L_post^-1/2. Then Ht(r) H_t^(r) =Xt(r)+λpostAttnhist(r)(Xt(r),Y<t,βthist,Mthist), =X_t^(r)+ _postAttn^(r)_hist (X_t^(r),Y_<t;β^hist_t,M^hist_t ), (S2.103) Mt,shist M^hist_t,s =0,s<t,−∞,s≥t, = cases0,&s<t,\\ -∞,&s≥ t, cases (S2.104) St(r) S_t^(r) =Ht(r)+λpostAttnself(r)(Ht(r),Ht(r),βtcand,MUt), =H_t^(r)+ _postAttn^(r)_self (H_t^(r),H_t^(r);β^cand_t,M^U_t ), (S2.105) Xt(r+1) X_t^(r+1) =St(r)+λpostFFNr(RMSr(St(r))),r=0,…,Lpost−1. =S_t^(r)+ _postFFN_r (RMS_ r(S_t^(r)) ), r=0,…,L_post-1. (S2.106) We set Attn(A,∅,β,M):=0Attn(A, ;β,M):=0. Hence at t=0t=0 the candidate-to-cover and history-attention residual sublayers are identities, while their FFNs and the remaining-candidate self-attention still run. Here Y<t=(y0,…,yt−1)Y_<t=(y_0,…,y_t-1); the strict history mask excludes the current and future queries, while MUtM^U_t permits all-to-all attention among the still-available candidates and masks every placed slot. The history bias contains the direct bidirectional edge pair between candidate i and the site selected at history position s, together with a dyadic-LCA term; βtcand _t^cand contains the candidate-pair edge features. Consequently, before the pointer chooses one site, the candidates may compare themselves conditional on the same prefix and on one another. Writing x¯i(t):=Xt,i(Lpost) x_i^(t):=X_t,i^(L_post), the pointer of Eq. (S2.107) pairs yty_t with these refined candidate states. For Hamilton-Zero, the active depths are four cover-query blocks, two candidate-to-cover blocks, and four post-prefix blocks, at width dh=512d_h=512 with 1616 logical heads and a 512512-dimensional pointer score. Each repeated stack uses its own depth gain: λcover=4−1/2=1/2 _cover=4^-1/2=1/2, λcand=2−1/2=1/2 _cand=2^-1/2=1/ 2, and λpost=4−1/2=1/2 _post=4^-1/2=1/2. There is no tokenwise causal backbone and no key–value attention cache in this model: previous ysy_s are retained as states and their keys and values are projected afresh in each post block. Only the support of the prefix query is logarithmic in N N; the candidate-to-cover and remaining-candidate stages cost (N^logN^)O( N N) and (N^2)O( N^2) per dense row, respectively. Across all N N rows, the latter makes the dense attention core (N^3)O( N^3); the conditional WL quotient is accounted for separately. Figure S10 shows the resulting computation. multiscale prefix query at t=7t=7selected prefix p<7p_<7remaining candidates i∈U7i∈ U_7v0v_0v1v_1v2v_2v3v_3v4v_4v5v_5v6v_6 z[0,4)z[0,4)z[4,6)z[4,6)v6v_6xi1x_i_1xi2x_i_2xi3x_i_3xi4x_i_4xi5x_i_5xi6x_i_6xi7x_i_7xi8x_i_8xi9x_i_9xi10x_i_10 dyadic cover 7C_7 and conditioned seed g¯+γ(7) g+γ(7)read-only cover query ×4× 4the query reads all three roots; the roots are not updatedy7y_7retained for later positionsfresh candidate row xi(7):i∈U7\x_i^(7):i∈ U_7\candidate–cover cross-attention ×2× 2candidates query 7C_7; pooled edge bias; FFNs receive g(7)g^(7)history cross-attention to the earlier query states y<7y_<7suffix self-attention among the remaining candidatesconditioned FFN post-prefixblock ×4× 4refined candidates x¯i(7) x_i^(7)tied dot-product pointerscore y7y_7 against every refined candidateconditional symmetry quotientcollapse the remaining candidates into WL classessample among class representatives (Gumbel pick) Figure S10: The multiscale prefix query at t=7t=7. The selected prefix p<7p_<7 is decomposed into its canonical dyadic cover 7=z[0,4),z[4,6),v6C_7=\z[0,4),z[4,6),v_6\, comprising complete subtrees of four, two, and one leaves. Same-level causal 2-FWL and attention refine the subtree roots during their construction. Four read-only query blocks attend from the conditioned seed g¯+γ(7) g+γ(7) to the three cover roots, without updating those roots, and produce the query state y7y_7, which is retained for later decoding positions. In parallel, each remaining candidate receives a fresh prefix-dependent embedding xi(7)x_i^(7). Two candidate–cover cross-attention blocks condition these embeddings on 7C_7, using the pooled edge bias and the routed global state g(7)g^(7). Four post-prefix blocks then apply history cross-attention to y<7y_<7, suffix self-attention among the remaining candidates, and a conditioned FFN, producing x¯i(7) x_i^(7). The tied pointer scores y7y_7 against every refined candidate; the conditional symmetry quotient collapses equivalent candidates into Weisfeiler–Leman classes before the Gumbel pick. Here |7|=popcount(7)=3|C_7|=popcount(7)=3, so only the support of the read-only prefix query grows logarithmically with prefix length. The two router refinement residuals retain their ordinary pre-norm skip connections. Every named update in the tree and global streams uses Eq. (S2.32), including the route-global pool update of Eq. (S2.98). (c) Pointer. [150] Let Wq,Wk∈ℝdh×dhW_q,W_k ^d_h× d_h be a single pair of bias-free maps shared across all decode steps, and let τ>0τ>0 be a temperature parameter that sets the exploration scale of the policy. The per-step logit at candidate slot i is the tied dot product of the projected decode state against the projected candidate token, ηt,i=1τ(Wqyt)⋅(Wkx¯i(t))dh+maskt,i, _t,i\;=\; 1τ\, (W_q\,y_t )· (W_k\, x_i^(t) ) d_h\;+\;mask_t,i, (S2.107) with the availability mask maskt,i=−∞if i∈p<t, or if t=0 and mi=0,0otherwise.mask_t,i\;=\; cases-∞&if i∈ p_<t, or if t=0 and m_i=0,\\ 0&otherwise. cases (S2.108) The availability mask first removes placed slots, holes remain eligible, except at the physical anchor step, and Eq. (S2.109) then converts the surviving raw pointer logits into the class logits whose softmax is the actual pick distribution. The query projection WqW_q is initialised at zero, so the raw pointer logits start equal. After the mean quotient of Eq. (S2.109), routing therefore starts uniform over valid WL classes (not necessarily over individual slots), and every early preference is learned rather than inherited from the initialisation. Figure S11 shows the complete pipeline. (d) Symmetry quotient. Before any sampling, each step’s logits are collapsed over WL colour classes of the remaining candidates. The pair refinement is the second-order update the trunk’s edge block already realises (Eqs. (S2.35)–(S2.36)): colours live on ordered pairs of vertices of the coloured coupling graph and update from the multiset of two-leg paths through every third vertex, run here with the placed prefix individualised, and vertex colours are read off the diagonal. Prefix-preserving automorphisms leave these colours invariant, so candidates in the same exact orbit cannot be separated; the converse is not a theorem for 2-WL on arbitrary graphs. At the empty prefix, the converged colour classes matched the exact orbits on all 247247 verifiable systems in the checked panel; prefix-conditioned checks additionally matched on a 2020-system chain-and-ladder spot sample. The per-system dispatch uses the cheaper first-order refinement when its empty-prefix partition agrees with the pair refinement and the system metadata permits it, and uses the pair version elsewhere. Let tQ_t be the resulting partition of UtU_t, and let Ct(i)∈tC_t(i) _t contain candidate i. Each class keeps a single representative whose logit averages its members, η~t,C=log1|C|∑i∈Ceηt,i, η_t,C\;=\; 1 C _i∈ Ce _t,i, (S2.109) and the step samples among representatives only (a Gumbel-max draw [50], which is exact ancestral sampling). The mean, rather than a sum, is what stops a class’s probability from scaling with its class size. All hole slots share one class at every step, so the (N^−N)!( N-N)! hole-exchange redundancy of Eq. (S2.89) is removed here. The same collapse is applied when scoring, so the factors of Eq. (S2.113) are understood over the collapsed logits. ℓ(L),e(L),g;p<t,v~<t,y<t ^(L),\,e^(L),\,g\ ; 18.49988ptp_<t,\, v_<t,\,y_<tfixed Hamiltonian fork and current prefix state Dynamic candidate composer streamwise RMS + raw hole ratios → additive fusion one residual candidate block Prefix tree and multiscale query retained leaves → dyadic tree → cover tC_t same-level causal refinement; read-only cover query ×4× 4 xi(t)∈ℝdhx_i^(t) ^d_h for every i∈Uti∈ U_tt(read-only roots),yt∈ℝdhC_t\ (read-only roots), 18.49988pty_t ^d_hCandidate–cover cross-attention ×2× 2candidate queries read tC_t; cover roots remain read-onlyFFNs alone receive g(t)=Update(g,pool(t))g^(t)=Update\! (g,pool(C_t) )Post-prefix candidate stack ×4× 4history cross-attention to y<ty_<t → suffix self-attention → FFNoutput x¯i(t) x_i^(t) for every remaining candidatequery history y<ty_<tTied dot-product pointerpairs yty_t with every x¯i(t) x_i^(t); raw logits ηt,i _t,idecode query yty_tConditional symmetry quotientWL classes of the remaining candidatesπθ(pt∣p<t,h,ℓ(L),e(L),m) _θ(p_t p_<t,h, ^(L),e^(L),m)sample among class representatives Figure S11: The route decoder (§ S2.5), read from top to bottom. The candidate composer RMS-normalises and adds the learned vector streams with the raw hole ratios, then applies one residual candidate block to obtain the row-dependent tokens xi(t)x_i^(t) (Eqs. S2.94–S2.97). In parallel, the prefix branch materialises the selected leaves as a dyadic tree, extracts its cover, and forms the read-only multiscale query yty_t. Candidates cross-attend to the cover, then each post-prefix block applies history cross-attention, suffix self-attention, and an FFN in that order. The fixed router global conditions the composer, tree, and cover query; the cover-refreshed g(t)g^(t) (Eq. (S2.98)) enters only the two candidate-to-cover FFNs. Finally, the tied pointer scores yty_t against every refined candidate (Eq. S2.107), and the conditional symmetry quotient collapses WL-indistinguishable choices before sampling (Eq. S2.109). We refresh the route at every training step. The refresh is cheap because nothing downstream is redrawn: sampling the new permutation costs one decoder pass, and the walker population is simply relabelled into the new frame, which is a measure-preserving coordinate change. After this, the route re-equilibrates in a few sampler sweeps (SM § S4, § S4.1) with no dedicated burn-in. S2.5.2 Sequential sampling and vectorised teacher scoring The policy has two execution schedules for one mathematical scoring row. Denote that row map by θ(p<t,h,ℓ(L),e(L),m)=(yt,X¯t,ηt),D_θ(p_<t;h, ^(L),e^(L),m)= (y_t,\, X_t,\, _t ), (S2.110) where X¯t X_t contains the fully refined states of the available candidates and ηt _t is obtained by the tied pointer; the symmetry quotient is applied immediately afterwards. In sequential sampling, the Hamiltonian-dependent projections, directed message tables, and structural attention biases that do not depend on the sampled prefix are computed once. At step t, the prefix, suffix, order and hole scans form every fresh xi(t)x_i^(t). The already selected states v~s=xps(s) v_s=x_p_s^(s) are retained and their spherical images vs=Sphere(v~s)v_s=Sphere( v_s) form the partial tree, its cover gives yty_t, and the candidate-to-cover and post-prefix blocks produce X¯t X_t before the tied pointer produces ηt _t. A quotient-aware Gumbel-max draw supplies ptp_t, after which the summary scans, raw selected-state history, and query-state history are updated. In Hamilton-Zero the selected tree is recomputed from the retained v~s v_s at the next row. The retained history consists of states, not projected attention keys and values, and thus there is (counter-intuitively) no tokenwise key–value recurrence behind Eq. (S2.110). In teacher-forced scoring, a complete stored route makes every prefix known. All step–candidate states are therefore constructed together, :=(xi(t))t,i∈ℝN^×N^×dh,v~t=t,pt,vt=Sphere(v~t).X:= (x_i^(t) )_t,i N× N× d_h, v_t=X_t,p_t, v_t=Sphere( v_t). (S2.111) The prefix and suffix terms are vector cumulative sums; the order-aware terms are triangular masked contractions. A single bottom-up tree scan builds every complete-subtree state at every level. For each row t, dynamic indexing gathers precisely the roots in tC_t from Eq. (S2.101); the cover-query, candidate-to-cover, strict-history, remaining-candidate, pointer and quotient operations then run over all rows in parallel with their row-specific masks. In particular, the raw leaf gathered at position t is t,ptX_t,p_t and is then sphered for the tree; it is not a site embedding shared across decode positions. The two schedules consequently satisfy, in exact arithmetic, ηt,isample=[θ(p<t,h,ℓ(L),e(L),m)]η,i=ηt,iteacher,i∈Ut.η^sample_t,i= [D_θ(p_<t;h, ^(L),e^(L),m) ]_η,i=η^teacher_t,i, i∈ U_t. (S2.112) The distinction is computational, sampling evaluates one changing row after another, whereas teacher forcing materialises every changing row and every dyadic cover in one masked batch. Figure S12 displays the two schedules. same θD_θ, parameters, masks, and row logits (up to floating-point evaluation order)Sequential ancestral samplingVectorised teacher-forced scoringprefix-independent Hamiltonian tablesand structural biasesbuild fresh row Xt=(xi(t))iX_t=(x_i^(t))_ifrom prefix/suffix/order/hole scansretain v~s=xps(s) v_s=x_p_s^(s); tree uses vs=Sphere(v~s)v_s=Sphere( v_s)tree scan →t→yt _t→ y_tcandidate–cover + post-prefix blockspointer → quotient → Gumbel ptp_tretain v~t v_t, yty_t, and summary scansno projected K/V cache; then advance to t+1t+1stored route p=(p0,…,pN^−1)p=(p_0,…,p_ N-1)construct all dynamic rows together∈ℝN^×N^×dhX N× N× d_hgather v~t=t,pt v_t=X_t,p_t and sphere it; one tree scanmaterialises every root at every leveldynamic gather of each tC_tall masked scoring rows in parallelcandidate-ID logits with a route-permuted triangular maskand exact quotiented route log-probability Figure S12: Two schedules for the same route scoring row. Left: sequential sampling constructs the dynamic candidates and prefix tree, samples one quotient class, and stores the selected row-conditioned leaf and query state before the next row. Right: teacher forcing constructs all N^2 N^2 dynamic candidate states, gathers and spheres the selected leaf in every row, performs one multilevel tree scan, dynamically gathers every row’s dyadic cover, and evaluates all masked scoring rows together. For teacher-forced scoring, the route log-probability therefore reads logπθ(p∣h,ℓ(L),e(L),m)=∑t=0N^−1[η~t,Ct(pt)−log∑C∈teη~t,C]. _θ(p h, ^(L),e^(L),m)\;=\; _t=0 N-1 [\, η_t,\,C_t(p_t)\;-\; _C _te η_t,C\, ]. (S2.113) Before detailing the training algorithm, we emphasise here that each route row contributes separately to the summed log-probability and to the Fisher statistics. The second-order tooling of § S4.3 aggregates those contributions over the route-position repeat axis into the shared policy-parameter blocks. This lets the model learn to route by the same second-order natural gradient that trains the wavefunction itself, without any extra added optimization machinery to calculate such gradients. S2.5.3 Training the route policy: the score-function gradient The route is discrete, so its derivative is estimated with the score-function identity [153], whereas the conditional wavefunction retains the standard VMC derivative. The sampling rule reads hierarchically as, p∼πθ,q∣p∼ρθ,p,ρθ,p(q):=|ψθ(q∣p)|2Zθ,p.p _θ, q p _θ,p, _θ,p(q):= | _θ(q p)|^2Z_θ,p. (S2.114) Here the fixed Hamiltonian-conditioning arguments of πθ _θ are suppressed, and Zθ,p:=∫|ψθ(q∣p)|2qZ_θ,p:= | _θ(q p)|^2\,dq is the squared norm that makes ρθ,p _θ,p a probability density. Thus the route and walker are not statistically independent. They are sampled by separate procedures: the policy first draws p, after which the MCMC kernel targets the routed density ρθ,p _θ,p. The procedures do not share transition steps, but the walker’s target does however depend on the route. For a fixed route, we can define Eloc(p)(q):=(H^ψθ(⋅∣p))(q)ψθ(q∣p),E¯(p):=q∼ρθ,p[Eloc(p)(q)],E:=p∼πθ[E¯(p)].E_loc^(p)(q):= ( H _θ(· p))(q) _θ(q p), E^(p):=E_q _θ,p[E_loc^(p)(q)], E:=E_p _θ[ E^(p)]. (S2.115) Differentiating the two factors gives ∇θE= _θE= p∼πθq∼ρθ,p[(Eloc(p)(q)−E¯(p))2Re∇θlogψθ(q∣p)] _p _θE_q _θ,p [ (E_loc^(p)(q)- E^(p) )2\,Re\, _θ _θ(q p) ] (S2.116) +p∼πθ[(E¯(p)−E)∇θlogπθ(p)]. +E_p _θ [ ( E^(p)-E ) _θ _θ(p) ]. The first line is the ordinary VMC gradient for each routed state and must be centred at that route’s own mean E¯(p) E^(p). Replacing it by the global mean E changes the finite-sample estimator and introduces a systematic cross-route term. The second line is the REINFORCE gradient from the policy-gradient theorem [141]. The following construction supplies robust energy scales and self-normalised importance sampling for this estimator. Let ese_s collect the local energies used by one policy update, over its route and walker indices, and let B be their total count. We first form the batch mean and mean (absolute) total variation (TV), μ1:=1B∑ses,TV1:=1B∑s|es−μ1|. _1:= 1B _se_s, _1:= 1B _s|e_s- _1|. (S2.117) We then clip before estimating the standard deviation, esclip e_s^clip :=clip(es,μ1−5TV1,μ1+5TV1), :=clip(e_s, _1-5\,TV_1, _1+5\,TV_1), (S2.118) e¯clip e^clip :=1B∑sesclip,σ1:=1B∑s(esclip−e¯clip)2+εσ. := 1B _se_s^clip, _1:= 1B _s (e_s^clip- e^clip )^2+ _σ. Here clip(x,a,b)clip(x,a,b) clamps x to [a,b][a,b], and εσ>0 _σ>0 is the numerical floor under the variance. We use the shared normalised reward Rs:=−(esclip−e¯clip)/σ1R_s:=-(e_s^clip- e^clip)/ _1. We write R(p,q)R(p,q) for this same normalised reward evaluated at walker q under route p. The clipping therefore protects the scale estimate itself (since σ1 _1 is not computed from the unclipped tail). Beam search is used to approximate the policy mode, p0:=argmaxp∈ℬbeamlogπθ(p).p_0:= _p _beam _θ(p). (S2.119) Here ℬbeamB_beam is the set of complete routes retained by the beam search. For a sampled action pip_i with BwB_w walkers qib∼ρθ,piq_ib _θ,p_i, indexed by b=1,…,Bwb=1,…,B_w, the mode expectation is estimated on those walkers with self-normalised weights, wib:=|ψθ(qib∣p0)|2|ψθ(qib∣pi)|2,w¯ib:=wib∑b′=1Bwwib′,b^imode:=∑b=1Bww¯ibR(p0,qib).w_ib:= | _θ(q_ib p_0)|^2| _θ(q_ib p_i)|^2, w_ib:= w_ib _b =1^B_ww_ib , b_i^mode:= _b=1^B_w w_ibR(p_0,q_ib). (S2.120) The unknown normalising constants cancel under self-normalisation. The estimator depends on the sampled action through the proposal density in the denominator. It is biased at finite BwB_w, but it is consistent and asymptotically unbiased as Bw→∞B_w→∞ under the usual support and finite-weight conditions. With K the number of sampled routes in the policy update, R¯i:=Bw−1∑bR(pi,qib) R_i:=B_w^-1 _bR(p_i,q_ib), and Jroute:=p∼πθq∼ρθ,p[R(p,q)]J_route:=E_p _θE_q _θ,p[R(p,q)] the expected routed reward, the policy channel is estimated by ∇θJ^route=1K∑i=1K(R¯i−b^imode)∇θlogπθ(pi). _θJ_route= 1K _i=1^K ( R_i- b_i^mode ) _θ _θ(p_i). (S2.121) All logπθ(pi) _θ(p_i) values are evaluated by the teacher-forced autoregressive scoring pass of Eq. (S2.113). Cost. Table S1 gives the compute distribution per system per training step. Operation When Cost ancestral dense candidate core K routes (KN^3dh)O(K N^3d_h) arithmetic sampled prefix-tree recomputation K routes (KN^4de)O(K N^4d_e) in slot count beam baseline, width KbeamK_beam every step KbeamK_beam times the sampled-route costs teacher scoring, dense core + one tree K routes (KN^3(dh+de))O(K N^3(d_h+d_e)) arithmetic, batched conditional quotient, one route row every row (RN^3)O(R N^3) (1-WL); (RN^4)O(R N^4) (2-FWL) advantage + policy KFAC factor every step (|policy params|)O( params ) Table S1: Per-step cost of the route policy machinery. N N is the routed slot count; dhd_h and ded_e are fixed feature widths; K=8K=8 is the per-physical-system route count, and Kbeam=16K_beam=16 the beam width in the reported checkpoint. The repeated sampled-tree bound comes from rerunning a cubic same-level 2-FWL tree scan at each route row. In the quotient row, R is the refinement-round limit (R=N^R= N by default); its per-row cost must additionally be multiplied by the route rows and routes actually evaluated. The compiled checkpoint maps the eight sampled routes to separate lanes, which changes elapsed critical path but not total arithmetic. The teacher schedule does not reduce the arithmetic order: it exposes the KN^K N scoring rows to masked batched kernels. Sampling and beam search remain sequential in route position, whereas the teacher-forced scoring that carries the policy gradient is vectorised over rows and routes. S2.5.4 An entropy floor for the optimal route policy We close the routing story with an aside. Nothing in this subsubsection is needed to build or train the architecture. We include it because it is a sharp theoretical observation about what routing can achieve, and because it doubles as a powerful sanity check on the implementation. We characterise the invariant policies through automorphism orbits and identify the minimum entropy among invariant policies supported on energy-minimising routes. Here we will show that the result depends on both route-orbit sizes, their stabilisers, and the group size. Define the Hamiltonian’s automorphism group as the permutations of the slot block that preserve J, h, and m simultaneously: Aut(J,h,m):=g∈SN^:Jg⋅t,g⋅u=Jt,u,hg⋅t=ht,mg⋅t=mt∀t,u.Aut(J,h,m)\;:=\; \\,g∈ S_ N\;:\;J_g· t,\,g· u=J_t,u,\;\;h_g· t=h_t,\;\;m_g· t=m_t\;\;∀\,t,u\, \. (S2.122) Aut(J,h,m)Aut(J,h,m) is finite and always contains the identity. Notice that it also always contains the (N^−N)!( N-N)! permutations of the holes among themselves: hole rows of J and h are zero and the mask is preserved, so exchanging holes changes nothing the network can see. On top of this padding symmetry sit the physical ones. For a Heisenberg ring, Aut(J,h,m)Aut(J,h,m) contains the lattice translations and reflections; for an irregular disordered system, it contains nothing more than the hole exchanges. Gauge equivalence under automorphism. For every g∈Aut(J,h,m)g (J,h,m) the routed wavefunction satisfies ψθ(q∣p)=ψθ(g⋅q∣g⋅p)⋅eiφ(g), _θ(q p)\;=\; _θ(g· q g· p)· e^i (g), (S2.123) up to a global phase φ(g) (g) that depends only on g. The proof is by induction up the readout of § S2.4: every Hamiltonian-side input, contexts, edges, clocks, masks, is built equivariantly from (J,h,m)(J,h,m), which g preserves, so the carrier path evaluates the same contraction on relabelled inputs. Consequently the routes p and g⋅pg· p yield the same observable energy at every walker: Eloc(p)(q)=Eloc(g⋅p)(g⋅q)E_loc^(p)(q)=E_loc^(g· p)(g· q). Optimal policy invariance. Among unbiased policies, the variance-minimising one is invariant under the Aut(J,h,m)Aut(J,h,m) action: πθ∗(p)=πθ∗(g⋅p)∀g∈Aut(J,h,m). _θ^*(p)\;=\; _θ^*(g· p) ∀\,g (J,h,m). (S2.124) The argument is symmetrisation. Take any non-invariant optimum ν and average it over the group, ν¯(p):=|Aut(J,h,m)|−1∑gν(g−1⋅p) ν(p):= (J,h,m) ^-1 _gν(g^-1· p). By (S2.123) the average has the same expected reward, and by Jensen’s inequality on the variance functional it has weakly smaller gradient variance. The policy entropy decomposes over route orbits as follows. Let G:=Aut(J,h,m)G:=Aut(J,h,m) act on the route set, and let k\O_k\ be its route orbits. Write ok:=|k|o_k:=|O_k| and let ωk _k be the total policy mass assigned to orbit k. A G-invariant policy is uniform within each occupied orbit, so, with H[π]:=−∑pπ(p)logπ(p)H[π]:=- _pπ(p) π(p) and H(ω):=−∑kωklogωkH(ω):=- _k _k _k denoting the corresponding Shannon entropies, H[π]=−∑kωklogωkok=H(ω)+∑kωklogok.H[π]=- _k _k _ko_k=H(ω)+ _k _k o_k. (S2.125) By orbit–stabiliser, ok=|G||StabG(pk)|,pk∈k.o_k= |G||Stab_G(p_k)|, p_k _k. (S2.126) where StabG(pk):=g∈G:g⋅pk=pkStab_G(p_k):=\g∈ G:g\!·\!p_k=p_k\ is the stabiliser subgroup of the representative route pkp_k. Let ⋆ O_ be the set of orbits containing energy-minimising routes. Because the energy objective is linear in the policy, there exists an optimal G-invariant policy supported uniformly on one orbit ⋆∈⋆O_ ∈ O_ . Among such optimal invariant policies, the minimum entropy is H⋆(J,h,m)=min∈⋆log||,H_ (J,h,m)= _O∈ O_ |O|, (S2.127) attained on a smallest energy-minimising orbit. Spreading mass across several optimal orbits adds the nonnegative mixing entropy H(ω)H(ω). Equation (S2.127) is an orbit-specific optimum, not the universal quantity log|G| |G| unless the action on the relevant route is free, meaning that its stabiliser is the identity subgroup. Whenever |⋆|<N^!|O_ |< N!, the explicit single-orbit construction yields a non-uniform optimum; other optima may be uniform. S3 Datasets and further detail on training We pretrain the foundation ansatz on a corpus of quadratic spin Hamiltonians, spanning a century of many body literature [5, 83, 59, 87, 62, 140, 122, 72, 25, 129, 130, 134, 95, 19, 57, 80, 73, 6, 11, 131, 12, 100, 1, 107, 84, 48, 154, 75, 76, 66, 18, 146, 46, 99, 9, 147, 15, 2, 158, 128, 71, 88, 32, 96, 77, 103]. In this section, we detail the core composition of this dataset in Sec. S3.1, and then walk through in detail the augmented training and per-round perturbation and held-out evaluation datasets that were mentioned in the main text. The pretraining data uses a quality-balanced mixture rather than empirical family frequencies. Rare regimes are over-represented, and each system’s neighbourhood is densified by the perturbative cover as training proceeds. During pretraining, curvature damping is set as low as the natural-gradient norm allows (SM § S4). Model selection uses a held-out out-of-distribution dataset, followed by evaluation on a separate unseen dataset. S3.1 Pretraining data core composition Our pretraining data’s core is a finite set of quadratic spin Hamiltonians (J(s),h(s))s=1S\(J^(s),h^(s))\_s=1^S in the Pauli-pair representation of SM § S2. There J is the N×N×3×3N× N× 3× 3 bond tensor carrying every Pauli-pair channel, and h is the per-site field. We group the systems first by where they commonly arise in the physics literature and refer to each group as a cell. Within a cell the Hamiltonians range over size, interaction graph, coupling channel, field and anisotropy. Construction is deterministic. A parametrised builder emits every system from a seed committed to the repository, and the exact systems can be rebuilt from source. The systems are grouped into physics cells, with three further coverage tiers33 3 the dataset employed is publically evailable at https://github.com/simulacra-research/HamiltonZero/blob/main/datasets/train/foundation_5000.jsonl. The defining property of each cell’s Hamiltonians are briefly given below. • Canonical: Heisenberg, XXZ, transverse-field Ising and Majumdar–Ghosh on regular lattices, through their exactly solvable points. • Topological / SPT: SSH, Rice–Mele, the Kitaev chain and Haldane, cluster and CZX proxies across topological transitions. • Critical-point fans: non-Ising quantum critical points, the 22D XXZ, J–Q and Ising–Dzyaloshinskii–Moriya families. • Hardware-native: Hamiltonians native to superconducting, trapped-ion, Rydberg, spin-qubit and molecular platforms. • Frustration / spin liquids: kagome, triangular, Kitaev-honeycomb, J1J_1–J2J_2 and Shastry–Sutherland magnets. • Disorder / MBL: random-field and random-bond chains across the many-body-localisation transition. • Extremal entanglement: volume-law ground states, from random Hermitians to the Lipkin–Meshkov–Glick, Page and GHZ parents. • Lattice gauge: ℤ2Z_2 and U(11) gauge theories in two-body form. • Quantum chemistry / Hubbard: Fermi–Hubbard and molecular Hamiltonians through the Jordan–Wigner map. • Goldstone: gapless symmetry-broken magnets with long-wavelength modes. • Many-body scars: athermal eigenstate towers, the PXP, AKLT and η-pairing families. • SYK / random graphs: Hamiltonians with no lattice structure, dense and sparse SYK and random-graph couplings. • Cross-channel J: off-diagonal Pauli couplings, the Dzyaloshinskii–Moriya, compass and per-bond-random families. • Coverage tiers: an easy low-correlation band, a space-filling random-Hamiltonian cover of the small-N regime, and graphs that defeat the routing quotient of SM § S2. Fig. S13 reports how many systems each cell contributes to the dataset. Cell Systems Canonical 1051 SYK & random graphs 396 Frustration & spin liquids 394 Topological & SPT 322 Cross-channel J 297 Hardware-platform-native 391 Disorder & MBL 215 Quantum many-body scars 509 Critical-point fans 153 Extremal entanglement 86 Quantum chemistry & Hubbard 79 Lattice gauge 42 Continuous symmetry & Goldstone 17 Coverage tiers Small-N random cover 664 Simple systems 287 WL-11-breaking 97 Total 5000 Figure S13: Composition and system-size distribution of the 5,000-system pretraining dataset Left: the number of systems contributed by each physics cell and coverage tier. Right: the number of systems at each spin count N, stacked by cell. Colours match the table colours and follow the table order from the bottom of each bar upwards. S3.2 Augmented training by per-round perturbation Every datapoint in the pretrain set is a single set of edge values on one interaction graph (J,h)(J,h). Trained on that point alone, the model can overfit to the exact edge values rather than the physics they carry. Therefore in our pretraining, we augment the training data with a perturbation routine that we detail in this subsection. Each round, we perturb the edge values of every system independently, so the optimiser sees a neighbourhood of values around each datapoint and averages over it instead of fitting one point. A round is one such perturbed version of the whole training dataset, swapped in at an epoch end. Round 0 is an unperturbed base, and the run settles back on it after having trained for a few epochs on the perturbed data. For each structured system and round, a subroutine draws one of five mutations with equal probability: (i) leave the system unchanged, (i) add one bond, (i) perturb one existing bond, (iv) remove one bond, or (v) add one orbit-symmetric random field. Sites, bonds, and orbits are sampled stratified over the Weisfeiler–Leman orbit partition of (J,h)(J,h) [10, 34]. In case (v), one random three-vector of magnitude ηsη s, with η∼(0.25,1)η (0.25,1) and s the system’s coupling scale, is added to every site in the selected orbit. For small random systems with N≤8N≤ 8, the worker additionally densifies towards all-to-all coupling; a field-free system receives a substantial random dense field with probability 1/21/2, together with occasional bond drops. Speculative re-gauging is evaluated after the MCMC update, when the routed walker population is already available. It applies Jij↦RiJijRj⊤,hi↦Rihi,qi↦qiu¯i,J_ij R_iJ_ijR_j , h_i R_ih_i, q_i q_i u_i, (S3.1) with one uiu_i per Weisfeiler–Leman orbit, u¯i:=ui−1 u_i:=u_i^-1, and Ri:=Rker(ui)∈SO(3)R_i:=R_ker(u_i) (3) the adjoint rotation expressed in the kernel frame. The transformed Hamiltonian is unitarily equivalent to the original and the transformed walkers target its routed density exactly. We emphasise that a round differs from the base in the values of (J,h)(J,h) and nothing else. As such, the perturbation is cheap, because it changes only the values in (J,h)(J,h), meaning shapes remain static, and thus we avoid any recompiling. To make sure none of this stalls training, we make the routine CPU-based such that it executes the perturbation run asynchronously. The routine stays 2 rounds in front of the inner training loop. For each upcoming round it perturbs the pretrain dataset and precomputes the reference data, the exact-diagonalisation energies to N≤22N≤ 22 and the per-system routing fields, so the next round is ready before the loop reaches it. The loop consumes one round per epoch and, at each epoch boundary, swaps in the next round if the routine has produced it. The look-ahead is the margin that absorbs the time the perturbation and the diagonalisations take, so the training loop has zero latency with respect to this routine. Figure S14 shows how the two processes run asynchronously with zero latency. base system(J,h)(J,h) uniformly draw one mutation keep ∣ add bond ∣ perturb bond remove bond ∣ add orbit-symmetric field build referencesand route fields round cachelook-ahead 22 epoch-boundaryhot-swap optimisationstep post-MCMC speculative re-gauging zero measured overhead GPU MCMCcurrent round Figure S14: Augmented training and asynchronous hot-swap. The CPU samples one of five system perturbations, builds its references and routing fields 22 rounds ahead, and caches the result. At each epoch end the GPU swaps in the next cached round. Post-MCMC speculative re-gauging reuses the current walkers and adds no measured overhead to the inner loop. The cover is local. Each round perturbs a system within a neighbourhood of its base, so augmentation fills in the surroundings of the curated physics rather than reaching for new phases. The critical points the model must reproduce are placed in the pretraining data itself, in the fans of § S3.1, since they are not reached by the perturbation in general. The exact diagonalisations the worker computes are a second payoff. Training is self-supervised and needs no reference energies, so they serve only as a convergence diagnostic, and because each is valid only within its own round, the gap to exact is read per round. Yet each round contributes 2882 of them at N≤22N≤ 22, and across a run they accumulate into a corpus of tens of thousands of exactly solved systems, a byproduct of the cover rather than a cost of it. S3.3 Evaluation and fine-tuning datasets In this section, we describe the three datasets we used to investigate how well Hamilton-Zero generalises to unseen interaction topologies and interaction types. To do this, we kept several types Hamiltonian families out of training, using them only here for evaluation and fine-tuning. First in the set of three is Hamiltonian’s corresponding to combinatorial optimisation problems. These are Hamiltonians whose bonding structure stays inside the training distribution (graph Ising, H=∑(i,j)∈EJijσizσjz+∑ihiσizH= _(i,j)∈ EJ_ijσ^z_iσ^z_j+ _ih_iσ^z_i) but whose problem semantics and interaction topologies are absent from training. Every Hamiltonian in this set encodes an NP-hard combinatorial optimisation problem whose ground state is the problem’s optimal solution. We emphasise that unlike many systems in training and the other evaluation sets, such solutins correspond to separable ground states, so the difficulty here not in the entanglement of ground states, but rather their single-direction support in an otherwise exponentially large Hilbert space. The sub-families are listed in Table S2. Sub-family #cells Graph / structure MaxCut on random regular graphs 20 k-regular (k∈3,5k∈\3,5\), BA scale-free, planted-cut MaxCut on Erdős–Rényi 8 Density p∈0.4,0.6p∈\0.4,0.6\ QUBO (dense + sparse) 16 Jij∈±1J_ij∈\± 1\, hi∼(0,1)h_i (0,1); dense KNK_N or sparse-ER Number partitioning 10 Rank-11 J=vv⊤J=v\,v , vi∼Uniformv_i or power-law Max-2-SAT 10 Random clauses at clause-to-variable ratio α∈0.8,2α∈\0.8,2\ Max independent set / min vertex cover 14 Penalty form H=∑v−σvz+λ∑(i,j)∈E(σiz+1)(σjz+1)/4H= _v-σ^z_v+λ _(i,j)∈ E(σ^z_i+1)(σ^z_j+1)/4 Graph 2-colouring (adversarial) 8 55-regular graphs with many odd cycles Hamiltonian-path proxy 6 22-body approximation of the 44-body Lucas encoding Table S2: Table of systems for the first evaluation set The second evaluation set covers unseen topologies and interaction types, including fractal, quasi-periodic, and higher-dimensional interaction geometries such as hypercubes. Its subfamilies are listed in Table S3. Sub-family #cells Graph Fractal: Sierpinski + T-fractal 14 Hausdorff dim log3/log2≈1.585 3/ 2≈ 1.585 (Sierpinski); branching tree (T-fractal) Penrose / Fibonacci quasi-periodic 10 Fibonacci chain ± pentagon ring Decorated lattices 16 Lieb (flat band), dice (T3T_3, AB caging), square-octagon (4.8.84.8.8) 3D high-coordination 10 FCC (frustrated), BCC (bipartite), diamond 4D hypercube N=16N=16 8 K24K_2^4; AFM / FM / XXZ / TFIM 5D hypercube N=32N=32 4 K25K_2^5, all out of scope for ED; the headline extrapolation cell Hyperbolic 7,3\7,3\ 10 Heptagon r=1r=1 + r=2r=2 patches (2 systems out of scope for ED) Möbius / complete-bipartite 7 Möbius-twisted chain, Ka,bK_a,b Table S3: The second evaluation set containing unseen interaction topology and unseen interaction channels. The training dataset has no fractal, no quasi-periodic-2D, no decorated, no 4D/5D hypercube, and no hyperbolic lattices. The third evaluation set contains Hamiltonians with physical structures absent from training: multi-spin-1/21/2-per-physical-site encodings, complex hoppings, and a gauge theory. Its subfamilies are listed in Table S4. Sub-family #cells Encoding SU(33) ULS chain (PROXY) 14 22 spin-1/21/2 per physical site; FM intra-bond →S=1→ S=1; Heisenberg + Z-boost approximation of the bilinear-biquadratic spin-11 chain. Kugel–Khomskii spin-orbital (PROXY) 14 22 spin-1/21/2 per site (spin + orbital); factorised as 22-body ⋅+⋅+S\!·\!S+ τ\!·\! τ+ cross-Ising (PROXY for the native 44-body (⋅)(⋅)(S\!·\!S)( τ\!·\! τ)). Hofstadter strip 24 22/33/44-leg ladders with complex hop t→teiϕt→ t\,e^iφ ⇒ DMz terms in real spin language; flux α per plaquette (incl. golden-mean quasi-periodic). Native 22-body. Spin-11 XXZ chain (encoded) 12 Bilinear-only spin-11 chain via 22 spin-1/21/2; Δ sweep covers Haldane / AKLT-like / easy-axis. Bethe-ansatz integrable reference points. 11D Hubbard JW 10 Lieb–Wu integrable line U=4U=4, Mott U=8U=8, weakly correlated U=2U=2 (truncated-JW PROXY). Schwinger θ-vacuum 11 11D U(11) gauge after staggered + JW + Gauss-law; θ∈0,π/2,πθ∈\0,π/2,π\ sweep; Coleman transition at θ=πθ=π. Native 22-body. Table S4: The third and hardest evaluation dataset. S3.3.1 Evaluation and fine-tuning configuration For evaluation, we let the pretrained router select the tree ordering (eight candidate trees decoded at temperature τ=1τ=1 from a width-64 beam, each raced over a short burn-in and measurement window before collapsing to the winner), compile the wavefunction once on that ordering, and sample the fixed model for 512 measurement steps with 256 walkers across eight tempered replicas and 24 adaptive Langevin moves per step, retaining 512×256×8≃1.05×106512× 256× 8 1.05× 10^6 samples per system. Energies are reported with the pooled standard deviation of per-walker means as the uncertainty. fine-tuning restarts from the same checkpoint but commits the route once, at initialisation, with a width-16 beam search: the wavefunction is compiled on that fixed ordering and the eager model and router are discarded, leaving a compiled ansatz of one shared kernel plus per-spin tree weights (∼ 4.7M parameters, against 547M for the full model) as the object being trained. Each system is optimised independently on a single A100 for 10410^4 KFAC steps with 256 walkers, four MCMC moves per step and a 128-step burn-in, with learning rate 0.01/(1+t/104)0.01/(1+t/10^4) (ending at ≈0.005≈ 0.005), damping 10−310^-3, curvature EMA 0.90.9 and curvature/inverse updates every 2/4 steps. After training, every system is re-evaluated with frozen weights: the final checkpoint is resumed at exactly zero learning rate for a further 512 measurement steps with the sampler state carried over, the freeze is certified per system by a byte-identical model checkpoint written inside the evaluation window, and the standard error carries an AR(1) autocorrelation correction. ED references are never supplied to evaluation or fine-tuning and enter only in the analysis of the main text. S4 Optimization and Engineering In this section, we detail our state-of-the-art sampling scheme for MCMC on the configuration manifold SU(2)NSU(2)^N, as well as our energy kernels, our extension of the KFAC optimizer to the trunk’s stacked weights and the rank-4 merge tensor, all of which we made shard-mappable and GEMM-optimised. We take them in the order the data flows; the sampler produces walker configurations, then energy kernels evaluate the energy and its gradient on them, and finally the KFAC extensions consume both. S4.1 MCMC on the SU(2) manifold Improving the energy and updating the weights both require samples drawn from the model’s own density |ψθ(q)|2| _θ(q)|^2 on (SU(2))N(SU(2))^N. The variational energy is the loss, and its gradient drives the optimiser, so the variance of these Monte Carlo estimates sets how noisy each training step is. That variance grows with the chain’s integrated autocorrelation time and shrinks with the number of samples, so a sampler that decorrelates faster reaches a given accuracy from fewer samples. The sampler is therefore worth as much engineering as the energy kernel. Ours goes well beyond the Metropolis–Hastings (MH) algorithm [94, 53] that variational Monte Carlo conventionally uses, whose limitations for electronic-structure sampling are analysed in [33]. MH algorithm proposes a move q→q′q→ q from a kernel g(q′∣q)g(q q) and accepts it with probability min(1,|ψθ(q′)|2g(q∣q′)/(|ψθ(q)|2g(q′∣q))) (1,\,| _θ(q )|^2\,g(q q )/(| _θ(q)|^2\,g(q q)) ). The usual random-walk proposal forces a bad trade-off in high dimensions. Large steps land in low-density regions and are almost all rejected; small steps are accepted but barely move. Either way, the chain decorrelates slowly, the autocorrelation time grows, and the effective sample size collapses [124]. Langevin dynamics breaks the trade-off by drifting the proposal up the density with gradient ascent. The Langevin diffusion dqt=∇log|ψθ(qt)|2dt+2dWtdq_t\;=\;∇ | _θ(q_t)|^2\,dt\;+\; 2\,dW_t (S4.1) has |ψθ|2| _θ|^2 as its stationary density, but a finite discretisation of it does not. A single Euler–Maruyama step of size σ gives the unadjusted Langevin proposal, and because it only approximates the diffusion, it samples a perturbed density whose bias away from |ψθ|2| _θ|^2 vanishes only as σ→0σ→ 0. The Metropolis-adjusted Langevin algorithm [126] removes this bias exactly by passing the Langevin proposal through the MH accept step, and so keeps the drift’s fast mixing while restoring |ψθ|2| _θ|^2 as the stationary density. Its optimal step size scales as σ∼D−1/3σ D^-1/3 rather than the random walk’s D−1/2D^-1/2 [125], a large gain at our dimension. MALA alone, however, is still not enough. Each rejected move discards a gradient evaluation, which is the most expensive quantity in the step, so reducing σ to cut rejections also shortens the steps and slows exploration. And a single local chain with a single scalar step size mixes slowly across the rugged, multimodal landscape of a frustrated or disordered Hamiltonian, where it can remain in one basin for many steps. The subsections below close these gaps in turn. First, a manifold-aware MALA proposal that respects the group geometry, then a ladder of tempered replicas [142, 60] that exchange to carry walkers between basins, followed by a global Haar-random refresh that erases stuck sites outright, and finally online adaptations that hold each piece at its optimal operating point [7]. Assembled, the scheme cuts the chain’s integrated autocorrelation time, and with it the variance of every energy and gradient estimate, so a target accuracy needs far fewer samples. And it returns batches differentiable in both the manifold coordinates and the model weights, the two derivatives the energy estimator and the natural-gradient step consume. Concretely, the sampler runs once per training step, in three stages. A setup stage runs first and once: it hoists the per-system quantities that do not depend on the configuration q, the featurizer and trunk outputs, and refreshes the cached log-density log|ψθ|2 | _θ|^2 and its gradient at the current q against the freshly updated model. A single global Haar refresh of m sites then fires, before any local move. Finally, a loop of inter-steps runs, each one a manifold MALA move followed by a cross-chain swap whose edge set alternates with the step parity. Three adaptations run on a slower cadence and condition the next pass: a multiplicative bang-bang rule drives the MALA step size to its optimal acceptance, a proportional rule tunes the number of Haar-refreshed sites to its own optimal acceptance, and a monotone-cubic interpolation re-spaces the temperature ladder to equalise swap rejection. Figure S15 shows this pipeline and where each subsection below fits. qt(r)q^(r)_tHoist per-system invariantsfeaturizer and trunk, independent of qqCache log|ψθ|2 | _θ|^2 and ∇alog|ψθ|2 _\!a | _θ|^2at the current q, refreshed modelHaar refresh of m sitesglobal, once before the loopMALA local movepropose from cached ∇, wrap-aware MH acceptDEO cross-chain swapeven or odd edges by step parityqt+1(r)q^(r)_t+1q′q loop ×X×\,Xparity flipsTune refresh count mmMultiplicative rule on σ cubic on β-gridHaar acceptanceupdated mmacceptance ata_tupdated σ rejectionupdated β-grid Figure S15: The per-chain MCMC pipeline at inverse temperature βr _r (§ S4.1), run once per training step. A setup stage runs first and once: it hoists the per-system, q-independent featurizer and trunk outputs, then refreshes the cached log-density and its gradient at the current q (§ S4.1.7). A single global Haar refresh of m sites (§ S4.1.5, dashed border for its stochastic accept) fires before any local move. The inner loop then runs a fixed number of inter-steps (the loop-back on the left), each a manifold MALA move (§ S4.1.2) followed by a deterministic even/odd cross-chain swap (§ S4.1.3) whose edge set alternates with the step parity, producing qt+1(r)q^(r)_t+1. Three adaptations (magenta, dashed) run on a slower cadence and condition the next pass: the refresh count m is tuned to the Haar acceptance rate (§ S4.1.5), the multiplicative rule of § S4.1.6 drives the MALA step size σ to its acceptance, and the Hyman monotone cubic (§ S4.1.4) re-spaces the inverse-temperature grid βr\ _r\ to the pair-rejection rates. All boxes are shades of the leaf-green hue used for the spin-configuration input in Figure S1, darkening from the setup stage along the loop; the colour match marks this figure as the subroutine that produces q. S4.1.1 Proposal density on the Lie group A proposal q→q′q→ q on SU(2)SU(2) is parameterised by a Lie algebra increment ξ∈ℝ3ξ ^3 via q′=q⋅exp(σξaea(q)),q =q· (σ\,ξ^a\,e_a(q) ), (S4.2) where ea(q)\e_a(q)\ is the left-invariant frame. The Gaussian density of ξ on the algebra is closed-form, but the induced density on the group is not: it must be wrapped around the periodicity of the exponential map. We use the wrap sum pwrap(ξ,σ)=∑n∈ℤ3,‖n‖≤Nwrap(ξ+2πn, 0,σ2I3),p_wrap(ξ;σ)\;=\; _n ^3,\,\|n\|≤ N_wrapN (ξ+2π n;\,0,\,σ^2I_3 ), (S4.3) truncated to Nwrap=8N_wrap=8 shells in the experiments reported here. For σ<7πσ<7π, vastly larger than any practical step, the truncation error falls below double-precision noise. The wrap sum costs about 3×3× a single Gaussian and takes under 0.3%0.3\% of MCMC wall time. S4.1.2 Manifold MALA on SU(2)NSU(2)^N Langevin and Hamiltonian proposals on manifolds have a developed literature [43, 20]; the construction below specialises to the product group SU(2)NSU(2)^N, where the exponential map and its wrapped proposal density are available in closed form. The MALA proposal of the introduction needs two changes on the Lie group. Vanilla addition q+ξq+ξ is not a group operation, so we exponentiate a Lie-algebra increment back onto the group, q′=q⋅exp(σξaea(q))q =q· (σ\,ξ^a\,e_a(q)), as in the proposal density above. And the Euclidean gradient becomes the left-invariant gradient ∇alog|ψθ|2=ea(q)⋅∇log|ψθ|2 _\!a | _θ|^2=e_a(q)·∇ | _θ|^2 along the frame ea(q)\e_a(q)\. The manifold-MALA proposal is then ξ=σ22∇alog|ψθ(q)|2+ση,η∼(0,I3).ξ\;=\; σ^22\, _\!a | _θ(q)|^2\;+\;σ\,η, η (0,I_3). (S4.4) The MH correction now needs both the wrap-sum proposal density (S4.3) and the volume Jacobian of the exponential map. Either correction alone breaks detailed balance, because the wrapped Gaussian is the proposal density on the group while the Jacobian accounts for the change of measure between the algebra and the group. S4.1.3 Replica exchange with deterministic even-odd swaps A single chain, however well tuned, still struggles to cross between well-separated modes, so we run a ladder of them. R chains run in parallel at decreasing inverse temperatures β1>β2>⋯>βR _1> _2>…> _R; the coldest replica, at β1 _1, sits near the modes of |ψθ|2| _θ|^2, while the hottest replica, at βR=0 _R=0, explores broadly. After each local MALA move the chain proposes swaps between adjacent replicas along the deterministic even/odd (DEO) schedule of 143, whose eligible edges alternate with the step parity. At even steps the even edges swap, (1,2),(3,4),…(1,2),(3,4),…; at odd steps the odd edges swap, (2,3),(4,5),…(2,3),(4,5),…; so every adjacent pair is offered a swap once in every two steps. Each pair accepts with the standard MH ratio, αr,r+1=min(1,exp((βr−βr+1)(log|ψθ(xr+1)|2−log|ψθ(xr)|2))). _r,r+1\;=\; \! (1,\; (( _r- _r+1)\,( | _θ(x_r+1)|^2- | _θ(x_r)|^2) ) ). (S4.5) The alternation is non-reversible. A pair swapped at an even step is not offered again until the next even step, two steps later; in between, the odd-step swap carries the walker on to its new neighbour rather than back. This directed transport along the ladder is what maximises information flow between the hot and cold chains [143]. Figure S16 shows the swap pattern over a few steps. β1 _1 (cold)β2 _2β3 _3βR _R (hot)t=0t=0event=1t=1oddt=2t=2event=3t=3oddt=4t=4event=5t=5oddt=6t=6evenchain indexstep parity → which edge set swaps Figure S16: Deterministic even/odd (DEO) replica-exchange schedule (§ S4.1.3). R chains run in parallel at decreasing inverse temperatures β1>β2>⋯>βR _1> _2>…> _R (rows, cold → hot colour gradient). At each MCMC time step (columns), pairs of adjacent chains are proposed for Metropolis swaps: even-edge pairs (1,2)(1,2) and (3,4)(3,4) at even t, odd-edge pair (2,3)(2,3) at odd t. The alternation is deterministic and exhaustive, every pair firing once in every two steps, and non-reversible in the sense that consecutive swaps cannot trivially undo each other, which maximises information flow along the temperature ladder. S4.1.4 Online β-adaptation via monotone cubic interpolation The ladder only carries walkers if adjacent rungs overlap enough to swap: too coarse a spacing rejects every swap, too fine a spacing wastes replicas. At stationarity the optimal schedule equalises the adjacent-pair rejection rates across the ladder [8]: for a target rate λ¯ λ, every adjacent pair (r,r+1)(r,r+1) should reject with probability λ¯ λ. This is the schedule-optimisation problem that Syed et al. [143] solve through the cumulative rejection (their communication barrier) Λ(β)=∫β1βλ(β′)|dβ′| (β)= _ _1^βλ(β )\, β : the optimal grid divides Λ into equal increments, βr=Λ−1((r−1)λ¯) _r= ^-1 ((r-1)\, λ ) with λ¯:=Λ(βR)/(R−1) λ:= ( _R)/(R-1). What we add is the online version: the running chain gives only the R samples (βr,Λ^(βr))r\( _r, ( _r))\_r, so recovering the grid means inverting a monotone interpolant of Λ , refreshed as the model, and with it the rejection profile, drifts during training. Piecewise cubic Hermite interpolation fixes both the value f(βr)f( _r) and the derivative f′(βr)f ( _r) at each node and fits each segment with a cubic. The choice of node derivative is what separates the methods. PCHIP, the piecewise cubic Hermite interpolating polynomial of 38, sets the derivative to zero wherever the local secant slope δr=(f(βr+1)−f(βr))/(βr+1−βr) _r=(f( _r+1)-f( _r))/( _r+1- _r) changes sign and to the harmonic mean of δr−1,δr _r-1, _r otherwise; it is monotone by construction, but the harmonic mean loses accuracy through a smooth inflection. Hyman’s 61 derivative is instead the centred secant f′(βr)=(f(βr+1)−f(βr−1))/(βr+1−βr−1)f ( _r)=(f( _r+1)-f( _r-1))/( _r+1- _r-1), which is exact for quadratic data, and applies the Fritsch–Carlson [38] monotonicity clamp only as a safety net at a detected sign change. On the quadratic test f(β)=β2f(β)=β^2 Hyman recovers f′f to double precision (∼10−12 10^-12 relative error) where PCHIP errs at ∼10−3 10^-3 near an inflection. Spin-manifold rejection curves carry smooth inflection regions where that harmonic-mean error, not monotonicity, is the binding constraint, so we use Hyman. Each adaptation interval then runs four steps. It updates Λ^(βr) ( _r) by an exponential moving average over the last T swap attempts at each pair, fits the Hyman cubic to (βr,Λ^(βr))r\( _r, ( _r))\_r, bisects it to recover the equal-rejection grid β~r\ β_r\, and blends that grid into the old one, βr←(1−μ)βr+μβ~r _r←(1-μ)\, _r+μ\, β_r, for stability. On synthetic cold-pair-compression benchmarks the end-to-end gain is +15+15 to +19%+19\% over a one-sided backward-Hermite baseline. S4.1.5 Global m-site Haar refresh and the 0.2340.234 optimum Local moves and swaps still change the configuration only gradually, so a basin that no replica has entered stays unvisited. A global move sidesteps that. Once per training step, before the local-move loop, the chain refreshes m randomly chosen sites to Haar-uniform values on SU(2)SU(2). This non-local move breaks the within-chain correlation that MALA builds up and reaches disconnected modes that the local Langevin proposal cannot bridge. The refresh acceptance rate falls as m grows, since a larger refresh lands further from the current state and is accepted less often. We tune m on the same slow cadence as the step size, by a proportional rule: increment it by one when the running Haar acceptance exceeds the target, decrement it by one when it falls short, clipped to [1,nactive][1,n_active]. A unit step is the right granularity because acceptance is discrete in m and the search space has only nactiven_active candidates, so the optimum is reached quickly. The hottest replica, at β=0β=0, is held at full erasure, m=nactivem=n_active, where every refresh is accepted and the mixing is free. The target α⋆α is the acceptance that maximises the expected squared-jump distance per move. For a k-site refresh on a compact group the expected squared-jump distance (ESJD) grows linearly in m, and so does the variance V of the log-density increment, which is a sum of m per-site contributions; the two are therefore proportional, and [ESJD]∝αVE[ESJD] α\,V. Because the increment is a m-term sum, the central limit theorem makes it approximately Gaussian whenever the site contributions are weakly correlated, and exactly Gaussian when the target factorises over sites, as for the rank-one ferromagnetic ground state. We therefore take Δlog|ψθ|2=M+Z,Z∼(0,V), | _θ|^2\;=\;M+Z, Z (0,V), (S4.6) with mean M<0M<0. The refresh draws its sites independently of the current state, so the proposal is symmetric; detailed balance at stationarity then gives [eΔ]=1E[e ]=1, which for a Gaussian increment implies M=−V/2,equivalentlyV=−2M.M\;=\;-V/2, V\;=\;-2M. (S4.7) . The expected MH acceptance probability is then the truncated Gaussian-tail integral α(V)=[min(1,eΔlog|ψθ|2)]= 2Φ(−V/2),α(V)\;=\;E [ (1,e | _θ|^2) ]\;=\;2\, \! (- V/2 ), (S4.8) where Φ is the standard-normal CDF. The expected squared-jump distance (ESJD) is proportional to α(V)⋅Vα(V)· V. Setting d/dV[α(V)V]=0d/dV[α(V)\,V]=0 and solving numerically: V⋆/2≈ 1.19,V⋆≈ 5.66,α⋆≈ 0.234. V /2\;≈\;1.19, V \;≈\;5.66, α \;≈\;0.234. (S4.9) The resulting optimum, α⋆≈0.234α ≈ 0.234, matches the random-walk Metropolis limit of 124. Both derivations reduce to maximising the expected squared-jump distance when the log-density increment is asymptotically Gaussian and stationarity fixes its mean to −V/2-V/2. The present derivation concerns an m-site Haar refresh on a compact group and uses neither a random walk, Euclidean geometry, nor a small-step limit. S4.1.6 Multiplicative step-size adaptation Every piece above assumes its proposal sits at the right acceptance rate. The step size σ of the manifold MALA proposal (§ S4.1.2) is the component that holds the local-move acceptance at its optimum. For MALA that optimum is 0.5740.574 in the high-dimensional limit [125]; a plain random-walk proposal would instead target 0.2340.234. The Haar refresh has no step size and reaches its own 0.2340.234 optimum through the site count m of § S4.1.5, not through σ. Each replica drives its own σ to the target with a multiplicative rule. After every adaptation window it measures the local acceptance A and moves one fixed factor toward the target, σ⟵clip(σfsign(A−A⋆), 10−4,π),f=1.1,A⋆=0.574.σ\; \;clip (σ\,f^\,sign(A-A ),\;10^-4,\;π ), f=1.1, A =0.574. (S4.10) The rule is bang-bang: σ rises by the factor f when acceptance runs hot and falls by it when acceptance runs cold, so at equilibrium it oscillates within a factor f of the value that holds A=A⋆A=A . That is all the precision the sampler needs, because the MALA efficiency curve is flat near its optimum, and the rule carries no schedule and no state. The clip keeps σ inside the range where the proposal is meaningful on the group: steps below 10−410^-4 are numerically idle, and π is the diameter-scale move on SU(2)SU(2). The natural alternative is a Robbins–Monro stochastic-approximation schedule on logσ σ [123], whose diminishing gains converge almost surely to the exact target. But a diminishing gain is built for a stationary target, and ours moves: training reshapes |ψθ|2| _θ|^2 continuously, so an adapter that never stops moving is the safer default. S4.1.7 Gradient cache between MCMC inter-steps A second saving falls out of the chain’s own structure. The gradient the chain computes when it accepts a move at q is the same gradient the next inter-step needs before it proposes from q. Threading this cached gradient through the scan saves one value_and_grad per inter-step and halves the per-step gradient cost. At the start of each training step the cache is refreshed against the freshly updated model, which is the caching step the figure marks before the Haar refresh. S4.2 Energy kernels The bottleneck of variational Monte Carlo on a Hamiltonian quadratic in the Pauli operators is the energy estimator. For a sampled configuration q∈(SU(2))Nq∈(SU(2))^N the local energy splits into a two-body exchange term and a linear field term. The field is the first order in the derivatives of logψ ψ and is computed as by-product of computing second order terms. The whole cost sits in the exchange term, which contracts the coupling tensor against the second derivatives of logψ ψ, Eexch(q)=−14∑i,j∑a,bJijab[LiaLjblogψ(q)+(Lialogψ(q))(Ljblogψ(q))],E_exch(q)\;=\;- 14 _i,j _a,bJ^ab_ij\, [L_i^aL_j^b ψ(q)+ (L_i^a ψ(q) )\, (L_j^b ψ(q) ) ], (S4.11) where LiaL_i^a is the left-invariant derivative along Pauli direction a at site i (§S1.3) and the prefactor 14 14 comes from S=12σS= 12σ. The first bracket is a Hessian of logψ ψ contracted against J; the second is an outer product of its gradient. The Hessian term is the expensive one. Our custom JAXPR interpreter evaluates it, together with every other term the estimator needs, in a single forward pass: a custom forward-mode second-order operator written for the tree, described next. S4.2.1 Custom forward-mode Second-order operator One JAXPR interpreter serves the whole local energy. Its function signature reads, custom_lap:q⟼(logψθ(q),∇logψθ(q),Tr(WHlogψθ(q))), custom\_lap:\;q\; \; (\, _θ(q),\;∇ _θ(q),\;Tr (W\,H_ _θ(q) )\, ), (S4.12) one forward pass per walker, from which the score, the linear field term and the exchange term of Eq. (S4.11) are all read off: no Hessian is ever materialised and no Hessian-vector product is ever formed on the energy path. The kernel propagates a jet through the network, so each intermediate carries its value, its Jacobian along the Lie-frame input directions, and a running weighted trace, and each primitive advances the three together. The trace advances as a quadratic form of the Jacobian it already carries, tr⟵tr+∑μ,ν,cμcΛcμννc,,tr\; \;tr\;+\; _μ,ν,cJ_μ c\; _cμν\;J_ν c,, (S4.13) where μcJ_μ c is the intermediate’s Jacobian (μ,νμ,ν over the Lie-frame input directions, c over channels),not to be confused with the coupling tensor J, and Λcμν _cμν is the second-order weight the chain rule assigns to the primitive at hand; for every linear primitive Λ vanishes, so the sum is paid only where genuine curvature enters. (S4.13) is the update for element-wise primitives, whose curvature is diagonal in the channel index. The merge tensor is the one primitive with cross-channel curvature: there the update couples the two child direction blocks through T itself, and it is the only point in the network where a cross-block entry enters the running trace. The natural alternative is a generic forward-mode Laplacian[82], as implemented in folx [40]. Its tracer treats the second-derivative tangent as densely coupled across sites, because the Pauli tangents hide the structure of J: it sees a dense 4N×3N4N× 3N Jacobian fan-out and pays for all of it. The tree of S2.4 makes that coupling sparse. Every q-dependent intermediate at merge level κ carries a Jacobian of the chunked shape [ 3k,N/k,∗S][\,3k,\,N/k,\,*S\,], where k=2κk=2^κ is the number of sites per chunk, N/kN/k the number of chunks, the factor 33 the Lie-frame directions per site, and ∗S*S the trailing feature axes; the only new cross-block coupling in the Jacobian is the one the merge itself creates between its two children. custom_lap is a forward-mode Second order tracer written for exactly this layout. Its tracer treats the second-derivative tangent as densely coupled across sites, because the Pauli tangents hide the structure of J: it sees a dense 4N×3N4N× 3N Jacobian fan-out and pays for all of it. custom_lap is our extension of this technique: a weighted-trace generalisation (Tr(WH)Tr(WH) rather than Tr(H)Tr(H), which is what lets one pass serve an arbitrary coupling tensor JabJ^ab), specialised to the tree’s layout, where chunk structure sets the cost. At level κ the N^/2κ N/2^κ chunks each carry 3⋅2κ3· 2^κ Lie-frame directions, so the quadratic-form updates at that level cost ((N^/2κ)⋅(3⋅2κ)2⋅S)=(N^ 2κS)O (( N/2^κ)·(3· 2^κ)^2· S )=O ( N\,2^κS ) with S the feature width. Summing the geometric series over κ≤Kκ≤ K gives (N^2S)O( N^2S) for the full pass, against the (N^3)O( N^3) of a dense forward-mode tracer. This claim is anticipated in § S2.4.5. We emphasize that our extension here cannot be used as a universal differentiator since it assumes the balanced binary tree, and a symmetric two-body weight W:ℝ3N→ℝ3NW:R^3N ^3N. Thus, the custom second order operator that we built for our model, can produce correct second order differential operator values only on the tree-like ansatz. This is the core reason we aren’t bound by the generic algorithmic complexity of evaluating the Hessian of a black box function. S4.3 Kronecker-factored curvature for the foundation ansatz KFAC (Kronecker-factored approximate curvature [91]) is the natural-gradient method we use for pretraining. Natural gradient multiplies the raw gradient by the inverse of the Fisher information matrix, the curvature of the model’s distribution; the resulting step is invariant to how the model is parameterised and reaches a given loss in far fewer iterations than plain gradient descent. The Fisher information is a P×P× P matrix for P parameters, far too large to form or invert at our scale, and KFAC’s contribution is to approximate it cheaply. For a single dense layer with weightW∈ℝdout×dinW ^d_out× d_in, input activation a∈ℝdina ^d_in and back-propagated output gradient δ∈ℝdoutδ ^d_out, the block of the Fisher information belonging to W is the covariance of the per-example gradient vec(δa⊤)vec(δ a ), FW=[vec(δa⊤)vec(δa⊤)⊤]=[aa⊤⊗δδ⊤].F_W\;=\;E [vec(δ a )\,vec(δ a ) ]\;=\;E [a δ ]. (S4.14) KFAC replaces the expectation of this Kronecker product by the product of expectations, which is exact when a and δ are independent, FW≈A⊗G,A:=[aa⊤],G:=[δδ⊤].F_W\;≈\;A G, A:=E[a ],\ G:=E[δ ]. (S4.15) The factorisation is what makes the method affordable. Inverting the two factors separately, (A⊗G)−1=A−1⊗G−1(A G)^-1=A^-1 G^-1, costs (d3)O(d^3) per layer in place of the (d6)O(d^6) of the full block. Before inversion, the KFAC implementation adds damping γ, F~−1≈(F^+γI)−1 F^-1≈( F+γ I)^-1, which trades fidelity for stability: a smaller γ keeps the preconditioner closer to the true Fisher information and each step closer to the natural-gradient direction, so fewer steps reach a given loss. Our model breaks three assumptions that KFAC’s standard blocks build in. The trunk’s L repeated blocks stack their weights through a loop, so the Fisher information of those weights couples across the loop iterations. The attention projections couple across the query, key, value and head channels. And the merge tensor of § S2.4 is cubic in its width, so its full Fisher information is far too large to form. The patches below extend the kfac_jax library [16] to each case. Together they remove from the model every dense Fisher-information block, and leave a curvature estimate faithful enough to run near the edge of stability. The Fisher information itself is registered on both output channels, the log-amplitude and the phase, as two normal-predictive registrations: the Fubini–Study metric of the complex output. In the variational Monte Carlo literature this natural gradient is stochastic reconfiguration [139], with the Fubini-Study metric playing the role of the Fisher information [56]. KFAC-preconditioned wavefunction optimisation was introduced for deep ansatze by [113], whose registration strategy ours extends to a complex output. S4.3.1 Fusing the query, key and value projections Each trunk block’s attention sublayer from § S2.3 projects the per-site stream ℓ∈ℝdℓ ^d_ into queries q_ h, keys k_ h and values v_ h for each of its n_ h heads. We compute all three projections with one weight matrix, W∈ℝ(3nd)×dℓW_ qkv ^(3n_ hd_ h)× d_ , rather than three separate ones. The fusion needs no custom curvature class: the fused weight registers with KFAC as one ordinary dense layer. The standard factorisation hands KFAC a single gradient-covariance factor G of size (3nd)2(3n_ hd_ h)^2, which carries the covariance between the query, key, value and head channels; three separate registrations would set that cross-covariance to zero by construction. S4.3.2 The stacked trunk and a third Kronecker factor A standard KFAC block factors a dense layer’s Fisher information as A⊗GA G. The trunk’s dense layers, however, are stacked: the same layer appears in all L trunk blocks, carried on an extra leading axis of length L. Sharing one A⊗GA G across the L copies treats them as independent and throws away any correlation between the blocks. Our block keeps that correlation, factoring the Fisher information as a three-way Kronecker product Fscan≈M⊗A⊗G,F_scan\;≈\;M A G, (S4.16) where M is the L×L× L covariance across the L blocks and A, G are the activation and gradient covariances as before. Here and below ++nS_++^n denotes the n×n× n symmetric positive-definite matrices, so M∈++LM _++^L, A∈++dinA _++^d_in and G∈++doutG _++^d_out. Two-factor KFAC is the special case M∝ILM I_L, which discards the cross-block correlation entirely. The third factor is cheap: L, the number of trunk blocks, is far smaller than the layer widths, so M adds only an L2L^2 term to factors that already cost d2d^2. Kronecker structure across a loop axis has precedent in the recurrent setting [90], where the shared weight’s Fisher couples across time steps. Carrying an explicit learned L×L× L factor across stacked, distinct block weights, capturing the cross-block correlation that layer-wise block-diagonal KFAC discards, is, as far as we are aware, novel. The block estimates M online, as a running exponentially-weighted average of the gradient correlations between the L blocks. This is the only statistic it keeps beyond what two-factor KFAC already tracks, and at (L2)O(L^2) per block its cost is negligible against the (d2)O(d^2) of the activation and gradient factors. S4.3.3 Stacked RMS scale parameters The RMS normalisation layers of SM § S2 carry a per-channel scale and no shift, and the scales are stacked across the L blocks in the same way. Their Fisher information takes the same form, with the gradient factor replaced by a diagonal: a block-coupling factor across the L copies, times a diagonal per-feature term D. With a single block it reduces, exactly, to the standard diagonal treatment. S4.3.4 The merge tensor The rank-4 merge tensor T of § 2.4 is the hardest block. With B sub-blocks of contraction width drd_r it carries Bdr3B\,d_r^3 parameters, of order 10610^6 at the widths used here, so its full Fisher information would have (Bdr3)2(B\,d_r^3)^2 entries, of order 101310^13, roughly ten terabytes, which cannot be formed. The custom block instead matricises the four axes (i,j,k,l)(i,j,k,l) into a balanced pair and factorises the Fisher information as the two-factor Kronecker product ΣL⊗ΣR _L _R over that split, storing and inverting two Gram factors of a few megabytes together; the factored state is about six orders of magnitude smaller than the dense block. The sub-block axis rides inside one of the two factors, a genuine Kronecker leg rather than a batch dimension. Without it the merge tensor would fall back to a dense block and, at this width, would not train; with it the tensor receives a genuine natural-gradient preconditioner. S4.3.5 Damping The faithful three-way Fisher information changes what limits the damping. With the approximation tightened, the floor on how small γ can go is the noise in the curvature estimate itself, not the approximation error. We hold γ at that noise floor: a constant identity-weight damping of 10−310^-3, with a hard floor of 10−410^-4 and a norm constraint of 10−310^-3 on the preconditioned step. Large-model pretraining has shown [27] that deep networks converge in an edge-of-stability regime, where the largest curvature eigenvalue sits just above 2/η2/η for learning rate η and the loss oscillates without diverging. Our schedule rides the analogous edge. A cumulative-sum change detector (CUSUM) on the natural-gradient norm classifies the run as stable, marginal or divergent; γ is raised only when the run crosses into the divergent regime and drifted back down afterwards, so the optimiser sits at the smallest stable damping rather than behind a conservative margin. S4.3.6 Benchmarking the third factor versus per-layer curvature The natural alternative to the third factor of Eq. (S4.16) is per-layer curvature: give each of the L stacked copies its own two-factor block and precondition with the block diagonal F~bd:=⨁k=1LA^k⊗G^k F_bd:= _k=1^L A_k G_k. Per diagonal block this is the more expressive family — every diagonal block of F~3=M^⊗A^⊗G F_3= M A G is the same A^⊗G A G up to the scalar M^kk M_k, whereas F~bd F_bd fits all L blocks independently — and L times more expensive in factor state and inverse work; for the fused W_ qkv projection alone the factor state is 8×8× larger. The extra freedom buys accuracy only at initialisation, and only on the fused QKV block, where per-layer curvature wins decisively, 0.9560.956 against 0.3430.343 in the damped-action cosine (S4.17); the three-factor block leads on the other three trunk modules already at initialisation, and at the trained point it wins on all four. We evaluate both families against the exact sampled Fisher of Hamilton-Zero. Curvature capture runs through the estimator’s own one-hot seeded vector–Jacobian products, so the per-layer activations aka_k and cotangents δk _k are precisely the tensors the curvature blocks consume; the exact block FscanF_scan is assembled from the per-walker gradients and inverted exactly through the Woodbury identity at damping γ=10−3γ=10^-3. A candidate F~ F is scored by its damped-action alignment on the model’s true gradient g, cos∠((F~+γI)−1g,(Fscan+γI)−1g), (\,( F+γ I)^-1g,\ (F_scan+γ I)^-1g\, ), (S4.17) evaluated at two points: random initialisation, on a 6464-site open-boundary Heisenberg chain, and a trained checkpoint (step 30 00030\,000) with equilibrated MCMC walkers, aggregated across the 1313 training Hamiltonians categories of §S3 as the estimator aggregates, i.e. factors averaged across systems as the exponential moving average averages them, gradients and damped actions per system. Initialisation Trained (step 30 00030\,000) Trunk module F~bd F_bd F~3 F_3 F~bd F_bd F~3 F_3 cond(M^)cond( M) Fused QKV (W_ qkv) 0.9560.956 0.3430.343 0.1120.112 0.1530.153 4.084.08 FFN in-projection 0.07170.0717 0.6990.699 0.06570.0657 0.2240.224 3.233.23 FFN out-projection 0.03090.0309 0.4220.422 0.04680.0468 0.1600.160 9.219.21 Attention output 0.6760.676 0.8330.833 0.07610.0761 0.2150.215 4.814.81 Table S5: Damped-action alignment of the two stacked-trunk curvature families: the cosine (S4.17) between the approximate damped natural gradient (F~+γI)−1g( F+γ I)^-1g and the exact (Fscan+γI)−1g(F_scan+γ I)^-1g on the model’s true gradient, at damping γ=10−3γ=10^-3 ; larger is better. F~bd F_bd is the per-layer block diagonal ⨁kA^k⊗G^k _k A_k G_k; F~3 F_3 is the deployed three-factor block M^⊗A^⊗G M A G of (S4.16). Initialisation: random parameters on a 6464-site open-boundary Heisenberg chain. Trained: checkpoint step 30 00030\,000 with equilibrated MCMC walkers; entries are medians of per-system cosines across the 1111 training Hamiltonians, with factors aggregated across systems using the same exponential moving average as in training. Last column: median cond(M^)cond( M) at the trained point over the three systems with retained per-walker captures. Table S5 fixes why the trained point belongs to the third factor. First, FscanF_scan is dominated by cross-layer coupling: 0.860.86–0.870.87 of its squared Frobenius mass lies off the layer-diagonal, at initialisation and at the trained point alike. F~bd F_bd is layer-diagonal by construction, so no choice of per-layer factors touches this mass, while the off-diagonal blocks M^klA^⊗G M_kl\, A G of F~3 F_3 model it directly. Second, the trained M M is far from cILc\,I_L: across the four trunk modules cond(M^)cond( M) has medians 3.233.23–9.219.21. The trained regime is three-factor shaped because the residual stream makes both ends of every block gradient shared across the stack. Each stacked copy reads the same normalised stream state, so the activations aka_k are nearly independent of k; the stream update is r↦r+fk(r)r r+f_k(r), so the cotangents propagate through near-identity Jacobians, δk=Jk⊤δk+1 _k=J_k _k+1 with Jk=I+∂fk/∂rJ_k=I+∂ f_k/∂ r, and stay near-collinear in turn. The per-sample gradient of the k-th copy, gk=δkak⊤g_k= _ka_k , therefore shares a rank-one direction with layer-dependent amplitude, and every block of FscanF_scan factorises accordingly, gk≈αkδa⊤+εk⟹(Fscan)kl≈[αkαl]A⊗G,g_k\;≈\; _k\,δ\,a + _k (F_scan)_kl\;≈\;E [ _k _l ]\,A G, (S4.18) with δ, a shared within each sample and εk _k the idiosyncratic remainder — exactly (S4.16), with M=[αα⊤]M=E [α ] far from a multiple of the identity. Trained attention concentrates the tangents onto a few specialised low-rank directions, strengthening the shared component; the per-layer family spends its L-fold extra freedom fitting the remainders εk _k, i.e. batch noise. Initialisation favours per-layer curvature on the fused QKV block for an equally structural reason. At random parameters the blocks are weakly input-sensitive, since attention is near-uniform and ∂fk/∂r∂ f_k/∂ r is small, so per-block gradient statistics are near-Gaussian and second moments suffice: on precisely this block, A^k⊗G^k A_k G_k reproduces its diagonal Fisher block to a relative error of 1.19×10−31.19× 10^-3, and winning the layer-diagonal is everything a layer-diagonal preconditioner can win. Training grows the input sensitivity and installs the shared low-rank stream of (S4.18), which simultaneously breaks per-block Gaussianity, so second moments stop being sufficient, and organises the cross-layer mass into the coherent correlation that only the M factor models. The trained regime is where the optimiser spends the run; there the deployed block is both the better preconditioner and the cheaper one, by a factor of L. Appendix A The Casimir regulariser for nonlinear ansätze This appendix constructs a variational principle for ansätze that satisfy the per-site oddness (S1.39) but are nonlinear in the quaternions, so their support spreads across the half-integer sectors. This is the setting of the phase-space programme of [55], where maintaining a variational principle was left open; the closed-form couplings below resolve it. What oddness does not do is single out the lowest half-integer. The network may still place weight on ji=32j_i= 32 and above, or on superpositions across them. Pinning it to a 12 12 variational principle requires a second ingredient. The second ingredient leans on the shape of the Casimir spectrum. By (S1.26) the per-site Casimir S^l2=−Δl S_l^2=- _l reads off the sector: it acts as jl(jl+1)j_l(j_l+1), which oddness restricts to the values 34,154,354,… 34, 154, 354,… at jl=12,32,52,…j_l= 12, 32, 52,…. The target value 34 34 is the floor, and every step away from it grows quadratically in jlj_l. Oddness is also what lets us charge that gap linearly: on the odd subspace every site sits at S^l2≥34 S_l^2≥ 34, so the bare per-site deviation is non-negative, and so is its sum C^=∑l=1N(S^l2−34),C[ψ]=⟨ψ|C^|ψ⟩⟨ψ|ψ⟩=∑l=1Nx∼|ψ|2[−Δlψ(x)−34], C\;=\; _l=1^N ( S_l^2- 34 ), C[ψ]\;=\; ψ|\, C\,|ψ ψ|ψ \;=\; _l=1^NE_x |ψ|^2\! [\, - _l\,ψ(x)- 34\, ], (A.1) which vanishes exactly on the target sector (Proposition S1.1). Were the integer sectors still present, the j=0j=0 site at S^2=0 S^2=0 would sit below the floor and this penalty would reward it; removing them is precisely what makes the bare deviation a sound penalty, and is why we did not need to square it. The reason this rescues the variational principle is a race between two rates. An excursion into a higher sector lowers the energy only through the spin magnitudes jl(jl+1) j_l(j_l+1), which grow linearly in jlj_l, whereas the penalty charges the Casimir gap jl(jl+1)−34j_l(j_l+1)- 34, which grows quadratically. The gap outpaces the gain, so there is a finite multiplier λ above which every unit of energy bought by drifting off-target is overpaid by the penalty, and the regularised objective Eλ[ψ]=⟨ψ|H^|ψ⟩⟨ψ|ψ⟩+λC[ψ]E_λ[ψ]\;=\; ψ| H|ψ ψ|ψ \;+\;λ\,C[ψ] (A.2) stays at or above the physical ground-state energy E0E_0. Because oddness has already removed every integer sector, the nearest a state can sit to the target without lying in it is jl=32j_l= 32, the tightest case and the one that fixes how large λ must be. Evaluating C[ψ]C[ψ] costs nothing beyond the energy autodiff scheme above, since the log-derivative identity (S1.18) of § S1.2 turns each per-site Casimir ratio −Δlψ/ψ- _l\,ψ/ψ into logψ ψ and its first two automatic derivatives. Indeed these are the same quantities the local energy already uses. What remains is to compute, from the couplings (J,h)(J,h), a λ large enough that Eλ[ψ]≥E0E_λ[ψ]≥ E_0 for every per-site-odd ψ. In this appendix, we prove there is a λ such that Eλ[ψ]≥E0E_λ[ψ]≥ E_0 for every per-site-odd ψ for any given Hamiltonian. We then find a tighter bound in which the scalar λ is promoted to a per-edge tensor μ⋆ _ that saturates the variational principle, i.e. our variational space contains the ground state energy on the tigher bound only. Both are a function of a given Hamiltonian’s interaction data, J,hJ,h. We use two quantities of interest that can be efficiently computed. First, at each edge (i,k)(i,k) the relevant size of the 3×33× 3 interaction block Jik⊂J_ik⊂ J is its spectral norm, i.e. the top singular value, Aik:=‖Jik‖2=σ1(Jik),A_ik\;:=\; \|J_ik \|_2\;=\; _1 (J_ik ), (A.3) From it we form the per-site operator-norm weighted degree and the field magnitude, wdi:=∑k≠iAik,bi:=‖hi‖2.wd_i\;:=\; _k≠ iA_ik, b_i\;:=\; \|h_i \|_2. (A.4) The weighted degree wdiwd_i collects how strongly site i is coupled to its neighbours through any quadratic Pauli channel, and bib_i is the size of its transverse field. Both are closed-form functions of (J,h)(J,h), and can be efficiently computed once. Theorem A.1 (Sufficient coupling). Let C^=∑l(S^l2−34) C= _l( S_l^2- 34) and H(λ)=H^+λC^H(λ)= H+λ\, C, and write φ=(1+5)/2 =(1+ 5)/2 for the golden ratio. For every λ≥λmin(J,h)λ≥ _ (J,h), where λmin(J,h)=(1+ϵ)maxi∈[N][φ2wdi+23bi], _ (J,h)\;=\;(1+ε)\, _i∈[N] [\, 2\,wd_i\;+\; 23\,b_i\, ], (A.5) and every per-site-odd ψ∈L2((SU(2))N)ψ∈ L^2((SU(2))^N), ⟨ψ|H(λ)|ψ⟩≥E0‖ψ‖2, ψ|\,H(λ)\,|ψ \;≥\;E_0\,\|ψ\|^2, with E0E_0 the physical spin-12 12 ground-state energy. The constants φ2 2 and 23 23 are the maxima of three gain-to-penalty ratios over the allowed sectors. The slack ϵ>0ε>0 absorbs numerical safety holes (normalisation conventions, fp32 drift in wdiwd_i, the isotropy test at the 10−910^-9 threshold); we set ϵ=0.1ε=0.1 in all reported runs. The proof is given in Appendix B.3. For orientation, the 1D Heisenberg chain at unit coupling has wdi=2wd_i=2 in the bulk and no field, so (A.5) reads λmin=(1+ϵ)φ≈1.618(1+ϵ) _ =(1+ε)\, ≈ 1.618\,(1+ε). The remaining gap to the true physical threshold is over-conservatism, not a safety hole, and the next refinement closes most of it. The coefficient φ2 2 in (A.5) is the universal one: it rests on the spectral norm (A.3) alone and holds for every JikJ_ik. It is loose for structured couplings, and the implementation tightens it edge by edge. The only place the proof touched JikJ_ik was its Cauchy–Schwarz step (Appendix B.3, 3b), through the top singular value, so using more of the matrix sharpens the per-edge coefficient. Call α≥0α≥ 0 a certificate for a pair tensor J if ‖S→i⊤JS→k‖op,(si,sj)+‖S→i⊤JS→k‖op,(12,12)≤α[q(si)+q(sj)] \| S_i J S_k \|_op,(s_i,s_j)+ \| S_i J S_k \|_op,( 12, 12)\;≤\;α\, [\,q(s_i)+q(s_j)\, ] (A.6) for every non-trivial half-integer pair (si,sj)≠(12,12)(s_i,s_j)≠( 12, 12). This is exactly the per-edge condition the proof of Theorem A.1 verified for α=φ2‖J‖opα= 2\|J\|_op; any certificate may take its place, and three are worth recording. Lemma A.2 (Per-edge certificates). Write ‖J‖op=σ1(J)\|J\|_op= _1(J) and ‖J‖∗=∑aσa(J)\|J\|_*= _a _a(J) for the spectral and nuclear norms. Each of αop=φ2∥J∥op,αnuc=12∥J∥∗,αorth=34|λ|(for J=λQ,Q∈(3)), _op= 2\,\|J\|_op, _nuc= 12\,\|J\|_*, _orth= 34\,|λ|\ \ (for J=λ Q,\ Q (3)), is a certificate in the sense of (A.6), and certificates add: α(J(1)+J(2))≤α(J(1))+α(J(2))α(J^(1)+J^(2))≤α(J^(1))+α(J^(2)). The three certificates and their additivity are proved in Appendix B.4. Additivity turns the three primitives into a family. Peeling off an isotropic part J=τI+KJ=τ I+K with τ=13tr(J)τ= 13tr(J) (so τIτ I takes the orthogonal certificate with Q=IQ=I) and bounding the remainder K by the operator or nuclear certificate, or peeling off the nearest orthogonal part J=λQ+KpolJ=λ Q+K_pol with Q=UV⊤Q=UV from the SVD and λ=13tr(Q⊤J)λ= 13tr(Q J), gives six certificates whose minimum is used for each edge, αfull(J)=minφ2‖J‖op,12‖J‖∗,34|τ|+φ2‖K‖op,34|τ|+12‖K‖∗,34|λ|+φ2‖Kpol‖op,34|λ|+12‖Kpol‖∗ _full(J)= \ 2\|J\|_op,\; 12\|J\|_*,\; 34|τ|+ 2\|K\|_op,\; 34|τ|+ 12\|K\|_*,\; 34|λ|+ 2\|K_pol\|_op,\; 34|λ|+ 12\|K_pol\|_* \ (A.7) Theorem A.3 (Tight per-edge threshold). With αfull _full from (A.7), the threshold μ⋆(J,h)=(1+ϵ)maxi∈[N][∑k≠iαfull(Jik)+23bi] _ (J,h)\;=\;(1+ε)\, _i∈[N] [\, _k≠ i _full(J_ik)+ 23\,b_i\, ] (A.8) obeys μ⋆≤λmin _ ≤ _ and the conclusion of Theorem A.1: for every λ≥μ⋆λ≥ _ and per-site-odd ψ, ⟨ψ|H(λ)|ψ⟩≥E0‖ψ‖2 ψ|H(λ)|ψ ≥ E_0\,\|ψ\|^2. The proof is given in Appendix B.5. On the canonical couplings the minimum lands on the physically right primitive. Heisenberg J=J0IJ=J_0I gives αfull=34|J0| _full= 34|J_0| from the isotropic entry, recovering the textbook μ⋆=34wdi _ = 34wd_i, that is 1.51.5 on the 1D unit chain rather than the universal φ≈1.618 ≈ 1.618; pure Ising or XY (J rank one) drops to 12|J| 12|J| through the nuclear entry; pure Dzyaloshinskii–Moriya stays at φ2|D| 2|D|, the operator entry, since its singular values are dense. Appendix B Proofs from Supp. Mat. S1 The five results stated in SM § S1 and Appendix A are proved here: Proposition S1.1, Theorem S1.2, Theorems A.1 and A.3, and Lemma A.2. Proposition S1.1 is elementary and can be checked by hand. B.1 Proof of Proposition S1.1 Proof. By the Peter–Weyl expansion above, ψ=∑,,c,,Φ,,ψ= _ J, m, nc_ J, m, n\, _ J, m, n, with −Δk- _k diagonal of eigenvalue jk(jk+1)j_k(j_k+1) by (S1.26). For the forward direction, ( ∗ ‣ S1.3) subtracts to 0=∑,,c,,(jk(jk+1)−34)Φ,,,0\;=\; _ J, m, nc_ J, m, n\, (j_k(j_k+1)- 34 )\, _ J, m, n, (B.1) and orthogonality of the basis forces c,,=0c_ J, m, n=0 whenever jk≠12j_k≠ 12. Holding at every site, only =(12,…,12) J=( 12,…, 12) survives, and T,:=c(12,…,12),,T_ m, n:=c_( 12,…, 12), m, n gives the stated form. Conversely, applying −Δk- _k to the tensor form, the factors at sites r≠kr≠ k pass through and the site-k factor is a spin-12 12 Wigner matrix on which −Δk- _k acts as 34 34, so −Δkψ=34ψ- _k\,ψ= 34\,ψ at every site. ∎ B.2 Proof of Theorem S1.2 Proof. Schur orthogonality gives, on one copy of G, ∫GDmn1/2(q)¯Dm′n′1/2(q)q=12δmm′δnn′. _G D^1/2_mn(q)D^1/2_m n (q)\,dq= 12 _m _n . (B.2) Taking the product over N sites and using ‖χ‖=1\|χ\|=1 shows ‖ιχ(Ψ)‖L2(GN)=‖Ψ‖ℋN\| _χ( )\|_L^2(G^N)=\| \|_H_N; hence ιχ _χ is an isometry and in particular is injective. Peter–Weyl identifies the spin-12 12 sector with N⊗ℋNK_N _N, which is exactly ℱlinF_lin. Since dimιχ(ℋN)=2Nanddimℱlin=4N, _χ(H_N)=2^N _lin=4^N, (B.3) the first inclusion in Eq. (S1.32) is strict for N≥1N≥ 1. The central element −I∈G-I∈ G acts on VjV_j as (−1)2jI(-1)^2jI. Consequently a Peter–Weyl coefficient is odd under qi↦−qiq_i -q_i if and only if its site-i spin jij_i is half-integer. This proves the direct sum in Eq. (S1.30) and gives ℱlin⊆ℱoddF_lin _odd. The inclusion is strict because, for example, D3/2,3/23/2(q1)∏i=2ND1/2,1/21/2(qi)∈ℱodd∖ℱlin,D^3/2_3/2,3/2(q_1) _i=2^ND^1/2_1/2,1/2(q_i) _odd _lin, (B.4) with the product over i≥2i≥ 2 omitted when N=1N=1. On V1/2∗⊗V1/2V_1/2^* V_1/2, the right-regular differential action generated by the left-invariant fields acts on the second, column/physical factor and leaves the first, row/multiplicity factor invariant. Tensoring over sites therefore gives H~|ℱlin≅IN⊗HℋN. H |_F_lin I_K_N H_H_N. (B.5) Its lowest eigenvalue is E0E_0, which proves Eq. (S1.35); equality is attained by ιχ(Ψ0) _χ( _0) for any physical ground state Ψ0 _0. Finally, on a per-site-odd Peter–Weyl block labelled by =(j1,…,jN) j=(j_1,…,j_N), C^=∑i=1N(ji(ji+1)−34)I. C= _i=1^N (j_i(j_i+1)- 34 )I. (B.6) Every summand is nonnegative for half-integer ji≥12j_i≥ 12, and the sum vanishes if and only if every ji=12j_i= 12. Thus C^⪰0 C 0 on ℱoddF_odd and kerC^=ℱlin C=F_lin. Theorem A.3 gives Eq. (S1.37) for λ≥μ⋆(J,h)λ≥ _ (J,h). Conversely the lifted ground state ιχ(Ψ0) _χ( _0) lies in kerC C and has regularised quotient E0E_0, so the infimum over ℱoddF_odd is exactly E0E_0. ∎ B.3 Proof of Theorem A.1 (sufficient coupling) Proof. Write q(j):=j(j+1)−34q(j):=j(j+1)- 34 for the Casimir gap and r(j):=j(j+1)r(j):= j(j+1) for the Casimir root, and recall the Peter–Weyl decomposition of the per-site-odd subspace ℋodd=⨁s:ji∈12,32,52,…ℋs,ℋs:=⨂i=1N(Vji∗⊗Vji).H_odd= _s:j_i∈\ 12, 32, 52,…\H_s, _s:= _i=1^N (V_j_i^* V_j_i ). (1) H H and C C are block-diagonal in s. The spin operators S^ia S_i^a act through the right-regular differential representation generated by the left-invariant fields. This action preserves every Peter–Weyl isotypic component Vji∗⊗VjiV_j_i^* V_j_i and acts on its second factor. Hence each S^ia S_i^a acts within ℋsH_s, and H^=∑JijabS^iaS^jb+∑hiaS^ia H=Σ J_ij^ab S_i^a S_j^b+Σ h_i^a S_i^a preserves ℋsH_s. The penalty C^=∑l(S^l2−34) C= _l( S_l^2- 34) is a function of the per-site Casimirs, so it acts on ℋsH_s as the scalar C^≡Q(s)I C≡ Q(s)\,I with Q(s):=∑lq(jl)Q(s):= _lq(j_l). (2) Sector-gap upper bound (variational + purification, derived inline). Fix a sector s with raised set R:=R(s)=i:ji≠12R:=R(s)=\i:j_i≠ 12\ non-empty, and Rc:=[N]∖R^c:=[N] R. Split the Hamiltonian by which edges/fields touch R: H^=H^off+H^inc+H^fieldR, H\;=\; H_off\;+\; H_inc\;+\; H^R_field, with H^off H_off collecting pair terms with both endpoints in RcR^c and the field on RcR^c; H^inc H_inc collecting pair terms with at least one endpoint in R; and H^fieldR H^R_field the field on R. By construction H^off H_off acts trivially on the R factor and lives entirely on RcR^c. Upper bound on E0(ℋ12)E_0(H_ 12) (variational). Let |ϕ0⟩Rc| _0 _R^c be the spin-12 12 ground state of H^off H_off on ⨂i∈Rc(V1/2∗⊗V1/2), _i∈ R^c(V_1/2^* V_1/2), with energy E0off,1/2E_0^off,1/2. Let |ξ⟩R|ξ _R be any unit vector in ⨂i∈R(V1/2∗⊗V1/2). _i∈ R(V_1/2^* V_1/2). The product |ψ12var⟩:=|ϕ0⟩Rc⊗|ξ⟩R|ψ^var_ 12 :=| _0 _R^c |ξ _R lies in ℋ12H_ 12, so E0(ℋ12) E_0(H_ 12) ≤⟨ψ12var|H^|ψ12var⟩ \;≤\; ψ^var_ 12| H|ψ^var_ 12 =⟨ϕ0|H^off|ϕ0⟩+⟨ψ12var|(H^inc+H^fieldR)|ψ12var⟩ \;=\; _0| H_off| _0 + ψ^var_ 12|( H_inc+ H^R_field)|ψ^var_ 12 ≤E0off,1/2+‖H^inc+H^fieldR‖op,12. \;≤\;E_0^off,1/2\;+\; \| H_inc+ H^R_field \|_op, 12. The choice of |ξ⟩R|ξ _R is free; the inequality holds for any choice and hence for the supremum, the operator-norm upper bound is independent of |ξ⟩|ξ . Lower bound on E0(ℋs)E_0(H_s) (purification / partial trace). Let |ψs⟩| _s be the ground state of H H on ℋsH_s, energy E0(ℋs)E_0(H_s). Define the reduced density matrix ρRc:=trR|ψs⟩⟨ψs| _R^c:=tr_R| _s _s| on ⨂i∈Rc(V1/2⊗V1/2∗) _i∈ R^c(V_1/2 V_1/2^*) (since ji=12j_i= 12 for i∈Rci∈ R^c by definition of R). Then E0(ℋs) E_0(H_s) =⟨ψs|H^off|ψs⟩+⟨ψs|(H^inc+H^fieldR)|ψs⟩ \;=\; _s| H_off| _s \;+\; _s|( H_inc+ H^R_field)| _s =tr(H^offρRc)+⟨ψs|(H^inc+H^fieldR)|ψs⟩ \;=\;tr\! ( H_off\, _R^c )\;+\; _s|( H_inc+ H^R_field)| _s ≥E0off,1/2−‖H^inc+H^fieldR‖op,s. \;≥\;E_0^off,1/2\;-\; \| H_inc+ H^R_field \|_op,s. The first equality uses that H^off H_off acts on the RcR^c factor only, so ⟨H^off⟩=tr(H^offρRc) H_off =tr( H_off\, _R^c). The inequality uses (i) tr(H^offρRc)≥E0off,1/2⋅tr(ρRc)=E0off,1/2tr( H_off _R^c)≥ E_0^off,1/2·tr( _R^c)=E_0^off,1/2 (spectral lower bound on H^off H_off, which has spectrum ≥E0off,1/2≥ E_0^off,1/2 on the spin-12 12 sector where ρRc _R^c is supported); and (i) |⟨ψs|M|ψs⟩|≤‖M‖op,s| _s|M| _s |≤\|M\|_op,s for any operator M acting on ℋsH_s. Subtract: E0(ℋ12)−E0(ℋs)≤‖H^inc+H^fieldR‖op,12+‖H^inc+H^fieldR‖op,s⏟=:B(s).E_0(H_ 12)-E_0(H_s)\;≤\; \| H_inc+ H^R_field \|_op, 12+ \| H_inc+ H^R_field \|_op,s_=:\;B(s). The E0off,1/2E_0^off,1/2 pieces cancel exactly. (3) Per-edge bounds, derived inline. By triangle inequality on the sum defining H^inc+H^fieldR H_inc+ H^R_field, each side t∈12,st∈\ 12,s\ satisfies ‖H^inc+H^fieldR‖op,t≤∑(i,k) inc. to R‖S→i⊤JikS→k‖op,t+∑i∈R‖h→i⊤S→i‖op,t. \| H_inc+ H^R_field \|_op,t\;≤\; _(i,k) inc. to R \| S_i J_ik\, S_k \|_op,t\;+\; _i∈ R \| h_i S_i \|_op,t. Two factor bounds are needed; both follow from one-line spin-algebra calculations. (3a) Field op-norm: ‖h→i⊤S→i‖op,Vj=bi⋅j\| h_i S_i\|_op,V_j=b_i· j. Pick the orthogonal frame in which h→i h_i points along z z, so h→i⊤S→i=biSiz h_i S_i=b_i\,S_i^z. On VjV_j, SizS_i^z is diagonal with eigenvalues −j,−j+1,…,j\-j,-j+1,…,j\; its operator norm equals the maximum absolute eigenvalue, j. Rotational invariance of the operator norm on the spin representation gives the result for any direction of h→i h_i. (3b) Cauchy–Schwarz pair op-norm: ‖S→i⊤JikS→k‖op,Vji⊗Vjk≤Aikr(ji)r(jk)\| S_i J_ik S_k\|_op,V_j_i V_j_k≤ A_ik\,r(j_i)\,r(j_k). Take the real SVD Jik=UΣV⊤J_ik=U V with Σ=diag(σ1,σ2,σ3) =diag( _1, _2, _3) and σ1=Aik _1=A_ik; rewrite S→i⊤JikS→k=∑a=13σa(u→a⋅S→i)(v→a⋅S→k)=∑aσaAaBa, S_i J_ik\, S_k\;=\; _a=1^3 _a\,( u_a· S_i)\,( v_a· S_k)\;=\; _a _a\,A_a\,B_a, with Aa:=u→a⋅S→iA_a:= u_a· S_i, Ba:=v→a⋅S→kB_a:= v_a· S_k. For any unit |ψ⟩∈Vji⊗Vjk|ψ ∈ V_j_i V_j_k, two successive Cauchy–Schwarz steps give |⟨ψ|∑aσaAaBa|ψ⟩|≤∑aσa‖Aa|ψ⟩‖‖Ba|ψ⟩‖≤∑aσa2‖Aa|ψ⟩‖2∑a‖Ba|ψ⟩‖2. | ψ| _a _aA_aB_a|ψ |\;≤\; _a _a\,\|A_a|ψ \|\,\|B_a|ψ \|\;≤\; _a _a^2\|A_a|ψ \|^2\, _a\|B_a|ψ \|^2. For the first factor: ∑aσa2‖Aa|ψ⟩‖2≤σ12∑a⟨ψ|Aa†Aa|ψ⟩=Aik2⟨ψ|S→i⊤(UU⊤)S→i|ψ⟩=Aik2ji(ji+1) _a _a^2\|A_a|ψ \|^2≤ _1^2 _a ψ|A_a A_a|ψ =A_ik^2\, ψ| S_i (U ) S_i|ψ =A_ik^2\,j_i(j_i+1), where UU⊤=I3U =I_3 (orthogonal) and S→i⋅S→i=ji(ji+1)I S_i· S_i=j_i(j_i+1)I on VjiV_j_i (Casimir). For the second factor: ∑a‖Ba|ψ⟩‖2=⟨ψ|S→k⋅S→k|ψ⟩=jk(jk+1) _a\|B_a|ψ \|^2= ψ| S_k· S_k|ψ =j_k(j_k+1) by the same argument with VV⊤=I3V =I_3. Product: |⟨ψ|S→i⊤JikS→k|ψ⟩|≤Aikr(ji)r(jk)| ψ| S_i J_ik S_k|ψ |≤ A_ik\,r(j_i)\,r(j_k), which is the operator-norm bound by taking the supremum over unit |ψ⟩|ψ . Packaging. Using jie:=jij_i^e:=j_i for i∈Ri∈ R and 12 12 for i∈Rci∈ R^c, and adding the t=12t= 12 and t=st=s sides: B(s)≤∑(i,k) inc.Aik[r(jie)r(jke)+r(12)2]+∑i∈Rbi[jie+12].B(s)\;≤\; _(i,k) inc.A_ik\! [\,r(j_i^e)\,r(j_k^e)+r( 12)^2\, ]\;+\; _i∈ Rb_i\! [\,j_i^e+ 12\, ]. (4) Per-site accounting. Assign each contribution in the packaged B(s)B(s) to a raised endpoint, measured against the gap q(j)q(j) that the linear penalty collects there. The reader can read each ratio straight off the packaging. Boundary edges. One endpoint i raised, the other k at 12 12, so the contribution is Aik[r(ji)r(12)+r(12)2]A_ik\,[r(j_i)r( 12)+r( 12)^2]. Assigning it to i and dividing by q(ji)q(j_i) defines γbdy(j):=r(12)(r(j)+r(12))q(j). _bdy(j)\;:=\; r( 12)\, (r(j)+r( 12) )q(j). Internal edges. Both endpoints raised, contribution Aik[r(ji)r(jk)+r(12)2]A_ik\,[r(j_i)r(j_k)+r( 12)^2]. Apply AM–GM, r(ji)r(jk)≤12(ji(ji+1)+jk(jk+1))r(j_i)\,r(j_k)≤ 12(j_i(j_i+1)+j_k(j_k+1)), split the constant r(12)2=12r(12)2+12r(12)2r( 12)^2= 12r( 12)^2+ 12r( 12)^2 symmetrically, and split half to each endpoint; the i-side share, over q(ji)q(j_i), defines γint(j):=j(j+1)+342q(j). _int(j)\;:=\; j(j+1)+ 342\,q(j). Field at i∈Ri∈ R. Assigning bi(ji+12)b_i(j_i+ 12) to i and dividing by q(ji)q(j_i) defines γh(j):=j+12q(j). _h(j)\;:=\; j+ 12q(j). With γJ(j):=max(γbdy(j),γint(j)) _J(j):= ( _bdy(j), _int(j) ), every incident edge contributes at most AikγJ(ji)q(ji)A_ik\, _J(j_i)\,q(j_i) to the i-side share, and summing the incident edges into wdiwd_i, B(s)≤∑i∈R[γJ(ji)wdi+γh(ji)bi]q(ji).B(s)\;≤\; _i∈ R [\, _J(j_i)\,wd_i+ _h(j_i)\,b_i\, ]\,q(j_i). (5) Monotonicity: the worst harmonic component sits at j=32j= 32. Treat j as a continuous variable on (12,∞)( 12,∞) and write u:=j(j+1)u:=j(j+1), so q(j)=u−34q(j)=u- 34 and u′=2j+1>0u =2j+1>0. Each ratio is strictly decreasing. For γh(j)=(j+12)/q(j) _h(j)=(j+ 12)/q(j), the quotient rule with q′(j)=2j+1q (j)=2j+1 and 2(j+12)=2j+12(j+ 12)=2j+1 gives ∂jγh=q(j)−(j+12)(2j+1)q(j)2=q(j)−12(2j+1)2q(j)2=−j2−j−54q(j)2< 0. _j _h\;=\; q(j)-(j+ 12)(2j+1)q(j)^2\;=\; q(j)- 12(2j+1)^2q(j)^2\;=\; -\,j^2-j- 54q(j)^2\;<\;0. For γint(j)=(u+34)/(2(u−34)) _int(j)=(u+ 34)/ (2(u- 34) ), differentiating in u (sign-preserving since u′>0u >0), ∂uγint=(u−34)−(u+34)2(u−34)2=−322(u−34)2< 0. _u _int\;=\; (u- 34)-(u+ 34)2(u- 34)^2\;=\; - 322(u- 34)^2\;<\;0. For γbdy(j)=32(u+32)/(u−34) _bdy(j)= 32 ( u+ 32 )/(u- 34), the quotient rule on u (with ∂u=1/(2u) _u u=1/(2 u)) carries the numerator u−3/42u−u−32=−u+3/42u−32<0 u-3/42 u- u- 32=- u+3/42 u- 32<0, so ∂uγbdy<0 _u _bdy<0. All three decrease on (12,∞)( 12,∞), so γJ=max(γbdy,γint) _J= ( _bdy, _int) does too. Per-site oddness (§ S1.4) leaves j=32j= 32 as the smallest allowed value above 12 12, where the ratios reach their worst case, γbdy(32)=1+54=φ2,γint(32)=34,γh(32)=23, _bdy\! ( 32 )= 1+ 54= 2, _int\! ( 32 )= 34, _h\! ( 32 )= 23, so γJ(ji)≤φ2 _J(j_i)≤ 2 (the boundary ratio wins, φ2>34 2> 34) and γh(ji)≤23 _h(j_i)≤ 23 at every raised site. Substituting into the per-site bound of step (4), B(s) B(s) ≤∑i∈R[φ2wdi+23bi]q(ji) \;≤\; _i∈ R [\, 2\,wd_i+ 23\,b_i\, ]\,q(j_i) ≤(maxi[φ2wdi+23bi])∑i∈Rq(ji) \;≤\; ( _i [ 2\,wd_i+ 23\,b_i ] ) _i∈ Rq(j_i) =λmin(J,h)1+ϵQ(s)≤λmin(J,h)Q(s). \;=\; _ (J,h)1+ε\,Q(s)\;≤\; _ (J,h)\,Q(s). (6) Assemble. Decompose any per-site-odd ψ=∑sψsψ= _s _s over the orthogonal sectors. Using block-diagonality (1) on H H and the scalar action of C C from (1), ⟨ψ|H(λ)|ψ⟩ ψ|H(λ)|ψ =∑s[⟨ψs|H^|ψs⟩+λQ(s)‖ψs‖2] \;=\; _s [\, _s| H| _s +λ\,Q(s)\,\| _s\|^2\, ] (block-diagonal in s) ≥∑s[E0(ℋs)+λQ(s)]‖ψs‖2 \;≥\; _s [\,E_0(H_s)+λ\,Q(s)\, ]\,\| _s\|^2 (within-sector variational principle) ≥∑s[E0−B(s)+λQ(s)]‖ψs‖2 \;≥\; _s [\,E_0-B(s)+λ\,Q(s)\, ]\,\| _s\|^2 (step 2: E0(ℋs)≥E0−B(s)E_0(H_s)≥ E_0-B(s)) ≥∑sE0‖ψs‖2 \;≥\; _sE_0\,\| _s\|^2 (step 5 with λ≥λminλ≥ _ ) =E0‖ψ‖2. \;=\;E_0\,\|ψ\|^2. For the all-12 12 sector s=(12,…,12)s=( 12,…, 12) the raised set is empty, B(s)=0B(s)=0 and Q(s)=0Q(s)=0, and the chain collapses to the physical variational bound ⟨ψ12|H^|ψ12⟩≥E0‖ψ12‖2 _ 12| H| _ 12 ≥ E_0\,\| _ 12\|^2. ∎ B.4 Proof of Lemma A.2 (per-edge certificates) Proof. Operator. Step (3b) of Appendix B.3 gives ‖S→i⊤JS→k‖(si,sj)≤‖J‖opr(si)r(sj)\| S_i J S_k\|_(s_i,s_j)≤\|J\|_op\,r(s_i)r(s_j), and ≤‖J‖op34≤\|J\|_op\, 34 on (12,12)( 12, 12), so the left side of (A.6) is at most ‖J‖op[r(si)r(sj)+34]\|J\|_op[r(s_i)r(s_j)+ 34]. The ratio [r(si)r(sj)+34]/[q(si)+q(sj)][r(s_i)r(s_j)+ 34]/[q(s_i)+q(s_j)] over non-trivial pairs is largest at (32,12)( 32, 12), where it equals (15232+34)/3=(5+1)/4=φ2( 152 32+ 34)/3=( 5+1)/4= 2. Nuclear. For rank one J=σuv⊤J=σ\,uv , S→i⊤JS→k=σ(u⋅S→i)(v⋅S→k) S_i J S_k=σ\,(u\!·\! S_i)(v\!·\! S_k); each factor has operator norm s (eigenvalues −s,…,s-s,…,s), so the pair norm is ≤|σ|sisj≤|σ|\,s_is_j, while on (12,12)( 12, 12) the two ±12± 12 factors give |σ|/4|σ|/4. Then [sisj+14]/[q(si)+q(sj)]≤12[s_is_j+ 14]/[q(s_i)+q(s_j)]≤ 12, because q(si)+q(sj)−2(sisj+14)=(si−sj)2+(si+sj)−2≥0q(s_i)+q(s_j)-2(s_is_j+ 14)=(s_i-s_j)^2+(s_i+s_j)-2≥ 0 for every non-trivial half-integer pair (si+sj≥2s_i+s_j≥ 2). Summing the rank-one pieces of the SVD gives 12‖J‖∗ 12\|J\|_*. Orthogonal. For J=λQJ=λ Q, Q∈(3)Q (3), the rotated operators S~k=QS→k S_k=Q S_k keep the Casimir ∑a(S~ka)2=sk(sk+1) _a( S_k^a)^2=s_k(s_k+1), so S→i⋅S~k S_i\!·\! S_k is unitarily (detQ=+1 Q=+1) or anti-unitarily (detQ=−1 Q=-1) equivalent to S→i⋅S→k S_i\!·\! S_k, whose pair operator norm is sisj+min(si,sj)s_is_j+ (s_i,s_j) (the extreme total-spin multiplets). The ratio [sisj+min+34]/[q(si)+q(sj)][s_is_j+ + 34]/[q(s_i)+q(s_j)] peaks at (32,32)( 32, 32), value (94+32+34)/6=34( 94+ 32+ 34)/6= 34. Additivity is the triangle inequality applied to each operator-norm term in (A.6). ∎ B.5 Proof of Theorem A.3 (tight per-edge threshold) Proof. By Lemma A.2 and its additivity each entry of (A.7) is a certificate (the isotropic and polar entries split J and sum a primitive on each part), and the minimum of certificates is a certificate, so αfull(Jik) _full(J_ik) satisfies (A.6). Assigning it to the endpoints (a boundary edge has q(12)=0q( 12)=0, so its whole weight lands on the raised site) and keeping the field bound j+12≤23q(j)j+ 12≤ 23q(j) of Appendix B.3 (its step 5), steps (2) and (6) there go through with ∑kαfull(Jik) _k _full(J_ik) in place of φ2wdi 2wd_i. Since every entry of the minimum is ≤φ2‖Jik‖op≤ 2\|J_ik\|_op, we have μ⋆≤λmin _ ≤ _ . ∎ References [1] D. A. Abanin, E. Altman, I. Bloch, and M. Serbyn (2019) Colloquium: many-body localization, thermalization, and entanglement. Reviews of Modern Physics 91 (2), p. 021001. External Links: Document Cited by: §S3, §4. [2] I. Affleck, T. Kennedy, E. H. Lieb, and H. Tasaki (1988) Valence bond ground states in isotropic quantum antiferromagnets. Communications in Mathematical Physics 115, p. 477–528. Note: The related PRL titled ”Rigorous results on valence-bond ground states in antiferromagnets” is 1987, DOI 10.1103/PhysRevLett.59.799 External Links: Document Cited by: §S3, §4. [3] D. Aharonov (1999) Quantum computation. Annual Reviews of Computational Physics VI, p. 259–346. Cited by: §6. [4] Amazon Web Services (2026) Amazon braket pricing. Note: https://aws.amazon.com/braket/pricing/Accessed 3 August 2026 Cited by: §6. [5] P. W. Anderson (1958) Absence of diffusion in certain random lattices. Physical Review 109 (5), p. 1492–1505. External Links: Document Cited by: §S3, §4. [6] P. W. Anderson (1973) Resonating valence bonds: a new kind of insulator?. Materials Research Bulletin 8 (2), p. 153–160. External Links: Document Cited by: §S3, §4. [7] C. Andrieu and J. Thoms (2008) A tutorial on adaptive mcmc. Statistics and computing 18 (4), p. 343–373. Cited by: §S4.1. [8] Y. F. Atchadé, G. O. Roberts, and J. S. Rosenthal (2011) Towards optimal scaling of metropolis-coupled markov chain monte carlo. Statistics and Computing 21 (4), p. 555–568. Cited by: §S4.1.4. [9] A. Auerbach (1994) Interacting electrons and quantum magnetism. Graduate Texts in Contemporary Physics, Springer, New York. External Links: Document, ISBN 978-0-387-94286-5 Cited by: §S3, §4. [10] L. Babel, I. V. Chuvaeva, M. Klin, and D. V. Pasechnik (2010) Algebraic combinatorics in mathematical chemistry. methods and algorithms. i. program implementation of the weisfeiler-leman algorithm. arXiv preprint arXiv:1002.1921. Cited by: §S3.2. [11] L. Balents (2010) Spin liquids in frustrated magnets. Nature 464 (7286), p. 199–208. External Links: Document Cited by: §S3, §4. [12] D. M. Basko, I. L. Aleiner, and B. L. Altshuler (2006) Metal–insulator transition in a weakly interacting many-electron system with localized single-particle states. Annals of Physics 321 (5), p. 1126–1205. External Links: Document, cond-mat/0506617 Cited by: §S3, §4. [13] B. Bauer, S. Bravyi, M. Motta, and G. K. Chan (2020) Quantum algorithms for quantum chemistry and quantum materials science. Chemical reviews 120 (22), p. 12685–12717. Cited by: §1. [14] F. Becca and S. Sorella (2017) Quantum monte carlo approaches for correlated systems. Cambridge University Press. Cited by: §S1.1, §S1.1, §5.1, §5, §6. [15] H. Bernien, S. Schwartz, A. Keesling, H. Levine, A. Omran, H. Pichler, S. Choi, A. S. Zibrov, M. Endres, M. Greiner, V. Vuletić, and M. D. Lukin (2017) Probing many-body dynamics on a 51-atom quantum simulator. Nature 551 (7682), p. 579–584. External Links: Document Cited by: §S3, §4. [16] A. Botev, H. Ritter, and D. Barber (2017) Practical gauss-newton optimisation for deep learning. In International Conference on Machine Learning, p. 557–565. Cited by: §S4.3. [17] JAX: composable transformations of Python+NumPy programs External Links: Link Cited by: §1. [18] S. B. Bravyi and A. Yu. Kitaev (2002) Fermionic quantum computation. Annals of Physics 298 (1), p. 210–226. External Links: Document, quant-ph/0003137 Cited by: §S3, §4. [19] A. Browaeys and T. Lahaye (2020) Many-body physics with individually controlled rydberg atoms. Nature Physics 16, p. 132–142. External Links: Document, 2002.07413 Cited by: §S3, §4. [20] S. Byrne and M. Girolami (2013) Geodesic Monte Carlo on embedded manifolds. Scandinavian Journal of Statistics 40 (4), p. 825–845. Cited by: §S4.1.2. [21] J. Cai, M. Fürer, and N. Immerman (1992) An optimal lower bound on the number of variables for graph identification. Combinatorica 12 (4), p. 389–410. Cited by: §S2.3. [22] G. Carleo and M. Troyer (2017) Solving the quantum many-body problem with artificial neural networks. Science 355 (6325), p. 602–606. External Links: Document Cited by: §S1.1. [23] G. Carleo and M. Troyer (2017) Solving the quantum many-body problem with artificial neural networks. Science 355, p. 602–606. Note: arXiv:1606.02318 Cited by: §1, §S2.4, §2. [24] M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, and P. J. Coles (2021) Variational quantum algorithms. Nature Reviews Physics 3, p. 625–644. External Links: Document Cited by: §1. [25] X. Chen, Z. Gu, and X. Wen (2011) Classification of gapped symmetric phases in one-dimensional spin systems. Physical Review B 83 (3), p. 035107. External Links: Document, 1008.3745 Cited by: §S3, §4. [26] R. W. Chien, M. Chiew, B. Harrison, J. Necaise, W. Wang, M. Mudassar, C. McLauchlan, T. M. Henderson, G. E. Scuseria, S. Strelchuk, et al. (2026) Simulating fermions with a digital quantum computer. Nature Reviews Physics 8 (3), p. 131–145. Cited by: §1. [27] J. M. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar (2021) Gradient descent on neural networks typically occurs at the edge of stability. arXiv preprint arXiv:2103.00065. Cited by: §S4.3.5. [28] T. S. Cubitt, A. Montanaro, and S. Piddock (2018) Universal quantum hamiltonians. Proceedings of the National Academy of Sciences 115 (38), p. 9497–9502. Cited by: §S1.2. [29] Y. N. DauphVaswaniin, A. Fan, M. Auli, and D. Grangier (2017) Language modeling with gated convolutional networks. In Proceedings of the 34th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 70, p. 933–941. External Links: Link Cited by: §S2.2. [30] S. De and S. Smith (2020) Batch normalization biases residual blocks towards the identity function in deep networks. In Advances in Neural Information Processing Systems, Vol. 33, p. 19964–19975. External Links: Link Cited by: §S2.3. [31] D. Du and P. M. Pardalos (2013) Handbook of combinatorial optimization. Springer Science & Business Media. Cited by: §1. [32] I. Dzyaloshinsky (1958) A thermodynamic theory of “weak” ferromagnetism of antiferromagnetics. Journal of Physics and Chemistry of Solids 4 (4), p. 241–255. External Links: Document Cited by: §S3, §4. [33] F. Falcioni, E. Orlova, T. Heightman, P. Mantrov, and A. Ustimenko (2025) Benchmarking simulacra ai’s quantum accurate synthetic data generation for chemical sciences. arXiv preprint arXiv:2511.07433. Cited by: §S4.1. [34] I. Faradzev, A. A. Ivanov, M. Klin, and A. Woldar (2013) Investigations in algebraic theory of combinatorial objects. Springer Science & Business Media. Cited by: §S3.2. [35] G. B. Folland (2016) A course in abstract harmonic analysis. 2 edition, CRC Press, Boca Raton. External Links: Document, ISBN 9781498727136 Cited by: §2. [36] A. Foster, Z. Schätzle, P. B. Szabó, L. Cheng, J. Köhler, G. Cassella, N. Gao, J. Li, F. Noé, and J. Hermann (2025) An ab initio foundation model of wavefunctions that accurately describes chemical bond breaking. arXiv preprint arXiv:2506.19960. Cited by: §1. [37] V. François-Lavet, P. Henderson, R. Islam, M. G. Bellemare, and J. Pineau (2018) An introduction to deep reinforcement learning. Foundations and Trends in Machine Learning 11 (3–4), p. 219–354. External Links: Document, 1811.12560 Cited by: §3. [38] F. N. Fritsch and R. E. Carlson (1980) Monotone piecewise cubic interpolation. SIAM Journal on Numerical Analysis 17 (2), p. 238–246. Cited by: §S4.1.4. [39] N. Gao and S. Günnemann (2023) Generalizing neural wave functions. International Conference on Machine Learning (ICML). Note: arXiv:2302.04168 Cited by: §1. [40] N. Gao, J. Köhler, and A. Foster (2023) folx – forward laplacian for JAX. Note: https://github.com/microsoft/folxVersion 0.2.5 Cited by: §1, §S4.2.1. [41] D. García-Heredia, A. Alonso-Ayuso, and E. Molina (2019) A combinatorial model to optimize air traffic flow management problems. Computers & operations research 112, p. 104768. Cited by: §1. [42] L. Gerard, M. Scherbela, P. Marquetand, and P. Grohs (2022) Gold-standard solutions to the Schrödinger equation using deep learning: how much physics do we need?. Advances in Neural Information Processing Systems (NeurIPS) 35. Note: arXiv:2205.09438 Cited by: §1. [43] M. Girolami and B. Calderhead (2011) Riemann manifold Langevin and Hamiltonian Monte Carlo methods. Journal of the Royal Statistical Society, Series B 73 (2), p. 123–214. Cited by: §S4.1.2. [44] F. Glover, G. Kochenberger, and Y. Du (2018) A tutorial on formulating and using qubo models. arXiv preprint arXiv:1811.11538. Cited by: §5. [45] F. Glover, G. Kochenberger, R. Hennig, and Y. Du (2022) Quantum bridge analytics i: a tutorial on formulating and using qubo models. Annals of Operations Research 314 (1), p. 141–183. Cited by: §5. [46] J. Goldstone (1961) Field theories with superconductor solutions. Il Nuovo Cimento 19, p. 154–164. External Links: Document Cited by: §S3, §4. [47] S. Gong, W. Zhu, D. N. Sheng, O. I. Motrunich, and M. P. A. Fisher (2014) Plaquette ordered phase and quantum phase diagram in the spin-1/21/2 J1J_1–J2J_2 square heisenberg model. Physical Review Letters 113 (2), p. 027201. External Links: Document, 1311.5962 Cited by: §5.3. [48] D. M. Greenberger, M. A. Horne, and A. Zeilinger (1989) Going beyond bell’s theorem. In Bell’s Theorem, Quantum Theory and Conceptions of the Universe, M. Kafatos (Ed.), p. 69–72. External Links: Document Cited by: §S3, §4. [49] S. Gu (2010) Fidelity approach to quantum phase transitions. International Journal of Modern Physics B 24 (23), p. 4371–4458. Cited by: §5. [50] E. J. Gumbel (1954) Statistical theory of extreme values and some practical applications: a series of lectures. Vol. 33, US Government Printing Office. Cited by: §S2.5.1. [51] D. Ha, A. M. Dai, and Q. V. Le (2017) HyperNetworks. In International Conference on Learning Representations, External Links: Link Cited by: §S2.3.1. [52] B. C. Hall (2015) Lie groups, lie algebras, and representations: an elementary introduction. 2 edition, Graduate Texts in Mathematics, Vol. 222, Springer, Cham. External Links: Document Cited by: §2. [53] W. K. Hastings (1970) Monte carlo sampling methods using markov chains and their applications. Biometrika 57 (1), p. 97–109. Cited by: §S4.1. [54] F. He and R. Qu (2014) A two-stage stochastic mixed-integer program modelling and hybrid solution approach to portfolio selection problems. Information Sciences 289, p. 190–205. Cited by: §1. [55] T. Heightman, E. Jiang, R. Mora-Soto, M. Lewenstein, and M. Płodzień (2025) Quantum machine learning in multi-qubit phase-space Part I: foundations. arXiv preprint arXiv:2507.12117. Cited by: Appendix A, §1, §2, §2. [56] T. Heightman and M. Płodzień (2025) Deep learning in classical and quantum physics. arXiv preprint arXiv:2508.10666. Cited by: §S4.3. [57] T. Hensgens, T. Fujita, L. Janssen, X. Li, C. J. Van Diepen, C. Reichl, W. Wegscheider, S. Das Sarma, and L. M. K. Vandersypen (2017) Quantum simulation of a fermi-hubbard model using a semiconductor quantum dot array. Nature 548 (7665), p. 70–73. External Links: Document, 1702.07511 Cited by: §S3, §4. [58] J. Hermann, Z. Schätzle, and F. Noé (2020) Deep-neural-network solution of the electronic Schrödinger equation. Nature Chemistry 12, p. 891–897. Note: arXiv:1909.08423 Cited by: §1. [59] J. Hubbard (1963) Electron correlations in narrow energy bands. Proceedings of the Royal Society of London. Series A. Mathematical and Physical Sciences 276 (1365), p. 238–257. External Links: Document Cited by: §S3, §4. [60] K. Hukushima and K. Nemoto (1996) Exchange monte carlo method and application to spin glass simulations. Journal of the Physical Society of Japan 65 (6), p. 1604–1608. Cited by: §S4.1. [61] J. M. Hyman (1983) Accurate monotonicity preserving cubic interpolation. SIAM Journal on Scientific and Statistical Computing 4 (4), p. 645–654. Cited by: §S4.1.4. [62] E. Ising (1925) Beitrag zur theorie des ferromagnetismus. Zeitschrift für Physik 31, p. 253–258. External Links: Document Cited by: §S3, §4. [63] J. James, W. Yu, and J. Gu (2019) Online vehicle routing with neural combinatorial optimization and deep reinforcement learning. IEEE Transactions on Intelligent Transportation Systems 20 (10), p. 3806–3817. Cited by: §1. [64] D. Jansen, D. Farina, L. Mortimer, T. Heightman, A. Leitherer, P. Mujal, J. Wang, and A. Acín (2026) Mapping phase diagrams of quantum spin systems through semidefinite-programming relaxations. Physical Review Letters 136 (5), p. 050401. Cited by: §6. [65] D. Jiang, X. Wen, Y. Chen, R. Li, W. Fu, H. Q. Pham, J. Chen, D. He, W. A. Goddard, L. Wang, and W. Ren (2025) Neural scaling laws surpass chemical accuracy for the many-electron Schrödinger equation. arXiv preprint arXiv:2508.02570. Cited by: §1. [66] P. Jordan and E. P. Wigner (1928) Über das paulische Äquivalenzverbot. Zeitschrift für Physik 47 (9–10), p. 631–651. Note: English title: About the Pauli exclusion principle External Links: Document Cited by: §S3, §4. [67] J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, et al. (2021) Highly accurate protein structure prediction with alphafold. nature 596 (7873), p. 583–589. Cited by: §S2.3. [68] L. P. Kadanoff (1966) Scaling laws for ising models near t c. Physics Physique Fizika 2 (6), p. 263. Cited by: §S2.4.4. [69] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei (2020) Scaling laws for neural language models. arXiv preprint arXiv:2001.08361. Cited by: §5.1, §6. [70] J. Kempe, A. Kitaev, and O. Regev (2006) The complexity of the local hamiltonian problem. Siam journal on computing 35 (5), p. 1070–1097. Cited by: §S1.2. [71] A. Kitaev and S. J. Suh (2018) The soft mode in the Sachdev–Ye–Kitaev model and its gravity dual. Journal of High Energy Physics 2018 (5), p. 183. External Links: Document, 1711.08467 Cited by: §S3, §4. [72] A. Yu. Kitaev (2001) Unpaired majorana fermions in quantum wires. Physics-Uspekhi 44 (10S), p. 131–136. External Links: Document, cond-mat/0010440 Cited by: §S3, §4. [73] A. Kitaev (2006) Anyons in an exactly solved model and beyond. Annals of Physics 321 (1), p. 2–111. External Links: Document, cond-mat/0506438 Cited by: §S3, §4. [74] A. W. Knapp (2001) Representation theory of semisimple groups: an overview based on examples. Princeton Mathematical Series, Vol. 36, Princeton University Press, Princeton, NJ. Note: Paperback reprint with a new preface; originally published in 1986 External Links: ISBN 9780691090894 Cited by: §2. [75] J. B. Kogut and L. Susskind (1975) Hamiltonian formulation of wilson’s lattice gauge theories. Physical Review D 11 (2), p. 395–408. External Links: Document Cited by: §S3, §4. [76] J. B. Kogut (1979) An introduction to lattice gauge theory and spin systems. Reviews of Modern Physics 51 (4), p. 659–713. External Links: Document Cited by: §S3, §4. [77] K. I. Kugel and D. I. Khomskii (1982) The jahn-teller effect and magnetism: transition metal compounds. Soviet Physics Uspekhi 25 (4), p. 231–256. External Links: Document Cited by: §S3, §4. [78] H. Lange, A. Van de Walle, A. Abedinnia, and A. Bohrdt (2024) From architectures to applications: a review of neural quantum states. Quantum Science and Technology 9 (4), p. 040501. Cited by: §1. [79] M. Larocca, S. Thanasilp, S. Wang, K. Sharma, J. Biamonte, P. J. Coles, L. Cincio, J. R. McClean, Z. Holmes, and M. Cerezo (2025) Barren plateaus in variational quantum computing. Nature Reviews Physics 7 (4), p. 174–189. Cited by: §1. [80] U. Las Heras, A. Mezzacapo, L. Lamata, S. Filipp, A. Wallraff, and E. Solano (2014) Digital quantum simulation of spin systems in superconducting circuits. Physical Review Letters 112 (20), p. 200501. External Links: Document, 1311.7626 Cited by: §S3, §4. [81] M. Ledoux, I. Nourdin, and G. Peccati (2015) Stein’s method, logarithmic sobolev and transport inequalities. Geometric and Functional Analysis 25 (1), p. 256–306. Cited by: §6. [82] R. Li, H. Ye, D. Jiang, X. Wen, C. Wang, Z. Li, X. Li, D. He, J. Chen, W. Ren, et al. (2024) A computational framework for neural network-based variational monte carlo with forward laplacian. Nature Machine Intelligence 6 (2), p. 209–219. Cited by: §S4.2.1. [83] E. Lieb, T. Schultz, and D. Mattis (1961) Two soluble models of an antiferromagnetic chain. Annals of Physics 16 (3), p. 407–466. External Links: Document Cited by: §S3, §4. [84] H. J. Lipkin, N. Meshkov, and A. J. Glick (1965) Validity of many-body approximation methods for a solvable model. I. exact solutions and perturbation theory. Nuclear Physics 62 (2), p. 188–198. External Links: Document Cited by: §S3, §4. [85] F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf (2020) Object-centric learning with slot attention. Advances in neural information processing systems 33, p. 11525–11538. Cited by: §S2.3. [86] I. Loshchilov, C. Hsieh, S. Sun, and B. Ginsburg (2025) nGPT: normalized transformer with representation learning on the hypersphere. In International Conference on Learning Representations, External Links: Link Cited by: §S2.3. [87] C. K. Majumdar and D. K. Ghosh (1969) On next-nearest-neighbor interaction in linear chain. I. Journal of Mathematical Physics 10 (8), p. 1388–1398. External Links: Document Cited by: §S3, §4. [88] J. Maldacena and D. Stanford (2016) Remarks on the sachdev-ye-kitaev model. Physical Review D 94 (10), p. 106002. External Links: Document, 1604.07818 Cited by: §S3, §4. [89] H. Maron, H. Ben-Hamu, H. Serviansky, and Y. Lipman (2019) Provably powerful graph networks. Advances in neural information processing systems 32. Cited by: §S2.3. [90] J. Martens, J. Ba, and M. Johnson (2018) Kronecker-factored curvature approximations for recurrent neural networks. In International Conference on Learning Representations, Cited by: §S4.3.2. [91] J. Martens and R. Grosse (2015) Optimizing neural networks with kronecker-factored approximate curvature. In International conference on machine learning, p. 2408–2417. Cited by: §1, §S4.3, footnote 1. [92] S. McArdle, S. Endo, A. Aspuru-Guzik, S. C. Benjamin, and X. Yuan (2020) Quantum computational chemistry. Reviews of Modern Physics 92 (1), p. 015003. Cited by: §1. [93] J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven (2018) Barren plateaus in quantum neural network training landscapes. Nature Communications 9, p. 4812. External Links: Document Cited by: §S1.1, §1. [94] N. Metropolis, A. W. Rosenbluth, M. N. Rosenbluth, A. H. Teller, and E. Teller (1953) Equation of state calculations by fast computing machines. The journal of chemical physics 21 (6), p. 1087–1092. Cited by: §S4.1. [95] C. Monroe, W. C. Campbell, L.-M. Duan, Z.-X. Gong, A. V. Gorshkov, P. W. Hess, R. Islam, K. Kim, N. M. Linke, G. Pagano, P. Richerme, C. Senko, and N. Y. Yao (2021) Programmable quantum simulations of spin systems with trapped ions. Reviews of Modern Physics 93 (2), p. 025001. External Links: Document, 1912.07845 Cited by: §S3, §4. [96] T. Moriya (1960) Anisotropic superexchange interaction and weak ferromagnetism. Physical Review 120 (1), p. 91–98. External Links: Document Cited by: §S3, §4. [97] V. Murg, Ö. Legeza, R. M. Noack, and F. Verstraete (2010) Simulating strongly correlated quantum systems with tree tensor networks. Physical Review B 82 (20), p. 205105. External Links: Document, 1006.3095 Cited by: §1, §3. [98] N. Nakatani and G. K. Chan (2013) Efficient tree tensor network states (TTNS) for quantum chemistry: generalizations of the density matrix renormalization group algorithm. Journal of Chemical Physics 138 (13), p. 134113. External Links: Document, 1302.2298 Cited by: §1, §3. [99] Y. Nambu (1960) Quasi-particles and gauge invariance in the theory of superconductivity. Physical Review 117 (3), p. 648–663. External Links: Document Cited by: §S3, §4. [100] R. Nandkishore and D. A. Huse (2015) Many-body localization and thermalization in quantum statistical mechanics. Annual Review of Condensed Matter Physics 6 (1), p. 15–38. External Links: Document, 1404.0686 Cited by: §S3, §4. [101] K. Nazaryan and L. Fu (2026) QERNEL: a scalable large electron model. arXiv preprint arXiv:2604.26018. Cited by: §1. [102] M. A. Nielsen I. L. Chuang et al. (2000) Quantum computation and quantum information. Vol. 1, Cambridge university press Cambridge. Cited by: §2, §6. [103] Z. Nussinov and J. van den Brink (2015) Compass models: theory and physical motivations. Reviews of Modern Physics 87 (1), p. 1–59. External Links: Document, 1303.5922 Cited by: §S3, §4. [104] K. Ohno (1964) Some remarks on the pariser–parr–pople method. Theoretica Chimica Acta 2, p. 219–227. External Links: Document Cited by: §5.3. [105] R. Oliveira and B. M. Terhal (2005) The complexity of quantum spin systems on a two-dimensional square lattice. arXiv preprint quant-ph/0504050. Cited by: §S1.2. [106] E. Orlova, A. Ustimenko, R. Jiang, P. Y. Lu, and R. Willett (2023) Deep stochastic mechanics. arXiv preprint arXiv:2305.19685. Cited by: §6. [107] D. N. Page (1993) Average entropy of a subsystem. Physical Review Letters 71 (9), p. 1291–1294. External Links: Document, gr-qc/9305007 Cited by: §S3, §4. [108] R. Pariser and R. G. Parr (1953) A semi-empirical theory of the electronic spectra and electronic structure of complex unsaturated molecules. i.. The Journal of Chemical Physics 21 (3), p. 466–471. Cited by: §5.6, Table 1, Table 1, §5. [109] R. Pariser and R. G. Parr (1953) A semi-empirical theory of the electronic spectra and electronic structure of complex unsaturated molecules. i. The Journal of Chemical Physics 21 (5), p. 767–776. Cited by: §5.6, Table 1, Table 1, §5. [110] E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. Courville (2018) FiLM: visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32, p. 3942–3951. External Links: Document Cited by: §S2.3.1. [111] A. Peruzzo, J. McClean, P. Shadbolt, M. Yung, X. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’Brien (2014) A variational eigenvalue solver on a photonic quantum processor. Nature Communications 5, p. 4213. External Links: Document Cited by: §1. [112] F. Peter and H. Weyl (1927) Die vollständigkeit der primitiven darstellungen einer geschlossenen kontinuierlichen gruppe. Mathematische Annalen 97, p. 737–755. External Links: Link Cited by: §2. [113] D. Pfau, J. S. Spencer, A. G. d. G. Matthews, and W. M. C. Foulkes (2020) Ab-initio solution of the many-electron Schrödinger equation with deep neural networks. Phys. Rev. Research 2, p. 033429. Note: arXiv:1909.02487 Cited by: §1, §S4.3. [114] J. A. Pople (1953) Electron interaction in unsaturated hydrocarbons. Transactions of the Faraday Society 49, p. 1375–1385. Cited by: §5.6, Table 1, Table 1, §5. [115] O. Press, N. A. Smith, and M. Lewis (2021) Shortformer: better language modeling using shorter inputs. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), p. 5493–5505. Cited by: §S2.4.4. [116] O. Press, N. A. Smith, and M. Lewis (2021) Train short, test long: attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409. Cited by: §S2.4.4, §S2.4.4, §S2.4.4. [117] M. Ragone, B. N. Bakalov, F. Sauvage, A. F. Kemper, C. Ortiz Marrero, M. Larocca, and M. Cerezo (2024) A lie algebraic theory of barren plateaus for deep parameterized quantum circuits. Nature Communications 15 (1), p. 7172. Cited by: §1. [118] R. Rende, L. L. Viteritti, F. Becca, A. Scardicchio, A. Laio, and G. Carleo (2025) Foundation neural-network quantum states as a unified ansatz for multiple Hamiltonians. Nature Communications 16, p. 7213. External Links: Document, 2502.09488 Cited by: §1, §1. [119] R. Rende, S. Goldt, F. Becca, and L. L. Viteritti (2024) Fine-tuning neural network quantum states. Phys. Rev. Research 6, p. 043280. Note: arXiv:2403.07795 Cited by: §1, §2. [120] R. Rende, A. Sinibaldi, L. L. Viteritti, R. Wiersema, A. Georges, and G. Carleo (2026) Scaling laws for neural-network quantum states. arXiv preprint arXiv:2606.02794. Cited by: §1. [121] R. Rende, L. L. Viteritti, F. Becca, A. Scardicchio, A. Laio, and G. Carleo (2025) Foundation neural-network quantum states as a unified ansatz for multiple hamiltonians. Nature Communications 16, p. 7213. Note: arXiv:2502.09488 Cited by: §1, §2. [122] M. J. Rice and E. J. Mele (1982) Elementary excitations of a linearly conjugated diatomic polymer. Physical Review Letters 49 (19), p. 1455–1459. External Links: Document Cited by: §S3, §4. [123] H. Robbins and S. Monro (1951) A stochastic approximation method. The annals of mathematical statistics, p. 400–407. Cited by: §S4.1.6. [124] G. O. Roberts, A. Gelman, and W. R. Gilks (1997) Weak convergence and optimal scaling of random walk metropolis algorithms. The annals of applied probability 7 (1), p. 110–120. Cited by: §S4.1.5, §S4.1. [125] G. O. Roberts and J. S. Rosenthal (1998) Optimal scaling of discrete approximations to langevin diffusions. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 60 (1), p. 255–268. Cited by: §S4.1.6, §S4.1. [126] G. O. Roberts and R. L. Tweedie (1996) Exponential convergence of langevin distributions and their discrete approximations. Bernoulli 2 (4), p. 341–363. Cited by: §S4.1. [127] S. Roca-Jerat, M. Gallego, F. Luis, J. Carrete, and D. Zueco (2024) Transformer wave function for quantum long-range models. Phys. Rev. B 110, p. 205147. Note: arXiv:2407.04773 Cited by: §1, §2. [128] S. Sachdev and J. Ye (1993) Gapless spin-fluid ground state in a random quantum heisenberg magnet. Physical Review Letters 70 (21), p. 3339–3342. External Links: Document, cond-mat/9212030 Cited by: §S3, §4. [129] S. Sachdev (1999) Quantum phase transitions. Cambridge University Press, Cambridge. External Links: ISBN 9780521582548, Document Cited by: §S3, §4. [130] A. W. Sandvik (2007) Evidence for deconfined quantum criticality in a two-dimensional heisenberg model with four-spin interactions. Physical Review Letters 98 (22), p. 227202. External Links: Document, cond-mat/0611343 Cited by: §S3, §4. [131] L. Savary and L. Balents (2017) Quantum spin liquids: a review. Reports on Progress in Physics 80 (1), p. 016502. Note: Published online 8 November 2016 External Links: Document Cited by: §S3, §4. [132] M. Scherbela, L. Gerard, and P. Grohs (2023) Towards a foundation model for neural network wavefunctions. arXiv preprint arXiv:2303.09949. Cited by: §1. [133] M. Scherbela, R. Reisenhofer, L. Gerard, P. Marquetand, and P. Grohs (2022) Solving the electronic Schrödinger equation for multiple nuclear geometries with weight-sharing deep neural networks. Nature Computational Science 2, p. 331–341. Note: arXiv:2105.08351 Cited by: §1. [134] T. Senthil, A. Vishwanath, L. Balents, S. Sachdev, and M. P. A. Fisher (2004) Deconfined quantum critical points. Science 303 (5663), p. 1490–1494. External Links: Document, cond-mat/0311326 Cited by: §S3, §4. [135] N. Shazeer (2020) GLU variants improve transformer. arXiv preprint arXiv:2002.05202. External Links: Link Cited by: §S2.2. [136] Y. Shi, L. Duan, and G. Vidal (2006) Classical simulation of quantum many-body systems with a tree tensor network. Physical Review A 74 (2), p. 022320. External Links: Document Cited by: §1, §S2.4.3, §S2.4, §3. [137] G. Singh and M. Rizwanullah (2022) Combinatorial optimization of supply chain networks: a retrospective & literature review. Materials today: proceedings 62, p. 1636–1642. Cited by: §1. [138] Z. G. Soos and S. Ramasesha (1984) Valence-bond theory of linear hubbard and pariser–parr–pople models. Physical Review B 29, p. 5410–5422. External Links: Document Cited by: §5.3. [139] S. Sorella (1998) Green function monte carlo with stochastic reconfiguration. Physical review letters 80 (20), p. 4558. Cited by: §S4.3. [140] W. P. Su, J. R. Schrieffer, and A. J. Heeger (1979) Solitons in polyacetylene. Physical Review Letters 42 (25), p. 1698–1701. External Links: Document Cited by: §S3, §4. [141] R. S. Sutton and A. G. Barto (1998) Reinforcement learning: an introduction. Adaptive Computation and Machine Learning, MIT Press, Cambridge, MA. External Links: ISBN 9780262193986 Cited by: §S2.5.3, §S2.5, §3. [142] R. H. Swendsen and J. Wang (1986) Replica monte carlo simulation of spin-glasses. Physical review letters 57 (21), p. 2607. Cited by: §S4.1. [143] S. Syed, A. Bouchard-Côté, G. Deligiannidis, and A. Doucet (2022) Non-reversible parallel tempering: a scalable highly parallel mcmc scheme. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 84 (2), p. 321–350. Cited by: §S4.1.3, §S4.1.3, §S4.1.4. [144] E. S. Sørensen, A. Catuneanu, J. S. Gordon, and H. Kee (2021) Heart of entanglement: chiral, nematic, and incommensurate phases in the kitaev-gamma ladder in a field. Physical Review X 11 (1), p. 011013. Cited by: §2. [145] N. Touati-Moungla V. Jost et al. (2012) Combinatorial optimization for electric vehicles management. Journal of Energy and Power Engineering 6 (5), p. 738–743. Cited by: §1. [146] A. Tranter, P. J. Love, F. Mintert, and P. V. Coveney (2018) A comparison of the bravyi–kitaev and jordan–wigner transformations for the quantum simulation of quantum chemistry. Journal of Chemical Theory and Computation 14 (11), p. 5617–5630. External Links: Document, 1812.02233 Cited by: §S3, §4. [147] C. J. Turner, A. A. Michailidis, D. A. Abanin, M. Serbyn, and Z. Papić (2018) Weak ergodicity breaking from quantum many-body scars. Nature Physics 14, p. 745–749. External Links: Document Cited by: §S3, §4. [148] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. Advances in neural information processing systems 30. Cited by: §S2.3, §S2.4.1. [149] G. Vidal (2007) Entanglement renormalization. Physical Review Letters 99 (22), p. 220405. External Links: Document Cited by: §S2.4.4. [150] O. Vinyals, M. Fortunato, and N. Jaitly (2015) Pointer networks. In Advances in Neural Information Processing Systems, Vol. 28. External Links: Link Cited by: §S2.5.1, §S2.5.1. [151] L. L. Viteritti, R. Rende, G. Bracci-Testasecca, J. Niedda, R. Moessner, G. Carleo, and A. Scardicchio (2025) Quantum spin glass in the two-dimensional disordered heisenberg model via foundation neural-network quantum states. arXiv preprint arXiv:2507.05073. Cited by: §1, §2. [152] L. L. Viteritti, R. Rende, S. Sachdev, and G. Carleo (2026) Approaching the thermodynamic limit with neural-network quantum states. arXiv preprint arXiv:2602.02665. Cited by: §1. [153] R. J. Williams (1992) Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 8 (3), p. 229–256. Cited by: §S2.5.3. [154] K. G. Wilson (1974) Confinement of quarks. Physical Review D 10 (8), p. 2445–2459. External Links: Document Cited by: §S3, §4. [155] K. G. Wilson (1975) The renormalization group: critical phenomena and the Kondo problem. Reviews of Modern Physics 47 (4), p. 773–840. External Links: Document Cited by: §S2.4.4. [156] K. G. Wilson (1983) The renormalization group and critical phenomena. Reviews of Modern Physics 55 (3), p. 583. Cited by: §S2.4. [157] D. Wu, R. Rossi, F. Vicentini, N. Astrakhantsev, F. Becca, X. Cao, J. Carrasquilla, F. Ferrari, A. Georges, M. Hibat-Allah, et al. (2024) Variational benchmarks for quantum many-body problems. Science 386 (6719), p. 296–301. Cited by: §5. [158] C. N. Yang (1962) Concept of off-diagonal long-range order and the quantum phases of liquid He and of superconductors. Reviews of Modern Physics 34 (4), p. 694–704. External Links: Document Cited by: §S3, §4. [159] H. Yang, J. Liang, and Q. Cui (2023) First-principles calculations for dzyaloshinskii–moriya interaction. Nature Reviews Physics 5 (1), p. 43–61. Cited by: §2. [160] G. Yu (2013) Industrial applications of combinatorial optimization. Springer Science & Business Media. Cited by: §1. [161] T. Zaklama, M. Geier, and L. Fu (2026) Large electron model: a universal ground state predictor. arXiv preprint arXiv:2603.02346. Cited by: §1. [162] T. Zaklama, D. Guerci, and L. Fu (2025) Attention-based foundation model for quantum states. arXiv preprint arXiv:2512.11962. Cited by: §1, §1, §2. [163] B. Zhang and R. Sennrich (2019) Root mean square layer normalization. Advances in neural information processing systems 32. Cited by: §S2.2. [164] S. Zhang, G. B. Halász, W. Zhu, and C. D. Batista (2021) Variational study of the kitaev-heisenberg-gamma model. Physical Review B 104 (1), p. 014411. Cited by: §2. [165] Y. Zhang and M. Di Ventra (2023) Transformer quantum state: a multi-purpose model for quantum many-body problems. Phys. Rev. B 107, p. 075147. Note: arXiv:2208.01758 Cited by: §1, §2.