Paper deep dive
Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superiority, in Homeostatic Compute Control
Francisco M. Arrabal-Campos, Ignacio Fernandez, Francisco G. Montoya, Alfredo Alcayde
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/29/2026, 4:23:07 AM
Summary
This paper investigates whether a dynamic internal field, specifically the Homeostatic Background Processor (HBP), can serve as a metacognitive governor for Transformer models by modulating computation (halting, gating) without performing reasoning. The HBP is defined by a family of PDEs on the module graph's Laplacian. The authors prove a discrete Schur-Cohn criterion for stability. Empirical results on NC1-hard tasks show that the specific physics (substance) of the field does not affect accuracy, and while second-order structure offers some robustness, a matched-interface GRU performs comparably. The key distinction of the HBP is its certifiable stability, not superior cognitive performance.
Entities (8)
Relation Signals (6)
Homeostatic Background Processor → governs → Transformer
confidence 95% · The HBP modulates a transformer’s cognition without performing it.
Homeostatic Background Processor → uses → Graph Laplacian
confidence 92% · HBP is a field defined over the module graph... governed by a family of partial differential equations on the graph Laplacian
Schur-Cohn criterion → certifies → Homeostatic Background Processor
confidence 90% · We certify the stability of the integrator... a discrete Schur-Cohn criterion... necessary and sufficient
Homeostatic Background Processor → uses → Verlet integrator
confidence 88% · discrete Schur-Cohn criterion for Verlet with velocity coupling... shows that a backward-differenced gyroscopic term does not inherit the neutrality
GRU → performscomparablyto → Homeostatic Background Processor
confidence 85% · a matched-interface GRU is indistinguishable in the first and nominally exceeds the field in the second
Homeostatic Background Processor → testedon → NC1-hard
confidence 85% · Empirically (S5, NC1-hard...)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:An intelligent system does not merely reason: it governs its own reasoning - how much to compute, when to stop, which module to activate. Can that role be played by a dynamic internal field - a low-dimensional homeostatic state with explicit physics and certified stability - that modulates cognition without performing it? Ours is a field on the module graph governed by a family of PDEs on the graph Laplacian, advancing with an adaptive-depth reasoner. We certify the stability of the integrator of the whole family - an integrator certificate, not a closed-loop one. New, and proved here: a discrete Schur-Cohn criterion for Verlet with velocity coupling, necessary and sufficient per latent root, with no commutation hypothesis. The answer is threefold: substance no, structure only in part, certifiability yes. The type of the field's physics is irrelevant for accuracy: wave, diffusion, gated mixtures and a 2D Navier-Stokes substrate tie. A twenty-seed preregistered deconfounding campaign bounds the structural claim: at equalized caps the second-order effect is strong in one family (+0.087 [+0.042, +0.132], t=4.0) but is not detected in the other (+0.014 [-0.013, +0.040], n.s.), so part of the original contrast was capacity, not order; and a matched-interface GRU is indistinguishable in the first and nominally exceeds the field in the second (-0.035 [-0.067, -0.002]). What distinguishes the field is not capability but that its one-step operator admits an exact runtime stability check - a difference of kind, not of existence: learned recurrences carry certificates too, sufficient and conservative ones. A kill-gate with a positive control finds no evidence for the field as evidence accumulator (Delta AUC +0.0007 [-0.0065, +0.0079] vs a 0.03 threshold). A dynamic internal field is a viable, certifiable compute governor, but not an enhancer of cognition: it modulates, it does not think.
Tags
Links
- Source: https://arxiv.org/abs/2608.24319v1
- Canonical: https://arxiv.org/abs/2608.24319v1
Trouble viewing inline? Open PDF directly →
Full Text
99,794 characters extracted from source content.
Expand or collapse full text
Can a Dynamic Internal Field Govern a Transformer’s Cognition? Certifiability, not Superiority, in Homeostatic Compute Control M. Arrabal-Campos∗ @ual.es Affiliation: of Chemistry and Physics, Research Centre CIAIMBITAL Affiliation: of Engineering, Research Centre CIAIMBITAL Affiliation: of Almería, 04120 Almería, Spain (∗corresponding author) G. Montoya @ual.es Affiliation: of Engineering, Research Centre CIAIMBITAL Affiliation: of Almería, 04120 Almería, Spain Alcayde @ual.es Affiliation: of Engineering, Research Centre CIAIMBITAL Affiliation: of Almería, 04120 Almería, Spain Fernández @ual.es Affiliation: of Chemistry and Physics, Research Centre CIAIMBITAL Affiliation: of Almería, 04120 Almería, Spain Abstract An intelligent system does not merely reason: it governs its own reasoning—how much to compute, when to stop, which module to activate. We ask whether that role of metacognitive governor can be played by a dynamic internal field: a low-dimensional homeostatic state, with explicit physics and certified stability, that modulates a transformer’s cognition without performing it. Our instance is the Homeostatic Background Processor (HBP): a field defined over the module graph of a transformer and governed by a family of partial differential equations on the graph Laplacian —damped wave (forced Klein–Gordon), its diffusive limit, and a KdV-type dynamics— that evolves along the iterations of an adaptive-depth reasoner and modulates its computation (halting threshold, block gain, memory gates; the interface also exposes a router-bias head, not consumed by the models reported here) through a universal interface of interoception and modulation. We characterize the stability of the integrator of the whole family, and we are explicit that this is an integrator certificate and not a closed-loop one. Two of the three ingredients are classical and we recall them with attribution: a placement dichotomy for antisymmetric operators (gyroscopic in the second-order branch, positional in the first), which is Kelvin–Tait–Chetaev 1879treatise,merkin1997 and, in its first-order half, exactly Anti-Symmetric DGN 2023adgn with γ generalized to K; and a coercivity bound giving an unconditionally contractive implicit kernel, which is the logarithmic-norm resolvent estimate 1958stability,soderlind2006logarithmic. The circulatory branch produces flutter with threshold βρ(A3)<2ζω02βρ(A^3)<2ζ _0^2, exact only when stiffness and damping are both multiples of the identity —a hypothesis none of our runs satisfies, and under which the threshold is Bottema’s criterion 1955stability. What is new is narrower, and we now prove it: a discrete Schur–Cohn criterion with complex coefficients for the Verlet integrator with velocity coupling, necessary and sufficient per latent root and requiring no commutation hypothesis, which shows that a backward-differenced gyroscopic term does not inherit the neutrality it has in continuous time. Empirically (S5S_5, NC1-hard; pre-registered protocol, n=10n=10, plus a twenty-seed preregistered deconfounding campaign that partly overturned it), the answer is threefold: substance no, structure only in part, certifiability yes. Substance (no): the type of the field’s physics is irrelevant for accuracy —wave, diffusion, gated mixtures, non-local Poisson-type coupling and even a 2D Navier–Stokes flow substrate all give the same accuracy; the evidence is consistent with the gradient laminating every modulation demand to a quasi-static set-point, and incompressible flow hits a physical ceiling (∇⋅u=0∇\!· u=0 forbids concentrating information). Structure (only in part): the second order of the dynamics endows out-of-distribution compute allocation with a robustness the end-to-end learned halting control (gating_wm) lacks, and —after fixing a BF16 bug that froze the physical parameters— replicates unchanged; but a preregistered deconfounding campaign bounds the claim. With twenty fresh seeds, with the first-order branch rebuilt at equalized caps (which the unconditionally contractive implicit kernel makes safe, so the original caps were an artifact of the explicit integrator), and with a matched-interface GRU replacing the physical integrator, the order effect is strong in one generator family (+0.087+0.087 [+0.042,+0.132][+0.042,+0.132], t=4.0t=4.0) and is not detected in the other (+0.014+0.014 [−0.013,+0.040][-0.013,+0.040], n.s.—an interval that excludes the original v3 estimate but not a small positive effect): part of the original contrast was capacity, not order. The GRU governor is statistically indistinguishable from the field in the first family (+0.006+0.006 [−0.051,+0.062][-0.051,+0.062], n.s.) and nominally exceeds it in the second (−0.035-0.035 [−0.067,−0.002][-0.067,-0.002]), with mutual accuracy non-inferiority throughout. Certifiability (yes): what distinguishes the field is therefore not capability but that its stability is provable —placement dichotomy, flutter threshold exact under =ω02K= _0^2I, unconditional contraction bound— and, above all, that the one-step operator admits an exact runtime check. That is a difference of kind, not of existence: learned recurrences do carry certificates, and for gated recurrent units in particular there are explicit ISS and incremental-ISS conditions on the weights 2021stability; they are sufficient and conservative, and obtaining them requires bespoke machinery, whereas the field’s spectral radius is read off directly. We conclude that a dynamic internal field is a viable and certifiable compute governor —the “brainstem” of a cognitive architecture— but not an enhancer of cognition: it modulates, it does not think. A kill-gate with a positively controlled probe finds no evidence for the remaining route—the field as temporal evidence accumulator—on twelve frozen solvers of a companion substrate (ΔAUC=+0.0007 =+0.0007 [−0.0065,+0.0079][-0.0065,+0.0079] against a 0.030.03 pass threshold); we read this as the recurrent state already integrating its own history, a mechanism we propose rather than demonstrate. The null, broad and with a mechanism, delimits which role a homeostatic field can and cannot play in an adaptive transformer. 1 Introduction A cognitive architecture is not exhausted by the module that reasons. Alongside perception, memory and inference, an intelligent system needs a layer that governs its own computation: how much to think before answering, when to stop, which module to activate, how to distribute effort across parts of the input of uneven difficulty. In the brain, that function is performed not by the cortex but by the neuromodulatory and homeostatic systems —brainstem, hypothalamus, norepinephrine/acetylcholine tone— which set the gain, the arousal and the budget of cognition without executing it. This article asks whether that role of metacognitive governor can be played, in an adaptive transformer, by a dynamic internal field: a low-dimensional homeostatic state, with explicit physics and certified stability, that modulates cognition without performing it. (Our substrate is a family of small decoder-only transformers—4.24.2–5.65.6M parameters, per-variant counts in the repository—on an NC1-hard task; the question and the formalism are architecture-generic, and “language model” below refers to the architectural class, not the scale.) The concrete motivation is adaptive computation. A standard transformer computes at fixed depth; under logarithmic precision it is confined to TC0TC^0 2023parallelism and only learns “shortcuts” that do not extrapolate 2023shortcuts. Matching depth to difficulty —Adaptive Computation Time 2016act, PonderNet 2021pondernet, Universal Transformers 2019universal, looped transformers 2025looped, depth-adaptive and early-exit decoding 2020depth,schwartz2020right, and token-level compute routing 2024mixture— is the classic remedy, but in all of them the halting policy is learned end-to-end and can overfit the training distribution. We propose to anchor that decision in a dynamic internal field and ask what is gained —and what is not— by doing so. Our instance of the governor is the Homeostatic Background Processor (HBP): a low-dimensional field over the model’s module graph, governed by a family of PDE operators on the graph Laplacian and advection —reaction, diffusion, wave, advection, dispersion and a KdV-type nonlinearity—, which evolves along the iterations of a recurrent reasoner and modulates its computation through a universal interface: each module exposes interoception sis_i (bottom-up) and receives modulation mim_i (top-down), inspired by biological predictive regulation and neuromodulation 2012allostasis,barrett2017allostasis,friston2010free, doya2002metalearning,vecoven2020neuromodulation (Fig. 1). The family lets us decompose the headline question into two that the adaptive-computation literature does not isolate: (i) does substance matter, i.e. the type of the field’s physics (wave vs. diffusion vs. KdV, local vs. non-local, modulator vs. substrate)? and (i) does structure matter, i.e. the order of the dynamics? We evaluate with a pre-registered protocol and adversarial verification, and the answer is threefold: substance does not —the physical regime is irrelevant for accuracy, a broad mechanism null that we explain by the temporal coarse-graining of the gradient and, for the flow substrate, by an incompressibility limit—; the second-order structure does, but only within one of two generator families once capacity is equalized, and a matched-interface GRU governor ties the field there and nominally exceeds it in the other (Sec. 5.4); what remains distinctive of the field is neither substance nor superiority but certifiability. The answer, in one sentence. A dynamic internal field can govern an LLM’s computation —it is a viable governor, with certifiable stability— but it cannot enhance its cognition: it modulates, it does not think. And it is not the only thing that can govern: the one learned alternative we tested—a GRU cell with this same interoceptive interface—does the job at least as well, matching the field in one generator family and nominally beating it in the other, so the field’s distinctive contribution is that its stability is provable rather than merely observed. We thus place the HBP in the metacognitive/autonomic layer of a cognitive architecture —the computational “brainstem”, not the cortex— and characterize precisely, via a broad null with a mechanism, which role a homeostatic field can occupy and which it cannot. Contributions. 1. An architectural positioning, backed by evidence: a dynamic internal field belongs to the metacognitive compute-governor layer of an LLM, not to the reasoning layer; the empirical pattern (neutral for accuracy; better compute allocation than the no-controller baseline, in one of two generator families, and matched there by a learned recurrent governor) is what suggests that placement (Sec. 3, 6). 2. A novel formalism: a family of homeostatic PDE fields over the module graph (wave, diffusion, advection, dispersion, saturated KdV nonlinearity), with continuous mixing between first- and second-order branches and optional gating of the physics by interoception (Sec. 3). 3. A stability characterization of the family’s integrator, assembled from classical results with attribution and one new criterion. Classical: the placement dichotomy (Kelvin–Tait–Chetaev 1879treatise,merkin1997; its first-order half is A-DGN 2023adgn with γ→ ), the flutter threshold (Bottema’s criterion 1955stability at isotropic stiffness and damping, and the internal-damping whirl threshold of 1933motion), and the unconditionally contractive implicit kernel (the logarithmic-norm resolvent bound 1958stability,soderlind2006logarithmic,bai2003hermitian). New, and proved here (App. B, Prop. 4): a discrete complex Schur–Cohn condition for the Verlet integrator with velocity coupling, necessary and sufficient per latent root and free of any commutation hypothesis, with the passage to an operator-level box certificate quantified rather than assumed. All are imposed as a differentiable penalty during training and audited on trained checkpoints with the exact spectral radius (Sec. 4, App. A). 4. Two pre-registered empirical studies on state tracking in S5S_5: a first campaign (n=10n=10) in which the OOD compute-robustness effect survives a paired re-run with learnable physics after a numerical-precision bug (BF16 freezing), and a twenty-seed deconfounding campaign that bounds the claim—at equalized caps the order effect holds in one generator family and not the other, and a matched-interface GRU governor attains it too (Secs. 5, 5.4). 5. A broad mechanism null, adversarially verified: none of the members of the family moves accuracy, not even in dual tasks designed to excite regime switching, nor under non-local coupling (Poisson enslavement on the graph), nor with the field as a substrate of a 2D Navier–Stokes flow —where an incompressibility limit (∇⋅u=0∇\!· u=0) prevents delivering information to a point. We propose the temporal coarse-graining of the gradient as the explanation for the modulator null (Sec. 6). 2 Related work Adaptive computation and learned halting. The idea of decoupling compute depth from architectural depth goes back to Adaptive Computation Time (ACT) 2016act; PonderNet 2021pondernet reformulates it probabilistically, and Universal Transformers 2019universal and looped transformers 2025looped,giannou2023looped carry it to depth recurrence. Our reasoner inherits this paradigm with PonderNet-style halting. The crucial distinction is that in these works the compute-allocation policy is learned end-to-end; our study provides direct evidence that such a policy can overfit (the adaptivity of the gating_wm control collapses out of distribution in one of the two generator sets) and that anchoring it in a stateful governor with an interoceptive interface makes it more robust—though the deconfounding of Sec. 5.4 shows that the explicit physics is not what buys the robustness: a GRU with the same interface is as robust in adjacent and nominally more so in cycle_transp. Expressivity and state tracking. Under logarithmic precision, a fixed-depth transformer sits in TC0TC^0 2023parallelism; the word problem over S5S_5 is NC1-complete by the theorem of 1989, and transformers learn shortcuts that do not extrapolate 2023shortcuts. Intermediate generation extends the expressive power 2024cot, and state-space models share the limitation 2024illusion. This justifies that our task requires iterative computation and motivates the extrapolation axis 2022length,jelassi2023arithmetic. External and working memory. Neural Turing Machines 2014ntm and the DNC 2016dnc enable algorithms with addressable memory. Our gating_wm control incorporates a working memory in this spirit; our experiments attribute the in-distribution accuracy gain to the recurrence rather than to the memory (gating ≥ gating_wm in both generator sets) and, in either case, not to the HBP: a negative that we make explicit. Certified learned dynamics, and where our certificate sits. An earlier version of this paper did not engage this literature, which is a serious omission because it is where the bar actually is. Stability guarantees for learned recurrent models come in two grades. By construction: coRNN and UnICORNN 2021cornn,rusch2021unicornn are structure-preserving discretizations of second-order oscillator networks with proven state and gradient bounds; AntisymmetricRNN 2019antisymmetricrnn and A-DGN 2023adgn obtain stability from antisymmetric parameterization; Recurrent Equilibrium Networks 2024recurrent give a free parameterization of all models that are contracting and satisfy prescribed incremental IQCs, trainable by unconstrained gradient descent; Lipschitz RNNs 2021lipschitz and the Lyapunov-projected models of 2019learning do the same by other routes; LinOSS 2025oscillatory and its damped extension 2025dissipate obtain stable oscillatory state-space layers with a non-negativity condition, the latter decoupling damping from frequency exactly as our learnable per-dimension ζ does; and CON 2024con proves global asymptotic stability and input-to-state stability for a coupled damped-oscillator network in closed loop. By certificate: 2021stability give explicit ISS and incremental-ISS conditions on the weights of a gated recurrent unit, checkable post hoc or imposable during training. Two consequences for this paper, both uncomfortable and both stated here rather than left for a referee. First, our guarantee is of the weaker grade: we impose a differentiable penalty and probe the operator at evaluation time, where coRNN, REN and D-LinOSS are stable by construction. Second, the closed-loop result we do not prove —the field, its host, and the interoception/modulation path considered as one system— is exactly what CON proves for its setting. What is genuinely unoccupied in this literature is the conjunction of second order, instance-gated (time-varying) coefficients, and closed loop: REN has the loop without the first two, CON has the first and third without gating, LinOSS has only the first. We do not fill that gap here; we mark it. PDE-governed neural networks on graphs. Interpreting layers as discretizations of continuous dynamics is the basis of Neural ODEs 2018node; on graphs, PDE-GCN 2021grand,eliasof2021pdegcn derives architectures from the diffusion and wave equations, GraphCON 2022graphcon couples oscillators at the nodes (the second-order dynamics closest to our wave branch), ADR-GNN 2023adr adds advection and reaction, and Anti-Symmetric DGN 2023antisymmetric obtains stability by construction with antisymmetric operators. We should be explicit about how close that last one is: the first-order half of our placement dichotomy is A-DGN, with their γ generalized to a symmetric ⪰ω02K _0^2I. Claiming the generalization is honest; not flagging the coincidence would not be. Our family differs on three points: (i) the field does not transport the task representation but a low-dimensional modulatory state; (i) the “time” of the PDE is the reasoner’s iterations, not the layers; (i) the wave↔ mixing and the odd operators (advection, KdV dispersion) require a stability analysis that the GNN literature does not cover: incorrect placement of an antisymmetric operator in a second-order dynamics produces circulatory flutter, a classical phenomenon of non-conservative mechanics 1952,merkin1997,kirillov2013 that, to our knowledge, had not been pointed out in this context. Homeostasis and neuromodulation. Differentiable plasticity and neuromodulation 2018plasticity,miconi2019backpropamine enable internal modulation of computation; homeostatic variables coupled to vulnerability signals confer adaptability under concept shift 2022homeostatic. The HBP articulates this intuition as a physically motivated and analyzable dynamical system, with certified stability and its effect on the compute policy isolated experimentally. 3 A family of homeostatic fields Figure 1: The Homeostatic Background Processor. An internal-state field h (top) lives over the model’s module graph —backbone blocks, the adaptive-depth reasoner and the working memory (bottom)— coupled by the graph Laplacian (−c2∇2h-c^2∇^2h). Each module exposes an interoception signal sis_i (bottom-up) and receives modulation mim_i (top-down: halting threshold, block gain, memory gates; the interface also exposes a router-bias head, not consumed by the experiments reported here). The field evolves under a damped-wave PDE (one instance of the family of Sec. 3), one tick per reasoner iteration. Let there be a graph of N nodes (modules: L backbone blocks, the reasoner and the working memory) with combinatorial Laplacian ∈ℝN×NL ^N× N (symmetric, PSD) and oriented advection matrix A (antisymmetric; spectrum ±iμk\± i _k\), defined over the chain backbone→reasoner→memory. Each node holds an internal state vector (VEI) hi∈ℝdhh_i ^d_h; we write u:=h−h∗u:=h-h (deviation from the learnable rest state). Per latent dimension, we define the operators :=ω02+c2,:=2ζω0+D,:=b+β3,K:= _0^2I+c^2L, :=2ζ _0I+D\,L, :=b\,A+β\,A^3, (1) (stiffness: reaction ++ spatial diffusion; dissipation: uniform ++ structural; antisymmetric: advection ++ third-order dispersion, the only odd spatial operator well defined over an oriented graph, with (iμ)3=−iμ3(iμ)^3=-iμ^3), and the KdV-type nonlinear perturbation r(u):=−νtanh(u)⊙tanh(u),r(u):=-ν\, (u) (Au), (2) the saturated u∂xuu\, _xu term: r(0)=0r(0)=0 with null Jacobian at the equilibrium (same quadratic order as KdV near u=0u=0) and globally bounded by ν. The family consists of two branches coupled by a convex mixture α∈[0,1]α∈[0,1]: (wave, 2nd order): u¨+(+)u˙+u=r(u)+fθ(h,s)+gϕ(h,x), u+(C+G)\, u+Ku=r(u)+f_θ(h,s)+g_φ(h,x), (3) (diffusion, 1st order): γu˙+(+)u=r(u)+fθ(h,s)+gϕ(h,x), γ\, u+(K+G)\,u=r(u)+f_θ(h,s)+g_φ(h,x), (4) where fθf_θ is a bounded self-check (tanh , gain ≤0.3≤ 0.3) that consumes the interoception s, gϕg_φ a bounded external forcing from the input, and γ a learnable rate of the diffusive branch (decoupled from ζ). Note the distinct placement of G: gyroscopic (on u˙ u) in (3), positional (on u) in (4); Lemma 1 shows it is the only stable choice in each branch. The state advances by hn+1=αΦwave(hn,hn−1)+(1−α)Φdiff(hn),h_n+1=α\, _wave(h_n,h_n-1)+(1-α)\, _diff(h_n), (5) with Φwave _wave a position Verlet step and Φdiff _diff an IMEX backward-Euler step (stiff part +K+G implicit; Prop. 2). This differs from the classical implicit–explicit split of 1995imex, where the stiff symmetric diffusive term is taken implicitly and the advective, antisymmetric term explicitly — with the conditional stability that entails. Here the antisymmetric operator goes inside the implicit kernel, which is what the coercivity argument of Prop. 2 makes unconditionally safe. The coefficient α can be (i) an architectural constant (α=1α=1: wave; α=0α=0 or order 1: diffusion), (i) imposed per instance (experimental oracles), or (i) gated by interoception together with D and b (“homeostasis chooses its physics”), the hypothesis that Section 6 evaluates and refutes for accuracy. Instantiated physics. We evaluate three members of the family: hbp_full (damped wave: α=1α=1, =0G=0, D=0D=0), hbp_first (diffusive relaxation: order 1, overdamped limit), and hbp_kdv (wave ++ dispersion ++ KdV nonlinearity: α=1α=1, βmax=0.1 _ =0.1, νmax=0.3 _ =0.3); plus the gated variant hbp_mix. At N=6N=6 the dispersion has only three pairs of modes: there is no solitonic regime, and we make it explicit as a limit of the testbed. “Time” is thinking. The HBP ticks once per reasoner iteration (PonderNet-style halting, up to Nmax=24N_ =24). At iteration n: (i) the field modulates the reasoner block with a soft perturbation 1+smodtanh(⋅)1+s_mod (·) around the identity; (i) per-node interoception sns_n is collected (progress, activation, effort, attention entropy, valid fraction); (i) one step of the family is integrated; (iv) the VEI biases the halting threshold and the write/forget gates of the working memory. The physical parameters (ω0,ζ,c,D,b,β,ν,γ _0,ζ,c,D,b,β,ν,γ) are learnable per latent dimension within safe ranges (squash), and are stored in FP32 (Sec. 5.1). 4 Stability analysis of the family We define the per-dimension energy E(u,u˙)=12∥u˙∥2+12u⊤uE(u, u)= 12 u ^2+ 12\,u Ku (kinetic ++ restoring well ++ graph Dirichlet energy). In the autonomous case of (3) with =0G=0, E˙=−u˙⊤u˙≤0 E=- u C u≤ 0 (strict dissipation with ζ≥ζmin>0ζ≥ _ >0). The nontrivial question is where the antisymmetric operators can enter without destroying this structure. Lemma 1 (Placement of antisymmetric operators). Let =⊤≻0K=K 0, =⊤≻0C=C 0 and =−⊤G=-G with spectrum iμk\i _k\. Then: (a) (gyroscopic: safe). For u¨+(+)u˙+u=0 u+(C+G) u+Ku=0, the energy satisfies E˙=−u˙⊤u˙≤0 E=- u C u≤ 0, since u˙⊤u˙=0 u G u=0 (the gyroscopic force does no work); by LaSalle, the origin is globally and asymptotically stable for all b,βb,β (Kelvin–Tait–Chetaev theorem 1997,kirillov2013). (b) (positional in first order: safe). For γu˙+(+)u=0γ u+(K+G)u=0, every eigenvalue λ of −(+)/γ-(K+G)/γ satisfies Reλ=−(v∗v)/γ≤−ω02/γ<0Reλ=-(v^*Kv)/γ≤- _0^2/γ<0 unconditionally in ∥ (numerical range: v∗v∈iℝv^*Gv∈ iR). With =0K=0, σ(−)⊂iℝσ(-G)⊂ iR: conservative transport and dispersion, with the KdV dispersion relation on the graph. (c) (circulatory: flutter, with a threshold that is exact only in the isotropic case). For u¨+u˙+(+)u=0 u+C u+(K+G)u=0 with =ω02K= _0^2I and =2ζω0C=2ζ _0I (equal frequencies and isotropic damping: the degenerate Merkin case 1997; this is the per-dimension structure of the field only when c=0c=0 and there is no structural diffusion, which no run of ours satisfies —see App. A, where we also note that the headline arm sets =0G=0 outright, making this a design-exclusion result rather than a certificate of a system we ran) and =β3G= ^3, the system decouples in the unitary basis of G and the Routh–Hurwitz criterion with complex coefficients gives asymptotic stability if and only if βρ(3)<2ζω02β\,ρ(A^3)<2ζ _0^2. Under these hypotheses this is Bottema’s criterion 1955stability evaluated at isotropic stiffness and damping, and it coincides with the internal-damping whirl threshold of 1933motion; we recall it rather than claim it. The threshold vanishes with the damping (as ζ→0ζ→ 0, every β>0β>0 destabilizes 1952) and grows only linearly with ζω02ζ _0^2. Note that the isotropic case is the one in which the celebrated destabilization paradox does not appear: the threshold tends to zero continuously with ζ, with none of the discontinuous drop that makes the general case interesting. (d) (nonlinear perturbation). r satisfies r(0)=0r(0)=0, Dr(0)=0Dr(0)=0 and ∥r(u)∥≤ν r(u) ≤ν: the linearization at the equilibrium —and with it (a)–(c) and the discrete certificates— is not altered; in (4) the origin is GAS if ν(1+ρ())<ω02ν(1+ρ(A))< _0^2 —the Jacobian of the saturating perturbation contributes a diagonal and an A-coupled term, so the Lipschitz constant is ν(1+ρ())ν(1+ρ(A)), not νρ()νρ(A) as an earlier version of this lemma stated; outside that condition we claim only local stability and do not exhibit the basin. On the invariant set induced by the state projection —a hard clamp in the implementation— r+fθ+gϕr+f_θ+g_φ acts as a bounded forcing and both branches are BIBS. Lemma 1 justifies the placement of (3) and (4) and invalidates the circulatory alternative: at our operating point (βmax=0.1 _ =0.1, ρ(3)=5.85ρ(A^3)=5.85) the threshold of (c), used here as the design guide it is for c>0c>0 (App. A), is violated over almost the entire admissible parameter box, and the free simulation confirms the flutter (divergence) against the gyroscopic decay. Proposition 1 (Discrete stability of the wave branch). For the position Verlet ut+1=(2−Δt2−Δt(+))ut−(−Δt(+))ut−1u_t+1=(2I- t^2K- t(C+G))u_t-(I- t(C+G))u_t-1, every latent eigenvalue z solves the complex scalar polynomial z2−(2−q−g−iμ~)z+(1−g−iμ~)=0z^2-(2-q-g-i μ)z+(1-g-i μ)=0 with q=Δt2x∗xq= t^2\,x^*Kx, g=Δtx∗xg= t\,x^*Cx real and μ~=ΔtIm(x∗x) μ= t\,Im(x^*Gx), |μ~|≤Δtρ()| μ|≤ t\,ρ(G), ρ()=maxk|bμk−βμk3|ρ(G)= _k|b _k-β _k^3|. For q>0q>0, Cohn’s criterion for the complex quadratic gives |z|<1|z|<1 if and only if (i)μ~2<g(2−g),(i)q(g2+μ~2)<2g(g(2−g)−μ~2).(i)\;\; μ^2<g(2-g), (i)\;\;q\,(g^2+ μ^2)<2g\, (g(2-g)- μ^2 ). (6) Two remarks on scope, both proved in App. B. First, the reduction to (6) is a latent-root argument and therefore needs no commutation hypothesis: each eigenvalue satisfies the scalar polynomial with the Rayleigh quotients of its own latent vector, whether or not K, C and G are simultaneously diagonalizable. Second, the equivalence is per latent root; certifying the whole operator requires the triples (q,g,μ~)(q,g, μ) to be controlled, and bounding them by the spectra of ,,K,C,G replaces their joint numerical range by a box that contains it, so the box form is sufficient but not necessary —and, we find, markedly conservative. At μ~=0 μ=0, (i) reduces to the classical damped-Verlet criterion q<4(1−ζω0Δt)q<4(1-ζ _0 t). Condition (i) is genuinely discrete: the gyroscopic term discretized with a backward difference does not inherit the continuous neutrality, but injects energy at the exact rate 1+μ~2 1+ μ^2 per step, which the dissipation must absorb; as Δt→0 t→ 0 the condition empties (μ~2=O(Δt2) μ^2=O( t^2) against g(2−g)=O(Δt)g(2-g)=O( t)) and the continuum is recovered. A Cayley discretization of the gyroscopic term 2000lie would restore exact neutrality at the cost of an implicit step. This is the one result in this paper for which we have not found a precedent; the placement dichotomy and the contraction bound are classical (App. B). We prove it in App. B and check it in experiments/verify_verlet_schurcohn.py: the symbolic identity, 302 400302\,400 cells of a (q,g,μ~)(q,g, μ) grid against exact roots with zero discrepancies, and 400400 draws of non-commuting (,,)(K,C,G) —none of the 400400 pairs commuted— with a maximum latent-polynomial residual of 6⋅10−146· 10^-14. Proposition 2 (IMEX diffusive branch with implicit antisymmetric kernel). With the antisymmetric part inside the implicit kernel, G:=+λ(+)M_G:=I+λ(K+G), λ=Δt/γλ= t/γ, one has Re⟨x,Gx⟩≥(1+λω02)∥x∥2Re x,M_Gx ≥(1+λ _0^2) x ^2 for all x, hence ∥G−1∥2≤(1+λω02)−1<1 _G^-1 _2≤(1+λ _0^2)^-1<1 unconditionally in γ,ω0,c,b,βγ, _0,c,b,β and with no commutation hypothesis between L and G. (With G explicit this claim is false; the Cayley variant on the antisymmetric part preserves exactly the norm of the conservative transport.) Remark 1. The force −u˙-G u in the wave branch is of Coriolis type: it preserves the stability structure of the antisymmetric operators, not their literal transport semantics, which survive only in the first-order branch. Remark 2. The stability of the convex mixture α of the discrete maps does not follow from the continuous lemma (the per-branch Lyapunov functions are distinct), and —correcting an earlier version of this remark— neither does it follow from Propositions 1 and 2: the spectral radius is not convex, so per-branch bounds do not transfer to the mixture, and a norm bound cannot be combined with a spectral-radius bound for a non-normal companion. The runtime probe builds the companion of the wave branch only. The gated mixture arm therefore has neither an analytic nor an exact numerical certificate, and we report its results on that footing. We attempted to repair this with the standard instrument —a common quadratic Lyapunov function obtained by semidefinite programming over the vertices of the coefficient box— and report the outcome in App. C: it succeeds for the wave branch on the region the trained models occupy, and it fails for the mixture. The per-branch conditions are imposed as a differentiable penalty during training (on the learned coefficients, or their caps if gated) and verified at evaluation time with the exact spectral radius of the 2N×2N2N×2N companion matrix per dimension (covers the non-normal case [,]≠0[L,A]≠ 0). Under bounded forcing, the VEI is expected to be ISS to a ball around h∗h by a standard argument 1989smooth, which we do not prove here and on which nothing below depends. The energy identity, the flutter threshold and the discrete criterion were validated numerically against the implemented operators at fixed parameter draws (energy identity to residual 4⋅10−164·10^-16; flutter threshold to ±2%± 2\%; 0/2⋅1050/2· 10^5 discrepancies of the discrete criterion against exact roots). Figure 2: Intrinsic dynamics of the field (simulation of the solver of Sec. 3). An interoceptive impulse at node 1 induces a damped oscillation of the VEI (left, ‖hi−h∗‖\|h_i-h \| per node) that propagates to neighboring modules through the Laplacian coupling (right, node×tick map). It is the second-order regime that the analogy with coupled spins in NMR describes (a local impulse spreads and precesses); the Verlet scheme integrates it stably under the certificate of Prop. 1. NMR analogy. The family is formally that of a system of coupled spins: ζ∼1/T2ζ 1/T_2 (transverse relaxation), ω0 _0 (precession), c2c^2L (dipolar coupling), D∼1/T1D 1/T_1 (longitudinal relaxation / spin diffusion), bb\,A (directed transport) and β3β\,A^3 (dispersion: mode-dependent group velocity; Fig. 2 shows the real propagation-oscillation). The normal modes of L are the spectrum of the system, and the flutter certificate of Lemma 1(c) is the analog of the instability threshold of a spin system pumped above its relaxation rate. 5 Experiments Task and protocol. Non-commutative state tracking: composition of K generators of S5S_5 (NC1-hard ⇒ requires recurrence), with dense supervision of the running permutation and two generator sets (adjacent transpositions; 5-cycle ++ transposition). Pre-registered protocol (declared before looking at any number): fresh seeds 10..1910..19, evaluation with a disjoint seed (no contamination), argmax restricted to the answer sub-vocabulary, Nmax=24≥KmaxN_ =24≥ K_ , primary analysis paired by seed (Fisher-z of the pure-OOD correlation, K≥14K≥ 14) with Holm correction over the 4 contrasts. Models of 4.24.2–5.65.6M parameters (vanilla 4.234.23M, gating 5.285.28M, gating_wm 5.545.54M, the field variants 5.605.60M; per-variant counts in the repository). OOD regime: training with K≤12K≤ 12, evaluation up to K≤24K≤ 24. 5.1 A numerical-precision bug and its A/B: the effect is structural During the audit we discovered that, in BF16, the field’s raw physical parameters (|θ|∼0.3|θ| 0.3–1.61.6) have ULP/2ULP/2 (1.0⋅10−31.0·10^-3 to 3.9⋅10−33.9·10^-3) larger than the typical Adam step (lr≈3⋅10−4lr≈ 3·10^-4): the update is rounded to zero at every step and ω0,ζ,c _0,ζ,c remain frozen at their initialization throughout training (verified: ζ≡0.5ζ≡ 0.5 exact after 2500 steps; with the fix —anchoring those parameters to FP32— they move). With the bug fixed, we re-ran the full protocol (90 cells). The seed-paired A/B between frozen and learnable physics gives ΔcorrOOD=+0.010±0.046 \,corr_OOD=+0.010± 0.046 (adjacent) and +0.019±0.054+0.019± 0.054 (5-cycle), both non-significant at n=10n=10 and with intervals wide enough to contain effects the size of the order effect itself: we find no evidence that the HBP effect depends on the fine tuning of its physical constants. This A/B re-runs the same cells and seeds under a different numerical precision, so it is a paired same-batch contrast rather than an independent replication —and Sec. 5.4 then bounds the structural reading to one generator family. We report both versions for transparency; all numbers that follow are from the fixed version (v3, learnable physics). 5.2 In distribution: recurrence—not the field, and not the working memory—explains accuracy Table 1: Accuracy by difficulty (in-distribution, mean over n=3n=3 seeds; per-seed values and paired contrasts in the repository). S5S_5 adjacent S5S_5 5-cycle++transp. Model short medium long short medium long vanilla 0.9840.984 0.6700.670 0.1470.147 0.9740.974 0.9760.976 0.5810.581 gating 0.9850.985 0.7610.761 0.2180.218 0.9740.974 0.9800.980 0.7250.725 gating_wm 0.9830.983 0.7470.747 0.1970.197 0.9740.974 0.9820.982 0.7140.714 hbp_first 0.9870.987 0.7760.776 0.2450.245 0.9740.974 0.9850.985 0.7300.730 hbp_full 0.9860.986 0.7710.771 0.2480.248 0.9740.974 0.9840.984 0.7410.741 The gain decomposes cleanly, and the working memory is not where it comes from: recurrence buys +0.071+0.071 (adjacent) and +0.144+0.144 (5-cycle) in the long stratum (vanilla→ ), adding the working memory buys nothing and is nominally negative (gating→ _wm: −0.021-0.021 and −0.011-0.011), and the HBP adds +0.03+0.03–0.050.05 in the long stratum, which does not reach significance at n=3n=3 (p≥0.11p≥ 0.11). The 2nd- vs 1st-order contrast is null in distribution. 5.3 Out of distribution (v3, before deconfounding): a second-order adaptivity signal We measure compute adaptivity corr(K,[niter])corr(K,E[n_iter]) restricted to the pure-OOD range (K≥14K≥ 14), paired by seed (Table 2, Fig. 3). Table 2: Pure-OOD compute adaptivity, protocol v3 (learnable physics), n=10n=10 fresh seeds (± ). Fisher-z paired contrasts against the gating_wm control. Boldface marks uncorrected p<0.05p<0.05: neither contrast survives Holm-4 individually, and the order contrast these numbers support is bounded by the deconfounding of Sec. 5.4. gating_wm (WM) hbp_first (diffusive) hbp_full (wave) S5S_5 adjacent 0.002±0.0870.002± 0.087 0.036±0.0900.036± 0.090 0.143±0.1340.143± 0.134 vs WM – t=0.77t=0.77, p=0.46p=0.46 =2.68t=2.68, p=0.025p=0.025 S5S_5 5-cycle++transp. 0.213±0.0680.213± 0.068 0.185±0.0770.185± 0.077 0.268±0.0850.268± 0.085 vs WM – t=−1.42t=-1.42, p=0.19p=0.19 =2.62t=2.62, p=0.028p=0.028 Figure 3: OOD compute adaptivity by variant and generator set (n=10n=10, ± ), protocol v3 (learnable physics). The hbp_first arm runs at its original caps: the twenty-seed replication of Sec. 5.4 shows that the 5-cycle gap displayed here is confounded with capacity and vanishes once the caps are equalized. Three observations. (1) A second-order signal, later bounded: the diffusive limit ties with the control in both generator sets, and the direct wave−-diffusion contrast gives t=1.90t=1.90 (p=0.09p=0.09) and t=3.49t=3.49 (p=0.007p=0.007). Both are exploratory (outside the pre-registered Holm-4 family) and both run with the hbp_first caps not equalized; Sec. 5.4 shows that once they are, the 5-cycle contrast disappears (+0.014+0.014 [−0.013,+0.040][-0.013,+0.040], n.s.)—part of what reads here as order was capacity. (2) Statistical honesty: none of the primary contrasts individually survives Holm-4 (p=0.025/0.028p=0.025/0.028 against αHolm=0.0125/0.0167 _Holm=0.0125/0.0167). We report no combined p: the two generator sets share the same ten seeds, so a Fisher combination would violate independence. We present the effect as directionally consistent evidence from a single confirmatory batch—not as definitive, and not as an independent replication, since v2 and v3 re-run the same cells and seeds. (3) The compute slope in pure OOD ([niter|K∈18..24]−[niter|K∈12..15]E[n_iter|K∈18..24]-E[n_iter|K∈12..15]) is maximal for the second order in both sets (0.450.45 and 2.452.45), though in 5-cycle the gating_wm control is a near-tie (2.292.29) and hbp_first (1.491.49) falls below the control, so the ordering is monotone only in adjacent. OOD accuracy does not separate across variants (≈0.06≈ 0.06–0.080.08; the task does not extrapolate for anyone): the effect is on the allocation of computation. Figure 4: The deconfounding campaign of this section (twenty fresh seeds, 2020–3939). Left: out-of-distribution compute adaptivity per arm and generator family; the first-order arm now runs at caps equalized to the second-order ones, which the unconditionally contractive implicit kernel of Proposition 2 makes safe. Right: the same evidence as seed-paired differences from the second-order field, which is where the verdict actually lives. The order effect clears zero only in adjacent; the matched-interface GRU governor ties the field there and its interval falls entirely below zero in cycle_transp. Error bars are ± (left) and 95%95\% CI (right). 5.4 Deconfounding order, and a matched-interface learned governor At review’s request we ran a preregistered deconfounding campaign (twenty fresh seeds 20–39, never used before; OOD protocol unchanged; paired by seed, one-sided, Holm-2 per hypothesis). The first-order arm was rebuilt with caps equalized to the second-order ones (cmax=0.7c_ =0.7, ω0 _0 up to 1.81.8). The equalization is of the coefficient caps only: the second-order branch still carries a velocity state, hence twice the effective controller state, and we did not run the two further controls that would separate order from state size and from underdamping (a first-order arm with 2dh2d_h state; an overdamped second-order arm at the same caps); the adjacent effect below must be read with that residual confound. The equalized caps are what the unconditionally contractive implicit kernel of Proposition 2 makes safe: the original hbp_first caps were an artifact of the explicit ζ-coupled integrator, so the earlier order contrast was confounded with capacity. The second new arm replaces the physical integrator with a per-node GRU cell of identical interface—same interoception, same external forcing, same modulation heads. The results execute the preregistration’s honest branches. For reference, mean pure-OOD adaptivity per arm (n=20n=20, ±1± 1 SE) is hbp_full 0.172±0.0220.172± 0.022 / hbp_first_eq 0.086±0.0240.086± 0.024 / hbp_gru 0.167±0.0230.167± 0.023 in adjacent, and 0.242±0.0130.242± 0.013 / 0.229±0.0140.229± 0.014 / 0.277±0.0150.277± 0.015 in cycle_transp—the GRU arm in cycle_transp is the best single arm of the study. The order effect is strong in adjacent (Δcorr=+0.087 _corr=+0.087 [+0.042,+0.132][+0.042,+0.132], t=4.0t=4.0, p=3.5×10−4p=3.5× 10^-4) and not detected in cycle_transp once capacity is equalized (+0.014+0.014 [−0.013,+0.040][-0.013,+0.040], n.s.): under Holm the order hypothesis is not confirmed as a family-general claim, and we report it as generator-bounded—part of the original contrast was capacity. The GRU governor is statistically indistinguishable from the field in adjacent (+0.006+0.006 [−0.051,+0.062][-0.051,+0.062], n.s.—an interval too wide to certify equivalence) and nominally exceeds it in cycle_transp (−0.035-0.035 [−0.067,−0.002][-0.067,-0.002]), with mutual accuracy non-inferiority everywhere (margin 0.020.02). We therefore downgrade structure yes to: a stateful recurrent governor with this interoceptive interface suffices. The field’s distinctive contribution is that it is an implementation whose stability is certifiable (Appendices A–B); the GRU carries no such certificate. 5.5 The zoo of physics: KdV dispersion neither helps nor hurts On the same protocol (n=10n=10), hbp_kdv (wave ++ dispersion β3 ^3 ++ saturated nonlinearity, gyroscopic placement of Lemma 1) matches the pure wave in OOD adaptivity: corrOOD=0.186corr_OOD=0.186 vs. 0.1430.143 (adjacent) and 0.2570.257 vs. 0.2680.268 (5-cycle); the paired kdv−-wave contrast is null in both (t=1.12t=1.12, p=0.29p=0.29; t=−0.59t=-0.59, p=0.57p=0.57). Since hbp_kdv is second order, it separates from the gating_wm control in adjacent (t=7.05t=7.05, p<0.001p<0.001) but not in 5-cycle (t=2.08t=2.08, p=0.07p=0.07, n.s.); both contrasts are exploratory, outside the pre-registered Holm-4 family. That is: what separates hbp_kdv from hbp_full —the third-order dispersive operator— shows no detectable effect at n=10n=10 in either direction, while what they share —the second-order inertia— is what covaries with adaptivity, and only in adjacent once Sec. 5.4 bounds it. The zoo runs with the corrected dynamics (gyroscopic placement, IMEX with antisymmetric kernel, complex Schur–Cohn certificate active; the exact spectral radius was checked <1<1 at initialization and at the β cap, not at every trained operating points). At N=6N=6 and ∼ 5–11 ticks there is no solitonic regime to exploit; the result bounds the role of dispersion in small module graphs and reinforces the reading we can defend: within this zoo, what covaries with adaptivity is the second-order inertial term rather than the richness of the spatial operator—a reading that Sec. 5.4 then confines to adjacent and shows is not exclusive to a physical integrator. 6 A mechanism null: the field’s physics does not move accuracy The family lets us ask directly: does the physical regime of the field matter for performance? Three independent lines of evidence, subjected to adversarial verification (refutation panels with pre-specified confounder controls), answer no: (1) Forced physics. Training with α imposed (pure wave vs. pure diffusion) over four tasks (composition in S5S_5, functional iteration, modular sum with distractors, recall), accuracy is identical in all (|Δ|≤0.012| |≤ 0.012, within binomial error), with both regimes certified stable. (2) Gated physics. Letting interoception choose α,D,bα,D,b per tick (hbp_mix), the gating converges to a constant physics (α→0.79α→ 0.79, deviation across inputs ∼10−5 10^-5, also per dimension: α∈[0.76,0.84]α∈[0.76,0.84] across the N×dhN×d_h components), with the head weights growing under weight decay (the gradient arrives; the loss prefers constant physics) and with a strongly input-sensitive initialization retrained back to the same attractor. Ablating D or b changes accuracy by <0.008<0.008. (3) Dual-regime tasks. We design tasks with two signaled modes whose temporal modulation demand differs (oscillatory retention vs. monotone settling), mounted on composition in S5S_5 to guarantee iteration (the reasoner reaches [niter]E[n_iter] up to 1111, increasing with K). Even so, wave and diffusion perform identically within each mode (|Δ|≤0.005| |≤ 0.005), which bounds to ≤0.005≤ 0.005 the gap a physics-switching oracle could have achieved on these tasks; we do not claim the bound transfers to demands we did not construct. (4) Beyond local physics: non-locality and a flow substrate. The three lines above vary the type of local physics. We further test two structural extensions that the set-point argument does not trivially dissolve. (a) Non-locality. We couple the modulation to a stream function enslaved by Poisson on the graph, ψ=+(h−h∗)ψ=L^+(h-h ), so that the modulation of each node depends on the whole field (the discrete analog of the non-local dipolar field in NMR, ∇2ψ=−ω∇^2ψ=-ω). On the same OOD protocol, non-locality does not improve accuracy (Δ≈0 ≈ 0) and degrades compute adaptivity in 5-cycle++transp. (Δz=−0.069 z=-0.069, t=−3.88t=-3.88) while showing no effect in adjacent (Δz=−0.023 z=-0.023, t=−0.66t=-0.66, n.s.). A plausible reading—which this exploratory contrast does not establish—is that the difficulty signal governing halting is local to the reasoner node, so globalizing it blurs it. (b) 2D flow substrate. We carry the field to a sheet with incompressible Navier–Stokes dynamics (vorticity ω, ψ=+ωψ=L^+ω, velocity u=∂yψ,v=−∂xψu= _yψ,\,v=- _xψ, advected passive scalar S), where the field computes transport instead of modulating. A fundamental physical limit closes it: by incompressibility (∇⋅u=0∇\!· u=0, conservation of material area) a flow can transport and mix but never concentrate information at a point —concentrating requires negative divergence, i.e. compressibility—; in a delivery task with local readout, the mass delivered to the destination is ≈0≈ 0 and the temporal dynamics makes it worse (mixes more). The incompressible flow substrate is a mixer, not a delivering computer. Both extensions confirm the null by routes that are not “another type of local physics”. Interpretation: temporal coarse-graining. The reasoner does not process one token per tick: it groups (∼ 5 ticks for ∼ 19 operations), and gradient descent systematically finds solutions in which the modulation demand is a quasi-static set-point per instance. Both branches of the family reach any forced set-point at equilibrium; they differ only in the transient (∼ 4 ticks), which is invisible to the loss. Physics switching would only pay off with oscillatory demands aligned tick by tick, which the architecture does not force and the gradient does not discover. Together with (4), the pattern is broad: neither the types of local physics we instantiated, nor non-locality, nor the substrate-computer role change the result. This null is, we believe, informative beyond our model: it suggests that dynamic modulatory substrates in adaptive architectures are exercised by their recurrent structure (statefulness, and—in one of our two generator families—the order of the dynamics) rather than by their type of physics, and that an incompressible field, despite its elegance, has a structural ceiling as a spatial computer. The field as an evidence accumulator: the niche is occupied. One employment survives the set-point argument on paper: a damped second-order field is, formally, a leaky integrator with inertia—the physical form of sequential evidence accumulation. If the reasoner’s per-tick posterior stream (halting mass, readout margin and entropy, prediction flips, trajectory speed) were noisy in a way that a temporal filter could exploit, the field would finally carry load. We tested the premise with a kill-gate whose primary cell and pass threshold (ΔAUC≥0.03 ≥ 0.03) were fixed in a dated design document before the verdict—not a formal pre-registration, and the probe was recalibrated after a first null (below)—run on the twelve frozen solvers of the integration program of the companion paper 2026cognition (six seeds × two independent runs, a reasoner-with-value substrate distinct from the n=10n=10 protocol above; we report the gate here because it closes this paper’s thesis, and the transfer of its conclusion rests on the shared recurrent-reasoner structure), using a probe whose sensitivity we certified with planted distributed signals (load 00 gives −0.001±0.001-0.001± 0.001, unbiased; load 0.200.20 gives +0.026+0.026 [+0.020,+0.031][+0.020,+0.031], i.e. sensitivity at the scale of—but not above—the 0.030.03 threshold; an uncalibrated version of the same probe paid an overfitting tax of ∼0.018 0.018 AUC that biased it toward the null, and the positive control caught it before any verdict). The result is flat: the full stream predicts success no better than the last tick alone, ΔAUC=+0.0007 =+0.0007 [−0.0065,+0.0079][-0.0065,\,+0.0079] in the predeclared primary cell—against a pass threshold of 0.030.03, which it does not meet—with every secondary cell within [−0.0012,+0.0073][-0.0012,+0.0073]; and the detrended stream carries almost no temporal structure (lag-22–66 autocorrelation ≈+0.05≈+0.05; non-DC spectral peaks in 0%0\% of trajectory-speed traces and 17%17\% of margin traces, none exploitable by the certified probe). The mechanism is structural rather than statistical: the recurrent state already integrates its own history, so the last tick is a sufficient statistic of the trajectory and the accumulator’s niche is occupied by construction. This closes the last route by which the field could have carried content, and sharpens the paper’s thesis into an audit: of the employments we were able to imagine and test—source of computation, load-bearing physics, evidence accumulator—none is performed better by the field than by a dedicated component of the architecture, except the one the field already performs. It modulates compute; that is the job, and it is a real one—but not an exclusive one: a matched learned recurrence modulates compute at least as well (Sec. 5.4), so what stays the field’s own is the certificate, not the job. 7 Discussion and limitations 7.1 What the homeostatic field does and does not do The results tease apart three questions. Does it improve accuracy? No: the recurrence explains in-distribution accuracy (the working memory adds nothing on top of it in our own table), and none of the members of the family we instantiated —forced, gated or from the zoo— moves it. Does the type of physics matter? No for performance (Sec. 6); the diffusion equation plays here the role of the contrast that gives meaning to the central claim, and the KdV operators bound the role of dispersion at small N. Does structure matter? Partly: the second-order dynamics produces a compute allocation more robust out of distribution than the control’s learned policy, and the effect is independent of the tuning of its constants (BF16-bug A/B) —but under deconfounding at equalized capacity it holds in one generator family and not the other, and a matched-interface GRU reaches the same place—nominally going past it in the family where the order effect vanishes (Sec. 5.4). The field is a compute controller, not an accuracy enhancer; inertia is an active ingredient in one family, and the class of viable governors is wider than the field. 7.2 Where it fits: the field as governor, not as cortex The pattern of results is not an ambiguous failure but a location signal. That a dynamic internal field is neutral for accuracy yet load-bearing only for compute allocation—in one generator family, and not exclusively (Sec. 5.4)—places it, on the evidence we have, in the metacognitive control / resource governor layer of a cognitive architecture —and not in the reasoning one. In biological terms, it is not the cortex (which computes the answer) but the neuromodulatory and homeostatic substrate (which sets gain, arousal and budget); in systems terms, it is the scheduler/governor, not the arithmetic unit; in control terms, the external supervisory loop, slow and certified, wrapped around the fast inference loop. Its universal interoception/modulation interface (Fig. 1) is precisely that of a governor: it observes the computational “vital signs” (progress, uncertainty, effort) and modulates how much and where to compute, without deciding what. The certified stability (Sec. 4) ceases to be a technical detail and becomes the design property a governor needs: a control loop that provably does not diverge. It is also, after the GRU comparison, the property that survives as distinctive: the field is not the strongest governor of its class, it is the one that comes with a proof. The substance null and the bounded structure effect, together, do not disqualify the field —they relocate it: it is a viable and certifiable compute governor, not a cognition module. 7.3 What this component is not, and what is missing Our thesis is bounded: the HBP is one piece —the regulatory one— of a broader intelligent system, not the system. It does not reason, it does not plan and it does not replace the recurrent reasoner, whose recurrence is what explains in-distribution accuracy. A complete architecture pursuing cognition would require, in addition to the governor characterized here, a deliberative module (the “cortex”), a world model/memory, and a system of drives and goals; all of them lie outside this work and are the subject of future development. This article’s contribution is to close, with rigor and honesty, the question about that piece: a dynamic internal field can govern an LLM’s computation in a certifiable way, and what is distinctively its own is the certificate—not capability, and not even the second-order structure, which is bounded to one generator family and matched by a learned recurrence. 7.4 Limitations (1) A single task family (S5S_5), though theoretically sharp. (2) The finding’s metric is compute allocation, not OOD accuracy (which does not separate), and the structural effect that remains is bounded to one of the two generator sets and is reproduced by a matched-interface GRU governor (Sec. 5.4). (3) Small scale (4.24.2–5.65.6M parameters; n=3n=3 in-dist and n=10n=10 OOD in v3, raised to n=20n=20 in the deconfounding campaign) and v3 primary contrasts that do not survive Holm-4 individually; the deconfounding equalizes coefficient caps but not controller state size. (4) N=6N=6 nodes: no solitonic regime for the dispersion; the zoo of physics at large N remains open. (5) The mechanism null is a result about what the gradient finds, not about what the architecture could express: a task with an architecturally forced tick-aligned oscillatory demand could reverse it. (6) The flow-substrate limit (Sec. 6, line 4) is specific to incompressibility: a field with learnable sinks (compressible advection–diffusion–reaction) could deliver, at the cost of abandoning the vorticity–stream-function formulation and its elegance; we did not evaluate it. 7.5 Future work (i) Larger module graphs (nodes per layer/head) where dispersion and advection have sufficient spectrum; (i) Cayley discretization of the gyroscopic term (exact discrete neutrality, Prop. 1); (i) tasks where the compute budget is the effective bottleneck, to turn allocation robustness into accuracy; (iv) interoceptive signals discriminative of the problem (the current ones are nearly input-invariant at fixed K, which limits any gating); (v) integration with NMR applications, where the formal analogy is transferable. 8 Conclusion We asked whether a dynamic internal field can modulate an LLM’s cognition. The answer, supported by proved results with their scope stated (App. A), numerical audits of the implementation, and two pre-declared empirical campaigns (the v3 protocol and the v4 deconfounding), is and threefold: substance no, structure only in part, certifiability yes. A second-order homeostatic field is a viable and certifiable compute governor —it confers on compute allocation an out-of-distribution robustness that the gating_wm control’s learned halting policy does not have, in one of two generator families, while a matched-interface GRU governor reaches the same place—and nominally goes past it in the other family—without a certificate—, but its type of physics is irrelevant for accuracy in every configuration we tested: we propose that the gradient laminates temporal demand to set-points, and the null is broad, resisting the sweep of local physics, Poisson-type non-locality and a 2D Navier–Stokes flow substrate (closed by an incompressibility limit). The field modulates, it does not think. We thus place the HBP in the metacognitive control layer of a cognitive architecture —its computational “brainstem”— and provide the stability tools (gyroscopic-placement lemma, complex Schur–Cohn certificates, unconditional IMEX scheme) that a governor of this kind needs —and that its learned competitors do not have. We believe the substance/structure/certifiability demarcation, the broad null with a mechanism, and the architectural location that follows from it are useful to anyone designing the layer that governs —rather than performs— a model’s computation. Characterizing one piece well, and its limits, is the necessary step before assembling the rest. Reproducibility. All the code is pure PyTorch, with no third-party framework dependencies, with fixed seeds and protocols frozen before observing the data. The field’s implementation and its certificates are in model/hbp.py (PDE family, gyroscopic placement, IMEX solver, stability_penalty and the exact spectral radius certificate_spectral_radius); the 2D flow substrate in model/flow2d.py. Reproducing the tables and figures: the pre-registered benchmark and its paired analysis (experiments/benchmark_v3.py, experiments/benchmark_report_v2.py), the v4 deconfounding campaign (PREREG_V4.md, experiments/benchmark_v4.py, results_benchmark_v4/report_v4.json), the accumulator kill-gate (mhbp/tasks/reasoner_g0/n4_g1.py), the KdV zoo (experiments/benchmark_zoo.py), the non-locality A/B (experiments/benchmark_elliptic.py), and the mechanism-null studies (experiments/_alpha_scan.py, experiments/_pde_study.py, experiments/_twin_pilot.py, experiments/_gates_flowroute.py). A regression test verifies that the field is actually wired in —ablating it changes the forward pass and the compute policy, though not accuracy— (experiments/_audit_hbp.py), a check of plumbing rather than of the functional load this paper’s null denies; and physical verifications of the 2D solver in experiments/_check_flow2d.py. The physical parameters are anchored to FP32 (pin_fp32); we document the BF16 precision bug and its A/B as part of the experimental record. Appendix A Scope and status of the stability results At a reviewer’s request we state precisely what is proven, under which hypotheses, and what is verified numerically. Proven. The discrete Verlet criterion of Prop. 1 is proved in App. B: it is necessary and sufficient per latent root, and the reduction uses no commutation hypothesis. The passage from latent roots to an operator-level certificate via the parameter box is sufficient only, and Remark 4 both proves the correct finite check and reports how conservative it is. The placement lemma’s structural claims—that antisymmetric coupling enters the second-order branch gyroscopically and the first-order branch positionally, and that the two placements have qualitatively different stability consequences—follow from the symmetric/antisymmetric decomposition given in the stability section, with the standard energy arguments sketched there —standard because they are classical: (a) is Kelvin–Tait–Chetaev and (b) is the accretive/numerical range argument. Part (c) of the dichotomy is not proved here: showing that one Lyapunov candidate stops decreasing does not establish instability, and the corresponding instability theorems are in 1999stability. The unconditional contraction of the IMEX kernel does not follow from Schur–Cohn applied mode by mode —that route needs simultaneous diagonalization, which fails here— but from the coercivity of the symmetric part, which is precisely what buys the absence of a commutation hypothesis (App. B). Exact under hypothesis, and the hypothesis is never met. The closed-form flutter threshold βρ(A3)<2ζω02βρ(A^3)<2ζ _0^2 is exact when both the stiffness and the damping operator are multiples of the identity: K=ω02IK= _0^2I (i.e. c=0c=0) and C=2ζω0IC=2ζ _0I (i.e. no structural diffusion). The second is a hypothesis the earlier version of this appendix left silent. Two corrections follow. First, the parenthetical equivalence “c=0c=0, or [,A3]=0[L,A^3]=0” was wrong: under commutation with c>0c>0 the system does decouple, but the per-mode threshold becomes β|νj|<2ζω0ω02+c2ℓjβ| _j|<2ζ _0 _0^2+c^2 _j, so the published threshold is then sufficient but not necessary, hence not exact. Second, and more consequentially: in the implementation c is parameterized as cmaxσ(⋅)c_ σ(·) and is therefore strictly positive in every run, so the exact regime covers none of the systems we report. Worse for the reader who checks, the headline arm instantiates no antisymmetric operator at all (=0G=0 by configuration), so the placement-and-flutter apparatus is a design-exclusion result about a variant we chose not to build, not a certificate of anything we ran. We use the threshold as the design guide it is, and claim nothing more. Verified numerically. The identities and thresholds of the stability section were validated against direct numerical experiments on the discrete system (residuals at machine precision for the identities; threshold location within the stated tolerance; no instability observed under the certified conditions in long-horizon integration). These are audits of the implementation at fixed parameter draws, not theorems. On trained runs. The operative guarantee during and after training is the runtime certificate: the exact spectral radius of the one-step operator, computed by probing the implementation (certificate_spectral_radius), together with the training-time stability penalty. Certificates are audited on trained checkpoints in dedicated scripts; they are not logged at every training step, and we scope the paper’s verification claims accordingly. Evidence-accumulator gate. The kill-gate of Section 6 ran on the twelve frozen solvers of the integration program of the companion paper 2026cognition—a reasoner-with-value substrate distinct from the n=10n=10 protocol of this paper; the transfer of its conclusion rests on the shared recurrent-reasoner structure, as stated in the main text. Appendix B Proofs of the certified results We give complete proofs of the three results whose scope Appendix A delimits: the placement dichotomy, the exact flutter threshold under the stated hypothesis, and the unconditional IMEX kernel bound. Throughout, K is symmetric with K⪰ω02IK _0^2I (in the field, K=ω02I+c2K= _0^2I+c^2L with L the graph Laplacian, positive semidefinite), and G is real antisymmetric (G⊤=−G =-G). Lemma 2 (Placement dichotomy). Let F(t)F(t) be a bounded forcing. (i) In the first-order branch u˙=−Ku−Gu+F u=-Ku-Gu+F, the antisymmetric term is norm-neutral: dt12‖u‖2=−u⊤Ku+u⊤F ddt 12\|u\|^2=-u Ku+u F, so stability is governed by K alone. (i) In the second-order branch with gyroscopic placement, x¨+2ζω0x˙+Kx+Gx˙=F x+2ζ _0 x+Kx+G x=F, the mechanical energy E=12‖x˙‖2+12x⊤KxE= 12\| x\|^2+ 12x Kx satisfies E˙=−2ζω0‖x˙‖2+x˙⊤F E=-2ζ _0\| x\|^2+ x F: the antisymmetric term does no work and E remains a Lyapunov function. (i) With circulatory placement, x¨+2ζω0x˙+Kx+Gx=F x+2ζ _0 x+Kx+Gx=F, one gets E˙=−2ζω0‖x˙‖2−x˙⊤Gx+x˙⊤F E=-2ζ _0\| x\|^2- x Gx+ x F; the cross term x˙⊤Gx x Gx is sign-indefinite and can pump energy, so stability is conditional (Lemma 3). Proof. All three identities follow by differentiating the stated functional along trajectories and using v⊤Gv=0v Gv=0 for every real v (by antisymmetry, v⊤Gv=(v⊤Gv)⊤=−v⊤Gv Gv=(v Gv) =-v Gv). In (i), u⊤u˙=−u⊤Ku−u⊤Gu+u⊤Fu u=-u Ku-u Gu+u F and the middle term vanishes. In (i), E˙=x˙⊤x¨+x˙⊤Kx E= x x+ x Kx; substituting x¨ x and cancelling x˙⊤Gx˙=0 x G x=0 gives the claim. In (i) the same substitution leaves −x˙⊤Gx- x Gx, which has no sign. ∎ Lemma 3 (Exact flutter threshold for K=ω02IK= _0^2I). Consider the unforced circulatory system x¨+2ζω0x˙+ω02x+βA3x=0 x+2ζ _0 x+ _0^2x+β A^3x=0 with ζ,ω0,β>0ζ, _0,β>0 and A real antisymmetric. The origin is asymptotically stable if and only if βρ(A3)<2ζω02β\,ρ(A^3)<2ζ _0^2, where ρ(⋅)ρ(·) is the spectral radius. Proof. A3A^3 is antisymmetric, hence normal with purely imaginary spectrum iνj\i _j\, νj∈ℝ _j , and ℂNC^N decomposes into orthogonal invariant eigenplanes. Because the remaining operators are multiples of the identity, the system block-diagonalizes over these planes, and on the mode with eigenvalue iνiν the characteristic polynomial is p(λ)=λ2+2ζω0λ+ω02+iβνp(λ)=λ^2+2ζ _0λ+ _0^2+iβν. At β=0β=0 both roots lie in the open left half-plane (real part −ζω0<0-ζ _0<0 for ζ<1ζ<1; two negative real roots −ζω0±ω0ζ2−1-ζ _0± _0 ζ^2-1 for ζ≥1ζ≥ 1). Roots depend continuously on β, so instability can only arise by a root crossing the imaginary axis. Setting λ=iωλ=iω (ω∈ℝω ) and separating real and imaginary parts: −ω2+ω02=0-ω^2+ _0^2=0 and 2ζω0ω+βν=02ζ _0ω+βν=0, i.e. ω=±ω0ω=± _0 and βν=∓2ζω02βν=∓ 2ζ _0^2. A boundary crossing therefore occurs exactly at β|ν|=2ζω02β|ν|=2ζ _0^2; below it no root can have reached the axis, above it (checking the sign of ∂Reλ/∂β∂\,Re\,λ/∂β at the crossing, which is positive) one root has positive real part. Taking the worst mode |ν|=ρ(A3)|ν|=ρ(A^3) gives the claim. For K=ω02I+c2K= _0^2I+c^2L with c>0c>0 the eigenplanes of A3A^3 are no longer invariant unless [,A3]=0[L,A^3]=0, and the threshold is not claimed exact (Appendix A). ∎ Proposition 3 (Unconditional IMEX kernel bound; standard, recalled for completeness). Let M=I+λ(K+G)M=I+λ(K+G) with λ>0λ>0, K symmetric with K⪰ω02IK _0^2I, and G real antisymmetric. Then M is invertible and ‖M−1‖2≤(1+λω02)−1<1\|M^-1\|_2≤(1+λ _0^2)^-1<1, with no condition on λ (hence on dtdt, ζ) and no commutation hypothesis. Consequently the backward-Euler update u′=M−1(u+λF)u =M^-1(u+λ F) is a strict contraction toward the forced equilibrium whenever F is independent of the state. When F depends on u —as it does in the implementation, where it collects the saturating perturbation and the learned drives— the composed map has Lipschitz constant at most (1+λLF)/(1+λω02)(1+λ L_F)/(1+λ _0^2), which is below one only if LF<ω02L_F< _0^2. That is a condition on the operating box, not a consequence of this proposition, and we do not assert it: at the low end of the admissible range (ω02 _0^2 small against the drive gains) it need not hold. Proof. For any x∈ℂNx ^N, x∗Gx^*Gx is purely imaginary: since G is real antisymmetric, (x∗Gx)∗=x∗G⊤x=−x∗Gx(x^*Gx)^*=x^*G x=-x^*Gx. Hence Rex∗Mx=‖x‖2+λx∗Kx≥(1+λω02)‖x‖2Re\,x^*Mx=\|x\|^2+λ\,x^*Kx≥(1+λ _0^2)\|x\|^2, using K⪰ω02IK _0^2I. By Cauchy–Schwarz, ‖Mx‖‖x‖≥|x∗Mx|≥Rex∗Mx\|Mx\|\,\|x\|≥|x^*Mx| \,x^*Mx, so ‖Mx‖≥(1+λω02)‖x‖\|Mx\|≥(1+λ _0^2)\|x\| for all x; thus M is injective (hence invertible) and σmin(M)≥1+λω02 _ (M)≥ 1+λ _0^2, which is the stated bound on ‖M−1‖2\|M^-1\|_2. ∎ Remark 3. Proposition 3 is the reason the first-order branch can run with equalized caps (cmaxc_ , ω0,max _0, at the second-order values) in the v4 deconfounding study: the implicit kernel is unconditionally contractive, so the caps that the explicit ζ-coupled Euler forced on hbp_first are not needed. When the antisymmetric coefficients are gated per instance, G cannot be folded into a fixed kernel and the certificate falls back to the runtime spectral-radius probe, as stated in Appendix A. Author contributions (CRediT). F.M.A.-C. (corresponding): conceptualization, methodology, software, formal analysis, investigation, visualization, project administration, writing—original draft. F.G.M.: conceptualization, methodology, formal analysis, validation, supervision, writing—review and editing. A.A.: software, data curation, resources, investigation, validation, writing—review and editing. I.F.: conceptualization, supervision, funding acquisition, resources, validation, writing—review and editing. Data and code availability. The full program—model code, preregistrations, run artifacts and the scripts that produce every table and figure in this paper—is covered by the MIT licence. It is available at https://github.com/fmarrabal/miuracognitive. Every number in the text enters as a macro transcribed from an archived artifact and checked against it, so each one can be traced back to the file it came from. Use of generative AI. The experiments were designed, executed and adjudicated by the authors. A large language model was used as a coding and drafting assistant throughout, including for implementation, for adversarial review of the manuscript’s internal consistency, and for literature search; all claims, numbers and verdicts reported here were verified by the authors against the archived artifacts. Proposition 4 (Discrete criterion for the wave branch; Prop. 1). Let =⊤K=K , =⊤C=C be real and =−⊤G=-G real, and consider the position-Verlet step with velocity coupling ut+1=(2−Δt2−Δt(+))ut−(−Δt(+))ut−1.u_t+1= (2I- t^2K- t(C+G) )u_t- (I- t(C+G) )u_t-1. Let z be an eigenvalue of its companion operator and x≠0x≠ 0 the corresponding latent vector, normalized to ∥x∥=1 x =1. Put q=Δt2x∗xq= t^2\,x^*Kx, g=Δtx∗xg= t\,x^*Cx and μ~=ΔtIm(x∗x) μ= t\,Im(x^*Gx), all real. Then z2+a1z+a0=0z^2+a_1z+a_0=0 with a0=(1−g)−iμ~a_0=(1-g)-i μ and a1=−(2−q−g)+iμ~a_1=-(2-q-g)+i μ, and, provided q>0q>0, |z|<1⟺μ~2<g(2−g)andq(g2+μ~2)<2g(g(2−g)−μ~2).|z|<1 μ^2<g(2-g)\ \ and\ \ q(g^2+ μ^2)<2g (g(2-g)- μ^2 ). No commutation between K, C and G is assumed. Proof. Reduction. Writing the recursion as a quadratic matrix polynomial, z is a latent root of P(z)=z2−z(2−Δt2−Δt(+))+(−Δt(+))P(z)=z^2I-z (2I- t^2K- t(C+G) )+ (I- t(C+G) ) with P(z)x=0P(z)x=0. Left-multiplying by x∗x^* and using ∥x∥=1 x =1 gives the scalar equation. Here x∗x^*Kx and x∗x^*Cx are real because ,K,C are real symmetric, and x∗x^*Gx is purely imaginary because G is real antisymmetric (x∗x¯=x∗⊤x=−x∗x x^*Gx=x^*G x=-x^*Gx). Hence a0,a1a_0,a_1 have the stated form. Note that this step uses only the latent vector attached to z; it does not diagonalize the three operators simultaneously, and therefore holds when they do not commute. Criterion. For a monic complex quadratic z2+a1z+a0z^2+a_1z+a_0, Cohn’s criterion states that both roots lie in the open unit disk if and only if |a0|<1|a_0|<1 and |a1−a¯1a0|<1−|a0|2|a_1- a_1a_0|<1-|a_0|^2. We evaluate both. For the first, |a0|2=(1−g)2+μ~2|a_0|^2=(1-g)^2+ μ^2, so |a0|<1⇔1−2g+g2+μ~2<1⇔μ~2<g(2−g)|a_0|<1 1-2g+g^2+ μ^2<1 μ^2<g(2-g), which is (i). Write D:=1−|a0|2=g(2−g)−μ~2D:=1-|a_0|^2=g(2-g)- μ^2, positive exactly under (i). For the second, set s:=2−q−gs:=2-q-g, so a1=−s+iμ~a_1=-s+i μ and a¯1=−s−iμ~ a_1=-s-i μ. Then a¯1a0=(−s−iμ~)((1−g)−iμ~)=−s(1−g)−μ~2+iμ~(s−(1−g)), a_1a_0=(-s-i μ) ((1-g)-i μ )=-s(1-g)- μ^2+i μ (s-(1-g) ), and therefore a1−a¯1a0=(μ~2−sg)+iμ~(2−s−g)=(qg−D)+iqμ~,a_1- a_1a_0= ( μ^2-sg )+i μ\,(2-s-g)= (qg-D )+i\,q μ, using 2−s−g=q2-s-g=q and μ~2−sg=μ~2−g(2−q−g)=qg−D μ^2-sg= μ^2-g(2-q-g)=qg-D. Squaring, the condition |a1−a¯1a0|2<D2|a_1- a_1a_0|^2<D^2 reads (qg−D)2+q2μ~2<D2⇔q2(g2+μ~2)<2qgD,(qg-D)^2+q^2 μ^2<D^2 q^2 (g^2+ μ^2 )<2qgD, and dividing by q>0q>0 gives exactly (i). Both steps are equivalences, so the criterion is necessary and sufficient, not merely sufficient. ∎ Remark 4 (From latent roots to the operator: the box, and its price). Certifying the whole operator requires controlling the triples (q,g,μ~)(q,g, μ), which range over the joint numerical range of (,,)(K,C,G) on unit vectors. Replacing that set by the box q∈Δt2[λmin(),λmax()]q∈ t^2[ _ (K), _ (K)], g∈Δt[λmin(),λmax()]g∈ t[ _ (C), _ (C)], |μ~|≤Δtρ()| μ|≤ t\,ρ(G) yields a sufficient condition, since the box contains the joint range. Verifying it over the box is a finite computation, but not simply a check at the corners, which an earlier version of this paper asserted without proof. In q and in μ~2 μ^2 the left side of (i) increases and the right side decreases, so the worst case is qmax,|μ~|maxq_ ,| μ|_ . In g it is not monotone: F(g):=2g(g(2−g)−μ~2)−q(g2+μ~2)F(g):=2g (g(2-g)- μ^2 )-q(g^2+ μ^2) is a cubic with negative leading coefficient, so its minimum over an interval is attained at an endpoint or at the interior critical point solving −6g2+(8−2q)g−2μ~2=0-6g^2+(8-2q)g-2 μ^2=0. Likewise (i) binds at whichever endpoint of the g-interval is farther from 11, since g(2−g)g(2-g) peaks there. The certificate is therefore the minimum of F over the endpoints together with any interior critical point. We record that this box bound is markedly conservative: over 400400 random non-commuting draws it certified only 88, all of them correctly. For that reason the operative check in the experiments is the exact spectral radius of the companion, not the box. Appendix C A common-Lyapunov certificate: what it buys, and where it stops The runtime probe of Sec. 4 returns a spectral radius, which is blind to non-normal transients and says nothing about convex mixtures or about time-varying (instance-gated) coefficients, since neither the spectral radius nor stability under products is implied by ρ(Φt)<1ρ( _t)<1. The standard repair is a common quadratic Lyapunov function: find P≻0P 0 with Φ⊤PΦ⪯ρ2P P ρ^2P simultaneously for every Φ in the family. If it exists, ∥P1/2ΦP−1/2∥≤ρ P^1/2 P^-1/2 ≤ρ is a norm bound; the operator norm is convex, so every convex mixture inherits it, and submultiplicative, so arbitrary products do too. That single object would close the mixture, the gated case and the transients at once. We solved it. One technical point makes the vertex argument valid: Φ is not affine in (ω0,ζ,c)( _0,ζ,c) —the entries carry ω02 _0^2, ζω0ζ _0 and c2c^2— so we reparameterize in (Δt2ω02, 2Δtζω0,Δt2c2,…)( t^2 _0^2,\,2 tζ _0,\, t^2c^2,…), in which it is affine; the derived box contains the physical one, so the certificate is conservative rather than optimistic. Feasibility is re-checked by explicit eigenvalue tests rather than trusted from the solver status. What it buys. On the envelope of the trained checkpoints (ω0∈[0.488,0.529] _0∈[0.488,0.529], ζ∈[0.475,0.538]ζ∈[0.475,0.538], c≤0.371c≤ 0.371, read off the weights) a common P exists with ρ=0.952ρ=0.952 and κ=λmax/λmin=2.51κ= _ / _ =2.51; the worst vertex spectral radius is 0.7320.732, so the trained models sit well inside. This is strictly stronger than the runtime probe: a norm bound, valid for mixtures and for gated products within the wave branch. Where it stops, in three places. First, the declared coefficient box is not certifiable because it is not stable: 4040 of its 6464 vertices are already divergent (worst ρ=5.24ρ=5.24), the culprit being c, which enters the stiffness as c2λmax()c^2 _ (L) and at cmax=0.7c_ =0.7 already contributes ≈1.83≈ 1.83. The certifiable frontier degenerates to c=0,ω0≤0.276c=0,\ _0≤ 0.276. What keeps the model away from those configurations is training, not the box; the honest reading is that the caps in the configuration are nominal and too loose, and the operative guarantee remains the runtime check. Second, the mixture still does not close: including the diffusive branch and the antisymmetric operators, max∥Φ∥P=1.57 _P=1.57 under the wave branch’s P, and no common P exists for the enlarged set either. Third, and most consequentially, the closed loop does not close by small-gain. The condition ρ+κΔtLF<1ρ+κ t\,L_F<1 requires LF<0.0192L_F<0.0192. The only state-dependent forcing is fgainfθ([h;s])f_gainf_θ([h;s]), whose Lipschitz constant in h is bounded exactly from the weights by fgain∥W2∥Lip(SiLU)∥W1h∥f_gain W_2 \,Lip(SiLU) W_1h ; measured on the trained checkpoints it lies between 0.3670.367 and 0.9180.918. The gap is a factor of roughly fifty, and it is structural rather than marginal: the contraction margin 1−ρ=0.0481-ρ=0.048 is an order of magnitude smaller than the learned forcing gain. Certifying a single operating point instead of the envelope would raise the threshold to ≈0.13≈ 0.13 and still leave a factor of seven. What that does and does not mean. Small-gain is a sufficient condition. Its failure does not show the loop is unstable —it does not diverge in training— but that this technique does not reach. We also do not attempt the loop through the host (h→h→ modulation → transformer → interoception →fθ→ f_θ), which would require a Lipschitz bound on the host with respect to its own modulation that we do not have. We state the boundary rather than blur it: the certificate is an integrator certificate, extended here from spectral radius to norm on the region actually used, and it stops at the mixture and at the closed loop. The scripts are experiments/certify_lmi.py, experiments/certify_lmi_ckpt.py and experiments/certify_lmi_iss.py. References [Anil et al.(2022)Anil, Wu, Andreassen, Lewkowycz, Misra, Ramasesh, Slone, Gur-Ari, Dyer, and Neyshabur] Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. Exploring length generalization in large language models. Advances in Neural Information Processing Systems (NeurIPS), 2022. arXiv:2207.04901. [Arrabal-Campos(2026)] Francisco M. Arrabal-Campos. Where cognition lives: Dissecting emergent from computed function in a minimal complete cognitive architecture. Companion paper, arXiv preprint, 2026. [Ascher et al.(1995)Ascher, Ruuth, and Wetton] Uri M. Ascher, Steven J. Ruuth, and Brian T. R. Wetton. Implicit-explicit methods for time-dependent partial differential equations. SIAM Journal on Numerical Analysis, 32(3):797–823, 1995. [Bai et al.(2003)Bai, Golub, and Ng] Zhong-Zhi Bai, Gene H. Golub, and Michael K. Ng. Hermitian and skew-hermitian splitting methods for non-hermitian positive definite linear systems. SIAM Journal on Matrix Analysis and Applications, 24(3):603–626, 2003. ISSN 0895-4798. doi: 10.1137/S0895479801395458. [Banino et al.(2021)Banino, Balaguer, and Blundell] Andrea Banino, Jan Balaguer, and Charles Blundell. PonderNet: Learning to ponder. In 8th ICML Workshop on Automated Machine Learning (AutoML), 2021. arXiv:2107.05407. [Barrett(2017)] Lisa Feldman Barrett. The theory of constructed emotion: An active inference account of interoception and categorization. Social Cognitive and Affective Neuroscience, 12(1):1–23, 2017. [Barrington(1989)] David A. Barrington. Bounded-width polynomial-size branching programs recognize exactly those languages in NC1. Journal of Computer and System Sciences, 38(1):150–164, 1989. [Bonassi et al.(2021)Bonassi, Farina, and Scattolini] Fabio Bonassi, Marcello Farina, and Riccardo Scattolini. On the stability properties of gated recurrent units neural networks. Systems & Control Letters, 157:105049, 2021. doi: 10.1016/j.sysconle.2021.105049. [Bottema(1955)] Oene Bottema. On the stability of the equilibrium of a linear mechanical system. Zeitschrift für angewandte Mathematik und Physik (ZAMP), 6(2):97–104, March 1955. doi: 10.1007/BF01607296. [Boyer et al.(2025)Boyer, Rusch, and Rus] Jared Boyer, T. Konstantin Rusch, and Daniela Rus. Learning to dissipate energy in oscillatory state-space models, 2025. Preprint; introduces the D-LinOSS (Damped Linear Oscillatory State-Space) model. [Bulatovic(1999)] R. M. Bulatovic. On the stability of linear circulatory systems. Zeitschrift für angewandte Mathematik und Physik (ZAMP), 50(4):669–674, 1999. doi: 10.1007/s000330050172. [Chamberlain et al.(2021)Chamberlain, Rowbottom, Gorinova, Webb, Rossi, and Bronstein] Benjamin P. Chamberlain, James Rowbottom, Maria Gorinova, Stefan Webb, Emanuele Rossi, and Michael M. Bronstein. Grand: Graph neural diffusion. In International Conference on Machine Learning (ICML), 2021. [Chang et al.(2019)Chang, Chen, Haber, and Chi] Bo Chang, Minmin Chen, Eldad Haber, and Ed H. Chi. AntisymmetricRNN: A dynamical system view on recurrent neural networks. In 7th International Conference on Learning Representations (ICLR 2019), 2019. URL https://openreview.net/forum?id=ryxepo0cFX. [Chen et al.(2018)Chen, Rubanova, Bettencourt, and Duvenaud] Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David K. Duvenaud. Neural ordinary differential equations. In Advances in Neural Information Processing Systems (NeurIPS), volume 31, p. 6571–6583, 2018. [Dahlquist(1959)] Germund Dahlquist. Stability and Error Bounds in the Numerical Integration of Ordinary Differential Equations. Number 130 in Transactions of the Royal Institute of Technology, Stockholm. Almqvist & Wiksell, Uppsala, 1959. Doctoral thesis, Stockholm University, defended December 1958. Widely cited with the year 1958 (thesis) or 1959 (series publication). [Dehghani et al.(2019)Dehghani, Gouws, Vinyals, Uszkoreit, and Kaiser] Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser. Universal transformers. In International Conference on Learning Representations (ICLR), 2019. arXiv:1807.03819. [Doya(2002)] Kenji Doya. Metalearning and neuromodulation. Neural Networks, 15(4-6):495–506, 2002. [Elbayad et al.(2020)Elbayad, Gu, Grave, and Auli] Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. Depth-adaptive transformer. In International Conference on Learning Representations (ICLR), 2020. [Eliasof et al.(2021)Eliasof, Haber, and Treister] Moshe Eliasof, Eldad Haber, and Eran Treister. PDE-GCN: Novel architectures for graph neural networks motivated by partial differential equations. In Advances in Neural Information Processing Systems (NeurIPS), 2021. [Eliasof et al.(2023)Eliasof, Haber, and Treister] Moshe Eliasof, Eldad Haber, and Eran Treister. ADR-GNN: Advection-diffusion-reaction graph neural networks. arXiv preprint arXiv:2307.16092, 2023. [Erichson et al.(2021)Erichson, Azencot, Queiruga, Hodgkinson, and Mahoney] N. Benjamin Erichson, Omri Azencot, Alejandro F. Queiruga, Liam Hodgkinson, and Michael W. Mahoney. Lipschitz recurrent neural networks. In 9th International Conference on Learning Representations (ICLR 2021), 2021. URL https://openreview.net/forum?id=-N7PBXqOUJZ. [Fan et al.(2025)Fan, Du, Ramchandran, and Lee] Ying Fan, Yilun Du, Kannan Ramchandran, and Kangwook Lee. Looped transformers for length generalization. In International Conference on Learning Representations (ICLR), 2025. arXiv:2409.15647. [Friston(2010)] Karl Friston. The free-energy principle: A unified brain theory? Nature Reviews Neuroscience, 11(2):127–138, 2010. [Giannou et al.(2023)Giannou, Rajput, Sohn, Lee, Lee, and Papailiopoulos] Angeliki Giannou, Shashank Rajput, Jy-Yong Sohn, Kangwook Lee, Jason D. Lee, and Dimitris Papailiopoulos. Looped transformers as programmable computers. In Proceedings of the 40th International Conference on Machine Learning (ICML), 2023. arXiv:2301.13196. [Graves(2016)] Alex Graves. Adaptive computation time for recurrent neural networks. arXiv preprint arXiv:1603.08983, 2016. [Graves et al.(2014)Graves, Wayne, and Danihelka] Alex Graves, Greg Wayne, and Ivo Danihelka. Neural turing machines. arXiv preprint arXiv:1410.5401, 2014. [Graves et al.(2016)Graves, Wayne, Reynolds, Harley, Danihelka, et al.] Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, et al. Hybrid computing using a neural network with dynamic external memory. Nature, 538(7626):471–476, 2016. [Gravina et al.(2023a)Gravina, Bacciu, and Gallicchio] Alessio Gravina, Davide Bacciu, and Claudio Gallicchio. Anti-symmetric DGN: a stable architecture for deep graph networks. In The Eleventh International Conference on Learning Representations (ICLR 2023), 2023a. URL https://openreview.net/forum?id=J3Y7cgZOOS. [Gravina et al.(2023b)Gravina, Bacciu, and Gallicchio] Alessio Gravina, Davide Bacciu, and Claudio Gallicchio. Anti-symmetric DGN: A stable architecture for deep graph networks. In International Conference on Learning Representations (ICLR), 2023b. [Iserles et al.(2000)Iserles, Munthe-Kaas, Nørsett, and Zanna] Arieh Iserles, Hans Z. Munthe-Kaas, Syvert P. Nørsett, and Antonella Zanna. Lie-group methods. Acta Numerica, 9:215–365, 2000. ISSN 0962-4929. doi: 10.1017/S0962492900002154. [Jelassi et al.(2023)Jelassi, d’Ascoli, Domingo-Enrich, Wu, Li, and Charton] Samy Jelassi, Stéphane d’Ascoli, Carles Domingo-Enrich, Yuhuai Wu, Yuanzhi Li, and François Charton. Length generalization in arithmetic transformers. arXiv preprint arXiv:2306.15400, 2023. [Kirillov(2013)] Oleg N. Kirillov. Nonconservative Stability Problems of Modern Physics. De Gruyter, 2013. [Kolter & Manek(2019)Kolter and Manek] J. Zico Kolter and Gaurav Manek. Learning stable deep dynamics models. In Advances in Neural Information Processing Systems 32 (NeurIPS 2019), p. 11126–11134, 2019. URL https://proceedings.neurips.c/paper/2019/hash/0a4bbceda17a6253386bc9eb45240e25-Abstract.html. [Liu et al.(2023)Liu, Ash, Goel, Krishnamurthy, and Zhang] Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. Transformers learn shortcuts to automata. In International Conference on Learning Representations (ICLR), 2023. arXiv:2210.10749. [Man et al.(2022)Man, Damasio, and Neven] Kingson Man, Antonio Damasio, and Hartmut Neven. Need is all you need: Homeostatic neural networks adapt to concept shift. arXiv preprint arXiv:2205.08645, 2022. [Merkin(1997)] David R. Merkin. Introduction to the Theory of Stability, volume 24 of Texts in Applied Mathematics. Springer, New York, NY, 1997. doi: 10.1007/978-1-4612-4046-4. Translated and edited by F. F. Afagh and A. L. Smirnov. [Merrill & Sabharwal(2023)Merrill and Sabharwal] William Merrill and Ashish Sabharwal. The parallelism tradeoff: Limitations of log-precision transformers. Transactions of the Association for Computational Linguistics (TACL), 11:531–545, 2023. [Merrill & Sabharwal(2024)Merrill and Sabharwal] William Merrill and Ashish Sabharwal. The expressive power of transformers with chain of thought. In International Conference on Learning Representations (ICLR), 2024. [Merrill et al.(2024)Merrill, Petty, and Sabharwal] William Merrill, Jackson Petty, and Ashish Sabharwal. The illusion of state in state-space models. In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024. arXiv:2404.08819. [Miconi et al.(2018)Miconi, Stanley, and Clune] Thomas Miconi, Kenneth O. Stanley, and Jeff Clune. Differentiable plasticity: Training plastic neural networks with backpropagation. In International Conference on Machine Learning (ICML), p. 3559–3568, 2018. [Miconi et al.(2019)Miconi, Rawal, Clune, and Stanley] Thomas Miconi, Aditya Rawal, Jeff Clune, and Kenneth O. Stanley. Backpropamine: Training self-modifying neural networks with differentiable neuromodulated plasticity. In International Conference on Learning Representations (ICLR), 2019. [Raposo et al.(2024)Raposo, Ritter, Richards, Lillicrap, Humphreys, and Santoro] David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, and Adam Santoro. Mixture-of-depths: Dynamically allocating compute in transformer-based language models. arXiv preprint arXiv:2404.02258, 2024. [Revay et al.(2024)Revay, Wang, and Manchester] Max Revay, Ruigang Wang, and Ian R. Manchester. Recurrent equilibrium networks: Flexible dynamic models with guaranteed stability and robustness. IEEE Transactions on Automatic Control, 69(5):2855–2870, 2024. doi: 10.1109/TAC.2023.3294101. [Rusch & Mishra(2021a)Rusch and Mishra] T. Konstantin Rusch and Siddhartha Mishra. Coupled oscillatory recurrent neural network (coRNN): An accurate and (gradient) stable architecture for learning long time dependencies. In 9th International Conference on Learning Representations (ICLR 2021), 2021a. URL https://openreview.net/forum?id=F3s69XzWOia. Oral presentation. [Rusch & Mishra(2021b)Rusch and Mishra] T. Konstantin Rusch and Siddhartha Mishra. UnICORNN: A recurrent model for learning very long time dependencies. In Marina Meila and Tong Zhang (eds.), Proceedings of the 38th International Conference on Machine Learning (ICML 2021), volume 139 of Proceedings of Machine Learning Research, p. 9168–9178. PMLR, 2021b. URL https://proceedings.mlr.press/v139/rusch21a.html. [Rusch & Rus(2025)Rusch and Rus] T. Konstantin Rusch and Daniela Rus. Oscillatory state-space models. In The Thirteenth International Conference on Learning Representations (ICLR 2025), 2025. URL https://openreview.net/forum?id=GRMfXcAAFh. Oral presentation; introduces the LinOSS (Linear Oscillatory State-Space) model. [Rusch et al.(2022)Rusch, Chamberlain, Rowbottom, Mishra, and Bronstein] T. Konstantin Rusch, Benjamin P. Chamberlain, James Rowbottom, Siddhartha Mishra, and Michael M. Bronstein. Graph-coupled oscillator networks. In International Conference on Machine Learning (ICML), 2022. [Schwartz et al.(2020)Schwartz, Stanovsky, Swayamdipta, Dodge, and Smith] Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A. Smith. The right tool for the job: Matching model and instance complexities. In Association for Computational Linguistics (ACL), 2020. [Smith(1933)] David Macleish Smith. The motion of a rotor carried by a flexible shaft in flexible bearings. Proceedings of the Royal Society of London. Series A, Containing Papers of a Mathematical and Physical Character, 142(846):92–118, October 1933. doi: 10.1098/rspa.1933.0158. [Söderlind(2006)] Gustaf Söderlind. The logarithmic norm. history and modern theory. BIT Numerical Mathematics, 46(3):631–652, 2006. ISSN 0006-3835. doi: 10.1007/s10543-006-0069-9. [Sontag(1989)] Eduardo D. Sontag. Smooth stabilization implies coprime factorization. IEEE Transactions on Automatic Control, 34(4):435–443, 1989. [Sterling(2012)] Peter Sterling. Allostasis: A model of predictive regulation. Physiology & Behavior, 106(1):5–15, 2012. [Stölzle & Della Santina(2024)Stölzle and Della Santina] Maximilian Stölzle and Cosimo Della Santina. Input-to-state stable coupled oscillator networks for closed-form model-based control in latent space. In Advances in Neural Information Processing Systems 37 (NeurIPS 2024), 2024. URL http://papers.nips.c/paper_files/paper/2024/hash/952ddaa9299d81c307427edab034784-Abstract-Conference.html. [Thomson & Tait(1879)Thomson and Tait] William Thomson and Peter Guthrie Tait. Treatise on Natural Philosophy, volume I. Cambridge University Press, Cambridge, new edition, 1879. Vol. I, Part I. Effect of dissipative and gyroscopic (“gyrostatic”) forces on stability treated in § 345 and its roman-numeral subsections, ending p. 415 (§ 346 begins p. 416). [Vecoven et al.(2020)Vecoven, Ernst, Wehenkel, and Drion] Nicolas Vecoven, Damien Ernst, Antoine Wehenkel, and Guillaume Drion. Introducing neuromodulation in deep neural networks to learn adaptive behaviours. PLOS ONE, 15(1):e0227922, 2020. [Ziegler(1952)] Hans Ziegler. Die Stabilitätskriterien der Elastomechanik. Ingenieur-Archiv, 20(1):49–56, 1952.